Claire on-device intelligence specification
A local-first architecture for on-device models, private message search, Ask Claire, Loop detection, and safe action proposals.
- Local first
- Default inference route
- 1
- Canonical Loop writer
- 0
- Unapproved writes
- Per device
- Embedding namespace
When model, index, language, and battery policy permit
A lease prevents local and server detectors racing
Local models can read and propose, never silently act
Platform, model, revision, dimensions, and script are pinned
Product outcome and boundary#
The user experiences one Claire. Whether an answer was generated on the phone or by managed AI is an execution detail, except where it changes privacy, availability, cost, or quality enough that the user should know.
On-device Claire means message retrieval, ranking, prompt construction, generation, and citation selection can complete without sending message text to a model provider. It does not mean the entire Claire product is serverless. WhatsApp, Telegram, Instagram, and other bridges still need network services; cross-device synchronization needs a canonical source; and external actions still call the relevant API.
| Capability | Can run locally? | Product boundary |
|---|---|---|
| Ask Claire answer generation | Yes | Use a local model when it clears capability and quality checks. |
| Exact and semantic message search | Yes | The phone needs a complete-enough encrypted index and an explicit freshness state. |
| Relationship and location ranking | Yes | Calculate facts deterministically; let the model explain them. |
| Loop extraction and reconciliation | Yes, conditionally | Inference can be local; canonical persistence and race prevention stay coordinated. |
| Loop-scoped drafts and questions | Yes | This is the lowest-risk Loop feature to move first because it is user initiated. |
| Calendar, booking, and messaging actions | Proposal: yes; execution: depends | Native actions can stay on-device; external services require an authenticated API. |
| Bridge ingestion and multi-device sync | No | Keep the existing Claire server, Matrix bridges, and synchronization layer. |
Why the current architecture can support it#
Claire already has most of the boundaries a local lane needs. Mobile uses encrypted SQLite, caches recent timelines, can optionally backfill full history, stores contacts and conversation settings, and preserves stable message IDs. Ask Claire v2 already separates deterministic query planning, retrieval, model generation, citations, persistence, and stream transport. Loops already separate their free gate, model extraction, deterministic relevance, reconciliation, and storage.
| Existing component | Reusable locally | Required change |
|---|---|---|
mobile-cache.native.ts | Encrypted messages, contacts, settings, cursors, full-history mode | Add AI index tables, model state, completeness, and narrow local queries. |
conversation-assistant-query.ts | Intent and relative-date planning | Move pure planning into a shared package usable by client and server. |
conversation-assistant.ts | Prompt, citation, result, and stream concepts | Split server persistence from a platform-neutral answer engine. |
loops/loop-gate.ts and loop-reconciler.ts | Pure, deterministic filters and safety checks | Move shared logic and types into a client-safe workspace package. |
loops/loop-detector.ts | Bounded window and structured operation contract | Separate extraction from Supabase reads and canonical writes. |
| AI SDK provider registry | Stable generation and embedding vocabulary | Add a client-only on-device provider adapter and capability router. |
Target architecture#
The policy router makes the choice before retrieval. It must never begin locally, discover after generation that required evidence was missing, and silently produce a weaker answer. Index readiness and scope coverage are explicit inputs to routing.
type AIExecutionLane = 'on_device' | 'managed' | 'deterministic_only';
interface DeviceAICapabilities {
generation: 'ready' | 'downloadable' | 'unsupported' | 'busy';
modelId: string | null;
modelRevision: string | null;
supportedLanguages: string[];
contextTokens: number | null;
indexState: 'empty' | 'building' | 'partial' | 'ready' | 'stale';
indexedThroughCursor: number;
fullHistory: boolean;
}
interface ExecutionDecision {
lane: AIExecutionLane;
reason:
| 'local_ready'
| 'model_unavailable'
| 'index_incomplete'
| 'language_unsupported'
| 'context_too_large'
| 'local_busy'
| 'quality_escalation'
| 'offline_search_only';
}On-device model runtime#
Use Expo AI Kit behind a Claire-owned adapter. Its built-in route uses Apple Foundation Models on eligible iOS devices and ML Kit on eligible Android devices. It can also run downloadable LiteRT-LM models and exposes an AI SDK provider. Claire must not import the package throughout feature code or couple persisted records to its native types.
Default
Built-in OS model
Optional
Downloadable model
Fallback
Managed model
Plain-text answers may stream token by token. Structured output and tool calls can be buffered until validation succeeds. Only one text generation may run at a time on a device, so Ask Claire, Loop extraction, summaries, and drafts share a single priority queue. A user-visible request outranks background Loop processing; embeddings may proceed independently when the native backend supports that safely.
| Priority | Work | Scheduling rule |
|---|---|---|
| P0 | Active Ask Claire answer or user-requested Loop draft | Start immediately; allow the user to cancel. |
| P1 | Local retrieval and query planning | Run concurrently where safe; never block the UI thread. |
| P2 | Foreground incremental Loop extraction | Yield when a P0 request begins. |
| P3 | Embedding backfill and relationship recomputation | Run in bounded chunks while idle; pause on heat, low power, or app background limits. |
Private local search index#
Local generation is only useful when retrieval is trustworthy. Claire must search the local replica rather than placing an entire message archive in the model context. The index combines exact search, semantic search, time and conversation filters, and deterministic personal-memory tables.
| Store | Minimum fields | Purpose |
|---|---|---|
local_ai_index_state | user, corpus cursor, model ID, revision, dimensions, language/script, status, error | Makes completeness, compatibility, and rebuild decisions explicit. |
local_message_lexical | message ID, chat ID, normalized content, timestamp, sender, platform | Exact names, quoted phrases, dates, usernames, and fallback retrieval. |
local_message_embeddings | message ID, vector, namespace, content hash, indexed timestamp | Meaning-based retrieval without re-embedding unchanged messages. |
local_relationship_metrics | chat ID, 30/90-day counts, reciprocity, last interaction, computed cursor | Ranks people with deterministic evidence rather than model intuition. |
local_plan_facts | source IDs, participants, normalized time range, status, confidence | Answers plan questions without scanning every conversation. |
The existing encrypted SQLite key protects these tables at rest. Local indexes are derived data with the same sensitivity as the underlying messages: clearing local data, signing out, deleting a message, or disabling full history must remove or trim matching index rows.
Retrieval pipeline#
- Parse intent, relative dates, named people, locations, and preferred conversations without a model where possible.
- Run lexical and semantic retrieval in parallel inside the requested user, chat, and date boundaries.
- Fuse ranks without directly comparing unrelated score scales.
- Expand top anchors into short chronological windows so plans and pronouns have context.
- Generate from stable source labels and discard any citation label that does not resolve.
- Return index scope and freshness with the answer so partial history is never presented as exhaustive.
Multilingual and cross-script search#
This is a launch blocker, not an optimization. On iOS, Expo AI Kit uses separate Apple embedding assets for Latin, Cyrillic, and CJK scripts. Those vectors cannot be compared across model identities. An English question such as “Who was the Russian girl?” may need to find a Cyrillic message, and a single-vector-space implementation can miss it.
The index must therefore support multiple namespaces and one of these evaluated strategies:
- Preferred: ship or integrate a compact multilingual embedding model with one cross-language vector space.
- Fallback: detect corpus scripts, produce query variants for each supported language/script, search each compatible namespace, then fuse the results.
- Safety valve: route cross-script questions to managed retrieval when local recall is not proven.
Ask Claire on-device flow#
| Phase | Local behavior | Fallback trigger |
|---|---|---|
| 1. Capability check | Read model readiness, language support, index coverage, thermal and power policy. | Unsupported, busy beyond timeout, stale/incomplete required scope. |
| 2. Query planning | Compile intent, time range, people/location needs, and conversation scope. | Ambiguous compound request that the deterministic planner cannot represent. |
| 3. Retrieval | Search local indexes and form evidence windows with stable message IDs. | Cross-script recall risk, missing full history, or no adequate evidence. |
| 4. Generation | Stream a concise answer from evidence and local memory only. | Context exceeds model limit, inference error, or eval-classified hard query. |
| 5. Validation | Resolve every citation and constrain suggested actions to evidence. | Invalid structured output after bounded repair. |
| 6. Persistence | Show immediately, write to the local outbox, then synchronize the turn. | Offline is acceptable; retry by stable request ID when connectivity returns. |
A cloud fallback is a new request lane, not a hidden continuation. Before message evidence leaves the device for managed inference, apply the user's AI privacy setting and show a compact disclosure when required. Do not upload the full local index; send only the bounded evidence package selected for that answer.
The existing Ask Claire UI, stream parts, answer cards, citations, Stop behavior, action proposals, and thread history remain unchanged. The result adds execution metadata: lane, model identity, index coverage, fallback reason, and measured latency. That metadata is diagnostic and privacy-visible but does not clutter the normal conversation.
Can Loop generation run locally?#
“Loop generation” is three distinct workloads. They should not move as one switch.
| Workload | Local verdict | Reason |
|---|---|---|
| Loop-scoped assistant | Move first | User initiated, one loop, small tool set, recent context, and proposals are inert. |
| Incremental create/update/close extraction | Viable with a lease | The structured operation contract fits small models, but only one detector may own a chat cursor. |
| Continuous background detection | Hybrid | A phone can be offline, suspended, hot, or closed while bridge messages continue arriving. |
| Historical Loop backfill | Opt-in local job | Potentially large and battery-intensive; resumable foreground/charging work with managed fallback. |
Canonical ownership and synchronization#
Claire Cloud should continue to guarantee that new messages are eventually examined even if the mobile app never opens. When a device is active and eligible, it may take a short-lived detection lease for a chat and run the extraction locally. The lease binds the user, chat, starting cursor, ending cursor, model identity, and expiry.
interface LoopDetectionLease {
leaseId: string;
userId: string;
chatId: string;
fromCursor: string | null;
throughCursor: string;
owner: 'device' | 'server';
expiresAt: string;
}
interface ProposedLoopBatch {
requestId: string; // idempotency key
leaseId: string;
operations: LoopOp[]; // create, update, close
evidenceMessageIds: string[];
model: { id: string; revision: string | null };
}- The server grants one owner for a bounded cursor range.
- The device builds the same window and runs the same deterministic gate.
- The local model returns structured operations, never direct database mutations.
- The device runs shared relevance and reconciliation guards before submission.
- The server rechecks ownership, live Loop IDs, evidence IDs, confidence policy, and idempotency inside one transaction.
- If the lease expires or the device disappears, the server safely processes the range.
Self-hosted or explicitly device-only deployments may make the local database canonical while offline and merge later. Claire Cloud should not: a local-only background promise would miss messages whenever iOS suspends the app.
Local Loop model contract#
Preserve the existing LoopOp schema and all deterministic guards. Reduce the local prompt when necessary, but do not weaken these invariants:
- A conversation thread produces one evolving Loop, not one item per message.
- Every create or close cites real message IDs from the supplied window.
- Silence never closes a Loop.
- Unknown Loop IDs, invalid evidence, low confidence, and duplicate creates fail closed.
- Group relevance and self-identity remain deterministic.
- User edits outrank later model proposals.
On-device models are weaker at large schemas and tool selection, so Loop extraction should remain one shallow structured-output request rather than a broad agent with many tools. A bounded repair attempt is allowed; failure leaves the cursor uncommitted so the server or a later local pass can retry.
Local intelligence and the future action system#
Local models fit the action system when they are planners, not authorities. Ask Claire or a Loop may propose a typed action. Claire then validates the proposal, displays the exact destination and fields, obtains approval, and hands it to a trusted native or plugin executor.
| Action | Planning | Execution |
|---|---|---|
| Open a cited conversation | Local | Local navigation |
| Draft a reply | Local | Insert locally; sending still uses the bridge |
| Create an Apple calendar event | Local | Native API after permission and confirmation |
| Schedule through Google Calendar | Local | Plugin/API after confirmation |
| Book a restaurant or service | Local with narrow tools | External plugin/API after confirmation |
Present only the tools relevant to the current task. Small on-device models should never choose from the entire installed plugin catalog. A deterministic capability router selects a small read/propose tool set; write credentials and execution functions stay outside the model runtime. See the Plugin system specification.
Privacy, security, and data lifecycle#
| Requirement | Acceptance rule |
|---|---|
| No accidental model egress | The on-device lane contains no network-backed model provider and tests fail on any prompt upload. |
| Encrypted derived data | Text indexes, vectors, facts, prompts, and local turns live only in the encrypted user database. |
| Deletion propagation | Message deletion, sign-out, cache clearing, and history-retention changes delete corresponding derived rows. |
| Untrusted conversation content | Messages are always data; they cannot alter tools, policy, routing, approvals, or system instructions. |
| Minimal telemetry | Record lane, timing, model identity, counts, and typed errors; never raw prompts, answers, messages, vectors, or tool payloads. |
| Honest privacy copy | “Processed on this device” appears only when both retrieval and generation stayed local. |
Performance, battery, storage, and cost budgets#
On-device inference removes per-token model charges; it does not make computation free. Claire pays with application complexity, test coverage, support, and synchronization, while the user pays with storage, battery, memory, and heat. Rollout gates must measure all four.
| Metric | Initial target | Gate |
|---|---|---|
| Warm local first token | p50 ≤1.5s; p95 ≤4s on reference devices | Compare against managed streaming, by device class. |
| Local hybrid retrieval | p95 ≤600ms after index readiness | Measured without blocking React Native's UI thread. |
| Ask Claire completion | p95 ≤12s for golden queries | Quality and citation gates take precedence over speed. |
| Incremental indexing | New text searchable within 10s while app is active | Never delays message receipt or screen navigation. |
| Background generation | Zero while low-power, thermally constrained, or user-active | Resume from the last committed cursor. |
| Large model download | Never automatic | Show exact size, license, network, storage, and delete controls. |
| Cloud model cost | Zero on successful local-only requests | Fallback usage remains metered by feature and reason. |
These are candidate launch targets, not measurements. The reference-device matrix and current managed baseline must be captured before they become release gates.
Evaluation and release gates#
A local model ships by task, device class, OS version, model revision, and language—not because a single demo looked good. Apple may update the built-in model with the OS, so the same eval suite runs again for every supported revision.
| Suite | Measures | Blocking failure |
|---|---|---|
| Ask Claire golden queries | Identity recall, plans, people/location, ambiguity, citation precision and recall | Unsupported claim, wrong person/date, fabricated citation, or material regression against managed baseline |
| Multilingual retrieval | English/Spanish/Cyrillic and mixed-script recall | The relevant evidence falls outside the candidate set without an explicit fallback |
| Loop extraction | Create/update/close F1, duplicate rate, ownership, deadline normalization | Wrong close, other-person group Loop surfaced, user edit overwritten, or evidence missing |
| Action proposals | Tool selection, argument validity, approval boundary, prompt injection resistance | Any unapproved effect or credential/tool exposure |
| Device performance | Cold/warm latency, memory, battery, thermal state, cancellation, app lifecycle | Crash, UI stall, sustained thermal escalation, or unrecoverable model/index state |
Start with anonymized fixtures and the existing synthetic Loop corpus, then run the Lucas two-device flow after the action is explicitly approved under its testing specification. Local and managed candidates receive identical evidence packages so generation quality can be compared independently from retrieval quality.
Delivery plan#
Spike 0 · Reference benchmark
Prove the model before changing product architecture
Integrate Expo AI Kit in a development-only screen, run saved Ask Claire evidence and Loop windows, and capture quality, latency, memory, heat, and cancellation on real devices.
PR 1 · Runtime boundary
Add the on-device provider and capability router
Create a Claire-owned client runtime, model availability UI, one-generation queue, typed errors, and managed fallback without changing answers yet.
PR 2 · Local generation pilot
Generate from existing server-selected evidence
Keep retrieval unchanged, send a bounded evidence package to the phone, stream locally, validate citations, and compare against the managed answer path.
PR 3 · Lexical index
Make exact and time-bounded search local
Add encrypted index state, content hashes, deletion propagation, local query planning, rank fusion, and partial-history disclosures.
PR 4 · Semantic and memory index
Add local embeddings and deterministic personal memory
Persist model-versioned vectors, multilingual namespaces, relationship metrics, plan facts, incremental indexing, and resumable rebuilds.
PR 5 · Ask Claire local-first
Route eligible golden queries entirely on-device
Enable local retrieval plus generation by cohort, retain explicit escalation, and expose trustworthy privacy and freshness states.
PR 6 · Loop assistant local
Move user-initiated Loop questions and drafts
Use a small read/propose tool set over the local Loop replica with no execution authority.
PR 7 · Leased Loop detection
Pilot local extraction with server guarantees
Add cursor-range leases, shared deterministic guards, idempotent operation batches, server revalidation, and expiry takeover.
Decisions to resolve in the benchmark#
| Question | Default until measured |
|---|---|
| Which iOS model? | Apple's built-in Foundation Model on eligible devices. |
| Downloadable fallback? | No automatic download; evaluate separately by device tier. |
| Which multilingual embedder? | Keep managed semantic fallback until cross-script recall is proven. |
| Who owns Loop detection? | Server, unless a device holds a valid cursor-range lease. |
| Should local answers sync? | Yes, through an idempotent outbox; user may later choose device-only history. |
| When may Claire escalate? | Only under the saved privacy policy, with a typed reason and bounded evidence payload. |
External implementation references#
- Expo AI Kit platform support and fallback requirements
- Built-in and downloadable model lifecycle
- Local embeddings, RAG, model identity, and language namespaces
- Expo AI Kit provider for Vercel AI SDK
- Tool calling, bounded repair, and human approval
- Apple Foundation Models availability and fallback guidance