ClaireDocs

Ask Claire

Search the project, not the web.

Answers are grounded in Claire’s published documentation and include the source pages used.

⌘/Ctrl + J opens Ask Claire anywhere in docs.

Claire on-device intelligence specification

A local-first architecture for on-device models, private message search, Ask Claire, Loop detection, and safe action proposals.

DraftresearchReviewed 2026-09-09View source ↗
Local first
Default inference route

When model, index, language, and battery policy permit

1
Canonical Loop writer

A lease prevents local and server detectors racing

0
Unapproved writes

Local models can read and propose, never silently act

Per device
Embedding namespace

Platform, model, revision, dimensions, and script are pinned

Product outcome and boundary#

The user experiences one Claire. Whether an answer was generated on the phone or by managed AI is an execution detail, except where it changes privacy, availability, cost, or quality enough that the user should know.

On-device Claire means message retrieval, ranking, prompt construction, generation, and citation selection can complete without sending message text to a model provider. It does not mean the entire Claire product is serverless. WhatsApp, Telegram, Instagram, and other bridges still need network services; cross-device synchronization needs a canonical source; and external actions still call the relevant API.

CapabilityCan run locally?Product boundary
Ask Claire answer generationYesUse a local model when it clears capability and quality checks.
Exact and semantic message searchYesThe phone needs a complete-enough encrypted index and an explicit freshness state.
Relationship and location rankingYesCalculate facts deterministically; let the model explain them.
Loop extraction and reconciliationYes, conditionallyInference can be local; canonical persistence and race prevention stay coordinated.
Loop-scoped drafts and questionsYesThis is the lowest-risk Loop feature to move first because it is user initiated.
Calendar, booking, and messaging actionsProposal: yes; execution: dependsNative actions can stay on-device; external services require an authenticated API.
Bridge ingestion and multi-device syncNoKeep the existing Claire server, Matrix bridges, and synchronization layer.

Why the current architecture can support it#

Claire already has most of the boundaries a local lane needs. Mobile uses encrypted SQLite, caches recent timelines, can optionally backfill full history, stores contacts and conversation settings, and preserves stable message IDs. Ask Claire v2 already separates deterministic query planning, retrieval, model generation, citations, persistence, and stream transport. Loops already separate their free gate, model extraction, deterministic relevance, reconciliation, and storage.

Existing componentReusable locallyRequired change
mobile-cache.native.tsEncrypted messages, contacts, settings, cursors, full-history modeAdd AI index tables, model state, completeness, and narrow local queries.
conversation-assistant-query.tsIntent and relative-date planningMove pure planning into a shared package usable by client and server.
conversation-assistant.tsPrompt, citation, result, and stream conceptsSplit server persistence from a platform-neutral answer engine.
loops/loop-gate.ts and loop-reconciler.tsPure, deterministic filters and safety checksMove shared logic and types into a client-safe workspace package.
loops/loop-detector.tsBounded window and structured operation contractSeparate extraction from Supabase reads and canonical writes.
AI SDK provider registryStable generation and embedding vocabularyAdd a client-only on-device provider adapter and capability router.

Target architecture#

One product, two inference lanes — The phone owns its encrypted replica and local AI index. A policy router chooses local inference or the existing managed path. The server remains authoritative for synchronized records and bridges.

The policy router makes the choice before retrieval. It must never begin locally, discover after generation that required evidence was missing, and silently produce a weaker answer. Index readiness and scope coverage are explicit inputs to routing.

ts
type AIExecutionLane = 'on_device' | 'managed' | 'deterministic_only';

interface DeviceAICapabilities {
  generation: 'ready' | 'downloadable' | 'unsupported' | 'busy';
  modelId: string | null;
  modelRevision: string | null;
  supportedLanguages: string[];
  contextTokens: number | null;
  indexState: 'empty' | 'building' | 'partial' | 'ready' | 'stale';
  indexedThroughCursor: number;
  fullHistory: boolean;
}

interface ExecutionDecision {
  lane: AIExecutionLane;
  reason:
    | 'local_ready'
    | 'model_unavailable'
    | 'index_incomplete'
    | 'language_unsupported'
    | 'context_too_large'
    | 'local_busy'
    | 'quality_escalation'
    | 'offline_search_only';
}

On-device model runtime#

Use Expo AI Kit behind a Claire-owned adapter. Its built-in route uses Apple Foundation Models on eligible iOS devices and ML Kit on eligible Android devices. It can also run downloadable LiteRT-LM models and exposes an AI SDK provider. Claire must not import the package throughout feature code or couple persisted records to its native types.

Default

Built-in OS model

No separate multi-gigabyte app asset. Prefer it when available, the language is supported, and device evals pass.

Optional

Downloadable model

User-initiated only, with size, license, Wi-Fi, storage, and deletion controls shown before downloading.

Fallback

Managed model

Preserve the current provider registry for unsupported hardware, difficult retrieval, larger context, and quality escalation.

Plain-text answers may stream token by token. Structured output and tool calls can be buffered until validation succeeds. Only one text generation may run at a time on a device, so Ask Claire, Loop extraction, summaries, and drafts share a single priority queue. A user-visible request outranks background Loop processing; embeddings may proceed independently when the native backend supports that safely.

PriorityWorkScheduling rule
P0Active Ask Claire answer or user-requested Loop draftStart immediately; allow the user to cancel.
P1Local retrieval and query planningRun concurrently where safe; never block the UI thread.
P2Foreground incremental Loop extractionYield when a P0 request begins.
P3Embedding backfill and relationship recomputationRun in bounded chunks while idle; pause on heat, low power, or app background limits.

Private local search index#

Local generation is only useful when retrieval is trustworthy. Claire must search the local replica rather than placing an entire message archive in the model context. The index combines exact search, semantic search, time and conversation filters, and deterministic personal-memory tables.

StoreMinimum fieldsPurpose
local_ai_index_stateuser, corpus cursor, model ID, revision, dimensions, language/script, status, errorMakes completeness, compatibility, and rebuild decisions explicit.
local_message_lexicalmessage ID, chat ID, normalized content, timestamp, sender, platformExact names, quoted phrases, dates, usernames, and fallback retrieval.
local_message_embeddingsmessage ID, vector, namespace, content hash, indexed timestampMeaning-based retrieval without re-embedding unchanged messages.
local_relationship_metricschat ID, 30/90-day counts, reciprocity, last interaction, computed cursorRanks people with deterministic evidence rather than model intuition.
local_plan_factssource IDs, participants, normalized time range, status, confidenceAnswers plan questions without scanning every conversation.

The existing encrypted SQLite key protects these tables at rest. Local indexes are derived data with the same sensitivity as the underlying messages: clearing local data, signing out, deleting a message, or disabling full history must remove or trim matching index rows.

Retrieval pipeline#

Local hybrid retrieval — A deterministic query plan fans out to lexical, vector, relationship, and plan indexes. Rank fusion chooses evidence windows, and the model receives only those windows.
  1. Parse intent, relative dates, named people, locations, and preferred conversations without a model where possible.
  2. Run lexical and semantic retrieval in parallel inside the requested user, chat, and date boundaries.
  3. Fuse ranks without directly comparing unrelated score scales.
  4. Expand top anchors into short chronological windows so plans and pronouns have context.
  5. Generate from stable source labels and discard any citation label that does not resolve.
  6. Return index scope and freshness with the answer so partial history is never presented as exhaustive.

Multilingual and cross-script search#

This is a launch blocker, not an optimization. On iOS, Expo AI Kit uses separate Apple embedding assets for Latin, Cyrillic, and CJK scripts. Those vectors cannot be compared across model identities. An English question such as “Who was the Russian girl?” may need to find a Cyrillic message, and a single-vector-space implementation can miss it.

The index must therefore support multiple namespaces and one of these evaluated strategies:

  1. Preferred: ship or integrate a compact multilingual embedding model with one cross-language vector space.
  2. Fallback: detect corpus scripts, produce query variants for each supported language/script, search each compatible namespace, then fuse the results.
  3. Safety valve: route cross-script questions to managed retrieval when local recall is not proven.

Ask Claire on-device flow#

PhaseLocal behaviorFallback trigger
1. Capability checkRead model readiness, language support, index coverage, thermal and power policy.Unsupported, busy beyond timeout, stale/incomplete required scope.
2. Query planningCompile intent, time range, people/location needs, and conversation scope.Ambiguous compound request that the deterministic planner cannot represent.
3. RetrievalSearch local indexes and form evidence windows with stable message IDs.Cross-script recall risk, missing full history, or no adequate evidence.
4. GenerationStream a concise answer from evidence and local memory only.Context exceeds model limit, inference error, or eval-classified hard query.
5. ValidationResolve every citation and constrain suggested actions to evidence.Invalid structured output after bounded repair.
6. PersistenceShow immediately, write to the local outbox, then synchronize the turn.Offline is acceptable; retry by stable request ID when connectivity returns.

A cloud fallback is a new request lane, not a hidden continuation. Before message evidence leaves the device for managed inference, apply the user's AI privacy setting and show a compact disclosure when required. Do not upload the full local index; send only the bounded evidence package selected for that answer.

The existing Ask Claire UI, stream parts, answer cards, citations, Stop behavior, action proposals, and thread history remain unchanged. The result adds execution metadata: lane, model identity, index coverage, fallback reason, and measured latency. That metadata is diagnostic and privacy-visible but does not clutter the normal conversation.

Can Loop generation run locally?#

“Loop generation” is three distinct workloads. They should not move as one switch.

WorkloadLocal verdictReason
Loop-scoped assistantMove firstUser initiated, one loop, small tool set, recent context, and proposals are inert.
Incremental create/update/close extractionViable with a leaseThe structured operation contract fits small models, but only one detector may own a chat cursor.
Continuous background detectionHybridA phone can be offline, suspended, hot, or closed while bridge messages continue arriving.
Historical Loop backfillOpt-in local jobPotentially large and battery-intensive; resumable foreground/charging work with managed fallback.

Canonical ownership and synchronization#

Claire Cloud should continue to guarantee that new messages are eventually examined even if the mobile app never opens. When a device is active and eligible, it may take a short-lived detection lease for a chat and run the extraction locally. The lease binds the user, chat, starting cursor, ending cursor, model identity, and expiry.

ts
interface LoopDetectionLease {
  leaseId: string;
  userId: string;
  chatId: string;
  fromCursor: string | null;
  throughCursor: string;
  owner: 'device' | 'server';
  expiresAt: string;
}

interface ProposedLoopBatch {
  requestId: string;       // idempotency key
  leaseId: string;
  operations: LoopOp[];    // create, update, close
  evidenceMessageIds: string[];
  model: { id: string; revision: string | null };
}
  1. The server grants one owner for a bounded cursor range.
  2. The device builds the same window and runs the same deterministic gate.
  3. The local model returns structured operations, never direct database mutations.
  4. The device runs shared relevance and reconciliation guards before submission.
  5. The server rechecks ownership, live Loop IDs, evidence IDs, confidence policy, and idempotency inside one transaction.
  6. If the lease expires or the device disappears, the server safely processes the range.

Self-hosted or explicitly device-only deployments may make the local database canonical while offline and merge later. Claire Cloud should not: a local-only background promise would miss messages whenever iOS suspends the app.

Local Loop model contract#

Preserve the existing LoopOp schema and all deterministic guards. Reduce the local prompt when necessary, but do not weaken these invariants:

  • A conversation thread produces one evolving Loop, not one item per message.
  • Every create or close cites real message IDs from the supplied window.
  • Silence never closes a Loop.
  • Unknown Loop IDs, invalid evidence, low confidence, and duplicate creates fail closed.
  • Group relevance and self-identity remain deterministic.
  • User edits outrank later model proposals.

On-device models are weaker at large schemas and tool selection, so Loop extraction should remain one shallow structured-output request rather than a broad agent with many tools. A bounded repair attempt is allowed; failure leaves the cursor uncommitted so the server or a later local pass can retry.

Local intelligence and the future action system#

Local models fit the action system when they are planners, not authorities. Ask Claire or a Loop may propose a typed action. Claire then validates the proposal, displays the exact destination and fields, obtains approval, and hands it to a trusted native or plugin executor.

Propose, approve, execute — The model can only create an inert proposal. Policy validation and the user approval boundary sit before native or plugin execution.
ActionPlanningExecution
Open a cited conversationLocalLocal navigation
Draft a replyLocalInsert locally; sending still uses the bridge
Create an Apple calendar eventLocalNative API after permission and confirmation
Schedule through Google CalendarLocalPlugin/API after confirmation
Book a restaurant or serviceLocal with narrow toolsExternal plugin/API after confirmation

Present only the tools relevant to the current task. Small on-device models should never choose from the entire installed plugin catalog. A deterministic capability router selects a small read/propose tool set; write credentials and execution functions stay outside the model runtime. See the Plugin system specification.

Privacy, security, and data lifecycle#

RequirementAcceptance rule
No accidental model egressThe on-device lane contains no network-backed model provider and tests fail on any prompt upload.
Encrypted derived dataText indexes, vectors, facts, prompts, and local turns live only in the encrypted user database.
Deletion propagationMessage deletion, sign-out, cache clearing, and history-retention changes delete corresponding derived rows.
Untrusted conversation contentMessages are always data; they cannot alter tools, policy, routing, approvals, or system instructions.
Minimal telemetryRecord lane, timing, model identity, counts, and typed errors; never raw prompts, answers, messages, vectors, or tool payloads.
Honest privacy copy“Processed on this device” appears only when both retrieval and generation stayed local.

Performance, battery, storage, and cost budgets#

On-device inference removes per-token model charges; it does not make computation free. Claire pays with application complexity, test coverage, support, and synchronization, while the user pays with storage, battery, memory, and heat. Rollout gates must measure all four.

MetricInitial targetGate
Warm local first tokenp50 ≤1.5s; p95 ≤4s on reference devicesCompare against managed streaming, by device class.
Local hybrid retrievalp95 ≤600ms after index readinessMeasured without blocking React Native's UI thread.
Ask Claire completionp95 ≤12s for golden queriesQuality and citation gates take precedence over speed.
Incremental indexingNew text searchable within 10s while app is activeNever delays message receipt or screen navigation.
Background generationZero while low-power, thermally constrained, or user-activeResume from the last committed cursor.
Large model downloadNever automaticShow exact size, license, network, storage, and delete controls.
Cloud model costZero on successful local-only requestsFallback usage remains metered by feature and reason.

These are candidate launch targets, not measurements. The reference-device matrix and current managed baseline must be captured before they become release gates.

Evaluation and release gates#

A local model ships by task, device class, OS version, model revision, and language—not because a single demo looked good. Apple may update the built-in model with the OS, so the same eval suite runs again for every supported revision.

SuiteMeasuresBlocking failure
Ask Claire golden queriesIdentity recall, plans, people/location, ambiguity, citation precision and recallUnsupported claim, wrong person/date, fabricated citation, or material regression against managed baseline
Multilingual retrievalEnglish/Spanish/Cyrillic and mixed-script recallThe relevant evidence falls outside the candidate set without an explicit fallback
Loop extractionCreate/update/close F1, duplicate rate, ownership, deadline normalizationWrong close, other-person group Loop surfaced, user edit overwritten, or evidence missing
Action proposalsTool selection, argument validity, approval boundary, prompt injection resistanceAny unapproved effect or credential/tool exposure
Device performanceCold/warm latency, memory, battery, thermal state, cancellation, app lifecycleCrash, UI stall, sustained thermal escalation, or unrecoverable model/index state

Start with anonymized fixtures and the existing synthetic Loop corpus, then run the Lucas two-device flow after the action is explicitly approved under its testing specification. Local and managed candidates receive identical evidence packages so generation quality can be compared independently from retrieval quality.

Delivery plan#

  1. Spike 0 · Reference benchmark

    Prove the model before changing product architecture

    Integrate Expo AI Kit in a development-only screen, run saved Ask Claire evidence and Loop windows, and capture quality, latency, memory, heat, and cancellation on real devices.

  2. PR 1 · Runtime boundary

    Add the on-device provider and capability router

    Create a Claire-owned client runtime, model availability UI, one-generation queue, typed errors, and managed fallback without changing answers yet.

  3. PR 2 · Local generation pilot

    Generate from existing server-selected evidence

    Keep retrieval unchanged, send a bounded evidence package to the phone, stream locally, validate citations, and compare against the managed answer path.

  4. PR 3 · Lexical index

    Make exact and time-bounded search local

    Add encrypted index state, content hashes, deletion propagation, local query planning, rank fusion, and partial-history disclosures.

  5. PR 4 · Semantic and memory index

    Add local embeddings and deterministic personal memory

    Persist model-versioned vectors, multilingual namespaces, relationship metrics, plan facts, incremental indexing, and resumable rebuilds.

  6. PR 5 · Ask Claire local-first

    Route eligible golden queries entirely on-device

    Enable local retrieval plus generation by cohort, retain explicit escalation, and expose trustworthy privacy and freshness states.

  7. PR 6 · Loop assistant local

    Move user-initiated Loop questions and drafts

    Use a small read/propose tool set over the local Loop replica with no execution authority.

  8. PR 7 · Leased Loop detection

    Pilot local extraction with server guarantees

    Add cursor-range leases, shared deterministic guards, idempotent operation batches, server revalidation, and expiry takeover.

Decisions to resolve in the benchmark#

QuestionDefault until measured
Which iOS model?Apple's built-in Foundation Model on eligible devices.
Downloadable fallback?No automatic download; evaluate separately by device tier.
Which multilingual embedder?Keep managed semantic fallback until cross-script recall is proven.
Who owns Loop detection?Server, unless a device holds a valid cursor-range lease.
Should local answers sync?Yes, through an idempotent outbox; user may later choose device-only history.
When may Claire escalate?Only under the saved privacy policy, with a typed reason and bounded evidence payload.

External implementation references#