ClaireDocs

Ask Claire

Search the project, not the web.

Answers are grounded in Claire’s published documentation and include the source pages used.

⌘/Ctrl + J opens Ask Claire anywhere in docs.

Ask Claire v2: streaming, retrieval, and cost specification

A staged implementation specification for a faster, cheaper, provider-portable Ask Claire with grounded streaming answers and a safe path to plugin actions.

Currentin progressReviewed 2026-09-15View source ↗
3
Golden personal-memory queries

Identity, plan recall, and people planning

1
Normal generation step

Escalate only for compound questions

4
Evidence windows maximum

Chronological context, not isolated snippets

0
Unapproved external writes

Every outreach action remains a proposal

Purpose and product boundary#

Ask Claire is the user-initiated research and assistance surface across connected conversations. It finds relevant evidence, answers with citations, and helps the person decide what to do next. It does not silently send messages, book meetings, edit external systems, or treat message text as authorization.

This specification covers the interactive Ask Claire tab, the conversation-scoped Ask Claire surface, and the shared server engine behind both. It extends the AI platform specification, the AI model and cost model, and the plugin system specification. Those documents remain authoritative for account billing, provider credentials, plugin permissions, approvals, and external execution.

GoalRequired outcome
FastShow useful progress immediately and stream answer text as it is generated.
GroundedEvery factual conversation claim links to the message evidence that supports it.
Low costUse no model when deterministic data is sufficient and one generation step for a normal answer.
PortableSelect providers and models by Claire task profile, not by imports in the product service.
SafeRead directly, propose writes, and require the plugin policy system to authorize every external effect.
DurableA reconnect, retry, server restart, or duplicate submission cannot lose or repeat a completed turn.

Non-goals for v2: autonomous background agents, unrestricted MCP execution, model-generated UI, silently sending messages, and storing a full model transcript as the database source of truth.

Baseline audited before the v2 implementation#

The September 8 baseline was a persisted RAG chat. New messages were embedded withtext-embedding-3-small and stored in conversation_message_embeddings. A question launches lexical and vector retrieval in parallel, merges up to 12 individual message hits, adds the last six assistant turns and selected conversation instructions, calls OpenAI Chat Completions for JSON, validates cited indices and two navigation actions, then stores the user and assistant turns. The table records the gaps that drove v2; the shipped state and remaining launch work are summarized indocs/ASK_CLAIRE_V2_LAUNCH_READINESS.md.

AreaAlready presentGap
RetrievalUser-scoped full-text search, pgvector HNSW, strict chat scope, preferred @ chatsRaw lexical rank and cosine similarity are compared directly; no threshold, neighborhood expansion, or retrieval evaluation.
GroundingThe model selects source indices; server caps and validates them; source cards open the original messageWhen no source index is returned, the current fallback still displays the first retrieved sources, which can imply support the model did not claim.
ThreadsPersisted global threads and one retained thread per conversationEvery answer reads the entire thread even though only six turns enter the prompt.
IndexingContent hashes, resumable backfill, and live fire-and-forget indexingBackfill embeds one message per provider call, stops after 1,000 rows, and only selects missing—not changed—embeddings.
Actionsopen_conversation and open_calendar suggestions are source-validated and inertopen_calendar only launches the native calendar; generated title/time are not written or prefilled, and this is not the plugin action system.
AI runtimeAI SDK 7 provider registry, structured generation, and a bounded loop agent already exist under services/aiAsk Claire still imports the OpenAI SDK directly and cannot use the shared providers, failover, usage, or timeout behavior.
ClientOptimistic user turn, persisted history, cached thread list, mobile citationsThe response is all-or-nothing; no token streaming, stop button, stream recovery, or shared mobile/desktop message-part renderer.

Representative questions and capability bar#

Ask Claire succeeds only if it can answer the questions people naturally ask when they remember part of a conversation but not where it happened. The following questions are release-defining workloads, not aspirational examples. They become versioned golden queries in the evaluation corpus and must work across every enabled messaging platform.

Partial today

Identity from a remembered detail

“Who was the Russian girl I was talking about in the past couple of days?”

Semantic search may find “Russian,” and open_conversation can navigate to a cited chat. It cannot reliably join the person, attribute, name, and time window across nearby messages.

Partial today

Plans on a relative date

“I made plans on Saturday, but with whom?”

Search may retrieve “Saturday,” “dinner,” or “see you.” It does not resolve the date in the user's timezone or distinguish tentative, confirmed, changed, and cancelled plans.

Partial today

People planning and outreach

“Who are my closest people in Mexico City or Brooklyn, and can we reach out?”

Ask Claire can use current contact profile facts and communication context, but coverage and freshness remain uneven. Approved, per-recipient outreach proposals remain future plugin work.

Required answer contracts#

Identity lookup

Identify a discussed person

Resolve the time range, search descriptor anchors, and expand each hit into a chronological window. Return the likely person, uncertainty, citations, and a trusted link to the best source conversation.

Temporal memory

Recall a plan

Resolve the local date, group nearby planning turns, and classify the plan as tentative, confirmed, changed, or cancelled. Preserve multiple credible candidates.

Relationship query

Build a people shortlist

Filter provenance-backed location facts, then rank deterministic communication metrics. Explain why each person appears and when their location was last confirmed.

Plugin proposal

Prepare outreach

Confirm recipients and generate one editable draft per destination. Show the exact proposal payload; send nothing until the user explicitly approves it.

Structured query plan#

Compile every question into a Claire-owned plan before retrieval. Deterministic code handles relative dates, explicit people, platform filters, trusted @ selections, and common workload phrases. Use at most one low-cost structured planner call only for compound or ambiguous questions that the deterministic compiler cannot represent. The planner emits constraints; it never receives database credentials or performs reads.

The personal-memory query path — Claire resolves trusted constraints before retrieval, assembles cited evidence, and only then generates an answer or proposal.
json
{
  "intent": "identity_lookup | plan_recall | people_ranking | conversation_qa | action_request",
  "timeRange": { "start": "ISO-8601", "end": "ISO-8601", "timezone": "IANA" },
  "people": [{ "name": null, "descriptors": ["Russian"], "chatId": null }],
  "location": { "city": "Mexico City", "asOf": "ISO-8601" },
  "planStatus": ["proposed", "confirmed", "changed", "cancelled"],
  "ranking": { "kind": "relationship_strength", "limit": 12 },
  "requestedAction": { "type": "draft_outreach", "requiresApproval": true }
}
  • Relative dates use the account timezone and an explicit request timestamp. Store the resolved range with the turn so a replay does not reinterpret “Saturday.”
  • “Past couple of days” defaults to the preceding 48 hours through the request time; product copy may display the concrete dates Claire searched.
  • If “Saturday” could reasonably mean two dates, prefer the nearest upcoming or most recent Saturday based on tense and show the chosen date. Ask a clarification only when candidates remain materially ambiguous.
  • All filters are enforced in database queries. A prompt instruction to consider only a date, person, or chat is not a security or correctness boundary.

People memory and relationship ranking#

Keep user-entered contact_profiles as authoritative overrides, but replace a single timeless location string as the only memory representation. Derived facts need provenance, time, and confidence so Claire can answer “still based in Brooklyn” without treating an old trip or previous home as current.

RecordMinimum fieldsRule
assistant_people and assistant_person_linksUser-scoped person ID, preferred display name, aliases, optional linked Claire contacts/platform identities, merge target, confidence, and source message IDsA person may be a contact or merely someone discussed inside another conversation. Never assume that the chat participant and the person being discussed are the same individual.
assistant_person_factsPerson/contact identity, predicate, normalized and displayed value, valid-from, valid-until, observed-at, last-confirmed-at, confidence, provenance type, source message IDsUser edits override inferred facts. Conflicting values coexist until resolved; expired or stale facts are shown with their date rather than silently discarded.
assistant_relationship_metricsPerson/chat, window start/end, sent/received counts, active days, last contact, response balance, one-to-one/group share, user relationship labelDerived asynchronously from metadata. Do not send message content to a model to calculate communication frequency.

“Closest” is an explainable product ranking, not a claim about the user's emotions. A default score may combine recency, active days, reciprocal exchanges, sustained contact, and a user-defined relationship label. Exclude raw message volume dominance, group-chat noise, automated accounts, and one-sided bursts. Return component explanations such as “you spoke on 9 days this month and last talked yesterday,” never the opaque score.

Normalize safe lookup aliases such as “CDMX,” “Ciudad de México,” and “Mexico City,” while retaining the displayed source wording. Nationality, precise location, health, religion, sexuality, and similar inferred attributes receive an explicit sensitivity class. Use them ephemerally for on-demand retrieval by default; persist them only under the configured memory policy, with provenance, retention, correction, and deletion.

Plan and event memory#

Maintain structured plan candidates separately from calendar events and Loops. A plan candidate is evidence that people discussed doing something; it is not proof that an event exists on a calendar.

Each assistant_plan_candidate stores the user, conversation, participants, normalized start/end when known, original date phrase, timezone, place/title, status, confidence, extraction version, source message IDs, and superseding candidate. Detect candidates asynchronously on changed conversation windows and support on-demand extraction for an uncovered date range. Cancellation or rescheduling creates a new state linked to the earlier evidence instead of overwriting history.

Target architecture#

Keep retrieval and persistence in Claire-owned services. The AI SDK is the provider and stream runtime, not the product architecture. Product code asks for a task profile and consumes Claire types; only apps/server/src/services/ai/ imports provider packages.

AI SDK adoption decision#

LayerDecisionReason
AI SDK CoreAdoptThe repository already uses it. streamText, generateText, Output.object, embedMany, tools, timeouts, usage, and provider-neutral models remove Ask Claire's duplicate OpenAI-only path.
Direct provider packagesAdoptUse @ai-sdk/openai, @ai-sdk/azure, and @ai-sdk/openai-compatible directly through Claire's registry. Add dedicated providers only with capability tests.
AI SDK UI stream protocolAdoptIt carries text, status data, sources, tool/proposal parts, terminal metadata, and errors over SSE and works with Node/Express.
@ai-sdk/reactPilot behind useAssistantStreamThe official Expo integration supports streaming with expo/fetch. Claire still owns server thread IDs, offline snapshots, citations, and persistence; the hook must remain replaceable.
AI SDK RSCDo not useClaire is Expo/Electron, not a React Server Components product, and the RSC API is not needed for data-only message parts.
Vercel AI GatewayOptional laterThe SDK does not require Gateway. Making it mandatory would conflict with BYOK, self-hosted, and local execution modes.
SDK tool approvalTransport aid onlyClaire's database proposal, payload hash, policy decision, approval, execution, and receipt remain authoritative. SDK approval parts cannot replace durable cross-device authorization.

Model portability is real for text generation and embeddings only when Claire tests the required capability. Maintain a capability record per configured model: streaming, structured output, tool calling, maximum context, embedding dimensions, usage reporting, and simulated-stream fallback. Never assume that every provider implements the abstraction equally well.

Answer request lifecycle#

  1. The client creates a UUID requestId and sends one request. A missing global thread is created inside this request.
  2. The server authenticates the user, checks AI policy and per-user budget, validates scope ownership, and idempotently inserts the user turn plus an empty assistant turn with status=streaming.
  3. The server immediately opens the stream and emits a sanitized retrieving state with the real persisted thread and turn IDs.
  4. The context router chooses deterministic sources. “Catch me up” reads a recent window; “what is unresolved?” reads Loops; scoped factual search uses hybrid retrieval. This choice does not require another model.
  5. Retrieval creates a bounded evidence pack. If no candidate passes the relevance floor, emit a deterministic no-evidence answer and spend no generation tokens.
  6. The budget service reserves the maximum permitted request cost. The model router selects an eligible model profile and provider candidate.
  7. streamText streams the grounded answer. Citations are stable source IDs, not array positions that can change between steps.
  8. On completion, validate citation markers and proposal parts, persist the final answer and usage atomically, settle the reservation, and emit finish.
  9. On provider failure before the first text delta, try the next eligible provider. After text has been shown, do not silently replay through another model; mark the turn failed/partial and offer an idempotent retry.

Streaming contract#

Add content negotiation to the existing message endpoints. Accept: text/event-streamselects the v2 AI SDK UI stream; Accept: application/json preserves the existing response during rollout. Do not create separate business logic for the two transports.

ts
type AskClaireRequest = {
  requestId: string;             // client UUID, unique per user
  threadId?: string;             // omitted to create a global thread
  question: string;
  scope:
    | { mode: 'global' }
    | { mode: 'preferred'; chatIds: string[] } // 1..5
    | { mode: 'strict_chat'; chatId: string };
};

type ClaireStreamData = {
  status: { phase: 'retrieving' | 'answering' | 'saving' };
  thread: { threadId: string; userTurnId: string; assistantTurnId: string };
  sources: { citations: AssistantCitation[]; indexStatus: AssistantIndexStatus };
  proposal: { proposalId: string; capabilityId: string; requiresApproval: boolean };
  finish: { finishReason: string; provider: string; model: string };
  error: { code: string; retryable: boolean };
};

Emit source metadata before answer text so citations can render without waiting for the final token. The text itself should be normal readable Markdown with stable citation marks such as [S1]. Do not stream serialized JSON into the answer bubble. Use custom data parts for Claire metadata and future proposals.

Persistence and disconnects#

  • Persist once before generation and once on terminal completion; never write one database update per token.
  • Store pending | streaming | completed | failed | cancelled on assistant turns and make (user_id, request_id) unique.
  • A duplicate request returns or resumes the existing turn instead of creating another model call.
  • Wire the client disconnect to an AbortSignal. Product policy decides whether a backgrounded mobile request gets a short grace period or is cancelled.
  • For v2 launch, reconnect loads the persisted terminal turn. Resumable mid-token streams are deferred until production evidence justifies the storage and protocol complexity.

Client behavior#

  • Add useAssistantStream as the only UI-facing transport adapter. It maps SDK UIMessage parts to the existing AssistantTurn renderer.
  • Use expo/fetch on iOS and Android; use the host fetch implementation in Electron. Confirm proxy buffering is disabled in staging and production.
  • Batch text deltas into UI updates approximately every 30–50ms to avoid a React Native render per token.
  • Show Stop while generating, keep the composer editable after cancellation, and auto-scroll only while the user remains near the bottom.
  • Do not announce every token to screen readers. Announce phases, then the completed answer.

Provider and model routing#

Generalize the existing loop-only provider registry into Claire's shared AI service. Ask Claire must request a profile, not a vendor model string. Provider credentials and model IDs remain configuration.

ProfileUseBudget behavior
assistant_groundedDefault cited synthesis and scoped chat questionsCheapest model that clears the grounding and citation eval gates; one generation step.
assistant_reasoningAmbiguous multi-chat relationship analysis or complex reconciliationOpt in through deterministic complexity signals or explicit user quality setting; never a retry merely because the answer was short.
assistant_agentFuture read tools and plugin proposalsStrong tool-capable model, maximum three steps for global Ask Claire.
embedding_messagesMessage and query vectorsFixed 1536-dimensional contract until a versioned re-index migration exists.

A provider fallback is eligible only when it matches the profile's capabilities and data policy. Failover may happen before the first visible delta. Mid-stream failure is surfaced honestly because concatenating two providers' answers can duplicate text, change claims, and invalidate citations.

Retrieval v2#

Deterministic context router#

Question shapePrimary contextModel calls before answer
Catch up / summarize this chatMost recent bounded chronological chat window0
Open commitments, questions, or plansLive Loops plus their cited evidence0
Find a remembered fact or phraseHybrid lexical + semantic retrieval0
Tone or relationship patternRecent windows plus selected older semantic evidence0
Explicit person/platform/date filtersApply structured metadata filters before ranking0
Discussed person identified by attributesTime-bounded descriptor retrieval plus conversation-window expansion0 normally; ≤1 planner call when ambiguous
Plan recall for a relative dateResolved local date range plus plan candidates and supporting message windows0
People ranked by relationship and locationTemporal person facts plus precomputed relationship metrics0

Routing should be conservative deterministic code. A cheap classifier adds latency and cost to every question; introduce one only if evals show that deterministic routing cannot reach the required recall.

Hybrid ranking and evidence assembly#

  1. Normalize the query and resolve explicit chat, platform, participant, and date filters from trusted UI IDs where possible.
  2. For relative time language, compile a concrete half-open timestamp range from the account timezone and request time. Apply it before text or vector ranking.
  3. Run full-text and vector retrieval concurrently. Add a GIN index over the full-text expression or a maintained tsvector; the current SQL recomputes to_tsvector without a matching index.
  4. Fuse result positions using reciprocal-rank fusion. Do not compare PostgreSQL text rank directly with cosine similarity.
  5. Apply calibrated relevance floors by route. Vector search must not always return “the best 12” when all 12 are poor.
  6. Diversify by chat and time, except in strict-chat mode. Avoid filling the prompt with near-duplicate adjacent messages.
  7. Expand each winning hit into a small chronological window around the message, preserving reply/thread metadata and sender identity.
  8. Fit at most four evidence windows into the context budget. The UI may show multiple messages inside one source window while citing stable message IDs.
  9. For identity and plan recall, cluster windows by person, chat, or plan candidate and preserve credible alternatives. Do not allow one high-scoring message to erase a second plausible answer.

Grounding contract#

  • The model can cite only source IDs present in the evidence pack.
  • Remove the current fallback that displays arbitrary retrieved sources when the model selects none.
  • Claims about messages require citations; advice clearly labeled as advice does not.
  • If evidence is insufficient, answer that plainly and suggest a narrower person, chat, or date. Do not manufacture a citation.
  • All message content is untrusted data. It cannot alter system rules, scopes, tools, approvals, or destinations.

Cost controls#

Measure cost from provider-reported usage and the versioned price book in integer micro-USD. Do not hard-code a permanent dollar estimate in this document. For requestr:

text
estimated_max_cost(r) =
  query_embedding_tokens × embedding_rate
  + uncached_input_tokens × input_rate
  + cache_read_tokens × cache_read_rate
  + max_output_tokens × output_rate
  + permitted_additional_steps × step_reserve

settled_cost(r) = provider-reported usage × price_book(version)
LeverSpecificationWhy it matters
No-call pathsNo generation for no-evidence search, index-disabled status, thread/navigation commands, or structured data Claire can render directly.The cheapest request is the one not sent.
One-step defaultNormal Ask Claire performs one query embedding and one streamed generation. No planner model and no tool loop.Every agent step repeats context and can multiply input cost.
Compound-query escalationRun the low-cost structured planner only when the deterministic compiler cannot represent a multi-constraint or action-oriented request. Permit one planner call, record why it was needed, and cap its input/output independently.Identity lookup and date parsing stay cheap while complex birthday-style queries remain expressible.
Metadata-first analyticsCompute communication frequency, active days, reciprocity, and recency with SQL or background jobs. Use incremental, changed-window extraction for person facts and plan candidates.Avoids repeatedly sending a person's full message history to a model.
Context budgetDefault ≤4,000 input tokens; p95 ≤8,000; four evidence windows; six recent turns plus an optional bounded rolling summary.Input growth is the dominant controllable interactive cost.
Output budgetDefault target ≤300 tokens and hard maximum 700. UI quick actions request shorter caps.Output tokens are normally more expensive and cannot use prefix caching.
Model routingUse the least expensive profile that clears eval gates; reserve strong tool models for complex or action-oriented requests.Provider portability should reduce spend without silently reducing grounding quality.
Embedding batchesUse AI SDK embedMany for 32–100 messages per logical batch with bounded parallelism and bulk upsert.Reduces HTTP and database overhead; provider batch pricing may further reduce offline backfill cost.
Prompt cachingPlace the stable Claire system prefix first and volatile user evidence last. Enable provider-specific cache controls only after measuring a real cache hit.Do not pad prompts merely to cross a provider cache minimum.
Query embedding cacheOptional short-TTL Redis cache keyed by an HMAC of normalized query + embedding model/version; never store raw query text in the cache key.Helps repeated quick actions; embedding spend itself is small, so complexity must earn its latency benefit.
Answer cacheDeferred. A safe key requires user, scope, thread summary, instruction version, evidence content hashes, prompt version, and model profile.A broad semantic answer cache risks stale or cross-context answers.
Budget reservationReserve worst-case managed cost before calling a provider; settle actual usage; release on failure. Enforce per-request, daily warning, monthly allowance, and opt-in overage caps.A rate limit alone does not prevent an expensive valid request or shared-IP unfairness.

Message indexing specification#

  • Replace process-local fire-and-forget backfill with a durable per-user job and bounded provider concurrency. Message ingestion must never wait for embedding.
  • Use embedMany; split by provider maximum and token budget, then bulk upsert vectors in the same order.
  • Select rows whose embedding is missing or whose stored content_hash differs. Content edits must become searchable.
  • Delete or exclude vectors when messages are soft-deleted, and keep index counts scoped to eligible live text rows.
  • Maintain index progress incrementally. Do not run exact counts over messages and embeddings during every answer request.
  • Store embedding_model, dimensions, and embedding_version. A model/dimension change creates a parallel versioned index and controlled cutover, not mixed vectors.
  • Retry transient provider failures with jitter and a ceiling; expose permanent failures without blocking lexical search.

Data model and API changes#

ChangePurpose
conversation_assistant_turns.statusRepresent streaming, completed, failed, and cancelled turns.
conversation_assistant_turns.request_idUnique per user for retry/idempotency.
conversation_assistant_turns.prompt_versionMake evals, cache invalidation, and regressions traceable.
conversation_assistant_turns.provider/model/finish_reasonPrivacy-safe operational provenance. Token/cost authority remains the usage ledger.
assistant_thread_summaries or bounded summary columnsOptional older-thread continuity without replaying every turn. Generate asynchronously only after a threshold.
conversation_message_embeddings version fieldsPrevent incompatible models or dimensions from sharing one index.
assistant_people and assistant_person_linksResolve aliases and cross-platform identities while supporting people mentioned in a conversation who are not themselves Claire contacts.
assistant_person_factsProvenance-backed, time-aware person facts for identity, location, aliases, and other durable memory.
assistant_relationship_metricsDeterministic rolling communication aggregates used for explainable people ranking.
assistant_plan_candidatesTentative, confirmed, changed, and cancelled plans linked to their source messages.
conversation_assistant_turns.query_planThe resolved intent, filters, timezone, and absolute date range used for a reproducible answer; exclude raw message content.

Persist the user turn and assistant placeholder in one database function. Finalize the assistant turn, update the thread title/timestamp, and settle the usage record atomically where practical. Return the updated thread in the terminal stream data so the client does not reload the entire thread list after every answer. Thread history uses keyset pagination; generation reads only the bounded recent slice.

Path to plugins and real actions#

Keep open_conversation as trusted Claire navigation. Rename the current calendar affordance to “Open calendar” so it does not imply that anything was created. When plugins land, Ask Claire receives two classes of tools:

Tool classExamplesExecution rule
ReadFree/busy, installed capabilities, selected file searchExecute only within installation and context grants; bounded result returned to the model.
ProposeCreate event, invite guests, create task, update CRMThe tool may only call PluginPolicyEngine.authorize() and create a durable proposal. It cannot call the provider adapter.
External writes remain user-controlled — Ask Claire may prepare a proposal, but the plugin policy service and the user approval step stand between the model and external execution.

The streamed proposal part contains only a Claire proposalId and display metadata. Approval reloads the proposal from the server, verifies the payload hash and destination, then queues idempotent execution and writes a receipt. External writes and invitations always follow the plugin specification even if the AI SDK marks a tool call approved.

Default Ask Claire remains one-step RAG. Enable tool calling only when an installed, granted capability is relevant or the question requires a structured Claire read tool. Cap global Ask Claire at three model steps, 20 seconds total, five seconds per read tool, 4KB per tool result, and one proposal per user intent. A model denial response must not cause it to retry the same proposal.

Privacy and security requirements#

  • Every service-key database query includes user_id; every chat and thread scope is ownership-checked server-side.
  • Telemetry sets recordInputs: false and recordOutputs: false. Metadata may include request ID, route, provider, model, token counts, timings, source count, cache status, and error class—never message text, question text, embeddings, tool inputs, or provider credentials.
  • Only citations selected for the answer leave the server in source parts. Retrieval candidates that were not used are not exposed to plugins.
  • Remote provider and plugin data handling follows the user's configured consent and the plugin manifest. BYOK is not described as local.
  • Sensitive inferred person attributes are ephemeral by default. Persisting them requires the applicable memory policy, purpose limitation, retention, correction, deletion, and source provenance; they are never exposed to a plugin merely because they appeared in retrieval.
  • Prompt injection fixtures must prove that message text cannot expand scope, select a hidden destination, grant a capability, lower approval, or cause direct execution.
  • Error events expose stable safe codes; raw provider errors remain in redacted server logs.

Observability and cost accounting#

Record one trace and one usage event per answer with privacy-safe spans:

text
assistant.request
  ├─ thread.load
  ├─ context.route
  ├─ retrieval.lexical
  ├─ retrieval.embedding
  ├─ retrieval.vector
  ├─ evidence.assemble
  ├─ model.first_attempt | model.fallback
  │    ├─ time_to_first_token
  │    └─ stream_duration
  └─ turn.finalize

Dashboards break down p50/p95 time to first token, completion latency, input/output/cache tokens, estimated and settled micro-USD, no-call rate, fallback rate, cancellation rate, stream failure rate, retrieval route, source count, and model profile. Alerts use error and budget rates, not conversation content.

Evaluation and test plan#

SuiteCoverage
Retrieval corpusFactual lookup, paraphrase, temporal query, multiple people with the same name, cross-platform duplicates, strict scope, preferred scope, tone, open loop, no-answer, deleted/edited messages.
Golden personal-memory queriesAttribute-based person identification in a recent time window; Saturday plan recall with proposals, confirmations, reschedules, and cancellations; Mexico City and Brooklyn location filtering; explainable relationship ranking; multiple plausible candidates; stale and conflicting facts.
Grounded answer evalCitation precision/recall, unsupported-claim rate, no-answer correctness, instruction adherence, useful concision, and source-card navigation.
Provider contract testsStreaming, abort, structured output, tools, usage fields, context limit, embedding dimensions, and simulated streaming fallback for every enabled profile.
Transport testsChunk fragmentation, UTF-8 boundaries, custom data parts, sanitized errors after headers, cancellation, duplicate request ID, reconnect, proxy buffering, and terminal event.
Client testsOptimistic/persisted ID replacement, batched deltas, Stop, background/foreground, offline retry, citations during stream, screen-reader behavior, mobile and Electron parity.
Action safetyInjection, mutated approval payload, missing grant, revoked account, duplicate execution, timeout, retry, denial, and receipt provenance.
People and plan memoryUser override precedence, fact provenance, validity intervals, last-confirmed copy, metric reproducibility, group-noise exclusion, tentative-versus-confirmed plans, timezone boundaries, and superseded plan state.

Reuse the Lucas two-device context-token scenario for real bridge grounding. Add a deterministic AI SDK mock stream so UI and server tests require no provider key and do not depend on token timing.

Release acceptance gates#

Interactive latency budget

Warm-staging launch targets; lower is better

  • Visible retrieval state≤250ms

    The stream is open and the UI acknowledges work immediately.

  • p50 first answer text≤1.5s
  • p95 first answer text≤3s
  • p95 answer completion≤10s
DimensionGate
CorrectnessGlobal and strict-chat integration tests pass; strict scope has zero cross-chat sources; no source is shown unless selected by the answer.
Personal-memory qualityGolden queries resolve absolute date ranges correctly, retrieve the supporting conversation window, preserve plausible alternatives, and attach provenance to every identity, location, relationship, and plan claim.
Action honestyPerson shortlists and outreach drafts may be produced, but zero recipient is contacted until the user approves the exact durable plugin proposal.
LatencyOn warm staging: p50 time to first text ≤1.5s, p95 ≤3s; p95 answer-only completion ≤10s. Emit a visible retrieval state within 250ms of opening the response stream.
Request countOne client request per question; normal answer uses at most one query embedding and one generation step.
ContextMedian generation input ≤4,000 tokens and p95 ≤8,000; hard output maximum 700 tokens.
CostMeasured cost per successful answer and per monthly active user is recorded. A candidate model must clear quality gates and reduce or hold settled cost; a >15% cost regression blocks rollout unless explicitly approved.
IndexingBackfill averages at least 32 texts per embedding provider request, resumes after restart, indexes edits, and never blocks message ingestion.
Reliability≥99% of opened streams produce a completed, failed, or cancelled terminal turn; retries with the same request ID produce no duplicate model call or turn.
SafetyZero tested path performs an external write from Ask Claire without a durable authorized plugin proposal; telemetry contains no prompt, output, embedding, or tool payload.

Numeric latency and quality gates are launch targets, not claims about the current system. Capture the current baseline before PR 1 and revise targets only from measured staging and device evidence.

Delivery plan#

  1. PR 0 · Now

    Correctness and baseline

    Fix conversation-thread access, add global/chat integration tests, and record current latency, tokens, request counts, and cost.

  2. PR 1

    Shared AI runtime

    Move generation and embeddings behind services/ai with task profiles, provider contracts, timeouts, and usage normalization.

  3. PR 2

    Hot path and index

    Add recent-turn queries, incremental index state, indexed full-text search, embedMany, atomic finalization, and idempotency.

  4. PR 3

    Server streaming

    Ship AI SDK UI streams over Express with typed parts, cancellation, finalization, and a legacy JSON fallback.

  5. PR 4

    Client streaming

    Add the shared Expo and Electron stream adapter, incremental rendering, Stop, recovery, and accessibility.

  6. PR 5

    Personal-memory retrieval v2

    Add structured query compilation, relative dates, message windows, person facts, relationship metrics, plan candidates, and golden-query evals.

  7. PR 6 · After plugins

    Durable action proposals

    Expose installed read and propose tools through the plugin policy engine without granting the model direct write access.

Roll out by account: internal fixtures → Lucas two-device account → 5% → 25% → 100%. Keep the JSON path and previous provider profile available for one release. Automatically roll back on grounding, cost, stream-error, or latency gate regression.

Primary implementation map#

  • apps/server/src/services/conversation-assistant.ts — split orchestration, retrieval, persistence, and transport responsibilities.
  • apps/server/src/services/ai/ — shared model profiles, provider capabilities, streamed generation, embeddings, usage, and test mocks.
  • apps/server/src/routes/ai.ts — content negotiation, stream lifecycle, abort, and legacy JSON fallback.
  • supabase/migrations/ — turn lifecycle/idempotency, embedding versioning, full-text index, and atomic database functions.
  • apps/client/services/conversationAssistant.ts — typed stream transport and legacy fallback.
  • apps/client/hooks/useAssistantStream.ts — shared mobile, in-chat, and Electron state adapter.
  • apps/client/components/claire/ — shared answer, source, proposal, status, stop, and error parts.

External implementation references#