ClaireDocs

Ask Claire

Search the project, not the web.

Answers are grounded in Claire’s published documentation and include the source pages used.

⌘/Ctrl + J opens Ask Claire anywhere in docs.

Loops: relevance and evaluation

How relevance scoring decides whether a group message concerns the user, and how to evaluate it.

CurrentReviewed 2026-08-18View source ↗

The contributor view of loops. For what a loop is, start at How Loops work.

States and status#

A loop spans messages and carries state through proposed → negotiating → pending_confirmation → agreed → resolved. Only agreed and later are actionable — no plugin is offered until the conversation actually agrees.

Status on the row is a separate, smaller vocabulary: open, waiting, snoozed, done, dropped, superseded. Overdue is derived, never stored — status IN (open, waiting) AND COALESCE(snoozed_until, deadline) < now(). Snoozing writes snoozed_until and never touches deadline, so a loop snoozed twice does not lose the date the user committed to.

Relevance scoring#

apps/server/src/services/loops/relevance.ts scores every candidate. Three properties, each deliberate:

  • Deterministic. No model call. It decides whether someone else’s business becomes your loop, which is privacy-adjacent, so it must be auditable and must not change when the model provider changes.
  • Pure. No I/O, so every branch is table-testable.
  • Explainable. Every decision returns the signals that produced it, stored on the loop, so “why didn’t Claire catch this?” has an answer.

Two hard passes bypass the threshold: a one-to-one conversation, and the user having committed themselves. sensitivity: off outranks both — “never make loops here” has to mean it.

SignalWeightNote
mention_exact+0.45Structured m.mentions preferred; text matching is a fallback
reply_to_me+0.40Requires persisted reply metadata
llm_assigned_to_me+0.35From the extraction stage
named_other−0.55The most important suppressor: the work is explicitly someone else’s
broadcast_mention−0.35@channel addresses everyone, so it addresses no one in particular
no_self_signal−0.25Nothing ties the message to the user

Platform differences are data, not branches#

The pipeline contains no if (platform === 'slack'). Every difference lives in loopSemantics on PlatformDefinition in packages/platform-catalog. Platforms absent from the table get safe defaults, so adding a bridge never requires editing detection logic.

PlatformHow a mention rendersThreading
WhatsAppThe phone number — never the display nameReplies only
TelegramA stable @handleReplies only
SlackDisplay name in the body; the real id only in m.mentionsNative threads
InstagramA @usernameReplies only

This is why structured mentions are the real mechanism: text matching works on Telegram, fails outright on WhatsApp, and is fragile on Slack. Self-identity is also per-workspace on Slack, so it is keyed (userId, platform, accountRef).

Evaluating it#

Terminal
bun run eval:loops                                  # report
cd apps/server && bun run eval:loops --show-passing # every scenario
cd apps/server && bun test                          # the eval runs here too

No database, network, or API key. The relevance stage is deterministic, so a failure is always a real regression rather than model variance — which is why it belongs in CI rather than a manual step.

CorpusPurpose
HAND_AUTHOREDAcceptance criteria, each stating why it matters
ADVERSARIALConflicting signals and messy language
generateScenariosSeeded breadth across platform × group size × mention style

Release gates: group-suppression accuracy ≥ 0.90, precision ≥ 0.85, recall ≥ 0.70. False positives are weighted harder than misses throughout — a wrong loop erodes trust in every other loop.

Where the code lives#

PathWhat
apps/server/src/services/loops/relevance.tsSignal scoring and thresholds
apps/server/src/services/loops/loop-{gate,context,prompts,reconciler,detector,queue}.tsThe detection pipeline
apps/server/src/services/loops/loop-agent.tsThe loop-scoped agent (read and propose only)
apps/server/src/services/ai/Provider registry and schema-validated output
apps/server/src/plugins/blocks/schema.tsPlugin block validation
apps/client/features/loops/List, details page, timeline, agent panel, block renderer
apps/server/src/services/loops/eval/Scenario types, generator, runner
apps/server/scripts/eval-loops.tsEvaluation CLI
apps/server/src/routes/loops.tsREST API
packages/platform-catalog/src/loop-semantics.tsPer-platform behaviour
supabase/migrations/20260817020*Schema

The detection pipeline#

One pass over a chat is gate → extract+reconcile → apply. Scheduling is perchat and debounced, not per message: a burst of twenty messages costs one pass over the whole exchange rather than twenty passes over twenty fragments. That is what lets one plan stay one loop as it evolves.

StageCostWhat it does
loop-gate.tsfreeRegex plus an open-loop check. A chat with a live loop always re-runs — otherwise resolutions are never noticed and nothing ever closes.
loop-context.tsfreeWindowing with overlap so a loop spanning the cursor stays coherent. Evidence attachment is idempotent in the database, so re-reading cannot duplicate.
loop-prompts.tsone callThe model returns operations against already-open loops, not a list of loops.
loop-reconciler.tsfreePure guards. Decides what each operation is allowed to do.

Closes are guarded hardest: explicit cited evidence, confidence ≥ 0.75, and the chat’s auto_close. Failing any of those, the close is recorded as a suggestion rather than applied. A missed loop disappoints; a wrongly-closed loop is a broken promise.

The loop agent#

The details page can ask Claire to help close a loop. Its safety property is structural, not prompted: the four tools are two reads and two proposals, and there is no tool that sends a message, writes externally, or mutates a row. A prompt injection that fully succeeds can make Claire say something wrong; it cannot make Claire do anything. Tests assert this against the source, so adding a sending tool later fails the build.

The step cap, the wall clock, and output truncation are one mechanism doing two jobs — “call search fifty times” is both an injection attempt and a bill.

Plugin blocks#

A plugin can contribute a small, fixed vocabulary of typed blocks to a loop. This is deliberately the opposite of unrestricted UI: no styling, no markup, no nesting, no code. The plugin supplies data; Claire owns rendering.

Validation runs server-side before persistence, never at render time — a renderer that trusted the shape would be trusting the plugin, and every client would have to re-implement the checks. Three rules carry the weight: requiresApproval is computed from the installation’s manifest risk and overwrites whatever the plugin supplied; link.url must be https with a host in the egress allowlist, and the host is re-derived rather than accepted; and an unknown block kind is rejected rather than ignored, so a newer plugin cannot smuggle a payload past an older server.

Not built yet#

Detection ships behind LOOP_DETECTION_MODE=off until the mining harness measures it against a real corpus — the per-call token count rises sharply and only measurement proves the call-count reduction more than compensates.

Still specified in the Loops revamp plan but not implemented: the plugin runtime (registry, policy engine, gateway, approvals, receipts) behind the block schema, the corpus-mining harness, cross-platform loop merging, and the transactional outbox that trigger fan-out will need before any plugin performs a background external write.