Files
gbrain/CLAUDE.md
T
662a6e27d4 v0.42.6.0 feat(enrich): gbrain enrich --thin — brain-internal grounded synthesis for stub pages (#1700) (#1757)
* feat(engine): listEnrichCandidates source-aware candidate selection (#1700)

One SQL query per engine: thin-filter + per-page source-correct inbound-link
count (to_page_id = p.id, mentions excluded) + enriched_at recency guard +
whitelisted ORDER BY (ENRICH_ORDER_SQL) + LIMIT, returning a lightweight
projection (no page bodies). EnrichCandidate/EnrichCandidatesOpts types.
pg + pglite parity, pinned by engine-parity.test.ts.

* feat(enrich): gbrain enrich --thin brain-internal grounded synthesis (#1700)

Develops stub pages at scale by consolidating scattered brain knowledge (search
+ backlinks + facts + raw_data) into one grounded gateway.chat call per page.
Resumable (op-checkpoint), budget-capped (best-effort under --workers), per-page
advisory lock, put_page write-through. CLI + thin-client refuse + Minion handler.

Includes codex-review fixes: sanitizeContext neutralizes the <context> envelope
delimiters (P1 injection escape); background fan-out idempotency key carries the
run fingerprint (P1); post-hoc budget-overage flag via new BudgetTracker.cap
getter (P1); checkpoint flush on budget exhaustion (P2). Accepts the documented
best-effort in-flight-cancel limitation (D5) with an explicit code note.

* feat(cycle): enrich_thin opt-in autopilot phase (#1700)

Default-OFF trickle around runEnrichCore: develops a few thin pages per source
per tick so the brain compounds over time. Per-source cap enforced as
min(per-source, brain-wide remaining) with brain-wide total + walltime caps
(P2 fix: per-source max_cost_usd was parsed but never enforced). Wired into
CyclePhase / ALL_PHASES / PHASE_SCOPE / NEEDS_LOCK + dispatch.

* chore: bump version and changelog (v0.42.6.0)

gbrain enrich --thin batch enrichment (#1700) + codex-review fixes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 22:13:19 -07:00

559 KiB
Raw Blame History

CLAUDE.md

GBrain is a personal knowledge brain and GStack mod for agent platforms. Pluggable engines: PGLite (embedded Postgres via WASM, zero-config default) or Postgres + pgvector

  • hybrid search in a managed Supabase instance. gbrain init defaults to PGLite; suggests Supabase for 1000+ files. GStack teaches agents how to code. GBrain teaches agents everything else: brain ops, signal detection, content ingestion, enrichment, cron scheduling, reports, identity, and access control.

Two organizational axes (read this first)

GBrain knowledge is organized along two orthogonal axes. Users AND agents must understand both, or queries misroute silently.

  • Brain — WHICH DATABASE. Your personal brain is host. You can mount additional brains (team-published, each with their own DB and access policy) via gbrain mounts add (v0.19+). Routing: --brain, GBRAIN_BRAIN_ID, .gbrain-mount dotfile.
  • Source — WHICH REPO INSIDE THE DATABASE. A brain can hold many sources (wiki, gstack, openclaw, essays). Slugs scope per source. Routing: --source, GBRAIN_SOURCE, .gbrain-source dotfile.

Both axes follow the same 6-tier resolution pattern. Read docs/architecture/brains-and-sources.md for topology diagrams (personal, team mount, CEO-class with multiple team brains) and skills/conventions/brain-routing.md for the agent-facing decision table.

Architecture

Contract-first: src/core/operations.ts defines ~47 shared operations (v0.29 adds get_recent_salience, find_anomalies, get_recent_transcripts). CLI and MCP server are both generated from this single source. Engine factory (src/core/engine-factory.ts) dynamically imports the configured engine ('pglite' or 'postgres'). Skills are fat markdown files (tool-agnostic, work with both CLI and plugin contexts).

Trust boundary: OperationContext.remote distinguishes trusted local CLI callers (remote: false set by src/cli.ts) from untrusted agent-facing callers (remote: true set by src/mcp/server.ts). Security-sensitive operations like file_upload tighten filesystem confinement when remote=true and default to strict behavior when unset.

Key files

  • src/core/operations.ts — Contract-first operation definitions (the foundation). Also exports upload validators: validateUploadPath, validatePageSlug, validateFilename, plus matchesSlugAllowList(slug, prefixes) (v0.23 glob matcher: <prefix>/* matches recursive children; bare <prefix> matches exact only). OperationContext.remote flags untrusted callers; OperationContext.allowedSlugPrefixes (v0.23) is the trusted-workspace allow-list set by the dream cycle. put_page enforces: when viaSubagent and allowedSlugPrefixes is set, slug must match the allow-list; else the legacy wiki/agents/<id>/... namespace check applies. Auto-link enabled for trusted-workspace writes (skipped only when remote=true && !trustedWorkspace). As of v0.26.0, every Operation also carries scope?: 'read' | 'write' | 'admin' + localOnly?: boolean. All ops are annotated; sync_brain, file_upload, file_list, and file_url are admin + localOnly (rejected over HTTP). OperationContext.auth?: AuthInfo is threaded through HTTP dispatch for scope enforcement in serve-http.ts before the op runs. v0.26.9 (D12 + F7b): OperationContext.remote is now a REQUIRED field in the TypeScript type — the compiler is the first defense against transports that forget to set it. Four trust-boundary call sites (put_page allowlist, file_upload trust-narrowing, submit_job protected-name guard, auto-link skip) flipped from falsy-default (!ctx.remote) to fail-closed semantics (ctx.remote === false for "trusted-only" sites and ctx.remote !== false for "untrust unless explicit-false"). Anything that isn't strictly false is now treated as remote. Closed an HTTP MCP shell-job RCE: a read+write-scoped OAuth token could submit shell jobs because the HTTP request handler's literal context skipped remote: true and submit_job's protected-name guard saw a falsy undefined. Stdio MCP set the field correctly via dispatch.ts; HTTP inlined a parallel context-builder for several releases and lost it. v0.34.1.0 (#861 + #876): new helper sourceScopeOpts(ctx) encodes the precedence ladder for source-scoped reads — federated array (ctx.auth.allowedSources) wins over scalar (ctx.sourceId / ctx.auth.sourceId) over nothing. Every read-side op handler routes through it so future ops can't silently drift from the canonical v0.31.8 thread. Closes the source-isolation leak on the read path: a read+write-scoped OAuth client bound to --source dept-x no longer sees rows from neighboring sources via search / query / list_pages / get_page / find_experts / query's image path. v0.41.30.0: put_page's v0.38 inline disk write-through (the bare writeFileSync into sync.repo_path) extracted into the shared writePageThrough helper (src/core/write-through.ts) that put_page now calls — behavior preserved (same render-from-row + frontmatter-override stamping) and upgraded to ATOMIC (temp-sibling + rename), so a crash or concurrent gbrain sync can no longer read a half-written .md. The same helper now also backs gbrain brainstorm/lsd --save.
  • src/core/engine.ts — Pluggable engine interface (BrainEngine). clampSearchLimit(limit, default, cap) takes an explicit cap so per-operation caps can be tighter than MAX_SEARCH_LIMIT. Exports LinkBatchInput / TimelineBatchInput for the v0.12.1 bulk-insert API (addLinksBatch / addTimelineEntriesBatch). As of v0.13.1, BrainEngine has a readonly kind: 'postgres' | 'pglite' discriminator so migrations (src/core/migrate.ts) and other consumers can branch on engine without instanceof + dynamic imports. v0.29: four new methods — batchLoadEmotionalInputs(slugs?) (CTE-shaped read with per-table aggregates so a page × N tags × M takes never produces N×M rows), setEmotionalWeightBatch(rows) (UPDATE FROM unnest($1::text[], $2::text[], $3::real[]) composite-keyed on (slug, source_id) for multi-source safety), getRecentSalience(opts), findAnomalies(opts). PageFilters extended with sort?: 'updated_desc' | 'updated_asc' | 'created_desc' | 'slug' + PAGE_SORT_SQL whitelist consumed by both engines (was hardcoded ORDER BY updated_at DESC). v0.32.8 (PR #860): new listAllPageRefs(): Promise<Array<{slug, source_id}>> ordered by (source_id, slug). Cheap cross-source enumeration for hot loops on large brains — replaces the getAllSlugs()→getPage(slug) N+1 pattern in extract-takes, extract, integrity, which silently defaulted to source_id='default' for non-default-source pages. Implementation parity across postgres-engine.ts + pglite-engine.ts. Pinned by test/e2e/multi-source-bug-class.test.ts. v0.34.1.0 (#861): SearchOpts + PageFilters add sourceIds?: string[] for the federated read axis; both engines apply WHERE source_id = ANY($N::text[]) when the array is set and preserve the scalar sourceId fast path when unset. traverseGraph(slug, depth, opts?) and traversePaths(slug, opts?) accept opts.sourceId / opts.sourceIds so graph walks respect the caller's scope. T8 wave (pgGraph-inspired CI infra, v0.37.4.0): traverseGraph opts gains frontierCap?: number (per-iteration cap on the recursive CTE — approximately per-BFS-layer). Return type stays Promise<GraphNode[]> for MCP wire stability. New export TraverseGraphOpts. Postgres path uses parenthesized LIMIT N ORDER BY (slug, id) inside the recursive term; PGLite mirrors with positional params + the same shape SQL. Pinned by test/regressions/v0_36_frontier_cap.test.ts (4 contracts: cap-unset back-compat, cap-hit bounds result to <= cap+1, MCP wire-shape preservation, concurrency independence). onTruncation callback designed but stripped pre-merge in /review — adversarial pass caught false-positive (organic count == cap) + false-negative (LIMIT-before-DISTINCT in diamond graphs) cases in the v1 algorithm. Restoring the signal requires a dedupe-then-cap SQL rewrite + Postgres parity E2E — see TODOS.md → "T8 truncation signal". v0.35.6.0: two new methods supporting the phantom-redirect cycle pass — refreshPageBody(slug, sourceId, compiled_truth, timeline, content_hash) narrow-UPDATEs three columns + updated_at, skipping soft-deleted rows (codex #7: content_hash refresh is required so gbrain sync sees the canonical as unchanged after fence merge); migrateFactsToCanonical(phantomSlug, canonicalSlug, sourceId) UPDATEs entity_slug + source_markdown_slug on every active fact row keyed on the phantom, preserving embedding/validUntil/kind/status/source_session/confidence — codex #3 fix for the writeFactsToFence lossy-migration trap. Both methods have engine parity tests at test/phantom-redirect-engine-parity.test.ts. v0.40.4.0: new getAdjacencyBoosts(pageIds): Promise<Map<number, AdjacencyRow>> method powers the per-query graph-signals stage in hybrid search. Single SQL query returns inbound-link counts among the top-K set plus a cross-source count (links from pages whose source_id differs from the target's). COALESCE(p.source_id, 'default') for null safety; HAVING >= 1 matches JSDoc; cross-source CASE-WHEN on the JOINed target row excludes the target's own source. Parity SQL between postgres-engine.ts + pglite-engine.ts. SearchResult gains 12 new optional fields (base_score, backlink_boost, salience_boost, recency_boost, exact_match_boost, graph_adjacency_boost, graph_cross_source_boost, session_demote_factor, reranker_delta, plus internal staging fields). Pinned by test/e2e/graph-signals-engine.test.ts (7 cases) + cross-engine parity in the same suite. v0.41.25.0 (supersedes PR #1538): two new REQUIRED methods on BrainEngine — deletePages(slugs, {sourceId}): Promise<string[]> (single-batch primitive returning slugs actually deleted, D6) and resolveSlugsByPaths(paths, {sourceId}): Promise<Map<path,slug>> (batch path→slug lookup). sourceId is REQUIRED on both at the type level (D5; asymmetric with single-row deletePage which keeps optional/'default' fallback — v0.42+ TODO to tighten). Both methods short-circuit on empty input and throw when called with > DELETE_BATCH_SIZE. Closes the 73K-delete-jams-sync bottleneck for both engines. v0.41.31.0 (stale-embedding-semantics wave): four engine-interface changes for real model/dims-swap stale detection. (1) countStaleChunks(opts?) gains an optional signature?: string — when set, the stale predicate widens from embedding IS NULL to ALSO include chunks whose JOINed page embedding_signature IS NOT NULL AND <> $signature (a model/dims swap). NULL signature is GRANDFATHERED (never counted) so the post-migration-v108 corpus isn't flagged en masse; omit signature for the legacy NULL-only count. (2) New sumStaleChunkChars(opts?: {sourceId?, signature?}): Promise<number>SUM(LENGTH(chunk_text)) over stale chunks (same stale predicate + embed_skip filter + optional sourceId scope as countStaleChunks); used by the gbrain sync --all cost preview to price the embedding backlog via estimateCostFromChars. (3) New setPageEmbeddingSignature(slug, {sourceId?, signature}) — stamps pages.embedding_signature for one page after its chunks are (re)embedded; idempotent, no-op when the page doesn't exist. (4) New invalidateStaleSignatureEmbeddings({signature, sourceId?}): Promise<number> — NULLs embedding + embedded_at on every chunk whose page embedding_signature is set AND differs from signature, returns the count invalidated; the embed-stale loop calls it BEFORE listStaleChunks so signature-drift pages flow through the existing NULL-embedding keyset cursor unchanged (GRANDFATHER: NULL signature never invalidated). Implementation parity across postgres-engine.ts + pglite-engine.ts. Also widens findOrphanPages(opts?: {sourceId?, sourceIds?}) per the v0.41.29.0 orphan-source-scoping wave (candidate-side scoping only; inbound links counted from any source). Pinned by test/sum-stale-chunk-chars.test.ts, test/embedding-signature-stale.test.ts, test/e2e/engine-parity.test.ts. v0.41.34.0 (retrieval-maxpool incident, T3): two new methods for the free-text alias layer — resolveAliases(aliasNorms, opts?): Promise<Map<string, Array<{slug, source_id}>>> (READ side; maps each normalized alias to the (slug, source_id) pairs that declare it, source-scoped via sourceId/sourceIds) and setPageAliases(slug, sourceId, aliasNorms) (WRITE side; replaces the full alias set for one (slug, source) — delete then insert, empty clears, idempotent on the unique triple). Called by the ingest projection in importFromContent and the reindex --aliases backfill so a page's declared aliases stay in lockstep with its frontmatter. Implementation parity across both engines; pinned by test/search/page-aliases-engine.test.ts. searchVector in both engines now injects the shared buildBestPerPagePoolCte per-page max-pool (T1) so a page surfaces on its strongest chunk, not whichever single chunk won the inner LIMIT.
  • src/core/engine-constants.ts (v0.41.25.0, NEW) — single source of truth for engine batch-sizing constants. Exports DELETE_BATCH_SIZE = 500 consumed by both engines' deletePages + resolveSlugsByPaths and by the sync delete + rename loops. Lives outside engine.ts (the interface module) to avoid circular-import worry. Same order-of-magnitude as addLinksBatch's effective per-call budget — bounded per-statement work for predictable lock hold time + write amplification.
  • src/core/search/graph-signals.ts (v0.40.4.0) — per-query graph-signals helper. applyGraphSignals(results, engine, opts) runs as the 4th post-fusion stage (after backlink/salience/recency). Three boosts: ADJACENCY_BOOST=1.05 (page is linked from 2+ OTHER top-K results — local hub for THIS query), CROSS_SOURCE_BOOST=1.10 (page is linked from 2+ DIFFERENT sources — corroborated across team brains, dormant in single-source brains), SESSION_DEMOTE=0.95 (3+ results from same chat session — keep the highest-scoring one at full score, demote the rest). All three inherit the v0.35.6.0 floor-ratio gate that prevents weak pages from getting boosted past strong ones via popularity. computeScoreDistribution(results) emits min/p25/p50/p75/p95/max + reorder_band_width — instrumentation for the v0.41+ magnitude calibration wave (TODOS T-todo-2). sessionPrefix(slug) extracts the chat-session anchor (chat/2026-05-15-...). Pure pairedBootstrapPValue(deltas, resamples, rng) exported for eval gates. Test seam via adjacencyFn DI. Fail-open: any error logs via logGraphSignalsFailure (JSONL audit via audit-writer) and returns the input array unchanged. Pinned by test/search/graph-signals.test.ts (24 cases including the IRON-RULE floor-gate regression).
  • src/core/search/hybrid.ts extension (v0.40.4.0) — runPostFusionStages extended with a 4th stage (graphSignalsEnabled, onGraphMeta, onScoreDistribution). base_score stamped at function entry idempotently (captured ONCE before any boost stage mutates score). Every existing post-fusion stage now stamps its multiplier on the result: applyBacklinkBoostbacklink_boost, applySalienceBoostsalience_boost, applyRecencyBoostrecency_boost. applyReranker (called earlier in the pipeline) stamps reranker_delta as a rank delta (positive = improved). applyExactMatchBoost in src/core/search/intent-weights.ts stamps exact_match_boost when fired. Per-stage attribution is the cathedral that powers gbrain search --explain — every boost surface carries its own field, so formatResultsExplain reads them all without coupling to internal stage ordering.
  • src/core/search/explain-formatter.ts (v0.40.4.0) — renders SearchResult[] as a multi-line per-result breakdown for gbrain search --explain. Reads every boost-stamping field. Handles the "no boosts applied" empty path. 4-decimal precision with trailing-zero strip. Pinned by test/search/explain-formatter.test.ts (19 cases).
  • src/core/search/mode.ts extension (v0.40.4.0) — new graph_signals: boolean knob in ModeBundle. Defaults: conservative=false, balanced=true, tokenmax=true. KNOBS_HASH_VERSION bumps 3→4 with a gs= parts entry appended per the CDX2-F13 cache-key contamination convention so a graph-on cache write can't be served to a graph-off lookup. SearchKeyOverrides + SearchPerCallOpts + loadOverridesFromConfig + SEARCH_MODE_CONFIG_KEYS + resolveSearchMode + attributeKnob all carry the new field. Opt-out per-brain: gbrain config set search.graph_signals false. Mid-deploy: existing query_cache rows from before the upgrade hash differently — natural row segregation, clears within cache.ttl_seconds (3600s default).
  • src/core/audit/audit-writer.ts (v0.40.4.0, NEW) — shared JSONL audit primitive consolidating the 5 hand-rolled audit modules. Exports createAuditWriter({kind, recordSchema}) returning {log, readRecent} plus shared helpers computeIsoWeekFilename(kind, now?) and resolveAuditDir() (honors GBRAIN_AUDIT_DIR). ISO-week file rotation; best-effort writes (stderr warn on failure, never throws); read-path scans current-week + previous-week files for boundary spans. Five existing audits refactored onto it for parity: src/core/rerank-audit.ts, src/core/audit-slug-fallback.ts, src/core/minions/handlers/shell-audit.ts, src/core/minions/handlers/supervisor-audit.ts, src/core/facts/phantom-audit.ts. Public API of each module preserved bit-for-bit — every existing test passes unchanged. The new graph-signals-failures audit (logGraphSignalsFailure) uses the same primitive. Codex F16 (TODOS) calls out one remaining hand-rolled audit at src/core/skillpack/audit.ts for a v0.41+ refactor. Pinned by test/audit/audit-writer.test.ts (22 cases pinning the contract).
  • src/core/cli-options.ts extension (v0.40.4.0) — CliOptions gains explain: boolean. parseGlobalFlags recognizes --explain anywhere in argv (stripped before command dispatch). src/cli.ts formatResult for search + query cases routes to formatResultsExplain from src/core/search/explain-formatter.ts when CliOptions.explain is set; falls through to the existing JSON / human formatters otherwise.
  • src/commands/search.ts:gbrain search stats extension (v0.40.4.0) — new graph_signals section (enabled/source/failures_count/failures_by_reason). JSON envelope adds a graph_signals sibling property to existing stats; _meta.metric_glossary adds two new keys (graph_signals.enabled, graph_signals.failures_by_reason). Human output prints the section after the existing block. Reads search.graph_signals config first, falls back to the mode default. Pinned by test/search/search-stats-graph-signals.test.ts (6 cases).
  • src/commands/doctor.ts extension (v0.40.4.0) — new graph_signals_coverage check wired into both runDoctor (local) and doctorReportRemote (HTTP/JSON thin-client path). Reads search.graph_signals config first, falls back to mode default; silent ok when disabled. Computes inbound link coverage on the page set; warns at <10% with gbrain extract all fix hint; ok at ≥30% ("fire on most queries") and 10-29% ("fire occasionally"), each with the percentage embedded in the message. Pinned by 7 new cases in test/doctor.test.ts.
  • src/core/engine-factory.ts — Engine factory with dynamic imports ('pglite' | 'postgres')
  • src/core/pglite-engine.ts — PGLite (embedded Postgres 17.5 via WASM) implementation, all 40 BrainEngine methods. addLinksBatch / addTimelineEntriesBatch use multi-row unnest() with manual $N placeholders. As of v0.13.1, connect() wraps PGlite.create() in a try/catch that emits an actionable error naming the macOS 26.3 WASM bug (#223) and pointing at gbrain doctor; the lock is released on failure so the next process can retry cleanly. v0.22.0: searchKeyword and searchKeywordChunks multiply ts_rank by the source-factor CASE expression at the chunk-grain level; searchVector becomes a two-stage CTE — inner CTE keeps ORDER BY cc.embedding <=> vec so HNSW stays usable, outer SELECT re-ranks by raw_score * source_factor. Inner LIMIT scales with offset to preserve pagination contract. As of v0.22.6.1, initSchema() calls applyForwardReferenceBootstrap() BEFORE replaying SCHEMA_SQL — probes for the specific forward-referenced state the embedded schema blob needs (pages.source_id, links.link_source, links.origin_page_id, content_chunks.symbol_name, content_chunks.language, sources FK target table) and adds only what's missing. Closes the upgrade-wedge bug class that bit users 10+ times across 6 schema versions over 2 years (#239/#243/#266/#357/#366/#374/#375/#378/#395/#396). No-op on fresh installs and modern brains. v0.35.5.0: probe set extended in parity with postgres-engine.ts — files.source_id, files.page_id, oauth_clients.source_id, oauth_clients.federated_read, sources.archived, sources.archived_at, sources.archive_expires_at. Bootstrap also threads the DDL connection from initSchema so probes run inside the advisory-lock scope. Closes #1018, #974, #820. v0.37.10.0: getBrainScore now returns 100/100 with full breakdown components (35/25/15/15/10) when pageCount === 0. Vacuous truth — an empty brain has no coverage problem to penalize. Pre-fix, fresh gbrain init --pglite users saw "Brain score 0/100" on first gbrain doctor run, which was structurally surprising AND triggered the v0.37.8.0 doctor-report-remote.test.ts flake (test expected health_score >= 70 on a freshly-initialized test brain). Mirrored in postgres-engine.ts. Pinned by test/brain-score-breakdown.test.ts's updated empty-brain assertion and the hermetic test/doctor-report-remote.serial.test.ts. v0.41.8.0 (#1247/#1269/#1290/#1340): two changes. (1) disconnect() rewritten with the snapshot+early-null pattern — snapshot _db/_lock refs and null the instance fields BEFORE any await so a concurrent connect() cannot observe a partial mid-close state — wrapped in a try/finally that guarantees lock-release even if db.close() throws (Codex eng-review finding #7 lock-leak guard). KEEPS the original close-then-release order; PR #1337's swap to release-then-close was explicitly rejected because it would widen the window where a sibling process could connect to a still-closing brain. Pinned by test/pglite-engine-disconnect.serial.test.ts (5 cases: ordering, snapshot, lock-leak-on-throw, double-disconnect idempotency, reconnect-after-disconnect). (2) New exported classifyPgliteInitError(message): 'bunfs' | 'macos-26-3' | 'unknown' + buildPgliteInitErrorMessage(verdict, original) route the PGlite.create() catch-block hint by failure shape. The bunfs verdict matches the literal $$bunfs marker OR ENOENT[\s\S]*pglite\.data co-occurrence (tightened from a generic pglite.data substring per Codex eng-review finding #9); it surfaces a paste-ready bun upgrade hint + Node fallback. The macos-26-3 verdict keeps the existing #223 link. unknown falls through to the prior generic hint. Closes the user-facing half of #1340 — pre-fix every PGlite.create() failure pointed at #223 regardless of root cause. Pinned by test/pglite-init-classifier.test.ts (12 cases including the #1340 reporter's exact error string round-trip). v0.41.25.0 (supersedes PR #1538): implements deletePages(slugs, {sourceId}) + resolveSlugsByPaths(paths, {sourceId}) via slug = ANY($1::text[]) array-param binding (the same idiom addLinksBatch already uses). Caller-chunking primitive — throws when input exceeds DELETE_BATCH_SIZE. Returns RETURNING slug rows for deletePages so callers can filter pagesAffected to confirmed deletes only. v0.41.31.0: implements the embedding-signature stale-detection quartet — sumStaleChunkChars({sourceId?, signature?}), setPageEmbeddingSignature(slug, {sourceId?, signature}), invalidateStaleSignatureEmbeddings({signature, sourceId?}), and the widened countStaleChunks({sourceId?, signature?}). The signature opt widens the stale predicate via JOIN pages p ... WHERE cc.embedding IS NULL OR (p.embedding_signature IS NOT NULL AND p.embedding_signature <> $signature) (NULL grandfathered). Parity SQL with postgres-engine.ts. — PGLite-specific DDL (pgvector, pg_trgm, triggers)
  • src/core/postgres-engine.ts — Postgres + pgvector implementation (Supabase / self-hosted). addLinksBatch / addTimelineEntriesBatch use INSERT ... SELECT FROM unnest($1::text[], ...) JOIN pages ON CONFLICT DO NOTHING RETURNING 1 — 4-5 array params regardless of batch size, sidesteps the 65535-parameter cap. As of v0.12.3, searchKeyword / searchVector scope statement_timeout via sql.begin + SET LOCAL so the GUC dies with the transaction instead of leaking across the pooled postgres.js connection (contributed by @garagon). getEmbeddingsByChunkIds uses tryParseEmbedding so one corrupt row skips+warns instead of killing the query. v0.22.0: searchKeyword, searchKeywordChunks, and searchVector apply source-aware ranking by inlining the source-factor CASE and NOT (col LIKE …) hard-exclude clause from src/core/search/sql-ranking.ts. searchVector switches to a two-stage CTE (HNSW-safe inner ORDER BY, source-boost re-rank in the outer SELECT) and carries p.source_id through inner→outer for v0.18 multi-source callers. v0.22.1 (#406): _savedConfig retains the connect config; reconnect() tears down + recreates the pool from saved config (called by supervisor watchdog after 3 consecutive health-check failures). executeRaw is a single-statement passthrough — no per-call retry (D3 dropped that as unsound for non-idempotent statements; recovery is supervisor-driven). v0.22.1 (#363, contributed by @orendi84): connect() applies resolveSessionTimeouts() from db.ts as connection-time startup parameters (statement_timeout, idle_in_transaction_session_timeout) so orphan pgbouncer backends can't hold locks for hours. v0.22.1 (#409, contributed by @atrevino47): countStaleChunks() + listStaleChunks() server-side-filter on embedding IS NULL for embed --stale, eliminating ~76 MB/call client-side pull on a fully-embedded brain; upsertChunks() resets both embedding AND embedded_at to NULL when chunk_text changes without a new embedding (consistency). As of v0.22.6.1, initSchema() calls applyForwardReferenceBootstrap() BEFORE replaying SCHEMA_SQL on the same forward-reference probe set as the PGLite engine, so old Postgres brains pinned at v0.13/v0.18/v0.19 walk forward cleanly instead of wedging on column "..." does not exist. v0.35.5.0: probe set extended for the column-only forward-reference cases the original v0.22.6.1 sweep missed — files.source_id, files.page_id (pre-v0.18 brains where idx_files_source_id was the choke point), oauth_clients.source_id, oauth_clients.federated_read (pre-v0.34 brains where v60+v61+v65 chain failed), and sources.archived + sources.archived_at + sources.archive_expires_at (pre-v0.26.5 brains where CREATE TABLE IF NOT EXISTS sources was a no-op on existing tables so the archive lifecycle columns never landed). Also (Codex P1 from pre-landing review): the entire probe path now runs on the DDL connection threaded down from initSchema — previously probes ran through the instance pool while the advisory lock sat on a different connection, opening a concurrent-bootstrap race for Supabase pooler users. Closes #1018, #974, #820. v0.28.1: disconnect() is now idempotent. New _connectionStyle instance field tracks whether the engine owns its pool (worker engines) or shares the module-level singleton; second call on an instance-pool engine is a no-op rather than falling through to db.disconnect() and clobbering the singleton. Pinned by test/e2e/postgres-engine-disconnect-idempotency.test.ts (2 cases). Closes the bug class where any test sharing an engine across multiple worker.start() / worker.stop() cycles silently broke its own DB connectivity. v0.37.10.0: getBrainScore empty-brain parity with PGLite — returns 100/100 with breakdown components 35/25/15/15/10 when pageCount === 0 instead of the pre-fix 0/100. Same vacuous-truth rationale as the PGLite side; both engines must agree to keep doctor-report-remote.serial.test.ts deterministic. v0.41.25.0 (supersedes PR #1538): implements deletePages(slugs, {sourceId}): Promise<string[]> via DELETE FROM pages WHERE slug = ANY($1::text[]) AND source_id = $2 RETURNING slug (single SQL round-trip per call; caller chunks). Pair resolveSlugsByPaths does the SELECT slug, source_path FROM pages WHERE source_path = ANY($1::text[]) AND source_id = $2 batch lookup. FK cascades through content_chunks/links/tags/raw_data/timeline_entries/page_versions; files.page_id + links.origin_page_id go SET NULL. Pair primitive throws when input exceeds DELETE_BATCH_SIZE (from src/core/engine-constants.ts); both methods short-circuit on empty input. v0.41.31.0: implements the embedding-signature stale-detection quartet in parity with pglite-engine.ts — sumStaleChunkChars, setPageEmbeddingSignature, invalidateStaleSignatureEmbeddings, and the widened countStaleChunks (all accept the optional signature to extend "stale" to model/dims-swap drift via the pages.embedding_signature JOIN; NULL signature grandfathered). The v0.22.1 countStaleChunks / listStaleChunks server-side embedding IS NULL filter is preserved as the no-signature fast path. Pinned by test/e2e/engine-parity.test.ts.
  • src/core/cjk.ts (v0.32.7 CJK wave) — Single source of truth for CJK detection across the codebase. Exports CJK_RANGES_REGEX, CJK_SLUG_CHARS (character-class fragment for embedding inside other regexes), CJK_SENTENCE_DELIMITERS (。!?), CJK_CLAUSE_DELIMITERS (;:,、), CJK_DENSITY_THRESHOLD = 0.30, hasCJK(s), countCJKAwareWords(s) (30% density threshold — English docs with one Japanese term stay whitespace-tokenized; Chinese-dominant docs get char-counted), and escapeLikePattern(s) (escapes %, _, \\ for ILIKE ... ESCAPE '\\'). Replaces the inline hasCJK regex previously duplicated at expansion.ts:58. BMP-only ranges (Han / Hiragana / Katakana / Hangul Syllables); widening to Unicode property escapes is a v0.33+ TODO. Consumers: expansion.ts, sync.ts:slugifySegment, operations.ts:validatePageSlug + validateFilename, chunkers/recursive.ts:countWords + DELIMITERS, pglite-engine.ts:searchKeyword + searchKeywordChunks.
  • src/core/audit-slug-fallback.ts (v0.32.7 CJK wave) — Weekly ISO-week-rotated audit JSONL at ~/.gbrain/audit/slug-fallback-YYYY-Www.jsonl. logSlugFallback(slug, sourcePath) fires when importFromFile falls back to a frontmatter slug because slugifyPath returned empty (emoji / Thai / Arabic / non-CJK exotic-script filenames). readRecentSlugFallbacks(days) reads the last N days for gbrain doctor's slug_fallback_audit check. Honors GBRAIN_AUDIT_DIR via the shared resolveAuditDir() from shell-audit.ts. Separate surface from sync-failures.jsonl per codex outside-voice review — that file carries bookmark-gating semantics that info events shouldn't trigger.
  • src/core/embedding-pricing.ts (v0.32.7 CJK wave) — EMBEDDING_PRICING map keyed provider:model for the post-upgrade reindex cost estimate. Sibling to anthropic-pricing.ts. Entries: OpenAI text-embedding-3-large ($0.13/1M), 3-small ($0.02/1M), ada-002 ($0.10/1M), Voyage 3-large ($0.18/1M), 3 ($0.06/1M). lookupEmbeddingPrice(modelString) returns a tagged union (known with price + unknown with provider name); estimateCostFromChars(charCount, pricePerMTok) uses 3.5 chars/token approximation. Unknown providers degrade gracefully to "estimate unavailable" instead of fabricating numbers.
  • src/core/post-upgrade-reembed.ts (v0.32.7 CJK wave) — Pure functions backing the gbrain upgrade chunker-bump cost prompt. computeReembedEstimate(engine, model) queries real SQL (COUNT(*) + COALESCE(SUM(LENGTH(compiled_truth)) + SUM(LENGTH(timeline)), 0)) on pages WHERE chunker_version < MARKDOWN_CHUNKER_VERSION. formatReembedPrompt(est, graceSeconds) is the stderr-line formatter. runPostUpgradeReembedPrompt(engine, model, opts) orchestrates the 10-second Ctrl-C window; TTY-only wait (non-TTY auto-proceeds for CI / cron); GBRAIN_NO_REEMBED=1 bails out with a doctor-warning marker; GBRAIN_REEMBED_GRACE_SECONDS=0 skips the wait.
  • src/commands/reindex.ts (v0.32.7 CJK wave) — gbrain reindex --markdown [--limit N] [--dry-run] [--json] [--no-embed] [--repo PATH]. Walks pages WHERE page_kind = 'markdown' AND chunker_version < MARKDOWN_CHUNKER_VERSION in 100-row batches, ordered by id. Rows with non-null source_path re-import via importFromFile; rows without fall back to importFromContent against the stored compiled_truth. Both paths pass forceRechunk: true to bypass importFromContent's content_hash short-circuit — without that flag (codex post-merge F1), the chunker version bump never reaches pages whose source content hasn't changed since last sync, AND master's v0.32.2 stripFactsFence privacy strip never applies to pre-strip chunks. Idempotent — partial-completion re-runs pick up where they left off via id-ordered batches. Wired into src/commands/upgrade.ts:runPostUpgrade after apply-migrations. v0.41.37.0 (#1621): the DB-only fallback (no source file on disk) no longer passes body-only compiled_truth to importFromContent — that path re-parsed with EMPTY frontmatter and OVERWROTE the page's real frontmatter / title / timeline (codex catch). It now getPage + getTags, reconstructs the FULL markdown via serializeMarkdown(frontmatter, compiled_truth, timeline, {type, title, tags}), and re-imports THAT, so re-chunking a DB-only page preserves everything while still bumping chunker_version. Pinned by test/reindex-preserve-tags.test.ts.
  • src/commands/reindex-code.ts (v0.21.0 Cathedral II E2, extended v0.37.3.0) — gbrain reindex --code [--source ID] [--dry-run] [--yes] [--json] [--force] [--no-embed]. Walks pages WHERE type = 'code' in 100-row batches, replays through importCodeFile for chunk + embed + content_hash folding. Idempotent unless --force bypasses the content_hash early-return. v0.37.3.0: cost-preview model field now reads getEmbeddingModelName() from the gateway instead of the back-compat EMBEDDING_MODEL constant — preview reflects what the gateway will actually embed with. Same wave adds an informational stderr nudge inside runReindexCode (not the CLI wrapper, so dry-run + execute both surface it): when the configured embedding model isn't code-tuned (allowlist currently {'voyage-code-3'}, case-insensitive bare match against gateway-returned name), prints a 4-line recommendation to switch to voyage:voyage-code-3. Suppress with GBRAIN_NO_CODE_MODEL_NUDGE=1, --no-embed, or --json. Pure shouldNudgeCodeModel(bareName) helper returns a tagged NudgeDecision union; the helper takes the bare model name (matches gateway return shape) and emits qualified voyage:voyage-code-3 for the paste-ready gbrain config set line. Pinned by test/ai/voyage-code-3-recipe.test.ts, test/reindex-code-nudge.serial.test.ts, and test/reindex-code-model-source.serial.test.ts (the latter is the IRON-RULE regression for the cost-preview fix).
  • src/commands/sync.ts:resolveSlugByPathOrSourcePath (v0.32.7 CJK wave, codex post-merge F4) — Resolves a slug by pages.source_path first (returns the stored slug for frontmatter-fallback pages whose path doesn't derive a slug), then falls back to resolveSlugForPath(path). Threaded into all 4 delete/rename call sites (performSync's un-syncable cleanup at ~:531, deletes at ~:603, rename oldSlug at ~:622). Without this, emoji-only / Thai / Arabic filenames whose slug came from frontmatter would orphan on delete/rename (the delete path would compute the wrong path-derived slug). Best-effort query — pre-migration brains fall through to the legacy path.
  • src/core/utils.ts — Shared SQL utilities extracted from postgres-engine.ts. Exports parseEmbedding(value) (throws on unknown input, used by migration + ingest paths where data integrity matters) and as of v0.12.3 tryParseEmbedding(value) (returns null + warns once per process, used by search/rescore paths where availability matters more than strictness). v0.26.9 (D14): adds isUndefinedColumnError(err) predicate — pattern-matches Postgres SQLSTATE 42703 / "column ... does not exist" with engine-driver shape variation tolerated. Replaces bare catch {} blocks in oauth-provider.ts so genuine errors (lock timeout, network blip, permission denied) propagate while column-missing falls through to the legacy fallback path. Reusable from any future code that needs the same column-existence probe semantics. v0.32.8 (PR #860): adds validateSourceId(id) that throws on anything outside ^[a-z0-9_-]+$. Used by the per-source disk-layout fix in patterns.ts/synthesize.ts before any join(brainDir, '.sources', source_id, slug+'.md') call so source_id can't traverse out of brainDir. rowToPage updated to populate the now-required Page.source_id field from the SELECT projection (scripts/check-source-id-projection.sh enforces that every projection feeding rowToPage includes the column).
  • src/core/db.ts — Connection management, schema initialization. v0.22.1 (#363, contributed by @orendi84): resolveSessionTimeouts() returns statement_timeout + idle_in_transaction_session_timeout (defaults: 5min each, env-overridable via GBRAIN_STATEMENT_TIMEOUT / GBRAIN_IDLE_TX_TIMEOUT / GBRAIN_CLIENT_CHECK_INTERVAL). Both connect() (module singleton) and PostgresEngine.connect() (worker pool) consume the result via postgres.js's connection option, sending GUCs as startup parameters that survive PgBouncer transaction mode (unlike the prior setSessionDefaults post-pool SET, kept as a back-compat no-op shim).
  • src/commands/migrate-engine.ts — Bidirectional engine migration (gbrain migrate --to supabase/pglite)
  • src/core/import-file.ts — importFromFile + importFromContent (chunk + embed + tags). v0.41.31.0: importFromContent and importCodeFile stamp pages.embedding_signature via setPageEmbeddingSignature(slug, {sourceId, signature: currentEmbeddingSignature()}) when the import actually embedded (not --no-embed) — covers the inline import/sync path so a model/dims swap is detectable as stale. importCodeFile only stamps when every chunk was freshly embedded this call (needsEmbedIndexes.length === chunks.length); mixed reuse-by-hash pages stay unstamped (reindex --code --force / embed --stale handle the swap for those). v0.41.37.0 (#1621): importFromContent's tag reconciliation is now ADD-ONLY. The pre-fix "remove every existing tag not in current frontmatter, then add current" logic wiped ALL DB-side enrichment tags (auto-tag / dream synthesize / signal-detector) on every re-import — most visibly under gbrain reindex --markdown, which re-imports every page with forceRechunk. The tags table has no provenance column and frontmatter tags are stripped from stored pages.frontmatter (markdown.ts:118), so at re-import time a frontmatter-origin tag can't be distinguished from a DB-enrichment tag — deletion is unsafe. Now it only addTag (idempotent, ON CONFLICT DO NOTHING). Accepted trade-off: removing a tag from frontmatter no longer removes it from the DB on next sync (needs a tag_source provenance column, deferred — TODOS.md #1621-followup). Pinned by test/reindex-preserve-tags.test.ts + test/import-file.test.ts.
  • src/core/sync.ts — Pure sync functions (manifest parsing, filtering, slug conversion). v0.35.5.0: new exported pruneDir(name: string): boolean helper is the single source of truth for descent-time directory exclusion across walkers. Blocks node_modules (no leading dot, so pre-v0.35.5 walkers slipped through and inflated MISSING_OPEN counts via vendor packages), dot-prefix dirs, ops/, and *.raw sidecars. isSyncable now applies it per path segment; walkMarkdownFiles in src/commands/extract.ts and listTextFiles in src/core/cycle/transcript-discovery.ts consult it BEFORE recursing so the IO cost of walking thousands of vendor files is saved. Closes #923 + #202. manageGitignore worktree fix in same wave: discriminator now matches the gitdir path segment (/modules/<name> = submodule, /worktrees/<name> = worktree, per Git's documented layout) instead of the legacy absolute-vs-relative check that misclassified absorbed submodules and worktrees both. Conductor worktrees are first-class repos and now get .gitignore management for storage-tiering. Closes #889. v0.22.12 (#500, foundation by @wintermute via #501): classifyErrorCode(errorMsg) regex-based classifier with 12 codes (SLUG_MISMATCH, YAML_PARSE, YAML_DUPLICATE_KEY, MISSING_OPEN, MISSING_CLOSE, NESTED_QUOTES, EMPTY_FRONTMATTER, NULL_BYTES, INVALID_UTF8, STATEMENT_TIMEOUT, FILE_TOO_LARGE, SYMLINK_NOT_ALLOWED) plus UNKNOWN fallback. summarizeFailuresByCode(failures) returns sorted [{code, count}]. code? optional field on SyncFailure; backfilled at ack time on pre-v0.22.12 entries. acknowledgeSyncFailures() returns AcknowledgeResult { count, summary }. Three regexes (MISSING_OPEN, MISSING_CLOSE, EMPTY_FRONTMATTER) broadened to match actual markdown.ts:159-244 validator message strings, not just the literal code-name prefix. FILE_TOO_LARGE covers all three production size sites in import-file.ts:199, 352, 401; SYMLINK_NOT_ALLOWED covers the rejection at :347. Closes the silent-skip pattern that motivated #500.
  • src/core/storage.ts — Pluggable storage interface (S3, Supabase Storage, local)
  • src/core/storage-config.ts (v0.22.11) — Storage tiering: loadStorageConfig reads gbrain.yml, normalizes deprecated keys (git_tracked / supabase_only) to canonical (db_tracked / db_only) with once-per-process deprecation warning, and runs normalizeAndValidateStorageConfig (auto-fixes missing trailing /, throws StorageConfigError on tier overlap). Path-segment matcher: media/x/ does NOT match media/xerox/foo. Replaces gray-matter (broken on delimiter-less YAML) with a dedicated parser for the gbrain.yml shape.
  • src/core/disk-walk.ts (v0.22.11) — walkBrainRepo(repoPath) returns Map<slug, {size, mtimeMs}> from one recursive readdirSync. Skips dot-dirs, node_modules, non-.md files. Used by gbrain storage status to replace per-page existsSync + statSync (~400K syscalls on 200K-page brains → tens).
  • src/core/git-head.ts (v0.41.27.0) — local git HEAD freshness probe for gbrain doctor. isSourceUnchangedSinceSync(localPath, lastCommit, opts?) returns true iff localPath is a git repo whose current HEAD matches lastCommit; when opts.requireCleanWorkingTree is true, also requires the working tree to have no uncommitted changes (mirrors gbrain sync's force-walk gate at sync.ts:1075 so doctor and sync agree on "is there work to do?"). Two probe seams (_setGitHeadProbeForTests, _setGitCleanProbeForTests) match the last-retrieved.ts precedent so unit tests stay R2-compliant (no mock.module). Uses execFileSync with array args — shell metachars in local_path cannot escape to a shell (the v0.41.27.0 superseded PR #1564 used execSync through /bin/sh -c with JSON.stringify for shell-escape, which is unsafe; the rebuild's regression test runs real execFileSync against '/nonexistent/$(touch <sentinel>)/repo' and asserts the sentinel file is never created). Fail-open on every error: missing path, not a git repo, git not installed, timeout, NULL inputs, dirty-probe errored → returns false, preserving the caller's prior time-based behavior. Chunker-version-match check lives in the caller (doctor.ts) because it depends on engine state (sources.chunker_version vs CHUNKER_VERSION from src/core/chunkers/code.ts). Designed for reuse: autopilot's per-source dispatch will want the same gate (filed as v0.41.27.1+ TODO in the plan). Pinned by test/core/git-head.test.ts (14 cases incl. the shell-injection regression guard). v0.41.32.0 (supersedes #1623): GitFreshnessOpts.requireCleanWorkingTree widened boolean → boolean | 'ignore-untracked'. In 'ignore-untracked' mode the clean probe runs git status --porcelain --untracked-files=no, so a quiet repo with stray untracked dirs (?? companies/, ?? media/) is still "unchanged" — sync's incremental path keys off the commit diff and never imports untracked files, so doctor agrees with sync. GitCleanProbe gains an ignoreUntracked? second arg. One-line headline fix for the false-SEVERE staleness bug.
  • src/core/source-health.ts (v0.40, extended v0.41.32.0) — per-source health metrics for gbrain sources status + doctor's federation_health. v0.41.32.0 (supersedes #1623): commit-relative staleness. New newestCommitMs(localPath) = HEAD committer time via git log -1 --format=%ct (fail-open null; NO working-tree mtime parsing — committed content only, robust against the porcelain-mtime bug farm). New pure lagFromContentMs(contentMs|null, lastSyncMs|null, nowMs) = remote/column comparator (null lastSync → null; negative wall-clock → skew passthrough; contentMs <= lastSync → 0; else/null-content → wall-clock). computeAllSourceMetrics(engine, sources, {probeContent?}): LOCAL (probeContent:true, gbrain sources status) → isSourceUnchangedSinceSync(..., {requireCleanWorkingTree:'ignore-untracked'}) ? 0 : wall-clock (live commit-hash, catches HEAD moving to an old-dated commit which a timestamp compare would miss); REMOTE (default, federation_health on the HTTP MCP path) → lagFromContentMs(row.newest_content_at, ...), NO git subprocess (v0.41.27.0 trust boundary). Dead isSourceStale(src, intervalMs) removed (only autopilot-fanout.ts's own variant is live). Pinned by test/source-health.test.ts.
  • src/core/git-remote.ts (v0.35.3.0) — SSRF-hardened git invocations for remote-source cloneRepo and pullRepo. Exports two distinct flag constants because git's argv grammar treats them differently: GIT_SSRF_FLAGS (3 -c config flags — protocol.allow=user, protocol.file.allow=never, http.allowRedirects=false) is global config, spread BEFORE the subcommand verb. New GIT_SSRF_SUBCOMMAND_FLAGS = ['--no-recurse-submodules'] is subcommand-scoped, spread AFTER the verb. Pre-v0.35.3 a single combined GIT_SSRF_FLAGS array spread --no-recurse-submodules before the verb where real git rejects it with exit 129 ("unknown option"); the fake-git test harness exited 0 regardless of argv shape, so CI missed it for ~7 months and every remote-source clone/pull was silently broken. cloneRepo argv: git <GIT_SSRF_FLAGS> clone <GIT_SSRF_SUBCOMMAND_FLAGS> --depth=1 [--branch X] -- <url> <dir>. pullRepo argv: git <GIT_SSRF_FLAGS> -C <dir> pull <GIT_SSRF_SUBCOMMAND_FLAGS> --ff-only. Pinned by test/git-remote.test.ts position-anchored regression guard (argv.indexOf('--no-recurse-submodules') > argv.indexOf(verb)).
  • src/commands/storage.ts (v0.22.11) — gbrain storage status [--repo P] [--json]. Split into pure data (getStorageStatus) + JSON formatter + human formatter (ASCII-only per D10) matching the orphans.ts pattern. PageCountsByTier and DiskUsageByTier are distinct nominal types so swaps fail at compile time.
  • gbrain.yml (brain repo root, v0.22.11) — Optional storage tiering config. Top-level storage: section with db_tracked: and db_only: array-valued keys. gbrain sync auto-manages .gitignore for db_only paths on successful sync (skips on dry-run, blocked-by-failures, submodule context, or GBRAIN_NO_GITIGNORE=1). gbrain export --restore-only [--repo P] [--type T] [--slug-prefix S] repopulates missing db_only files from the database.
  • src/core/supabase-admin.ts — Supabase admin API (project discovery, pgvector check)
  • src/core/file-resolver.ts — File resolution with fallback chain (local -> .redirect.yaml -> .redirect -> .supabase)
  • src/core/chunkers/ — 3-tier chunking (recursive, semantic, LLM-guided). v0.19.0 adds code.ts — tree-sitter-based semantic chunker for 30 languages (v0.41 added SQL via DerekStride/tree-sitter-sql) with embedded-asset WASMs (src/assets/wasm/), @dqbd/tiktoken cl100k_base tokenizer, small-sibling merging. CHUNKER_VERSION constant folded into importCodeFile's content_hash so chunker shape changes force clean re-chunks across releases. v0.41 D2 SQL wave (#1173): extractSymbolName gains an inline SQL branch (extractSqlSymbolName) that dives through DerekStride's statement wrapper into the inner DDL child (create_table/create_function/create_view/create_index/create_procedure/create_type/create_schema/create_database/create_trigger/alter_table/alter_view) and extracts the target identifier via the name field with identifier-shaped fallback. DML kinds (select/insert/update/delete/merge/with) deliberately return null so chunks emit unnamed — code-def is a DDL signal. normalizeSymbolType gains parallel SQL branches mapping create_table → 'table', create_view → 'view', etc. src/commands/code-def.ts:DEF_TYPES was extended in the same wave with 'table' | 'view' | 'index' | 'procedure' | 'schema' | 'database' | 'trigger' so the new chunks surface in gbrain code-def <name> queries.
  • src/core/errors.ts (v0.19.0) — StructuredAgentError + buildError + serializeError. Every new v0.19.0 agent-facing surface (code-def, code-refs, usage errors) uses this envelope; matches v0.17.0 CycleReport.PhaseResult.error shape.
  • src/assets/wasm/ (v0.19.0, extended v0.41) — 37 tree-sitter grammar WASMs + tree-sitter runtime. Committed to the repo so bun --compile embeds them deterministically via import path from ... with { type: 'file' }. The CI guard scripts/check-wasm-embedded.sh fails the build if the compiled binary ever silently falls through to recursive chunks. v0.41 D2 wave (#1173): tree-sitter-sql.wasm (DerekStride/tree-sitter-sql @ c2e1e08db1ea20dc23bdb8d228a81a8756e9c450, built with tree-sitter-cli@v0.26.3 --abi 14) adds SQL coverage at 11 MB — substantially larger than peers because the grammar covers PostgreSQL + MySQL + SQLite + T-SQL basics (40 MB of generated parser.c). The compiled gbrain binary grows ~6% as a result; tracked as a follow-up TODO whether to switch to a smaller-coverage grammar.
  • src/commands/code-def.ts + src/commands/code-refs.ts (v0.19.0) — symbol definition + references lookup. Query content_chunks.symbol_name or chunk_text ILIKE with page_kind='code' filter. Auto-JSON when stdout is not a TTY (gh-CLI convention). Bypass the standard searchKeyword DISTINCT ON (slug) collapse so multiple call-sites from the same file surface.
  • src/core/search/ — Hybrid search: vector + keyword + RRF + multi-query expansion + dedup. As of v0.22.0, searchKeyword / searchKeywordChunks / searchVector apply source-aware ranking at the SQL layer (curated content like originals/, concepts/, writing/ outranks bulk content like wintermute/chat/, daily/, media/x/). searchVector uses a two-stage CTE so source-boost re-ranking doesn't kill the HNSW index. Hard-exclude prefixes (test/, archive/, attachments/, .raw/ by default) filter at retrieval, not post-rank. Both gates honor detail !== 'high' so temporal queries surface chat pages normally.
  • src/core/search/intent.ts — Query intent classifier (entity/temporal/event/general → auto-selects detail level)
  • src/core/search/eval.ts — Retrieval eval harness: P@k, R@k, MRR, nDCG@k metrics + runEval() orchestrator
  • src/core/search/source-boost.ts (v0.22.0) — Source-type boost map keyed by slug prefix. DEFAULT_SOURCE_BOOSTS (originals/ 1.5, concepts/ 1.3, writing/ 1.4, people/companies/deals/ 1.2, daily/ 0.8, media/x/ 0.7, wintermute/chat/ 0.5) and DEFAULT_HARD_EXCLUDES (test/, archive/, attachments/, .raw/). parseSourceBoostEnv / parseHardExcludesEnv parse comma-separated prefix:factor pairs from GBRAIN_SOURCE_BOOST / GBRAIN_SEARCH_EXCLUDE env vars. resolveBoostMap and resolveHardExcludes merge defaults + env + caller SearchOpts.exclude_slug_prefixes/include_slug_prefixes.
  • src/core/search/sql-ranking.ts (v0.22.0, extended v0.41.34.0) — Pure SQL string builders. buildSourceFactorCase(slugColumn, boostMap, detail) emits a CASE expression with longest-prefix-match wins (returns literal '1.0' when detail === 'high' for temporal-bypass parity with COMPILED_TRUTH_BOOST). buildHardExcludeClause(slugColumn, prefixes) emits NOT (col LIKE 'p1%' OR col LIKE 'p2%') — OR-chain wrapped in NOT, NOT NOT LIKE ALL/ANY (those quantifiers don't express set-exclusion). LIKE meta-character escape covers all three of %, _, AND \ (backslash matters because it's Postgres LIKE's default escape char). Single-quote doubling on SQL string literals so injection-style inputs are inert text. v0.41.34.0 (retrieval-maxpool incident, T1): new buildBestPerPagePoolCte(...) is the shared per-page max-pool CTE both engines' searchVector inject — instead of returning the single best chunk per page from an inner ORDER BY embedding <=> vec LIMIT N (which let a page lose to a neighbor on ONE weak chunk while its strong chunk sat just below the inner cut), the CTE pools the BEST chunk score per (source_id, slug) composite key, so a page surfaces on its strongest evidence. Composite key (not bare slug) keeps multi-source brains correct. Single source of truth so the two engines can't drift on the pooling SQL.
  • src/core/search/title-match.ts (v0.41.34.0, retrieval-maxpool incident, T2) — pure, zero-I/O title-phrase matcher shared by the production title boost AND NamedThingBench (no drift). isTitlePhraseMatch(query, title) returns true when the normalized query is a contiguous token run inside the title with >= MIN_CONTENT_TOKENS=2 non-stopword tokens, OR an exact full-title match (covers deliberate 1-word chosen names like "Mingtang"). Token-boundary matching (never raw substring, so "art" doesn't match "Bartholomew"); small conservative English stopword set excluded from the content-token floor (Codex#10 guard against promoting generic pages on stopword-y queries); NFKC normalize so CJK / width variants converge. Exports tokenizeTitle + __test__ internals. Closes the direct regression: "Greek amphitheater" matched a weak body chunk instead of the page whose title contained the phrase.
  • src/core/search/alias-normalize.ts (v0.41.34.0, retrieval-maxpool incident, T3) — ONE normalizer shared by the WRITE path (ingest projects frontmatter aliases: into page_aliases) and the READ path (search matches query against page_aliases), so stored aliases can't silently fail to match queries via divergent normalization (same single-source posture as cjk.ts). normalizeAlias(raw) does NFKC + lowercase + whitespace-collapse + trim + strip one layer of wrapping quotes/brackets; returns '' for empty (callers MUST skip empty aliases). normalizeAliasList(value) coerces a frontmatter scalar / array / comma-list / garbage into a deduped list of normalized non-empty aliases — used by both the ingest projection and the reindex --aliases backfill.
  • src/core/search/evidence.ts (v0.41.34.0, retrieval-maxpool incident, T4) — the agent-facing why-it-matched contract that closes the incident's root behavior (agent read one blended score 0.64, decided "no strong match, safe to create", wrote a duplicate over a fully-developed page). classifyEvidence(r) names the strongest signal (precedence: alias_hit > exact_title_match > high_vector_match (base ≥ HIGH_MATCH_FLOOR=0.85) > keyword_exact (base ≥ SOLID_MATCH_FLOOR=0.6) > weak_semantic). createSafetyFor(evidence) derives the don't-duplicate hint (exists / probable / unknown) the agent keys off INSTEAD of a raw threshold (a blended RRF/cosine score is not a calibrated probability, Codex#5). stampEvidence(results) stamps evidence + create_safety in place once at pipeline end (after the alias hop, before slice); idempotent.
  • src/core/search/mode.ts extension (v0.41.34.0, retrieval-maxpool incident, T2) — new title_boost: number | undefined knob in ModeBundle (default 1.25 for all three modes; multiplier for the post-fusion title-phrase boost). Override chain: per-call SearchOptssearch.title_boost config (clamped [1.0, 5.0]) → bundle. KNOBS_HASH_VERSION bumps 6→7 with a tib= parts entry so a title-boost-on cache write can't be served to a title-boost-off lookup. SEARCH_MODE_CONFIG_KEYS gains search.title_boost.
  • src/commands/search-diagnose.ts (v0.41.34.0, retrieval-maxpool incident, T0) — gbrain search diagnose "<query>" --target <slug> [--json] [--source <id>]: Phase-0 retrieval diagnostic. Traces WHERE a target page surfaces (or fails to) across keyword / vector (per-page max-pool) / alias / hybrid layers and names the layer responsible for an incident, so an operator can pin whether the fix is max-pool/innerLimit (vector) vs title/alias. The verdict names the layer that DOES surface the target (or "none"). Pinned by test/search/search-diagnose.test.ts.
  • src/commands/reindex-aliases.ts (v0.41.34.0, retrieval-maxpool incident, T8) — gbrain reindex --aliases [--limit N] [--dry-run] [--json] [--source <id>]: backfills the free-text alias layer for EXISTING pages whose frontmatter aliases: predate v110 (the import-time projection covers new + changed pages). Reads each page's frontmatter aliases:, writes via engine.setPageAliases. Idempotent + convergent (setPageAliases replaces a page's alias set) so no op-checkpoint needed; walks listAllPageRefs (cheap cross-source enumeration), --source narrows. Pinned by test/search/reindex-aliases.test.ts.
  • src/eval/retrieval-quality/harness.ts + src/commands/eval-retrieval-quality.ts + test/fixtures/retrieval-quality/namedthing.jsonl (v0.41.34.0, retrieval-maxpool incident, T6) — NamedThingBench, the retrieval-quality eval that makes the incident impossible to reintroduce silently. Seven query families, each a distinct failure class: title-substring (the direct regression), generic-to-named (tourist label → named thing), alias-synonym (declared alias / romanization → canonical), multi-chunk-dilution (one strong chunk among many weak — stresses max-pool), short-vs-rich, graph-relationship (guardrail), hard-negative (precision guard, must NOT return a page). gbrain eval retrieval-quality <fixture.jsonl> runs it; hard gates (e.g. title-substring Hit@1 ≥ 0.95, alias Hit@1 ≥ 0.98, multi-chunk-dilution Hit@3 = 1.0). Pure: caller injects a SearchFn (CLI uses hybridSearch, tests stub) so it's engine-agnostic. Metric glossary entries (hit@1 / hit@3) added to src/core/eval/metric-glossary.ts. Pinned by test/eval-retrieval-quality.test.ts + test/retrieval-quality-harness.test.ts.
  • docs/architecture/RETRIEVAL.md + docs/architecture/RETRIEVAL_MAXPOOL_INCIDENT.md (v0.41.34.0) — retrieval-pipeline architecture reference + the named-thing-miss incident write-up (root cause, the five-layer fix, the eval that pins it).
  • src/core/types.ts extension + src/core/operations.ts:search + src/core/import-file.ts + src/cli.ts + src/core/search/telemetry.ts (v0.41.34.0, retrieval-maxpool incident) — the wiring layer for the cathedral. SearchResult gains evidence, create_safety, title_match_boost, and alias_hit (all optional; evidence/create_safety reference the union types in evidence.ts). The search MCP op (T4/D4/D15) switches from keyword-only to a cheap-hybrid path by default and accepts a per-call mode (conservative|balanced|tokenmax) honored ONLY for trusted/local callers (resolvePerCallMode(ctx, ...) — remote callers use the configured mode so a remote provider can't force tokenmax spend); every search path stamps evidence fail-soft. importFromContent projects frontmatter aliases: into page_aliases via normalizeAliasList + engine.setPageAliases (T3 WRITE side) so new + changed pages register aliases at ingest. src/cli.ts adds the gbrain search diagnose dispatch (lazy import) and reconciles the search CLI path with the cheap-hybrid op. src/core/search/telemetry.ts extends the rollup with the T7 rank-1 base_score drift signal (sum/count + 3 coarse buckets, aggregate not per-query per D10) surfaced via gbrain search stats, backed by migration v111's search_telemetry columns. Tests: test/cli-search-dispatch.test.ts, test/search/per-call-mode.test.ts, test/search/telemetry-rank1.test.ts, test/search/title-boost-stage.test.ts, test/search/alias-hop.test.ts, test/search/evidence.test.ts, test/search/searchvector-maxpool.test.ts, test/search/pre-migration-failopen.test.ts.
  • src/commands/eval.tsgbrain eval command: single-run table + A/B config comparison. v0.25.0 adds sub-subcommand dispatch on args[0] so gbrain eval export + gbrain eval prune + gbrain eval replay route into session-capture handlers; bare gbrain eval --qrels … fall-through preserves the legacy IR-metrics flow. v0.27.x adds gbrain eval cross-modal to the dispatch (the user-facing path is the cli.ts no-DB branch — src/commands/eval.ts:cross-modal only fires when callers re-enter with an existing engine).
  • src/commands/eval-cross-modal.ts (v0.27.x, extended v0.40.1.0 Track D) — multi-model quality gate. Three different-provider frontier models score the OUTPUT against the TASK on a 5-dim list. Verdict pass (exit 0) / fail (exit 1) / inconclusive (exit 2; <2/3 model successes per Q3=A in plans/radiant-napping-lerdorf.md). Reuses src/core/ai/gateway.ts:chat() so config/auth/aliasing comes from the gateway recipe registry — no parallel provider stack. Self-configures the gateway (configureGateway(loadConfig() + process.env)) since the cli.ts dispatch bypasses connectEngine(). Default cycles 3 in TTY, 1 in non-TTY (T11=B partial cost guardrail). Receipts land at gbrainPath('eval-receipts')/<slug>-<sha8-of-output>.json. The full --budget-usd cap is a v0.27.x follow-up TODO. v0.40.1.0 Track D (T3+T4, per D1/D5/D6/D10): new --batch <jsonl> [--limit N] [--concurrent N] [--max-usd FLOAT] [--yes] flags fan out cross-modal scoring across a LongMemEval-shape JSONL. Mutually exclusive with --task (fail-fast usage error if both set). Filters kind: "by_type_summary" rows from the input (Codex #6). Pre-flight cost estimate refuses if > --max-usd without --yes; default cap 5.00 USD. Semaphore-bounded fan-out via inline runWithLimit<T>(items, limit, fn) (~15 LOC, exported for unit tests): max N questions in-flight at once × 3 model slots = ceiling of 3N parallel API calls (default --concurrent 3 → 9). Per-question receipts land in a per-batch tempdir and are deleted at end of run (per D10 — keeps ~/.gbrain/eval-receipts/ clean); the summary receipt inlines per-question verdicts as JSON, not file paths. Exit precedence (NEW batch-level policy, NOT inherited from aggregate.ts): ERROR > FAIL > INCONCLUSIVE > PASS. DI seam: runEvalCrossModal(args, opts?: {runEval?: typeof runEval}) mirrors runEvalLongMemEval(args, {client?}) at eval-longmemeval.ts:299; tests pass opts.runEval to bypass real LLM calls AND the gateway availability check. Pinned by 15 cases in test/eval-cross-modal-batch.test.ts.
  • src/core/cross-modal-eval/json-repair.ts (v0.27.x) — parseModelJSON(raw) named export with a 4-strategy fallback chain (direct parse → fence-strip → trailing-comma + single-quote + embedded-newline repair → regex nuclear option). Adversarial input throws rather than fabricating scores — the aggregator treats a throw as "this model contributed nothing this cycle" so the gate stays correct at >=2/3 successes.
  • src/core/cross-modal-eval/aggregate.ts (v0.27.x) — pure verdict logic. Pass criterion: (successes >= 2) AND (every dim mean >= 7) AND (every dim min across models >= 5) (Q2=A floor). Inconclusive when <2/3 models returned parseable scores (Q3=A regression guard for the v1 .mjs Object.values({}).every(...) === true empty-array PASS bug).
  • src/core/cross-modal-eval/runner.ts (v0.27.x) — orchestrator. Each cycle runs Promise.allSettled([gwChat(slotA), gwChat(slotB), gwChat(slotC)]) (T4=A — bare allSettled, no rate-leases for the CLI path; minion-integration TODO recovers cross-process concurrency). Stops early on PASS or INCONCLUSIVE; runs up to 3 cycles. Default slots: openai:gpt-4o / anthropic:claude-opus-4-7 / google:gemini-1.5-pro. estimateCost() exports a small per-model pricing table (drifts; refresh alongside model-family bumps).
  • src/core/cross-modal-eval/receipt-name.ts (v0.27.x) — receipt filename binds (slug, SKILL.md sha-8). findReceiptForSkill(skillPath, receiptDir) returns 'found' | 'stale' | 'missing' (T10=A). Skillify-check item 11 surfaces the status as informational (T7=C); the audit does NOT fail on missing/stale receipts.
  • src/core/cross-modal-eval/receipt-write.ts (v0.27.x) — wraps fs.writeFileSync with mkdirSync({recursive:true}) ahead of every write (T5 correction; gbrainPath() does NOT auto-mkdir).
  • src/commands/eval-export.ts (v0.25.0) — streams eval_candidates rows as NDJSON to stdout with schema_version: 1 prefix on every line. EPIPE-safe, progress heartbeats on stderr, stable id-desc tiebreaker so --since windows never dupe/miss rows.
  • src/commands/eval-prune.ts (v0.25.0) — explicit retention cleanup. Requires --older-than DUR. --dry-run reports would-delete count.
  • src/commands/eval-replay.ts (v0.25.0, extended v0.41.0.0) — contributor-facing replay tool. Reads NDJSON from gbrain eval export, re-runs each captured query / search op against the current brain, computes set-Jaccard@k between captured + current retrieved_slugs, top-1 stability rate, and latency Δ. Stable JSON shape (schema_version: 1) for CI gating; human mode prints a regression table. Pure Bun, zero new deps. The dev-loop half of BrainBench-Real that closes the gap between "data captured" and "data used to gate a PR." See docs/eval-bench.md for the workflow. v0.41.0.0 changes: parseNdjson now skips lines where _kind === 'baseline_metadata' so gbrain bench publish-written baselines parse cleanly without the metadata header polluting row counts. New exported replayCore(engine, opts): Promise<{summary, results}> programmatic entrypoint + exported ReplaySummary type so gbrain eval gate can call replay in-process (NOT spawn subprocess — avoids gbrain-version-drift for source-tree CI runs). The existing CLI runEvalReplay now wraps replayCore.
  • src/core/bench/baseline-file.ts + src/core/bench/qrels-file.ts + src/core/bench/correctness-gate.ts + src/commands/bench-publish.ts + src/commands/eval-gate.ts (v0.41.0.0) — the eval-loop wave. Two new CLI verbs: gbrain bench publish --from <captured.ndjson> --to <X.baseline.ndjson> writes a baseline file from gbrain eval export output (stamps stable query_hash on every row; metadata header carries _kind: 'baseline_metadata' + thresholds + source_hash + baseline_mean_latency_ms; deterministic sort by (tool_name, query_hash); D4 strict posture — empty=fail, dupes=fail with paste-ready hint, --to exists=refuse without --force); gbrain eval gate [--baseline X] [--qrels Y] is the two-gate dispatcher (regression gate via in-process replayCore, correctness gate via bare hybridSearch for determinism per eng-D6, both must pass when both flags set, exit 0 PASS / 1 FAIL / 2 USAGE). Source-id-aware: bench publish dedup key is (tool_name, source_ids, query_hash); qrels compare keys are ${source_id}::${slug} everywhere; closes the multi-source-bug-class (canonical gbrain pitfall) at the file-shape layer. Latency math: (baseline + delta) / baseline <= multiplier (the original delta / baseline formula would have let 2.5x slowdowns pass at multiplier=2.0). D3 fail-closed: ANY in-process throw flips verdict to fail with named breach in breaches[] — never silently exit 0. .qrels.json shape preserves the existing 12-row test/fixtures/eval-baselines/qrels-search.json fixture (slug-only relevant_slugs + first_relevant_slug auto-promote to source_id='default') AND supports the federated shape (explicit relevant: [{source_id, slug}] + expected_top1). correctness-gate.ts runs each qrels query via bare hybridSearch; per-query throw recorded as errored: true and flagged as gate failure (Finding 2D). Audit JSONL for bench-publish at ~/.gbrain/audit/bench-publish-YYYY-Www.jsonl. Plan + 23 decisions + 2 codex outside-voice rounds at ~/.claude/plans/system-instruction-you-are-working-rustling-peacock.md. Pinned by 65 cases across test/bench/baseline-file.test.ts (9) + test/bench/qrels-file.test.ts (19) + test/bench/correctness-gate.test.ts (6) + test/bench-publish.test.ts (10) + test/eval-gate.test.ts (10) + test/eval-replay-metadata-skip.test.ts (2) + test/cycle/nightly-probe-adapters.test.ts (6) + test/autopilot-nightly-probe-wiring.test.ts (8) + test/e2e/eval-loop.test.ts (4).
  • src/core/cycle/nightly-probe-adapters.ts (v0.41.0.0) — bridges the autopilot's object-shape NightlyProbeDeps to the existing argv-shape runEvalLongMemEval + runEvalCrossModal CLI functions. Cross-modal adapter argv MUST include --output summaryPath (codex round-2 #1 — without it the summary lands at the default receipt path and the adapter would read nothing from summaryPath). In-process invocation (NOT subprocess) — avoids gbrain-version-drift class for source-tree CI runs. ~125 LOC. Pinned by test/cycle/nightly-probe-adapters.test.ts (6 cases including argv-shape regression for the --output fix).
  • src/commands/autopilot.ts extension (v0.41.0.0) — tick body invokes runNightlyQualityProbe when cfg.autopilot.nightly_quality_probe.enabled === true (default OFF — opt-in to protect API spend on fresh installs). Per eng-D10 (codex round-1 #11): NO scheduler-side rate-limit check — runNightlyQualityProbe's internal shouldRunNightly (reading the audit JSONL) is the single source of truth. Probe call wrapped in try/catch that logs via logError and DOES NOT bump consecutiveErrors (probe failure is informational, never crashes the loop). Default max_usd cap = 5. Source-shape pinned by test/autopilot-nightly-probe-wiring.test.ts (8 regression assertions).
  • test/eval-replay-gate.test.ts + test/fixtures/eval-baselines/qrels-search.json (v0.40.1.0 Track D / T5, per D8 reshape) — hermetic retrieval qrels gate that runs in the standard PR unit-shard CI matrix (.github/workflows/test.yml, NOT the fixed-file E2E workflow). Structural replacement for the original "replay against captured eval_candidates baseline" design (deferred to v0.41+: contributor-mode CI capture in TODOS.md; Codex outside-voice caught three fatal flaws — select-e2e.ts is local-only, eval-export bypasses op-layer capture on PGLite-seed tests, replay re-embeds via gateway which needs an API key CI doesn't have). Uses the canonical PGLite block from CLAUDE.md test-isolation rules (R3+R4) and the basis-vector embedding pattern from test/e2e/search-quality.test.ts:23-28 for fully hermetic retrieval. The qrels fixture (12 queries) is hand-curated with PLACEHOLDER names only (alice-example, widget-co-example, etc. — per Codex #9 + CLAUDE.md privacy rule) and embeds each query at a deterministic basis dimension so retrieval is reproducible. Each query lists relevant_slugs[] + first_relevant_slug; for each, the test computes top1_match_rate (top-1 == first_relevant) and recall@10 (fraction of relevant_slugs in top-10), asserting both meet floors (defaults >= 0.80 and >= 0.85). Env-overridable floors via GBRAIN_REPLAY_GATE_TOP1_FLOOR / GBRAIN_REPLAY_GATE_RECALL_FLOOR (with withEnv() per CLAUDE.md R1). Refresh discipline (per D4): when ranking changes intentionally move expected slugs, edit qrels-search.json directly and include a Why: line in the commit body so future maintainers can read the audit trail. Without Why:, the gate degrades to rubber-stamp within months. Pinned by 5 cases in test/eval-replay-gate.test.ts including a privacy-grep regression guard against real-name reintroduction.
  • src/core/cycle/nightly-quality-probe.ts + src/core/audit-quality-probe.ts + test/fixtures/longmemeval-nightly.jsonl + test/nightly-quality-probe.test.ts (v0.40.1.0 Track D / T6+T7+T8, per D12) — opt-in nightly cross-modal quality probe. The phase runs gbrain eval longmemeval --by-type against the committed 10-question placeholder fixture, pipes the output through gbrain eval cross-modal --batch --max-usd 5 --yes, and writes one event per run to ~/.gbrain/audit/quality-probe-YYYY-Www.jsonl (ISO-week-rotated, mirrors audit-slug-fallback.ts; honors GBRAIN_AUDIT_DIR). Default DISABLED — opt-in via gbrain config set autopilot.nightly_quality_probe.enabled true (prevents surprise API spend on gbrain init). 24h rate limit (the pure shouldRunNightly(now, recentEvents, windowMs?) function) skips with audit row outcome: rate_limited when a recent run exists. Embedding-key short-circuit: longmemeval needs gateway.embedQuery(), so the phase exits early with outcome: no_embedding_key + stderr warn when no provider is configured. Full DI surface via NightlyProbeDeps (isEnabled, hasEmbeddingProvider, resolveMaxUsd, resolveRepoRoot, runLongMemEval, runCrossModalBatch, now) so the unit test stubs every external effect — no PGLite, no real LLM calls, no env mutation outside withEnv(). Cost ceiling: $5/run × 30 nights ≈ $150/month worst-case; expected real cost ~$0.35/night × 30 ≈ $10.50/month. Autopilot scheduler wiring deferred to v0.41+ follow-up (filed in TODOS.md) — the phase is callable in isolation today; the cycle-loop dispatcher hasn't been wired to invoke it on the 24h cadence yet. New nightly_quality_probe_health doctor check (in src/commands/doctor.ts right after slug_fallback_audit) reads the last 7 days of audit events: SKIPPED when feature flag is off (with paste-ready enable command); OK when enabled + all PASS; WARN on any FAIL / ERROR / BUDGET_EXCEEDED in the window with per-outcome counts. Pinned by 14 cases in test/nightly-quality-probe.test.ts (rate-limit pure-function unit tests + DI-stubbed end-to-end phase tests across every outcome branch).
  • src/commands/eval-trajectory.ts + src/commands/founder-scorecard.ts + src/core/trajectory.ts (v0.35.7) — temporal trajectory + founder scorecard. The wave that turns the v0.35.3.1 date-aware contradiction probe into a useful temporal substrate. gbrain eval trajectory <entity> shows the chronological typed-claim history (mrr/arr/team_size/etc) with regressions auto-flagged inline; gbrain founder scorecard <entity> rolls up claim_accuracy / consistency / growth_trajectory / red_flags into one JSON. Pure-function math lives in trajectory.ts: detectRegressions(points, threshold) walks consecutive metric-value pairs per metric (10% drop default, env override GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD); computeDriftScore(points) returns 1 - mean(cosine(emb[i], emb[i-1])) over existing embeddings (null when <3 embedded points). Backed by BrainEngine.findTrajectory(opts) — both Postgres and PGLite, single SQL query, deterministic ORDER BY valid_from ASC, id ASC (R3). Source-scoped via the v0.34.1.0 sourceId scalar / sourceIds array dual pattern (D-CDX-6); visibility-filtered for remote callers (D-CDX-1) — recall-equivalent posture. MCP op find_trajectory (read scope, NOT localOnly) registered after find_experts. Migration v67 adds four optional typed-claim columns (claim_metric, claim_value, claim_unit, claim_period) + a partial index on (entity_slug, claim_metric, valid_from) WHERE claim_metric IS NOT NULL. Fence widens from 10 to 14 cells when any row has typed data; renderer stays at 10 cells when none do (no churn diff on existing fences). Metric labels normalize to lowercase snake_case via normalizeMetricLabel (15-entry seed map for common founder metrics). The consolidate cycle phase gains semantic upsert keyed on (page_id, claim, since_date) — fixes the pre-existing F4 duplicate-takes bug where re-running the full cycle after extract_facts cleared consolidated_at would silently append duplicate takes via MAX(row_num)+1. Also writes chronological valid_until on each cluster's older facts. The extract_facts cycle phase batch-embeds via gateway.embed() before insert AND threads pages.effective_date as the pageEffectiveDate fallback for valid_from (precedence chain: fence-row > pageEffectiveDate > now()). The contradiction probe MUST NOT write valid_until — R1+R8 grep guard at test/eval-contradictions/no-valid-until-write.test.ts pins this. Codex outside-voice round caught F1 (v66 collision → v67), F2 (Haiku lives in facts/extract.ts not extract-facts.ts cycle phase), F3 (cycle didn't embed before insert), F4 (idempotency bug), F5+F6 (missed fence-write.ts caller + no Page object there → pageEffectiveDate is OPTIONAL), F7 (privacy regression — visibility filter added), F8 (ParsedFact needed typed-field extension for markdown system-of-record), F9 (dual scalar+federated sourceId). Plan: ~/.claude/plans/system-instruction-you-are-working-curious-jellyfish.md. Tests: 258 across 12 files. v0.40.2.0 (migration v89 facts_event_type_column): facts table gains a nullable event_type TEXT column so the typed-claim substrate can carry event-shaped rows (event_type='meeting', 'job_change', 'location_change') alongside metric-shaped rows. TrajectoryPoint.event_type: string | null projected by both PGLite and Postgres findTrajectory paths. TrajectoryOpts.kind?: 'metric' | 'event' | 'all' filter added (default 'all'). Existing callers (founder-scorecard, eval-trajectory) pass kind: 'metric' explicitly for call-site clarity — no behavior change since both already defensively skipped NULL-metric rows in their per-metric math. Back-compat pinned by test/regressions/v0_40_2_0-trajectory-backcompat.test.ts (4 cases: byte-identical computeFounderScorecard + computeTrajectoryStats output with and without event rows). Engine parity in test/engine-parity-event-type.test.ts (6 cases). Plan: ~/.claude/plans/system-instruction-you-are-working-crystalline-owl.md.
  • src/core/trajectory-format.ts (v0.40.2.0) — shared formatTrajectoryBlock(points, entitySlug, opts) consumed by both gbrain think (production) and the LongMemEval harness (benchmark). Groups by (metric ?? event_type), per-metric cap 20, total cap 100, knowledge_update intent annotates value-change rows with (superseded prior). Emits a <trajectory entity="..."> XML envelope — INJECTION_PATTERNS in src/core/think/sanitize.ts extended to escape </trajectory>, <trajectory ...> open tags, and attribute injection so adversarial fact text can't break out. Pinned by test/trajectory-format.test.ts (17 cases: grouping, caps, sanitization, supersession annotation, determinism, provenance, text-cap, adversarial </trajectory> escape).
  • src/core/think/intent.ts + src/core/think/entity-extract.ts (v0.40.2.0) — pure classifyIntent(question) returns 'temporal' | 'knowledge_update' | 'other' (regex-first, no LLM call, 'other' fast path short-circuits with zero SQL). extractCandidateEntities(question, retrievedSlugs) pulls high-precision candidates from retrieved entity-prefix slugs (people/, companies/, organizations/) and medium-precision noun phrases from the question. Stop-word boundaries + leading-verb stripper handle "When did I last meet Marco" → marco cleanly. Both consumed by runThink and by the LongMemEval harness so the two paths cannot drift. Pinned by test/think-intent.test.ts (14) and test/think-entity-extract.test.ts (10).
  • src/commands/eval-suspected-contradictions.ts + src/core/eval-contradictions/{judge,runner,types,date-filter,cost-tracker,cache,severity-classify,cross-source,trends,calibration,judge-errors,auto-supersession,fixture-redact}.ts (v0.32.6) — gbrain eval suspected-contradictions [run|trend|review]. Probe samples top-K retrieval pairs per query (cross-slug + intra-page chunk-vs-take), date pre-filters (3-rule layered — same-paragraph-dual-date overrides separation rule), LLM judge (query-conditioned per Codex; UTF-8-safe truncation; C1 confidence-floor double-enforcement; resolution_kind output drives M7 paste-ready commands), persistent cache keyed on (chunk_a_hash, chunk_b_hash, model_id, prompt_version, truncation_policy) (Codex outside-voice fix — prompt edits cleanly invalidate prior verdicts), Wilson 95% CI calibration on the headline percentage with small_sample_note when n<30, judge_errors as first-class typed counters (parse_fail/refusal/timeout/http_5xx/unknown — Codex fix to bias from silent skip), M5 trend writes to eval_contradictions_runs, M6 source-tier breakdown reuses DEFAULT_SOURCE_BOOSTS prefix logic, deterministic sampling (combined_score DESC + lex tiebreaker — stable cache hit-rate across re-runs). Hermetic via judgeFn + searchFn DI in the runner; never touches the real gateway in tests. Engine surface: BrainEngine.listActiveTakesForPages (P1 batched), writeContradictionsRun + loadContradictionsTrend (M5), getContradictionCacheEntry + putContradictionCacheEntry + sweepContradictionCache (P2). Schema migrations v51 + v52. MCP op find_contradictions (read scope, NOT localOnly, NOT in subagent allowlist — user-initiated only). M1 doctor check surfaces high-severity findings with paste-ready resolution commands. M2 synthesize phase pre-fetches latest probe's top-5-by-severity findings and threads them into buildSynthesisPrompt as an informational block. 226 hermetic unit tests + 12 real-Postgres E2E. Plan: ~/.claude/plans/system-instruction-you-are-working-hashed-dewdrop.md. Architecture doc: docs/contradictions.md.
  • src/core/think/index.ts (v0.35.5.0 — gateway adapter, extended v0.40.2.0) — runThink no longer instantiates new Anthropic() directly. The internal LLMClient instance is now built by a small adapter that wraps gateway.chat() from src/core/ai/gateway.ts, the canonical AI seam v0.31.12 established for chat/embed/expansion. Closes #952: stdio MCP launches (Claude Desktop, Cursor) don't inherit shell env, so the Anthropic SDK's env-only key resolution lost the key any user had set via gbrain config set anthropic_api_key. The gateway reads from ~/.gbrain/config.json AND from env, so both paths work. Test seam preserved: opts.client?: ThinkLLMClient injection still works for the 12+ existing tests (test/think-pipeline.serial.test.ts, test/think-gateway-adapter.test.ts, etc.); opts.stubResponse continues to short-circuit before any LLM call. When neither key nor client is available, the graceful "no LLM available" stub still fires with the same NO_ANTHROPIC_API_KEY warning. v0.36.x TODO: drop ThinkLLMClient indirection entirely, migrate tests to __setChatTransportForTests seam from src/core/ai/gateway.ts. v0.40.2.0 trajectory injection (default ON): runThink orchestrates classifyIntent(question)extractCandidateEntities(question, retrievedSlugs)findTrajectory (5s Promise.race timeout per candidate, concurrency cap 3) → formatTrajectoryBlock from src/core/trajectory-format.ts. buildThinkUserMessage (in src/core/think/prompt.ts) extended with a trajectory?: ThinkTrajectoryBlockOpts slot that honors BOTH existing prompt orderings (calibration mode: retrieval → calibration → trajectory → question; default mode: question → retrieval → trajectory → instruction). No third ordering invented. The MCP think op handler extracts sourceScopeOpts(ctx) to scalar sourceId / allowedSources / remote fields on RunThinkOpts so federated-read OAuth clients can't see trajectory rows outside their source scope. New config key think.trajectory_enabled (default true) flips the entire path off. Any error in the trajectory path degrades to "no block injected" + TRAJECTORY_INJECTION_FAILED warning — the think call itself never crashes from trajectory. Production path skips fallback_slugify resolutions deliberately (avoid querying invented slugs); the LongMemEval harness accepts them because the extractor and lookup go through the same slugify path. Pinned by test/think-trajectory-injection.test.ts (7 cases: temporal injection with superseded-prior annotation, 'other' short-circuit, withTrajectory: false bypass, think.trajectory_enabled=false bypass, empty-trajectory skip, findTrajectory throw caught by Promise.allSettled, TRAJECTORY_INJECTED warning count). Debug: GBRAIN_THINK_DEBUG=1 gbrain think "..." prints the spliced prompt to stderr.
  • src/core/operations.ts extension (v0.35.5.0 orphans fix) — findOrphanPages (both engines) now filters p.deleted_at IS NULL on the candidate side AND adds JOIN pages src ON src.id = l.from_page_id WHERE src.deleted_at IS NULL to the EXISTS subquery on the link-source side. Pre-v0.35.5 the query filtered nothing on deleted_at, so soft-deleted pages (v0.26.5 soft-delete shipped without updating this query) appeared as orphans AND links from soft-deleted source pages still suppressed live pages from orphan results. Closes #1021. Pinned by test/orphans.test.ts's soft-delete cases.
  • src/commands/eval-longmemeval.ts + src/eval/longmemeval/{harness,adapter,sanitize}.ts (v0.28.1, extended v0.40.1.0 Track D and v0.40.2.0) — gbrain eval longmemeval <dataset.jsonl> runs the public LongMemEval benchmark against gbrain's hybrid retrieval. Architecture: one in-memory PGLite per benchmark run created via createBenchmarkBrain + withBenchmarkBrain (NO EphemeralBrain class). Between questions, TRUNCATE over runtime-enumerated pg_tables so future schema migrations don't silently leak data across questions; infrastructure tables (sources, config, gbrain_cycle_locks, subagent_rate_leases) are preserved. cli.ts has a pre-dispatch bypass so eval longmemeval skips connectEngine() — the user's ~/.gbrain brain is never opened. --expansion defaults to OFF (deterministic, no per-query Haiku call); pass --expansion to opt in. Default model resolves through resolveModel() 6-tier chain with models.eval.longmemeval as the new config key. Sanitization parity: harness.ts re-uses INJECTION_PATTERNS from src/core/think/sanitize.ts (now exported, line 22) so adding a pattern automatically covers takes AND benchmarks. Retrieved chat content is wrapped in <chat_session id="..." date="..."> framing; the answer-gen system prompt declares the content UNTRUSTED. LLM injection seam: runEvalLongMemEval(args, {client?: ThinkLLMClient}) lets tests stub the client so the full pipeline runs without an Anthropic API key. p50 25.9ms / p99 30.3ms warm reset+import+search on Apple Silicon (per test/eval-longmemeval.test.ts perf gate). Hand the JSONL output to LongMemEval's evaluate_qa.py to score (their published evaluator, not bundled — needs OpenAI gpt-4o per their spec). v0.40.1.0 Track D (T1+T2, per D9): per-question JSONL row gains question: string (additive — evaluate_qa.py ignores unknown fields) so the gbrain eval cross-modal --batch consumer has the task text without joining back against the source dataset. Each row also carries question_type: string and recall_hit?: boolean so a --resume-from run can rebuild the cumulative recallByType from the file alone. New --by-type flag emits a {schema_version:1, kind:"by_type_summary", recall_by_type:{...}, aggregate:{...}} line as the FINAL line of the output; resume-replace strips any prior summary at the tail so 5 resumed runs produce 1 summary, not 5 (Codex #7). Empty-bucket guard: aggregate.rate is null (not NaN) when no questions had ground truth. Optional --by-type-floor F (0..1) exits non-zero with a stderr line per breached question_type; default informational only. Pure buildByTypeSummary(buckets) + emitByTypeSummary(path, summary) + seedRecallByTypeFromFile(path, bucket) helpers exported for unit tests. Pinned by 5 new cases in test/eval-longmemeval.test.ts. v0.40.2.0 inline Haiku extractor + trajectory routing (methodology change — read carefully): new src/eval/longmemeval/extract.ts runs extractAndInsertClaims() over each haystack session before retrieval, populating the benchmark brain's facts table inline at import time. Single Haiku call per session with content-hash cache (cuts a 3-iteration benchmark run from $1.50 to $0.50 when sessions repeat across questions). Per-question alias map (fresh per question, never leaks across) — "Marco" + "Marco Smith" + "marco" collapse to one canonical slug via first-mention-wins. Fail-open on every error path (malformed JSON, Haiku throw, insert collision, empty array → inserted: 0 without throwing). getCacheStats() writes empirical hit rate to stderr per run. New src/eval/longmemeval/intent.ts prefers the dataset's question_type label (temporal-reasoning, knowledge-update, etc.) before falling back to the SHARED regex set imported from src/core/think/intent.ts — single source of truth means think and longmemeval cannot drift. runOneQuestion routes temporal/knowledge_update intents through the shared extractCandidateEntitiesfindTrajectory → splice into the answer-gen prompt before the retrieved-sessions block. New CLI flag --no-trajectory bypasses BOTH the extractor and the intent routing (used by the measurement protocol to baseline default-on vs no-trajectory across 3 seeds per condition with paired-bootstrap CI). JSON envelope adds 5 per-question fields when trajectory routing is on: intent, trajectory_points, entity_resolved, resolution_source, methodology_note. The methodology_note also writes to stderr at run completion (extractor=haiku-preprocess-full-haystack-v1) — honest disclosure that the published gbrain LongMemEval number is "gbrain + Haiku-preprocess pipeline" vs "gbrain alone" so it's NOT directly comparable to baseline LongMemEval scores without that note. Pinned by test/longmemeval-extract.test.ts (13), test/longmemeval-intent.test.ts (9), test/longmemeval-trajectory-routing.test.ts (4 end-to-end through runEvalLongMemEval with both clients stubbed).
  • docs/eval-bench.md (v0.25.0) — contributor guide for using captured data to benchmark retrieval changes before merging. Linked from CONTRIBUTING.md under "Running real-world eval benchmarks (touching retrieval code)".
  • src/core/eval-capture.ts (v0.25.0) — op-layer capture wrapper called from src/core/operations.ts query + search handlers. Catches MCP + CLI + subagent tool-bridge from one site. Fire-and-forget; failures route to engine.logEvalCaptureFailure so gbrain doctor sees drops cross-process. Capture is off by defaultisEvalCaptureEnabled resolution: explicit config.eval.capture (true/false) wins, else process.env.GBRAIN_CONTRIBUTOR_MODE === '1', else off. Production users get a quiet brain; contributors set export GBRAIN_CONTRIBUTOR_MODE=1 in .zshrc to enable the dev loop. PII scrubber gate is independent and defaults to true regardless of CONTRIBUTOR_MODE.
  • src/core/eval-capture-scrub.ts (v0.25.0) — zero-deps PII scrubber: emails, phones, SSN, Luhn-verified credit cards, JWT-shaped tokens, bearer tokens.
  • src/core/search/hybrid.ts — Cathedral II Promise<SearchResult[]> return shape unchanged in v0.25.0. Adds onMeta?: (m: HybridSearchMeta) => void callback so op-layer capture can record what hybridSearch actually did. Existing callers leave it undefined. v0.33: HybridSearchOpts.types?: PageType[] (defined on SearchOpts) threads a multi-type filter through to per-engine searchKeyword + searchVector + searchKeywordChunks, where it lands as AND p.type = ANY($N::text[]). Primary consumer is gbrain whoknows (filters to ['person','company']). AND-applies alongside the existing single-value type filter; either or both can be used. v0.36.3.0: hybridSearch now resolves the embedding column at the boundary via resolveColumn(loadRegistry(cfg), opts.embedding_column, cfg) from src/core/search/embedding-column.ts, threads the ResolvedColumn descriptor into per-engine searchVector (not a raw string), and uses isCacheSafe(resolved, cfg) for the cache-skip decision (replaces the prior name-based isDefaultColumn check that leaked across vector spaces when a user repointed the embedding builtin). cosineReScore calls engine.getEmbeddingsByChunkIds(ids, resolved.name) so rerank uses vectors from the active column, not the hardcoded OpenAI embedding. The query MCP op accepts embedding_column for per-call A/B benchmarking; search (keyword-only) deliberately rejects the param. v0.41.34.0 (retrieval-maxpool incident): two new post-fusion stages plus the evidence stamp. applyTitleBoost(results, query, titleBoost, floorThreshold) (T2) multiplies a result's score by the resolved title_boost when isTitlePhraseMatch fires, stamps title_match_boost, and inherits the floor-ratio gate so a title match can't shove a much-stronger page below it. applyAliasHop(engine, results, query, opts) (T3) normalizes the query, calls engine.resolveAliases, and when the normalized query exactly matches a page's declared alias, surfaces that page at top-of-organic + epsilon with alias_hit=true. stampEvidence(...) (T4) runs LAST (after the alias hop, before slice) on every path — keyword-only, no-embed, and full hybrid — so MCP callers and --explain read the same evidence + create_safety contract. title_boost is resolved from the mode bundle and threaded in.
  • docs/eval-capture.md (v0.25.0) — stable NDJSON schema reference for gbrain-evals consumers.
  • test/public-exports.test.ts (v0.25.0 / R2) — runtime contract test. Imports each of the 17 public subpaths via package name and pins a canary symbol per module. Paired with scripts/check-exports-count.sh.
  • src/core/embedding.ts — OpenAI text-embedding-3-large, batch, retry, backoff. v0.28.7: BATCH_SIZE reverted 50→100 — the original Voyage safety guard halved OpenAI throughput on every page. Per-recipe pre-split + recursive halving + adaptive shrink-on-miss now live in the gateway, so the outer paginator goes back to its original purpose: progress-callback granularity, not batch protection. v0.41.31.0 (cost-model + stale-semantics wave): estimateEmbeddingCostUsd(tokens) now prices against the CURRENTLY-CONFIGURED model's rate instead of the hardcoded OpenAI EMBEDDING_COST_PER_1K_TOKENS. Pre-fix, a brain on a cheaper model (e.g. ZeroEntropy) saw a sync --all preview that named OpenAI and over-stated spend ~2.6x. New currentEmbeddingPricePerMTok() resolves the per-1M-token rate via lookupEmbeddingPrice(gatewayGetModel()) from embedding-pricing.ts, falling back to the OpenAI 3-large rate (0.13) only when the gateway is unconfigured (unit tests, pre-connect preview) or the model is unknown to the pricing table. EMBEDDING_COST_PER_1K_TOKENS is retained for back-compat with direct importers/tests. New currentEmbeddingSignature(): string returns the embedding-provenance signature <provider:model>:<dims> (e.g. openai:text-embedding-3-large:1536) stamped onto pages.embedding_signature at every embed-write site; DELIBERATELY excludes the chunker version (that's tracked separately via pages.chunker_version) — this signature is strictly the EMBEDDING space, so a model OR dimension swap makes the stored signature differ from the current one and a page becomes stale. Same unconfigured-gateway fallback as the cost helpers. New willEmbedSynchronously({v2Enabled, serialFlag, noEmbed}): SyncEmbedMode is the single source of truth for whether gbrain sync --all embeds at sync time ('inline') or defers to per-source embed-backfill minion jobs ('deferred') — mirrors sync.ts's effectiveNoEmbed resolution exactly (v2Enabled && !serialFlag && !noEmbed → deferred) so the cost gate and the actual embed decision can never drift. New shouldBlockSync(costUsd, floorUsd, mode): boolean is the pure cost-gate decision: blocks ONLY when mode === 'inline' && costUsd > floorUsd — deferred mode never blocks (the backfill's $X/source/24h cap is the real money gate). Pinned by test/sync-cost-preview.test.ts + test/embedding-signature-stale.test.ts.
  • src/core/ai/dims.ts (v0.33.1.1, PR #962 + #866) — per-provider providerOptions resolver for embed-time dimension passthrough. The single source of truth for "which provider needs which knob to produce vector(N)". Exports dimsProviderOptions(implementation, modelId, dims) (called by embed() in gateway.ts), VOYAGE_OUTPUT_DIMENSION_MODELS (private const — the 7 hosted Voyage models that accept output_dimension: voyage-4-large, voyage-4, voyage-4-lite, voyage-3-large, voyage-3.5, voyage-3.5-lite, voyage-code-3 — nano deliberately excluded), VOYAGE_VALID_OUTPUT_DIMS = [256, 512, 1024, 2048] as const, supportsVoyageOutputDimension(modelId), and isValidVoyageOutputDim(dims). Voyage path uses the SDK-supported dimensions field ({ openaiCompatible: { dimensions: N } }), NOT Voyage's output_dimension wire-key — the existing voyageCompatFetch shim in gateway.ts:541 translates dimensions → output_dimension before the HTTP body is built. The reverse (sending output_dimension from here) was the v0.33.1.0 bug class: the AI SDK's openai-compatible adapter doesn't recognize the wire-key so it was silently dropped, Voyage returned its default 1024-dim, and the gateway dimension check threw on every embed call. Runtime guard: when a Voyage flexible-dim model is configured with dims outside VOYAGE_VALID_OUTPUT_DIMS, throws AIConfigError with a paste-ready gbrain config set embedding_dimensions <256|512|1024|2048> hint at the embed boundary — fail-loud instead of opaque Voyage HTTP 400. Most common trigger: embedding_model: voyage:voyage-4-large configured without embedding_dimensions (falls back to DEFAULT_EMBEDDING_DIMENSIONS=1536, an OpenAI default not a Voyage one). Eva (@100yenadmin) shipped the wire-key fix in #866; Codex P3 follow-up landed the validator + valid-dims allowlist in #962.
  • src/core/ai/types.ts — provider/recipe types. v0.28.7 (#680): EmbeddingTouchpoint extended with optional chars_per_token (default 4 chars/token, matching OpenAI tiktoken on English) and safety_factor (default 0.8, budget-utilization ceiling). Both consulted only when max_batch_tokens is also set. Voyage declares chars_per_token=1 + safety_factor=0.5 to handle dense payloads (CJK/JSON/base64) that overshoot tiktoken. The pre-split budget is max_batch_tokens × safety_factor / chars_per_token. v0.28.11 (#719): EmbeddingTouchpoint.multimodal_models?: string[] model-level allow-list for recipes that mix text-only + multimodal models under one touchpoint (Voyage's 12 models share supports_multimodal: true but only voyage-multimodal-3 accepts /multimodalembeddings). When omitted, recipe-level supports_multimodal is sufficient. AIGatewayConfig.embedding_multimodal_model?: string lets embedMultimodal() route to a different model than embedding_model — brains using OpenAI for text can use Voyage for images without flipping the primary embedding pipeline. v0.37.6.0 (#1210): new Recipe.default_headers?: Record<string, string> (static) and Recipe.resolveDefaultHeaders?(env) (env-templated) seam for per-recipe headers that ride alongside auth on every openai-compat touchpoint. Mutually exclusive (declaring both throws AIConfigError at gateway-configure time); keys conflicting with the resolved auth header (Authorization, the resolver's custom header) are rejected at applyResolveAuth call time so defaults cannot accidentally shadow auth. Used by OpenRouter for the HTTP-Referer + X-OpenRouter-Title + X-Title attribution triple; usable by any future recipe (Together/Groq) that wants attribution.
  • src/core/ai/gateway.ts — unified seam for every AI call. v0.36.3.0: embedQuery(text, opts?) and isAvailable(touchpoint, modelOverride?) accept a model override so the resolved-column path can embed via the column's provider (Voyage / ZeroEntropy / OpenAI) instead of the global default. The hybrid path passes {embeddingModel: resolved.provider, dimensions: resolved.dimensions}; the gateway resolves the matching recipe and routes through its instantiateEmbedding() branch. isAvailable('embedding', 'voyage:voyage-3-large') checks the override's recipe (not the default) so hybrid skips vector search only when the active column's provider is actually down — fixes the CDX-4 bug where a healthy Voyage column would skip vector retrieval because OpenAI happened to be unconfigured. v0.35.0.0: ZeroEntropy support lands. New zeroEntropyCompatFetch shim (sibling to voyageCompatFetch) handles ZE's non-OpenAI-compatible wire shape — rewrites the request URL from /embeddings to /models/embed, injects input_type (default 'document'; 'query' when threaded via providerOptions.openaiCompatible.input_type) and explicit encoding_format: 'float', and rewrites the response from {results: [{embedding}], usage: {total_bytes, total_tokens}} to {data: [{embedding, index}], usage: {prompt_tokens, total_tokens}} so the SDK's openai-compatible Zod schema validates (Voyage's shim hit the same prompt_tokens requirement at :655). Layer 1 (Content-Length) + Layer 2 (per-embedding) OOM caps via a new tagged ZeroEntropyResponseTooLargeError class (kept separate from VoyageResponseTooLargeError because test/voyage-response-cap.test.ts does structural source-text greps pinning the Voyage name — class unification is a deferred cleanup). Wired in instantiateEmbedding() via the same recipe.id === 'zeroentropyai' branch pattern Voyage uses. New gateway.rerank() native HTTP path (no AI-SDK reranking abstraction): resolves the configured reranker model via getRerankerModel(), posts to ${recipe.base_url}/models/rerank with bearer auth, returns RerankResult[] sorted by relevance score. RerankError.reason classifier: auth | rate_limit | network | timeout | payload_too_large | unknown. 5s default timeout (search hot path). Pre-flight payload guard rejects bodies over recipe.touchpoints.reranker.max_payload_bytes with reason: 'payload_too_large' so callers can fail-open without an HTTP call. _rerankTransport test seam mirrors _embedTransport. New gateway.embedQuery(text) companion threads inputType: 'query' through dimsProviderOptions() (now 4-arg). getRerankerModel() accessor + isAvailable('reranker') branch added. configureGateway + reconfigureGatewayWithEngine thread reranker_model through the same path as embedding/expansion/chat. applyResolveAuth + defaultResolveAuth widen touchpoint param to include 'reranker'. v0.34.1.0 (#875): new embedMultimodalOpenAICompat() routes recipes with implementation: 'openai-compatible' (LiteLLM, Anyscale, vLLM, Gemini multimodal via proxy) through the standard /embeddings endpoint with content arrays carrying image_url entries. The pre-existing Voyage /multimodalembeddings path is unchanged; the gateway selects by recipe implementation tag. Runtime dimension validation throws AIConfigError (with model id + observed + expected) before the vector reaches storage when the provider returns a width that doesn't match the recipe's default_dims or the brain's embedding_dimensions config — no more cryptic vector dimension mismatch at INSERT time. Pinned by 11 cases in test/openai-compat-multimodal.test.ts. v0.28.7 (#680): module-scoped _embedTransport defaulting to AI SDK embedMany, with __setEmbedTransportForTests(fn) test seam so tests drive the public embed() function with a stubbed transport instead of probing private helpers. splitByTokenBudget and isTokenLimitError are now exported @internal — pure functions reused directly by the test file. Module-level _shrinkState: Map<recipeId, {factor, consecutiveSuccesses}> halves the recipe's effective safety_factor on token-limit miss (floor 0.05) and heals back ×1.5 toward the ceiling after SHRINK_HEAL_AFTER=10 consecutive successes. configureGateway() walks every registered recipe at construction time and emits a once-per-process stderr warning for any embedding touchpoint missing max_batch_tokens (excluding the canonical OpenAI fast-path recipe). resetGateway() clears _shrinkState, the warned-set, and restores the real transport. ASCII flow diagram embedded in the embed() JSDoc covers the routing decision, recursion + halving, and shrinkState lifecycle. v0.28.11 (#719): embedMultimodal() reads cfg.embedding_multimodal_model first (falls back to cfg.embedding_model for single-model setups). After the existing recipe-level supports_multimodal fast-fail, validates the resolved model against touchpoint.multimodal_models when declared — closes the Voyage-text-only-model-into-multimodal-endpoint footgun before any HTTP call (Codex F1 from PR review). New getMultimodalModel() accessor mirrors getEmbeddingModel / getChatModel so doctor and integration tests can read the gateway state. v0.33.1.1 (#962, Codex P3 follow-up): new exported VoyageResponseTooLargeError tagged class at the top of the file. voyageCompatFetch's two OOM-defense caps (Layer 1 Content-Length check at :595, Layer 2 per-embedding base64 cap at :619) now throw VoyageResponseTooLargeError instead of a generic Error. The inbound response-rewriter's surrounding try/catch (which intentionally swallows parse failures so misshaped Voyage responses fall through to the SDK's JSON parser) checks instanceof VoyageResponseTooLargeError and rethrows. Pre-fix, the Layer 2 throw was silently swallowed and the oversized response returned to the AI SDK anyway — Layer 2 was theatrical. Source-shape regression assertion in test/voyage-response-cap.test.ts pins the instanceof ⇒ throw err line.
  • src/core/ai/recipes/zeroentropyai.ts (v0.35.0.0) — ZeroEntropy openai-compatible recipe declaring BOTH embedding (zembed-1, 7 Matryoshka dims: 2560/1280/640/320/160/80/40) AND reranker (zerank-2 flagship + zerank-1 + zerank-1-small, 5MB payload cap) touchpoints. implementation: 'openai-compatible' (NOT the misspelled 'openai-compat' the original plan draft had — pinned by F1 regression in test/ai/zeroentropy-recipe.test.ts). base_url_default: 'https://api.zeroentropy.dev/v1' already ends with /v1, so the zeroEntropyCompatFetch URL rewrite /embeddings → /models/embed produces …/v1/models/embed (NOT …/v1/v1/… — pinned by F2 regression). chars_per_token: 1 + safety_factor: 0.5 match Voyage's dense-content hedge.
  • src/core/ai/recipes/llama-server-reranker.ts (v0.40.6.1) — sibling of llama-server (the embedding recipe) for llama.cpp running in --reranking mode. Distinct recipe rather than dual-touchpoint extension because --reranking and --embeddings are mutually exclusive at server-launch time, so the two backends need independent base URLs (default 8081 here vs 8080 there). Declares reranker touchpoint with models: [] (user-provided id matching the --alias the user launched llama-server with), path: '/rerank' (leaf-only; consumes the new RerankerTouchpoint.path override on the touchpoint; gateway concatenates with base_url_default which already ends in /v1, producing …/v1/rerank — original draft of this recipe set path: '/v1/rerank' which doubled the prefix, caught by post-impl /codex review and corrected before merge), default_timeout_ms: 30_000 (consumed by src/core/search/mode.ts's reranker timeout resolution chain — gives CPU-only first-call warmup headroom; the 5s mode-bundle default would fail-open silently as timeout), cost_per_1m_tokens_usd: 0 (recognized by FREE_LOCAL_RERANK_PROVIDERS in src/core/budget/budget-tracker.ts so --max-cost callers don't TX2 hard-fail on local rerank). Setup hint emphasizes --alias because llama-server's /v1/models defaults model id to the gguf file path without it, which makes provider:model config strings ugly. Covers Qwen3-Reranker via llama.cpp AND self-hosted ZE weights via llama.cpp — same recipe, different --model at launch (quality parity with ZE-hosted NOT guaranteed; user pins their own eval). Pinned by test/ai/recipe-llama-server-reranker.test.ts. Voyage / Cohere / vLLM rerankers stay out of scope — different wire shapes need adapter hooks designed in a follow-up plan. Same wave (v0.40.6.1) adds: path?: string + default_timeout_ms?: number on RerankerTouchpoint in src/core/ai/types.ts; consumed by the URL build at src/core/ai/gateway.ts:rerank() and by mode-resolution at src/core/search/mode.ts:resolveSearchMode (precedence: per-call > config-key > recipe touchpoint default > mode bundle); LLAMA_SERVER_RERANKER_BASE_URL env passthrough in src/cli.ts:buildGatewayConfig; FREE_LOCAL_RERANK_PROVIDERS set in src/core/budget/budget-tracker.ts:lookupPricing (rerank-kind-only zero-pricing for the local provider prefix); doctor-fix at src/commands/models.ts:probeRerankerConfig to read search.reranker.model via loadSearchModeConfig + resolveSearchMode (closes file-plane / DB-plane divergence where doctor said "not configured" while live search was actively reranking — pre-fix bug was getRerankerModel() reading GBrainConfig.reranker_model, a file-plane field nothing writes); probeRerankerReachability upgraded to read the recipe's default_timeout_ms so CPU-only cold-start doesn't false-fail.
  • src/core/ai/recipes/openrouter.ts (v0.37.6.0, #1210) — OpenRouter openai-compatible recipe: single key, many providers via openrouter:<provider>/<model> strings. base_url_default: 'https://openrouter.ai/api/v1'. Embedding touchpoint: openai/text-embedding-3-small at 1536 dims with Matryoshka dims_options: [512, 768, 1024, 1536] (native breakpoints from Weaviate's MRL analysis); max_batch_tokens: 300_000 = OpenAI's aggregate-per-request token cap (NOT per-input — Codex caught the semantic in pre-merge review). Chat touchpoint declares 8 curated entry points (gpt-5.2, gpt-5.2-chat, gpt-5.5, claude-haiku-4.5, claude-sonnet-4.6, claude-opus-4.7, gemini-3-flash-preview, deepseek-chat) but openai-compat tier accepts any model ID; deliberately no max_context_tokens because OR's catalog spans 128K to 1M+. supports_subagent_loop: false is INFORMATIONAL — the real gate is isAnthropicProvider() in src/core/model-config.ts which hard-pins gbrain's subagent infra to Anthropic-direct. The recipe declares resolveDefaultHeaders(env) (the env-templated variant of the new default_headers seam) returning OR's three attribution headers: HTTP-Referer (required for OR to create an app-attribution entry), X-OpenRouter-Title (current preferred name per OR docs), X-Title (documented back-compat alias). Defaults to https://gbrain.ai / gbrain; forks (downstream agent deployments) override via OPENROUTER_REFERER / OPENROUTER_TITLE env vars so their traffic gets their attribution on OR's leaderboard. Smoke-tested by test/ai/recipe-openrouter.test.ts (11 cases including the D5 shape-test regression guard: every model in the chat list matches ^[a-z0-9-]+\/[a-z0-9._-]+$, catching typos and malformed IDs without pinning the catalog's churn).
  • src/core/rerank-audit.ts (v0.35.0.0) — failure-only JSONL audit at ~/.gbrain/audit/rerank-failures-YYYY-Www.jsonl (ISO-week rotation, mirrors src/core/audit-slug-fallback.ts). Exports logRerankFailure({reason, model, query_hash, doc_count, error_summary}) + readRecentRerankFailures(days). Deliberately no logRerankSuccess (CDX2-F22 in plan review): writing once per tokenmax search is hot-path I/O churn AND success events leak query volume + timing into a local audit file. gbrain doctor's reranker_health check reads search.reranker.enabled first so "no events in window" is interpreted correctly (disabled → ok; enabled → ok). Query text is SHA-256-prefix-hashed (8 hex chars) for privacy. GBRAIN_AUDIT_DIR env override honored via the shared resolveAuditDir().
  • src/core/search/embedding-column.ts (v0.36.3.0) — single source of truth for "which content_chunks.* column does this query rank against?" Pure functions, no engine I/O: loadRegistry(cfg) walks the embedding_columns config (DB plane, JSON map keyed by column name with {provider, dimensions, type} entries), seeds the OpenAI embedding builtin when unset, validates everything before it lands (column-name regex, type ∈ vector | halfvec, dims in [1, 8192], provider format) using Object.create(null) + Object.hasOwn so a key like constructor rejects instead of resolving to Object.prototype.constructor. resolveColumn(registry, override?, cfg) is the boundary call: returns a ResolvedColumn descriptor (frozen {name, provider, dimensions, type}) honoring per-call override → search_embedding_column config → 'embedding' default. Throws UnknownEmbeddingColumnError with the list of registered names on miss. isCacheSafe(resolved, cfg) compares the full embedding SPACE (provider + dimensions + name) against cfg's default — a user who repointed the embedding builtin at Voyage doesn't accidentally serve OpenAI-shaped cache rows. validateResolvedColumn(descriptor) re-validates hand-rolled descriptors that bypass the registry (internal-SDK passthrough path) so the SQL-injection escape hatch through the descriptor field is closed. Consumed by hybridSearch (resolves once at the boundary, threads a ResolvedColumn into per-engine searchVector instead of a raw string), gateway.embedQuery(text, {embeddingModel, dimensions}) (resolved column's provider drives the query-time embed call), cosineReScore (engine pulls vectors from the active column, not the hardcoded embedding), and the query MCP op (per-call embedding_column param). 538 lines of pure resolver, 511 unit cases in test/search/embedding-column.test.ts covering the 16 codex-flagged corners (prototype-pollution, descriptor passthrough, env-only Postgres install, empty-brain coverage gate, cache-space comparison).
  • src/core/search/rerank.ts (v0.35.0.0) — the call-site abstraction. applyReranker(query, results, opts) slots between dedupResults() and enforceTokenBudget() in src/core/search/hybrid.ts. Slices opts.topNIn (default 30) by current RRF order, sends to gateway.rerank(), reorders by relevanceScore desc, and appends the un-reranked tail unchanged (recall protection). Fail-open on every RerankError.reason: any error logs via logRerankFailure and returns the input array unchanged. Stamps rerank_score onto reordered items so downstream telemetry sees the new ordering signal. topNOut: null is the explicit "don't truncate" signal — semantically distinct from undefined which means "fall through to mode bundle" (CDX2-F16). Test seam: opts.rerankerFn lets tests stub gateway.rerank without touching the network.
  • src/core/search/return-policy.ts (v0.41.33.0, NEW — default OFF) — intent-aware adaptive return-sizing. Pure, dependency-light module that trims the final ranked candidate set to an intent-driven cap instead of always returning the full top-K (noisy) list. entity intent (single-answer-ish) gets a tight cap; temporal/event/general (enumeration-ish) get a recall-preserving cap. A minKeep failsafe (≥1) guarantees a human never gets a silent blank when candidates exist. WHY a cap, not a score-cliff detector: PrecisionMemBench Phase-1 instrumentation (gbrain-evals) measured that the rank1→rank2 RRF score gap is ~identical whether rank-1 is correct (0.602) or wrong (0.569) — mechanical decay, not a trustworthy separatrix; rank-1 is right in 94% of single-answer cases, so "return a tight set" is the whole win and cliff-cutting just adds noise. Exports AdaptiveReturnConfig, DEFAULT_ADAPTIVE_RETURN (frozen: enabled=false, entityMax=2, otherMax=6, minKeep=1), AdaptiveReturnDecision ({applied, intent, cap, kept, total}), AdaptiveReturnInput (boolean | Partial<AdaptiveReturnConfig> | undefined), adaptiveReturnFromConfig(cfg) (reads the config plane), resolveAdaptiveReturn(perCall, fromConfig) (defaults → config → per-call merge), adaptiveReturnEnabled(...) (gate check for the cache skip), and applyAdaptiveReturn(results, intent, cfg) (the trim). Config knobs (DB or file plane): search.adaptive_return (master switch, boolean), search.adaptive_return_entity_max, search.adaptive_return_other_max, search.adaptive_return_min_keep (each clamped to ≥1). Precision-sensitive agents tune entity=1, other=1 ≈ "return only the top result." Wired into hybridSearch (src/core/search/hybrid.ts): applied AFTER applyReranker, BEFORE the limit slice, and ONLY on the first page (offset===0) — paginating a confidence-gated set is incoherent, so paginated calls fall through to the fixed limit. Stamps the decision onto HybridSearchMeta.adaptive_return for gbrain search --explain. hybridSearchCached SKIPS the cache when the gate is on (a trimmed set must not be served to a gate-off lookup and vice versa); folding the gate params into KNOBS_HASH so adaptive-on calls cache safely is a v0.42+ follow-up before any mode-default flip. SearchOpts.adaptiveReturn + HybridSearchMeta.adaptive_return declared in src/core/types.ts. Agent-facing surface (v0.41.33.0): the query op (src/core/operations.ts) exposes an adaptive_return boolean param whose description instructs the agent WHEN to set it (single-answer questions → on; breadth/exploration → off; pass limit:1 for a hard single-answer cap), threaded into hybridSearchCached. This is the path that matters — end users never touch the config knob; their agent decides per query (same pattern as salience/recency). Pinned by test/search/return-policy.test.ts (mechanism) + test/search/query-op-adaptive-return.test.ts (agent surface: param exists + description teaches both directions + the never-empty contract).
  • src/core/search/autocut.ts (v0.42.3.0, NEW — default ON in reranked modes) — Weaviate-style autocut: score-discontinuity result-sizing on the cross-encoder rerank separatrix. Fixes part of #1663 (the "20 vs 1" precision problem). applyAutocut(results, scoreOf, cfg) normalizes the reranker scores, finds the largest consecutive gap, and cuts there when it clears jumpRatio (default 0.20); robust to unsorted provider output (cuts on a sorted copy, keeps items in INPUT order via a score threshold), guards top<=0/non-finite, never returns empty, and no-ops when <2 results carry a finite rerank_score (covers the reranker fail-open path). WHY rerank_score and NOT RRF/cosine: gbrain measured (see return-policy.ts) that the RRF rank1→rank2 gap is ~flat whether rank-1 is right or wrong — not a separatrix; the cross-encoder score IS. So autocut runs ONLY where the reranker ran (the floor reaches balanced+tokenmax; conservative is a documented no-op). Exports AutocutConfig, DEFAULT_AUTOCUT (frozen: enabled=true, jumpRatio=0.20, minKeep=1), AutocutDecision ({applied, signal:'rerank'|'none', cut, kept, total, gapRatio}), AutocutInput, autocutFromConfig, resolveAutocut, applyAutocut. Cache-key integration (clean path, not the adaptive-return cache-skip hack): enable+sensitivity flow through ModeBundleResolvedSearchKnobsknobsHash exactly like graph_signals. mode.ts adds autocut/autocut_jump (conservative false, balanced/tokenmax true@0.20) AND sets reranker_top_n_in = searchLimit for reranked modes (D4 — so the reranker scores the full returned set; there is no un-scored tail for autocut to wrongly drop — closes the Codex review's load-bearing recall finding). KNOBS_HASH_VERSION 7→8 (master title_boost claimed v=7; autocut appends v=8 — one-time global cache cold-miss on upgrade). Wired into hybridSearch AFTER adaptive-return, BEFORE the limit slice, first page only; emits HybridSearchMeta.autocut. BOTH the cache-miss finalMeta and cache-HIT cachedMeta rebuilds carry autocut+adaptive_return+mode+embedding_column (Codex P2; also fixes a pre-existing adaptive_return meta-drop). Preserves alias-hop exact matches (Codex P1): applyAutocut takes an optional preserve predicate; hybrid passes r => r.alias_hit === true so a canonical page injected by applyAliasHop after reranking (no rerank_score) is never cut. Agent surface: query op autocut boolean (ceiling override — false forces full top-K); SearchOpts.autocut; --explain shows per-result rerank_score (cut visible by which results survive), formatAutocutSummary renders the decision when search meta is threaded; gbrain search modes attribution; metric glossary autocut.signal/autocut.gap_ratio. Config: search.autocut, search.autocut_jump. Default-ON is backed by an in-repo eval gatetest/search/autocut-eval.test.ts (also bun run eval:autocut) measures precision-lift-without-recall-regression over labeled qrels fixtures with modeled cross-encoder distributions (no API key, no sibling repo; runs in CI): mean precision 0.33→0.94, recall 1.00→0.95, ZERO recall regression on enumeration queries (autocut declines on flat curves by construction). Env-overridable floors. A live-corpus PrecisionMemBench run (gbrain-evals) is an optional empirical confirmation, not a blocker. Pinned by test/search/autocut.test.ts (pure-fn), test/search/query-op-autocut.test.ts (agent surface), test/search/autocut-integration.serial.test.ts (IRON-RULE behavioral via rerankerFn DI seam: cliff trims, flat doesn't, no-reranker no-ops, autocut:false ceiling, composes with adaptive-return), test/search/autocut-eval.test.ts (the precision/recall gate), and the v=8 knobsHash assertions in test/search-mode.test.ts. One wave of the #1663 multi-wave redesign; remaining waves (router, tiered tier, CRAG escalation) keep the issue open.
  • src/core/ai/recipes/voyage.ts — Voyage AI openai-compatible recipe. v0.28.7 (#680): declares chars_per_token=1 + safety_factor=0.5 so the gateway pre-splits Voyage batches at a 60K-character budget (50% of 120K-token cap with the dense-tokenizer ratio). Closes the v0.27 backfill loop where ~26% of the corpus stayed un-embedded because tiktoken-grounded budgeting silently undercounted Voyage's actual token usage. v0.28.11 (#719): declares multimodal_models: ['voyage-multimodal-3'] so the gateway rejects text-only Voyage models pointed at the multimodal endpoint with a clear AIConfigError instead of waiting for Voyage's HTTP 400. v0.33.1.1 (#962, fixup): recipe docstring at :7-16 tightened to name the seven hosted flexible-dim models that accept output_dimension explicitly (voyage-4-large, voyage-4, voyage-4-lite, voyage-3-large, voyage-3.5, voyage-3.5-lite, voyage-code-3) and call out that voyage-4-nano is the open-weight variant listed separately by Voyage as fixed 1024-dim — does NOT accept the parameter. The "all v4 variants are flexible" misread is what caused the original PR to include nano in VOYAGE_OUTPUT_DIMENSION_MODELS; the negative regression assertion in test/ai/gateway.test.ts (dimsProviderOptions returns undefined for voyage-4-nano) pins the contract. v0.37.3.0: voyage-code-3 is the recommended embedding model for gstack per-worktree code brains (Topology 3 in docs/architecture/topologies.md). Registration was already in the models list since pre-v0.33; the v0.37.3.0 wave adds discoverability surfaces — decision-tree branch in docs/integrations/embedding-providers.md, Topology 3 "Recommended embedding model" subsection, runtime nudge from gbrain reindex --code against non-code-tuned models. Recipe-shape regression pinned by test/ai/voyage-code-3-recipe.test.ts.
  • src/core/ai/recipes/anthropic.ts — Anthropic recipe (chat + expansion touchpoints). v0.31.12: chat and expansion models: lists drop the v0.31.6 phantom claude-sonnet-4-6-20250929 date suffix — canonical id is claude-sonnet-4-6. The wrong-direction alias claude-sonnet-4-6 → claude-sonnet-4-6-20250929 is removed; a reverse alias claude-sonnet-4-6-20250929 → claude-sonnet-4-6 keeps stale user configs working (rescues facts.extraction_model and models.dream.synthesize set by v0.31.6 installs). Recipe-shape regression pinned by test/anthropic-model-ids.test.ts (6 cases, verbatim cherry-pick of PR #830 plus the reverse-alias rescue case).
  • src/core/anthropic-pricing.ts — Single source of truth for Anthropic model pricing (per-MTok input/output). v0.31.12: Opus 4.7 corrected from $15/$75 to $5/$25 (the old number was from Opus 4 generation, never refreshed when 4.7 shipped); Opus 4.6 also corrected. Consumed by src/core/budget-meter.ts and src/core/cross-modal-eval/runner.ts — the cross-modal estimator now reads ANTHROPIC_PRICING for Anthropic models instead of duplicating the table, killing the v0.31.6 drift bug class.
  • src/core/budget/budget-tracker.ts (v0.37.x) — keystone primitive for the brainstorm cost-cathedral wave. One typed error (BudgetExhausted with reason: 'cost' | 'runtime' | 'no_pricing'), one schema-stable audit JSONL at ~/.gbrain/audit/budget-YYYY-Www.jsonl. Contracts pinned by 18 unit cases: TX1record() throws when cumulative spend exceeds cap (the cap is a real ceiling, not a suggestion); TX2reserve() hard-fails with reason: 'no_pricing' when maxCostUsd is set AND the model is missing from pricing maps (warn-once preserved when cap is unset); A3 amendedextractUsageFromError(err, fallback) returns err.usage when SDK provides it, else the pessimistic fallback (caller passes maxOutputTokens, not the optimistic pre-call estimate). onExhausted(cb) callback fires once synchronously BEFORE the throw propagates so callers can persist checkpoints. Replaces three parallel copies (inline brainstorm class, cycle/budget-meter, eval-contradictions). Adapts the old BudgetMeter via T5 (public shape preserved + schema_version: 1 stamped on every dream-budget audit line).
  • src/core/audit-week-file.ts (v0.37.x, Q1) — single source of truth for ISO-week audit JSONL filename math. Exports isoWeek(d), isoWeekFilename(prefix, now?), resolveAuditDir() (honors GBRAIN_AUDIT_DIR). Year-boundary correctness pinned by tests at 2020-W53 (the 53-week year), 2025-W01 rolling in from 2024-12-30 (Monday), 2026-W01. Four call sites migrated in T4: src/core/minions/handlers/shell-audit.ts, src/core/facts/phantom-audit.ts, src/core/audit-slug-fallback.ts, src/core/cycle/budget-meter.ts. Each call site keeps its compute<X>AuditFilename thin wrapper for back-compat with existing tests.
  • src/core/ai/gateway.ts:withBudgetTracker (v0.37.x, T3 / TX5) — gateway-layer enforcement via AsyncLocalStorage<BudgetTracker>. withBudgetTracker(tracker, fn) installs the tracker on the module-internal store; every gateway.chat / embed / rerank call inside the scope auto-composes (reserve before, record in try/finally). Outside-scope calls are budget no-ops (current behavior preserved). Nested scopes restore the outer tracker on exit. getCurrentBudgetTracker() is the test seam. The chat path uses A3-amended pessimistic fallback on error paths; the embed path estimates input tokens from char count × recipe's chars_per_token because the AI SDK doesn't surface per-batch embed token usage; the rerank path estimates char count of query+docs. 6 unit cases pin the contract.
  • src/core/diarize/payload-fitter.ts (v0.37.x, P6 / Q3) — generic fit-arbitrarily-large-items-into-per-call-token-budget utility. 'batch' strategy is deterministic token-budgeted chunking with no LLM calls. 'summarize' strategy embed-clusters into ceil(items/4) groups via cheap deterministic nearest-neighbor on cosine, Haiku-summarizes each cluster via Promise.allSettled at parallelism=4 (Perf1). Each Haiku call composes the active BudgetTracker via T3's AsyncLocalStorage. The quality gate (codex outside-voice finding #4): when success_ratio < min_success_ratio (default 0.75), result is flagged degraded: true — the fitter preserves the successful subset; the caller decides whether to surface a partial result or abort.
  • src/core/brainstorm/checkpoint.ts (v0.37.x, P7 / TX3+TX4+A5 amended) — crash-resilient checkpoint for gbrain brainstorm and gbrain lsd. Persists FULL idea bodies (~50KB per run) so resume can MERGE the pre-crash ideas with the post-resume ideas before the judge runs (codex's load-bearing finding — a resume that produces only second-run output is silent partial output). run_id = sha256(question + profile + sort(close_slugs) + sort(far_slugs)).slice(0,16) — NO embedding bits, stable across embedding-model swaps. Atomic write via .tmp + rename. ONE resume flag (--resume <run_id> — the proposed --retry-failed was dropped per TX4: failed AND never-attempted crosses both go through --resume). --list-runs prints saved run_ids mtime-newest-first. --force-resume bypasses the 7-day staleness gate. The cycle purge phase (gbrain dream --phase purge) GCs checkpoints older than 7 days via gcStaleCheckpoints(7). Pinned by 20 unit cases + 3 E2E cases in test/e2e/brainstorm-resume.test.ts including the load-bearing merge contract.
  • src/core/remediation-checkpoint.ts (v0.37.x, T7 / A4 amended) — doctor --remediate checkpoint at ~/.gbrain/remediation/<plan_hash>.json. plan_hash = sha256(JSON.stringify(sorted recommendation ids)).slice(0,16). Schema-versioned. Atomic write via .tmp + rename. gbrain doctor --remediate --resume <plan_hash> (or with no arg — picks the newest matching checkpoint) loads it and skips already-completed steps. Mismatched plan_hash refuses with a paste-ready message. Cleared on clean completion. Pinned by 13 unit cases.
  • src/core/model-config.ts — Model-string resolution (the seam every internal LLM call walks through). v0.31.12: four-tier system (ModelTier = 'utility' | 'reasoning' | 'deep' | 'subagent') with TIER_DEFAULTS (utility→haiku-4-5, reasoning→sonnet-4-6, deep→opus-4-7, subagent→sonnet-4-6) and tier?: ModelTier on ResolveModelOpts. Resolution chain is now 8 steps: cliFlag → deprecated key → config key → models.defaultmodels.tier.<tier> → env var → TIER_DEFAULTS[tier] → caller fallback. Two new exports — isAnthropicProvider(modelString) checks provider:model prefix OR claude- bare-id pattern, and enforceSubagentAnthropic() is the layer-2 runtime guard: when tier === 'subagent' resolves to a non-Anthropic provider, it emits a once-per-(source, model) stderr warn AND falls back to TIER_DEFAULTS.subagent instead of letting the Anthropic Messages API tool-loop attempt to run on OpenAI/Gemini. _resetDeprecationWarningsForTest() now also clears _subagentTierWarningsEmitted so tests re-emit. v0.41.21.0: isAnthropicProvider routes through the new splitProviderModelId from src/core/model-id.ts so slash-form ids (anthropic/claude-sonnet-4-6) classify correctly. Pre-fix the colon-only check silently returned false on slash form, so a user who set models.tier.subagent to the slash form had enforceSubagentAnthropic fall through to TIER_DEFAULTS.subagent AND skipped the warn — the explicit config was honored as if it had never been set. Now both shapes classify the same. Pinned by 2 new cases in test/model-config.serial.test.ts.
  • src/core/ai/model-resolver.ts — Recipe-touchpoint validator. v0.31.12: assertTouchpoint(recipe, touchpoint, modelId, extendedModels?) gains an optional 4th extendedModels: ReadonlySet<string> argument. When the modelId is in that set, the native-recipe allowlist throw is bypassed — the user explicitly opted into this model via config so we let provider rejection surface as model_not_found at HTTP call time (and gbrain models doctor catches it earlier). Default code paths with hardcoded model strings MUST NOT pass extendedModels — typos in source code still fail fast. Replaces the earlier plan to soften the validator wholesale (Codex F4/F5 in plan review flagged that as too broad — it would have removed the fail-fast contract for chat + expand + embed all three).
  • src/core/ai/gateway.ts extension (v0.31.12) — new module-scoped _extendedModels: Map<providerId, Set<modelId>> registry feeds assertTouchpoint's 4th-arg path. New reconfigureGatewayWithEngine(engine) async function is called from cli.ts after engine.connect() (and before every command except CLI_ONLY no-DB commands) — re-resolves expansion + chat defaults through resolveModel() so models.tier.* and models.default overrides apply to expansion + chat both. DEFAULT_CHAT_MODEL corrected to anthropic:claude-sonnet-4-6 (was the v0.31.6 phantom -20250929). New __setChatTransportForTests seam mirrors __setEmbedTransportForTests so tests drive chat() with a stubbed transport.
  • src/core/minions/queue.ts extension (v0.31.12) — MinionQueue.add() now rejects subagent jobs whose data.model resolves through isAnthropicProvider() to a non-Anthropic provider. Lazy-imports model-config.ts to avoid pulling engine types into queue's eager-load surface. Layer 1 of the three-layer subagent provider enforcement (Codex F1+F2 in plan review). Layers 2 + 3 live in src/core/model-config.ts (enforceSubagentAnthropic runtime fallback) and src/commands/doctor.ts (subagent_provider check). Pinned by 3 cases in test/agent-cli.test.ts.
  • src/commands/models.ts (v0.31.12) — gbrain models [--json] read-only routing dashboard: prints tier defaults (utility/reasoning/deep/subagent), the resolved value for each (re-walking the resolution chain to attribute properly), every per-task override (11 PER_TASK_KEYS entries — models.dream.synthesize, models.dream.patterns, models.drift, models.auto_think, models.think, models.subagent, facts.extraction_model, models.eval.longmemeval, models.expansion, models.chat, models.dream.synthesize_verdict), the alias map (defaults + user overrides), and a source-of-truth column showing default / config: <key> / env: <VAR>. gbrain models doctor [--skip=<provider>] [--json] fires a 1-token gateway.chat() probe against each configured chat + expansion model and classifies failures into {model_not_found, auth, rate_limit, network, unknown} — the structural fix for the v0.31.6 silent-no-op bug class. Wired into cli.ts dispatch table + CLI_ONLY set. v0.33.1.1 (#962, Codex P3 follow-up): doctor gains a zero-token embedding_config probe that runs FIRST, before any chat/expansion probes spend money. probeEmbeddingConfig() reads getEmbeddingModel() + getEmbeddingDimensions() from the gateway, parses the model id, and (for Voyage flexible-dim models) checks isValidVoyageOutputDim(dims) against VOYAGE_VALID_OUTPUT_DIMS. New ProbeStatus variant 'config' and optional fix?: string field on ProbeResult — surfaced in both human output (paste-ready gbrain config set ... line under the bad probe) and JSON output. New touchpoint label 'embedding_config' joins 'chat' and 'expansion' in the probe-row taxonomy. Closes the Voyage flexible-dim bug class at config time, not first-embed.
  • src/commands/doctor.ts extension (v0.31.12) — new subagent_provider check (layer 3 of 3 — Codex F13). Warns when models.tier.subagent is explicitly set to a non-Anthropic provider (fail-loud since the user clearly meant it — message names the bad value and prints the paste-ready fix command gbrain config set models.tier.subagent anthropic:claude-sonnet-4-6); also warns when models.default would sneak subagent into a non-Anthropic provider via tier inheritance. OK status when subagent tier resolves to Anthropic. Tests cover all three paths in test/doctor.test.ts.
  • src/core/skill-trigger-index.ts (v0.41.14.0, closes #1451) — Shared loader that unions per-skill SKILL.md frontmatter triggers: with curated RESOLVER.md / AGENTS.md rows from skillsDir AND the parent directory (preserves the OpenClaw workspace-root layout). UNION semantics: explicit RESOLVER.md rows ADD to frontmatter triggers (don't replace). Dedup keyed on (skillPath, trigger.trim().toLowerCase()) so case/whitespace drift collapses to one entry. Three consumers fold through this primitive — checkResolvable, runRoutingEvalCli, mounts-cache.composeResolvers — closing the v0.41 drift bug class (#1451) where per-consumer loaders meant fixing frontmatter for doctor wouldn't reach the routing-eval CLI or cross-brain composed dispatchers. Exports loadSkillTriggerIndex(skillsDir): SkillTriggerEntry[], entriesToResolverContent(entries): string (synthesizes a markdown-table resolver-content string for runRoutingEval's string-content API), findPrimaryResolverPath(skillsDir): string | null, plus the FRONTMATTER_SECTION constant and _resetWarnedSkillsForTests test seam. Skip rules: non-directory entries, _* / .* prefixes, conventions/ + migrations/ subdirs, skills with no SKILL.md (deprecated install/ graceful-skipped), skills with no triggers: array, malformed YAML (warn-once + skip). Reuses parseSkillFrontmatter from src/core/skill-frontmatter.ts (regex-based, not full YAML — adequate for the uniform shape every bundled skill uses, follow-up TODO if drift surfaces). Pinned by 18 hermetic cases in test/skill-trigger-index.test.ts covering frontmatter auto-registration, RESOLVER.md/AGENTS.md merge, OpenClaw ../AGENTS.md parent-dir scan, case-insensitive dedupe, deprecated install skip, missing-skillsDir resilience, malformed-frontmatter warn-once, and entriesToResolverContent round-trip through parseResolverEntries. CI gate: bun run check:resolver (= bun src/cli.ts check-resolvable --strict --skills-dir skills/) wired into bun run verify so future drift between frontmatter and routing-eval fixtures fails PR CI.
  • src/core/skill-catalog.ts (v0.41.36.0, PR1) — host-repo skill catalog backing the MCP list_skills / get_skill ops. Lets a thin MCP client (Codex desktop, Claude Code, Claude Cowork, Perplexity) DISCOVER + FOLLOW the agent repo's fat-markdown skills over gbrain serve — a skill is prose, not code, so "using" one = fetching its body then calling the same gbrain MCP tools the server already exposes. Read-scope, NOT localOnly (defensible only via the full mitigation stack documented in the file's trust-boundary memo). Mitigations: (1) publish gateassertPublishEnabled(ctx, publishSkills); remote callers require mcp.publish_skills === true, default-OFF at runtime so an upgrade never silently grants existing read tokens host-skill read; local callers (ctx.remote === false) always pass. (2) path confinementassertSkillNameShape rejects separators/../null/space before any FS access; the client name is a manifest LOOKUP KEY (via loadOrDeriveManifest), never a raw path segment; confineManifestPath does realpath + relative-containment + SKILL.md-regular-file check on EVERY entry (defeats poisoned manifest.json path, symlink/.. escape). (3) frontmatter allowlistGetSkillResult.frontmatter projects a safe subset; private writes_to (brain taxonomy) + sources (absolute paths) dropped. (4) prose-only + 256KB cap (MAX_SKILL_MD_BYTES, env GBRAIN_MAX_SKILL_MD_BYTES) — no source code (tarballs deferred to PR2), size-checked twice (statSync + UTF-8 byte length). (5) no install_path serve for remote — remote callers use autoDetectSkillsDir (no install-path tier) so a hosted gbrain with no agent repo returns storage_error, never gbrain's own bundled dev skills; local callers use autoDetectSkillsDirReadOnly. (6) MCP rate-limiter caps call rate. Config reads honor BOTH planes: readMcpPublishSkills / readMcpSkillsDir prefer the DB plane (engine.getConfig, what gbrain config set + gbrain init write) over the file plane (ctx.config.mcp). D7 tool-honesty: crossReferenceTools(declared, ctx) splits a skill's declared tools: into usable_tools (in this server's op set AND callable by the caller's scope) vs unavailable_tools, and buildSkillCatalog's instructions envelope (SKILL_CATALOG_INSTRUCTIONS from operations-descriptions.ts) carries the load-bearing "these are prose, follow-then-call-tools" use protocol. Skills are host-filesystem repo-global — sourceScopeOpts(ctx) / ctx.brainId deliberately do NOT apply (one server host = one skills dir = one catalog). buildSkillCatalog is resilient (one malformed/escaping skill is skipped, never throws the call). Config keys land in src/core/config.ts: GBrainConfig.mcp?: { publish_skills?, skills_dir? } + KNOWN_CONFIG_KEYS entries mcp.publish_skills / mcp.publish_skills_prompted / mcp.skills_dir + the mcp. prefix in KNOWN_CONFIG_KEY_PREFIXES. src/commands/init.ts writes config.mcp = { publish_skills: true, ... } for new PGLite + Postgres installs (existing config wins on re-init so a prior opt-out survives). src/commands/upgrade.ts:runPostUpgrade adds a one-time consent prompt (gated by mcp.publish_skills_prompted; existing installs stay OFF until the owner opts in; TTY readline Y/n with non-TTY paste-ready hint). The two ops register in src/core/operations.ts (list_skills with optional section filter + cliHints: { name: 'skills' }; get_skill taking name + cliHints: { name: 'skill', positional: ['name'] }) and dynamically import this module to avoid the import cycle (skill-catalog statically imports the operations array for D7). Descriptions + envelopes in src/core/operations-descriptions.ts (LIST_SKILLS_DESCRIPTION, GET_SKILL_DESCRIPTION, SKILL_CATALOG_INSTRUCTIONS, SKILL_CLIENT_GUIDANCE), pinned by test/operations-descriptions.test.ts. CLI surfaces: gbrain skills / gbrain skill <name> (via cliHints). Pinned by test/skill-catalog.test.ts, test/skill-catalog-security.test.ts (path-confinement / poisoned-manifest / symlink-escape), test/skill-catalog-transports.test.ts (publish-gate + remote-vs-local resolution) over fixtures at test/fixtures/skill-catalog/. PR2 (downloadable/installable skillpack tarballs) is deferred and tracked in TODOS.md.
  • src/core/check-resolvable.ts — Resolver validation: reachability, MECE overlap, DRY checks, structured fix objects. v0.14.1: CROSS_CUTTING_PATTERNS.conventions is an array (notability gate accepts both conventions/quality.md and _brain-filing-rules.md). New extractDelegationTargets() parses > **Convention:**, > **Filing rule:**, and inline backtick references. DRY suppression is proximity-based via DRY_PROXIMITY_LINES = 40. v0.41.7.0: parseResolverEntries accepts a second compact list format alongside the original markdown table (- **skill-name**: trigger1 | trigger2 | trigger3 or - skill-name: trigger1 | trigger2). Both shapes can mix in one file; the v0.31.7 multi-resolver merge folds them into one entry stream. Skill name MUST be kebab-lowercase (regex [a-z][a-z0-9-]+) so prose bullets like - **Note**: …, - **Convention**: …, - **TODO**: … in real-world AGENTS.md files don't false-match as skill rows (codex F2 / D4). skillPath is ALWAYS derived as skills/<name>/SKILL.md: an optional → \skills/path`(or ASCII->) suffix is stripped from the trigger string but NOT honored as the path — two downstream consumers (routing-eval.ts:skillSlugFromPath, the manifest lookup at :367) assume the convention, and honoring the explicit path would silently break their coverage. For non-conventional paths, use the table format. Multi-trigger rows fan out to one entry per trigger sharing the same skillPath; checkResolvablededupes downstream so the reachability count counts each skill once. Closes the OpenClaw bug class where a 306-skill agent's list-format resolver registered 238 FAILs ("skill not reachable from RESOLVER.md") on everygbrain doctorrun — productionized from upstream PR #1370 by @garrytan-agents. Pinned by 11 new unit cases intest/check-resolvable.test.ts(bold + plain forms, Unicode + ASCII suffix strip, ellipsis filter, empty pipe segments, mixed shapes, prose-bullet rejection regression) and the 8-case integration suite intest/check-resolvable-openclaw-compact.test.tsover two fixtures:test/fixtures/openclaw-compact-resolver/(10 skills in pure compact form) +test/fixtures/openclaw-mixed-merge/(table-formatskills/RESOLVER.md+ compact../AGENTS.mdfor the v0.31.7 merge contract). User-facing tutorial:docs/guides/scaling-skills.md` walks through when and why to switch to the compact format (the three-tier scaling architecture that gets a 300-skill agent down to ~4K tokens per turn from ~25K with zero capability loss).
  • src/core/repo-root.ts — Shared findRepoRoot(startDir?) (v0.16.4): walks up from startDir (default process.cwd()) looking for skills/RESOLVER.md. Zero-dependency module imported by both doctor.ts and check-resolvable.ts. Parameterized startDir makes tests hermetic. v0.31.7: read-path / write-path split. autoDetectSkillsDir (shared, read+write-safe) gains tier-0 $GBRAIN_SKILLS_DIR explicit operator override (Docker mounts, CI, monorepo subdirs) ahead of the existing 4-tier chain. New autoDetectSkillsDirReadOnly wraps it with a tier-5 install-path fallback that walks up from fileURLToPath(import.meta.url) and gates on isGbrainRepoRoot so unrelated repos can't false-positive. Read-path callers (doctor, check-resolvable, routing-eval) use the read-only variant; write-path callers (skillpack install, skillify scaffold, post-install-advisory) deliberately stay on the shared function so gbrain skillpack install from ~ cannot silently retarget the bundled gbrain repo's skills/ instead of the user's actual workspace. Two new SkillsDirSource variants: 'env_explicit', 'install_path'. New AUTO_DETECT_HINT_READ_ONLY documents the extra tier. The D6 --fix safety gate in doctor.ts + check-resolvable.ts refuses auto-repair when detected.source === 'install_path' so gbrain doctor --fix from ~ cannot silently rewrite the bundled install tree.
  • src/commands/check-resolvable.ts — Standalone CLI wrapper (v0.16.4) over checkResolvable(). Exports parseFlags, resolveSkillsDir, DEFERRED, runCheckResolvable. Exit rule: 1 on any issue (warnings OR errors), stricter than doctor's ok flag — honors README:259. Stable JSON envelope {ok, skillsDir, report, autoFix, deferred, error, message} — same shape on success and error paths. --fix path runs autoFixDryViolations BEFORE checkResolvable (same ordering as doctor). scripts/skillify-check.ts subprocess-calls gbrain check-resolvable --json (cached per process) and fails loud on binary-missing — no silent false-pass. v0.19: AGENTS.md workspaces now resolve natively (see src/core/resolver-filenames.ts) — gbrain inspects the 107-skill OpenClaw deployment whether the routing file is RESOLVER.md or AGENTS.md. DEFERRED[] is empty — Checks 5 + 6 shipped as real code, not issue URLs. v0.31.7: the resolver lookup switched from first-match-wins to the multi-file merge in src/core/check-resolvable.ts — entries collected from every RESOLVER.md / AGENTS.md across the skills dir AND its parent, deduped by skillPath (first occurrence wins). Lifted reachable skills on the reference OpenClaw layout from 37/224 to 200/224 — the deployment ships a thin skills/RESOLVER.md (~40 entries from skillpack) plus a fat ../AGENTS.md (200+ entries, the real dispatcher), and the previous code only saw the first one. The CLI also switched to autoDetectSkillsDirReadOnly so cd ~ && gbrain check-resolvable finds the bundled skills via the install-path fallback. --fix carries the same D6 safety gate as gbrain doctor --fix: refuses to write when detected.source === 'install_path'.
  • src/core/resolver-filenames.ts (v0.19) — central list of accepted routing filenames (RESOLVER.md, AGENTS.md). Shared by findRepoRoot, check-resolvable, and skillpack install so every code path walks the same fallback chain.
  • src/commands/skillify.ts + src/core/skillify/{generator,templates}.ts (v0.19) — gbrain skillify scaffold <name> creates all stubs for a new skill in one command: SKILL.md, script, tests, routing-eval.jsonl, resolver entry, filing-rules pointer. gbrain skillify check <script> runs the 10-step checklist (LLM evals, routing evals, check-resolvable gate, filing audit) against a candidate skill before it lands.
  • src/commands/skillify-check.ts (v0.19) — gbrain skillpack-check agent-readable health report. Exit 0/1/2 for CI pipeline gating; JSON for debugging. Wraps check-resolvable --json, doctor --json, and migration ledger into one payload so agents can decide whether a human action is required. v0.37.1.0: required item 12 (brain_first_compliance) joins the 11-item checklist. Calls analyzeSkillBrainFirst() on the candidate SKILL.md; exits 1 when the verdict is missing_brain_first (external-lookup pattern present, no callout, no brain_first: exempt). The corresponding scaffold path in src/core/skillify/templates.ts pre-inserts the canonical Convention callout into new SKILL.md files so freshly-scaffolded skills pass item 12 by default.
  • src/commands/book-mirror.ts (v0.25.1) — gbrain book-mirror --chapters-dir <path> --slug <slug> [flags]. Flagship of the v0.25.1 skills wave. Submits N read-only subagent jobs (one per chapter; allowed_tools: ['get_page', 'search']), waits for all via waitForCompletion, reads each child's job.result, assembles two-column markdown CLI-side, writes a single operator-trust put_page to media/books/<slug>-personalized.md. Codex HIGH-1 fix applied: trust narrowing happens at the tool-allowlist layer (subagents can't call put_page) instead of allowedSlugPrefixes — untrusted EPUB content cannot prompt-inject any people page. Cost-estimate prompt before launching; refuses to spend in non-TTY without --yes. Per-chapter idempotency keys (book-mirror:<slug>:ch-<N>) for retry-friendly re-runs. Partial-failure handling: assembles with completed chapters and a ## Failed chapters section listing retries. Test surface: test/book-mirror.test.ts (9 cases — CLI registration + source invariants).
  • src/commands/skillpack.ts + src/core/skillpack/{bundle,scaffold,reference,migrate-fence,scrub-legacy,harvest,harvest-lint,copy,apply-hunks,diff-text,installer}.ts (v0.19 → v0.36) — v0.36 contract change: managed-block install model retired. install and uninstall removed (clean break, no alias; both exit non-zero with a hint pointing at the replacement command). New surface: scaffold (one-time additive copy via shared copyArtifacts helper in copy.ts; refuses to overwrite existing files; partial-state fills missing paired sources declared in SKILL.md frontmatter sources:), reference (read-only diff lens with agent-readable framing line + --apply-clean-hunks two-way auto-apply via pure-JS unified-diff parser/applier in apply-hunks.ts + diff-text.ts), migrate-fence (one-shot strip of legacy fence; cumulative-slugs receipt → row-parsing fallback; preserves rows verbatim as user-owned routing), scrub-legacy-fence-rows (opt-in row cleanup with skill-present + non-empty-triggers gate), harvest (host→gbrain inverse with symlink-reject + canonical-path containment via validateUploadPath-style gate + default-on privacy linter in harvest-lint.ts against ~/.gbrain/harvest-private-patterns.txt plus built-in \bWintermute\b + email + Slack-channel patterns; rollback on match). Paired-source declarations moved from openclaw.plugin.json to each SKILL.md's frontmatter sources: array (D2; validated by loadSkillSources in bundle.ts). autoDetectSkillsDir (in src/core/repo-root.ts) gains a cwd_walk_up tier ahead of ~/.openclaw/workspace (D3; non-OpenClaw hosts like ~/git/wintermute auto-detect; R5 regression preserves $OPENCLAW_WORKSPACE precedence). gbrain skillpack check --strict exits non-zero on drift (CI gate); top-level gbrain skillpack-check keeps exit-1-on-issues for cron compat. Companion editorial skill skills/skillpack-harvest/SKILL.md drives the genericization checklist before the CLI runs. Design + workflow doc: docs/guides/skillpacks-as-scaffolding.md. ~600 LOC of managed-block machinery deleted; ~400 LOC of new modules + ~1000 LOC of new test coverage across test/skillpack-{copy,scaffold,reference,reference-apply,apply-hunks,migrate-fence,scrub-legacy,harvest,harvest-lint,frontmatter-sources}.test.ts + 9-case E2E in test/e2e/skillpack-flow.test.ts. installer.ts + test/skillpack-install.test.ts survive for now — gbrain skillpack diff still uses diffSkill from there; slated for v0.37 cleanup. Historical (v0.19-v0.35.1): managed-block model with <!-- gbrain:skillpack:begin -->/end --> fence, cumulative-slugs="..." receipt, content-hash gates, lockfile; install --all prune; uninstall with D8 receipt gate + D11 atomic-refusal content-hash pre-scan. Replaced wholesale in v0.36.
  • src/core/skillpack/{manifest-v1,tarball,state,remote-source,trust-prompt,bootstrap-display,scaffold-third-party,registry-schema,registry-client,rubric,doctor,init-scaffold,pack-publish,endorse,audit}.ts + examples/skillpack-reference/ + docs/skillpack-anatomy.md + scripts/build-skillpack-anatomy.ts (v0.37.0.0) — third-party skillpack ecosystem layered on the v0.36 scaffolding contract. gbrain skillpack scaffold <owner/repo|https-url|./tgz|./local-dir> resolves the spec via classifySpec, fetches through SSRF-hardened git-remote.ts (git) or extracts the tarball into ~/.gbrain/skillpack-cache/<host>/<owner>/<repo>/<sha>/, validates skillpack.json (api_version gbrain-skillpack-v1), checks gbrain_min_version, surfaces a TOFU first-install identity-confirm prompt (codex G4: author + source + pinned commit + tarball SHA + tier; non-TTY requires --trust), records the pin in machine-owned ~/.gbrain/skillpack-state.json (codex G1; schema gbrain-skillpack-state-v1, atomic .tmp + rename, isAlreadyTrusted skips re-prompt on author+pin match), runs through the existing enumerateScaffoldEntriescopyArtifacts pipeline (one-time additive copy; refuses to overwrite), then DISPLAYS runbooks/bootstrap.md without executing (codex T1 npm-postinstall lesson — "gbrain deliberately does NOT auto-execute these steps"). Registry catalog at garrytan/gbrain-skillpack-registry (separate repo) split into registry.json (PR-able, schema gbrain-registry-v1) + endorsements.json (Garry-only overlay, gbrain-endorsements-v1); effectiveTier merges. registry-client.ts fetches both URLs via If-None-Match etag with 1h soft-TTL + stale-fallback (origins: fresh_fetch | cache_warm | cache_soft_stale | cache_hard_stale); hard-fail only on no-cache + no-network. gbrain skillpack {search,info,registry,doctor,init,pack,endorse} CLI surface. Doctor walks SKILLPACK_RUBRIC_V1 (10 binary dimensions, codex T4 split into 5 required CORE — manifest_valid, skills_have_skill_md, routing_evals_present ≥5 intents, skills_have_unique_triggers MECE, changelog_present_and_current — and 5 quality BADGES — unit_tests_present, e2e_tests_present, llm_eval_present ≥3 cases, bootstrap_runbook_present, license_present); tier eligibility: endorsed needs all 10, community needs core + ≥3 badges, experimental needs core only, blocked when any core fails. --quick ~5s structural sweep; --fix --yes auto-scaffolds auto_fixable: true dimensions and refuses to overwrite files whose mtime is newer than skillpack.json. gbrain skillpack init <name> lands the full cathedral (11 files: skillpack.json, SKILL.md, routing-eval.jsonl, test/example.test.ts, e2e/example.e2e.test.ts, evals/example.judge.json, runbooks/{bootstrap,uninstall,upgrade-template}.md, CHANGELOG, README, LICENSE); freshly-init'd pack scores 10/10. --minimal skips test/e2e/evals. gbrain skillpack pack packs a deterministic tarball via GNU tar (--sort=name --mtime=@0 --owner=0 --group=0 --numeric-owner + GZIP=-n + TZ=UTC); refuses on tier_eligibility === 'blocked'. Extract caps (5000 files / 100MB total / 1MB per file / 255-char paths / 100:1 compression ratio); rejects symlinks/hardlinks/devices/FIFOs. gbrain skillpack endorse <name> [--tier ...] [--push] [--dry-run] runs in a clone of the registry repo: validates pack exists in registry.json, mutates endorsements.json via pure applyEndorsement, stable-key-orders the write, stages + commits endorse: <name> -> <tier>, optionally pushes. JSONL audit at ~/.gbrain/audit/skillpack-YYYY-Www.jsonl (ISO-week rotated, honors GBRAIN_AUDIT_DIR). examples/skillpack-reference/ is a real 10/10 reference pack pinned by test/e2e/skillpack-third-party.test.ts. docs/skillpack-anatomy.md is auto-generated via scripts/build-skillpack-anatomy.ts with --check mode for CI drift detection. CLI dispatch in src/commands/skillpack.ts disambiguates third-party (contains /, ://, .tgz) from bundled-skill kebab; kebab routes bundled-first, registry-fallback. Test coverage: test/skillpack-{manifest-v1,tarball,state,remote-source,trust-prompt,registry-schema,registry-client,rubric,doctor,init-scaffold,pack-publish,endorse,audit,scaffold-third-party}.test.ts + e2e at test/e2e/skillpack-third-party.test.ts. Plan + spec at docs/designs/SKILLPACK_REGISTRY_V1_SPEC.md. Deferred to follow-up waves: subprocess sandbox for publish-gate, garrytan/gbrain-skillpack-registry repo creation + CI workflow split (codex G3), Printing Press cross-list at mvanhorn/printing-press-library, generated gbrain-cli via Printing Press, W4.5 retrofit of ~25 bundled skills to 10/10.
  • src/core/archive-crawler-config.ts (v0.25.1) — D12 + codex HIGH-4 safety gate for the archive-crawler skill. Refuses to run unless archive-crawler.scan_paths: is explicitly set in the brain repo's gbrain.yml. Mirrors the storage-config.ts parsing pattern (sibling file; separate concern from storage tiering). loadArchiveCrawlerConfig(repoPath) throws ArchiveCrawlerConfigError(missing_section | empty_scan_paths | invalid_path | parse_error). normalizeAndValidateArchiveCrawlerConfig rejects relative paths and .. traversal; ~ is expanded; trailing-slash normalized for unambiguous prefix matching. isPathAllowed(candidate, config) is the runtime per-file gate (scan_paths prefix-match with directory-boundary correctness; deny_paths overrides). Tests in test/archive-crawler-config.test.ts (19 cases).
  • test/helpers/cli-pty-runner.ts (v0.25.1) — generic real-PTY harness ported from gstack and trimmed to ~470 lines. Uses pure Bun.spawn({terminal:}) (Bun 1.3.10+; engines.bun pin in package.json). Generic primitives only — no plan-mode orchestrators. Exports: launchPty, resolveBinary, stripAnsi, parseNumberedOptions, optionsSignature, isNumberedOptionListVisible, isTrustDialogVisible. Self-tests in test/cli-pty-runner.test.ts (24 cases).
  • src/core/skill-manifest.ts (v0.19) — parser for skill-manifest.json records. Used by skillpack installer to detect drift between the shipped bundle and the user's local edits, so updates merge instead of overwriting.
  • src/commands/routing-eval.ts + src/core/routing-eval.ts (v0.19) — gbrain routing-eval catches user phrasings that route to the wrong skill. Reads skills/<name>/routing-eval.jsonl fixtures ({intent, expected_skill, ambiguous_with?}). Structural layer runs in check-resolvable by default (zero API cost). The --llm flag is accepted as a placeholder for a future LLM tie-break layer; in v0.24.0 it emits a stderr notice and runs structural only. False positives surface before users hit them. v0.31.7: switched to autoDetectSkillsDirReadOnly and the same multi-file resolver merge as check-resolvable, so on OpenClaw layouts (skills/RESOLVER.md + ../AGENTS.md) all three commands see the same trigger index — previously routing-eval read only the first resolver file it found. The v0.25.1 wave skills' RESOLVER.md rows were also synced to include the full frontmatter triggers: arrays (was only the first trigger), so the structural matcher actually sees the realistic phrasings; ambiguous-fixture annotations cover deliberate skill chains like enrich → article-enrichment.
  • src/core/filing-audit.ts + skills/_brain-filing-rules.json (v0.19) — Check 6 of check-resolvable. Parses new writes_pages: / writes_to: frontmatter on skills and audits their filing claims against the filing-rules JSON. Warning-only in v0.19, upgrades to error in v0.20. v0.37.1.0: internal parseFrontmatter reduced to a thin wrapper over the new shared src/core/skill-frontmatter.ts parser so both filing-audit and skill-brain-first read the same shape (tools?, triggers?, brain_first?: 'exempt', typed brain_first_typo) from one source of truth.
  • src/core/skill-frontmatter.ts (v0.37.1.0) — shared content-based SKILL.md frontmatter parser. Replaces filing-audit's private path-based parser. Recognizes the new brain_first: 'exempt' declarative opt-out and surfaces near-miss declarations (brain-first, BrainFirst, quoted values, unknown values) as a typed brain_first_typo field so doctor can emit a paste-ready hint rather than fail silently. Single canonical form: snake_case brain_first: exempt, lowercase, unquoted.
  • src/core/skill-brain-first.ts (v0.37.1.0) — pure analyzer. analyzeSkillBrainFirst(skillPath, content): SkillBrainFirstResult walks the compliance ladder for every SKILL.md: (1) absent external-lookup pattern → no_external; (2) brain_first: exempt frontmatter → exempt_frontmatter; (3) canonical > **Convention:** see [conventions/brain-first.md](...) callout → compliant_callout; (4) explicit ## Phase 1: Brain heading → compliant_phase; (5) first gbrain search / query / get_page reference precedes first external pattern in the BODY (frontmatter stripped) → compliant_position; (6) else missing_brain_first warn. External pattern set: word-boundary regex over web_search, web_fetch, exa, perplexity, happenstance, crustdata, captain_api, firecrawl. Position scan is BODY-ONLY (A4) so a tools: [web_search] frontmatter declaration doesn't false-flag the skill. PR #1206's 40-name EXEMPT_SKILLS list is preserved as FORMERLY_HARDCODED_EXEMPT purely so doctor can emit a "this used to be auto-exempt, declare brain_first: exempt if still appropriate" hint (CMT1). Consumed by 3 surfaces: doctor check, skillify-check item 12, dry-fix MISSING_RULE_PATTERNS.
  • src/core/skill-fix-gates.ts (v0.37.1.0) — shared safety primitives extracted from dry-fix.ts. getWorkingTreeStatus(file) 3-state ('clean' | 'dirty' | 'not_a_repo'); isInsideCodeFence(content, offset); findAfterH1Paragraph(content) (canonical insertion offset for the auto-inserted Convention callout). Both REPLACE expanders (existing v0.14.1 DRY violations) and the new INSERT expander (MISSING_RULE_PATTERNS) consume from here so the install-path D6 refusal and dirty-tree gates apply uniformly. src/core/dry-fix.ts re-exports for back-compat with existing call sites.
  • src/core/audit-skill-brain-first.ts (v0.37.1.0) — snapshot+diff JSONL audit at ~/.gbrain/audit/skill-brain-first-YYYY-Www.jsonl (ISO-week rotated, honors GBRAIN_AUDIT_DIR via the shared resolveAuditDir()). recordBrainFirstRun(results) reads the previous snapshot at ~/.gbrain/audit/skill-brain-first-snapshot.json, diffs against the current results, writes transition events (detected | resolved | fixed) one line per change, then atomically overwrites the snapshot via .tmp + rename. Transition-only writes — a stable brain produces 0 audit lines per doctor run, so tail -20 shows real signal instead of noise. readRecentBrainFirstEvents(days) is the readback path used by the future skill_brain_first_trend doctor check (filed in TODOS.md). Snapshot file is last-writer-wins under concurrent doctor runs; subsequent runs reconcile.
  • src/core/dry-fix.tsgbrain doctor --fix engine. autoFixDryViolations(fixes, {dryRun}) rewrites inlined rules to > **Convention:** see [path](path). callouts via three shape-aware expanders (bullet / blockquote / paragraph). Five guards: working-tree-dirty (getWorkingTreeStatus() returns 3-state 'clean' | 'dirty' | 'not_a_repo'), no-git-backup, inside-code-fence, already-delegated (40-line proximity, consistent with detector), ambiguous-multi-match, block-is-callout. execFileSync array args (no shell — no injection surface). EOF newline preserved. v0.37.1.0: safety primitives extracted to src/core/skill-fix-gates.ts (back-compat re-exports preserved). New MISSING_RULE_PATTERNS INSERT pattern type lives alongside the existing REPLACE patterns — same auto-fix entry point, same git-safety gates, but instead of rewriting an existing block, INSERT patterns place a canonical callout at a target offset (today: after-h1-paragraph only; designed to extend). The first INSERT pattern is brain_first, which auto-inserts > **Convention:** see [conventions/brain-first.md](../conventions/brain-first.md) for the lookup chain (search → query → get_page → external). on any flagged SKILL.md whose analyzer verdict is missing_brain_first. Idempotent — re-runs detect the existing callout and skip.
  • src/core/backoff.ts — Adaptive load-aware throttling: CPU/memory checks, exponential backoff, active hours multiplier
  • src/core/retry.ts (v0.41.19.0, extended v0.41.27.0) — canonical retry primitive for transient connection errors. Exports withRetry<T>(fn, opts) execution wrapper + BULK_RETRY_OPTS constant ({maxRetries:3, delayMs:1000, delayMaxMs:10000, jitter:'decorrelated'}) tuned for Supabase Supavisor's 5-10s circuit-breaker recovery window + BATCH_AUDIT_SITES typed const (closed enum of every audit-emission site) + resolveBulkRetryOpts(env) (reads GBRAIN_BULK_MAX_RETRIES / GBRAIN_BULK_RETRY_BASE_MS / GBRAIN_BULK_RETRY_MAX_MS with >=0 validation, throws on bad input with paste-ready hint) + abortableSleep(ms, signal?) (AbortSignal-aware setTimeout) + RetryAbortError (tagged error for clean shutdown) + computeNextDelay() (pure-fn for 3 jitter modes: 'none', 'full', 'decorrelated'). The execution wrapper is consumed by postgres-engine.ts + pglite-engine.ts batch primitives (addLinksBatch / addTimelineEntriesBatch / upsertChunks) so every caller — current AND future — inherits retry as part of the data-primitive's contract. CI guard scripts/check-no-double-retry.sh fails the build on withRetry(...engine.batch...) patterns (prevents 3×3=9 retry amplification). CI guard scripts/check-batch-audit-site.sh validates every string-literal auditSite: '...' against the closed BATCH_AUDIT_SITES enum. Closes the v0.41.17 production incident where ~3,000 rows were silently lost per dream cycle on a 16K-page brain because the prior single-500ms-retry shape couldn't survive Supavisor's 5-10s circuit-breaker recovery. Decorrelated jitter (AWS-style: uniform(base, prevDelay*3) capped at delayMaxMs) replaces naive 'full' jitter which allowed near-zero retries that re-hit the still-recovering breaker. Pinned by test/core/retry.test.ts (37 cases) + test/core/retry-stress.slow.test.ts (5 cases simulating 100 batches × 30% blip rate, asserts zero row loss). v0.41.27.0 (#1570): WithRetryOpts gains optional reconnect?: () => Promise<void> callback awaited in the catch branch AFTER isRetryableConnError classification but BEFORE the inter-attempt sleep — lets engine-level callers rebuild a dead pool/singleton between attempts. PostgresEngine.batchRetry injects () => this.reconnect() so the v0.22.1 race-safe _reconnecting guard kicks in transparently. Fail-loud posture (per codex outside-voice finding 3): a reconnect throw PROPAGATES as the new error, replacing the symptomatic "No database connection" so operators see the real cause (auth failure, EHOSTUNREACH). onRetry callbacks now also awaited (back-compat-safe: existing sync arrows work identically; async callbacks now correctly delay the sleep). Pinned by test/core/retry-reconnect.test.ts (5 cases: ordering, back-compat, fail-loud propagation, signal-abort short-circuit, awaited onRetry timing) + test/e2e/db-singleton-shared-recovery.test.ts (3 DB-gated cases pinning shared-singleton survival via retry reconnect + diagnostic audit fires on disconnect + instance-pool disconnect doesn't touch module singleton).
  • src/core/audit/batch-retry-audit.ts (v0.41.19.0, extended v0.41.26.1) — JSONL audit primitive for batch-retry events. Built on audit-writer.ts cathedral. Schema: {ts, site, batch_size, attempt, outcome: 'success' | 'exhausted', delay_ms, error_message_summary, error_code?}. Privacy posture: NEVER logs slugs / page IDs / content (mirrors shell-audit.ts from v0.20+). logBatchRetry fires per successful retry recovery; logBatchExhausted fires when retries exhaust and rows are lost. readRecentBatchRetryEvents(hours=24) returns {events, corrupted_lines, files_scanned, files_unreadable} — corruption + permission errors surface to doctor, not silently swallowed (codex H-9). pruneOldBatchRetryAuditFiles(daysToKeep=30) deletes old files (codex H-8 — implements the v0.41 plan's "30-day pruning convention" for real). Called from gbrain dream --phase purge. File: ~/.gbrain/audit/batch-retry-YYYY-Www.jsonl (honors GBRAIN_AUDIT_DIR). v0.41.26.1 privacy backfill: summarizeError now routes error messages through the shared redactConnectionInfo helper from src/core/audit/redact-connection-info.ts BEFORE truncation, so DSNs / hostnames / credentials / IPv4 octets can't leak from a Postgres connection-failure error into operator-shared JSONL dumps. Same risk class as lock-renewal-audit; closed in the same wave. Pinned by test/audit/batch-retry-audit.test.ts (12 cases) + new test/audit/batch-retry-redaction.test.ts (3 privacy regressions).
  • src/core/audit/lock-renewal-audit.ts (v0.41.26.1) — JSONL audit primitive for per-job lock-renewal faults. Sibling of batch-retry-audit.ts; built on audit-writer.ts. Four outcomes: failure (single renewLock throw, counter incremented), success_after_failure (recovery; emits the recovery count), gave_up (time-based deadline exceeded; abort fired), executeJob_rejected (the SECOND unhandledRejection vector from D7 — the stored executeJob(...).finally(...) promise itself rejected, e.g. failJob threw during the same DB outage). Schema: {ts, job_id, job_name, attempt?, outcome, error_message_summary?, error_code?}. Privacy: NEVER logs lock_token or job.data. Error summaries route through redactConnectionInfo BEFORE truncation. Defense-in-depth: every audit call inside the lock-renewal tick's catch block is wrapped in its own inner try/catch so a misbehaving audit-writer can't re-introduce the unhandledRejection bug class via a new surface. readRecentLockRenewalEvents(hours=24) walks current + previous ISO week with corrupted-line tolerance. pruneOldLockRenewalAuditFiles(daysToKeep=30) is ready for future dream-cycle purge wiring (filed as TODO-LR-3). File: ~/.gbrain/audit/lock-renewal-YYYY-Www.jsonl. Operator UX is tail -F until the doctor check lands (TODO-LR-2). Pinned by test/audit/lock-renewal-audit.test.ts (11 cases).
  • src/core/audit/redact-connection-info.ts (v0.41.26.1) — Shared pure helper. redactConnectionInfo(text: string): string strips Postgres connection info before any audit JSONL write: postgres:///postgresql:// URLs, host=foo, user=foo, password=foo, pwd=foo, IPv4 octets. Each match becomes <REDACTED:kind>. Negative-lookbehind/lookahead [\w.@-] on the IPv4 pattern defeats version-string false positives like v3.1.4.0 or tree-sitter@0.26.3.1 while still matching real IPs in PG errors ((192.168.1.42)). Order-sensitive pattern set: URL forms first so substrings inside URLs don't get double-redacted. Idempotent (running twice produces same output), pure (no I/O), hot-path-safe (regex compiled at module load). Wired into BOTH lock-renewal-audit.ts (new) AND batch-retry-audit.ts (privacy backfill — same risk class). Risk model documented in module header: Postgres errors during connection failures embed DSNs / credentials / hostnames into error messages; operators routinely paste audit dumps into GitHub issues / Slack to debug; the audit channel must be safe to share by construction. Known limitations: bare-quoted hostnames (at "db.example.com") and bare-quoted usernames (for user "postgres.foo") are NOT caught — the IP in those PG error shapes is the highest-value leak and IS caught. Quoted-string patterns filed as TODO-LR-5. Pinned by test/audit/redact-connection-info.test.ts (15 cases covering all 5 patterns + real Supabase fixture + ENOTFOUND fixture + version-string false-positive defense).
  • src/core/minions/lock-renewal-tick.ts (v0.41.26.1) — Pure extracted function from MinionWorker.launchJob's setInterval body. The structural fix that closes the v0.41.22.1 production unhandledRejection crash class. Exports runLockRenewalTick(deps, state) → Promise<TickResult> plus resolveLockRenewalKnobs(env, lockDuration) → LockRenewalKnobs. Three env knobs all parsed as positive integers with stderr-warn-once-per-process on bad input + default fallback: GBRAIN_LOCK_RENEWAL_MAX_FAILURES (default 3, audit-labeling only — abort triggering uses time-based instead), GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS (default lockDuration/3, bounds the hung-renewLock vector via Promise.race), GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS (default lockDuration/6, ensures we release before stall-detector reclaim). The tick checks state.cancelled() at three guard points (entry, after await resolves, after await throws) so a long renewLock that resolves after the job ended bails cleanly. Result type is a tagged union {kind: 'ok' | 'cancelled' | 'lock_lost' | 'should_abort', reason?} that the worker switches on; all abort.abort() / clearInterval side effects live in the worker's closure scope, not the pure function. Audit calls wrapped in inner try/catch (codex C4 defense-in-depth) so the audit module can never re-introduce the bug class through a new surface. Pinned by test/worker-lock-renewal.test.ts (18 hermetic state-machine cases — no fs, no PGLite, no real setInterval). v0.42.5.0 (#1678): LockRenewalDeps gains an optional reconnect? callback; the renewal-tick's catch branch (after the retryable-error classification + the time-based deadline check, before the inter-tick sleep) calls it ONCE, wrapped in its own try/catch so a reconnect throw can't re-introduce the unhandledRejection class. Recovers the mid-process-disconnect case (nulled _sql) where postgres.js's auto-reconnect can't help because the engine's pool reference is gone. Deliberately NOT a background withRetry (Codex #2: a 12s background retry could refresh a lock AFTER the worker already decided lock-renewal-failed and another worker reclaimed it → two holders); the reconnect is bounded by the same callTimeoutMs race + abort signal as renewLock itself.
  • src/core/minions/worker-exit-codes.ts (v0.42.5.0, #1678, NEW) — single source of truth for reserved worker process exit codes, shared by the worker (sets them) and supervisor/CLI (classify them). Exports WORKER_EXIT_RSS_WATCHDOG = 12. Before this, the RSS watchdog drained via gracefulShutdown('watchdog') which set running=false and exited code 0 — indistinguishable from a healthy queue-drain, so the supervisor's code===0 → clean_exit classifier never counted it and a ~400×/24h respawn loop stayed invisible (its only symptom was the downstream DB-connection cascade from the worker being shot mid-cycle). A distinct code makes the drain self-identifying as likely_cause=rss_watchdog. Code 12 is deliberately outside {0 clean, 1 runtime_error} and the 128+N signal range.
  • src/core/minions/rss-default.ts (v0.42.5.0, #1678, NEW) — cgroup-aware auto-sized default for the worker RSS watchdog cap. resolveDefaultMaxRssMb(opts?) / describeDefaultMaxRss(opts?) (provenance for the startup log) / readCgroupMemLimitBytes(readFile?). Replaces the flat ?? 2048 MB default at every spawn site (jobs work, jobs supervisor, autopilot, MinionSupervisor) — 2048 was absurdly low for any brain doing embeddings (working set ~10GB), so the watchdog drained legit work on every heavy cycle. Formula clamp(round(0.5 × basisMB), 4096, 16384) where basis = min(cgroupLimit, totalmem). THE LOAD-BEARING NUANCE (Codex #5): plain os.totalmem() reports HOST RAM, so in a 4GB cgroup on a 126GB host it would pick 16GB, the watchdog would never fire, and the kernel OOM-killer would SIGKILL at 4GB — straight back to the opaque death this fixes. The cap MUST sit below the real ceiling so the graceful drain (distinct exit code, loud log) beats the kernel's silent kill; the 4096 floor is applied only when it stays below the basis. Reads cgroup v2 /sys/fs/cgroup/memory.max (literal max = unlimited) then v1 /sys/fs/cgroup/memory/memory.limit_in_bytes. Explicit --max-rss (including 0 to disable) always wins. Pinned by test/rss-default.test.ts.
  • src/core/cycle/extract-atoms-drain.ts (v0.42.5.0, #1678, NEW) — pure single-hold bounded drain for the silent lens-phase backlog. runExtractAtomsDrain(deps, opts) over injected deps (withLock, runBatch, countRemaining, now, optional onBatch) loops bounded batches under ONE continuous lock hold, rediscovering eligibility each batch (idempotent NOT-EXISTS-on-source_hash, so content mutated by a concurrent process simply doesn't match — Codex #8: no cross-window stale cursor), until the backlog is empty OR the time window elapses. Returns {phase, status, extracted, skipped, remaining, batches, stopped}. Backs gbrain dream --phase extract_atoms --drain. Takes the SAME cycleLockIdFor(sourceId) the routine cycle takes (Codex #9: a concurrent autopilot tick genuinely defers with cycle_already_running); NO release/reacquire-between-windows primitive (Codex #10). Pinned by test/extract-atoms-drain.test.ts.
  • scripts/check-worker-lock-renewal-shape.sh (v0.41.26.1, D4) — CI guard wired into bun run verify. Two invariants on src/core/minions/worker.ts: (1) the v0.41.22.1 bug pattern lockTimer = setInterval(async ...) must NOT appear (narrowed via lockTimer = prefix so unrelated setInterval(async) calls elsewhere in the file — like the stall detector at line ~269, codex C13 territory — don't false-fire), (2) runLockRenewalTick must remain referenced so the pure-function test seam survives refactors. Bug-pattern-specific by design (codex C12) — a future refactor to setTimeout-recursion or AbortController-based scheduling passes as long as the bug pattern stays absent. Uses POSIX ERE + [[:space:]] for BSD-grep portability. Honors GBRAIN_LOCK_RENEWAL_SHAPE_TARGET env override for fixture-based meta-tests. Pinned by test/scripts/check-worker-lock-renewal-shape.test.ts (5 cases including the C13 false-positive defense for the stall-detector pattern elsewhere in the file).
  • src/commands/doctor.ts:checkBatchRetryHealth (v0.41.19.0, extended v0.41.27.0) — batch_retry_health check surfacing Supavisor circuit-breaker incidents. Wired into both runDoctor (local) and doctorReportRemote (thin-client). Reads last 24h (not 7d — codex H-9: avoid permanent noise from one historical blip). States: ok (zero exhausted in 24h OR <3 from a single site), warn (>=3 same-site OR >=5 cross-site), fail (>=20 sustained breaker). Surfaces bad GBRAIN_BULK_* env at doctor startup (codex M-10) instead of first-retry. Corrupt-JSONL tolerant. Paste-ready fix hints in every warn/fail message. Pinned by test/doctor-batch-retry.test.ts (10 cases). v0.41.27.0 (#1570, codex finding 11): extended (NOT a new check) to read readRecentDbDisconnects(24) and append Disconnect-call audit: N call(s) in 24h (most recent caller: <frame>). to ALL three message paths (ok / warn / fail). The single-place surface keeps connection-incident signal greppable from one gbrain doctor --json call so operators correlating the v0.41.27 retry-reconnect symptom fix with the offending caller don't have to grep two audit files. Module-import wrapped in try/catch so older brains without the v0.41.27 audit file degrade silently.
  • src/core/audit/db-disconnect-audit.ts (v0.41.27.0, #1570, NEW) — JSONL audit for every call to db.disconnect() and PostgresEngine.disconnect(). Built on audit-writer.ts cathedral (no parallel subsystem; codex finding 11). Schema: {ts, engine_kind: 'postgres'|'pglite'|'unknown', connection_style: 'module'|'instance'|'unknown', caller_stack, command, pid}. caller_stack captured via new Error().stack truncated to ~20 frames so operators identify the offending caller without inflating JSONL forever. Privacy: stack frames carry file paths but NO SQL content / row data / user strings — matches the v0.20 shell-audit posture. File: ~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl (honors GBRAIN_AUDIT_DIR). readRecentDbDisconnects(hours=24) walks current + previous ISO week and returns {count, most_recent_caller, files_scanned}. The instrument half of the "instrument first, fix later" pivot — v0.41.27.0 ships the symptom fix (retry reconnect callback + facts:absorb drain) AND this diagnostic surface; v0.41.28+ patches the specific ownership boundary once production data tells us which code path is calling disconnect() mid-process. Wired into src/core/db.ts:disconnect and src/core/postgres-engine.ts:disconnect (logs BEFORE the early-return branches so even no-op disconnects on never-connected engines are recorded — that case may itself be a caller-side bug). Pinned by test/db-disconnect-audit.test.ts (6 cases: round-trip, stack truncation, sort order, empty-dir nulls, stable feature name, EROFS best-effort).
  • src/core/facts/queue.ts:FactsQueue.drainPending (v0.41.27.0, #1570) — new method drainPending({timeout?: number}): Promise<{drained, unfinished}>. Semantically distinct from shutdown() (which calls this.internalAbort.abort() and would abort the very facts:absorb worker trying to log its post-completion event — preserving the bug class we're fixing per codex finding 9). Drain lets in-flight finish; only the wait is bounded. Default timeout 1000ms so commands that don't enqueue facts pay only one fast 0ms check before exit (per codex finding 10). src/cli.ts op-dispatch finally block awaits getFactsQueue().drainPending({timeout: 1000}) BEFORE engine.disconnect(). Lazy-import keeps the facts-queue module off the hot path for ops that never touch it. Closes the trailing 'No database connection' line that fired on stderr after gbrain capture because the post-page-write facts:absorb queued work outlived the CLI process. Pinned by test/facts-queue-drain-pending.test.ts (4 cases: empty fast-path, in-flight settled without abort, unfinished count on timeout, default timeout = 1000ms).
  • scripts/check-no-double-retry.sh + scripts/check-batch-audit-site.sh (v0.41.19.0) — CI lint guards wired into bun run verify. The former greps src/ for withRetry(...engine.{addLinksBatch|addTimelineEntriesBatch|upsertChunks}) patterns and fails the build on hit (Eng-D6 migration ordering hazard — prevents 3×3=9 retry amplification on incomplete reverts). The latter extracts every string-literal auditSite: '...' from src/ and validates each appears in the BATCH_AUDIT_SITES const in src/core/retry.ts (codex H-7 typo guard — prevents fragmented doctor output).
  • src/core/fail-improve.ts — Deterministic-first, LLM-fallback loop with JSONL failure logging and auto-test generation
  • src/core/transcription.ts — Audio transcription: Groq Whisper (default), OpenAI fallback, ffmpeg segmentation for >25MB
  • src/core/enrichment-service.ts — Global enrichment service: entity slug generation, tier auto-escalation, batch throttling
  • src/commands/enrich.ts + src/core/enrich/thin.ts + src/core/cycle/enrich-thin.ts (v0.42.6.0, #1700) — gbrain enrich --thin: batch-develops stub (thin) pages via brain-internal grounded synthesis. gbrain's own model tooling can only see brain-internal context (search / get_page / facts / backlinks), not the web, so enrich consolidates what the brain ALREADY knows about an entity (scattered across meetings, other pages, deals, facts) into one cited page via ONE gateway.chat call per page; web research stays the agent-driven enrich SKILL's job. runEnrichCore(engine, opts, signal) (strict per-source; multi-source iteration is the caller's job) drives enrichOne per candidate: withRefreshingLock('enrich:<src>:<slug>')getPage → deterministic retrieve (hybridSearch + getBacklinks + facts + raw_data, source-scoped, sanitized via INJECTION_PATTERNS) → assessGrounding gate (skip < MIN_CONTEXT_CHARS, no LLM) → buildEnrichPrompt (grounded dossier, [Source: slug] citations, SKIP sentinel) → synth → put_page handler (remote:false, auto-link + write-through) stamping enriched_at + enriched_by:'cli:enrich'. Candidate selection is the SQL-native engine.listEnrichCandidates(opts) (src/core/engine.ts interface + EnrichCandidate/EnrichCandidatesOpts/ENRICH_ORDER_SQL in src/core/types.ts + pg/pglite impls): thin-filter + per-page source-correct inbound count (to_page_id = p.id, mentions excluded) + enriched_at recency guard + whitelisted ORDER BY + LIMIT, lightweight projection (NO bodies). Resume via src/core/op-checkpoint.ts (local enrichFingerprint); budget via BudgetTracker + withBudgetTracker (best-effort under --workers > 1 per D5 — runSlidingPool aborts new claims on BUDGET_EXHAUSTED but does NOT cancel in-flight gateway.chat; pin --workers 1 for a hard ceiling). codex-review fixes folded in: sanitizeContext (thin.ts) neutralizes the <context>…</context> data-envelope delimiters (P1 injection escape, mirrors the </trajectory> convention); the --background multi-source fan-out idempotency key carries the run fingerprint via exported backgroundIdempotencyKey(sid, args) (P1, was a bare enrich:${sid} that returned stale completed jobs); runEnrichCore flags budget_exhausted post-hoc when tracker.totalSpent > tracker.cap even when the gateway swallowed the final-call throw (P1, via new read-only BudgetTracker.cap getter — no gateway.ts change); body() flushes the checkpoint on BudgetExhausted before it propagates so resume doesn't re-charge (P2). The opt-in enrich_thin cycle phase (default OFF via cycle.enrich_thin.enabled) trickles max_pages_per_tick (default 3) per source with per-source cost cap enforced as min(per_source_cap, brain_wide_remaining) (P2 fix) + brain-wide total + walltime caps. Wired into cycle.ts (CyclePhase/ALL_PHASES between conversation_facts_backfill and skillopt/embed; PHASE_SCOPE='source'; NEEDS_LOCK; dispatch), cli.ts (CLI_ONLY + CLI_ONLY_SELF_HELP + THIN_CLIENT_REFUSED_COMMANDS + dispatch), jobs.ts (Minion enrich handler, strict per-source, NOT in PROTECTED_JOB_NAMES). DI seam opts.synthesizeFn keeps tests hermetic (no API key, no mock.module). Pinned by test/enrich/thin.test.ts, test/enrich/idempotency.test.ts, test/enrich-cycle-phase.test.ts, test/e2e/enrich-pglite.test.ts (grew-cited, skip, ordering, multi-source, recency, resume, budget abort + checkpoint flush, final-call overage, lock-skip, provenance), test/e2e/engine-parity.test.ts (listEnrichCandidates pg↔pglite parity). Plan: ~/.claude/plans/pick-up-this-context-composed-neumann.md.
  • src/core/data-research.ts — Recipe validation, field extraction (MRR/ARR regex), dedup, tracker parsing, HTML stripping
  • src/commands/embed.tsgbrain embed [--stale|--all] [--slugs ...]. v0.22.1 (#409, contributed by @atrevino47): --stale path now starts with engine.countStaleChunks() (single SELECT count(*) WHERE embedding IS NULL, ~50 bytes wire). On a fully-embedded brain that's a 1-line short-circuit — no further reads. When stale chunks exist, engine.listStaleChunks() returns just the chunks needing embeddings (slug + chunk_index + chunk_text + metadata, no vector(1536) payload). Caller groups by slug, embeds via OpenAI, re-upserts via upsertChunks. Replaces the prior page-walk that pulled every chunk's embedding column over the wire and discarded most. v0.40.3.0: all 16 console.log / console.error call sites migrated to slog / serr from src/core/console-prefix.ts so that when runEmbedCore runs inside a per-source withSourcePrefix scope (which the gbrain sync --all worker pool installs), every line carries the [<source-id>] prefix. Standalone gbrain embed --stale callers see identical output to v0.40.2 because slog/serr fall through to bare console fns outside the wrap. v0.41.31.0: every embed-write path now stamps pages.embedding_signature via engine.setPageEmbeddingSignature(slug, {sourceId, signature: currentEmbeddingSignature()}) so a later model/dims swap is detectable as stale. The per-slug path (embedPage, used by gbrain embed <slug> AND sync's post-import embed step) and the full-re-embed path (embedAll) stamp unconditionally per page. The stale path (embedAllStale) first calls invalidateStaleSignatureEmbeddings on a live run so signature-drifted pages flow through the NULL cursor, then stamps each page — but ONLY when EVERY chunk was stale this pass (a partially-stale page keeps preserved chunks of unknown provenance, so it stays unstamped rather than falsely marked current; embed --all fully re-embeds + stamps those). dry-run never mutates: it counts signature-drift via the widened countStaleChunks({signature}) predicate without NULLing anything.
  • src/core/conversation-parser/ (v0.41.16.0, v0.41.24.0, v0.41.29.0) — 14-pattern built-in chat-format registry + opt-in LLM polish/fallback. Replaces the single inline regex that PR #1461 was trying to extend with one more shape. v0.41.29.0 added the 14th pattern bold-name-no-time (regex /^\*\*(?!\[)(.+?):\*\*\s*(.*)$/, declared at index 3 after bold-paren-time): parses **Speaker:** text with NO per-line timestamp — the shape Circleback / Granola / Zoom emit. Anchors every message at T00:00:00Z of the frontmatter date (no fabricated wall-clock; line order preserves sequence), same no-time convention as irc-classic. The non-shadow guarantee is the colon-INSIDE-bold regex (NOT declaration order — parse.ts scores every candidate independently, order is only the tie-break); the (?!\[) lookahead rejects telegram-bracket **[18:37] Name:** so disabling telegram-bracket yields an honest no_match instead of speaker="[18:37] Name". Because **Label:** text is a common prose idiom, the pattern sets the new optional PatternEntry.score_full_body: true: parse.ts recomputes the winner's acceptance score over the FULL body (not just the 10-line head window) before the SCORING_MIN_ACCEPTANCE floor, so a notes page with bold labels clustered in its first 10 lines (head score 0.3, which is NOT < SCORING_HEAD_TRIGGER_THRESHOLD so the head fallback never fires) drops to its true low density and stays no_match instead of mis-parsing as a conversation. The same wave scrubbed the pre-existing real names in bold-paren-time's test_positive. Pinned by the bold-name-no-time + clustered-head no_match cases in test/conversation-parser/parse.test.ts + bold-name-no-time.jsonl / adversarial.jsonl fixtures. v0.41.24.0 (closes #1533) added the 13th pattern bold-paren-time (**Speaker** (HH:MM): text and **Speaker** (HH:MM:SS): text — the time-only shape an OpenClaw meeting-ingestion pipeline produces from Circleback transcripts; date_source: frontmatter) AND tightened the parser's fallback gates in parse.ts: SCORING_HEAD_TRIGGER_THRESHOLD = 0.3 triggers a full-body re-score when the head pass scores below that (closes the bug class where a stray head match at 0.1 — e.g. blockquote that accidentally matches an unrelated pattern — suppressed the fallback entirely); SCORING_MIN_ACCEPTANCE = 0.05 blocks essay false-positives (a 300-line essay with one stray chat-shape line scores ~0.003, below the floor, stays no_match instead of returning a 1-message regex_match). Both gates pinned by 11 new unit cases in test/conversation-parser/parse.test.ts. New exported scorePatternFull(body, entry) for direct full-body scoring; private getNonBlankLines(body, headCap?) + scoreFromLines(lines, entry) helpers DRY the quick_reject+regex loop. Empirical against all Circleback meetings in ~/git/brain/meetings/: 113 of 367 parse correctly post-fix (was 0); 20,167 messages flow through to the fact extractor (was 0). The other 254 are notes-only meetings whose transcripts live in separate files. Plan + Codex absorption at ~/.claude/plans/system-instruction-you-are-working-starry-frost.md. Modules: types.ts (PatternEntry + ParseResult + DateContext + CaptureMap + TimezonePolicy), builtins.ts (13 hand-vetted patterns sourced from public format docs — iMessage/Slack, Telegram bracket + text-export, bold-paren-time, Discord classic + export, WhatsApp ISO + US, Signal, Matrix/Element, IRC classic + weechat, Teams export; module-load validation runs every test_positive[] + test_negative[] sample at startup, so a typo in any built-in regex makes gbrain refuse to start; DEFAULT_SPEAKER_CLEAN exported as the PR #1461 helper promoted to a module-level default), parse.ts (orchestrator with D18 pattern-priority scoring across the first 10 lines + D8 date derivation chain explicit > frontmatter.date > effective_date > '1970-01-01' + D5 multi-line continuation + D19 timezone warning), llm-base.ts (shared runLlmCall<T> with content-hash cache in-process + DB-persistent via migration v97 + 4-strategy JSON repair + Anthropic-key probe; both polish and fallback are thin wrappers per D6 DRY), llm-polish.ts (per D15 opt-IN; headroom guard skips when tracker within $0.10 of cap; pure applyPolish for merge/drop/edit ops), llm-fallback.ts (per D15 opt-IN; per D17 NO regex inference + NO persistence — silent corruption avoided), eval.ts (scoreFixture + aggregateScores + parseFixtureJsonl for the fixture-corpus CI gate), nightly-probe.ts (DI-stubbed; mode-gated default per D10 — tokenmax=ON, conservative/balanced opt-in; adversarial false-positive detection). CLI surfaces at src/commands/eval-conversation-parser.ts (gbrain eval conversation-parser <fixture.jsonl> exit 0/1/2 — wired into bun run verify via check:conversation-parser) and src/commands/conversation-parser.ts (scan <slug> for debug, list-builtins for operator discoverability, validate <file> emits "deferred to v0.42+" notice). Doctor checks: conversation_format_coverage (per-pattern hit count + unmatched %), progressive_batch_audit_health (abort-verdict counts from last 7d), conversation_parser_probe_health (opt-in nightly probe state). Closes PR #1461 — contributor's BRACKET_TIME_RX + cleanSpeaker survive verbatim as the telegram-bracket built-in + DEFAULT_SPEAKER_CLEAN export, Co-Authored-By preserved. Plan + 23 decisions + codex outside-voice absorption at ~/.claude/plans/system-instruction-you-are-working-cuddly-hollerith.md. Pinned by 74 unit cases across test/conversation-parser/{parse,llm-base,llm-fallback,llm-polish,nightly-probe}.test.ts + the 27-case PR #1461 baseline at test/extract-conversation-facts.test.ts (back-compat invariant). Migration v97 (conversation_parser_llm_cache_table). Fixtures at test/fixtures/conversation-formats/{imessage,telegram-bracket,whatsapp-iso,whatsapp-us,signal-export,irc-classic,irc-weechat,matrix-element,teams-export,all,adversarial}.jsonl with scripts/check-fixture-privacy.sh banning real-name leaks.
  • src/core/progressive-batch/ (v0.41.16.0) — shared ramp-up + cost-cap + verification primitive extracted after rule-of-three was comfortably past (12+ ad-hoc cost-prompt patterns scattered across reindex, reindex-multimodal, book-mirror, brainstorm, eval-suspected-contradictions, post-upgrade-reembed, eval-contradictions/cost-prompt). Wintermute-inspired ramp shape (trial 10 → ramp 100 → ramp 500 → full with verification at each stage), productionized as a runtime primitive with verifier+policy injection (callers describe HOW TO MEASURE SUCCESS, not WHEN TO WAIT FOR CTRL-C). Modules: types.ts (Stage, StageVerdict, AbortReason, discriminated Verifier union per D20: OutputCountVerifier | IdempotentMutationVerifier | NoopVerifier, Policy, StageReport), orchestrator.ts (runProgressiveBatch(items, verifier, policy, runner) — reads getCurrentBudgetTracker() ahead of Policy.maxCostUsd per D3 fail-closed; null both ways triggers abort_cost_cap reason='no_budget_safety_net'), audit.ts (ISO-week JSONL at ~/.gbrain/audit/progressive-batch-YYYY-Www.jsonl via the shared audit-writer primitive), stage-report.ts (ASCII stage-report formatter for the default Policy.onStageReport handler). Env knobs: GBRAIN_PROGRESSIVE_BATCH_DISABLED=1 (skip ramp + stderr warn), GBRAIN_PROGRESSIVE_BATCH_AUTO=1 (skip Ctrl-C grace), GBRAIN_PROGRESSIVE_BATCH_STAGES=10,100,500 (override). v0.41.16.0 ships with ONE proven consumer (the parser cathedral via extract-conversation-facts retrofit); the 9-site retrofit of reindex.ts, reindex-multimodal.ts, reindex-code.ts, post-upgrade-reembed.ts, book-mirror.ts, brainstorm/orchestrator.ts, eval-suspected-contradictions.ts, eval-contradictions/cost-prompt.ts, contextual-reindex-per-chunk.ts lands as the v0.41.14.0+ follow-up wave per D2's bisectability requirement (each retrofit gets its own commit; behavior parity tested before/after migration). Per D21: sites that previously "jumped straight to full" keep doing so by default; ramp is opt-in per-site via Policy.interactiveAbortMs > 0. Pinned by 35 unit cases in test/progressive-batch/orchestrator.test.ts covering every verdict path.
  • src/commands/extract-conversation-facts.ts + src/core/cycle/conversation-facts-backfill.ts (v0.41.11.0) — bulk fact extraction for long-form conversation pages. The recall-miss class on long iMessage/Slack imports (60K+ msg history; a chunk that reads only "Locker 93 code 9494" has no topical anchor because "cabin" was established 50K messages earlier) gets fixed by walking conversation/meeting/slack/email pages, splitting into time-windowed segments (30-min gap or 30-msg cap), prepending a topical/temporal header ("Conversation between A and B from to "), and running through the existing extractFactsFromTurn() so the resulting anchor-rich facts surface in gbrain search (which already blends facts). Architecture decisions absorbed from 2 Codex outside-voice rounds + 2 eng review passes: strict per-source core (runExtractConversationFactsCore({sourceId, ...}) always takes one sourceId; CLI + cycle phase each do their own multi-source iteration because PHASE_SCOPE='source' is taxonomy-only per cycle.ts:131); two-phase enumeration (paginated listPages({type, sourceId, limit:10}) so worst case is 10×25MB=250MB resident vs PR #1406's unbounded; per-page body cap MAX_PAGE_BODY_BYTES=25MB with pages_skipped_too_large counter surfaced in doctor); page-global row_num accumulator (facts unique index is (source_id, source_markdown_slug, row_num) per migration v51 — per-segment row_num would collide on segment 2); page-level TERMINAL audit row written to facts table after all segments commit (source='cli:extract-conversation-facts:terminal'; doctor's NOT EXISTS query matches the terminal row, NOT any-fact-for-slug, so partial-extraction pages stay in backlog); optional opts.budgetTracker? (when present, core uses it as-is — nested withBudgetTracker REPLACES per gateway.ts:2011; when absent, core auto-wraps with BudgetTracker({maxCostUsd})); body read covers compiled_truth + timeline (iMessage importers put history in timeline half; PR's compiled_truth ?? '' would silently drop it); honors facts.extraction_enabled kill-switch with --override-disabled escape; --types LIST allowlist (conversation,meeting,slack,email) with CLI default reading cycle.conversation_facts_backfill.types config as single source of truth; fingerprint on sourceId only so widening types config doesn't invalidate completed-page state; string-encoded op-checkpoint entries "<sourceId>|<slug>|<endIso>" for resume only (op_checkpoints stores string[] and is 7-day GC'd — durable audit is the facts table itself via terminal row); --background via maybeBackground (Minion handler extract-conversation-facts re-creates BudgetTracker from data.max_cost_usd; on BudgetExhausted mid-job, catches + persists + marks completed with result.budget_exhausted=true — NOT a failure). The companion cycle phase conversation_facts_backfill (default OFF; gbrain config set cycle.conversation_facts_backfill.enabled true opts in) iterates listSources(engine) directly, creates ONE brain-wide tracker per tick + wraps the loop in withBudgetTracker + passes the tracker into every per-source invocation via opts.budgetTracker so the core doesn't nest-replace. Two-layer cost AND walltime protection: per-source caps (max_cost_usd=$1, max_walltime_min=20) AND brain-wide caps (max_total_cost_usd=$5, max_total_walltime_min=30). Pinned by 27 unit cases in test/extract-conversation-facts.test.ts (parse/segment/render/checkpoint/fingerprint/terminal-audit/row_num/F2-kill-switch). Schema migration v94 in src/core/migrate.ts adds a partial index idx_facts_extract_conversation_session ON facts(source_id, source_session) WHERE source LIKE 'cli:extract-conversation-facts%' (v14 precedent: transaction:false + invalid-index pre-drop on Postgres; plain CREATE INDEX on PGLite) so the doctor backlog query stays fast on million-fact brains. src/commands/doctor.ts:computeConversationFactsBacklogCheck is 3-state (SKIPPED when cycle.conversation_facts_backfill.enabled=false — no opt-out noise debt per Eng-v2 C9; OK when caught up; WARN when >10 pages lack the terminal audit row, with paste-ready remediation step composable via gbrain doctor --remediate). src/commands/sources.ts:runAudit extended with facts_backfill_estimate: {pages, est_segments, est_cost_usd, types} so gbrain sources audit <id> previews cost before any run. Schema-pack (v0.41.11.0): gbrain-base.yaml promotes conversation (temporal, extractable) + atom (annotation, NOT extractable — atoms ARE the extracted form) into the base seed; concept.extractable flips true semantically (cosmetic on backstop path because backstop uses hardcoded ELIGIBLE_TYPES in src/core/facts/eligibility.ts:51, not pack extractable — the original concept-grandfather migration was solving a phantom, dropped per Codex T3). ALL_PAGE_TYPES in src/core/types.ts extended with the two new types. Plan + all decisions persisted at ~/.claude/plans/system-instruction-you-are-working-linear-unicorn.md and the CEO plan at ~/.gstack/projects/garrytan-gbrain/ceo-plans/2026-05-25-conversation-retrieval-upgrade.md.
  • src/commands/extract.tsgbrain extract links|timeline|all [--source fs|db] [--source-id <id>]: batch link/timeline extraction. fs walks markdown files, db walks pages from the engine (mutation-immune snapshot iteration; use this for live brains with no local checkout). As of v0.12.1 there is no in-memory dedup pre-load — candidates are buffered 100 at a time and flushed via addLinksBatch / addTimelineEntriesBatch; ON CONFLICT DO NOTHING enforces uniqueness at the DB layer, and the created counter returns real rows inserted (truthful on re-runs). v0.22.1 (#417): ExtractOpts.slugs?: string[] enables incremental extract — when set, extractForSlugs() reads ONLY those slugs' files (single combined links+timeline pass) instead of the full directory walk. CLI gbrain extract keeps full-walk behavior; the cycle path threads sync's pagesAffected through. walkMarkdownFiles(brainDir) still runs at line 455 to build allSlugs for link resolution — see TODOS.md for replacing it with engine.getAllSlugs(). v0.37.7.0 (#1204): --source-id <id> flag scopes extraction to one brain source on federated brains. Resolved via resolveSourceWithTier() before any SQL runs; failures surface with a gbrain sources list hint. Closes the silent-collapse-to-default bug class for extract.
  • src/core/extract/receipt-writer.ts + src/core/extract/rollup-writer.ts + src/commands/extract-status.ts + src/commands/extract-explain.ts + src/commands/extract-benchmark.ts + src/core/schema-pack/scaffold-extractable.ts (v0.41.23.0) — unified extract operator-surface wave. Every shipped extractor (deterministic facts.conversation in src/commands/extract-conversation-facts.ts + three LLM-backed cycle phases at src/core/cycle/{extract-atoms,synthesize-concepts,propose-takes,extract-facts}.ts) now writes ONE receipt page per run (writeReceipt) + UPSERTs a row to extract_rollup_7d (upsertExtractRollup). Receipt slug per D-EXTRACT-17: extracts/{date}/{kind}/{source_id}/{run_id_short}/round-{N}.md; frontmatter stamps BOTH type: extract_receipt AND dream_generated: true per D-EXTRACT-19 (belt+suspenders against extraction-loop guard drift). extract_receipt joins ALL_PAGE_TYPES in src/core/types.ts; extracts/ prefix gets a 0.3x source-boost demote in src/core/search/source-boost.ts so receipts surface in search but never dominate user content. Migration v104 adds extract_rollup_7d (kind, source_id, day, cost_usd, halt_count, eval_pass_count, eval_fail_count, round_completed_count, rollup_write_failures, updated_at) with PK (kind, source_id, day) + idx_extract_rollup_7d_day. Rollup writes are best-effort with process-scoped error-dedup so transient DB failures bump rollup_write_failures instead of crashing the cycle. New extract_health doctor check reads the rollup for last 7 days, warns at halt-rate > 10% AND when rollup_write_failures > 0; pre-v104 brains report ok silently. Operator CLI surface: gbrain extract status [--source-id ID] [--kind X] [--verbose] [--json] (7-day rollup, sorted by halt_rate desc + cost desc, kubectl-style top-5 + "more rows" hint, stable schema_version: 1 JSON envelope); gbrain extract --explain <kind> (resolution chain: pack-declared vs built-in cycle phase, prompt_template + fixture_corpus paths with /(missing) markers, last 7d rollup); gbrain extract benchmark --pack X --kind Y (loads pack fixture corpus through strict D-EXTRACT-21 path validation — rejects absolute paths, .. traversal, null bytes, AND symlinks resolving outside pack root; v0.41.23 ships as a stub-reporter — LLM dispatch deferred to a follow-up release). Pack-author authoring loop: src/core/schema-pack/manifest-v1.ts widens extractable from z.boolean() to z.union([z.boolean(), ExtractableSpecSchema]) (carries prompt_template, fixture_corpus, eval_dimensions, benchmark_min_recall, plus reserved verifier_path per D-EXTRACT-37 — parses but refuses at runtime in v0.41.23); new extractableSpecsFromPack + getExtractableSpec + refuseVerifierPathInV042 helpers in src/core/schema-pack/extractable.ts; gbrain schema scaffold-extractable <type> --pack <pack> declares the type extractable in one verb, generates 5 placeholder fixtures + a prompt template stub under packs/<pack>/{fixtures,prompts}/extract/, refuses to overwrite without --force. Tests: 22 cases in test/extractable-spec-widening.test.ts (back-compat boolean shape + struct shape + D-EXTRACT-37 verifier_path refuse), 12 cases in test/extract/receipt-writer.test.ts (uses canonical PGLite block per CLAUDE.md R3+R4 — one engine per file with resetPgliteState in beforeEach), 17 cases in test/extract/benchmark.test.ts (path validation rejections + JSONL contract), 15 cases in test/extract/status.test.ts (pure aggregation + formatting), 15 cases in test/schema-pack/scaffold-extractable.test.ts (scaffold mechanics + explicit privacy-rule guards against real-name leakage), 8 cases in test/doctor-extract-health.test.ts. Plan persisted at ~/.claude/plans/system-instruction-you-are-working-stateless-dragonfly.md. Deferred to a follow-up release: replay versioning + gbrain extract replay --since v<sha>; unified LLM dispatcher (deliberately respects v0.41.13 head comment that extraction is regex + cost-cap value-add lives at embed step).
  • src/commands/import.tsgbrain import <path> [--source-id <id>]: page import with the v0.34.2.0 path-set checkpoint described above. v0.37.7.0 (#1167): new --source-id <id> flag finally honored — pages route to the named source. Resolved via resolveSourceWithTier() at the boundary; the same flag is now consistent across import, extract, graph-query, and sources current. Pinned by test/import-source-id.test.ts.
  • src/commands/graph-query.tsgbrain graph-query <slug> [--type T] [--depth N] [--direction in|out|both] [--include-foreign]: typed-edge relationship traversal (renders indented tree). v0.37.7.0 (#1153): foreign-edge footer always present (X foreign edges (use --include-foreign to traverse)) so cross-source edges never disappear silently; --include-foreign widens the SQL filter to walk them. Pinned by test/graph-query.test.ts.
  • src/commands/sources.tsgbrain sources {list,add,remove,archive,restore,archived,purge,current,status,audit}. v0.37.7.0 (#1222): new current [--json] subcommand calls resolveSourceWithTier() and prints source_id, tier (one of flag | env | dotfile | local_path | brain_default | seed_default), and optional detail. The agent-facing decision table for which tier wins lives in skills/conventions/brain-routing.md. v0.40.3.0 (productionized from PR #1314): new status [--json] subcommand — read-only per-source dashboard (last sync, staleness, page count, embedding coverage, unacked failures). Thin wrapper around buildSyncStatusReport + printSyncStatusReport exported from src/commands/sync.ts. --json emits the stable {schema_version: 1, sources, ...} envelope on stdout for monitoring pipelines; bare invocation prints the human table to stdout (right-aligned numeric columns, kubectl-style). Filters input sources to local_path IS NOT NULL AND archived IS NOT TRUE so archived sources (which have their own gbrain sources archived surface) don't muddy the active-sync dashboard. Lives under sources (not sync --status) per D3 from the v0.40.3.0 plan-eng-review — reads and writes don't share a verb. v0.40.9.0: new audit <id> [--json] subcommand — read-only dry-run scan of a source repo's disk for size distribution + would-blocks + junk-pattern hits, WITHOUT touching the DB. Catches scraper junk and oversized content BEFORE sync. Walks sources.local_path, reads each markdown file, runs assessContent() from src/core/content-sanity.ts, aggregates by verdict (ok | warn_oversize | hard_block_junk_pattern). JSON envelope is stable for monitoring pipelines. Pinned by test/sources-audit.test.ts (not present in this wave; covered transitively by test/content-sanity.test.ts + test/import-file-content-sanity.test.ts). v0.41.31.0 (P0a): gbrain sources status (the live runStatus health table path, NOT just the buildSyncStatusReport dashboard) gains a BACKFILL column between EMBED and FAILS: active(N) beats queued(N) beats idle, sourced from SourceMetrics.backfill_active / backfill_queued in src/core/source-health.ts. Under federated_v2, sync --all defers embedding to per-source embed-backfill minion jobs, so an operator needs to SEE that deferred work after sync exits 0 instead of assuming it was lost. The data layer (jobCountsBySource in source-health.ts) widens its one minion_jobs SQL with two COUNT(*) FILTER (WHERE name = 'embed-backfill' AND ...) aggregates (active vs waiting/delayed/waiting-children) — best-effort, all-0 on pre-minions brains. Pinned by test/source-health.test.ts.
  • src/commands/reindex-frontmatter.tsgbrain reindex-frontmatter. v0.37.7.0 (#1225): wrapped the query path in the standard withEngine(...) lifecycle so engine.connect() runs before the first SQL call. Pre-fix the command process.exit(1)'d with a TypeError on first invocation. Pinned by test/reindex-frontmatter-connect.test.ts.
  • src/core/source-resolver.ts — 6-tier source resolution. v0.37.7.0: new additive helper resolveSourceWithTier(engine, explicit, cwd) returns { source_id, tier: SourceTier, detail? } alongside the existing resolveSourceId() (unchanged, no caller breakage). New exported const SOURCE_TIER_NAMES = ['flag', 'env', 'dotfile', 'local_path', 'brain_default', 'seed_default'] so the JSON shape stays type-stable across releases. Order matches the 1-6 priority of resolveSourceId(). Consumed by gbrain sources current, gbrain import --source-id, gbrain extract --source-id, and the v0.37.7.0 source_routing_health doctor check. Pinned by test/source-resolver-with-tier.test.ts (uses withEnv() wrapper per the test-isolation lint). v0.41.13.0 (#1434): new tier 5.5 sole_non_default slots between brain_default and seed_default. When NO sources.default config is set AND exactly one registered source has local_path AND isn't 'default', auto-route to it. SOURCE_TIER_NAMES extended to 7 entries. New private pickSoleNonDefaultSource(engine) shared by both resolver entry points (cannot drift). Archived sources excluded (try/catch for pre-v34 brains). New exported formatSoleNonDefaultNudge(sourceId): string | null builds the user-facing stderr nudge (returns null when GBRAIN_NO_SOLE_NON_DEFAULT_NUDGE=1). src/commands/sync.ts:1497-1519 rewired to call resolveSourceWithTier unconditionally so the new tier actually fires (codex caught the original plan shipping dead code). src/commands/import.ts:96-128 mirrors the same wiring with the tier-gated nudge. Pinned by test/source-resolver-sole-non-default.test.ts (14 cases) + test/sync-sole-non-default-routing.test.ts (3 PGLite cases driving real runSync with stderr capture).
  • src/core/sync.ts extension (v0.41.13.0, #1433) — isSyncable factored through a private classifySync(path, opts): SyncableReason | null helper. New exported unsyncableReason(path, opts) companion returns the same tagged reason or null when syncable. SYNC_SKIP_FILES lifted to a named export (the four canonical metafile basenames: schema.md, index.md, log.md, README.md). New SyncableReason union: 'metafile' | 'strategy' | 'pruned-dir' | 'include-glob-miss' | 'exclude-glob-hit'. src/commands/sync.ts:772 cleanup loop now guards on unsyncableReason(path) === 'metafile' so previously-indexed metafile pages (typically log.md from an older gbrain that didn't filter it) survive every re-sync. Honest scope: does NOT cover manifest.deleted (the upstream filter already strips metafiles, so rm log.md doesn't delete the page either — v0.42+ TODO for a gbrain pages remove <slug> operator surface). Pinned by test/sync-isSyncable-shape.test.ts (15 cases pinning the duality contract) + test/sync-metafile-skip.serial.test.ts (3 PGLite cases including the negative case — renamed .md → .txt still gets cleaned up properly).
  • src/core/import-file.ts extension (v0.41.13.0, #1309) — identity-based dedup pre-check at :427-490. Calls the new engine.findDuplicatePage?.(sourceId, {hash, frontmatterId}) (optional ? so test doubles compile unchanged). Posture per codex review: SKIP when frontmatter.id matches (true external duplicate from overlapping ingest roots), WARN-ALWAYS on content_hash collision with different/missing frontmatter.id (templates and daily logs may legitimately share text), FAIL CLOSED on lookup error (silent fallthrough would mask the very bug this fixes), bypass via --force-rechunk. Soft-deleted pages excluded at the engine layer so tombstones don't block legitimate re-imports under new slugs. Pinned by test/import-dedup-frontmatter-id.test.ts (11 cases).
  • src/core/engine.ts extension (v0.41.13.0) — two interface changes: (1) #1309 — new optional findDuplicatePage?(sourceId, {hash, frontmatterId?}): Promise<{slug, id} | null> method; identity precedence is content_hash OR frontmatter->>'id' both with deleted_at IS NULL; (2) #1436 — resolveSlugs(partial, opts?) signature extended with {sourceId?, sourceIds?} so the MCP fuzzy get_page path scopes by source (closes the source-bleed bug class that mirrored the v0.34.1 P0 seal for exact lookups). Field names match sourceScopeOpts(ctx) output so handlers can spread directly. Both signatures back-compatible — callers passing no opts get the pre-fix behavior. Plus a stable tiebreaker ORDER BY score DESC, page_id ASC, chunk_id ASC added to searchVector in both engines: when two chunks tie on score (basis-vector eval fixtures), older page_id wins. Closes the planner-non-determinism regression class — adding any new index to pages could previously flip ranking on tied scores between branches.
  • src/core/migrate.ts v95 (v0.41.13.0, #1309) — pages_dedup_partial_index adds CREATE INDEX pages_dedup_idx ON pages (source_id, content_hash) WHERE deleted_at IS NULL. Postgres path uses CREATE INDEX CONCURRENTLY with transaction: false + pre-drops any invalid remnant from a prior failed attempt; PGLite uses plain CREATE INDEX. Powers findDuplicatePage hot path on 100K-page brains — query becomes O(log n) instead of O(n).
  • src/cli.ts extension (v0.41.13.0, #1422) — dream dispatch at lines 1063-1080 binds the caught engine-connect error and emits [dream] WARNING: could not connect to DB (...) to stderr before falling through to filesystem-only phases. Pre-fix the catch swallowed silently and the user saw gbrain dream exit 0 with no clue why DB phases were skipped. The runDream(null, ...) no-DB fallback is preserved — filesystem phases still run honestly. Pinned by test/cli-dream-engine-warn.test.ts (2 subprocess cases against good + bad DATABASE_URL).
  • src/commands/autopilot.ts extension (v0.37.7.0) — three changes for federated-brain co-existence and launchd hygiene. (1) #1226 lockfile scope: LOCK_PATH resolves via gbrainPath('autopilot.lock') so it honors GBRAIN_HOME. Two brains can run autopilot simultaneously without lock-stealing. Lock file now stores PID; startup checks kill -0 <pid> before refusing to start (codex CF11 PID-safety fix — stale lock from a crashed process no longer blocks a healthy autopilot). (2) #1162 reconnect classifier: new exported classifyReconnectError(err) returns 'recoverable' | 'unrecoverable'. Unrecoverable causes the daemon to process.exit(0) and let launchd back off instead of the v0.37.6 loop that logged config.database_url undefined every 5s forever. (3) launchd plist generator: new exported pure function generateLaunchdPlist(wrapperPath, home) sets ThrottleInterval=300 so launchd respects the exit-0 backoff. Both helpers pinned by test/autopilot-lock-path.test.ts + test/autopilot-reconnect-classifier.test.ts.
  • src/core/oauth-provider.ts + src/commands/serve-http.ts extension (v0.37.7.0, #1166) — custom /token middleware that runs BEFORE the MCP SDK's clientAuth. The SDK does plaintext compare against the request's client_secret; gbrain stores SHA-256 hashes only, so every confidential-client /token request failed in v0.37.0v0.37.6. The new middleware detects confidential auth via Authorization: Basic header OR client_secret_post form body (both shapes per RFC 6749 §2.3.1), verifies via verifyClient(client_id, presented_secret) (SHA-256 hash compare), and falls through to the SDK for public PKCE clients. Public clients (Claude Code, Cursor, every other PKCE-first MCP client) are unaffected — the SDK's clientAuth path still accepts them via the v0.34.1.0 NULL-client_secret_hash normalization. Pinned by test/oauth-confidential-client.test.ts (both client_secret_basic and client_secret_post paths).
  • src/core/sync.ts:pruneDir extension (v0.37.7.0, #1169) — pruneDir(name, parentDir?) signature extended with optional parentDir. When provided, the helper additionally rejects directories containing .git as a FILE — the git submodule gitfile pattern (regular repos have .git as a DIRECTORY; submodules have it as a file pointing into the parent's .git/modules/). Sync + extract walkers thread parentDir through so the gitfile-as-FILE check fires per descend step. Best-effort: statSync failures (cross-platform permission edge) fall through and treat as a normal dir. Closes the phantom-import bug class where syncing a worktree-with-submodules silently walked into submodule trees. Pinned by test/sync-walker-submodule.test.ts.
  • src/core/minions/handlers/subagent.ts extension (v0.37.7.0, #1151) — terminal-state short-circuit on resume. When a stored message thread already ends in stop_reason: 'end_turn', the handler returns { ok: true } immediately instead of issuing another messages.create call. Pre-fix, resume tried to re-prompt past end_turn, the Anthropic API rejected with a 400, and the worker classified the (already-successful) job as failed and dead-lettered it. Pinned by test/subagent-handler.test.ts.
  • src/commands/doctor.ts extension (v0.37.7.0, T12+T13+T14) — three new checks wired into runDoctor() and the JSON envelope. (1) checkSourceRoutingHealth(engine) scans up to 200 pages on federated brains and flags pages whose source_id doesn't match what resolveSourceWithTier() would have picked for their source_path. Single-source brains short-circuit to ok ("no federation to check"). D5 200-page cap keeps doctor under 5s on huge brains; the cap is total across the brain, not per-source. (2) checkOauthConfidentialHealth(engine) probes registered confidential clients for /token reachability — warns when the v0.37.0v0.37.6 plaintext-compare bug class would have rejected them. (3) checkAutopilotLockScope() is a pure-function check (no engine) that compares the resolved lock path to $GBRAIN_HOME; warns when $GBRAIN_HOME is set but the lock lives elsewhere, with a PID-safe inspection hint per codex CF11 (lock file holds the daemon PID; check kill -0 <pid> before considering deletion). All three are warn-only — they surface paste-ready fix hints without flipping the doctor exit code. Pinned by test/doctor-v0_37_7_checks.test.ts.
  • skills/conventions/brain-routing.md (v0.37.7.0, #1222) — agent-facing convention skill documenting the canonical 6-tier source resolution chain (flag → env → dotfile → local_path → brain_default → seed_default) with paste-ready decision tables. Linked from CLAUDE.md's "Two organizational axes" section and from gbrain sources current's hint output.
  • src/commands/doctor.ts extension (v0.40.4.1) — buildChecks(engine, args, dbSource): Promise<Check[]> exported as a test seam. runDoctor is now a thin wrapper: buildChecks → computeDoctorReport → render + process.exit. All 10 process.exit sites stay in the wrapper; the two early-return paths (no engine, connection failure) return partial check lists instead of inline exits. No behavior change — observable output identical because the wrapper renders the same partial list. Pinned by test/doctor-behavioral.test.ts (13 cases: pure aggregation math over computeDoctorReport, orchestrator cases for --fast skip set + --json flag + no-engine partial path + snapshot of load-bearing check names) and test/doctor-cli-smoke.serial.test.ts (1 subprocess case spawning bun run src/cli.ts doctor --json against a fresh PGLite tempdir, asserting schema_version=2 envelope, status enum, non-empty checks array — the render-path coverage that buildChecks-only tests miss). Quarantined as .serial.test.ts because PGLite write-locks don't play with parallel runners.
  • src/core/cycle.ts extension (v0.40.4.1) — runPhaseLint + runPhaseBacklinks gain the export keyword so behavioral tests can drive them directly. No body changes; documented as internal helpers exposed for test-only consumption (downstream code should NOT take a dependency). Pinned by test/cycle-legacy-phases.test.ts (11 cases across both phases: clean run → status='ok', partial fix → status='warn' with dryRun in details, dry-run path doesn't write, throw-from-lib → status='fail' with the wrapper's try/catch envelope populated). Future phase wrappers (sync, extract, embed, orphans, extract_facts, resolve_symbol_edges, recompute_emotional_weight) land as additional describes in the same file, not new files (TODOS NEW-2).
  • test/operations-trust-boundary.test.ts + scripts/check-operations-filter-bypass.sh (v0.40.4.1) — operations trust-boundary contract coverage. Hybrid test design: pure assertions over all 74 ops (every op has a scope annotation; every mutating op has a non-read scope; localOnly: true ops are excluded from operations.filter(op => !op.localOnly); the seven historically-sensitive localOnly ops snapshot-pinned by name) plus targeted handler-invocation regressions for the two historically-broken HTTP-callable classes: submit_job with name='shell' + ctx.remote=true MUST reject (the F7b HTTP MCP shell-job RCE class), and search_by_image with image_path + ctx.remote=true MUST reject (the D18 P0 image-leak class). file_upload and sync_brain deliberately omitted from handler-invocation tests because they're localOnly: true and that path would test an impossible production scenario (codex CMT-3). The shell guard greps src/ for any module importing the operations value outside the canonical filter site at src/commands/serve-http.ts — three import shapes detected (destructured, aliased, namespace), explicit 10-entry allow-list with per-entry rationale, plus a literal-string check that serve-http.ts still contains operations.filter(op => !op.localOnly). Wired into bun run verify.
  • src/core/content-sanity.ts (v0.40.9.0, NEW) — pure assessor for the content-sanity defense wave. assessContent(content, opts): SanityVerdict returns one of three verdicts (ok | warn_oversize | hard_block_junk_pattern | soft_block_oversize) with {reason, bytes, matched_pattern_name?} detail. Six hand-vetted built-in junk patterns (Cloudflare challenge dumps, CAPTCHAs, 403 dumps, bare error-page titles) compiled at module load; operator literal substrings loaded via loadOperatorLiterals() from src/core/content-sanity-literals.ts. ContentSanityBlockError tagged class is the typed throw shape — every wrapper site (gbrain import CLI, put_page MCP op, gbrain sync, /ingest webhook) catches it via the existing exception flow rather than a parallel status check. The bytes-parity contract (D2) pins Buffer.byteLength(content, 'utf8') against the embedder's actual byte count so a 499K-byte page can't be soft-blocked on assessment then overflow on embed. Knob resolution chain: env > file (~/.gbrain/config.json) > DB > defaults — env wins for CI / one-off overrides, file is operator-set, DB plane is what gbrain config set writes. Four knobs: content_sanity.bytes_warn (default 50_000), content_sanity.bytes_block (default 500_000), content_sanity.junk_patterns_enabled (default true), content_sanity.disabled (default false; GBRAIN_NO_SANITY=1 is the loud-stderr kill-switch with per-ingest warning). Pinned by test/content-sanity.test.ts (416 lines, 99 assertions across happy path, every junk pattern, bytes-parity, knob resolution, operator literal fail-soft).
  • src/core/content-sanity-literals.ts (v0.40.9.0, NEW) — operator literal-substring loader. Reads ~/.gbrain/junk-substrings.txt, one literal per non-comment non-blank line. Optional # name=<id> header pairs an identifier with the following literal so audit JSONL groups by site (linkedin_auth_wall, reddit_blocked, etc.). Fail-soft on ENOENT (missing file = empty array, no error). Loaded on every ingest. Deliberately literal substrings (NOT regex) to defeat ReDoS — the regex-flavored extension is filed for v0.41+ once a real ReDoS budget exists. Pinned by test/content-sanity-literals.test.ts (110 lines).
  • src/core/embed-skip.ts (v0.40.9.0, NEW) — 5-site shared predicate for the soft-block embed-skip filter. Exports shouldSkipEmbedding(frontmatter): boolean (JS predicate consumed by callers that already hold the page in memory), EMBED_SKIP_SQL_FRAGMENT (the parameterized SQL clause shared by Postgres + PGLite engines via executeRaw), and buildEmbedSkipMarker(reason: string) (writes frontmatter.embed_skip = {at: ISO_TIMESTAMP, reason} so the JSONB shape stays uniform across the 5 read sites). The 5 sites are: embed.ts --stale, embed.ts --all, the embed-stale Minion helper, plus both engines' listStaleChunks + countStaleChunks. Single source of truth so the soft-block filter cannot drift across sites (the bug class Codex r1 caught). Pinned by test/embed-skip.test.ts (cross-site invariant + JSONB shape).
  • src/core/audit/content-sanity-audit.ts (v0.40.9.0, NEW) — ISO-week JSONL audit at ~/.gbrain/audit/content-sanity-YYYY-Www.jsonl built on the v0.40.4.0 audit-writer.ts primitive. Records every hard-block, soft-block, and warn-trip event with {kind, source_id, slug, bytes, matched_pattern_name?, reason, ts}. Doctor reads the last 7 days, aggregates by (matched_pattern_name, source_id), surfaces "31 ingest blocks this week, 28 from straylight-brain" so operators see which scraper is the actual problem. Honors GBRAIN_AUDIT_DIR for shared-filesystem multi-host setups (documented caveat in the doctor message for ops that don't share the dir). Pinned by test/audit/content-sanity-audit.test.ts (219 lines, 219 assertions).
  • src/commands/doctor.ts extension (v0.40.9.0) — three new checks wired into runDoctor() and the JSON envelope: oversized_pages (warns on pages exceeding content_sanity.bytes_warn), scraper_junk_pages (warns on pages that match any junk pattern despite being live in the DB — these escaped pre-v0.40.9.0 ingest), and content_sanity_audit_recent (reads the last 7 days of audit events, aggregates by pattern+source). Default scans the 1000 most-recent pages; new --content-audit flag opts into a full scan for the cleanup wave. All three are warn-only with paste-ready fix hints (junk page → gbrain sources audit <id> + git rm source-of-truth, oversize → split or accept).
  • src/commands/lint.ts extension (v0.40.9.0) — two new lint rules: huge-page (flags pages exceeding content_sanity.bytes_warn threshold) and scraper-junk (flags pages matching any junk pattern). Both reuse assessContent() from src/core/content-sanity.ts so lint, doctor, and ingest share one assessor — adding a junk pattern automatically covers all three surfaces. lint.ts lifts DB config when ~/.gbrain/ is reachable (matches what gbrain config set writes); falls back to file/env on CI. Pinned by test/lint-content-sanity.test.ts (161 lines).
  • src/commands/embed.ts extension (v0.40.9.0) — applies the embed-skip filter at all 5 stale-chunk sites: runEmbedCore --stale, runEmbedCore --all, the embed-stale Minion helper, plus both engines' listStaleChunks + countStaleChunks via EMBED_SKIP_SQL_FRAGMENT. A soft-blocked page is queryable by title and slug but its chunks never enter the embed sweep. The shared helper from src/core/embed-skip.ts is the regression guard — no per-site ad-hoc filter is allowed. Pinned by test/embed-skip.test.ts.
  • src/core/import-file.ts extension (v0.40.9.0) — importFromContent is the narrow waist that every ingest path passes through (gbrain import, gbrain sync, put_page MCP, /ingest webhook). It now calls assessContent() BEFORE chunking; verdict hard_block_junk_pattern throws ContentSanityBlockError (which every wrapper site already catches via its exception flow); verdict warn_oversize OR oversize-without-junk sets frontmatter.embed_skip via buildEmbedSkipMarker() AND deletes any pre-existing chunks for the page in the same transaction so search can't surface stale chunks against content that's now soft-blocked. gbrain import honors errors > 0 for non-zero exit (was silently exit-0 on failed files). classifyErrorCode in src/core/sync.ts recognizes the new PAGE_JUNK_PATTERN code so sync-failures.jsonl grouping bins these correctly. Pinned by test/import-file-content-sanity.test.ts (206 lines). — shared library for the v0.12.0 graph layer. extractEntityRefs (canonical, replaces backlinks.ts duplicate) matches both [Name](people/slug) markdown links and Obsidian [[people/slug|Name]] wikilinks as of v0.12.3. extractPageLinks, inferLinkType heuristics (attended/works_at/invested_in/founded/advises/source/mentions), parseTimelineEntries, isAutoLinkEnabled config helper. DIR_PATTERN covers people, companies, deals, topics, concepts, projects, entities, tech, finance, personal, openclaw. Used by extract.ts, operations.ts auto-link post-hook, and backlinks.ts.
  • src/core/zombie-reap.ts (v0.28.1) — idempotent installSigchldHandler() so JS-spawned children get reaped via Bun's internal waitpid(). Bun (like Node) only auto-reaps when a SIGCHLD listener is registered; without it, every child the worker spawns (shell jobs, embed batches, sub-agents) becomes a zombie on exit and holds connection slots. Called once at module load from src/cli.ts (with Windows platform guard — SIGCHLD doesn't exist on Windows). Cross-file leak guard via _uninstallSigchldHandlerForTests() for tests. Layer 1 of the three-layer zombie defense; Layer 2 is tini-as-PID-1 wrapping the worker subtree (via src/core/minions/spawn-helpers.ts); Layer 3 is the container's own tini for hard Bun crashes.
  • src/core/minions/ — Minions job queue: BullMQ-inspired, Postgres-native (queue, worker, backoff, types, protected-names, quiet-hours, stagger, handlers/shell).
  • src/core/minions/queue.ts — MinionQueue class (submit, claim, complete, fail, stall detection, parent-child, depth/child-cap, per-job timeouts, cascade-kill, attachments, idempotency keys, child_done inbox, removeOnComplete/Fail). add() takes a 4th trusted arg (separate from opts to prevent spread leakage); protected names in PROTECTED_JOB_NAMES require {allowProtectedSubmit: true} and the check runs trim-normalized (whitespace-bypass safe). v0.14.1 #219: add() plumbs max_stalled through with a [1, 100] clamp; omitted values let the schema DEFAULT (5) kick in. v0.19.0: handleWallClockTimeouts(lockDurationMs) is Layer 3 kill shot for jobs where FOR UPDATE SKIP LOCKED stall detection and the timeout sweep both fail to evict (wedged worker holding a row lock via a pending transaction). v0.19.1: maxWaiting coalesce path now uses pg_advisory_xact_lock keyed on (name, queue) to serialize concurrent submits for the same key, and filters on queue in addition to name so cross-queue same-name jobs don't suppress each other.
  • src/core/minions/worker.ts — MinionWorker class (handler registry, lock renewal, graceful shutdown, timeout safety net). v0.14.0 abort-path fix: aborted jobs now call failJob with reason (timeout/cancel/lock-lost/shutdown) instead of returning silently. shutdownAbort (instance field) fires on process SIGTERM/SIGINT and propagates to ctx.shutdownSignal — shell handler listens to it; non-shell handlers don't. v0.22.1 (#403): per-job timeout fires abort.abort(new Error('timeout')) then a 30-second grace-then-evict safety net force-evicts the job from inFlight and marks it dead in DB if the handler ignores the abort signal — frees the slot even when a handler wedges (the 98-waiting-0-active prod incident driver). v0.41.26.1 (supersedes #1567): the lock-renewal cathedral wave. The launchJob setInterval block — which previously was setInterval(async () => await renewLock(...)) — is the v0.41.22.1 production crash class (~39 worker exits/day from unhandledRejection at renewLock during PgBouncer connection rotation). Replaced with a thin sync wrapper around the new pure runLockRenewalTick function from src/core/minions/lock-renewal-tick.ts. Five additional gaps closed in the same wave: (1) cancelled flag captured in the timer closure stops in-flight IIFEs from writing misleading audit events after the job ended; (2) re-entrancy guard (tickInFlight) plus per-call Promise.race timeout (GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS, default lockDuration/3) so a hung renewLock can't wedge the worker indefinitely; (3) time-based abort (Date.now() - lastSuccessfulRenewalAt >= lockDuration - GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS) replaces count-based, so we release the lock BEFORE another worker can reclaim; (4) explicit .catch() on the stored executeJob(...).finally(...) promise closes the SECOND unhandledRejection vector (if failJob throws during the same DB outage); (5) INFRASTRUCTURE_ABORT_REASONS = new Set(['lock-renewal-failed', 'lock-lost']) is exported; executeJob's catch skips failJob for these reasons so PgBouncer blips no longer dead-letter healthy jobs (the stall detector reclaims cleanly with the lock already expired). The 30s grace-evict safety net moved out of the job.timeout_ms-only branch into a generic abort listener that fires for ANY abort reason. CI guard scripts/check-worker-lock-renewal-shape.sh (wired into bun run verify) asserts the bug pattern lockTimer = setInterval(async ...) stays absent AND launchJob continues to call runLockRenewalTick so the test seam survives refactors. Behavior pinned by test/worker-lock-renewal.test.ts (18 hermetic state-machine cases over the pure function), test/audit/lock-renewal-audit.test.ts (11 cases), test/scripts/check-worker-lock-renewal-shape.test.ts (5 fixture-driven meta-tests). v0.28.1 engine-ownership invariant: start() no longer calls engine.disconnect() on shutdown — that was a leaky abstraction (the worker disconnected an engine it didn't own). The CLI handler in src/commands/jobs.ts case 'work' now owns engine lifecycle via try/finally with loud error logging on disconnect failure. Pinned by test/worker-shutdown-disconnect.test.ts asserting the inverse (disconnectSpy).not.toHaveBeenCalled()). v0.34.3.0: RSS watchdog metric switched to non-file-backed pages on Linux. New exports parseRssFromProcStatus(status) (pure parser, exported for unit tests) and getAccurateRss(readStatus?) (reads /proc/self/status for RssAnon + RssShmem, falls back to process.memoryUsage().rss on macOS / restricted containers / kernel <4.5). The default getRss injected into WorkerOpts is now getAccurateRss instead of process.memoryUsage().rss. Closes the prod incident where VmRSS inflated to 7GB on a 96K-page brain (file-backed git packfile mmaps) while heap stayed at ~100MB; the watchdog was firing every autopilot cycle. M1 parser fix uses field-presence regex checks so RssAnon: 0 + RssShmem: 512 (shmem-only worker case) parses correctly instead of falling through to VmRSS. Pinned by test/worker-rss.test.ts (11 cases). v0.42.5.0 (#1678): checkMemoryLimit now tracks peak RSS, fires an 80%-of-cap soft-warn (once per crossing, carrying peak + in-flight job kinds), and on exceed sets _rssWatchdogTriggered=true (exposed via get rssWatchdogTriggered()) so jobs work's finally can process.exit(WORKER_EXIT_RSS_WATCHDOG) after disconnect — making the drain self-identifying instead of an opaque code-0 exit. The poll loop wraps claim in try/catch: on a retryable conn error it reconnects ONCE and continues to the next tick rather than blind-retrying (Codex #1: a retry after the UPDATE...RETURNING committed but the socket died would double-claim). LockRenewalDeps is wired with reconnect when the engine supports it.
  • src/core/minions/supervisor.ts — MinionSupervisor process manager. Spawns gbrain jobs work as a child, restarts on crash with exponential backoff, periodic health check. v0.22.1 (#406): consecutiveHealthFailures counter; on 3 consecutive failures emits health_warn with reason: 'db_connection_degraded' and calls engine.reconnect() to swap in a fresh pool, then resets the counter. Worker exit classifier emits likely_cause field on worker_exited events: oom_or_external_kill (SIGKILL), graceful_shutdown (SIGTERM), runtime_error (code 1), clean_exit (code 0), unknown. v0.28.1: consumes detectTini() + buildSpawnInvocation() from src/core/minions/spawn-helpers.ts to wrap the worker subtree in tini-as-PID-1 when tini is on PATH (handles native-addon zombie reaping that the in-process SIGCHLD reaper can't reach). Exposes isTiniDetected read-only accessor for tests. v0.34.3.0: spawn-and-respawn loop extracted into the shared ChildWorkerSupervisor core (see entry below). MinionSupervisor now composes the inner class via runSuperviseLoop()new ChildWorkerSupervisor({...}) and maps ChildSupervisorEvent shapes back through the existing emit() SupervisorEvent channel — JSONL audit consumers see byte-compatible output across the rename. PID lock, signal handlers, health check, and process.exit on max-crashes stay in MinionSupervisor (standalone-daemon concerns). The pre-shipped reset-to-0-on-code=0 hunk that originally fixed the prod crash-counter incident is gone; the same fix lives in the shared core under the D1 amendment (code=0 leaves crashCount untouched, so a worker alternating real crashes + watchdog drains still trips max_crashes). D2 cleanRestartBudget (default 10 restarts per 60s) caps the macOS/non-Linux-fallback tight-loop by emitting health_warn { reason: 'clean_restart_budget_exceeded' } plus backoff after the threshold trips. shutdown() drains via childSupervisor.killChild('SIGTERM') + awaitChildExit(35_000) instead of reaching into this.child directly. Pinned by test/supervisor.test.ts (16 cases; existing tests that previously relied on clean-exit-as-crash semantics now use exit-1 workers since clean exits no longer count) and test/supervisor-tini.test.ts.
  • src/core/minions/child-worker-supervisor.ts (v0.34.3.0) — shared spawn-and-respawn core extracted from MinionSupervisor so it can be reused by both MinionSupervisor (standalone gbrain jobs supervisor daemon) and src/commands/autopilot.ts (autopilot daemon). Pre-v0.34.3.0 the two consumers maintained parallel spawn loops that drifted into the same bug class — Codex caught it during plan-eng-review on PR #1003. Pure class: NO PID file, NO signal handlers, NO process.exit, NO health check. Lifecycle events fire via injected onEvent: (ChildSupervisorEvent) => void callback so each composer routes to its own log/audit channel. D1 exit classifier: code === 0 leaves crashCount UNCHANGED (preserves flap detection across mixed exit sequences — a worker that alternates exit 1 / exit 0 / exit 1 / exit 0 correctly trips max_crashes after 10 real crashes regardless of intervening clean exits). code != 0 follows the existing runDuration > stableRunResetMs ? 1 : ++crashCount rule. D2 clean-restart budget: sliding window tracks code=0 exits; when count exceeds cleanRestartBudget (default 10) inside cleanRestartWindowMs (default 60s), emits health_warn { reason: 'clean_restart_budget_exceeded' } and applies cleanRestartBudgetBackoffMs (default 1s) before the next spawn. Caps the worst-case tight-loop on macOS / restricted containers / kernel <4.5 where the worker's RSS watchdog falls back to VmRSS. Public read-only accessors childAlive, inBackoff, crashCount for composer health checks; killChild(signal) + awaitChildExit(timeoutMs) for shutdown paths. awaitChildExit short-circuits when child.exitCode !== null || child.signalCode !== null (regression caught in pre-landing /review: pre-fix, fast-SIGTERM responders caused a 35-second shutdown hang because the late once('exit', ...) listener never fired). Test hooks: _backoffFloorMs skips the real backoff curve, _now injects a fake clock. Pinned by test/child-worker-supervisor.test.ts (7 cases: D1 classifier with code=0 not counted, interleaved exits still trip max_crashes, stable-run + clean-exit interaction across faked 6-minute run, D2 budget triggers backoff + health_warn, budget config is per-instance, awaitChildExit short-circuit, event-shape regression). Plan that produced the design lives at ~/.claude/plans/this-is-a-real-sleepy-sketch.md. v0.42.5.0 (#1678): the exit classifier special-cases WORKER_EXIT_RSS_WATCHDOGlikely_cause='rss_watchdog', and that exit does NOT bump crashCount (so it routes to its own breaker, not the generic crash path). A dedicated _watchdogExitTimestamps sliding window (mirroring _cleanRestartTimestamps) trips a loud rss_watchdog_loop health_warn naming the cap when N watchdog exits land inside the window — INDEPENDENT of the stable-run reset that hid the original 400×/day loop (a >5-min run that watchdog-drains would otherwise never trip max_crashes). New opts watchdogLoopBudget (3), watchdogLoopWindowMs (600000), watchdogBackoffMs (30000); ChildSupervisorEvent extended. supervisor-audit.ts adds rss_watchdog as a non-clean cause + its own CrashSummary.by_cause bucket (denylist preserved so future causes still route to legacy).
  • src/core/minions/spawn-helpers.ts (v0.28.1) — pure detectTini() + buildSpawnInvocation() helpers consumed by both supervisor.ts and autopilot.ts. Resolves the DRY violation between the two spawn sites and makes the tini wrapping testable without mock.module() (rule R2 of scripts/check-test-isolation.sh). detectTini() calls execFileSync('which', ['tini']) with explicit env: process.env so Bun sees runtime PATH mutations (the env-snapshot bug fix). buildSpawnInvocation(tiniPath, cmd, args) returns {cmd, args} with tini prepended when present, or the bare invocation otherwise. Pinned by test/spawn-helpers.test.ts (5 cases) and test/supervisor-tini.test.ts (4 cases).
  • src/core/minions/types.tsMinionJobInput + MinionJobStatus + handler context types. MinionJobInput.max_stalled (new in v0.14.1) is optional; omitted values let the schema DEFAULT (5) kick in, provided values are clamped to [1, 100].
  • src/core/minions/protected-names.ts — side-effect-free constant module exporting PROTECTED_JOB_NAMES + isProtectedJobName(). Kept pure so queue core can import without loading handler modules.
  • src/core/minions/handlers/shell.tsshell job handler. Spawns /bin/sh -c cmd (absolute path, PATH-override-safe) or argv[0] argv[1..] (no shell). Env allowlist: PATH, HOME, USER, LANG, TZ, NODE_ENV + caller env: overrides + (v0.36.5.0) inherit:-resolved keys. UTF-8-safe stdout/stderr tail via string_decoder.StringDecoder. Abort (either ctx.signal or ctx.shutdownSignal) fires SIGTERM → 5s grace → SIGKILL on child. Requires GBRAIN_ALLOW_SHELL_JOBS=1 on worker (gated by registerBuiltinHandlers). v0.36.5.0: ShellJobParams.inherit?: string[] is a free-form list of snake_case config-key names. The worker resolves each via loadConfig() and injects the value under the derived env key (database_urlGBRAIN_DATABASE_URL; everything else uppercased). Names persist in minion_jobs.data (and the shell-audit JSONL); values never do. The canonical validator validateShellJobParams (sibling file shell-validate.ts) runs pre-enqueue in both submit surfaces — gbrain jobs submit shell (jobs.ts:271) AND submit_job op for name='shell' (operations.ts:2085). The handler-entry re-validation here is defense-in-depth. Closes the codex F-CDX-1 load-bearing bug class where validation in the handler ran AFTER queue.add() persisted the row. The validator does NOT police which config keys the agent inherits — same-uid trust model treats the agent as a peer of the worker.
  • src/core/minions/handlers/shell-inherit.ts (v0.36.5.0, NEW) — three small helpers, no closed enum. INHERIT_NAME_RE (/^[a-z][a-z0-9_]*$/) is the snake_case shape guard used by the validator; rejects __proto__, leading-underscore, uppercase, and path-traversal shapes so audit logs stay readable and prototype-pollution lookups can't smuggle through. deriveEnvKey(name) maps config-key → child-env-key (name.toUpperCase() with one override: database_urlGBRAIN_DATABASE_URL because plain DATABASE_URL is ambiguous). resolveInheritValue(cfg, name) is the value lookup; uses Object.hasOwn to defeat prototype-pollution lookups, returns undefined for missing / non-string / empty-string values. An earlier closed-enum design (hardcoded INHERITABLE record with shadow-keys per name) was abandoned because the agent and worker share a uid — refusing to let the agent inherit arbitrary config keys defends nothing in that trust model.
  • src/core/minions/handlers/shell-validate.ts (v0.36.5.0, NEW) — validateShellJobParams(data, opts?) shared pre-enqueue validator. Throws UnrecoverableError with paste-ready operator hints on every failure path. Three rules: (1) existing cmd/argv/cwd/env shape, (2) inherit array shape + snake_case regex per element (prototype-pollution defense), (3) fail-fast on missing config value with gbrain config set <key> hint. Also accepts optional redact_secrets?: boolean for output-side scrubbing. The validator deliberately does NOT police WHICH secrets the agent passes — single-uid trust model. Test seam: opts.config lets unit tests drive the validator hermetically without mocking the module. The defense-in-depth re-call at shell.ts handler entry catches pre-existing rows submitted before v0.36.5.0.
  • src/core/minions/handlers/shell-redact.ts (v0.36.5.0, NEW) — opt-in output-side scrubbing for shell-job stdout/stderr. Pure redactSecretsInText(text, secrets) function: string-mode replaceAll so regex metacharacters in values stay literal. When the caller passes redact_secrets: true (or --redact-secrets on the CLI), the handler builds a Map of inherit-name → resolved-value and post-processes both tails before throw/return, so the persisted result.stdout_tail / result.stderr_tail / error_text carry <REDACTED:name> instead of the value. Only inherit:-resolved values are scrubbed; caller-supplied env: values stay through (those are the agent's "fine in the row" channel). Heuristic — defeats the common-case echo "$GBRAIN_DATABASE_URL" echo, not adversarial encode-then-print. Default false for back-compat.
  • src/core/config.ts:ensureGitignore (v0.36.5.0) — idempotent retroactive writer of ~/.gbrain/.gitignore (single line *). Called from saveConfig() so every config-writing path lays it down, AND from runPostUpgrade() so existing users pick it up on next gbrain upgrade. Never clobbers a user-customized .gitignore (checks file exists + content non-empty before writing). Honest scope, named in CHANGELOG: blocks casual git add ~/.gbrain from inside an enclosing worktree, but does NOT cover already-tracked files, screenshots, backups (Time Machine / iCloud / Dropbox), or git add -f. The doctor check home_dir_in_worktree surfaces what .gitignore can't.
  • src/commands/doctor.ts:home_dir_in_worktree (v0.36.5.0) — filesystem check walking up from gbrainPath() toward $HOME looking for either a .git directory (main repo) or .git file (linked worktree pointer; Conductor + git-worktrees topology). Walk terminates at $HOME so a .git above the user's home doesn't false-positive. Honors GBRAIN_HOME (gbrain appends .gbrain to the override). Warn (not fail) with worktree-root path + paste-ready fix pointing at GBRAIN_HOME override or moving the brain.
  • src/core/minions/handlers/shell-audit.ts — per-submission JSONL audit trail at ~/.gbrain/audit/shell-jobs-YYYY-Www.jsonl (ISO-week rotation; override via GBRAIN_AUDIT_DIR). Best-effort: mkdirSync(recursive) + appendFileSync; failures logged to stderr, submission not blocked. Logs cmd (first 80 chars) or argv (JSON array). Never logs env values.
  • src/core/minions/handlers/supervisor-audit.ts — supervisor lifecycle JSONL audit at ~/.gbrain/audit/supervisor-YYYY-Www.jsonl (ISO-week rotation; shares computeIsoWeekName() helper with shell-audit.ts). writeSupervisorEvent(emission, supervisorPid) appends one line per supervisor event (started, worker_spawned, worker_exited, backoff, health_warn, health_error, max_crashes_exceeded, shutting_down, stopped, worker_spawn_failed). readSupervisorEvents({sinceMs}) is the readback path for gbrain doctor. v0.35.5.0: new exports isCrashExit(event), summarizeCrashes(events), CrashSummary type, and CLEAN_EXIT_CAUSES denylist ('clean_exit' | 'graceful_shutdown'). Single regression point — both gbrain doctor (Lane D supervisor check at doctor.ts:1011-1043) and gbrain jobs supervisor status (jobs.ts:803-826) import from here so the two CLI surfaces cannot drift. isCrashExit classifies a single worker_exited event against the denylist: clean_exit / graceful_shutdown are NON-crashes; everything else (runtime_error, oom_or_external_kill, unknown, AND any future likely_cause value added upstream in child-worker-supervisor.ts) is a crash. Pre-v0.34 audit lines lacking likely_cause fall back to code !== 0. summarizeCrashes returns {total, by_cause: {runtime_error, oom_or_external_kill, unknown, legacy}, clean_exits} so dashboards bind to named buckets — the legacy bucket catches BOTH pre-v0.34 fallback entries AND future unrecognized likely_cause values, fail-loud instead of silent underreport. Denylist-over-allowlist was a codex outside-voice catch during /plan-eng-review — the bug being fixed (read sites counting every worker_exited as a crash, inflating to 120+/day on healthy brains after v0.34.3.0 watchdog drains) was itself an allowlist-of-event-names. Pinned by test/supervisor-audit.test.ts (14 cases: 9-case isCrashExit branch matrix including denylist regression guard for unrecognized future causes + non-exit-event defensive case, 5-case summarizeCrashes aggregator including unrecognized-cause routing to legacy + null-code edge case) and 4 source-grep wiring assertions in test/doctor.test.ts guarding both surfaces against drift.
  • src/core/minions/backpressure-audit.ts (v0.19.1) — sibling of shell-audit.ts for maxWaiting coalesce events. JSONL at ~/.gbrain/audit/backpressure-YYYY-Www.jsonl. Fires one line per coalesce with (queue, name, waiting_count, max_waiting, returned_job_id, ts). Closes the silent-drop vector the v0.19.0 maxWaiting guard introduced.
  • src/core/minions/handlers/subagent.ts (v0.15) — LLM-loop handler. Two-phase tool persistence (pending → complete/failed), replay reconciliation for mid-dispatch crashes, dual-signal abort (ctx.signal + ctx.shutdownSignal), Anthropic prompt caching on system + tool defs. makeSubagentHandler({engine, client?, ...}) factory; MessagesClient is an injectable interface the real SDK implements structurally. Throws RateLeaseUnavailableError (renewable) when rate-lease capacity is full. v0.30.2: Anthropic 400 prompt is too long responses (status 400 + body matches /prompt is too long|prompt_too_long|context.*length/i) classify as UnrecoverableError so the job goes straight to dead on first attempt instead of stalling three times before dead-lettering. Catches both initial-prompt overflow and turn-N tool-loop accumulation that the chunker in synthesize.ts can't bound ahead of time.
  • src/core/minions/handlers/subagent-aggregator.ts (v0.15) — subagent_aggregator handler. Claims AFTER all children resolve (queue changes guarantee every terminal child posts a child_done inbox message with outcome). Reads inbox via ctx.readInbox(), builds deterministic mixed-outcome markdown summary. No LLM call in v0.15.
  • src/core/minions/handlers/subagent-audit.ts (v0.15) — JSONL audit + heartbeat writer at ~/.gbrain/audit/subagent-jobs-YYYY-Www.jsonl. Events: submission (one line per submit) + heartbeat (per turn boundary: llm_call_started | llm_call_completed | tool_called | tool_result | tool_failed). Never logs prompts or tool inputs. readSubagentAuditForJob(jobId, {sinceIso}) is the readback path for gbrain agent logs.
  • src/core/minions/rate-leases.ts (v0.15) — lease-based concurrency cap for outbound providers (default key anthropic:messages, max via GBRAIN_ANTHROPIC_MAX_INFLIGHT). Owner-tagged rows with expires_at auto-prune on acquire; pg_advisory_xact_lock guards check-then-insert; CASCADE on owning job deletion. renewLeaseWithBackoff retries 3x (250/500/1000ms).
  • src/core/minions/wait-for-completion.ts (v0.15) — poll-until-terminal helper for CLI callers. TimeoutError does NOT cancel the job; AbortSignal exits without throwing. Default pollMs: 1000 on Postgres, 250 on PGLite inline.
  • src/core/minions/transcript.ts (v0.15) — renders subagent_messages + subagent_tool_executions to markdown. Tool rows splice under their owning assistant tool_use by tool_use_id. UTF-8-safe truncation; unknown block types fall through to fenced JSON.
  • src/core/minions/plugin-loader.ts (v0.15) — GBRAIN_PLUGIN_PATH discovery. Absolute paths only, left-wins collision, gbrain.plugin.json with plugin_version: "gbrain-plugin-v1", plugins ship DEFS only (no new tools), allowed_tools: validated at load time against the derived registry.
  • src/core/minions/tools/brain-allowlist.ts (v0.15, extended v0.23, v0.29, v0.35.3.0) — derives subagent tool registry from src/core/operations.ts. 13-name allow-list as of v0.29 (was 11). By default put_page schema is namespace-wrapped per subagent (^wiki/agents/<subagentId>/.+). v0.23 trusted-workspace path: when BuildBrainToolsOpts.allowedSlugPrefixes is set, the put_page schema instead describes the prefix list to the model and the OperationContext is threaded with allowedSlugPrefixes. Trust comes from PROTECTED_JOB_NAMES gating subagent submission — MCP cannot reach this field. Only cycle.ts (synthesize/patterns) and direct CLI submitters set it. v0.29: get_recent_salience + find_anomalies added to the allow-list. get_recent_transcripts deliberately NOT added — all subagent calls run with ctx.remote === true, and the v0.29 trust gate rejects remote callers, so adding it would always reject (footgun). The cycle synthesize phase already calls discoverTranscripts directly. v0.35.3.0: paramsToInputSchema() now consumes paramDefToSchema from src/mcp/tool-defs.ts instead of its own inline destructure. Required-aggregation at the tool-def level stays here (out of scope for the shared helper, which is per-param). Closes the third drift site in the ParamDef→JSON Schema bug class.
  • src/mcp/tool-defs.ts (v0.15, v0.35.3.0) — extracted buildToolDefs(ops) helper. MCP server + subagent tool registry both call it; byte-for-byte equivalence pinned by test/mcp-tool-defs.test.ts. v0.35.3.0: exports the new recursive paramDefToSchema(p: ParamDef) helper — single source of truth for ParamDef→JSON Schema mapping. Three consumers now share one mapper: buildToolDefs (stdio MCP), src/commands/serve-http.ts:837 (HTTP MCP tools/list), and src/core/minions/tools/brain-allowlist.ts:84 (subagent tool registry). Pre-v0.35.3, three inline destructures had drifted across the surface — the live HTTP MCP path dropped items on every array param after a v0.32 review caught only the stdio side. Recursive on items so nested array-of-arrays preserves inner shape on the wire. Key ordering (type, description, enum, default, items) is intentional — matches the pre-v0.35.3 inline mappers so JSON.stringify output stays byte-stable. test/mcp-tool-defs.test.ts adds a findArrayWithoutItems walker that fails the suite with a property path on any future type: 'array' lacking items.type.
  • src/core/minions/attachments.ts — Attachment validation (path traversal, null byte, oversize, base64, duplicate detection)
  • src/commands/agent.ts (v0.16) — gbrain agent run <prompt> [flags] CLI. Submits subagent (or N children + 1 aggregator) under {allowProtectedSubmit: true}. Single-entry --fanout-manifest short-circuits. Children get on_child_fail: 'continue' + max_stalled: 3. --follow is the default on TTY; streams logs + polls waitForCompletion in parallel. Ctrl-C detaches, does not cancel.
  • src/commands/agent-logs.ts (v0.16) — gbrain agent logs <job> [--follow] [--since]. Merges JSONL heartbeat audit + subagent_messages into a chronological timeline. parseSince accepts ISO-8601 or relative (5m, 1h, 2d). Transcript tail renders only for terminal jobs.
  • src/commands/jobs.tsgbrain jobs CLI subcommands + gbrain jobs work daemon. v0.28.1: case 'work' now wraps worker.start() in try/finally and owns engine lifecycle — calls engine.disconnect() on shutdown with loud error logging on failure. Replaces the prior call inside MinionWorker.start() (which violated engine ownership: the worker disconnected an engine it didn't own, and clobbered the module-level singleton on PostgresEngine via the now-fixed idempotency bug). Pool slots now free immediately on shutdown instead of waiting for TCP keepalive (~minutes). v0.13.1 surfaces the full MinionJobInput retry/backoff/timeout/idempotency surface as first-class CLI flags on jobs submit: --max-stalled, --backoff-type fixed|exponential, --backoff-delay, --backoff-jitter, --timeout-ms, --idempotency-key. jobs smoke --sigkill-rescue is the opt-in regression guard for #219. v0.16 wires registerBuiltinHandlers to always register subagent + subagent_aggregator (no env flag — ANTHROPIC_API_KEY is the natural cost gate, trust is via PROTECTED_JOB_NAMES) and loads GBRAIN_PLUGIN_PATH plugins at worker startup with a loud startup-line per plugin. shell handler still gated by GBRAIN_ALLOW_SHELL_JOBS=1 (RCE surface, separate concern). v0.22.10 (#521): the autopilot-cycle handler now forwards job.data.phases to runCycle (was previously discarded — caller-supplied phase selection silently became a full cycle). Phases are validated against ALL_PHASES from src/core/cycle.ts; invalid names are filtered out and an empty/missing array falls back to the default 6-phase cycle. v0.22.13 (PR #490 CODEX-1+CODEX-4): sync handler now resolves sourceId at entry by looking up sources.local_path (mirrors cycle.ts:480's autopilot fix from PR #475) so multi-source brains read the per-source last_commit anchor instead of the global config key. Concurrency routed through the shared autoConcurrency() policy in src/core/sync-concurrency.ts instead of the prior hardcoded 4; PGLite stays serial. noEmbed default is true (embed is a separate job — submit gbrain embed --stale after sync, or rely on the autopilot cycle's embed phase). v0.35.5.0: gbrain jobs supervisor status at jobs.ts:803-826 now consumes summarizeCrashes() from src/core/minions/handlers/supervisor-audit.ts for cross-surface parity with gbrain doctor. JSON output adds crashes_by_cause: {runtime_error, oom_or_external_kill, unknown, legacy} + clean_exits_24h fields so dashboards bind to named buckets; human output gains a per-cause line under Crashes (24h) plus a Clean exits (24h) line. Pre-fix the read site at jobs.ts:805 counted every worker_exited event as a crash regardless of likely_cause — the same bug class the v0.35.5.0 doctor fix closes. Pinned by the 4 source-grep wiring assertions in test/doctor.test.ts that require the per-cause breakdown substrings (crashes_by_cause, clean_exits_24h=) to appear in BOTH doctor.ts and jobs.ts.
  • src/commands/features.tsgbrain features --json --auto-fix: usage scan + feature adoption salesman
  • src/commands/autopilot.tsgbrain autopilot --install: self-maintaining brain daemon (sync+extract+embed). v0.28.1: consumes detectTini() from src/core/minions/spawn-helpers.ts and resolves it once at startup instead of per worker respawn (was paying an execFileSync cost on every restart). v0.34.3.0: inline spawn-and-respawn loop replaced with a ChildWorkerSupervisor instance. Drops crashCount, lastWorkerStartTime, STABLE_RUN_RESET_MS, startWorker, and the inline child.on('exit') block — all consolidated into the shared core. --max-rss 2048 and maxCrashes: 5 preserved from the legacy loop. onMaxCrashesExceeded now routes through autopilot's own shutdown('max_crashes') so the autopilot lockfile gets cleaned up (pre-refactor the inline loop called process.exit(1) directly and bypassed cleanup). shutdown() drains via childSupervisor.killChild('SIGTERM') + awaitChildExit(35_000) instead of workerProc.kill(). Pinned by test/autopilot-supervisor-wiring.test.ts (6 static-shape regression guards: composes ChildWorkerSupervisor not the legacy inline names, --max-rss 2048 in argv, maxCrashes: 5 literal, shutdown-via-callback wiring, no workerProc reference). Closes the parallel-supervisor bug class Codex flagged during plan-eng-review.
  • src/mcp/server.ts — MCP stdio server (generated from operations). v0.22.7: tool-call handler delegates to dispatchToolCall from src/mcp/dispatch.ts so stdio + HTTP transports share one validation, context-build, and error-format path. v0.34.1.0 (#870): stdin 'end' / 'close' shutdown hooks are skipped when process.env.MCP_STDIO === '1'. Gateway-piped stdio MCP wrappers (OpenClaw's bundle-mcp, similar) pipe the JSON-RPC handshake then close their stdin half; pre-fix this killed the server before the first tool call landed. Signal handlers (SIGTERM / SIGINT / SIGHUP) and the parent-process watchdog still cover legitimate disconnects. src/commands/serve.ts exposes ServeOptions.mcpStdio?: boolean as a test seam so the runtime guard is exercisable without process.env mutation. Pinned by test/serve-stdio-lifecycle.test.ts.
  • src/mcp/dispatch.ts (v0.22.7) — Shared tool-call dispatch consumed by both stdio (server.ts) and HTTP transports. Exports dispatchToolCall(engine, name, params, opts), buildOperationContext(engine, params, opts), and validateParams(op, params). Single source of truth for (ctx, params) handler arg order and the 5-field OperationContext shape (engine + config + logger + dryRun + remote). Defaults to remote: true (untrusted); local CLI callers pass remote: false. Closed F1/F2/F3 drift bugs in the original v0.22.5 HTTP transport. v0.26.9 (F8): adds summarizeMcpParams(opName, params) — privacy-preserving redactor for mcp_request_log and the admin SSE feed. Returns {redacted, kind, declared_keys, unknown_key_count, approx_bytes}. Intersects submitted top-level keys against the operation's declared params allow-list (declared keys preserved as a sorted array for debug visibility; unknown keys counted but never named, closing the attacker-controlled-key-name leak). Byte counts bucketed up to nearest 1KB so an attacker can't binary-search secret-content sizes via repeated probes. Operators on a personal laptop who want raw payload visibility opt back in with gbrain serve --http --log-full-params (loud stderr warning at startup). Canonical helper — new logging code paths route through it rather than JSON.stringify(params).
  • src/mcp/rate-limit.ts (v0.22.7) — Bounded-LRU token-bucket limiter. buildDefaultLimiters() returns the two-bucket pipeline: pre-auth IP (30/60s, fires BEFORE the DB lookup so brute-force load against access_tokens is actually capped) + post-auth token-id (60/60s). Tracks lastTouchedMs separately from lastRefillMs so an exhausted key can't be reset by hammering past the TTL. LRU cap bounds memory under attacker-controlled key growth.
  • src/commands/serve-http.ts (v0.26.0) — Express 5 HTTP MCP server with OAuth 2.1, admin dashboard, and SSE live activity feed. Started via gbrain serve --http [--port N] [--token-ttl N] [--enable-dcr] [--public-url URL] [--log-full-params]. Supersedes the v0.22.7 src/mcp/http-transport.ts simple bearer-auth path. Combines MCP SDK's mcpAuthRouter (authorize / token / register / revoke endpoints), a custom client_credentials handler (SDK's token endpoint throws UnsupportedGrantTypeError for CC; the custom handler runs BEFORE the router and falls through for auth_code / refresh_token), requireBearerAuth middleware for /mcp with scope enforcement before op dispatch, localOnly rejection, and express-rate-limit at 50 req / 15 min on /token. Serves the built admin SPA from admin/dist/ with SPA fallback. /admin/events SSE endpoint broadcasts every MCP request to connected admin browsers. cookie-parser middleware wired (Express 5 has no built-in). Startup logging prints port, engine, configured issuer URL (honors --public-url), registered-client count, DCR status, and admin bootstrap token. v0.26.9 hardening pass: F7 sets remote: true explicitly on the /mcp request handler's OperationContext literal (closes the HTTP shell-job RCE — without this, submit_job's protected-name guard at operations.ts:1391 saw a falsy undefined and skipped, letting a read+write-scoped OAuth token submit shell jobs). F8 wires summarizeMcpParams from src/mcp/dispatch.ts into both mcp_request_log writes and the admin SSE feed by default (raw payloads opt-in via --log-full-params with stderr warning). F9 sets cookie Secure flag when behind HTTPS or a public-URL proxy. F10 caps the magic-link nonce store with an LRU bound. F12 routes DCR disable through the GBrainOAuthProvider constructor's dcrDisabled option instead of the prior monkey-patch on the express router. F14 wraps transport.handleRequest in try/catch so SDK throws return a JSON-RPC 500 envelope instead of express's default HTML error page. F15 unifies OperationError + unexpected exceptions through buildError / serializeError so /mcp always returns the same envelope shape. v0.28.1: /health endpoint extracted into pure probeHealth(engine) async function with HEALTH_TIMEOUT_MS = 3000 exported constant — drops the timeout from 5s to 3s so Fly.io's 5s health-check deadline gets 2s of headroom for TCP, response framing, and clock skew. Races engine.getStats() against the timeout via Promise.race; saturated pool returns 503 with Health check timed out (database pool may be saturated) instead of hanging. clearTimeout in finally block prevents pending-timer pile-up under high probe rates (race-leak fix from adversarial review). v0.28.10: /health is now liveness-only via the new probeLiveness(sql, engineName, version, timeoutMs) helper that races sql\SELECT 1`againstHEALTH_TIMEOUT_MSand returns the sameProbeHealthResulttagged-union asprobeHealth(single timer-cleanup site, single 503 envelope). Body shape:{status, version, engine}only — engine stats are no longer spread on the public route. Full stats moved to a new admin endpoint/admin/api/full-stats(sibling to/admin/api/statsand/admin/api/health-indicators) gated by the existing requireAdminmiddleware; that route callsprobeHealth(engine, ...)and returns the original spread-stats body.?full=truequery param removed entirely. Closes the original DoS surface wheregetStats()'s 6× count(*) on 96K-page brains through PgBouncer exceeded HEALTH_TIMEOUT_MSand triggered orchestrator restart cascades (Fly.io / k8s seeing 503 → restart loop → advisory-lock pile-up on the migration lock). Outside-voice review (Codex) caught that/admin/api/health-indicatorsis NOT a full-stats endpoint (returns only{expiring_soon, error_rate}), and that an alternative loopback-IP gate would have depended on app.set('trust proxy', 'loopback')semantics holding under proxy/XFF misconfiguration; the shipped admin-cookie design avoids both. **v0.31.3 (#681):** every OAuth/admin/audit SQL call routes throughsqlQueryForEngine(engine)fromsrc/core/sql-query.tssogbrain serve --httpworks against PGLite brains. The fourmcp_request_log.paramsINSERT sites (success path, auth_failed path, scope_denied path, server-error path) all go throughexecuteRawJsonb(engine, ...)so the JSONB column stores real objects, not JSON-encoded strings — closes the bug whereparams->>'op'returned the encoded string"search"(with quotes) instead ofsearch. Migration v46 normalizes any pre-v0.31.3 string-shaped backlog rows on first start. **v0.34.1.0 (#864):** new --bind HOSTCLI flag with default127.0.0.1. Personal-laptop installs no longer publish the brain to the LAN by accident. Self-hosted operators pass --bind 0.0.0.0(or a specific interface IP) once to accept remote connections. A stderr WARN fires when--public-urlis set without--bindso the operator sees the binding before the first request (common cause of "ngrok forwards to me but the agent can't reach the upstream" misconfigurations). The startup banner prints aBind:line. **v0.34.1.0 (#861):** drops the(authInfo as AuthInfo & {sourceId?: string}).sourceId ?? env ?? 'default'cast chain —AuthInfo.sourceIdandAuthInfo.allowedSourcesare now the typed source of truth, populated byoauth-provider.ts:verifyAccessTokenfrom theoauth_clientsrow. **v0.35.3.0:** the inline ParamDef→schema mapper at:837-849(HTTP MCPtools/listhandler) is replaced withparamDefToSchema(v)fromsrc/mcp/tool-defs.ts. Pre-fix this site silently dropped itemson every array param so strict-mode OAuth clients (Gemini Pro structured outputs, OpenAI strict tool defs) rejected the whole tool list. Single mapper now serves stdio MCP, HTTP MCPtools/list`, and the subagent registry.
  • src/core/sql-query.ts (v0.31.3) — Engine-aware tagged-template SQL adapter for OAuth/admin/auth infrastructure. sqlQueryForEngine(engine) returns a SqlQuery ((strings, ...values) => Promise<rows[]>) that walks the template, builds $N positional SQL, asserts every value is a SqlValue (string | number | bigint | boolean | Date | null), and routes through engine.executeRaw(sql, params) so Postgres goes via postgres.js's unsafe(sql, params) path and PGLite via its embedded db.query(sql, params). Deliberately narrower than postgres.js's sql tag: no nested fragments, no sql.json(), no sql.unsafe(), no sql.begin(), no array binding. The narrow surface is the feature — codex finding #7 from the v0.31 plan review argued the adapter should stay scalar-only or it drifts into a partial postgres.js clone. JSONB writes go through the separate executeRawJsonb(engine, sql, scalarParams, jsonbParams) helper that composes positional $N::jsonb casts and passes JS objects through; the v0.12.0 double-encode bug class doesn't apply because positional binding through unsafe() reaches the wire protocol with the correct type oid (verified by test/sql-query.test.ts on PGLite and test/e2e/auth-permissions.test.ts:67 on Postgres). scripts/check-jsonb-pattern.sh doesn't fire because executeRawJsonb(...) is a method call, not the banned literal-template-tag interpolation pattern. Consumed by src/commands/auth.ts, src/commands/serve-http.ts, src/core/oauth-provider.ts, src/commands/files.ts, and src/mcp/http-transport.ts so all five sites work against PGLite and Postgres uniformly. Closes the bug where gbrain auth + gbrain serve --http were silently Postgres-only because they routed every SQL through the postgres.js singleton (community PR #681).
  • src/commands/serve.ts (v0.31.3) — gbrain serve stdio MCP entrypoint with idempotent shutdown across every parent-disconnect signal. Stdio EOF, SIGTERM, SIGINT, SIGHUP, and parent-process death (every reparent case — PID 1, launchd subreaper, systemd, tmux, or a parent shell with PR_SET_CHILD_SUBREAPER) all funnel into one cleanup(reason) path that releases the engine and the PGLite write-lock dir within 5 seconds. Pre-v0.31.3 the stdio MCP server held the lock indefinitely after Claude Desktop / Cursor / launchd-managed gateways disconnected, forcing a 5-minute stale-lock wait on the next start. Watchdog reparent check is getParentPid() !== initialParentPid (capturing the initial ppid once at install time and firing on any change); the previous === 1 check missed the subreaper case under launchd / systemd. Bun's process.ppid cache is stale across reparenting (see oven-sh/bun#30305) so getParentPid() runs spawnSync('ps', ['-o', 'ppid=', '-p', PID]) per tick to read the live kernel PPID. Startup probe verifies ps is on PATH; if not (stripped containers, busybox without procps), the watchdog skips installing AND emits a loud [gbrain serve] watchdog disabled: ps unavailable, parent-death detection unavailable — child will rely on stdin EOF / signals only stderr line so operators see the degraded mode at boot. Pinned by test/serve-stdio-lifecycle.test.ts (22 cases). Closes #413, #446. Credit @Aragorn2046 (origin features in #591) and @seungsu-kr (rebased submitter, Bun ppid workaround).
  • src/core/oauth-provider.ts (v0.26.0) — GBrainOAuthProvider implementing the MCP SDK's OAuthServerProvider + OAuthRegisteredClientsStore interfaces. Backed by raw SQL (works on both PGLite and Postgres — OAuth is infrastructure, not a BrainEngine concern). Full OAuth 2.1 spec: authorize + exchangeAuthorizationCode with PKCE (for ChatGPT), client_credentials (for Perplexity / Claude), refresh_token with rotation, revokeToken, registerClient (DCR path validates redirect_uri must be https:// or loopback per RFC 6749 §3.1.2.1). All tokens + client secrets SHA-256 hashed before storage. Auth codes single-use with 10-minute TTL via atomic DELETE...RETURNING (closes RFC 6749 §10.5 TOCTOU race). Refresh rotation also DELETE...RETURNING (closes §10.4 stolen-token detection bypass). pgArray() escapes commas/quotes/braces in elements so a comma-bearing redirect_uri can't smuggle a second array element. Legacy access_tokens fallback in verifyAccessToken grandfathers pre-v0.26 bearer tokens as read+write+admin. sweepExpiredTokens() runs on startup wrapped in try/catch. v0.26.9 RFC 6749/7009 hardening pass: F1+F2 fold client_id atomically into the DELETE WHERE clauses for both auth-code exchange and refresh rotation — pre-fix the post-hoc client compare burned the row on wrong-client paths so the legitimate client couldn't retry. F3 enforces refresh-scope-subset against the original grant on the row (RFC 6749 §6), not the client's currently-allowed scopes — fixes the case where revoking a scope from a client wouldn't shrink the agent's existing refresh tokens. F4 binds client_id on revokeToken so a client can only revoke its own tokens (RFC 7009 §2.1). F7c validates the /token request's redirect_uri against the value stored at /authorize (RFC 6749 §4.1.3) — empty-string treated as missing rather than wildcard match (adversarial-review fix). F5 swaps bare catch {} blocks in verifyAccessToken and getClient for isUndefinedColumnError from src/core/utils.ts — only SQLSTATE 42703 falls through to legacy fallback; lock timeouts and network blips throw and surface. F6 makes sweepExpiredTokens() actually return the count via RETURNING 1 + array length, not a fire-and-forget zero. F12 adds dcrDisabled constructor option so serve-http.ts can disable the /register endpoint without monkey-patching the router. v0.26.2: module-private coerceTimestamp() boundary helper at the top of the file normalizes postgres-driver-as-string BIGINT columns to JS numbers at every read site (5 call sites: getClient L112+L113 for DCR /register RFC 7591 §3.2.1 numeric timestamps, exchangeRefreshToken L274 + verifyAccessToken L296+L303 for the SDK's typeof === 'number' bearerAuth check). Throws on non-finite input (NaN/Infinity) so corrupt rows fail loud at the boundary instead of riding through as expiresAt: NaN; returns undefined for SQL NULL so callers decide NULL semantics explicitly (refresh + access token paths treat NULL as expired). Helper intentionally NOT promoted to src/core/utils.ts — codex review flagged repo-wide BIGINT precision-loss risk for a generic helper. v0.34.1.0 (#909): registerClient honors token_endpoint_auth_method: "none" (RFC 7591 §3.2.1) — public PKCE clients (Claude Code, Cursor, every other PKCE-first MCP client) store client_secret_hash = NULL and the response payload omits client_secret entirely. Confidential clients (default client_secret_post and explicit client_secret_basic) keep their one-time-reveal shape. getClient correctly normalizes a NULL client_secret_hash to JS undefined so the SDK's clientAuth path accepts the public client at /token. v0.34.1.0 (#861 + #876): verifyAccessToken JOINs oauth_clients.source_id (write scope, scalar) + oauth_clients.federated_read (read scope, TEXT[]) and surfaces both on the returned AuthInfo. Pre-v60 / pre-v61 brains degrade gracefully via isUndefinedColumnError fallback so the upgrade chain is non-blocking on legacy DBs.
  • admin/ (v0.26.0) — React 19 + Vite + TypeScript admin SPA embedded in the binary via admin/dist/ served by serve-http.ts. 7 screens: Login (bootstrap token → session cookie), Dashboard (metrics + SSE feed + token health), Agents (sortable table + sparklines + Register button), Register (modal with scope checkboxes + grant type selector), Credentials reveal (full-screen modal with Copy + Download JSON + yellow one-time-only warning), Request Log (filterable paginated), Agent Detail drawer (Details / Activity / Config Export tabs + Revoke). Design tokens: #0a0a0f bg, Inter for UI, JetBrains Mono for data, 4-32px spacing scale, rounded pill badges. HTTP-only SameSite=Strict cookie auth. 65KB gzip. Build: cd admin && bun install && bun run build; output at admin/dist/ is committed for self-contained binaries.
  • src/commands/auth.ts — Token management. gbrain auth create/list/revoke/test for legacy bearer tokens (v0.22.7 wired as a first-class CLI subcommand) plus gbrain auth register-client (v0.26.0) and gbrain auth revoke-client <client_id> (v0.26.2) for OAuth 2.1 client lifecycle. revoke-client runs an atomic DELETE...RETURNING on oauth_clients; FK ON DELETE CASCADE on oauth_tokens.client_id and oauth_codes.client_id purges every active token + authorization code in a single transaction. process.exit(1) on no-such-client (idempotent — re-running on the same id produces the same exit-1 message). Legacy tokens stored as SHA-256 hashes in access_tokens; OAuth clients in oauth_clients. As of v0.26.0, legacy tokens grandfather to read+write+admin scopes on the OAuth HTTP server, so pre-v0.26 deployments keep working with no migration. v0.31.3 (#681): every SQL site routes through sqlQueryForEngine(engine) from src/core/sql-query.ts (and executeRawJsonb for the takes-holders permissions JSONB column) so gbrain auth works against PGLite brains. Pre-fix, every call hit the postgres.js singleton via getConn() and silently failed (or wrote to the wrong DB) when the active engine was PGLite. The takes-holders write goes through executeRawJsonb(engine, sql, [name, hash], [{takes_holders:[...]}]) which round-trips with jsonb_typeof = 'object' instead of the pre-v0.31.3 quoted-string shape. v0.34.1.0 (#876): register-client accepts --source <id> (write authority, scalar) and --federated-read <S1,S2,...> (read scope, array). The output prints the resolved Write source and Federated reads for the registered client. Pre-v0.34 clients backfill to source_id='default' via migration v60 so existing deployments keep their v0.33 effective behavior verbatim. v0.42.2.0: the bare gbrain auth create <name> form (no --takes-holders flag) was silently dropping the name and printing usage instead of minting a token. The positional-vs-flag parse is extracted into the exported pure parseAuthCreateArgs(rest) so it's unit-testable — the pre-fix inline version used rest[takesIdx + 1] which resolved to rest[0] (the name) when takesIdx === -1, excluding the name from the positional search. Pinned by test/auth-create-args.test.ts.
  • src/commands/connect.ts + src/core/connect-probe.ts (v0.42.2.0) — gbrain connect <mcp-url> [--token <bearer>] one-command coding-agent onboarding from a bearer token. Turns an MCP URL + token into a paste-ready claude mcp add ... -H "Authorization: Bearer ..." block (default) or, with --install, runs it directly and smoke-tests the token. Direct HTTP MCP — Claude Code talks straight to a remote gbrain serve --http; no local install or thin-client OAuth config needed for the connection itself (the local CLI over a bearer token is deferred T7 in TODOS). Token resolution: --token > $GBRAIN_REMOTE_TOKEN > (print mode) a <paste-your-token> placeholder / (install mode) error. The generated block tells the agent to call get_brain_identity + list_skills (the LEARN_INSTRUCTION export) so it self-orients on what the brain is and can do, with a core-tools fallback list for hosts that haven't enabled skill publishing. URL normalization appends /mcp to a bare host but REJECTS a scheme-less host (so you can't silently point at the wrong thing); pure helpers (isLinkLocalOrMetadata, URL parse, render) are unit-tested. Flags: --token, --name <id> (default gbrain, validated against NAME_RE), --agent claude-code|codex|perplexity|generic, --install, --yes (required for --install in a non-TTY), --force (replace an existing same-name server), --json (token redacted unless --show-token), --timeout-ms. connect is in CLI_ONLY + CLI_ONLY_SELF_HELP (prints its own --help); dispatched in cli.ts:handleCliOnly with no local DB connect (print mode touches nothing; --install talks to the remote). Multi-agent (v0.42.2.0): AGENT_SPECS drives per-agent rendering + --install dispatch. claude-codebuildClaudeMcpAddArgv (literal -H "Authorization: Bearer <tok>"). codexbuildCodexMcpAddArgv = codex mcp add <name> --url <url> --bearer-token-env-var GBRAIN_REMOTE_TOKEN (Codex reads the token from the env var at runtime — never written to its config; --install runs codex mcp add and prints an export GBRAIN_REMOTE_TOKEN hint when the live env doesn't already carry it). perplexity + generic are installable:false (perplexity prints Settings → Connectors GUI steps; generic prints URL + header) and reject --install. OAuth (--oauth, supportsOAuth:true agents = perplexity/generic only): emits an OAuth 2.1 client-credentials connector block (Issuer URL via issuerFromMcpUrl = mcp-url minus /mcp, Client ID, Client Secret) instead of a bearer header — the correct path for a third-party cloud connector (least-privilege scopes + short-lived rotating tokens vs a long-lived full-access secret). Creds from --client-id/--client-secret (BYO, runs anywhere) or --register (host-side: deps.registerOAuthClient shells gbrain auth register-client <name> --grant-types client_credentials --scopes <DEFAULT_SCOPES="read write"> --token-endpoint-auth-method client_secret_post and parses Client ID:/Client Secret:). --oauth is rejected for claude-code/codex and is incompatible with --install. buildJson oauth shape: auth:'oauth', issuer_url, client_id, client_secret (redacted unless --show-token), scopes. Full OAuth chain (register → connect --oauth → OAuth discovery → /token client_credentials mint → get_brain_identity) is E2E-proven in connect-bearer.test.ts (client registered in beforeAll before serve takes the PGLite single-writer lock). cmdString(binary, argv) POSIX-single-quotes args (the codex export <token> line is single-quoted too). buildJson is a generic shape (agent, command/command_argv null for perplexity/generic, header, env_var); the codex command carries only the env-var name, never the token. ConnectDeps = {isTTY, promptYesNo, hasBinary(bin), runBinary(bin, argv), probe, env(name)} — binary-generic so claude and codex share the path; env injectable so tests control GBRAIN_REMOTE_TOKEN. Per-agent docs: docs/mcp/CODEX.md (new), docs/mcp/PERPLEXITY.md (connect fast-path added). Security posture: the rendered command single-quotes the token so shell metacharacters in a token can't run code when pasted; the token is validated before it lands in an HTTP header; link-local / cloud-metadata addresses (incl. IPv4-mapped IPv6 ::ffff:169.254.x.x and AWS IMDSv2-over-IPv6 fd00:ec2::254) are refused as a token-exfil guard while localhost / RFC1918 / LAN stay allowed (self-hosted is a supported topology); the token is redacted from all error output. src/core/connect-probe.ts is the raw-bearer MCP smoke probe backing --install: connects the official MCP SDK Client over StreamableHTTPClientTransport with a STATIC Authorization header (no OAuth, no discovery — distinct from mcp-client.ts:callRemoteTool which is OAuth-only, and from remote-mcp-probe.ts:smokeTestMcp which only sends initialize), runs the full initialize handshake via client.connect(), then calls get_brain_identity (read-scope, non-localOnly) to prove a tool call actually round-trips. Never throws — every failure maps to a discriminated { ok: false, reason: 'auth' | 'unreachable' | 'timeout' | 'tool_error' | 'unknown', message } so the caller renders a precise warning (a wrong/expired token fails at setup, not silently on the agent's first request). DEFAULT_PROBE_TIMEOUT_MS = 15_000 is the single source of truth shared with connect.ts. Pinned by test/connect.test.ts (pure-helper + render coverage, all four agents) + test/e2e/connect-bearer.test.ts: the raw-bearer probe against a live gbrain serve --http AND real-CLI coverage that drives the actual claude + codex binaries through connect --install against the live server (sandboxed HOME / CODEX_HOME, asserts claude mcp get / codex mcp get registered it + the token never lands in Codex config; skips when a binary is absent, e.g. CI). Perplexity is a GUI connector so it's print-asserted only (no CLI to wire E2E). Surfaced in docs/mcp/CLAUDE_CODE.md, docs/mcp/CODEX.md, docs/mcp/PERPLEXITY.md, and the README. v0.42.2.0 follow-up (two-funnel onboarding): LEARN_INSTRUCTION no longer names capture — it's a CLI-only command, NOT an MCP tool, so an agent that followed the block hit "unknown tool"; the list now names put_page (the real MCP write tool), pinned by a connect.test.ts regression block AND by the new test/e2e/serve-stdio-roundtrip.test.ts which spawns real gbrain serve (stdio) against a fresh init --pglite brain and drives the official MCP SDK client through initializetools/listtools/call (get_brain_identity + search), asserting the advertised core-tool set matches what the server exposes (and that capture is NOT advertised) — the local stdio funnel's first e2e coverage. src/commands/serve-http.ts adds an exported pure skillPublishStatus(publishSkills) consumed by the startup banner (Skills: published / not published line) + a one-line stderr nudge (gbrain config set mcp.publish_skills true) when publishing is OFF, so an operator wiring a coding agent to an existing brain sees their skill catalog is invisible to connected agents instead of discovering it via an empty list_skills (pinned by test/serve-skills-publish-nudge.test.ts). The user-facing walkthrough is docs/tutorials/connect-coding-agent.md (Path A: connect to an existing brain; Path B: start from nothing, local stdio) with the brain-first protocol + the four translatable habits; README has a "Quick start: Claude Code or Codex" fork and INSTALL.md shows the one-command wire-up at the standalone-CLI section. The pre-existing test/audit/batch-retry-audit.test.ts ENOENT case was made hermetic (it had read the real ~/.gbrain/audit).
  • src/commands/upgrade.ts — Self-update CLI. runPostUpgrade() enumerates migrations from the TS registry (src/commands/migrations/index.ts) and tail-calls runApplyMigrations(['--yes', '--non-interactive']) so the mechanical side of every outstanding migration runs unconditionally.
  • src/commands/migrations/ — TS migration registry (compiled into the binary; no filesystem walk of skills/migrations/*.md needed at runtime). index.ts lists migrations in semver order. v0_11_0.ts = Minions adoption orchestrator (8 phases). v0_12_0.ts = Knowledge Graph auto-wire orchestrator (5 phases: schema → config check → backfill links → backfill timeline → verify). phaseASchema has a 600s timeout (bumped from 60s in v0.12.1 for duplicate-heavy brains). v0_12_2.ts = JSONB double-encode repair orchestrator (4 phases: schema → repair-jsonb → verify → record). v0_14_0.ts = shell-jobs + autopilot cooperative (2 phases: schema ALTER minion_jobs.max_stalled SET DEFAULT 3 — superseded by v0.14.3's schema-level DEFAULT 5 + UPDATE backfill; pending-host-work ping for skills/migrations/v0.14.0.md). All orchestrators are idempotent and resumable from partial status. As of v0.14.2 (Bug 3), the RUNNER owns all ledger writes — orchestrators return OrchestratorResult and apply-migrations.ts persists a canonical {version, status, phases} shape after return. Orchestrators no longer call appendCompletedMigration directly. statusForVersion prefers complete over partial (never regresses). 3 consecutive partials → wedged → --force-retry <version> writes a 'retry' reset marker. v0.14.3 (fix wave) ships schema-only migrations v14 (pages_updated_at_index) + v15 (minion_jobs_max_stalled_default_5 with UPDATE backfill) via the MIGRATIONS array in src/core/migrate.ts — no orchestrator phases needed. v0.41.37.0 (#1605): new in-process.ts exports runMigrateOnlyCore({timeoutMs?}) — the single source of truth for "bring schema to head" (configureGatewaycreateEngineconnectinitSchemadisconnect, idempotent, 600s wall-clock guard MIGRATE_ONLY_TIMEOUT_MS, throws MigrateOnlyError on no-config / timeout). The migration orchestrators' 9 schema phases AND init.ts:initMigrateOnly (the gbrain init --migrate-only CLI path) both delegate to it, so the schema bring-up can't drift between them. Previously each phase shelled out to a child gbrain init --migrate-only via execSync; on Windows + bun + Supabase pooler the SPAWNED child died with getaddrinfo ENOTFOUND before it could connect (a bun-on-Windows child-process DNS failure, NOT an env-propagation bug — the parent connects fine with env: process.env). Running in-process removes the spawn entirely. runGbrainSubprocess is the diagnostic wrapper for the REMAINING non-schema spawns (extract/repair/stats): captures child stderr (64MB buffer) and folds it into the thrown error so a Windows failure shows the real getaddrinfo ENOTFOUND line instead of a bare Command failed. Pinned by test/migration-in-process.serial.test.ts. v0.41.37.0 (#1581): v0_13_1.ts:phaseCGrandfather rewritten from a per-page getPage + putPage loop (which hung CPU-bound 70+ min on an 82K-page PGLite brain) to a CHUNKED bulk SQL pass. Three correctness properties the per-page loop and a naive single bulk UPDATE both got wrong: (1) keyed on pages.id (globally unique PK), NOT slug — slug uniqueness is (source_id, slug), so a slug-batched UPDATE would mutate same-slug pages across other sources; (2) filters deleted_at IS NULL so tombstones aren't grandfathered; (3) chunked in CHUNK_SIZE batches (DELETE_BATCH_SIZE convention) so lock-hold stays bounded instead of one giant transaction. The rollback log carries source identity ({id, slug, source_id, pre_frontmatter}) so a rollback is unambiguous across sources. Idempotent + resumable (each UPDATE flips its rows out of GRANDFATHER_WHERE). Completes ~1-2s on the 82K-page brain. Pinned by test/migrations-v0_13_1-grandfather.test.ts.
  • src/commands/repair-jsonb.tsgbrain repair-jsonb [--dry-run] [--json]: rewrites jsonb_typeof='string' rows in place across 5 affected columns (pages.frontmatter, raw_data.data, ingest_log.pages_updated, files.metadata, page_versions.frontmatter). Fixes v0.12.0 double-encode bug on Postgres; PGLite no-ops. Idempotent.
  • src/commands/orphans.tsgbrain orphans [--json] [--count] [--include-pseudo] [--source <id>]: surfaces pages with zero inbound wikilinks, grouped by domain. Auto-generated/raw/pseudo pages filtered by default. Also exposed as find_orphans MCP operation. Shipped in v0.12.3 (contributed by @knee5). v0.41.29.0 source-scoping wave: findOrphans/getOrphansData (the canonical pure fn shared with doctor's orphan_ratio) takes { sourceId?, sourceIds? }; BrainEngine.findOrphanPages(opts?) (both engines) scopes ONLY the candidate side (p.source_id = $1 scalar, or = ANY($1::text[]) federated) while still counting inbound links from ANY source — a page in source X linked FROM source Y is reachable, so NOT an orphan of X (the deliberate, less-surprising definition; the stricter intra-source-only reading is rejected). --source is an explicit raw-flag parse (NOT resolveSourceWithTier, which would scope bare invocations to a default). Also corrected the total_linkable denominator: it now enumerates ALL live pages (scoped) and subtracts every excluded-by-slug page (templates/, scratch/, etc.) — pre-fix total - excludedOrphans left excluded NON-orphan pages with inbound links in the denominator, inflating it and suppressing warnings (changes orphan_ratio output for every brain, in the accurate direction). The find_orphans MCP op threads sourceScopeOpts(ctx) so a source-bound OAuth client no longer sees brain-wide orphans (closes a cross-source read leak in the v0.34.1 class). gbrain doctor --source <id> scopes orphan_ratio (entity-count gate + getOrphansData + messages) and, under explicit --source below 100 entity pages, reports the ratio with a low-scale caveat instead of a vacuous "ok". Thin-client doctor --source orphan_ratio remains brain-wide (TODO). Pinned by test/orphans-source-scope.test.ts (PGLite) + test/e2e/engine-parity.test.ts (Postgres↔PGLite scalar + federated parity).
  • src/commands/salience.ts (v0.29) — gbrain salience [--days N] [--limit N] [--kind PREFIX] [--json]: pages ranked by emotional + activity salience over a recency window. Mirrors orphans.ts shape (pure data fn + JSON formatter + human formatter). Calls engine.getRecentSalience(opts). Score formula: (emotional_weight × 5) + ln(1 + active_take_count) + 1/(1 + days_since_update).
  • src/commands/anomalies.ts (v0.29) — gbrain anomalies [--since YYYY-MM-DD] [--lookback-days N] [--sigma N] [--json]: cohort-level activity outliers. Calls engine.findAnomalies(opts). Two cohort kinds in v1: tag, type. Year cohort deferred to v0.30.
  • src/commands/whoknows.ts (v0.33) — gbrain whoknows <topic> [--explain] [--limit N] [--json]: expertise + relationship-proximity routing. Mirrors v0.29 salience/anomalies shape (pure rankCandidates() + findExperts() orchestrator + runWhoknows() CLI dispatch + thin-client routing). MCP op = find_experts (scope: read, localOnly: false) per ENG-D5. Ranking formula (ENG-D1 locked): score = log(1 + raw_match) × max(0.1, exp(-days/180)) × (0.5 + 0.5 × salience) where raw_match is hybridSearch's RRF+source-boost score. Filters at SQL via the new SearchOpts.types: ['person', 'company'] (no post-filter waste). hybridSearch's internal salience+recency boosts are intentionally disabled — the locked formula applies on a clean signal. Floors prevent multiplicative-zero edge cases (cold-start people stay visible); ties break alphabetically by slug for determinism. 16 unit tests in test/whoknows.test.ts pin the math.
  • src/commands/eval-whoknows.ts (v0.33, v0.33.1.3 thin-client wiring) — gbrain eval whoknows <fixture.jsonl> [--json] [--skip-replay]: two-layer eval gate (ENG-D2). Layer 1 quality (hand-labeled fixture, top-3 hit rate ≥ 0.8). Layer 2 regression (eval_candidates replay set-Jaccard@3 ≥ 0.4). Sparseness fallback: < 20 replay-eligible rows → Layer 2 auto-skips with stderr warning. Stable JSON envelope with schema_version: 1. Exit 0/1/2 for pass/fail/usage so CI can gate. Mirrors v0.27.x cross-modal + v0.28.1 longmemeval dispatch shape under src/commands/eval.ts. v0.33.1.3: WhoknowsFn callable abstraction lets the gates be impl-agnostic. runEvalWhoknows(engine: BrainEngine | null, args) picks the impl at entry — thin-client mode (isThinClient(cfg)) routes per-query through callRemoteTool(cfg, 'find_experts', {topic, limit}) via the v0.31.1 seam; local mode calls findExperts(engine, ...) directly. cli.ts adds a thin-client bypass before connectEngine for gbrain eval whoknows, matching the longmemeval/cross-modal no-DB pattern. Regression gate auto-skips in thin-client mode (no DB access to eval_candidates). Public exports jaccardAtK, topKHit, readFixture, WhoknowsFn, threshold constants are pinned by test/eval-whoknows.test.ts (25 cases, +2 for the null-engine signature contract).
  • test/fixtures/whoknows-eval.jsonl (v0.33) — 10-row synthetic placeholder demonstrating the eval-fixture schema ({query, expected_top_3_slugs, notes?} JSONL). End users replace with their own real queries before shipping; the placeholder uses obviously-example slugs (wiki/people/example-alice) so production data isn't conflated with the test fixture. Drives test/e2e/whoknows.test.ts (which seeds a matching synthetic brain and asserts the >=80% gate) and the whoknows_health doctor check.
  • src/core/skillopt/ + src/commands/skillopt.ts + skills/skill-optimizer/ (v0.42.0.0 SkillOpt wave, closes #1481) — self-evolving skill optimization grounded in the SkillOpt paper (arXiv 2605.23904, Microsoft Research, May 2026). Top-level CLI gbrain skillopt <skill> treats SKILL.md as trainable parameters of a frozen agent; validation-gated (D12 median-of-3 + epsilon=0.05), budget-capped (D3 preflight estimator), per-skill DB-locked (D14 tryAcquireDbLock('skillopt:<name>', 60min)), atomic-versioned (D8 history-intent-first 5-step commit), body-only mutations (D5 — frontmatter forbidden). Rollouts use gateway.toolLoop directly with no-op persistence callbacks (D2 — zero subagent_messages pollution) + a read-only tool allowlist derived from BRAIN_TOOL_ALLOWLIST minus put_page/submit_job/file_upload (D13). Two reflect calls per step (D7 paper-faithful). Rejected-edit buffer LRU-bounded to 100 entries. Bundled-skill gate (D16). Bootstrap workflow (D15 sentinel + --bootstrap-reviewed). D17 D_sel floor (>=5 with --split override). Audit JSONL via shared audit-writer.ts. Added to ALL_PHASES after patterns (default OFF; opt-in via gbrain config set cycle.skillopt.enabled true); cycle phase wrapper at src/core/skillopt/cycle-phase.ts walks stale skills with per-skill ($0.50) + brain-wide ($2.00) caps. Added to PROTECTED_JOB_NAMES. Full cathedral ships: F1 dream-cycle phase wrapper, F2 41-case adversarial test suite (6 files), F3 E2E PGLite serial test (3 cases), F4 --all batch mode (src/core/skillopt/batch.ts:runBatchAll), F5 --target-models fleet (runFleet parallel per-model receipts under skillopt/fleet/<slug>/), F6 MCP op run_skillopt (admin scope + per-skill skillopt.allowed_skills allowlist, NOT localOnly), F7 Minion skillopt handler + --background CLI flag with allowProtectedSubmit: true, F8 reflect-prompt quality eval at evals/skillopt-reflect/, F9 judge LLM accuracy eval at evals/skillopt-judge/, F10 write-flavored optimization via src/core/skillopt/write-capture.ts:buildWriteCaptureRegistry (virtual put_page/submit_job/file_upload captured in-memory; --write-capture flag), F11 held-out real-user test set scaffold via src/core/skillopt/held-out.ts (capture infra at ~/.gbrain/skillopt-captures/<skill>/<run>.jsonl, --held-out <path> flag, runHeldOutGate candidate >= baseline). Pinned by 152 tests across 18 files (88 foundation + 41 adversarial + 23 v2 surface + 3 E2E PGLite serial); hermetic via DI seams (opts.chatFn for optimizer + judge; opts.toolLoopFn for rollouts; no mock.module, R2-compliant). Plan + 17 review decisions + outside-voice codex absorption at ~/.claude/plans/system-instruction-you-are-working-drifting-falcon.md. v0.42+ remaining: admin UI Calibration-style dashboard tab for optimizer history; sweep all 47 bundled skills with their own skillopt-benchmark.jsonl fixtures (now unblocked by --bootstrap-from-skill, below). v0.42.1.0 (--bootstrap-from-skill): runBootstrapFromSkill in src/core/skillopt/bootstrap-benchmark.ts is the second bootstrap generator — reads SKILL.md directly (no routing-eval.jsonl) and makes ONE LLM call that emits a full starter benchmark (tasks + rule judges) as JSONL, parsed line-by-line with skip-bad-line salvage (D5 — a truncated final line drops, the rest survive) and a min-2-valid-checks-per-task drop (D6 — never a task judged by one weak check). Provider/transport errors PROPAGATE (not collapsed to bootstrap_empty) so the CLI surfaces the real failure. --bootstrap-tasks N (default 15, CLI-capped at 50) tunes the count; maxTokens scales min(8000, max(4000, N*220)). The generator's stderr REVIEW line prints the literal paste-ready gbrain skillopt <name> --bootstrap-reviewed --split 1:1:1 next command — load-bearing because the optimizer's default 4:1:5 split makes a 15-task starter's D_sel = floor(15/10) = 1, below the >=5 floor, so a 15-task benchmark is unusable without --split 1:1:1. Both bootstrap generators now share extracted helpers assertBenchmarkAbsent + readSkillBodyOrThrow (D1); runBootstrap's routing-eval parse hardened to skip malformed lines. CLI: --bootstrap-from-skill short-circuits after the routing-bootstrap block, mutually exclusive with --bootstrap-from-routing/--benchmark/--all/--target-models/--resume (--background/--follow already rejected by the unknown-flag guard — a pre-existing CLI gap NOT fixed here: those two flags are never parsed so the documented --background path is currently unreachable; filed for a separate fix). parseFlags exported for CLI unit tests. Generated rule judges are explicitly WEAK DRAFTS — skill-optimizer/SKILL.md + docs/tutorials/improving-skills-with-skillopt.md reposition from-skill as the PRIMARY no-benchmark path and tell the agent to STRENGTHEN the judges during review (closes the codex "trains toward benchmark artifacts" concern via the D15 review gate). Pinned by 20 cases in test/skillopt/bootstrap-from-skill.test.ts (happy-path + deterministic task_ids, round-trip into loadBenchmark + splitBench(1:1:1), JSONL salvage, min-2-checks drop, provider-error propagation, bootstrap_empty, benchmark_exists/--force, no_skill_md, fenced output, maxTokens scaling, sub-15 warning + REVIEW line, parseFlags cap + 6-way mutual exclusion). Plan + eng-review + codex absorption at ~/.claude/plans/system-instruction-you-are-working-recursive-sutton.md.
  • src/core/brainstorm/{domain-bank,orchestrator,judges}.ts + src/commands/{brainstorm,lsd,eval-brainstorm}.ts + src/core/last-retrieved.ts (v0.37.0 Open Collider wave) — bisociation-grounded idea generation pair: gbrain brainstorm <question> (defensible, cite-heavy, 4 close × 6 far, judge threshold 4.0/5, save by default) and gbrain lsd <question> (Lateral Synaptic Drift — inverted judge that rejects ideas with resistance >4.5 ("too obvious"), stale-page bias via pages.last_retrieved_at, 2 close × 12 far, axiomatic inversions required, ephemeral by default). The architecture corrects Open Collider's training-data-grounded approach: gbrain has the user's actual cross-domain knowledge already, so the "domain bank" is prefix-stratified sampling from the user's own brain (SELECT DISTINCT substring(slug from '^[^/]+/[^/]+') cached 1h-TTL in config table per source) tiebroken by JOIN page_links connection_count, with corpus-sampling fallback when fewer prefixes than M exist. Distance scores normalized to [0,1] via 1 - clamp(cosine_distance, 0, 2) / 2 (1=opposite, 0=identical). The judge is a single judges.ts with runJudge(config, ideas) + two exported configs (BRAINSTORM_JUDGE_CONFIG weighted originality/resistance/thesis_density/concrete_grounding/cognitive_load 0.25/0.20/0.20/0.20/0.15 vs LSD_JUDGE_CONFIG with cognitive_load 0.50 + inversion rule). Calibration cold-start fallback (D4 + codex #8): when calibration_profiles.active_bias_tags is empty, judge runs without anti-bias context AND stderr-warns. Op-layer write-back at src/core/operations.ts search/query/get_page handlers fires bumpLastRetrievedAt(engine, pageIds) (fire-and-forget, 5-min throttled via SQL clause, default-on with search.track_retrieval config escape hatch per D13) — internal callers (sync, migrations, dream cycle) bypass the op layer so the LSD stale signal stays clean. v0.41.8.0 (#1247/#1269/#1290): the IIFE is now tracked in a module-scoped Set<Promise<unknown>> (mirrors the v0.36.1.x awaitPendingSearchCacheWrites precedent for #1090). New exported awaitPendingLastRetrievedWrites(timeoutMs?: number): Promise<{outcome, pending}> resolves once all tracked promises settle, bounded by a 5s Promise.race timeout that stderr-warns with the pending count if it fires. src/cli.ts awaits the drain unconditionally for every op in the op-dispatch finally block BEFORE engine.disconnect(), then narrows a fallback process.exit(0) to fire ONLY when outcome === 'timeout' AND shouldForceExitAfterMain(argv) (excludes serve so daemons stay alive). Closes the PGLite CLI search/query/get-hang class: pre-fix, the IIFE raced disconnect, PGLite's WASM kept Bun's event loop alive, CLI hung at ~95-98% CPU until SIGKILL. Three community-validated reports; PR #1259 (jehoon) supplied the structural drain pattern; PR #1337 (matt-dean-git) supplied the snapshot+early-null disconnect pattern AND the force-exit guard we narrowed to fire only on the timeout path. Pinned by test/last-retrieved.test.ts (6 unit cases: empty/single/multi-pending settle, throw-in-IIFE still settles, permanently-pending hits timeout, empty pageIds untracked), test/e2e/pglite-cli-exit.serial.test.ts (IRON-RULE behavioral regression: real bun src/cli.ts subprocess against a hermetic PGLite tempdir, asserts search/get/query exit 0 in <15s + daemon-survival case), test/fix-wave-structural.test.ts (behavioral-positioning assertion that the drain await appears textually BEFORE engine.disconnect in the local-engine path — survives variable-rename refactors). Migration v79 adds pages.last_retrieved_at TIMESTAMPTZ NULL + full B-tree index (NOT partial — covers both NULL and range branches per codex r2 #6); full forward-reference bootstrap probe on both engines. Frontmatter mode: lsd makes the dream-cycle synthesize phase skip LSD output (noise-by-design — isLsdOutput() check in src/core/cycle/transcript-discovery.ts short-circuits isDreamOutput()). gbrain eval brainstorm <fixture.jsonl> is a three-axis evaluation gate (distance + usefulness + grounding, conjunctive) per codex r2 #11 — distance alone is gameable. gbrain doctor gains brainstorm_health check surfacing (a) migration v79 applied, (b) search.track_retrieval setting, (c) calibration cold-start status. 38 unit tests across test/brainstorm/{distance,lsd-mode-skip,eval-brainstorm}.test.ts. Plan: ~/.claude/plans/system-instruction-you-are-working-staged-coral.md. Open Collider source: github.com/CL-ML/open-collider. v0.41.21.0: judges.ts replaces maxTokens: 4000 with computeJudgeMaxTokens(ideaCount, modelId). Named constants at top of file: TOKEN_BUDGET_PER_IDEA, TOKEN_BUDGET_ENVELOPE, LEGACY_MIN_MAX_TOKENS, MAX_OUTPUT_TOKENS_CEIL. New ANTHROPIC_OUTPUT_CAPS map (Opus 4.7 = 32K, Sonnet 4.6 / Haiku 4.5 = 64K, legacy Claude 3.5 = 8K) so legacy 8K-cap models bind at 8K instead of failing mid-judge. When the caller passes no modelOverride, the cap routes through the gateway's actual configured chat model via getChatModel() so the formula matches what chat() will use (not whatever a stale override hints at). Closes the headline v0.41.21.0 bug: pre-fix, a 72-idea brainstorm chunked the judge into ~3 calls of ~24 ideas each; each call needed ~7.2K output tokens; the hard-coded 4K cap truncated every call mid-JSON; the parser threw; the whole run came back judge_failed: true with 0/72 scored. Post-fix: same fixture returns ~39/72 passing. Pinned by 16 cases in test/brainstorm/judges-maxtokens.test.ts. v0.41.30.0: --save for both brainstorm and lsd rewritten to persist through the canonical ingestion path instead of the lightweight DB-only write that never produced the advertised .md file. New exported persistSavedIdea(engine, {slug, content, provenanceVia}) calls importFromContent({noEmbed:true, sourcePath}) (chunked + tagged + content_hash so gbrain search finds the idea; no embedding cost at save time) THEN renders the saved row to disk via the shared writePageThrough helper (see src/core/write-through.ts) — file rendered FROM the row so the two sinks can't diverge and a later gbrain sync doesn't churn it. New exported pure formatSaveOutcome(outcome, ctx) returns an honest per-branch message: both-sinks success, DB-only (no sync.repo_path / repo-not-a-dir → file write skipped, exit 0), DB-saved-but-file-errored (exit 0, sync reconciles), and the total-failure case (neither sink → loud save FAILED … NOT persisted on stderr + nonzero exit). Closes the silent-false-success bug class where --save printed "Saved" unconditionally even when the PgBouncer-transaction-mode DB write failed and gbrain get <slug> then returned page_not_found. buildIdeaSlug(question, label, nonce?) gains a random nonce suffix (injectable for deterministic tests) so two same-day runs whose questions share the first 60 slug chars no longer clobber each other's page + file. --json callers stay DB-only (unchanged). New exported buildBrainstormFrontmatterObject(result) in orchestrator.ts returns the object form for serializeMarkdown (the string buildBrainstormFrontmatter is left untouched). Pinned by test/brainstorm/save.test.ts (covers persistSavedIdea both-sinks / DB-only / total-failure paths, formatSaveOutcome per-branch messages + exit codes, and buildIdeaSlug nonce collision-resistance).
  • src/core/write-through.ts (v0.41.30.0, NEW) — shared atomic disk write-through for the canonical ingestion path. writePageThrough(engine, slug, {sourceId?, frontmatterOverrides?, logger?}) reads sync.repo_path, re-reads the just-written DB row (getPage), renders it via serializePageToMarkdown + resolvePageFilePath, and writes the .md under the repo so the brain has a committable artifact that round-trips through gbrain sync. Rendering FROM the row means the file and the DB row cannot diverge. Extracted from the v0.38 put_page write-through in src/core/operations.ts (which the op now calls instead of hand-rolling its own copy) AND upgraded to ATOMIC: writes to a unique temp sibling (<file>.tmp.<pid>.<rand>) + renameSync into place, cleaning up the temp on any failure, so a crash or a concurrently-running gbrain sync/autopilot walking the live git tree can never read a half-written .md (matches the .tmp + rename convention in import-checkpoint.ts / op-checkpoint.ts). Never throws — returns WriteThroughResult { written, path?, skipped?: 'no_repo_configured' | 'repo_not_found' | 'page_not_found_after_write', error? } so the caller decides messaging + exit codes. Trust gating (subagent sandbox, dry-run) stays at the CALLER. Consumers: put_page op (write-through behavior preserved, now atomic) and gbrain brainstorm/lsd --save via persistSavedIdea. Pinned by test/write-through.test.ts (no-repo / repo-not-found / page-not-found-after-write skips, atomic-rename happy path + no stray .tmp on write failure, frontmatter-override stamping).
  • src/core/model-id.ts (v0.41.21.0, NEW) — splitProviderModelId(input: string | null | undefined): {provider: string | null, model: string} shared parser for the pricing side. Splits on : first, then /. Defensive contract: null/undefined/empty/whitespace returns {provider: null, model: ''}. Five sites consume it (src/core/anthropic-pricing.ts:estimateMaxCostUsd, src/core/budget/budget-tracker.ts:lookupPricing, src/core/eval-contradictions/cost-tracker.ts:pricingFor, src/core/minions/batch-projection.ts at two call sites, src/core/model-config.ts:isAnthropicProvider) so the pricing + classification surface no longer has 5 parallel re-implementations of provider:model splitting. Closes the bug class where slash-form ids (anthropic/claude-sonnet-4-6 — the form CLI flags accept and OpenRouter recipes emit) silently fell through to "unknown model" at every site. Distinct from the existing gateway-side parseModelId in src/core/ai/model-resolver.ts: that one throws on bare names because routing needs an explicit provider; this one returns {provider: null, model: 'bare'} because pricing lookups happen against bare model ids. Pinned by 16 cases in test/model-id.test.ts.
  • src/core/ai/model-resolver.ts:parseModelId extension (v0.41.21.0) — gateway-side resolver extended to accept slash form alongside colon. Pre-fix the colon-only check threw AIConfigError: model id must be in format provider:model at every gateway entry point (chat / embed / rerank) the moment a slash-form id was passed. So even with the v0.41.21.0 pricing fix, a --judge-model anthropic/claude-sonnet-4-6 invocation would clear BudgetTracker but then fail mid-judge inside gateway.chat(). Now both shapes resolve to the same recipe. Bare names without ANY separator still throw — gateway routing always needs an explicit provider. Pinned by 10 cases in test/ai/model-resolver-slash.test.ts including a resolveRecipe round-trip asserting slash form resolves to the same recipe object as colon form.
  • src/commands/transcripts.ts (v0.29) — gbrain transcripts recent [--days N] [--full] [--json]: recent raw .txt transcripts from the dream-cycle corpus dirs. Imports listRecentTranscripts from src/core/transcripts.ts (the same library the gated get_recent_transcripts MCP op uses). Local-only by construction — the CLI always runs with ctx.remote=false.
  • src/commands/integrity.tsgbrain integrity check|auto|review|extract: bare-tweet detection, dead-link detection, three-bucket repair (auto-repair / review-queue / skip). scanIntegrity() is the shared library function called from gbrain doctor (sampled at limit=500) and cmdCheck (full scan). v0.22.8: batch-load fast path on Postgres uses a single SQL query to fix the PgBouncer round-trip timeout (60s → ~6s). Gated by engine.kind === 'postgres' at the call site so PGLite never enters batch; fallback catch logs at GBRAIN_DEBUG=1 so real Postgres errors are diagnosable. v0.32.8 (PR #860): batch projection switched from SELECT DISTINCT ON (slug) to SELECT ... ORDER BY source_id, slug so multi-source brains scan each (source, slug) row independently (pre-fix the DISTINCT collapsed same-slug-different-source pages into one scan, the same bug class this PR fixes). Sequential and auto-repair loops use listAllPageRefs() to enumerate (slug, source_id) pairs and thread sourceId to getPage. Batch + sequential paths now report the same page count on multi-source brains.
  • src/commands/doctor.tsgbrain doctor [--json] [--fast] [--fix] [--dry-run] [--index-audit]: health checks. v0.12.3 added jsonb_integrity + markdown_body_completeness reliability checks. v0.14.1: --fix delegates inlined cross-cutting rules to > **Convention:** see [path](path). callouts (pipes DRY violations into src/core/dry-fix.ts); --fix --dry-run previews without writing. v0.14.2: schema_version check fails loudly when version=0 (migrations never ran — the #218 bun install -g signature) and routes users to gbrain apply-migrations --yes; new opt-in --index-audit flag (Postgres-only) reports zero-scan indexes from pg_stat_user_indexes (informational only, no auto-drop). v0.15.2: every DB check is wrapped in a progress phase; markdown_body_completeness runs under a 1s heartbeat timer so 10+ min scans are observable on 50K-page brains. v0.19.1 added queue_health (Postgres-only) with two subchecks: stalled-forever active jobs (started_at > 1h) and waiting-depth-per-name > threshold (default 10, override via GBRAIN_QUEUE_WAITING_THRESHOLD). Worker-heartbeat subcheck intentionally deferred to follow-up B7 because it needs a minion_workers table to produce ground-truth signal. Fix hints point at gbrain repair-jsonb, gbrain sync --force, gbrain apply-migrations, and gbrain jobs get/cancel <id>. v0.22.12 (#500): sync_failures check shows [CODE=N, ...] breakdown for both unacked entries (warn) and acked-historical entries (ok), surfacing systemic failure modes (SLUG_MISMATCH=2685) instead of a bare count. v0.26.7 (#612): rls_event_trigger check (post-install drift detector for migration v35's auto-RLS event trigger). Lives outside the // 5. RLS slice that the structural doctor.test.ts guards anchor on, so the existing test guards stay intact. Healthy evtenabled set is ('O','A') only — R is replica-only and would not fire in normal sessions; D is disabled. Fix hint is gbrain apply-migrations --force-retry 35. v0.30.2: queue_health gains a fourth subcheck — surfaces dead-lettered subagent jobs with last_error matching the prompt_too_long classifier within the last 24h. Fix hint points at gbrain dream --phase synthesize --dry-run --json to identify the offending transcript and gbrain jobs prune --status dead --queue default to clean up. Postgres-only. v0.31.7: runDoctor switches to autoDetectSkillsDirReadOnly (from src/core/repo-root.ts) so bun install -g github:garrytan/gbrain && cd ~ && gbrain doctor finds the bundled skills/ via the install-path fallback instead of warning "Could not find skills directory" + docking the health score. --fix carries a D6 safety gate: when detected.source === 'install_path', the command refuses auto-repair with a stderr message pointing at $GBRAIN_SKILLS_DIR / $OPENCLAW_WORKSPACE / --skills-dir, because autoFixDryViolations writes to SKILL.md files and would otherwise silently rewrite the install tree. The graph_coverage check now short-circuits to ok: 'No entity pages — graph_coverage not applicable (markdown-only brain)' when SELECT COUNT(*) FROM pages WHERE type IN ('entity','person','company','organization') returns 0 (closes #530); the entity count is woven into the warn message and the WARN hint switches from the long-deprecated gbrain link-extract && gbrain timeline-extract (gone since v0.16) to the canonical gbrain extract all. Pinned by an IRON-RULE regression assertion in test/doctor.test.ts that bans the stale verb names from the source string. v0.32.4: new sync_freshness check (exported checkSyncFreshness at the same file) added to both runDoctor (local) and doctorReportRemote (thin-client). Pure staleness probe — queries sources.last_sync_at only, no filesystem access. Warns at 24h, fails at 72h (or never-synced). Future-last_sync_at warns ("clock skew or corrupted timestamp") instead of silently falling through as ok — codex outside-voice caught the negative-ageMs bug pre-merge. Env-var overrides GBRAIN_SYNC_FRESHNESS_WARN_HOURS / GBRAIN_SYNC_FRESHNESS_FAIL_HOURS; invalid values fall back to defaults with a once-per-process stderr warn (_resolveSyncFreshnessHours). Failure messages embed source.id (not source.name) so the printed fix command gbrain sync --source <id> matches what the user copy-pastes. Filesystem-vs-DB page drift detection was deliberately stripped from the v0.32.4 scope — doctorReportRemote runs in the HTTP MCP server (src/commands/serve-http.ts), and walking DB-supplied local_path from a remote-callable endpoint crosses a trust boundary (OAuth write scope could mutate sources.local_path). Drift detection will resurface in a separate PR routed through multi_source_drift's existing guard infrastructure (GBRAIN_DRIFT_LIMIT / GBRAIN_DRIFT_TIMEOUT_MS) with slug normalization tests and a meta-file allow-list. Pinned by 12 cases in test/doctor.test.ts ("v0.32.4 — sync_freshness check" describe block): empty sources, never-synced fail, >72h fail, exact 72h boundary, 24h-72h warn, exact 24h boundary, <24h ok, future-timestamp warn, mixed sources (highest severity wins), executeRaw throws → outer-catch warn, GBRAIN_SYNC_FRESHNESS_FAIL_HOURS=6 override fires at 7h, source.id-in-message regression. v0.36.3.0: new embedding_column_registry check probes each declared column via Postgres format_type(atttypid, atttypmod) so a registry entry claiming 1024d Voyage against an actual 1536d OpenAI column surfaces with a paste-ready gbrain config set embedding_columns '{...}' ALTER hint instead of mysterious "vector dimension mismatch" errors at search time. On Postgres the check also probes HNSW index presence (pg_indexes lookup keyed by column name) and warns when missing (search will still work via seq scan but won't hit the index). The active default column's population coverage is computed via COUNT(*) FILTER (WHERE <col> IS NOT NULL) / COUNT(*) and warns below 90% — except empty brains (chunk_count = 0) where the gate short-circuits to ok so fresh gbrain init runs don't see "Active column 'embedding' is 0.0% populated" (CDX-5 codex fix). PGLite parity via the same SQL through executeRaw — registry validation happens on both engines. v0.35.5.0: the Lane D supervisor check at doctor.ts:1011-1043 now consumes summarizeCrashes(events) from src/core/minions/handlers/supervisor-audit.ts instead of the pre-fix events.filter(e => e.event === 'worker_exited').length. The warn threshold drops from >3 to >=1 (any real crash is signal now that the counter is calibrated against clean exits). The ok message gains clean_exits_24h=N; the warn message gains runtime=A oom=B unknown=C legacy=D per-cause breakdown so an operator triages OOM vs runtime-error vs unknown-future-cause at a glance without grep'ing the JSONL audit. Closes the "Supervisor crashes: 120x/24h, was 62x — nearly doubled" alarm class that bit users on healthy brains after v0.34.3.0's RSS-watchdog work added more code=0 worker drains — both doctor and gbrain jobs supervisor status were counting every worker_exited event as a crash regardless of cause. Cross-surface parity is the regression guard: 4 source-grep wiring assertions in test/doctor.test.ts ban the ad-hoc filter pattern, pin the >=1 threshold, and require the per-cause breakdown substrings (runtime=, oom=, unknown=, legacy=, clean_exits_24h=, crashes_by_cause) to appear in BOTH doctor.ts and jobs.ts. v0.37.1.0: new skill_brain_first check. Walks every SKILL.md under the configured skills dir (autoDetectSkillsDirReadOnly so cd ~ && gbrain doctor finds the bundled skills via the install-path fallback), calls analyzeSkillBrainFirst() from src/core/skill-brain-first.ts per file, aggregates verdicts into a single check with structured Check.issues[] for JSON tooling. Warn states: missing_brain_first (external-lookup pattern present, no canonical callout, no brain_first: exempt), brain_first_typo (near-miss declaration like brain-first or BrainFirst — paste-ready hint surfaces the correct snake_case form). Ok states: compliant_callout, compliant_phase, compliant_position, exempt_frontmatter, no_external. --fix routes through dry-fix.ts MISSING_RULE_PATTERNS to auto-insert the canonical > **Convention:** see [conventions/brain-first.md](...) callout (D6 install-path safety gate enforced — --fix from ~ refuses to write to the bundled tree). Snapshot+diff audit at ~/.gbrain/audit/skill-brain-first-YYYY-Www.jsonl records detected / resolved / fixed transitions only (stable brains: 0 lines/run). Motivated by the 2026-05-19 tweet-shield incident: cross-modal eval flagged Garry's Palantir tweet as risky because no model knew he built it, but the brain already had "designed the entire Finance product UI" and "150+ PSDs from April-December 2006." Static check catches the AUTHORSHIP miss class; v0.37+ runtime gate (filed in TODOS.md) closes the dispatch side. v0.41.27.0: checkSyncFreshness gains a localOnly-gated git short-circuit (D4 trust-boundary preservation per Codex P0-1): runDoctor passes localOnly: true, doctorReportRemote keeps the default false so the HTTP MCP path doesn't walk DB-supplied local_path values via subprocess. The narrowed predicate (D7) mirrors sync.ts:1057+1075's actual "do work?" gate: HEAD == last_commit AND working tree clean AND sources.chunker_version === String(CHUNKER_VERSION). Inline SELECT widens to carry last_commit + chunker_version (columns already exist; no schema migration). Three-bucket count math (D6) populates Check.details = {unchanged_count, synced_recently_count, stale_count} with the invariant sum === sources.length pinned in the test suite. OK message reshape (D2): all-unchanged → "All N up to date (no new commits since last sync)"; mixed → "N source(s): X synced recently, Y unchanged since last sync"; all-synced-recently keeps the prior message verbatim. checkCycleFreshness is INTENTIONALLY NOT touched (Codex P0-2 / D5): last_commit == HEAD answers "are there new commits to sync?" but cannot answer "did the full cycle complete?" — a sync can succeed while later cycle phases fail, and silencing that warn would hide real cycle staleness. Pinned by 9 new cases in test/doctor.test.ts ("v0.41.27.0 — sync_freshness git short-circuit" describe block) including a load-bearing D4 regression guard that verifies probes are NEVER called when localOnly is unset or false, plus the three-bucket invariant explicitly asserted in the mixed 3-source case. Supersedes PR #1564 (Co-Authored-By preserved). v0.41.32.0 (supersedes #1623): checkSyncFreshness short-circuit now passes requireCleanWorkingTree: 'ignore-untracked' (was true) — the headline fix: a quiet repo whose only dirt is untracked dirs is unchanged, not SEVERE. The SELECT widens to carry newest_content_at; the REMOTE (non-localOnly) path computes lag via lagFromContentMs(newest_content_at, lastSync, now) from the stored column — NO git subprocess on a DB-supplied local_path (trust boundary intact). LOCAL fall-through and the < 0 clock-skew check stay on raw wall-clock (A1). checkCycleFreshness deliberately NOT content-relativized (CM2 — different axis, last_full_cycle_at; filed in TODOS as a probe-phase follow-up). Pinned by the "v0.41.32.0 — commit-relative staleness" describe in test/doctor.test.ts (T1 untracked-folders headline bug + T2/T2b/T2c remote-never-shells-out trust boundary).
  • src/core/migrate.ts — schema-migration runner. Owns the MIGRATIONS array (source of truth for schema DDL). v40 (v0.29): pages_emotional_weight adds pages.emotional_weight REAL NOT NULL DEFAULT 0.0. Column-only (no index). On Postgres 11+ and PGLite, ADD COLUMN with a constant DEFAULT is metadata-only — instant on tables of any size. v0.14.2 extended the Migration interface with sqlFor?: { postgres?, pglite? } (engine-specific SQL overrides sql) and transaction?: boolean (set to false for CREATE INDEX CONCURRENTLY, which Postgres refuses inside a transaction; ignored on PGLite since it has no concurrent writers). Migration v14 (fix wave) uses a handler branching on engine.kind to run CONCURRENTLY on Postgres (with a pre-drop of any invalid remnant via pg_index.indisvalid) and plain CREATE INDEX on PGLite. v15 bumps minion_jobs.max_stalled default 1→5 and backfills existing non-terminal rows. v0.22.6.1: migration v24 (rls_backfill_missing_tables) uses sqlFor: { pglite: '' } to no-op on PGLite — PGLite has no RLS engine and is single-tenant by definition, and the v24 ALTERs target subagent tables that don't exist in pglite-schema.ts. Closes #395 (contributed by @jdcastro2). v30 (v0.23): creates dream_verdicts (file_path TEXT, content_hash TEXT, worth_processing BOOL, reasons JSONB, judged_at TIMESTAMPTZ, PK(file_path, content_hash)). RLS-enabled when running as a BYPASSRLS role. The synthesize phase reads/writes this table to avoid re-judging on backfill re-runs. v35 (v0.26.7): auto-RLS event trigger + one-time backfill. auto_rls_on_create_table fires on ddl_command_end for WHEN TAG IN ('CREATE TABLE','CREATE TABLE AS','SELECT INTO') and runs ALTER TABLE … ENABLE ROW LEVEL SECURITY on every new public.* table — no FORCE (matches v24/v29/schema.sql posture so non-BYPASSRLS apps can still read their own tables). The same migration backfills RLS on every existing public.* base table whose comment doesn't match the doctor regex (^GBRAIN:RLS_EXEMPT\s+reason=\S.{3,}). Per-table failure aborts the offending CREATE TABLE (event triggers fire inside the DDL transaction); no EXCEPTION wrap — that would convert loud rollback into silent permissive default. PGLite no-op via sqlFor.pglite: ''. Breaking change: operators with intentionally-RLS-off public tables must add the GBRAIN:RLS_EXEMPT comment BEFORE upgrade or the backfill will flip them on. v46 (v0.31.3): mcp_request_log_params_jsonb_normalize rewrites pre-v0.31.3 rows where mcp_request_log.params was stored as a JSON-encoded string (jsonb_typeof = 'string') up to a real JSONB object via UPDATE ... SET params = params::text::jsonb WHERE jsonb_typeof(params) = 'string'. Single statement, idempotent — second-run finds no string-shaped rows and is a no-op. Closes the bug where /admin/api/requests returned a quoted string instead of the parsed object. v0.36.3.0 (v68): eval_candidates_embedding_column adds eval_candidates.embedding_column TEXT NULL. Per-row provenance for gbrain eval replay: captured rows record which column the live query ran against so replay reproduces the same retrieval space (Voyage rows replay against Voyage; OpenAI rows against OpenAI). NULL-tolerant — pre-v0.36 rows fall back to the current default during replay rather than failing. Column-only migration, metadata-only on both engines. v0.34.1.0 (#861 + #876, v60-v65): six-migration chain wires source-scoping into the OAuth client table. v60 (oauth_clients_source_id_fk) adds oauth_clients.source_id TEXT with NULL→'default' backfill and an FK to sources(id) ON DELETE SET NULL. v61 (oauth_clients_federated_read_column) adds federated_read TEXT[] NOT NULL DEFAULT '{}'. v62 (oauth_clients_federated_read_backfill) explicit-CASE backfills so source_id IS NULL produces '{}' not an array-containing-NULL. v63 (oauth_clients_federated_read_validate) is the fail-loud check that every row's source_id is in its federated_read array post-backfill. v64 (oauth_clients_source_id_fk_restrict) flips the FK to ON DELETE RESTRICT now that federated_read provides the alternative scope-loss path — source delete is refused if any client references it. v65 (oauth_clients_federated_read_gin_index) is the GIN index for the array-containment queries the read paths run. PGLite parity via sqlFor.pglite where needed. v108 (v0.41.31.0): pages_embedding_signature adds pages.embedding_signature TEXT NULL = <provider:model>:<dims> stamped when a page's chunks are embedded (setPageEmbeddingSignature). A later model/dimension swap makes the stored signature differ from the current one so countStaleChunks/sumStaleChunkChars (with the signature opt) and invalidateStaleSignatureEmbeddings can detect and re-embed those pages. GRANDFATHER (critical): the stale predicate is embedding_signature IS NOT NULL AND embedding_signature <> $current, so a NULL signature is NEVER stale — after this migration every existing page has NULL, none are flagged, and the next embed --stale does NOT re-embed the whole corpus. Signatures only get stamped going forward. No index (read only via a JOINed pages row in the chunk-grain stale queries; no standalone lookup hot path). ADD COLUMN with no DEFAULT (NULL) is metadata-only on Postgres 11+ / PGLite 17.5. v109 (v0.41.32.0, supersedes #1623): sources_newest_content_at adds sources.newest_content_at TIMESTAMPTZ — durable newest-COMMIT timestamp (HEAD committer time) written at sync time by writeSyncAnchor. The REMOTE staleness path (federation_health, get_status_snapshot MCP op) reads it instead of shelling out to git on a DB-supplied local_path. Renumbered 108→109 on the master merge that landed v0.41.31 pages_embedding_signature at v108. Metadata-only ADD COLUMN; mirror in pglite-schema.ts + schema.sql + bootstrap probe coverage. v0.41.34.0 (retrieval-maxpool incident, v110 + v111): v110 (page_aliases) creates the free-text alias table — (id, source_id, alias_norm, slug, ...) with UNIQUE (source_id, alias_norm, slug) and lookup indexes on (source_id, alias_norm) + (source_id, slug). alias_norm is the output of normalizeAlias() so the WRITE side (ingest projection) and READ side (search) key on the same normalized form. Also lands in src/core/pglite-schema.ts for fresh PGLite installs. v111 (search_telemetry_rank1_columns) is ADD COLUMN IF NOT EXISTS on both engines for the rank-1 base_score drift signal: sum_rank1_score, count_rank1, and three coarse distribution buckets (rank1_lt_solid / rank1_solid / rank1_high) on the existing search_telemetry rollup — aggregate columns (NOT per-query rows, D10) so a downward drift in the median rank-1 match score is computable with bounded growth. search_telemetry lives only in migration v57, so these ALTER right after v57 runs.
  • src/core/progress.ts — Shared bulk-action progress reporter. Writes to stderr. Modes: auto (TTY: \r-rewriting; non-TTY: plain lines), human, json (JSONL), quiet. Rate-gated by minIntervalMs and minItems. startHeartbeat(reporter, note) helper for single long queries. child() composes phase paths. Singleton SIGINT/SIGTERM coordinator emits abort events for every live phase. EPIPE defense on both sync throws and stream 'error' events. Zero dependencies. Introduced in v0.15.2. v0.40.3.0: emitHumanLine is prefix-aware — when called inside a withSourcePrefix(id, ...) scope from src/core/console-prefix.ts, prepends [id] to the line content. TTY-rewrite mode (\r\x1b[2K) gets the prefix inside the clear-to-EOL escape so the rewritten line carries the prefix too. JSON mode (emitJson) is intentionally NOT prefixed — NDJSON consumers parse the JSON envelope and would choke on a [id] {...} shape.
  • src/core/console-prefix.ts (v0.40.3.0) — AsyncLocalStorage<string>-backed per-source line-prefix helper. Exports withSourcePrefix(id, fn) (runs fn with id as the active prefix; nested wraps replace the active prefix and restore on exit), getSourcePrefix() (read-only accessor; test seam), and slog(...) / serr(...) (prefix-aware console.log / console.error replacements). Embedded-newline-safe: a multi-line string under prefix [foo] emits [foo] line1\n[foo] line2. Outside a wrap, slog / serr fall through to bare console.log / console.error so single-source callers see identical output (back-compat invariant). Use src.id (slug-validated by sources add) NOT src.name (free-form text) to defeat log-injection through newline / control-character names. Coverage in v0.40.3.0: src/commands/sync.ts performSync + in-file callees (38 call sites migrated), src/commands/embed.ts runEmbedCore + helpers (16 call sites), src/core/progress.ts emitHumanLine. Lines emitted from outside these modules will NOT get the prefix under parallel sync — file an issue if you find a missed migration target.
  • src/core/cli-options.ts — Global CLI flag parser. parseGlobalFlags(argv) returns {cliOpts, rest} with --quiet / --progress-json / --progress-interval=<ms> stripped. getCliOptions() / setCliOptions() expose a module-level singleton so commands reach the resolved flags without parameter threading. cliOptsToProgressOptions() maps to reporter options. childGlobalFlags() returns the flag suffix to append to execSync('gbrain ...') calls in migration orchestrators. OperationContext.cliOpts extends shared-op dispatch for MCP callers.
  • src/core/db-lock.ts (v0.22.13) — generic tryAcquireDbLock(engine, lockId, ttlMinutes) over the existing gbrain_cycle_locks table. Parameterized lock id so different scopes can nest cleanly: gbrain-cycle for the broad cycle (held by cycle.ts) and gbrain-sync (SYNC_LOCK_ID constant) for performSync's narrower writer window. Same UPSERT-with-TTL semantics as the prior cycle-only helper, just generalized. Survives PgBouncer transaction pooling (unlike session-scoped pg_try_advisory_lock); crashed holders auto-release once their TTL expires.
  • src/core/sync-concurrency.ts (v0.22.13) — single source of truth for the parallel-sync policy. Exports autoConcurrency(engine, fileCount, override?) (PGLite always serial; explicit override clamped to >=1; auto path returns DEFAULT_PARALLEL_WORKERS=4 when fileCount > AUTO_CONCURRENCY_FILE_THRESHOLD=100), shouldRunParallel(workers, fileCount, explicit) (Q1: explicit --workers bypasses the >50-file floor), and parseWorkers(s) (rejects '0', '-3', 'foo', '1.5', trailing chars — replaces the prior parseInt-with-no-validation in both sync.ts and import.ts). Used by performSync, performFullSync, runImport, and the Minion sync handler so the three sites can no longer drift. v0.40.3.0: new sibling constant DEFAULT_PARALLEL_SOURCES = 4 for the per-source fan-out under gbrain sync --all. Kept SEPARATE from DEFAULT_PARALLEL_WORKERS because they cover two different axes: total live Postgres connections per fan-out wave is approximately DEFAULT_PARALLEL_SOURCES × DEFAULT_PARALLEL_WORKERS × 2 (per-worker pool) = 32 connections with both at the default 4. Conflating them as one constant was a v0.40.2 footgun that Codex flagged during the v0.40.3.0 plan review; keeping them separate makes the multiplication visible at the constant level. The × 2 per-file pool factor is because each per-file worker opens its own PostgresEngine with poolSize = min(2, resolvePoolSize(2)). sync.ts emits a stderr warning when parallel × workers × 2 > 16 so operators size pgbouncer / Postgres max_connections accordingly. v0.41.17.0 (T2, D9): new sibling resolveWorkersWithClamp(engine, override, commandName, fileCount) wrapper composes autoConcurrency with a per-command stderr clamp warning on PGLite. NOT a modification to shared autoConcurrency itself (which stays silent + used by sync/import). The wrapper is the canonical surface for every bulk-command --workers N flag landed in v0.41.17.0 (extract-conversation-facts, extract, edges-backfill, reindex-multimodal, reindex, reindex-code). embed.ts deliberately bypasses the wrapper and keeps GBRAIN_EMBED_CONCURRENCY || 20 (codex #13). Per-(command, requested) stderr dedup via module-scoped warned-once set with _resetWorkersClampWarningsForTest() seam. Pinned by 12 cases in test/pglite-workers-clamp.test.ts.
  • src/core/worker-pool.ts (v0.41.17.0 NEW, T1) — Canonical sliding-pool + bounded-semaphore primitive. Extracted from src/commands/embed.ts (both sliding-pool sites) and src/commands/eval-cross-modal.ts (runWithLimit semaphore). Two exports: runSlidingPool<T>({items, workers, onItem, signal?, onError?, failureLabel?, onProgress?}) + runWithLimit<TIn, TOut>({items, limit, fn, signal?}). Atomicity invariant (const idx = nextIdx++ is one synchronous JS statement — no await between read and write — guaranteed by JS single-threaded event loop) documented in module header AND enforced by scripts/check-worker-pool-atomicity.sh (D5) wired into bun run verify. Guard rejects two failure modes: importing worker_threads in any file that imports the helper (would cross kernel threads + break the event-loop guarantee), and inserting await between the nextIdx read and write inside the helper. MUST_ABORT_ERROR_TAGS set (D13) seeded with BUDGET_EXHAUSTED from src/core/budget/budget-tracker.ts; tagged errors bypass the helper's onError and hard-abort the pool, propagating an AbortController.abort() to in-flight workers' onItem. Budget cap becomes a structural ceiling under concurrency, not a per-caller convention. Property-tag match (err.tag === 'BUDGET_EXHAUSTED') avoids cross-module import. failures[] shape (D7 + codex #10): {idx, label, error} records, NOT full items; callers supply failureLabel(item) => string projector. Bounded memory under 197K-page brains. Pinned by 23 cases in test/worker-pool.test.ts + 9 cases in test/scripts/check-worker-pool-atomicity.test.ts (on the CI guard itself). The shared helper drives every --workers N bulk command in v0.41.17.0.
  • src/commands/embed.ts extension (v0.41.17.0, T3 REGRESSION per IRON RULE) — both inline sliding-pool sites (embedAll simple at :458-467 and embedAllStale paginated + AbortSignal at :586-632) replaced with calls to runSlidingPool from the shared worker-pool helper. Byte-equivalent contract preserved at the invariant level (counts + cost + AbortSignal propagation + per-batch rate-limit retry firing through embedBatchWithBackoff); per codex #16/#17 byte-equality on progress event ORDERING is NOT promised. Existing GBRAIN_EMBED_CONCURRENCY || 20 default preserved per codex #13 — embed bypasses the new resolveWorkersWithClamp wrapper because its 20-worker default pre-dates autoConcurrency and routing through it would silently change every existing brain's embed hot path. Pinned by 8 structural cases in test/embed-helper-migration.test.ts (asserts helper is wired in AND pre-migration let nextIdx = 0 + Promise.all(Array.from({length: numWorkers}, ...)) shapes are gone).
  • src/commands/extract-conversation-facts.ts extension (v0.41.17.0, T5) — --workers N for LLM-bound fact extraction over conversation pages. Combined with the per-page advisory lock from src/core/db-lock.ts:withRefreshingLock (D2 + D12 — lock id extract-conversation-facts:<source>:<slug>, TTL PER_PAGE_LOCK_TTL_MINUTES=2 with 20s refresh via Math.max(15s, 120s/6); LockUnavailableError triggers D6 skip-and-continue with rate-limited log per (source, minute) + pages_lock_skipped counter in result + CLI exits 3 when non-zero AND no hard failures), the deleteOrphanFactsForPage(engine, sourceId, slug) delete-orphans-first replay safety (D11 — wipes any facts left by a prior crashed/killed run for this (sourceId, slug) before re-extracting; closes the "terminal audit row written after partial insertFacts failure" bug class codex caught in eng review), and the assertFactsEmbeddingDimMatchesConfig(engine) startup preflight (D15 — throws FactsEmbeddingDimMismatchError with paste-ready ALTER hint BEFORE the first fact insert; cached per engine via WeakMap). Result type extended with pages_lock_skipped + orphan_facts_cleaned counters. Checkpoint state migrated from per-page-mutated cpEntries: string[] array to shared cpMap: Map<slug, endIso> so JS-single-thread atomic Map.set survives parallel workers (codex #5/#6 fix). Minion handler extract-conversation-facts in src/commands/jobs.ts round-trips workers via job.data.workers for --background --workers 20. Cycle phase cycle.conversation_facts_backfill.workers config key (default 1; opt-in concurrency for cycle paths under brain-wide cost + walltime caps). Pinned by 17 cases in test/extract-conversation-facts-workers.test.ts + 27 existing extract-conversation-facts behavioral tests still green.
  • src/core/embedding-dim-check.ts extension (v0.41.17.0, T5+T6) — facts.embedding dim drift surface. New readFactsEmbeddingDim(engine): Promise<FactsColumnDimResult> covers both vector(N) and halfvec(N) shapes (codex #19 — migration v40 falls back to vector on pgvector < 0.7). Regex ordering halfvec-before-vector pinned by tests (substring "vec" appears in "halfvec"; naive /vector/i would shadow). New buildFactsAlterRecipe(dims, configured, type) emits the paste-ready DROP INDEX IF EXISTS idx_facts_embedding_hnsw; ALTER TABLE facts ALTER COLUMN embedding TYPE halfvec(N) USING embedding::halfvec(N); CREATE INDEX idx_facts_embedding_hnsw ON facts USING hnsw (embedding halfvec_cosine_ops) WHERE ... flow (codex #18 — NOT bare REINDEX which doesn't rewrite the index after a column-type change). New assertFactsEmbeddingDimMatchesConfig(engine) is the D15 preflight — throws FactsEmbeddingDimMismatchError (tagged class with tag: 'FACTS_EMBEDDING_DIM_MISMATCH' for parity with the worker-pool's MUST_ABORT semantics) when configured dim doesn't match the column width. Result cached per-engine via WeakMap. PGLite engines silently skip. New doctor check facts_embedding_width_consistency (registered in runDoctor after embedding_width_consistency) reuses the same helpers — surfaces drift with paste-ready ALTER recipe identical to the preflight error. Pinned by 18 cases in test/embedding-dim-check-facts.test.ts.
  • src/core/postgres-engine.ts extension (v0.41.17.0, T6, codex #20) — insertFact + insertFacts no longer hardcode tx.unsafe(\'${embedLit}'::vector`)for the embedding column. NewresolveFactsEmbeddingCast()private method probespg_attributeonce per engine instance (cached in_factsEmbeddingCastSuffix) and returns '::halfvec'when migration v40 created the column as halfvec, else'::vector'. Both insert paths use the cached suffix so the cast matches the actual column type. Pre-fix all three insert sites hardcoded ::vector; works on pgvector >= 0.7 via implicit auto-cast but fails on older pgvector. Test seam __resetFactsEmbeddingCastCacheForTest()` clears per-engine cache.
  • src/core/cycle.ts + src/core/cycle/extract-atoms.ts + src/core/cycle/synthesize-concepts.ts + src/commands/extract.ts + src/commands/doctor.ts + src/core/op-checkpoint.ts extensions (v0.41.20.0 ops-fix-wave) — six daily-driver ops pains in one bisectable wave. (1) Batch idempotency for extract_atoms: new atomsExistingForHashes(engine, sourceId, hashes[]) exported from src/core/cycle/extract-atoms.ts replaces the per-hash loop that did 7K individual queries at the start of every cycle on brains with conversation-transcript corpora (5-10 min silent overhead). One batched SQL roundtrip returns the set of content_hash16 values already extracted as atoms for this source. Fail-open: SQL error logs to stderr and returns an empty set so extraction proceeds. Powered by migration v104 pages_atom_source_hash_idx (partial expression index on frontmatter->>'source_hash' for atom rows where deleted_at IS NULL; Postgres uses CREATE INDEX CONCURRENTLY with invalid-remnant pre-drop mirroring v97 pages_dedup_partial_index, PGLite uses plain CREATE INDEX). Without this index the batch idempotency check would seq-scan pages on big brains and defeat the perf win. (2) Shorter cycle lock TTL + active in-phase refresh: LOCK_TTL_MINUTES dropped from 30 → 5 in src/core/cycle.ts. New exported buildYieldDuringPhase(lock, outer) closure calls lock.refresh() AND any external hook on every fire; throttled to 30s inside each phase via maybeYield. Fires both inside the main work loop AND immediately after every await chat(...) LLM call so long Haiku/Sonnet calls don't sit past TTL. LockHandle and buildYieldDuringPhase exported for test seam access. synthesize_concepts no longer fires yieldDuringPhase per-concept-group — same hook, throttled via the new shared maybeYield helper. A crashed cycle releases its lock 6x faster while a healthy long-running cycle keeps the lock alive. Known residual under shorter TTL: a single await chat() past 5 min wallclock can expire the lock mid-await without the original phase noticing — DbLockHandle.refresh() throwing on 0 rows affected is filed as P2 follow-up TODO-OPS-2. (3) Progress wiring through long phases: new progress?: ProgressReporter opt on ExtractAtomsOpts and SynthesizeConceptsOpts. Cycle.ts passes its phase-level reporter down (NOT a child reporter — that would produce a path collision cycle.extract_atoms.extract_atoms.work). Phases only call tick() and heartbeat(); cycle.ts owns start() and finish(). You see [cycle.extract_atoms] N (atoms_created) ticks every ~1s during both long phases instead of "start" then silence for 10+ min. (4) by-mention resume: new mentionsFingerprint({source, type, since, gazetteerHash}) in src/core/op-checkpoint.ts. The gazetteer hash is the load-bearing field — adding new entity pages mid-pause shifts the hash, gets a new fingerprint, and triggers a fresh scan against the new gazetteer instead of silently skipping previously-scanned pages. gbrain extract links --by-mention now resumes from where it died via the existing op_checkpoints framework with a flushAndCheckpoint ordering — links flush to the DB FIRST, page keys commit to the checkpoint SECOND, persist THIRD. A crash between batch.push() and flush leaves the page un-checkpointed so resume re-scans it. Persist cadence: every 1000 items OR every 30s, whichever first. Clean exit clears the checkpoint. --dry-run deliberately skips both load and write so it stays an inspection mode. (5) sync_consolidation doctor check: multi-source brains see a paste-ready gbrain sync --all --parallel 4 --workers 4 --skip-failed recommendation. Single-source brains get "not applicable." SQL errors return warn via the check's own try/catch — outer doctor catch isn't a safe assumption. Companion "Multi-source brains" recipe block in skills/cron-scheduler/SKILL.md documents the sync --all pattern as preferred over per-source entries. (6) Test-isolation fixes: test/cycle-last-full-cycle-at.test.ts + test/schema-cli.test.ts migrated to per-test GBRAIN_HOME=tempdir isolation. Pinned by 44 new unit/PGLite cases across 9 files: test/cycle/extract-atoms-batch.test.ts (5 — batch idempotency), test/cycle/cycle-lock-ttl.test.ts (1 — regression pin on LOCK_TTL_MINUTES === 5), test/op-checkpoint-mentions-fingerprint.test.ts (7 — fingerprint sensitivity including gazetteer-hash regression guard), test/cycle/extract-atoms-progress.test.ts (4 — phase doesn't call start/finish, ticks fire per item), test/cycle/synthesize-concepts-progress.test.ts (3 — same shape), test/cycle/yield-during-phase-refresh.test.ts (7 — buildYieldDuringPhase actually calls lock.refresh() + outer hook, throws non-fatal), test/cycle/yield-during-phase-throttle.test.ts (3 — 30s throttle gate), test/extract-by-mention-resume.test.ts (5 — checkpoint persistence ordering, dry-run skips persist, gazetteer change invalidates, filtered pages get checkpointed), test/doctor-sync-consolidation.test.ts (6 — edge case matrix for source counts + archived filtering + SQL error path). Two follow-up TODOs filed in TODOS.md under "v0.41.19.0 ops-fix-wave follow-ups": gbrain sync print-cron subcommand (TODO-OPS-1) and lock-loss detection via DbLockHandle.refresh() throwing on 0 rows affected (TODO-OPS-2).
  • src/commands/sync.tsgbrain sync CLI + the performSync / performFullSync library entrypoints (consumed by the autopilot cycle and the Minion sync handler). v0.22.13 (PR #490): performSync wraps its body in a gbrain-sync writer lock so two concurrent syncs (manual + autopilot, two terminals, two Conductor workspaces) cannot both write last_commit and let the last writer win. Head-drift gate after the import phase re-checks git rev-parse HEAD; if HEAD moved (someone ran git checkout / git pull mid-sync), the bookmark refuses to advance. Vanished files now record a failedFiles entry instead of silent-skip — the silent-skip-then-advance pathology that survived prior hardening passes is dead. Worker engines wrap in try/finally so disconnect always fires (panic-path leak fix). Both PGLite-detection sites use engine.kind === 'pglite'. CLI accepts --workers N (alias --concurrency N), validated via parseWorkers. Explicit --workers bypasses the auto-path file-count floor; auto path defers to autoConcurrency(). Banner moved to stderr. v0.34.2.0: the inline .sort() over add/mod paths is replaced with sortNewestFirst(addsAndMods) from src/core/sort-newest-first.ts, so the newest-first descending-lex policy lives in one helper shared with gbrain import instead of drifting across two files. v0.40.3.0 (productionized from PR #1314 by @garrytan-agents): the load-bearing change is the lock-identity invariant — performSync now defaults to a per-source lock id (gbrain-sync:<sourceId>) whenever opts.sourceId is set, AND wraps the writer window in withRefreshingLock from src/core/db-lock.ts so long-running sources (the PR motivator: media-corpus / 250K+ chunks) don't lose their lock at the 30-min TTL mid-run. Legacy single-default-source brains keep the bare tryAcquireDbLock(SYNC_LOCK_ID) path for back-compat. SyncOpts.lockId?: string is the explicit override (escape hatch; production callers don't set it). The fix closes the bug class Codex's outside-voice review caught in the original PR — a sync --all worker on per-source lock racing against sync --source foo on global lock would have silently corrupted the same source. gbrain sync --all got a continuous worker pool: parseWorkers-validated --parallel N (default min(sourceCount, --workers, DEFAULT_PARALLEL_SOURCES=4) from src/core/sync-concurrency.ts), parallel long-lived async workers pulling from a shared FIFO queue (no head-of-line blocking from wave dispatch), per-source withSourcePrefix(src.id, ...) wrap from src/core/console-prefix.ts so every slog/serr line carries [<source-id>] for kubectl-style greppable parallel output. --skip-failed and --retry-failed reject with a paste-ready error when combined with --parallel > 1 (the brain-global sync-failures.jsonl has no per-source scope; v0.41+ TODO source-scopes it). Connection-budget stderr warning fires when parallel × workers × 2 > 16 (the × 2 per-file pool factor: each per-file worker opens its own PostgresEngine with poolSize=2, so --parallel 4 --workers 4 is actually 32 connections, not 16 — Codex caught the 2× understatement in the original plan). New exports resolveParallelism, syncOneSource, buildSyncStatusReport, printSyncStatusReport, SyncStatusReport back the gbrain sources status dashboard surface (D3 → sources subcommand, not a sync flag, so reads and writes don't share a verb). Stable --json envelope {schema_version: 1, sources, parallel, ok_count, error_count} on stdout under --json; human banners route to stderr via the humanSink helper so jq parses cleanly. Exit matrix: 0 = all ok, 1 = any error, 2 = cost-prompt-not-confirmed (unchanged). The dashboard SQL is the canonical content_chunks ch JOIN pages pg ON pg.id = ch.page_id WHERE pg.deleted_at IS NULL shape with archived = false filter at the caller — the original PR shipped chunks ch JOIN ON page_slug, which would have crashed on PGLite parse and silently zeroed on Postgres via the swallow-catch. Embedding column resolved via resolveEmbeddingColumn(undefined, cfg) from src/core/search/embedding-column.ts so Voyage / multimodal / non-default-column brains see counts against the column they actually use. Errors propagate from the dashboard SQL — no swallow-catch (Q2 sub-fix). The IRON RULE regression lives in test/e2e/sync-status-pglite.test.ts: real PGLite seeds 2 sources × pages × chunks, soft-deletes 1 page, archives 1 source, validates the SQL excludes both AND the active embedding column is the one used. 38 in-file console.log/console.error call sites inside performSync* migrated to slog/serr (top-level runSync orchestrator console calls intentionally stay outside the prefix scope). v0.41.25.0 (supersedes PR #1538): sync delete loop rewritten as interleaved per-batch resolve+delete using engine.resolveSlugsByPaths + engine.deletePages from src/core/engine.ts. 73K-delete commit drops from ~146K SQL round-trips (~5h on Postgres+pgbouncer) to ~292 round-trips (~2min) — closes the cascade-staleness bug class where one big-delete commit jammed every other source's sync for hours. Per-batch try-catch decomposes batch DELETE failures to per-slug deletePage fallback so a transient blip on batch 73 doesn't lose 500 deletes; unrecoverable per-slug failures land in failedFiles (matches the existing import-loop pattern). pagesAffected filters to D6 confirmed-deleted slugs (downstream extract/embed stop wasting work on phantoms). Same batched slug-resolve treatment for the rename loop (gbrain sync renames — Phase 3 in the plan; the per-file importFile stays per-file, only the slug-resolve N+1 gets batched). The resolveSlugByPathOrSourcePath single-call helper at sync.ts:267 now delegates to engine.resolveSlugsByPaths when sourceId is set (D8 DRY — one owner of the SQL + fallback semantics), keeping legacy executeRaw fallback for the no-sourceId path which is functionally dead post-v0.34.1's source-resolution wiring. Legacy no-sourceId branch in the delete loop preserves correctness on a slow path; the engine batch methods require sourceId so the new fast path only fires when production threading is intact. failedFiles declaration hoisted to top of performSyncInner so both delete decompose and import loops feed the same sync-bookmark gate. v0.41.31.0 (mode-aware cost gate): the v0.20.0 unconditional sync --all ConfirmationRequired gate is superseded by a mode-aware gate driven by willEmbedSynchronously + shouldBlockSync from src/core/embedding.ts. federated_v2 is resolved ONCE up front (shared with the fan-out below). On the DEFERRED path (v2 on, parallel) embedding goes to per-source embed-backfill jobs that carry their own $X/source/24h cap (default $25, SPEND_CAP_CONFIG_KEY from embed-backfill-submit.ts), so the gate is INFORMATIONAL — it prints an FYI naming the backfill cap + current backlog (sumStaleChunkChars({signature}) priced via estimateCostFromChars at currentEmbeddingPricePerMTok()) + the count of already-queued embed-backfill jobs (TODO-2 visibility), and NEVER exits 2. On the INLINE path (v2 off, or --serial without --no-embed) the BLOCKING gate fires only when the NEW-CONTENT estimate exceeds the sync.cost_gate_min_usd floor (default $0.50). The new-content estimate (estimateInlineNewTokens) contributes ZERO for sources provably unchanged since last sync (isSourceUnchangedSinceSync HEAD==last_commit + clean tree, AND chunker_version === CURRENT) — mirrors doctor's sync_freshness + sync's own do-work gate — and the full-tree ceiling for changed sources. F2 fix: the pre-existing stale backlog (NULL embeddings + signature drift) is shown informationally but NEVER gated on, because sync doesn't sweep it inline (gbrain embed --stale does); gating on it would block every inline cron after a model swap. New helpers resolveCostGateFloorUsd(engine) (reads sync.cost_gate_min_usd, fail-open $0.50, accepts 0) + resolveBackfillCapUsd(engine). The sync --all SELECT widens to carry last_commit + chunker_version (both columns predate v0.41; no migration). JSON envelopes gain mode + gate discriminators (dry_run | deferred_notice | below_floor | confirmation_required). SyncStatusReportSource gains backfill_queued / backfill_active / backfill_last_completed_at (best-effort; 0/null on pre-minions brains). All embedding-cost preview math reads the configured model name via getEmbeddingModelName() (no hardcoded OpenAI). Pinned by test/sync-cost-gate.serial.test.ts + test/sync-cost-preview.test.ts. v0.41.37.0 (#1569): new SyncOpts.noSchemaPack (CLI --no-schema-pack, threaded through performSync AND syncOneSource so sync --all honors it) — skips loadActivePack entirely so no user-supplied pack page-type regex runs during import; pages fall back to legacy prefix typing. Escape hatch for completing a sync when a suspect pack regex is the suspected cause of a wedge (re-run extraction later). Same release adds a per-file BEGIN heartbeat: if (process.env.GBRAIN_SYNC_TRACE) serr('[sync] begin import: <path>') emitted BEFORE importFile (the existing progress.tick fires only AFTER importFile returns — useless when one file wedges). Off by default to avoid a line per file on huge brains; serr is source-prefix-aware so under --workers >1 / --all the stuck file is the begin-line with no matching completion. Triage doc: docs/architecture/serve-sync-concurrency.md (PGLite single-writer serve↔sync contention — stop gbrain serve before a large PGLite sync — plus the GBRAIN_SYNC_TRACE + --no-schema-pack hang-triage recipes).
  • src/commands/import.tsgbrain import CLI + runImport library entrypoint. v0.34.2.0 replaces the prior positional-index checkpoint (processedIndex: N into a sorted file list) with a path-set checkpoint via src/core/import-checkpoint.ts. The walk still applies sortNewestFirst() for embed-cost ordering, but checkpoint correctness no longer depends on sort order. A file enters completed: Set<relativePath> only when its processFile returns success (including content-hash short-circuit no-ops); failed files never enter the set, so the next run retries them automatically with no manual ~/.gbrain/import-checkpoint.json delete. Three bug classes died: parallel-import-with-slow-worker drops the slow file on crash-resume (closed — the slow file isn't in completed until its own processFile resolves), failed-file-bumps-counter-past-itself (closed — failures don't add to completed), and v0.33.x sort-flip-drops-newest-N-on-cross-version-resume (closed — order is no longer part of the checkpoint). Old positional checkpoints are detected and discarded with a stderr line on first resume; re-walking is cheap because content_hash short-circuits unchanged files. Checkpoint persists every 100 successful adds, not every 100 processed files, so a long failure tail doesn't churn the JSON. Pinned by test/import-checkpoint.test.ts (18 unit cases over the helpers) + test/import-resume.test.ts (5 integration cases under PGLite, including the SLUG_MISMATCH retry regression codex caught during plan-eng-review).
  • src/core/import-checkpoint.ts (v0.34.2.0) — loadCheckpoint(brainDir), saveCheckpoint(brainDir, completed), resumeFilter(files, completed, brainDir), clearCheckpoint(), plus the ImportCheckpoint type. Path-set checkpoint format ({schema_version, brainDir, completed: string[]}) replaces the v0.33.x positional {processedIndex: N} format. Atomic write via .tmp + rename() so a mid-write crash never leaves a partial JSON. loadCheckpoint returns null on: missing file, malformed JSON, brainDir mismatch (you ran import against a different brain), and the old positional format (logged to stderr before being discarded). resumeFilter returns {toProcess, skippedCount} — pure, no I/O, deterministic. clearCheckpoint is no-op-on-missing for clean-exit cleanup. Honors GBRAIN_HOME via gbrainPath() so test isolation via withEnv({GBRAIN_HOME: tmpdir}) works without monkey-patching the fs layer. Best-effort persistence — saveCheckpoint logs warnings on write errors but never throws, so import keeps making progress even if disk is full.
  • src/core/sort-newest-first.ts (v0.34.2.0) — single source of truth for the descending-lex sort that gbrain import and gbrain sync both apply. Mutates in place (Array.prototype.sort semantics), returns the same array reference for fluent chaining. Empty/single-element inputs short-circuit. Future ordering changes flip one line in this helper instead of touching two CLI commands. Pinned by test/sort-newest-first.test.ts (5 hermetic cases: descending order, mixed prefixes, empty input, single-element input, in-place-mutation contract).
  • src/core/cycle.ts — v0.17 brain maintenance cycle primitive (extended to 9 phases in v0.29). runCycle(engine: BrainEngine | null, opts: CycleOpts): Promise<CycleReport> composes phases in semantically-driven order: lint → backlinks → sync → synthesize → extract → patterns → recompute_emotional_weight → embed → orphans. v0.29 adds the recompute_emotional_weight phase between patterns and embed; it sees the union of syncPagesAffected + synthesizeWrittenSlugs for incremental mode, or all pages when neither anchor is set (full backfill via gbrain dream --phase recompute_emotional_weight). v0.29 also extends CycleReport.totals with pages_emotional_weight_recomputed (additive, schema_version stays "1"). v0.23's synthesize phase runs after sync (cross-references see fresh brain) and before extract (auto-link materializes its writes); patterns runs after extract so it reads a fresh graph (codex finding #7 — subagent put_page sets ctx.remote=true and skips auto-link/timeline by default; extract is the canonical materialization). Three callers: gbrain dream CLI, gbrain autopilot daemon's inline path, and the Minions autopilot-cycle handler. Coordination via gbrain_cycle_locks DB table + ~/.gbrain/cycle.lock file lock with PID-liveness for PGLite. CycleReport.schema_version: "1" is stable; totals additively grew in v0.23 (transcripts_processed, synth_pages_written, patterns_written). yieldBetweenPhases runs between phases. v0.23 added yieldDuringPhase for in-phase keepalive — synthesize/patterns call it during long waits to renew the cycle-lock TTL. Engine nullable; lock-skip on read-only phase selections. v0.22.1 (#403): CycleOpts.signal?: AbortSignal propagates the worker's abort signal; checkAborted() fires between every phase. v0.22.1 (#417): runPhaseSync returns pagesAffected via SyncPhaseResult; runCycle captures it and threads to runPhaseExtract as the 4th arg. v0.22.1 (Codex F2): runPhaseSync takes willRunExtractPhase: boolean and sets noExtract: phases.includes('extract') so gbrain dream --phase sync doesn't silently lose extraction. v0.22.5 (#475): resolveSourceForDir(engine, brainDir) threads sourceId to performSync() so sync reads the per-source sources.last_commit anchor instead of the drift-prone global config.sync.last_commit key. v0.41.38.0 (postgres support): CycleOpts.brainDir is now string | null; when null (checkout-less postgres/Supabase brain) the 6 filesystem phases (lint/backlinks/sync/synthesize/extract/patterns) skip with details.reason: 'no_brain_dir' and the DB-only phases run. resolveSourceForDir is null-tolerant (returns undefined). cycleSourceId = opts.sourceId ?? resolveSourceForDir(engine, brainDir) is the canonical per-source scope for extract_facts/extract_atoms/calibration — so gbrain dream --source repo-a reconciles repo-a's facts even with no checkout (pre-fix they scoped to 'default' while the cycle stamped repo-a fresh — a freshness stamp that lied). deriveStatus counts edges_resolved/edges_ambiguous as work so an edges-only cycle reports ok not clean. The jobs.ts autopilot-cycle + phase-wrapper handlers pass null (not cwd '.') when no repo is configured. Pinned by test/dream-postgres.serial.test.ts + test/jobs-autopilot-cycle-braindir.serial.test.ts.
  • src/core/cycle/synthesize.ts (v0.23) — Synthesize phase: conversation-transcript-to-brain pipeline. Reads from dream.synthesize.session_corpus_dir, runs cheap Haiku verdict (cached in dream_verdicts), then fans out one Sonnet subagent per worth-processing transcript with allowed_slug_prefixes (sourced from skills/_brain-filing-rules.json dream_synthesize_paths.globs). Orchestrator collects slugs from subagent_tool_executions (NOT pages.updated_at — codex finding #2) and reverse-renders DB → markdown via serializeMarkdown. Cooldown via dream.synthesize.last_completion_ts, written ONLY on success. Idempotency key dream:synth:<file_path>:<content_hash>. Auto-commit deferred to v1.1 (codex #5). --dry-run runs Haiku, skips Sonnet (codex #8). Subagent never gets fs-write access. v0.23.2: renderPageToMarkdown (now exported) stamps dream_generated: true and dream_cycle_date into every reverse-write's frontmatter; writeSummaryPage does the same on the dream-cycle summary index. The marker is the explicit identity surface checked by isDreamOutput in transcript-discovery.ts — replaces the v0.23.1 content-prefix heuristic that could miss real output (serializeMarkdown doesn't embed slugs in body) and false-positive on user transcripts citing brain pages. judgeSignificance and JudgeClient are exported; judgeSignificance accepts a verdictModel parameter (default claude-haiku-4-5-20251001) loaded from dream.synthesize.verdict_model via loadSynthConfig. v0.30.2: model-aware chunker splitTranscriptByBudget(content, contentHash, maxChars) splits oversized transcripts at paragraph boundaries (## Topic:---\n ladder) using a deterministic offset seeded from the first 32 bits of contentHash so retries chunk identically. Per-chunk char budget computed from MODEL_CONTEXT_TOKENS[resolvedModel] × 0.9 × 3.5 chars/token; non-Anthropic ids fall back to a 180K-token safe default with a once-per-process stderr warning. Operator overrides: dream.synthesize.max_prompt_tokens (floor 100K, wins when set) and dream.synthesize.max_chunks_per_transcript (default 24). Per-chunk idempotency keys dream:synth:<filePath>:<hash16>:c<i>of<n>; single-chunk transcripts preserve the legacy dream:synth:<filePath>:<hash16> key byte-for-byte (D8 lookup), so existing brains skip with already_synthesized_legacy_single_chunk instead of re-spending Sonnet on upgrade. collectChildPutPageSlugs raw-fetches every (job_id, slug) pair (not SELECT DISTINCT) and rewrites bare-hash6 slugs to <hash6>-c<idx> for chunked children (D6 — orchestrator-side, zero Sonnet trust). Cap-hit skips don't write to dream_verdicts, so raising the cap on next run re-attempts cleanly. D7 scope: bounds INITIAL prompt size only; tool-loop turn-N accumulation is caught by the v0.30.2 terminal-error classification in subagent.ts, not bounded ahead of time. v0.41+ community-PR-wave (reworked PR #1349): makeHaikuClient() replaced by gateway-routed makeJudgeClient(verdictModel) exported helper — mirrors tryBuildGatewayClient in src/core/think/index.ts:579-637 (the v0.35.5.0 #952 pattern). Construction-time provider/key probe returns null on a clear miss (unknown provider id via resolveRecipe AIConfigError, or Anthropic provider with no key via hasAnthropicKey() matching the same env + config-file pattern think uses). Verdict loop wraps judgeSignificance in try/catch for AIConfigError so mid-run provider failures (revoked key, recipe misconfig) surface as per-transcript worth=false, reasons=['gateway error: ...'] instead of crashing the phase. Canonical config key models.dream.synthesize_verdict (per PER_TASK_KEYS in src/core/model-config.ts); contributor's parallel DEEPSEEK_API_KEY fetch adapter dropped in favor of recipe-registered providers (DeepSeek, OpenRouter, Voyage, Ollama, llama-server, ...). JudgeClient interface signature preserved verbatim for test-seam stability. CI guard at scripts/check-gateway-routed-no-direct-anthropic.sh prevents future contributors from re-introducing new Anthropic() here or in think/index.ts. Pinned by 13 cases in test/cycle/synthesize-gateway-adapter.test.ts (A1-A9 unit + R3 parsed-verdict semantic parity + R3 corollary) and 1 E2E case in test/e2e/dream-synthesize-pglite.test.ts ('gateway-adapter mid-run AIConfigError catch') plus 4 R1-R4 critical regression pins in test/cycle/regression-pr-wave-r1-r2-r4.test.ts. v0.41.5.0: narrow anthropic: prefix fix at the queue.add boundary (lines 395-404). resolveModel returns bare ids from TIER_DEFAULTS/DEFAULT_ALIASES (e.g. claude-sonnet-4-6); the subagent validator requires provider:model form and was rejecting with unknown provider, dropping synthesize to status: fail with SYNTH_PHASE_FAIL. Conditional prefix at the call site (only when no colon AND starts with claude-) avoids changing the shared constants which would ripple across every resolveModel caller.
  • scripts/check-gateway-routed-no-direct-anthropic.sh (v0.41+ community-PR-wave) — CI guard that fails the build if src/core/cycle/synthesize.ts or src/core/think/index.ts reintroduces a runtime new Anthropic() constructor call or a value-shaped import Anthropic from '@anthropic-ai/sdk' import. Type-only imports (import type Anthropic from '@anthropic-ai/sdk') stay allowed for adapter types. Comment lines (// or * prefixes) are excluded so historical JSDoc references don't false-fire. Mirrors scripts/check-jsonb-pattern.sh. Wired into bun run verify and bun run check:all. Extend GUARDED_FILES in the script when migrating another file off direct SDK construction.
  • src/core/cycle/patterns.ts (v0.23) — Patterns phase: cross-session theme detection over reflections within dream.patterns.lookback_days (default 30). Names a pattern only when ≥dream.patterns.min_evidence (default 3) reflections support it. Single Sonnet subagent; same allow-list path as synthesize. Runs AFTER extract so the graph is fresh.
  • src/core/cycle/extract-facts.ts (v0.32.2, extended v0.35.6.0) — extract_facts cycle phase. v0.32.2 contract: fence is canonical; per-page wipe (deleteFactsForPage) + reinsert from parseFactsFence + extractFactsFromFenceText + engine.insertFacts. Empty-fence guard refuses when v0.31 legacy rows (row_num IS NULL AND entity_slug IS NOT NULL) pend the v0_32_2 backfill (status: warn, hint: gbrain apply-migrations --yes). v0.35.6.0 adds a phantom-redirect pre-pass that runs AFTER the legacy-row guard, BEFORE the main reconcile loop. When opts.brainDir is set, runPhantomRedirectPass(engine, brainDir, sourceId, dryRun) walks unprefixed-slug pages capped by GBRAIN_PHANTOM_REDIRECT_LIMIT (default 50). The pass returns touched_canonicals — canonical slugs whose disk fence was merged with phantom rows; runExtractFacts UNIONs them into the main reconcile slug set so canonical's DB facts derive from the merged fence in the same cycle (round-14 scenario-B fix: phantom had only-on-disk fence, no DB facts). ExtractFactsResult gains six phantom fields: phantomsScanned, phantomsRedirected, phantomsAmbiguous, phantomsSkippedDrift, phantomsLockBusy, phantomsMorePending. Three of those bubble to CycleReport.totals (phantoms_redirected, phantoms_ambiguous, phantoms_skipped_drift).
  • src/core/entities/resolve.ts (v0.30+, extended v0.35.6.0) — Free-form entity name → canonical slug resolution. resolveEntitySlug(engine, source_id, raw): exact slug → fuzzy (pg_trgm @ 0.4 threshold) → bare-name prefix expansion (people/<token>-% then companies/<token>-% using correlated-subquery connection_count for tiebreaker) → deterministic slugify fallback. v0.35.6.0 exports two new helpers for the phantom-redirect pass: resolvePhantomCanonical(engine, sourceId, phantomSlug) — variant that SKIPS the exact-slug step (codex #1: phantom slug 'alice' exact-matches itself, would make the redirect handler a no-op); returns the canonical only when result is non-null AND contains /. findPrefixCandidates(engine, sourceId, token) — standalone SQL query returning ALL candidates across PREFIX_EXPANSION_DIRS (currently hardcoded ['people', 'companies']) using slug LIKE ANY($N::text[]) over patterns dir/token + dir/token-%; cap of 10 ordered by connection_count DESC, slug ASC. NOT a wrapper around tryPrefixExpansion because that path returns per-dir top-1 and suppresses ambiguity by design (codex #11). Pinned by test/phantom-redirect.test.ts resolvePhantomCanonical describe (3 cases) + findPrefixCandidates describe (6 cases including multi-dir ambiguity and the people/aliceberg-doesn't-match-alice false-positive guard).
  • src/core/cycle/phantom-redirect.ts (v0.35.6.0) — Phantom-redirect orchestrator. Exports runPhantomRedirectPass(engine, brainDir, sourceId, dryRun): Promise<PhantomPassResult> (the per-cycle wrapper that acquires gbrain-sync writer lock once for the entire pass, 30s bounded retry, walks up-to-GBRAIN_PHANTOM_REDIRECT_LIMIT unprefixed phantoms) + tryRedirectPhantom(engine, page, sourceId, brainDir, dryRun): Promise<RedirectResult> (single-phantom orchestrator) + stripFenceAndFrontmatterAndLeadingH1 (pure helper for the body-shape gate — strips facts fence including the preceding ## Facts heading and the leading H1; zero residue = phantom). Handler order: body-shape gate → resolvePhantomCanonical (codex #1: bypasses exact-self-match) → findPrefixCandidates ambiguity check (codex #11: standalone query, not per-dir top-1) → fenceDbDrift bi-directional check (rounds 27/29/30) → dry-run early exit → materialize canonical via serializeMarkdown if DB-only (codex #6) → append phantom fence rows to canonical's disk fence with (claim, valid_from) dedup-guard + row_num continuation → engine.refreshPageBody with SHA-256 content_hash recomputed via the import-file shape (codex #7) → engine.migrateFactsToCanonical (codex #3/#4/#12 lossless preservation) → engine.rewriteLinks (DB FK rewrite; wiki-link text rewrite is a documented follow-up per codex #5) → engine.softDeletePage + engine.deleteFactsForPage(phantom) + fs.unlinkSync(phantomPath) (rounds 19/20). RedirectResult.canonical populated on outcome 'redirected' (incl. dry-run preview) so the caller can populate touched_canonicals. Idempotent on re-run: phantom soft-deleted → predicate fails (deleted_at IS NULL filter); migrate UPDATE matches no rows; dedup-guard prevents double-append.
  • src/core/facts/phantom-audit.ts (v0.35.6.0) — JSONL audit at ${resolveAuditDir()}/phantoms-YYYY-Www.jsonl. Pattern copy of src/core/audit-slug-fallback.ts (ISO-week rotation, honors GBRAIN_AUDIT_DIR). Exports logPhantomEvent(record) + readRecentPhantomEvents(days) + computePhantomAuditFilename(now?). Records every outcome: redirected | ambiguous | drift | no_canonical | not_phantom_has_residue | pass_skipped_lock_busy. Best-effort writes — stderr warn on failure, never throws. Separate file from stub-guard-audit.ts because the consumer + lifecycle are distinct (stub-guard logs PREVENTIVE blocks; phantom-audit logs CLEANUP decisions, will be read by a future T9 doctor phantoms_pending check).
  • src/core/cycle/emotional-weight.ts (v0.29) — Pure function computeEmotionalWeight({tags, takes}, {highEmotionTags?, userHolder?}). Deterministic 0..1 score: tag-emotion boost (max 0.5, case-insensitive match against HIGH_EMOTION_TAGS seed list), take density (0.1/take, capped at 0.3), take avg weight (0..0.1), user-holder ratio (0..0.1 over active takes; default holder = 'garry'). Total clamped to [0..1]. Anglocentric / personal-life-biased seed list intentional; users override via config key emotional_weight.high_tags (JSON array). userHolder overridable via emotional_weight.user_holder.
  • src/core/cycle/anomaly.ts (v0.29) — Pure stats helpers for find_anomalies. meanStddev returns sample stddev (n-1 denominator) and (0,0) for empty input. computeAnomaliesFromBuckets(baseline, today, sigma, limit) takes densified daily-count buckets + today's counts per cohort, returns AnomalyResult[]. Zero-stddev fallback: cohort fires when count > mean + 1, with sigma_observed = count - mean as a finite sort proxy (no NaN). Brand-new cohorts (no baseline) have mean=0, stddev=0 so the fallback fires at count >= 2. Sorted by sigma_observed desc, top limit (default 20). page_slugs capped at 50 per cohort.
  • src/core/cycle/recompute-emotional-weight.ts (v0.29) — Cycle phase orchestrator. Two SQL round-trips total: engine.batchLoadEmotionalInputs(slugs?)computeEmotionalWeight (per-row pure function) → engine.setEmotionalWeightBatch(rows). Reads config keys emotional_weight.high_tags (JSON array, falls back to default seed list on parse error) and emotional_weight.user_holder. Empty affectedSlugs array short-circuits with zero-work success. dry-run mode reports the would-write count without touching the DB. Engine throw bubbles into status: 'fail' with code RECOMPUTE_EMOTIONAL_WEIGHT_FAIL so the cycle continues.
  • src/core/transcripts.ts (v0.29) — listRecentTranscripts(engine, opts) library reused by both the gbrain transcripts recent CLI and the get_recent_transcripts MCP op. Reads dream.synthesize.session_corpus_dir + dream.synthesize.meeting_transcripts_dir config keys (same as discoverTranscripts); walks for .txt files within days; applies isDreamOutput guard from transcript-discovery.ts (skips dream-generated files); returns {path, date, mtime, length, summary}[] sorted newest-first. Summary mode (default true) returns first non-empty line + ~250 trailing chars. Full mode caps at 100KB/file. Missing/non-existent corpus dirs return [], not error. Trust gate lives in the op handler, not here: the op throws permission_denied for ctx.remote === true; this library is a trusted library function used by both the gated op and the local CLI.
  • src/core/operations-descriptions.ts (v0.29) — Constants module for tool descriptions. Pinned via test/operations-descriptions.test.ts. Houses GET_RECENT_SALIENCE_DESCRIPTION, FIND_ANOMALIES_DESCRIPTION, GET_RECENT_TRANSCRIPTS_DESCRIPTION plus the redirect-edited LIST_PAGES_DESCRIPTION, QUERY_DESCRIPTION, SEARCH_DESCRIPTION. Stable surface for the Tier-2 LLM routing eval — extracting them keeps the test from binding to whatever was in operations.ts at test-run time.
  • src/core/cycle/transcript-discovery.ts (v0.23) — Pure filesystem walk for synthesize. discoverTranscripts(opts) filters .txt files by date range, min_chars, and word-boundary regex excludePatterns (Q-3: medical matches "medical advice" but NOT "comedical"; power users may pass full regex). readSingleTranscript(path) is the gbrain dream --input <file> ad-hoc path. v0.23.2 self-consumption guard: DREAM_OUTPUT_MARKER_RE (anchored at frontmatter open ---\n, optional BOM + CRLF tolerance, scans first 2000 chars for dream_generated: true with case-insensitive value and word boundary on true) drives isDreamOutput(content, bypass=false). Both discoverTranscripts and readSingleTranscript skip matching files and emit a [dream] skipped <basename>: dream_generated marker stderr log (no more silent skips). bypassGuard?: boolean on DiscoverOpts and readSingleTranscript's opts disables the guard for the explicit --unsafe-bypass-dream-guard escape hatch only — never auto-applied for --input. Replaces v0.23.1's DREAM_OUTPUT_SLUGS content-prefix list.
  • src/commands/dream.ts — v0.17 gbrain dream CLI; ~80-line thin alias over runCycle. brainDir resolution requires explicit --dir OR sync.repo_path config. Flags: --dry-run, --json, --phase <name>, --pull, --dir <path>. v0.23 added --input <file> (ad-hoc transcript, implies --phase synthesize), --date YYYY-MM-DD, --from <d> --to <d> (backfill range). Conflict detection: --input + --date exits 2. ISO date validation. --dry-run runs Haiku significance verdict but skips Sonnet synthesis (codex finding #8 — NOT zero LLM calls). Exit code 1 on status=failed. v0.23.2 added --unsafe-bypass-dream-guard (long-form intentional, plumbed through runCycle.synthBypassDreamGuardSynthesizePhaseOpts.bypassDreamGuarddiscoverTranscripts({bypassGuard}) and readSingleTranscript({bypassGuard})). Loud stderr warning fires at synthesize-phase entry when set. Never auto-applied for --input so any caller can't silently re-trigger the loop bug. v0.41.38.0: resolveBrainDir now returns string | null (order: --dir → the resolved --source's local_path → global sync.repo_path → null) instead of process.exit(1) on "No brain directory found". A checkout-less postgres/Supabase brain runs the DB-only phases (incl. resolve_symbol_edges, the call-graph builder) and skips the 6 filesystem phases with details.reason: 'no_brain_dir'; runDream owns the only hard error (no checkout AND no engine). When --source resolves but that source has no on-disk checkout, returns null (DB-only) rather than borrowing a different source's global sync.repo_path (would mix scopes — codex P1 review). Pinned by test/dream-postgres.serial.test.ts. v0.42.5.0 (#1678): new --drain [--window <seconds>] for --phase extract_atomsrunDrain() bypasses the pack-gate and runs the single-hold bounded drain from src/core/cycle/extract-atoms-drain.ts under the same cycleLockIdFor(sourceId) the routine cycle uses (so a concurrent autopilot tick defers with cycle_already_running), reporting {extracted, skipped, remaining}. Exits EXIT_DRAIN_INCOMPLETE=3 while remaining > 0 so a cron/agent loop knows to continue. A null backlog count (the count query FAILED) is also treated as incomplete (exit 3), never as a drained success (Codex pre-landing #4). LockUnavailableErrorcycle_already_running skip (also exit 3). The extract_atoms_backlog doctor check (computeExtractAtomsBacklogCheck) surfaces the silent pack-gated backlog with the exact --drain command; pack-gated cycle skips carry a greppable pack_gated:true marker.
  • src/commands/friction.ts + src/core/friction.ts (v0.23) — gbrain friction {log,render,list,summary} reporter. Append-only JSONL under $GBRAIN_HOME/friction/<run-id>.jsonl. Schema is a flat extension of StructuredAgentError (D20). Render groups by severity → phase, defaults to --redact for md output (strips $HOME/$CWD to placeholders so reports paste safely in PRs). Run-id resolves from --run-id > $GBRAIN_FRICTION_RUN_ID > standalone.jsonl. Skills the claw-test exercises gain a _friction-protocol.md callout so agents know when to log friction.
  • src/commands/claw-test.ts + src/core/claw-test/ (v0.23) — gbrain claw-test [--scenario <name>] [--live --agent openclaw]. End-to-end "fresh user" friction harness. Two modes: scripted (CI gate, agent-free) and live (real openclaw subprocess, $12 in tokens). Sets GBRAIN_HOME=<tempdir> for hermeticity and captures gbrain's --progress-json events from each child's stderr to verify expected phases ran (import.files, extract.links_fs, doctor.db_checks). Phases for scripted mode: setup → install_brain (gbrain init --pglite) → import (--no-embed) → query → extract → verify (gbrain doctor --json, asserts status: 'ok') → render. Live mode hands BRIEF.md from test/fixtures/claw-test-scenarios/<name>/ to the agent runner. v1 ships with the OpenClaw runner only (src/core/claw-test/runners/openclaw.ts, invokes openclaw agent --local --agent <name> --message <brief>); hermes runner deferred to v1.1. Transcript capture (transcript-capture.ts) uses fs.createWriteStream with 'drain'-event backpressure — D17 fix for the 256KB-burst child-stall scenario. v0.18 upgrade scenario seeded via seed-pglite.ts SQL replay.
  • skills/_friction-protocol.md (v0.23) — shared cross-cutting convention skill (like _brain-filing-rules.md). Tells agents when to call gbrain friction log and how to choose a severity. Routes to friction CLI from any skill the claw-test exercises.
  • scripts/check-progress-to-stdout.sh — CI guard against regressing to \r-on-stdout progress. Wired into bun run test via scripts/check-progress-to-stdout.sh && bun test in package.json.
  • docs/progress-events.md — Canonical JSON event schema reference. Stable from v0.15.2, additive only.
  • src/core/markdown.ts — Frontmatter parsing + body splitter. splitBody requires an explicit timeline sentinel (<!-- timeline -->, --- timeline ---, or --- immediately before ## Timeline/## History). Plain --- in body text is a markdown horizontal rule, not a separator. inferType auto-types /wiki/analysis/ → analysis, /wiki/guides/ → guide, /wiki/hardware/ → hardware, /wiki/architecture/ → architecture, /writing/ → writing (plus the existing people/companies/deals/etc heuristics).
  • scripts/check-jsonb-pattern.sh — CI grep guard. Fails the build if anyone reintroduces (a) the ${JSON.stringify(x)}::jsonb interpolation pattern (postgres.js v3 double-encodes it), or (b) max_stalled INTEGER NOT NULL DEFAULT 1 in any schema source file (v0.15.1 #219 regression guard — must be DEFAULT 5 to preserve SIGKILL-rescue). Wired into bun test.
  • scripts/check-source-id-projection.sh (v0.32.8, PR #860) — CI grep guard for the multi-source bug class. Greps src/core/postgres-engine.ts + src/core/pglite-engine.ts for SELECT.*FROM pages projections matching the rowToPage feeder shape (id + slug + type + title) and fails if source_id is missing. After v0.32.8 Page.source_id is required at the type level; a projection that drops the column produces Page rows with source_id: undefined while TypeScript's : string lies about it. Codex's outside-voice review caught two pre-existing projections (getPage, putPage RETURNING) that lacked the column. Wired into bun run verify + bun run check:all.
  • docker-compose.ci.yml + scripts/ci-local.sh (v0.23.1) — Local CI gate. bun run ci:local spins up pgvector/pgvector:pg16 + oven/bun:1 with named volumes (gbrain-ci-pg-data, gbrain-ci-node-modules, gbrain-ci-bun-cache), runs gitleaks on host, smoke-tests scripts/run-e2e.sh argv handling, runs unit tests with DATABASE_URL unset (matches GH Actions structure), then runs all 29 E2E files sequentially. --diff swaps in the diff-aware selector; --no-pull skips upstream pulls; --clean nukes named volumes. Postgres host port defaults to 5434 (avoids 5432 manual gbrain-test-pg and 5433 sibling-project conflict); override with GBRAIN_CI_PG_PORT=NNNN. Stronger gate than current PR CI's 2-file Tier 1 set — closes the "push-and-wait" feedback loop pre-push.
  • scripts/select-e2e.ts + scripts/e2e-test-map.ts (v0.23.1) — Diff-aware E2E test selector. Reads three git sources (committed origin/master...HEAD, working-tree HEAD, and git ls-files --others --exclude-standard for untracked, NOT-gitignored files), classifies as EMPTY / DOC_ONLY / SRC. Fail-closed by design: EMPTY → all 29 files (clean branch shouldn't run nothing), DOC_ONLY (every path matches the README/CLAUDE/AGENTS/CHANGELOG/TODOS allowlist) → empty stdout, SRC → escape-hatch paths (schema, package.json, skills/) trigger all; otherwise the hand-tuned E2E_TEST_MAP glob → tests narrows; an unmapped src/ change still emits ALL files, never silently nothing. Pure-function exports (selectTests, classify, matchGlob) so it's trivial to test and fork. bun run ci:select-e2e prints the current selection on stdout, pipe-friendly. test/select-e2e.test.ts covers all 4 branches plus 3 codex regression guards (skills/, untracked files, unmapped src/) — 24 cases.
  • scripts/run-e2e.sh (v0.23.1 update) — Sequential E2E runner. Now accepts an optional argv-driven file list (used by ci:local:diff to pipe in selector output) and a --dry-run-list flag that prints the resolved file list and exits (used by ci-local.sh's startup smoke-test). Falls back to test/e2e/*.test.ts when invoked with no args.
  • scripts/llms-config.ts + scripts/build-llms.ts — Generator for llms.txt (llmstxt.org-spec web index) + llms-full.txt (inlined single-fetch bundle). Curated config drives both. Run bun run build:llms after adding a new doc. LLMS_REPO_BASE env var lets forks regenerate with their own URL base. FULL_SIZE_BUDGET (600KB) caps the inline bundle; generator WARNs if exceeded. Committed output is not analogous to schema-embedded.ts (no runtime consumer); we commit for GitHub browsing and fork-safe fetching.
  • AGENTS.md — Local-clone entry point for non-Claude agents (Codex, Cursor, OpenClaw, Aider). Mirrors CLAUDE.md intent via relative links. Claude Code keeps using CLAUDE.md.
  • docs/UPGRADING_DOWNSTREAM_AGENTS.md — Patches for downstream agent skill forks to apply when upgrading. Each release appends a new section. v0.10.3 includes diffs for brain-ops, meeting-ingestion, signal-detector, enrich.
  • src/core/schema-embedded.ts — AUTO-GENERATED from schema.sql (run bun run build:schema)
  • src/schema.sql — Full Postgres + pgvector DDL (source of truth, generates schema-embedded.ts)
  • src/commands/integrations.ts — Standalone integration recipe management (no DB needed). Exports getRecipeDirs() (trust-tagged recipe sources), SSRF helpers (isInternalUrl, parseOctet, hostnameToOctets, isPrivateIpv4). Only package-bundled recipes are embedded=true; $GBRAIN_RECIPES_DIR and cwd ./recipes/ are untrusted and cannot run command/http/string health checks.
  • src/core/search/expansion.ts — Multi-query expansion via Haiku. Exports sanitizeQueryForPrompt + sanitizeExpansionOutput (prompt-injection defense-in-depth). Sanitized query is only used for the LLM channel; original query still drives search.
  • recipes/ — Integration recipe files (YAML frontmatter + markdown setup instructions)
  • docs/guides/ — Individual SKILLPACK guides (broken out from monolith)
  • docs/integrations/ — "Getting Data In" guides and integration docs
  • docs/architecture/infra-layer.md — Shared infrastructure documentation
  • docs/ethos/THIN_HARNESS_FAT_SKILLS.md — Architecture philosophy essay
  • docs/ethos/MARKDOWN_SKILLS_AS_RECIPES.md — "Homebrew for Personal AI" essay
  • docs/guides/repo-architecture.md — Two-repo pattern (agent vs brain)
  • docs/guides/sub-agent-routing.md — Model routing table for sub-agents
  • docs/guides/skill-development.md — 5-step skill development cycle + MECE
  • docs/guides/idea-capture.md — Originality distribution, depth test, cross-linking
  • docs/guides/quiet-hours.md — Notification hold + timezone-aware delivery
  • docs/guides/diligence-ingestion.md — Data room to brain pages pipeline
  • docs/designs/HOMEBREW_FOR_PERSONAL_AI.md — 10-star vision for integration system
  • docs/mcp/ — Per-client setup guides (Claude Desktop, Code, Cowork, Perplexity)
  • BrainBench (benchmark suite + corpus): lives in the separate gbrain-evals repo. Not installed alongside gbrain.
  • skills/_brain-filing-rules.md — Cross-cutting brain filing rules (referenced by all brain-writing skills)
  • skills/RESOLVER.md — Skill routing table (based on the agent-fork AGENTS.md pattern)
  • skills/conventions/ — Cross-cutting rules (quality, brain-first, model-routing, test-before-bulk, cross-modal)
  • skills/_output-rules.md — Output quality standards (deterministic links, no slop, exact phrasing)
  • skills/signal-detector/SKILL.md — Always-on idea+entity capture on every message
  • skills/brain-ops/SKILL.md — Brain-first lookup, read-enrich-write loop, source attribution
  • skills/idea-ingest/SKILL.md — Links/articles/tweets with author people page mandatory
  • skills/media-ingest/SKILL.md — Video/audio/PDF/book with entity extraction
  • skills/meeting-ingestion/SKILL.md — Transcripts with attendee enrichment chaining
  • skills/citation-fixer/SKILL.md — Citation format auditing and fixing
  • skills/repo-architecture/SKILL.md — Filing rules by primary subject
  • skills/skill-creator/SKILL.md — Create conforming skills with MECE check
  • skills/daily-task-manager/SKILL.md — Task lifecycle with priority levels
  • skills/daily-task-prep/SKILL.md — Morning prep with calendar context
  • skills/cross-modal-review/SKILL.md — Quality gate via second model
  • skills/cron-scheduler/SKILL.md — Schedule staggering, quiet hours, idempotency
  • skills/reports/SKILL.md — Timestamped reports with keyword routing
  • skills/testing/SKILL.md — Skill validation framework
  • skills/soul-audit/SKILL.md — 6-phase interview for SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.md
  • skills/webhook-transforms/SKILL.md — External events to brain signals
  • skills/data-research/SKILL.md — Structured data research: email-to-tracker pipeline with parameterized YAML recipes
  • skills/minion-orchestrator/SKILL.md — Unified background-work skill (v0.20.4 consolidation of the former minion-orchestrator + gbrain-jobs split). Two lanes: shell jobs via gbrain jobs submit shell --params '{"cmd":"..."}' (operator/CLI only; MCP throws permission_denied for protected names) and LLM subagents via gbrain agent run (user-facing entrypoint). Shared Preconditions block, parent-child DAGs with depth/cap/timeouts, child_done inbox for fan-in, PGLite --follow inline path for dev. Triggers narrowed from bare "gbrain jobs" to "gbrain jobs submit" + "submit a gbrain job" so stats/prune/retry questions fall through to gbrain --help.
  • templates/ — SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.md templates
  • skills/migrations/ — Version migration files with feature_pitch YAML frontmatter
  • src/commands/publish.ts — Deterministic brain page publisher (code+skill pair, zero LLM calls)
  • src/commands/backlinks.ts — Back-link checker and fixer (enforces Iron Law)
  • src/commands/lint.ts — Page quality linter (catches LLM artifacts, placeholder dates)
  • src/commands/report.ts — Structured report saver (audit trail for maintenance/enrichment)
  • src/core/destructive-guard.ts (v0.26.5) — three-layer protection against accidental data loss in gbrain. assessDestructiveImpact(engine, sourceId) counts pages/chunks/embeddings/files for a source. checkDestructiveConfirmation(impact, opts) is the fail-closed gate (--confirm-destructive required when data is present; --yes alone is rejected). softDeleteSource / restoreSource / listArchivedSources / purgeExpiredSources drive the source-level archive lifecycle via the column shape introduced in migration v34 (sources.archived BOOLEAN, archived_at TIMESTAMPTZ, archive_expires_at TIMESTAMPTZ). v0.26.5 added the page-level analog through BrainEngine.softDeletePage / restorePage / purgeDeletedPages plus pages.deleted_at TIMESTAMPTZ and a partial purge index. The MCP delete_page op rewires to softDeletePage; new ops restore_page (scope: write) and purge_deleted_pages (scope: admin, localOnly: true) round out the surface. Search visibility (buildVisibilityClause in src/core/search/sql-ranking.ts) hides soft-deleted pages and archived sources from searchKeyword / searchKeywordChunks / searchVector in both engines. The autopilot cycle's new 9th purge phase calls purgeExpiredSources + engine.purgeDeletedPages(72) so the 72h TTL is real, not honor-system.
  • src/commands/pages.ts (v0.26.5) — gbrain pages purge-deleted [--older-than HOURS|Nd] [--dry-run] [--json] operator escape hatch. Mirror of gbrain sources purge for the page-level lifecycle. Hard-deletes pages whose deleted_at is older than the cutoff; cascades to content_chunks/page_links/chunk_relations.
  • src/core/op-checkpoint.ts (v0.36.4.0) — DB-backed checkpoint primitive for long-running ops. Migration v67 introduces op_checkpoints (op TEXT, fingerprint TEXT, completed_keys JSONB, updated_at TIMESTAMPTZ, PK(op, fingerprint)). Per-op fingerprint helpers (embedFingerprint, extractFingerprint, reindexFingerprint, integrityFingerprint, purgeFingerprint) each compute sha8(canonical-JSON(relevant-params)) so re-running with the same params resumes from completed_keys and re-running with different params (e.g. --limit 100 vs --limit 200) starts fresh. Cross-worker safe on Postgres (DB row, no file-lock race); PGLite degrades gracefully. Replaces per-op file-backed JSON checkpoints scattered across import.ts, embed.ts, reindex.ts. The 7-day TTL GC runs in the cycle's purge phase. Pinned by test/op-checkpoint.test.ts (~15 cases including per-op fingerprint scoping, codex outside-voice review #10-#16). import-checkpoint.ts from v0.34.2.0 was NOT migrated to this primitive in v0.36.4.0 — both checkpoint systems coexist without conflict; the migration requires async-propagating 4 sync call sites in src/commands/import.ts and rewriting 18 tests, deferred to a follow-up wave.
  • src/core/brain-score-recommendations.ts (v0.36.4.0) — pure data layer consumed by both gbrain doctor --remediation-plan / --remediate and gbrain features. computeRecommendations(checks, opts) returns Remediation[] with stable id, content-hash idempotency_key, severity, est_seconds, est_usd_cost, depends_on (D14: references stable ids, not check names — so plan order is reproducible across runs). classifyChecks(report) triages every doctor check into remediable | human_only | blocked (D13: three-state, not boolean — human_only covers RLS warnings and other human-judgment-required gates; blocked covers dependency chains where a parent check failed). maxReachableScore(checks) computes the ceiling for empty / under-configured brains (no entity pages → graph_coverage caps at 70; no embedding key → embedding_coverage caps at 60). Cost estimates pull from anthropic-pricing.ts for synthesize / patterns / consolidate and embedding-pricing.ts for embed jobs. Pinned by test/brain-score-recommendations.test.ts (~27 cases) including D6 #5 determinism, D9 content-hash idempotency, D12 DB-backed checkpoint provenance, and D13 three-state triage.
  • src/commands/doctor.ts extension (v0.36.4.0) — new --remediation-plan and --remediate CLI surfaces. --remediation-plan [--json] [--target-score N] prints what would run (stable id, idempotency_key, severity, est_seconds, est_usd_cost, depends_on); --remediate [--yes] [--target-score N] [--max-usd N] actually submits each plan step as a Minion job, in dependency order, re-checking score between every step. --target-score N defaults to 90; refuses to start when target exceeds maxReachableScore() and lists what's missing. --max-usd N is the cron-safety guard — submission refuses when the plan's est_total_usd_cost exceeds the cap (prevents synthesize loops from burning $100 of Anthropic credits while you're at lunch). JSON envelope adds a Check.remediation field (additive, schema_version unchanged). Pinned by tests in test/doctor.test.ts.
  • src/commands/jobs.ts extension (v0.36.4.0) — registers 11 new Minion handlers: reindex, repair-jsonb, orphans, integrity, purge, synthesize (PROTECTED), patterns (PROTECTED), consolidate (PROTECTED), extract_facts, resolve_symbol_edges, recompute_emotional_weight. Phase wrappers delegate to runCycle({phases:[name]}) so src/core/cycle.ts remains the single source of truth for phase semantics. Same fix wave: the standalone sync handler now passes noExtract: true to match runPhaseSync's contract — pre-fix, doctor's remediation plan emitting [sync, extract] double-extracted (codex #5).
  • src/core/minions/protected-names.ts extension (v0.36.4.0) — PROTECTED_JOB_NAMES extended with synthesize, patterns, consolidate. These phases internally submit subagent children with allowProtectedSubmit=true and can therefore spend Anthropic credits; treating them as routine "data-quality maintenance" was a misread caught by codex outside-voice (#6). Only trusted local callers (CLI, autopilot, doctor --remediate) can submit them; MCP requests are rejected by submit_job's protected-name guard.
  • src/commands/autopilot.ts extension (v0.36.4.0) — targeted-submit loop replaces blanket autopilot-cycle dispatch. Each tick: cheap engine.getHealth() (single SQL count) + computeRecommendations(), then route by shape — score >= 95 AND no plan AND <60min since last full → sleep; score >= 95 AND >=60min → submit autopilot-cycle (60-min floor exercises phase-coupling invariants on healthy brains); plan <= 3 steps AND est <5min → submit individual handlers (targeted); plan large OR score < 70 → submit full autopilot-cycle. The gbrain-cycle lock ensures targeted submissions and the full cycle can't run concurrently. maxWaiting: 1 per submit closes the queue-fan-out vector codex flagged (#17). Pre-fix autopilot ran a full 6-phase cycle every 5 minutes regardless of brain state; healthy brains burned synthesize+patterns+embed cycles for zero work.
  • src/core/cycle.ts extension (v0.36.4.0) — purge phase (the cycle's 9th phase, introduced in v0.26.5 for soft-delete TTLs) extended to GC stale op_checkpoints rows older than 7 days. Non-fatal on pre-v67 brains (DROP-target-table check before DELETE).
  • src/commands/embed.ts extension (v0.36.4.0) — wires --background as the reference integration for the new maybeBackground() helper. gbrain embed --stale --background submits as a Minion job and prints job_id=N to stdout, exits 0. Composable in shell pipelines: JOB=$(gbrain embed --stale --background | grep -oE 'job_id=[0-9]+' | cut -d= -f2); gbrain jobs follow $JOB. The other six commands (extract, lint, backlinks, reindex, integrity, pages) adopt the same 4-line pattern in a follow-up wave (T7 deferred).
  • src/core/cli-options.ts extension (v0.36.4.0) — new maybeBackground(opName, fingerprintArgs, runDirect) helper. Same semantics in TTY and cron (D9 — no --no-tty-detect flag, no surprise behavior change between contexts): when --background is passed, submits the op as a Minion job via op_checkpoints for resumability and returns the job_id. --background --follow execs gbrain jobs follow <id> so the user sees the same stderr stream they'd get from a direct call. PGLite degrades to inline execution with a clear stderr note ("PGLite worker pool not yet supported; running inline"). Returns a tagged union the caller dispatches on.
  • openclaw.plugin.json — ClawHub bundle plugin manifest
  • src/commands/capture.ts + src/commands/serve-http.ts + src/core/{operations,import-file,types,utils,facts/absorb-log,brainstorm/{orchestrator,error-classify},scope,postgres-engine,pglite-engine}.ts extensions (v0.39.3.0 smoke-test wave) — productionization of the v0.38.0.0 ingestion cathedral release after a real production smoke test against Supabase+PgBouncer (PR #1299, contributed by @garrytan-agents). 12 atomic commits closing 2 bugs + 10 warnings: BUG-1 capture frontmatter merge via mergeCaptureFrontmatter (uses gray-matter directly, NOT the lossy parseMarkdown); BUG-2 /ingest null-guard + outer try/catch envelope with !res.headersSent guard; WARN-1 dedup via separate normalize-for-hash (normalizeForHash strips BOM/CRLF/whitespace/NFKC) + body-after-frontmatter-strip on the DB hash (excludes captured_at + ingested_at from the existing.content_hash check so capture-cli timestamp variations don't invalidate the chunk cache); WARN-3 friendly pages_source_id_fk rewrite via maybeRewriteSourceFkError applied to BOTH local + thin-client callRemoteTool catch blocks; WARN-4 facts:absorb 'No database connection' suppression via typed instanceof GBrainError && e.problem check + first-occurrence stack-trace info log (module-scoped _hasLoggedDisconnectedFactsAbsorb flag, test seam _resetFactsAbsorbDisconnectedFlagForTests); WARN-5 + WARN-6 CLI help discoverability (capture added to CLI_ONLY_SELF_HELP set + pre-engine-bind --help short-circuit in handleCliOnly + new BRAIN section in printHelp text); WARN-7 binary file guard via detectBinaryNullByte(buf) first-8KB NUL scan on both --file (Buffer-read, no encoding) and --stdin (readStdinBuffer Buffer accumulator); WARN-8 provenance write-through with put_page accepting 3 new optional params (source_kind, source_uri, ingested_via — ingested_at is server-stamped) + CV6 trust gate (when ctx.remote !== false IGNORE client params, server stamps mcp:put_page, fail-closed posture matches v0.26.9 F7b) + CV12 COALESCE-preserve UPDATE semantics (omitting params on a later put_page preserves prior values; first-write-wins audit trail); WARN-9 /admin/api/register-client scopes normalization via new normalizeScopesInput(raw: unknown) helper in src/core/scope.ts (accepts string, string[], missing; rejects ['read write'] space-in-element bug shape, non-string elements, empty array, unknown scope names; deduped + sorted output); WARN-10 brainstorm timeout surfacing via orchestrator-level try/catch wrap at runBrainstorm entry point — covers EVERY internal SQL site by being a single-point wrap, classifies SQLSTATE 57014 (via postgres.js .code, alternate .sqlState, or message fallback) into StructuredAgentError with code='brainstorm_timeout' and hint covering all 3 PG cancel sub-causes (statement_timeout, lock_timeout, user-cancel); CV5 read-path extension surfaces all 4 provenance columns via getPage SQL projection + rowToPage 3-state optional read + Page interface; CV15 canonical source resolver routes capture through resolveSourceWithTier(engine, parsed.source, cwd) matching every other CLI op (was parsed.source ?? 'default' pre-fix, silently ignoring env/dotfile/local_path tiers); CV7 thin-client --source rejection (server-side OAuth client registration owns source scope, per-call override would reopen CV6 spoofing surface); CV3 source_kind taxonomy is now closed (capture-cli | put_page | mcp:put_page | webhook | file-watcher | inbox-folder | cron-scheduler), --source flag maps to source_id only. Tests: 8 new test files + 3 extended files (test/capture-build-content.test.ts, test/capture-runcapture.test.ts, test/put-page-provenance.test.ts, test/scope-normalize.test.ts, test/cli-help-discoverability.test.ts, test/brainstorm-timeout.test.ts; extended test/facts-absorb-log.test.ts, test/import-file.test.ts, test/e2e/engine-parity.test.ts, test/e2e/serve-http-ingest-webhook.test.ts). Smoke-test report ships verbatim at docs/v0.38-smoke-test-report.md with Editor's Note prepending the two factual corrections (BUG-2 actual line + WARN-5 root cause). Plan + 15 decisions + codex outside voice trace at ~/.claude/plans/system-instruction-you-are-working-async-popcorn.md. v0.39 follow-ups in TODOS.md: SQL-shape rewrite of listPrefixSampledPages for PgBouncer, magic-byte allowlist for binary detection, --source-kind override flag, ingest_capture handler architecture migration, provenance-history table, facts:absorb root-cause trace.

BrainBench — in a sibling repo (v0.20+)

BrainBench — the public benchmark for personal-knowledge agent stacks — lives in github.com/garrytan/gbrain-evals. It depends on gbrain as a consumer; gbrain never pulls in the ~5MB eval corpus or the pdf-parse dev dep at install time.

gbrain's public API surface (the exports map in package.json) is what gbrain-evals consumes: gbrain/engine, gbrain/types, gbrain/operations, gbrain/pglite-engine, gbrain/link-extraction, gbrain/import-file, gbrain/transcription, gbrain/embedding, gbrain/config, gbrain/markdown, gbrain/backoff, gbrain/search/hybrid, gbrain/search/expansion, gbrain/extract. Removing any of these is a breaking change for the gbrain-evals consumer.

v0.36.1.0 Hindsight calibration wave (key files cluster)

The wave that taught gbrain to know how the user tends to be wrong + use that knowledge at every advice surface. Six-migration schema (v67-v72), three new cycle phases, eight expansions, one admin tab. Plan persisted at ~/.claude/plans/system-instruction-you-are-working-rippling-knuth.md. Convention skill at skills/conventions/calibration.md has the agent- facing rules.

v0.37.2.0 hotfix (2026-05-20, migration v80)takes_resolution_consistency CHECK widened to accept quality='unresolvable' AND outcome=NULL as the 4th valid resolution state. Column-level CHECK on resolved_quality renamed to takes_resolved_quality_values and widened to enumerate all 4 states. Unblocks production grading scripts that write the judge's 4th verdict type. Take.resolved_quality, TakeResolution.quality, and takes-fence.ts:TakeQuality all widen to 4-state. TakesScorecard gains sibling fields unresolvable_count + unresolvable_rate; resolved stays 3-state (correct+incorrect+partial) so historical comparisons hold. finalizeScorecard formula: `unresolvable_rate = unresolvable_count / (resolved

  • unresolvable_count), NULL when both 0. Spec doc preserved at docs/architecture/calibration-quality-gate-spec.md(from closed PR #1191) since the follow-up minor (forthcoming) ships the falsifiability + per-category calibration on top. Migration renumbered v74→v79→v80 during successive master merges — v0.37.0.0's autonomous-remediation wave claimed v68-v78, then v0.37.1.0 (brainstorm/lsd) claimed v79. Pinned by R1-R5 intest/takes-resolution.test.tsandtest/migrate.test.ts`'s v80 structural + PGLite round-trip suite (CHECK admits unresolvable+NULL, still rejects partial+true and unresolvable+true|false, pre-v80 NULL/NULL rows survive).
  • src/core/cycle/base-phase.ts — abstract BaseCyclePhase class. Enforces sourceScopeOpts(ctx) threading at the type level; closes the v0.34.1 source-isolation leak class structurally for every new phase. Inherits source-scope, budget meter, error envelope, progress reporter. propose_takes / grade_takes / calibration_profile all extend it.
  • src/core/cycle/propose-takes.ts — LLM scans markdown prose, proposes gradeable claims to take_proposals queue. Idempotency cache on (source_id, page_slug, content_hash, prompt_version) composite unique index. F2 fence-dedup: existing canonical takes passed to extractor as context. v0.36.1.0 ships a stub prompt; tuned prompt arrives via the T19 synthetic corpus build.
  • src/core/cycle/grade-takes.ts — walks unresolved takes older than 6 months, retrieves evidence, asks judge model, caches verdict. Auto-resolve DISABLED by default (D17). Conservative thresholds:

    =0.95 single OR >=0.85 ensemble 3/3 unanimous. T5 ensemble (aggregateEnsemble) reuses v0.27.x cross-modal substrate; fires on borderline 0.6-0.95 band. Writes to take_grade_cache.

  • src/core/cycle/calibration-profile.ts — aggregates resolved takes into 2-4 narrative pattern statements + active bias tags. Voice-gated via gateVoice(). Cold-brain skip when <5 resolved. Writes to calibration_profiles with audit columns (voice_gate_passed, voice_gate_attempts, grade_completion).
  • src/core/calibration/voice-gate.ts — single gateVoice() function (D24), mode parameter (pattern_statement | nudge | forecast_blurb | dashboard_caption | morning_pulse). 2 regens then hand-written template fallback from src/core/calibration/templates.ts. Haiku judge with mode-specific rubrics; all rubrics structurally forbid clinical/preachy voice.
  • src/core/calibration/cross-brain.ts — D18 4-rule contract for cross-brain calibration reads. Local-first → mount-fallback (only with canReadMountsForCtx(ctx) true) → cross-brain attribution via source_brain_id + from_mount → subagent prohibition closes the OAuth-token-to-cross-brain-leak surface. All 4 rules pinned in test/cross-brain-calibration.test.ts.
  • src/core/calibration/nudge.ts — E7 real-time pattern surfacing. evaluateAndFireNudge(opts) is the full pipeline: threshold check (conviction > 0.7, holder match, slug-derived domain hint matches active bias tag), cooldown probe (14d via take_nudge_log), fire + log. STDERR-only output for v0.36.1.0; multi-channel deferred.
  • src/core/calibration/take-forecast.ts — E5 Brier-trend at write time. Pure math over existing TakesScorecard; no LLM. Returns predicted_brier, bucket_n, overall_brier. Insufficient-data branch at MIN_BUCKET_N = 5. batchForecast memoizes per (holder, domain) tuple.
  • src/core/calibration/gstack-coupling.ts — E4 outcome-driven learnings coupling. writeIncorrectResolution(opts) shells out to gstack-learnings-log binary. Config gate (cycle.grade_takes.write_gstack_learnings, default false for external users). Namespace prefix gbrain:calibration:v0.36.1.0: so --undo-wave can scrub.
  • src/core/calibration/svg-renderer.ts — D23 server-rendered SVG for the admin SPA Calibration tab. Pure functions: data → SVG string. Inlines design tokens; XSS-safe via escapeXml(). Four chart renderers: renderBrierTrend, renderDomainBars, renderAbandonedThreadsCard, renderPatternStatementsCard. SPA renders via <TrustedSVG> wrapper behind requireAdmin.
  • src/core/calibration/undo-wave.ts — D18 CDX-3 resolution. undoWave reverses the wave's mutations: unsets takes.resolved_* for wave-applied resolutions (cross-checks resolved_by so manual writes persist), deletes calibration_profiles, purges nudge logs, marks grade-cache rows applied=false. --dry-run shows counts without writing. Idempotent on wave_version match.
  • src/core/calibration/think-ab.ts — D19 A/B harness. runAbTrial calls thinkRunner twice (baseline + with-calibration), records preference to think_ab_results. buildAbReport aggregates over 30-day window; flags calibration_net_negative when n>=20 + win rate < 45% on decisive trials.
  • src/core/calibration/recall-footer.ts — formatter for the morning pulse calibration block. Cold-brain branch when <5 resolved. v0.36 ship state: opt-in via the wiring layer; auto-on in v0.37+.
  • src/core/eval-contradictions/calibration-join.ts — E3 cross- reference. tagFindingWithCalibration(finding, profile) returns bias-tag context for contradictions that match active patterns. Returns null when profile missing (R2 regression — output byte-identical to v0.32.6).
  • src/core/think/prompt.ts extension — E1 anti-bias rewrite. withCalibration option on buildThinkSystemPrompt adds anti-bias rules. New buildCalibrationBlock() emits the <calibration> XML. buildThinkUserMessage has TWO shapes: default (question first) for R1 regression, with-calibration (retrieval → calibration → question per D22) when opt-in. Wired into runThink via opts.withCalibration + opts.calibrationHolder.
  • src/commands/calibration.ts — CLI: gbrain calibration (read + print), --regenerate, --undo-wave <ver> (T17), ab-report (T18). MCP op get_calibration_profile (scope: read) backs the same data path. Source-scoped via sourceScopeOpts(ctx).
  • src/commands/serve-http.ts extension — three new admin routes: /admin/api/calibration/profile, /admin/api/calibration/charts/:type (image/svg+xml; type in {brier-trend, domain-bars, pattern-statements, abandoned-threads}), and /admin/api/calibration/pattern/:id (TD3 drill-down).
  • src/commands/takes.ts extension — gbrain takes revisit <slug> (TD4 / D30) opens $EDITOR on the source page with a <!-- gbrain:revisit --> cursor marker.
  • src/commands/doctor.ts extension — 4 new checks: abandoned_threads, calibration_freshness, grade_confidence_drift (CDX-11 mitigation surface; math arrives v0.37+), voice_gate_health.
  • admin/src/pages/Calibration.tsx — Calibration tab. Single-column Linear-calm-clarity layout matching the approved variant-B mockup. <TrustedSVG> wrapper handles dangerouslySetInnerHTML for the server-rendered SVG.
  • admin/src/index.css extension — --text-muted: #555 → #777 (TD2, WCAG AA contrast bump from 4.0 to ~5.5 on the #0a0a0f bg).
  • test/fixtures/calibration/extract-takes-corpus/ — synthetic prompt- tuning corpus. v0.36.1.0 ships 5 representative pages; full 50-page
    • 10-page holdout generated by gbrain calibration build-corpus (v0.37+ subcommand). All anonymized per CLAUDE.md placeholder list.
  • scripts/check-synthetic-corpus-privacy.sh — CDX-14 mitigation. CI guard in bun run verify. Greps for explicit dollar amounts + verifies non-essay fixtures reference at least one placeholder name.
  • test/regressions/v0.36.1.0-iron-rule.test.ts — R1-R5 regression inventory test file. Pins all 5 IRON-RULE regressions in one place for future bisects.
  • DESIGN.md — repo-root design system. Formalizes the de facto admin tokens that landed v0.26.0. Calibration target for future /plan-design-review and /design-review.

Thin-client routing (v0.31.1, Issue #734)

gbrain init --mcp-only (v0.29.2) sets up a thin-client install: no local brain content, just an OAuth client pointing at a remote gbrain serve --http. v0.29.2/v0.30.0 only refused 9 obvious local-only commands; the other ~25 silently fell through to connectEngine() and opened the empty local PGLite, returning "No results." against a populated remote brain. v0.31.1 fixes the silent-empty-results bug class for every operation surface.

Key files:

  • src/cli.ts — Routing seam INSIDE the existing op-dispatch path (CDX-1: no parallel src/core/thin-client/ module; routing is a ~80-line conditional in runThinClientRouted). Detects isThinClient(cfg) BEFORE connectEngine so thin-client installs never open the empty PGLite. localOnly ops on thin-client refuse via refuseThinClient (with pinpoint hint table THIN_CLIENT_REFUSE_HINTS). Banner via printIdentityBannerBestEffort before each routed call (suppressed by --quiet, GBRAIN_NO_BANNER=1, non-TTY default). Exhaustive TS never switch on RemoteMcpError.reason for canned, actionable error messages. ENG-2 renderer parity: local-engine path runs JSON.parse(JSON.stringify(result)) so renderers see the same shape on both paths (kills Date/bigint/Buffer drift class).
  • src/core/mcp-client.tscallRemoteTool(config, toolName, args, opts). Hardened in v0.31.1 (CDX-4): all transport errors normalized to RemoteMcpError via the toRemoteMcpError funnel. New CallRemoteToolOptions {timeoutMs, signal}; buildAbortController composes external signal with timeout. New RemoteMcpErrorReason stable union, RemoteMcpErrorDetail.kind ('timeout' | 'aborted' | 'unreachable') sub-tag, RemoteMcpErrorDetail.code field carrying server-supplied error codes (e.g. missing_scope). extractToolErrorCode parses JSON envelopes first, falls back to substring detection for legacy server messages. unpackToolResult<T>(res) unchanged (parses tool-call JSON content). _clearMcpClientTokenCache() test escape.
  • src/core/cli-options.tsparseGlobalFlags adds --timeout=Ns (accepts 30s, 2m, 500ms, plain ms). Default null = per-command default (30s for most ops, 180s for think). parseTimeout(s) exported helper.
  • src/core/doctor-remote.tsgbrain remote doctor adds the oauth_client_scopes_probe check (CDX-5). Probes the read tier via get_brain_identity and admin tier via get_health; reports per-tier status with pinpoint remediation when admin is missing. buildScopeCheck
    • ScopeProbeResult exported for test access. Skippable via GBRAIN_DOCTOR_SKIP_SCOPE_PROBE=1 for fixtures that mock /mcp at JSON-RPC initialize level only (MCP SDK Client hangs on shape mismatch).
  • src/core/ssrf-validate.ts (v0.36 Commit 0) — DNS-rebinding-defended URL validation. validateAndResolveUrl(url) resolves the hostname via dns.lookup({all: true, family: 0}), checks EVERY A AND AAAA record against the internal-IP deny list, returns the resolved IP so callers fetch by IP (defeats DNS rebinding: validation IP === fetch IP). fetchWithSSRFGuard(url, opts) does redirect-aware fetching with per-hop re-validation, max 3 hops by default. Reusable across all URL-fetching features. Test seam __setDnsLookupForTests for hermetic tests.
  • src/core/search/query-intent.ts extension (v0.36 cross-modal wave) — new suggestedModality: 'text' | 'image' | 'both' axis on QuerySuggestions. Module-scope CROSS_MODAL_PATTERNS regex array (compiles once at module load). isAmbiguousModalityQuery(query) heuristic gate fires when a visual noun + reference marker combination indicates genuinely ambiguous routing — used by the Commit 4 LLM tie-break to bound LLM calls to <1% of queries.
  • src/core/search/mode.ts extension (v0.36 cross-modal wave) — ModeBundle extended with 7 cross-modal knobs: cross_modal_both_text_weight / cross_modal_both_image_weight (D6 weighted RRF for 'both' mode, defaults 0.6/0.4), image_query_text_refinement_weight / image_query_image_refinement_weight (D13 hybrid intersect for searchByImage query refinement, defaults 0.4/0.6), unified_multimodal + unified_multimodal_only (Phase 3 unified column routing flags), cross_modal_llm_intent (Commit 4 opt-in escalation). SEARCH_MODE_CONFIG_KEYS extended with 7 corresponding config keys. KNOBS_HASH_VERSION bumped 2→3 (D2 — closes the silent cache-hit class where a cached text-mode result could leak to an image-mode caller).
  • src/core/search/hybrid.ts extension (v0.36 cross-modal wave) — cross-modal routing branch at the embed step. Resolves effectiveModality from per-call opts.crossModal (normalized: literal 'auto' → undefined per D22-1) → suggestions.suggestedModality'text' default. Image route: embedQueryMultimodal + searchVector({embeddingColumn: 'embedding_image'}), skip expansion + keyword (D9 mode-bundle override). 'both' route: parallel text + image vector searches merged via rrfFusionWeighted with effectiveRrfK(baseRrfK, weight) from the configured cross-modal weights. Phase 3 unified routing fires when cfg.search.unified_multimodal === true — bypasses dual-column branching, runs embedQueryMultimodal + searchVector({embeddingColumn: 'embedding_multimodal'}), D8 fail-open on zero rows + not strict-mode falls through to dual-column. Commit 4 LLM escalation fires only when (no explicit per-call opt) AND (regex returned 'text') AND (cfg.search.cross_modal.llm_intent is true) AND (isAmbiguousModalityQuery returns true). Fail-open on every error.
  • src/core/search/image-loader.ts (v0.36 Phase 2) — loadImageInput(input, opts) accepts local path, data: URI, or http(s):// URL. Magic-byte sniff for PNG/JPEG/WebP. Hard size cap (default 10 MB, configurable via search.image_query.max_bytes). For URLs: routes through fetchWithSSRFGuard so DNS rebinding + redirect chains are defeated. Pre-flight Content-Length check + post-fetch size guard for lying servers. ImageLoadError with discriminated code (INVALID_FORMAT / OVERSIZED / INVALID_URL / FETCH_FAILED / TIMEOUT / SSRF_BLOCKED / NOT_FOUND).
  • src/core/search/by-image.ts (v0.36 Phase 2) — searchByImage(engine, input, opts). Always runs image branch (embedQueryMultimodalImage + searchVector(embedding_image)). D13 hybrid intersect: when caller provides optional query, runs parallel text branch via embedQueryMultimodal(query) and merges via rrfFusionWeighted with weights from resolved mode. Phase 3 widens to unified column once search.unified_multimodal=true (transparently upgrades the retrieval quality post-reindex).
  • src/core/spend-log.ts (v0.36 Phase 2 D23-#6) — per-OAuth-client paid-API spend tracking against the mcp_spend_log table (migration v74). checkBudget(engine, clientId, capCents) is the pre-flight gate; throws BudgetExceededError when today's spend has hit the cap. recordSpend(engine, entry) is best-effort post-call. UTC day-aligned aggregation so caps roll over deterministically regardless of server timezone. Local CLI callers (no clientId) bypass the gate. Pre-v0.36 brains without the table fail open to spend=0. VOYAGE_MULTIMODAL_3_PER_IMAGE_CENTS = 0.12 cents per image embed.
  • src/core/search/llm-intent.ts (v0.36 Commit 4) — opt-in LLM tie-break. classifyModalityWithLLM(query, fallback) routes through gateway.chat() with a fixed single-word-output system prompt. 1s timeout via AbortController. parseModality(raw, fallback) is the pure parser — tolerates trailing punctuation + casing. Fail-open on every error (gateway unavailable, timeout, parse failure, unrecognized output) — returns fallback so a misbehaving LLM can never break search. Cost-bounded by the ambiguity heuristic in query-intent.ts (fires <1% of queries when on).
  • src/commands/reindex-multimodal.ts (v0.36 Phase 3) — gbrain reindex --multimodal [--limit N] [--dry-run] [--cost-estimate] [--no-embed] [--yes] [--json]. Walks content_chunks WHERE embedding_multimodal IS NULL, batches via embedMultimodalSafe (Commit 0 partial-failure-aware), persists. D7 lock acquisition via tryAcquireDbLock('gbrain-reindex-multimodal', 360min). Cost prompt + 10s Ctrl-C grace window in TTY. GBRAIN_NO_REEMBED=1 bypass. Checkpoint at ~/.gbrain/reindex-multimodal-checkpoint.json for resume. D23-#2 auto-flip prompt at coverage=100% completion (TTY: interactive; non-TTY: stderr hint with paste-ready command).
  • src/core/backfill-registry.ts extension (v0.36) — new modality backfill kind. SQL filter requires chunk_source='image_asset' AND embedding_image IS NOT NULL AND (modality IS NULL OR modality != 'image'). D22-7 defensive guard: never flag a non-image chunk that happens to have embedding_image populated. Idempotent — second run finds zero rows.
  • src/core/migrate.ts v74 (mcp_spend_log) + v75 (embedding_multimodal_column) — Phase 2 spend-log table + Phase 3 unified column ALTER. v75 is column-only (no HNSW index — deferred to post-reindex per pgvector best practice). v74 uses BTREE on (client_id, created_at) + (token_name, created_at)date_trunc('day', TIMESTAMPTZ) is NOT IMMUTABLE so can't appear in index expressions; range scan on created_at covers the per-day rollup query.
  • src/core/operations.tsget_brain_identity op (read scope, no params, banner-only): cheap counter packet {version, engine, page_count, chunk_count, last_sync_iso} for the thin-client identity banner. Reuses engine.getStats(); banner's 60s client-side TTL bounds frequency to ≤1/60s per CLI process (well below the Fly.io health-check cadence that motivated the original getStats cost warning).
  • src/commands/{salience,anomalies,graph-query,think}.ts — Per-command thin-client routing branches. These commands bypass the operation-layer dispatch in cli.ts (call engine.foo() directly), so each gets its own if (isThinClient(cfg)) { callRemoteTool(...) } branch that maps CLI flags to op params. think is a special case: the server's think op intentionally disables --save/--take for remote callers (operations.ts:1103-1135 trust-boundary gate); thin-client think warns loudly when those flags are set.

Schema Cathedral v3 (v0.40.7.0)

The schema-pack mutation surface shipped in v0.40.7.0 as the production rebuild of closed community PR #1321 (@garrytan-agents). Six new foundation modules + a mutate skeleton + stats/sync data plane + 14 CLI verbs + 9 MCP ops + a first-class agent skill. See ~/.claude/plans/system-instruction-you-are-working-recursive-thacker.md for the full plan + 21 captured design decisions.

Key files (v0.40.7.0 additions):

  • src/core/schema-pack/pack-lock.ts — Atomic O_CREAT|O_EXCL per-pack lock. DELIBERATELY NOT the existsSync + writeFileSync TOCTOU shape from src/core/page-lock.ts (codex C8 caught the bug class). Default 60s TTL, refresh every 10s while withPackLock(fn) runs, --force semantics = "steal stale lock" NOT "skip locking." Lock path per-pack so two packs never block each other.
  • src/core/schema-pack/mutate-audit.ts — ISO-week JSONL at ~/.gbrain/audit/schema-mutations-YYYY-Www.jsonl. Privacy-redacted per D20: type names → sha8, prefixes → first slug segment only, matches candidate-audit.ts privacy posture. Logs BOTH success AND failure events so the v0.40.7+ schema_pack_writability doctor check has signal to read. summarizeMutations() is the cross-surface parity primitive.
  • src/core/schema-pack/registry.ts extensions — invalidatePackCache(name?) walks the extends-chain reverse-graph (codex C6 fix; pre-v0.40.7, editing a parent pack silently left children stale). tryCachedPack(name) TTL-gated fast path: inside STAT_TTL_MS (default 1000ms, env GBRAIN_PACK_STAT_TTL_MS) returns cached without statting. Outside the window: stats every file in the chain; cascade-invalidates on mtime change (D11 cross-process detection).
  • src/core/schema-pack/best-effort.tsloadActivePackBestEffort(ctx) returns ResolvedPack | null. Single source of truth for the T1.5 wiring sites. null means EMPTY FILTER (NOT hardcoded defaults — D4 contract closing the silent-violation bug class).
  • src/core/schema-pack/lint-rules.ts — 12 pure rule functions. withMutation's pre-write validation gate composes the 10 file-plane rules; the 2 DB-aware rules (extractable_empty_corpus, mutation_count_anomaly) need an engine. Single source of truth consumed by CLI lint + MCP schema_lint + the pre-write validation gate. v0.41.37.0 (#1569): new file-plane rule link_regex_catastrophic_backtrack — advisory ReDoS pre-screen that flags the classic nested-quantifier shapes ((a+)+, (a*)*, (a+)*, (\w+)+) in a link_type's inference.regex via NESTED_QUANTIFIER_RE. WARNING not error: a hard reject would disable the whole pack on upgrade (pages fall back to legacy typing). The runtime input-length cap in redos-guard.ts is the actual safety net; this rule tells the pack author to fix the pattern.
  • src/core/schema-pack/redos-guard.ts + src/core/schema-pack/link-inference.ts (v0.41.37.0, #1569) — ReDoS hardening for pack inference regexes. redos-guard.ts adds MAX_REGEX_INPUT_CHARS (default 64_000, env override GBRAIN_MAX_REGEX_INPUT_CHARS) — a hard input-length cap, the real runtime safety net (catastrophic backtracking needs a long input to blow up; a link-extraction context is normally a sentence or short paragraph). Over the cap, runRegexBounded throws the new tagged RegexInputTooLargeError and the regex is skipped (degrade-to-mentions) without entering the node:vm. link-inference.ts:inferLinkTypeFromPack no-budget branch (test contexts) previously ran new RegExp(pattern).test(context) UNBOUNDED — the one ReDoS hole with no timeout; it now routes through runRegexBounded so the input-length cap + per-regex vm timeout (PER_REGEX_TIMEOUT_MS = 50) apply on every path. Defensive hardening + diagnostics; the deterministic ~3100-file sync-wedge root cause remains open (no repro). Pinned by test/redos-hardening.test.ts + test/schema-pack-lint-rules.test.ts.
  • src/core/schema-pack/query-cache-invalidator.tsinvalidateQueryCache(engine, sourceId?) DELETEs query_cache rows so cached search results bound to old page types don't survive a schema mutation. Codex C9 fix.
  • src/core/schema-pack/mutate.ts — 8-step withMutation skeleton (bundled-guard → lock → read → mutator → validate → atomic write → audit → invalidate). 11 mutation primitives: addTypeToPack, removeTypeFromPack (with codex C14 reference check), updateTypeOnPack, addAliasToType, removeAliasFromType, addPrefixToType, removePrefixFromType, addLinkTypeToPack, removeLinkTypeFromPack, setExtractableOnType, setExpertRoutingOnType. Atomic write via .tmp + fsync + rename — pack file on disk is NEVER partial. Inline minimal JSON→YAML emitter so YAML packs stay YAML (does NOT preserve comments — pin pack.json if you care about layout).
  • src/core/schema-pack/stats.tsrunStatsCore(engine, opts) returns per-source + aggregate page counts + coverage % + dead_prefixes (declared prefixes with zero matching pages — agent's drilldown signal). Multi-source aware (sourceIds[] federated, sourceId single, or whole-brain). PGLite + Postgres parity via executeRaw. Empty brain → coverage:1.0 (vacuous truth).
  • src/core/schema-pack/sync.tsrunSyncCore(engine, opts) chunked UPDATE in 1000-row batches per declared prefix (D14). Concurrent writers never block on a single row >100ms. Codex C5 write-side scoping via ctx.sourceId directly (NOT sourceScopeOpts which inherits OAuth read federation). Idempotent on --apply re-run.
  • src/commands/schema.ts extension — 14 new CLI verbs in the dispatch table: add-type, remove-type, update-type, add-alias, remove-alias, add-prefix, remove-prefix, add-link-type, remove-link-type, set-extractable, set-expert-routing, stats, sync, reload. withConnectedEngine defensive fix from closed PR #1321 retained. Lifecycle-grouped help text (Inspection / Activation / Authoring / Discovery+repair).
  • src/core/operations.ts extension — 9 new MCP ops: get_active_schema_pack, list_schema_packs, schema_stats, schema_lint, schema_graph, schema_explain_type, schema_review_orphans (all read-scope, NOT localOnly), plus schema_apply_mutations (admin scope, NOT localOnly per D2 so remote agents like your OpenClaw can author packs over HTTPS MCP — batched per D10, one MCP tool taking an mutations[] array atomically inside ONE withPackLock, audit log captures actor: mcp:<clientId8>) and reload_schema_pack (admin, NOT localOnly). Trust posture: per-call schema_pack opt STAYS rejected for remote callers via op-trust-gate.ts (R2 regression preserved).
  • src/commands/whoknows.ts + src/core/operations.ts:find_experts — T1.5 wiring sites. Pack-aware via expertTypesFromPack(pack.manifest) from best-effort.ts. Pack-load failure → EMPTY filter (NOT hardcoded ['person', 'company'] defaults per D4). A researcher type declared --expert now surfaces in whoknows results; pre-v0.40.7 it silently never matched.
  • skills/schema-author/SKILL.md — Agent dispatcher for "evolve the schema pack." Triggers: 15+ phrasings including "add a page type", "my brain has untyped pages", "propose new types from my corpus", "backfill page types". Explicit Non-goals callout to brain-taxonomist (files one page) and eiirp (schema-check during iteration) so agents pick the right surface. 7-phase workflow: brain → assess → propose → apply → sync → verify → commit. Lists every gbrain schema CLI verb + every MCP op the skill uses. brain_first: exempt frontmatter. Required v0.40.7+ conformance sections: Contract, Anti-Patterns, Output Format.
  • skills/conventions/schema-evolution.md — Canonical convention: "when to add a type vs alias vs prefix." Decision tree: <20 pages → don't pack-codify; 20-100 → alias or narrow prefix on existing type; 100+ → first-class type. Don'ts section + "when to remove a type" + "when to commit the pack" all answered in one place.
  • skills/RESOLVER.md + skills/manifest.json — schema-author wired into the dispatcher with full functional-area trigger list (compressed routing pattern per v0.32.3 dispatcher convention).

T1.5 wiring is partial in v0.40.7.0. Three follow-ups filed in TODOS.md under "v0.40.7.0 Schema Cathedral v3 follow-ups (v0.40.7+)" — enrichment-service.ts union widening ('person' | 'company'string), facts/eligibility.ts pack-aware ELIGIBLE_TYPES wiring, and 3 doctor checks (schema_pack_coverage, schema_pack_writability, schema_pack_mutation_audit).

Commands

Run gbrain --help or gbrain --tools-json for full command reference.

Key commands added in v0.7:

  • gbrain init — defaults to PGLite (no Supabase needed), scans repo size, suggests Supabase for 1000+ files
  • gbrain migrate --to supabase / gbrain migrate --to pglite — bidirectional engine migration

Key commands added for Minions (job queue):

  • gbrain jobs submit <name> [--params JSON] [--follow] [--dry-run] — submit a background job. v0.13.1 adds first-class flags for every MinionJobInput tuning knob: --max-stalled N, --backoff-type fixed|exponential, --backoff-delay Nms, --backoff-jitter 0..1, --timeout-ms N, --idempotency-key K.
  • gbrain jobs list [--status S] [--queue Q] — list jobs with filters
  • gbrain jobs get <id> — job details with attempt history
  • gbrain jobs cancel/retry/delete <id> — manage job lifecycle
  • gbrain jobs prune [--older-than 30d] — clean old completed/dead jobs
  • gbrain jobs stats — job health dashboard
  • gbrain jobs smoke [--sigkill-rescue] — health smoke test. --sigkill-rescue is the v0.13.1 regression guard for #219: simulates a killed worker and asserts the stalled job is requeued instead of dead-lettered on first stall.
  • gbrain jobs work [--queue Q] [--concurrency N] — start worker daemon (Postgres only)

Key commands added in v0.40.3.0 (contextual retrieval + cache gate + 4 CLI verbs):

  • gbrain mounts enable <id> / disable <id> — toggle a mount without removing it.
  • gbrain mounts trust-frontmatter <id> / untrust-frontmatter <id> — let a mounted brain's per-page contextual_retrieval_mode frontmatter override the source default. Off by default for mounts (security posture per D15); host source is always trusted.
  • gbrain sources set-cr-mode <id> <none|title|per_chunk_synopsis> — per-source contextual retrieval mode override. Pass unset or default to clear (NULL falls through to global). Missing source ID fails loudly with paste-ready gbrain sources list hint.
  • gbrain config set search.mode tokenmax triggers a mode-switch banner explaining the per-chunk Haiku synopsis backfill cost, and (on TTY + active Minion worker) offers to submit gbrain reindex --markdown as a Minion job. Non-TTY callers get a paste-ready hint to stderr instead of a silent stall. Suppress with GBRAIN_NO_MODE_SWITCH_UX=1 for CI fixtures.
  • Cache invalidation gate is now two-layer (per D11 codex finding): Layer 1 cheap MAX(generation) bookmark via pages_generation_idx, Layer 2 per-page snapshot via jsonb_each + LEFT JOIN pages. Pre-v0.40.3.0 cache rows with empty {} snapshot are vacuously valid (IRON-RULE backward compat).
  • KNOBS_HASH_VERSION bumped 3 → 5 (skipped past 4 reserved by salem's v0.40.4 graph-signals work). One-time cache-row invalidation on upgrade; refills within TTL.
  • Three new Minion handlers wired into RemediationStep consumer pattern: lint-fix, integrity-auto, sync-retry-failed. Thin wrappers around already-shipping CLI commands; NOT in PROTECTED_JOB_NAMES (idempotent, no shell exec, MCP-safe). sync-skip-failed deliberately NOT in this set per codex D12 Bug 3.
  • src/core/remediation-step.ts (NEW canonical module) exports RemediationStep type + makeRemediationStep factory + canonical-JSON idempotencyKey() per codex D12 Bug 2. Future doctor checks emit RemediationSteps via the factory instead of hand-rolling the shape.
  • src/core/search/query-cache-gate.ts (NEW) exports buildPageGenerationsSnapshot(engine, pageIds) + CACHE_GATE_WHERE_CLAUSE SQL fragment + validateCacheRowAgainstPages() pure validator. v0.41.25.0 (codex outside-voice on /plan-eng-review): Layer 1 bookmark read source switched from MAX(generation) FROM pages to SELECT value FROM page_generation_clock WHERE id = 1 at both store + lookup sites. Closes two pre-existing silent stale-cache bug classes that were independent of any sync work: (1) CDX-2 — UPDATE to a non-max page set NEW.generation = OLD.generation + 1 which didn't advance MAX, so cache silently served stale on every non-max UPDATE; (2) CDX-1 — DELETE didn't fire the row-level trigger at all, and even an AFTER DELETE wouldn't move MAX because surviving rows are untouched. The new clock is bumped per-statement by bump_page_generation_clock_trg (statement-level trigger created in migration v106) so every INSERT/UPDATE/DELETE statement advances the bookmark exactly once regardless of row cardinality (D19 — per-row would turn a 73K-row batch DELETE into 73K UPDATEs on the same counter). Also CDX-6/D20 fix: empty-result cache rows used to be "vacuously valid" via Layer 2's qc.page_generations = '{}'::jsonb shortcut, silently serving stale empty results across subsequent matching INSERTs. The shortcut is gone — empty snapshots now require Layer 1 to pass. CACHE_GATE_WHERE_CLAUSE enforces qc.page_generations <> '{}'::jsonb AND the per-page check (not OR). One-time post-upgrade cache miss spike on legacy {} rows is acceptable — the cache fills back up correctly and the clock seed COALESCE(MAX(pages.generation), 0) keeps non-empty legacy rows serving until the next write. Pinned by 5 new cases in test/page-generation-counter.test.ts (CDX-1/CDX-2/CDX-6 regressions + statement-level-trigger-fires-once contract) and 2 new e2e cases in test/e2e/cache-gate-pglite.test.ts. The existing pre-v0.41.25.0 vacuously valid legacy row IRON-RULE assertion in both files is intentionally inverted: that path was the CDX-6 bug.
  • src/core/search/mode-switch-ux.ts (NEW) exports summarizeTransition() 5-cell matrix + probeWorkerAvailable() worker liveness proxy via minion_jobs activity + buildReindexIdempotencyKey() content-stable key + runModeSwitchUx() orchestrator.

Key commands added in v0.41.31.0 (embedding cost-model + stale-semantics wave):

  • gbrain sync --all cost gate is now mode-aware. When embedding is DEFERRED (federated_v2 parallel path → per-source embed-backfill jobs), the gate is INFORMATIONAL — it prints the backfill cap, the current backlog cost, and the count of queued backfill jobs, and NEVER exits 2. When embedding runs INLINE (v2 off, or --serial without --no-embed), the gate BLOCKS only when the NEW-CONTENT estimate exceeds the floor. Cost previews now name the ACTUALLY-configured embedding model and price at its real rate (no more hardcoded OpenAI over-statement on cheaper models).
  • gbrain config set sync.cost_gate_min_usd <USD> — inline-path cost-gate floor (DB plane, default $0.50). Below this estimate the inline sync --all gate proceeds without blocking, killing the noise where a tiny incremental sync prompted on every cron. Set to 0 to block on any nonzero inline cost. Fail-open to the default on a missing/invalid value.
  • gbrain embed --stale now re-embeds pages whose embedding model/dimensions changed, not just pages with NULL embeddings. After a model or dims swap, the stored pages.embedding_signature differs from the current one, so the next embed --stale (or per-source embed-backfill job) invalidates and re-embeds those pages. GRANDFATHER: pages embedded before v0.41.31.0 have a NULL signature and are NEVER flagged stale, so the upgrade does NOT trigger a whole-corpus re-embed. embed --stale --dry-run counts signature-drift without mutating.
  • gbrain sources status gains a BACKFILL column (active(N) / queued(N) / idle) between EMBED and FAILS so an operator can see deferred embedding work after sync --all exits 0.

Key commands added in v0.36.4.0 (brain-health-100 wave):

  • gbrain doctor --remediation-plan [--target-score N] [--json] — preview the dependency-ordered plan that would drive the brain to target. JSON envelope is stable: each Remediation carries id, idempotency_key (content-hash for cron-safe retries), severity, est_seconds, est_usd_cost, and depends_on (referencing other ids). Empty recommendations array when the brain is already at target.
  • gbrain doctor --remediate [--yes] [--target-score N] [--max-usd N] — actually submit the plan. Walks dependency order, submits one Minion job per step, re-checks score between steps, refuses to spend past --max-usd (defaults: target=90, max-usd=infinite — but cron callers should always pass --max-usd). Bails when target exceeds maxReachableScore() for the brain (empty / under-configured brains) with a clear list of what's missing.
  • gbrain embed --stale --background — submit the embed sweep as a Minion job; print job_id=N to stdout; exit. Composable in shell pipelines. Add --background --follow to attach to the job's stderr stream (same UX as a direct call).
  • Eleven new Minion job types submittable via gbrain jobs submit <name>: reindex, repair-jsonb, orphans, integrity, purge, synthesize (PROTECTED), patterns (PROTECTED), consolidate (PROTECTED), extract_facts, resolve_symbol_edges, recompute_emotional_weight. PROTECTED ones reject MCP submission and require --allow-protected from a trusted local caller (CLI, autopilot, doctor --remediate).
  • gbrain autopilot (existing daemon) is now health-aware. Tick cost on a healthy brain drops from "full 6-phase cycle every 5 minutes" to "one SQL count, then sleep". Degraded brains get targeted handlers ([sync], [embed], [backlinks]) instead of the full cycle when the plan is small; large plans still get autopilot-cycle. The "60-minute full-cycle floor" runs the full phase set on a healthy brain at least every hour so phase-coupling invariants (lint-first, synthesize-before-patterns, embed-after-consolidate) keep getting exercised.

Key commands added in v0.32.7 (CJK fix wave):

  • gbrain reindex --markdown [--limit N] [--dry-run] [--json] [--no-embed] [--repo PATH] — operator-facing markdown re-chunk sweep. Walks pages with chunker_version < MARKDOWN_CHUNKER_VERSION (currently 2) and re-imports each with forceRechunk: true so the new chunker shape actually applies. Run automatically by gbrain upgrade's post-upgrade hook; available manually for triage.
  • gbrain doctor learns a new slug_fallback_audit check: surfaces info-severity entries from ~/.gbrain/audit/slug-fallback-YYYY-Www.jsonl (last 7 days) as an ok count when CJK / emoji / exotic-script filenames imported via the frontmatter-slug fallback path.
  • gbrain search "<CJK substring>" on PGLite brains now uses an ILIKE-based fallback with bigram-frequency-count ranking when the query contains Han / Hiragana / Katakana / Hangul Syllables. ASCII queries continue through websearch_to_tsquery('english') unchanged. Postgres-side CJK FTS still requires an extension (pgroonga / zhparser) — see v0.33+ TODO.
  • gbrain upgrade post-upgrade flow now prints a cost estimate before re-embedding: [chunker-bump] Will re-embed ~N markdown pages via <provider:model>, est. ~$X.XX, ~Ymin. Press Ctrl-C within 10s to abort. Sourced from real SQL counts + char totals; TTY-only wait (non-TTY auto-proceeds for CI / cron). Env overrides: GBRAIN_NO_REEMBED=1 bails out entirely with a doctor-warning marker; GBRAIN_REEMBED_GRACE_SECONDS=0 skips the wait.

Key commands added in v0.33.1.1 (Voyage 2048-dim correctness wave):

  • gbrain models doctor learns a new zero-token embedding_config probe that runs FIRST, before any chat/expansion probes spend money. Catches Voyage flexible-dim misconfigs at config time, not first-embed: embedding_model: voyage:voyage-4-large with embedding_dimensions outside {256, 512, 1024, 2048} (most commonly: embedding_dimensions left unset, falling back to the OpenAI default 1536 which Voyage rejects with an opaque HTTP 400). Surfaces a paste-ready gbrain config set embedding_dimensions <256|512|1024|2048> fix in both human and JSON output. New probe status 'config' joins {ok, model_not_found, auth, rate_limit, network, unknown}; new touchpoint label 'embedding_config' joins 'chat' and 'expansion'.
  • Voyage 2048-dim brains now actually embed at 2048 dims. embedding_model: voyage:voyage-4-large + embedding_dimensions: 2048 routes through the SDK-supported dimensions field, which voyageCompatFetch translates to Voyage's output_dimension on the wire. Same fix covers voyage-3-large, voyage-3.5, voyage-3.5-lite, voyage-4, voyage-4-lite, voyage-code-3. voyage-4-nano (open-weight, fixed 1024-dim) intentionally NOT in the flexible-dim allowlist — sending output_dimension to nano's endpoint produces an error.
  • Runtime validator: dimsProviderOptions() throws AIConfigError at the embed boundary with a paste-ready fix hint when a Voyage flexible-dim model is configured with an invalid dim — fail-loud even if you skipped gbrain models doctor.
  • VoyageResponseTooLargeError (new tagged class exported from src/core/ai/gateway.ts): the 256 MB per-response cap inside voyageCompatFetch was previously throwing a generic Error that the surrounding parse-error try/catch silently swallowed, returning the oversized response to the AI SDK anyway. Now thrown at both cap sites (Content-Length Layer 1, per-embedding base64 Layer 2) and rethrown from the catch via instanceof check — the cap is now actually effective.

Key commands added in v0.31.12 (model tier system + routing CLI):

  • gbrain models [--json] — read-only routing dashboard. Prints the four tier defaults (utility/reasoning/deep/subagent), the resolved value for each (after re-walking models.defaultmodels.tier.<tier> → env → TIER_DEFAULTS), every per-task override (models.dream.synthesize, models.dream.patterns, models.drift, models.auto_think, models.think, models.subagent, facts.extraction_model, models.eval.longmemeval, models.expansion, models.chat, models.dream.synthesize_verdict), the alias map (defaults + user overrides), and a source-of-truth column (default / config: <key> / env: <VAR>).
  • gbrain models doctor [--skip=<provider>] [--json] — 1-token reachability probe against each configured chat + expansion model. Classifies failures into {model_not_found, auth, rate_limit, network, unknown}. The structural fix for the bug class that motivated v0.31.12 (v0.31.6's claude-sonnet-4-6-20250929 chat default 404'd silently on every install).
  • Power-user model routing via config keys:
    • gbrain config set models.default opus — route every internal call (chat, expansion, synthesis, classification) through Opus 4.7. Subagent loop still falls back to claude-sonnet-4-6 automatically (Anthropic-only by construction).
    • gbrain config set models.tier.<tier> <model> — override one tier independently (utility / reasoning / deep / subagent).
    • gbrain config set models.aliases.frontier anthropic:claude-opus-4-7 — define an alias, then gbrain config set models.default frontier.
    • Per-task keys (e.g. gbrain config set models.dream.synthesize <model>) still beat tier overrides because they are more specific.
  • New subagent_provider check in gbrain doctor surfaces config drift if models.tier.subagent or models.default would route the Anthropic Messages API tool-loop to a non-Anthropic provider.
  • The skill at skills/conventions/model-routing.md was rewritten to cover both the new tier system AND the existing subagent spawn routing in one canonical doc (power-user recipes, three-layer enforcement explanation, override priority chain).

Key commands added in v0.28.1 (LongMemEval in the box):

  • gbrain eval longmemeval <dataset.jsonl> — run the public LongMemEval benchmark against gbrain hybrid retrieval. Flags: --limit N, --model M, --retrieval-only, --keyword-only, --expansion, --top-k K, --output FILE. One in-memory PGLite per benchmark run; TRUNCATE between questions over runtime-enumerated pg_tables (schema-migration-safe); ~/.gbrain never opened. --expansion defaults OFF (deterministic, no per-query Haiku). Default model resolves through resolveModel() 6-tier chain with new models.eval.longmemeval config key. gbrain eval longmemeval --help works without a configured brain (hermeticity gate).
  • Sanitization parity with takes: INJECTION_PATTERNS exported from src/core/think/sanitize.ts. The benchmark harness re-uses the same pattern set so adding a new injection pattern automatically covers takes AND benchmarks.
  • Hand the resulting JSONL to LongMemEval's published evaluate_qa.py to score (not bundled — needs OpenAI gpt-4o per their spec). Dataset: https://huggingface.co/datasets/xiaowu0162/longmemeval.

Key commands added in v0.26.5 (destructive-guard, end-to-end):

  • gbrain sources archive <id> — soft-delete a source. Hides from search via the new sources.archived column + cascading visibility filter. Preserves data for 72h. (PR #595 cherry-pick.)
  • gbrain sources restore <id> [--no-federate] — un-archive a soft-deleted source. Re-federates by default.
  • gbrain sources archived [--json] — list soft-deleted sources with their TTL.
  • gbrain sources purge [<id>] [--confirm-destructive] — permanent delete; with no id, purges all sources whose TTL expired.
  • gbrain sources remove <id> [--confirm-destructive] [--dry-run]--yes alone no longer enough on populated sources. Boxed impact preview before destruction.
  • gbrain pages purge-deleted [--older-than HOURS|Nd] [--dry-run] [--json] — operator escape hatch for page-level soft-delete cleanup. Mirror of gbrain sources purge. The autopilot cycle's new purge phase calls the same library function automatically every run.
  • MCP delete_page op semantically shifts from hard-delete to soft-delete. New ops: restore_page (scope: write), purge_deleted_pages (scope: admin, localOnly: true).
  • get_page and list_pages extended with include_deleted: boolean (default false).
  • New autopilot cycle phase purge (9th, runs after orphans). gbrain dream --phase purge runs only the purge sweep.
  • Index strategy note: the partial index pages_deleted_at_purge_idx ON pages (deleted_at) WHERE deleted_at IS NOT NULL supports the autopilot purge query. Search filters (WHERE deleted_at IS NULL) do NOT need their own index — soft-deleted cardinality stays low and Postgres won't use the partial index for the negative predicate. Don't add a regular (deleted_at) index without measuring.
  • Schema migration v34 (destructive_guard_columns) adds pages.deleted_at + the partial purge index; promotes archived from sources.config JSONB to real columns; backfills any pre-v0.26.5 JSONB shape.

Key commands added in v0.25.0:

  • gbrain eval export [--since DUR] [--limit N] [--tool query|search] — stream captured eval_candidates rows as NDJSON to stdout. Every line starts with "schema_version": 1 per the stable contract in docs/eval-capture.md. EPIPE-safe, progress heartbeats on stderr, deterministic ordering. Primary consumer is the sibling gbrain-evals repo for BrainBench-Real replay.
  • gbrain eval prune --older-than DUR [--dry-run] — explicit retention cleanup for eval_candidates. Requires --older-than (never deletes without a window). Duration strings: 30d, 7d, 1h, 90m, 3600s.
  • gbrain eval replay --against FILE.ndjson [--limit N] [--top-regressions K] [--json] [--verbose] — contributor-facing dev loop. Reads a captured NDJSON snapshot, re-runs each query / search op against the current brain, computes mean set-Jaccard@k between captured + current retrieved_slugs, top-1 stability rate, and latency Δ. JSON mode (schema_version: 1) for CI gating; human mode prints a regression table sorted worst-first. Closes the gap between "data captured" and "data used to gate a PR." See docs/eval-bench.md for the workflow.
  • gbrain eval cross-modal --task "..." --output <path> [--cycles N] [--slot-a-model ID] [--slot-b-model ID] [--slot-c-model ID] [--receipt-dir DIR] [--json] (v0.27.x) — multi-model quality gate. Three different-provider frontier models score the OUTPUT against the TASK on 5 documented dimensions. Pass criterion: every dim mean >=7 AND no model scored any dim <5. Exit codes: 0 PASS, 1 FAIL, 2 INCONCLUSIVE (<2/3 models returned parseable scores). Default cycles=3 in TTY, cycles=1 in non-TTY (limits accidental scripted bulk spend). Default slots: openai:gpt-4o / anthropic:claude-opus-4-7 / google:gemini-1.5-pro — refresh alongside model-family bumps. Receipts land at ~/.gbrain/.gbrain/eval-receipts/<slug>-<sha8-of-output>.json (gbrainPath honors GBRAIN_HOME). Bypasses connectEngine() via the cli.ts no-DB branch — runs cleanly before gbrain init. Reuses src/core/ai/gateway.ts:chat() for config/auth (no parallel provider stack). Cost-estimate prints to stderr before each cycle (T11=B partial cost guardrail; full --budget-usd N is a follow-up TODO).
  • gbrain doctor gains an eval_capture check: reads eval_capture_failures for the last 24h, groups by reason, warns when non-zero. Cross-process visibility (doctor runs in a separate process from MCP). Pre-v31 brains get Skipped (table unavailable) — non-fatal.
  • Config addition: eval: { capture?: boolean, scrub_pii?: boolean } in ~/.gbrain/config.json. File-plane onlygbrain config set writes the DB plane and does NOT control capture.
  • GBRAIN_CONTRIBUTOR_MODE=1 env var is the contributor-facing toggle. Capture is off by default as of v0.25.0; production users get a quiet brain. Resolution order: explicit eval.capture config wins both directions, then env var, then off. Documented in README.md, CONTRIBUTING.md, and docs/eval-bench.md.

Key commands added in v0.12.2:

  • gbrain repair-jsonb [--dry-run] [--json] — repair double-encoded JSONB rows left over from v0.12.0-and-earlier Postgres writes. Idempotent; PGLite no-ops. The v0_12_2 migration runs this automatically on gbrain upgrade.

Key commands added in v0.12.3:

  • gbrain orphans [--json] [--count] [--include-pseudo] — surface pages with zero inbound wikilinks, grouped by domain. Auto-generated/raw/pseudo pages filtered by default. Also exposed as find_orphans MCP operation. The natural consumer of the v0.12.0 knowledge graph layer: once edges are captured, find the gaps.
  • gbrain doctor gains two new reliability detection checks: jsonb_integrity (v0.12.0 Postgres double-encode damage) and markdown_body_completeness (pages truncated by the old splitBody bug). Detection only; fix hints point at gbrain repair-jsonb and gbrain sync --force.

Key commands added in v0.14.2:

  • gbrain sync --skip-failed — acknowledge the current set of failed-parse files recorded in ~/.gbrain/sync-failures.jsonl so the sync bookmark advances past them. Doctor's sync_failures check shows previously-skipped as "all acknowledged" instead of warning.
  • gbrain sync --retry-failed — re-walk the unacknowledged failures and re-attempt parsing. If the files now succeed, they clear from the set and the bookmark advances naturally.
  • gbrain apply-migrations --force-retry <version> — reset a wedged migration (3 consecutive partials with no completion) by appending a 'retry' marker. Next apply-migrations --yes treats the version as fresh. complete status never regresses to partial either before or after a retry marker.
  • GBRAIN_POOL_SIZE env var — honored by both the singleton pool (src/core/db.ts) and the parallel-import worker pool (src/commands/import.ts). Default is 10; lower to 2 for Supabase transaction pooler to avoid MaxClients crashes during gbrain upgrade subprocess spawns. Read at call time via resolvePoolSize().
  • gbrain doctor gains two new checks: sync_failures (surfaces unacknowledged parse failures with exact paths + fix hints) and brain_score (renders the 5-component breakdown when score < 100: embed coverage / 35, link density / 25, timeline coverage / 15, orphans / 15, dead links / 10 — sum equals total).

Key commands added in v0.26.0 (OAuth 2.1 + HTTP server + admin dashboard):

  • gbrain serve --http [--port 3131] [--token-ttl 3600] [--enable-dcr] [--log-full-params] — HTTP MCP server with OAuth 2.1, admin dashboard at /admin, SSE activity feed at /admin/events, health check at /health. Prints admin bootstrap token on first start. Alongside (not replacing) stdio gbrain serve. As of v0.26.9, mcp_request_log.params and the SSE feed default to a redacted summary ({redacted, kind, declared_keys, unknown_key_count, approx_bytes}); pass --log-full-params to log raw payloads on a personal laptop with a startup warning.
  • OAuth client registration — three paths:
    1. CLI: gbrain auth register-client <name> --grant-types <types> --scopes <scopes> (wired into src/commands/auth.ts as a thin wrapper over GBrainOAuthProvider.registerClientManual). Default grant types: client_credentials. Default scopes: read.
    2. Admin dashboard: Register client modal → credential reveal with Copy + Download JSON.
    3. SDK: oauthProvider.registerClientManual(name, grantTypes, scopes, redirectUris) for programmatic wrappers. --enable-dcr on serve --http opens the /register endpoint for RFC 7591 self-service registration (off by default).
  • gbrain auth create|list|revoke|test — legacy bearer tokens still work and grandfather to read+write+admin scopes on the OAuth server. auth is wired as a first-class gbrain subcommand in v0.26.0 (previously only invokable via bun run src/commands/auth.ts). No migration required to keep pre-v0.26 clients working.

Key commands added in v0.14.3 (fix wave):

  • gbrain doctor --index-audit — opt-in Postgres-only check reporting zero-scan indexes from pg_stat_user_indexes. Informational only; never auto-drops.
  • gbrain doctor schema_version check fails loudly when version=0 — catches bun install -g github:... postinstall failures (#218) and routes users to gbrain apply-migrations --yes.
  • gbrain jobs submit gains --max-stalled, --backoff-type, --backoff-delay, --backoff-jitter, --timeout-ms, --idempotency-key — exposing existing MinionJobInput fields as first-class CLI flags.
  • gbrain jobs smoke --sigkill-rescue — opt-in regression smoke case simulating a killed worker; asserts the v0.14.3 schema default (max_stalled=5) actually rescues on first stall.

Key commands added in v0.22.13 (PR #490):

  • gbrain sync --workers N (alias --concurrency N) — parallelize the import phase using per-worker Postgres engines (small pool of 2 each) with an atomic queue index. Auto-concurrency: defaults to 4 workers when the diff exceeds 100 files. Smaller diffs stay serial. Explicit --workers always wins (even on a 30-file diff). PGLite forces serial regardless. Validation rejects 0, negatives, non-integers loud (replaces the prior silent fall-through to auto-concurrency).
  • gbrain import --workers N — same parseWorkers() validation as sync; same try/finally worker-engine cleanup. Behavior surface unchanged.

Key commands added in v0.22.16 (claw-test friction loop):

  • gbrain claw-test [--scenario fresh-install|upgrade-from-v0.18] [--keep-tempdir] — scripted-mode CI gate that runs the full canonical first-day flow against a fresh tempdir. Asserts every expected --progress-json phase fired and doctor's status === 'ok'. ~30s, no API keys.
  • gbrain claw-test --live --agent openclaw — friction-discovery mode. Spawns real openclaw, hands it BRIEF.md, captures stdin/stdout/stderr to <run>/transcript.jsonl, lets the agent log friction via the friction CLI. Run on demand; ~510 min and ~$12 in tokens.
  • gbrain claw-test --list-agents — reports which agent runners are registered + their detection state (binary path or unavailable reason).
  • gbrain friction log --severity {confused|error|blocker|nit} --phase <name> --message <text> [--hint ...] [--kind {friction|delight}] [--run-id ...] — append a friction or delight entry to the active run JSONL.
  • gbrain friction render --run-id <id> [--json] [--transcripts] [--no-redact] — markdown report grouped by severity + phase; --redact is the default for md output (strips $HOME/$CWD placeholders so reports paste safely in PRs/issues).
  • gbrain friction list [--json] — recent run-ids with friction/delight counts; interrupted runs marked (interrupted).
  • gbrain friction summary --run-id <id> [--json] — two-column friction + delight summary.
  • GBRAIN_HOME env override is now honored uniformly across every gbrain write site (config, audit, friction, sync-failures, import checkpoint, integrity log, integrations heartbeat, migration rollback, etc.) — gbrainPath(...) from src/core/config.ts is the canonical helper. Read-side host-fingerprint detection (~/.claude/~/.openclaw etc.) intentionally NOT confined in v1; that's a v1.1 follow-up.

Testing

Test command tiers (v0.26.4 — parallel fast loop)

Five tiers of test commands, each with a clear scope:

Command What it runs Wallclock When to use
bun run test Parallel unit-test fast loop. 8-shard fan-out via scripts/run-unit-parallel.sh, then a serial pass over *.serial.test.ts. Excludes *.slow.test.ts and test/e2e/*. No pre-checks, no typecheck. ~85s on a Mac dev box (3650+ tests) Inner edit loop. Default.
bun run verify CI's authoritative pre-test gate set: check:privacy && check:jsonb && check:progress && check:wasm && bun run typecheck. The 4 checks .github/workflows/test.yml runs on shard 1 + typecheck. Single source of truth — CI literally calls bun run verify. ~12s (wasm-compile dominates) Before pushing; before /ship.
bun run test:full verify && bun run test && bun run test:slow && [smart e2e]. The local equivalent of "everything CI runs." Smart e2e: runs e2e only when DATABASE_URL is set; else loud skip notice to stderr. ~3-5min depending on slow + e2e Pre-merge sanity, before opening a PR.
bun run test:slow Just the *.slow.test.ts set (intentional cold-path correctness checks). seconds-to-minutes When touching slow-path code.
bun run test:serial Just the *.serial.test.ts set (cross-file-contention quarantine; runs at --max-concurrency=1). ~1s per quarantined file Debugging a specific quarantined file.
bun run test:e2e Real Postgres E2E. Requires Docker + DATABASE_URL. Sequential (template-DB parallelization is a v0.27+ TODO). ~5-10min Pre-ship; nightly.
bun run check:all All 7 historical pre-checks (privacy + jsonb + progress + no-legacy-getconnection + trailing-newline + wasm + exports-count). Superset of verify. ~10s Local-only sweep. The 4 not in verify are nice-to-haves.

CI vs local: intentionally divergent file sets

  • CI matrix (.github/workflows/test.yml) runs scripts/test-shard.sh 4-way, which uses FNV-1a hash bucketing and INCLUDES *.slow.test.ts. As of v0.31.4.1, CI EXCLUDES *.serial.test.ts from the hash buckets and runs them on shard 1 via bun run test:serial at --max-concurrency=1. Before that, serial files were hashed in alongside parallel files, which broke the mock.module quarantine (top-level mocks in serial files leaked into the parallel files they shared a shard process with — most visibly, eval-takes-quality-runner.serial.test.ts stubbed gateway.ts and broke every gateway.embedMultimodal test in voyage-multimodal.test.ts on shard 2). CI is the ground truth for "did everything pass."
  • Local fast loop (scripts/run-unit-shard.sh via the parallel wrapper) uses round-robin-by-index sharding and EXCLUDES *.slow.test.ts AND *.serial.test.ts. Local trades coverage for inner-loop speed; CI catches what local skips.

This divergence is intentional. Don't try to make them equal — the two scripts deliberately solve different problems. The regression test at test/scripts/run-unit-shard.test.ts pins what the local fast loop should and shouldn't include.

Failure-first logging

When bun run test finds any failure, the wrapper:

  1. Writes failure blocks (each prefixed with --- shard N: <test name> ---) to .context/test-failures.log (workspace-local, gitignored). On systems without a writable .context/, falls back to /tmp/gbrain-test-failures.log.
  2. Prints a loud stderr banner with the absolute log path, plus the last 30 lines of the failure log inlined. Banner survives | head / | tail / agent-side log truncation.
  3. Writes a one-line-per-shard summary to .context/test-summary.txt (shard N/M: pass=X fail=Y skip=Z rc=W).
  4. Exits non-zero. Empty failure log + non-zero exit = infrastructure problem (wedged shard, killed child); the banner says so.

If a shard wedges (per-shard GBRAIN_TEST_SHARD_TIMEOUT cap, default 600s), the wrapper writes --- shard N: WEDGED after ${SHARD_TIMEOUT}s --- to the failure log, includes the last 50 lines of the shard log, and proceeds with other shards' results.

File taxonomy

  • *.test.ts → fast loop (parallel 8-shard fan-out).
  • *.slow.test.ts → run via bun run test:slow only (intentional cold-path tests; would dominate the fast loop's wallclock).
  • *.serial.test.ts → run via bun run test:serial after the parallel pass completes; uses --max-concurrency=1. Quarantine for tests that share file-wide state and race when run alongside other files in the same bun test process. Currently: test/brain-registry.serial.test.ts, test/reconcile-links.serial.test.ts, test/core/cycle.serial.test.ts, test/embed.serial.test.ts (the latter two added in v0.26.7 — they use mock.module(...) which leaks across files in the shard process). Do not put the parallelism back on a serial file unless you've fixed the contention root cause (it just re-introduces the flake).
  • test/e2e/*.test.ts → real-Postgres E2E. Skipped when DATABASE_URL is unset.
  • tests/heavy/*.sh → ops-shape shell scripts. Cost minutes per run; NOT in default bun test. Run via bun run test:heavy or scheduled nightly via .github/workflows/heavy-tests.yml. Examples: pg_upgrade matrix (boot legacy brain → walk to head), RSS budget gate (measure peak worker RSS vs committed baseline), read-latency-under-sync (p50/p95/p99 under concurrent writer load), sync lock regression (N concurrent syncs assert 1 winner + N-1 lock-busy + zero leaked gbrain_cycle_locks rows). See tests/heavy/README.md for when to add a script here vs *.slow.test.ts. Files prefixed with _ (e.g. tests/heavy/_build_legacy_fixtures.sh) are helpers/libs invoked by sibling tests — the runner skips them.
  • test/fuzz/*.test.ts → property-based fuzz harness. Pure-validator targets in pure-validators.test.ts are guarded by scripts/check-fuzz-purity.sh (in bun run verify), which bun build --target=bun bundles each target and greps the resulting bundle for banned transitive imports (node:fs, node:child_process, engine modules). Anything that fails the guard moves to mixed-validators.test.ts (still property-tested, but no purity guarantee) or filesystem-validators.test.ts (fs-backed, uses temp dirs). Fuzz tests run in the default bun test loop because they're fast (~3s for ~12 properties × 1000 runs each).

The intra-file parallelism project (turn bun test into bun test --concurrent after sweeping shared-state contention sites) is sliced across v0.26.7 (foundation), v0.26.8 (env-mutation sweep), and v0.26.9 (PGLite sweep + codemod + measurement). v0.26.4 ships file-level parallelism only.

Test-isolation lint and helpers (v0.26.7)

The cross-file flake class is enforced statically by scripts/check-test-isolation.sh, wired into bun run verify and bun run check:all. Rules (non-serial unit files only; *.serial.test.ts and test/e2e/* are skipped):

Rule What it bans Fix
R1 process.env.X = ..., bracket assignment, delete process.env.X, Object.assign(process.env, ...), Reflect.set(process.env, ...) Use withEnv() from test/helpers/with-env.ts, OR rename file to *.serial.test.ts
R2 mock.module(...) anywhere in the file Rename file to *.serial.test.ts (no DI on production code for testability)
R3 new PGLiteEngine( outside ~50 lines after a beforeAll( line Use the canonical block (below) inside beforeAll(
R4 Files creating new PGLiteEngine( without engine.disconnect( inside an afterAll( block Add afterAll(() => engine.disconnect())

Files that violated these rules at the v0.26.7 baseline are listed in scripts/check-test-isolation.allowlist. The allow-list MUST shrink over time — never add new entries. v0.26.8 (env sweep) and v0.26.9 (PGLite sweep) remove entries as files get fixed.

Canonical PGLite block (R3 + R4 compliant)

Every test file that needs a PGLite engine should use this exact pattern:

import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';

let engine: PGLiteEngine;

beforeAll(async () => {
  engine = new PGLiteEngine();
  await engine.connect({});
  await engine.initSchema();
});

afterAll(async () => {
  await engine.disconnect();
});

beforeEach(async () => {
  await resetPgliteState(engine);
});

Why this exact shape: beforeAll creates a single engine per file (PGLite WASM cold-start + initSchema is ~20s); beforeEach truncates user data via resetPgliteState ("two orders of magnitude faster" than fresh-engine-per-test); afterAll disconnects so the engine doesn't leak across file boundaries within a shard process.

withEnv pattern (R1 fix)

import { withEnv } from './helpers/with-env.ts';

test('reads OPENAI_API_KEY', async () => {
  await withEnv({ OPENAI_API_KEY: 'sk-test' }, async () => {
    expect(loadConfig().openai_key).toBe('sk-test');
  });
});

// Delete a var (override is undefined):
await withEnv({ GBRAIN_HOME: undefined }, fn);

// Multiple keys:
await withEnv({ A: '1', B: '2', C: undefined }, fn);

withEnv saves the prior value of every key it touches and restores via try/finally — including when the callback throws. It is cross-test safe but NOT intra-file concurrent-safe. process.env is process-global; two test.concurrent() calls in the same file both touching the same key will race. Files using withEnv stay outside the future test.concurrent() codemod's eligibility filter.

When to quarantine instead of fix

Rename to *.serial.test.ts when:

  • The file uses mock.module(...) (R2 — there's no clean fix without changing production code).
  • The file is genuinely env-coupled (e.g. gbrain-home-isolation.test.ts, claw-test-cli.test.ts) — module-load env readers + ESM caching defeat dynamic-import-after-env tricks.
  • The file's tests intentionally share state across it() boundaries.

Quarantine count cap: 10 (informational). Beyond that, push back on the design.

Inventory (legacy)

bun test runs all tests. After the v0.12.1 release: ~75 unit test files + 8 E2E test files (1412 unit pass, 119 E2E when DATABASE_URL is set — skip gracefully otherwise). Unit tests run without a database. E2E tests skip gracefully when DATABASE_URL is not set.

Unit tests: test/markdown.test.ts (frontmatter parsing), test/chunkers/recursive.test.ts (chunking), test/parity.test.ts (operations contract parity), test/cli.test.ts (CLI structure), test/config.test.ts (config redaction), test/files.test.ts (MIME/hash), test/import-file.test.ts (import pipeline), test/upgrade.test.ts (schema migrations), test/file-migration.test.ts (file migration), test/file-resolver.test.ts (file resolution), test/import-resume.test.ts (import checkpoints), test/migrate.test.ts (migration; v8/v9 helper-btree-index SQL structural assertions + 1000-row wall-clock fixtures that guard the O(n²)→O(n log n) fix + v0.13.1 assertions on v12/v13 SQL shape, sqlFor + transaction:false runner semantics, the max_stalled DEFAULT 1 regression guard, and v0.22.6.1 v24 sqlFor.pglite: '' no-op assertion), test/bootstrap.test.ts (v0.22.6.1 — bootstrap contract: no-op on fresh install, idempotent across two initSchema() calls, no-op on modern brain that already has every probed column, full bootstrap path on simulated pre-v0.18 brain, fresh-install regression guard, pre-v0.13 links shape coverage), test/schema-bootstrap-coverage.test.ts (v0.22.6.1 CI guard — REQUIRED_BOOTSTRAP_COVERAGE lists every forward reference in PGLITE_SCHEMA_SQL; the test fails loudly if applyForwardReferenceBootstrap skips one. When you add a column-with-index to the embedded schema blob, you extend both arrays or this guard fails. The pattern that broke gbrain ten times in two years is now structurally prevented. v0.35.5.0: test now also parses src/core/migrate.ts source text for every ALTER TABLE ... ADD COLUMN (top-level sql:, sqlFor.{postgres,pglite} overrides, AND handler-body engine.runMigration(N, \ALTER TABLE ...`)), and asserts each (table, column) pair is covered by the bootstrap OR by the schema blob's CREATE TABLE bodies. Catches the column-only forward-reference class (e.g. sources.archivedshape from v0.26.5,oauth_clients.source_idfrom v0.34.1) that the pre-existing CREATE INDEX parser couldn't see. Pre-existing parser bug fixed in same wave:parseBaseTableColumnsnow strips SQL line + block comments before identifying column names so commented-out lines no longer hide adjacent columns from coverage.),test/helpers/schema-diff.ts+test/helpers/schema-diff.test.ts+test/e2e/schema-drift.test.ts(v0.26.6 #588 — cross-engine schema parity gate. Helper exports puresnapshotSchema(query)/diffSnapshots(pg, pglite, opts)/formatDiffForFailure(diff)/isCleanDiff(diff) over a four-tuple per column (data_type, udt_name, is_nullable, column_default). E2E test spins up fresh PGLite + Postgres, runs engine.initSchema()on each (bootstrap + schema replay + migrations), snapshotsinformation_schema.columns, then diffs. 2-table allowlist (files, file_migration_ledger) — every other Postgres table must reach PGLite via PGLITE_SCHEMA_SQL or a migration's sqlFor.pglitebranch. Sentinels foroauth_clients, mcp_request_log, access_tokens, eval_candidatesgive tighter blame messages. Skip-gracefully withoutDATABASE_URL. Wired into scripts/e2e-test-map.tsso changes tosrc/schema.sql, src/core/pglite-schema.ts, or src/core/migrate.tstrigger it. The failure message names every drift with a paste-ready hint pointing atsrc/core/pglite-schema.ts.), test/setup-branching.test.ts(setup flow),test/slug-validation.test.ts(slug validation),test/storage.test.ts(storage backends),test/supabase-admin.test.ts(Supabase admin),test/yaml-lite.test.ts(YAML parsing),test/check-update.test.ts(version check + update CLI),test/pglite-engine.test.ts(PGLite engine, all 40 BrainEngine methods including 11 cases foraddLinksBatch/addTimelineEntriesBatch: empty batch, missing optionals, within-batch dedup via ON CONFLICT, missing-slug rows dropped by JOIN, half-existing batch, batch of 100 + v0.13.1 connect()error-wrap assertion (original error nested, #223 link in message, lock released)),test/engine-factory.test.ts(engine factory + dynamic imports),test/integrations.test.ts(recipe parsing, CLI routing, recipe validation),test/publish.test.ts(content stripping, encryption, password generation, HTML output),test/backlinks.test.ts(entity extraction, back-link detection, timeline entry generation),test/lint.test.ts(LLM artifact detection, code fence stripping, frontmatter validation),test/report.test.ts(report format, directory structure),test/skills-conformance.test.ts(skill frontmatter + required sections validation),test/resolver.test.ts(RESOLVER.md coverage, routing validation + v0.20.4 round-trip: every quoted RESOLVER.md trigger must match a frontmattertriggers:entry in the target skill, and everyname=""reference in any SKILL.md must resolve to a declared op insrc/core/operations.tsor a Minions handler inPROTECTED_JOB_NAMES), test/search.test.ts(RRF normalization, compiled truth boost, cosine similarity, dedup key),test/sql-ranking.test.ts(v0.22.0 source-boost helpers: 39 cases covering longest-prefix-match in SQL CASE, detail=high temporal-bypass, three-meta-char LIKE escape (%, _, \\), single-quote SQL-literal doubling, env override parsing for GBRAIN_SOURCE_BOOST + GBRAIN_SEARCH_EXCLUDE, resolveBoostMap / resolveHardExcludes merge semantics),test/dedup.test.ts(source-aware dedup, compiled truth guarantee, layer interactions),test/intent.test.ts(query intent classification: entity/temporal/event/general),test/eval.test.ts(retrieval metrics: precisionAtK, recallAtK, mrr, ndcgAtK, parseQrels),test/check-resolvable.test.ts(resolver reachability, MECE overlap, gap detection, DRY checks + v0.14.1 proximity-based DRY detection +extractDelegationTargetscoverage — 13 DRY cases),test/dry-fix.test.ts(v0.14.1 auto-fix: three shape-aware expander pure-function tests, five guards — working-tree-dirty, no-git-backup, inside-code-fence, already-delegated within 40 lines, ambiguous-multi-match, block-is-callout — 28 cases),test/doctor-fix.test.ts(v0.14.1gbrain doctor --fixCLI integration: dry-run preview, apply path, JSON output shape — 3 cases),test/backoff.test.ts(load-aware throttling, concurrency limits, active hours),test/fail-improve.test.ts(deterministic/LLM cascade, JSONL logging, test generation, rotation),test/transcription.test.ts(provider detection, format validation, API key errors),test/enrichment-service.test.ts(entity slugification, extraction, tier escalation),test/data-research.test.ts(recipe validation, MRR/ARR extraction, dedup, tracker parsing, HTML stripping),test/minions.test.ts(Minions job queue v7: CRUD, state machine, backoff, stall detection, dependencies, worker lifecycle, lock management, claim mechanics, depth/child-cap, timeouts, cascade kill, idempotency, child_done inbox, attachments, removeOnComplete/Fail + v0.13.1max_stalledclamp/default/plumbing coverage),test/extract.test.ts(link extraction, timeline extraction, frontmatter parsing, directory type inference),test/extract-db.test.ts(gbrain extract --source db: typed link inference, idempotency, --type filter, --dry-run JSON output),test/extract-fs.test.ts(gbrain extract --source fs: first-run inserts + second-run reports zero, dry-run dedups candidates across files, second-run perf regression guard — the v0.12.1 N+1 dedup bug),test/link-extraction.test.ts(canonical extractEntityRefs both formats, extractPageLinks dedup, inferLinkType heuristics, parseTimelineEntries date variants, isAutoLinkEnabled config),test/graph-query.test.ts(direction in/out/both, type filter, indented tree output),test/features.test.ts(feature scanning, brain_score calculation, CLI routing, persistence),test/file-upload-security.test.ts(symlink traversal, cwd confinement, slug + filename allowlists, remote vs local trust),test/query-sanitization.test.ts(prompt-injection stripping, output sanitization, structural boundary),test/search-limit.test.ts(clampSearchLimit default/cap behavior across list_pages and get_ingest_log),test/repair-jsonb.test.ts(v0.12.2 JSONB repair: TARGETS list, idempotency, engine-awareness),test/migrations-v0_12_2.test.ts(v0.12.2 orchestrator phases: schema → repair → verify → record),test/markdown.test.ts(splitBody sentinel precedence, horizontal-rule preservation, inferType wiki subtypes),test/orphans.test.ts(v0.12.3 orphans command: detection, pseudo filtering, text/json/count outputs, MCP op),test/postgres-engine.test.ts(v0.12.3 statement_timeout scoping:sql.begin+SET LOCALshape, source-level grep guardrail against reintroduced bareSET statement_timeout), test/sync.test.ts(sync logic + v0.12.3 regression guard asserting top-levelengine.transactionis not called),test/sync-concurrency.test.ts(v0.22.13 PR #490: 17 cases coveringautoConcurrency()thresholds + PGLite-forces-serial + explicit-override clamping,shouldRunParallel()Q1 explicit-bypasses-floor contract, andparseWorkers()validation that rejects'0'/'-3'/'foo'/'1.5'/trailing chars), test/sync-parallel.test.ts(v0.22.13 PR #490: PGLite-routed coverage of the bookmark gate under concurrency request, head-drift gate, vanished-file failure capture, PGLite-stays-serial, and thegbrain-syncwriter-lock contract — 7 cases),test/sync-failures.test.ts(v0.22.12: 28 cases pinningclassifyErrorCoderegex coverage for all 12 codes against literal production message strings frommarkdown.ts:159-244andimport-file.ts:199, 347, 352, 401; summarizeFailuresByCodesort + pre-classified-honor;recordSyncFailurescode-field persistence;acknowledgeSyncFailuresAcknowledgeResult shape + backfill on pre-v0.22.12 entries),test/doctor.test.ts(doctor command + v0.12.3 assertions thatjsonb_integrityscans the four v0.12.0 write sites andmarkdown_body_completenessis present),test/utils.test.ts(shared SQL utilities +tryParseEmbeddingnull-return and single-warn semantics),test/build-llms.test.ts(llms.txt/llms-full.txt generator: path resolution, idempotence, spec shape, regen-drift guard, content contract, AGENTS.md install-path mirror, size-budget enforcement — 7 cases),test/oauth.test.ts(v0.26.0 OAuth 2.1 provider — 27 cases: register, getClient,client_credentialsgrant exchange,authorization_codeflow with PKCE challenge / verifier, refresh token rotation,verifyAccessTokenwith both OAuth + legacyaccess_tokensfallback,revokeToken, sweepExpiredTokens, and a contract test asserting scope+localOnlyannotations are set correctly on all 30 operations; **v0.26.2** adds 5coerceTimestamp unit cases (null/undefined/string/number/throw-on-NaN), NULL-expires_at-as-expired contract tests for both refresh + access token paths, and a cascade-delete contract test asserting revoke-clientpurgesoauth_tokens+oauth_codesrows via FK CASCADE; **v0.26.9** adds 14 cases pinning the F1/F2/F3/F4/F5/F6/F7c/F12 invariants, including the F1/F4 cross-client isolation pattern (wrong-client attempt MUST reject AND rightful owner MUST still succeed atomically afterward) and the empty-stringredirect_uribypass guard surfaced during adversarial review),test/mcp-dispatch-summarize.test.ts(v0.26.9 — 7 cases pinning F8summarizeMcpParamsinvariants: declared-keys allow-list intersection, attacker-key-name leak guard (unknown keys counted not named), 1KB byte bucketing for size-probe defense, missing op falls through to fully-redacted shape, declared-keys sorted for deterministic output),test/trust-boundary-contract.test.ts(v0.26.9 — 4 cases pinning F7b fail-closed semantics under cast bypass:ctx.remote === undefinedtreated as remote/untrusted at every flipped call site,as anyandPartial<>spreads can't downgrade trust by accident),test/check-resolvable-cli.test.ts(v0.19 CLI wrapper: exit codes, JSON envelope shape, AGENTS.md fallback chain),test/regression-v0_16_4.test.ts(findRepoRoot regression guard — hermetic startDir parameterization),test/repo-root.test.ts(v0.16.4 / v0.19 / v0.31.7 — 20 cases:findRepoRootwalk semantics + default-arg parity, the 4-tierautoDetectSkillsDir fallback chain ($OPENCLAW_WORKSPACE~/.openclaw/workspace→ repo-root →./skills), W1 RESOLVER.md/AGENTS.md filename precedence, D-CX-4 explicit-env-wins-over-repo-root, and 8 new v0.31.7 D3+D5 cases pinning tier-0 $GBRAIN_SKILLS_DIRvalid/invalid/precedence-over-OPENCLAW_WORKSPACE, the install-path walk inautoDetectSkillsDirReadOnly, no-drift on primary success, AUTO_DETECT_HINT+AUTO_DETECT_HINT_READ_ONLYcontent, and the D5 regression guard asserting the sharedautoDetectSkillsDirMUST NEVER return'install_path'source — that's how the read-path/write-path split stays safe),test/resolver-merge.test.ts(v0.31.7 — 8 cases pinning the multi-file resolver merge:findAllResolverFilesempty / RESOLVER.md-only / AGENTS.md-only / both-present (RESOLVER.md first), andcheckResolvablemerge semantics acrossskills/RESOLVER.md+../AGENTS.mdfor the OpenClaw layout where the skillpack ships a thin RESOLVER.md and the real dispatcher lives at the workspace root — dedup byskillPath(first occurrence wins), AGENTS.md-at-workspace-root works alone, and the previously-unreachable 187/224 OpenClaw skills become reachable),test/filing-audit.test.ts(v0.19 Check 6:writes_pages/writes_tofrontmatter, filing-rules JSON validation),test/skill-brain-first.test.ts(v0.37.1.0 — 56 cases: shared frontmatter parser,analyzeSkillBrainFirstcompliance ladder across 9 fixtures undertest/fixtures/brain-first-skills/(compliant-callout, compliant-phase, compliant-position, exempt-frontmatter, missing-brain-first, multi-pattern, negation-prose, no-external, typo-frontmatter), offset helpers, external-lookup regex shape, audit snapshot+diff transition logic, PR #1206FORMERLY_HARDCODED_EXEMPTregression absorption),test/e2e/skill-brain-first.test.ts(v0.37.1.0 — 12 E2E cases: doctor reportsskill_brain_firstcheck with structured issues;--fix --dry-runpreviews insertion without writing;--fixapplies the canonical Convention callout idempotently;brain_first: exemptfrontmatter resolves the warn;brain_first_typosurfaces paste-ready hint; audit JSONL recordsdetected/resolved/fixedtransitions; stable brain emits 0 audit lines/run),test/routing-eval.test.ts(v0.19 Check 5: fixture parsing, structural routing, ambiguous_with, Haiku tie-break layer),test/skill-manifest.test.ts(v0.19 skill manifest parser: drift detection, managed-block markers),test/skillify-scaffold.test.ts(v0.19gbrain skillify scaffoldstubs: SKILL.md, script, tests, routing-eval fixtures),test/skillpack-install.test.ts(v0.19gbrain skillpack installmanaged-block install / update / no-clobber semantics),test/skillpack-sync-guard.test.ts(v0.19 sync-guard: bundled skills stay byte-identical toskills/source),test/http-transport.test.ts(v0.22.7 HTTP transport: 23 unit cases covering bearer auth + missing/no-Bearer/unknown/revoked +/healthbypass, F1+F2 round-trip via dispatch.ts, F3 invalid_params, application/json response shape (not SSE), CORS default-deny + allowlist, body cap on Content-Length AND chunked, two-bucket rate limit (refill, exhaust+Retry-After, LRU eviction, TTL prune, pre-auth IP fires before DB), andmcp_request_logaudit on success + auth_failed),test/restart-sweep.test.ts(v0.28.3 — 27 bun:test cases for therecipes/restart-sweep.mdinlined script: sentinel-anchored fenced-block extraction with salted tmp filenames to bypass ESM cache; constructor-time env reads (proves no module-load snapshot); idempotency layer load/save/atomic-tmp-rename/corrupt-JSON-recovery/30-day-prune;(sessionKey, lastAlertedAt)cooldown gate with 6h threshold (the C1 fix that survives synthesized restartTime); AGGRESSIVE-gate two-state tests; execFile argv shape proving shell metachars inOPENCLAW_TELEGRAM_GROUPcannot reach/bin/sh; real-\n-not-literal alert formatting; GBRAIN_HOMEstate path override),test/eval-longmemeval.test.ts(v0.28.8 LongMemEval harness — 12 hermetic cases with noDATABASE_URLand no API keys: PGLite create + reset over runtime-enumeratedpg_tables, infrastructure-table preservation across resets, JSONL question parsing, retrieval-only and answer-gen modes via stubbed ThinkLLMClient, --limitcutoff,--keyword-onlyvs hybrid, default--expansion=offbehavior, perf gate (p50 < 30ms / p99 < 50ms warm reset+import+search on Apple Silicon),--helpworks without a configured brain, fixture round-trip viatest/fixtures/longmemeval-mini.jsonl), test/longmemeval-sanitize.test.ts(v0.28.8 sanitization parity: 12 cases pinning thatINJECTION_PATTERNSfromsrc/core/think/sanitize.tsis the single source of truth — adding a pattern there must cover bothframing and<chat_session>` framing, no per-surface regex drift).

E2E tests (test/e2e/): Run against real Postgres+pgvector. Require DATABASE_URL.

  • bun run test:e2e runs Tier 1 (mechanical, all operations, no API keys). Includes 9 dedicated cases for the postgres-engine addLinksBatch / addTimelineEntriesBatch bind path — postgres-js's unnest() binding is structurally different from PGLite's and gets its own coverage.
  • test/e2e/search-quality.test.ts runs search quality E2E against PGLite (no API keys, in-memory)
  • test/e2e/graph-quality.test.ts runs the v0.10.3 knowledge graph pipeline (auto-link via put_page, reconciliation, traversePaths) against PGLite in-memory
  • test/e2e/postgres-jsonb.test.ts — v0.12.2 regression test. Round-trips all 5 JSONB write sites (pages.frontmatter, raw_data.data, ingest_log.pages_updated, files.metadata, page_versions.frontmatter) against real Postgres and asserts jsonb_typeof='object' plus ->>'key' returns the expected scalar. The test that should have caught the original double-encode bug.
  • test/e2e/integrity-batch.test.ts (v0.22.8) — parity tests for scanIntegrity's batch-load fast path vs sequential. Four cases (dedup, hits, validate, topPages) seed a fixture and assert both paths return identical results. Dedup case uses raw SQL via getConn().unsafe() to seed a (test-source-2, people/alice) row alongside the default-source row, since engine.putPage doesn't take a source_id. Pins the codex-caught multi-source overcounting regression.
  • test/e2e/jsonb-roundtrip.test.ts — v0.12.3 companion regression against the 4 doctor-scanned JSONB sites. Assertion-level overlap with postgres-jsonb.test.ts is intentional defense-in-depth: if doctor's scan surface ever drifts from the actual write surface, one of these tests catches it.
  • test/e2e/sync.test.ts (v0.22.12 — --skip-failed failure-loop test, alongside the existing 13 happy-path tests): exercises the full chain — broken file → performSync returns blocked_by_failures with grouped breakdown → performSync({skipFailed: true}) advances bookmark and returns AcknowledgeResult with code summary → second broken file → second cycle. Saves and restores the user's real ~/.gbrain/sync-failures.jsonl so the test is hermetic on a developer machine. Asserts bookmark gating, JSONL state, dedup across paths, summary aggregation, and the literal doctor-rendering string format. This is the integration test that proves the v0.22.12 chain holds together — unit tests cover the pure functions in isolation, this covers the integration.
  • test/e2e/upgrade.test.ts runs check-update E2E against real GitHub API (network required)
  • test/e2e/minions-shell-pglite.test.ts (v0.20.4) exercises the PGLite --follow inline shell-job path (in-memory, no DATABASE_URL required) — the path the consolidated minion-orchestrator skill documents for dev use
  • test/e2e/openclaw-reference-compat.test.ts (v0.19) — exercises check-resolvable + skillpack install against a minimal AGENTS.md workspace fixture (test/fixtures/openclaw-reference-minimal/), regression guard for the 107-skill OpenClaw deployment shape
  • test/e2e/search-swamp.test.ts (v0.22.0) — reproduces the headline source-swamp case. Seeds a curated originals/talks/article-outline-fat-code page against two wintermute/chat/ pages stuffed with the same multi-word phrase. Asserts the article wins keyword AND vector ranking, that detail=high lets the chat swamp re-surface (temporal-query workflow preserved), and that source_id passes through the two-stage CTE intact. PGLite in-memory.
  • test/e2e/search-exclude.test.ts (v0.22.0) — verifies test/ + archive/ pages are hidden by default, that include_slug_prefixes opts back in, and that caller-supplied exclude_slug_prefixes adds to defaults. Both keyword and vector search paths covered.
  • test/e2e/engine-parity.test.ts (v0.22.0) — Postgres ↔ PGLite top-result and result-set parity for searchKeyword + searchVector. Codex flagged that Postgres ranks pages then picks best chunk while PGLite returns chunks directly — without parity coverage the source-boost fix could pass on PGLite and fail on Postgres. Skips gracefully when DATABASE_URL is unset.
  • test/e2e/postgres-bootstrap.test.ts (v0.22.6.1) — exercises PostgresEngine.initSchema() directly against a fresh real Postgres database. Asserts the bootstrap path is no-op on fresh installs and that SCHEMA_SQL replays cleanly through the engine path (not via the standalone db.initSchema from src/core/db.ts, which would have produced false-positive coverage). Codex caught the E2E-shape gap during plan review.
  • test/e2e/http-transport.test.ts (v0.22.7) — 8 cases against real Postgres covering gbrain serve --http end-to-end: bearer auth round-trip, last_used_at SQL-level debounce semantics, mcp_request_log row insertion on success and auth_failed paths, /health DB-down → 503 (DB-probing health check), and the F1+F2+F3 dispatch round-trip with a real operation. Skips gracefully when DATABASE_URL is unset.
  • test/e2e/serve-http-oauth.test.ts (v0.26.0, expanded v0.26.2, expanded v0.26.9) — real-Postgres E2E against gbrain serve --http with full OAuth 2.1. Spawns a subprocess server, registers a client via the CLI, mints client_credentials tokens, exercises the /mcp JSON-RPC pipeline. v0.26.2 adds: real DCR /register HTTP-level response-shape test (asserts typeof body.client_id_issued_at === 'number' over the wire — RFC 7591 §3.2.1 spec compliance, not just internal-store shape); real CLI subprocess test for revoke-client (registers → mints token → revokes via execSync → asserts token rejected at /mcp → asserts re-run exits 1); server fixture flips on --enable-dcr so /register is reachable. bun execSync env-inheritance fix: bun's execSync does NOT inherit env mutations done via process.env.X = ..., only OS-level env from before bun started. helpers.ts loads .env.testing and sets DATABASE_URL via process.env mutation, which is invisible to subprocesses unless env: { ...process.env } is passed explicitly — every subprocess call in this file passes env: { ...process.env } for that reason. Reference fix for the next maintainer hitting the same failure mode in sibling sync/cycle/dream/claw-test E2Es. afterAll cleanup is guarded on clientId (won't throw if beforeAll failed before registration); cleanup errors surface to stderr without throwing so real test failures aren't masked. Tracks DCR-registered clients alongside the manual one. v0.26.9 adds 2 regressions for the F7 trust-boundary fix: an HTTP MCP submit_job for name: "shell" MUST reject with a permission error (proving the request handler now sets remote: true and submit_job's protected-name guard fires), and the same guard rejects subagent submission. Closes the OAuth-token-to-RCE escalation path. Skips gracefully when DATABASE_URL is unset.
  • test/e2e/sync-parallel.test.ts (v0.22.13 PR #490) — DATABASE_URL-gated. T2: 60-file Postgres sync at concurrency=4 imports all + no connection leak (probes pg_stat_activity before/after to confirm worker engines disconnected). P4: 120-file serial-vs-parallel benchmark prints SYNC_PARALLEL_BENCH N files | serial=Xms | parallel(4)=Yms | speedup=Zx for CHANGELOG quoting. Asserts parallel ≤ serial × 1.5 (CI-noise tolerant; not a strict speedup gate).
  • test/e2e/multi-source-bug-class.test.ts (v0.32.8, PR #860) — 7-case PGLite in-memory regression suite pinning every bug site fixed in this PR: listAllPageRefs ordering by (source_id, slug) (F11), getPage with sourceId picks the right (source, slug) row (F2), extract-takes processes both overlapping people/alice rows independently, listPages filters correctly with PageFilters.sourceId, addLinksBatch with from/to_source_id targets the right rows (F4), validateSourceId rejects path traversal (F6), reverse-write disk layout uses brainDir/.sources/<id>/<slug>.md for non-default sources (F6). No DATABASE_URL needed. Wired into scripts/e2e-test-map.ts so changes to extract-takes / patterns / synthesize / embed / extract / migrate-engine auto-trigger this test. Companion: test/e2e/integrity-batch.test.ts's "multi-source duplicate slugs scan once" case was pinning the pre-fix bug — assertion flipped in v0.32.8 to expect both batch + sequential paths report 2.
  • test/e2e/source-isolation-pglite.test.ts (v0.34.1.0, #861) — 14-case PGLite in-memory regression suite pinning the source-isolation P0 seal at two layers. Engine layer: searchKeyword / searchVector / searchKeywordChunks / listPages / getPage / traverseGraph / traversePaths apply sourceId (scalar fast path) and sourceIds (array path) correctly across both engines. Op-handler layer: routes through sourceScopeOpts(ctx) so a read+write-scoped OAuth client bound to --source dept-x cannot see rows from neighboring sources via search, query, list_pages, get_page, or find_experts. Covers both ctx.sourceId (single-source clients) and ctx.auth.allowedSources (federated_read clients) precedence; federated array wins over scalar wins over nothing. No DATABASE_URL needed.
  • test/openai-compat-multimodal.test.ts (v0.34.1.0, #875) — 11-case unit suite for the gateway's openai-compatible multimodal path: happy-path single + multi-input embedding, unauthenticated proxy mode, dimension-mismatch guard (D12; throws AIConfigError with model id + observed + expected pre-storage), default-dim fallback when recipe declares default_dims, HTTP 401 / 400 / malformed-JSON / non-array error paths, plus a regression test that the existing Voyage /multimodalembeddings recipe still routes through its dedicated path (not the openai-compatible one). Hermetic via the __setEmbedTransportForTests seam.
  • test/serve-stdio-lifecycle.test.ts (extended v0.34.1.0, #870) — adds 3 new cases for the MCP_STDIO=1 env guard: stdin EOF does NOT trigger shutdown when the env is set, SIGTERM still does (guard scope is correct), unset env preserves the pre-v0.34 CLI lifecycle. Exercises the ServeOptions.mcpStdio?: boolean test seam directly so tests don't mutate process.env.
  • test/oauth.test.ts (extended v0.34.1.0, #909) — 5 new cases for the PKCE DCR public-client gate: registerClient with token_endpoint_auth_method: "none" returns no client_secret field on the public client, default client_secret_post clients still get the one-time-reveal secret, getClient NULL→undefined normalization so the SDK's clientAuth path accepts public clients, full PKCE /authorize/token round-trip against a public client (no client_secret presented), and a regression test that the public-vs-confidential branch doesn't break confidential client client_secret_post exchange.
  • Tier 2 (skills.test.ts) requires OpenClaw + API keys, runs nightly in CI
  • If .env.testing doesn't exist in this directory, check sibling worktrees for one: find ../ -maxdepth 2 -name .env.testing -print -quit and copy it here if found.
  • Run E2E tests without asking permission. When you want to verify behavior, there's a relevant E2E test, or you're shipping touching anything covered by an E2E suite — just spin up the test DB, run the tests, and tear down. Don't ask, don't propose it, don't defer. The lifecycle is short (~2-30s startup, sub-minute tests, instant teardown) and the gate value is high. Skipping with "DATABASE_URL unset" is silent regression, not caution.

API keys and running ALL tests

ALWAYS source the user's shell profile before running tests:

source ~/.zshrc 2>/dev/null || true

This loads OPENAI_API_KEY and ANTHROPIC_API_KEY. Without these, Tier 2 tests skip silently. Do NOT skip Tier 2 tests just because they require API keys — load the keys and run them.

When asked to "run all E2E tests" or "run tests", that means ALL tiers:

  • Tier 1: bun run test:e2e (mechanical, sync, upgrade — no API keys needed)
  • Tier 2: test/e2e/skills.test.ts (requires OpenAI + Anthropic + openclaw CLI)
  • Always spin up the test DB, source zshrc, run everything, tear down.

E2E test DB lifecycle (ALWAYS follow this)

You are responsible for spinning up and tearing down the test Postgres container. Do not leave containers running after tests. Do not skip E2E tests, do not ask permission to run them — see the "run without asking" rule above.

  1. Check for .env.testing — if missing, copy from sibling worktree. Read it to get the DATABASE_URL (it has the port number).
  2. Check if the port is free: docker ps --filter "publish=PORT" — if another container is on that port, pick a different port (try 5435, 5436, 5437) and start on that one instead.
  3. Start the test DB:
    docker run -d --name gbrain-test-pg \
      -e POSTGRES_USER=postgres -e POSTGRES_PASSWORD=postgres \
      -e POSTGRES_DB=gbrain_test \
      -p PORT:5432 pgvector/pgvector:pg16
    
    Wait for ready: docker exec gbrain-test-pg pg_isready -U postgres
  4. Bootstrap the schema (required — fresh containers have no oauth_clients, mcp_request_log, pages etc.; tests like serve-http-oauth.test.ts will fail with relation "oauth_clients" does not exist if you skip this):
    DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test \
      bun run src/cli.ts doctor --json > /dev/null 2>&1
    
    gbrain doctor triggers initSchema() on first connect, which is the canonical way to bring a fresh DB to head. apply-migrations --yes alone does NOT seed the base schema — it runs ALTER-style migrations on top of initSchema. Tests that bypass the engine (raw execSync-spawned auth register-client) hit the schema directly and need this step to have run first.
  5. Run E2E tests: DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test bun run test:e2e
  6. Tear down immediately after tests finish (pass or fail): docker stop gbrain-test-pg && docker rm gbrain-test-pg

Never leave gbrain-test-pg running. If you find a stale one from a previous run, stop and remove it before starting a new one.

Search Mode (v0.32.3)

GBrain ships three named search modes that bundle the search-lite knobs from PR #897 into a single config key. Pick one at install time; the rest of the project resolves through src/core/search/mode.ts.

Knob conservative balanced tokenmax
cache.enabled true true true
cache.similarity_threshold 0.92 0.92 0.92
cache.ttl_seconds 3600 3600 3600
intentWeighting true true true
tokenBudget 4000 12000 off
expansion (LLM multi-query) false false true
searchLimit default 10 25 50

Cost anchors (downstream agent input cost — gbrain itself is rounding error). The corner-to-corner spread is 25x once you pair mode with downstream model. Chunks ~400 tokens avg. Per-query cost @ 10K queries/month (typical single-user volume), full search payload, no cache savings:

Mode \ Downstream Haiku 4.5 ($1/M) Sonnet 4.6 ($3/M) Opus 4.7 ($5/M)
conservative (~4K) $40/mo $120/mo $200/mo
balanced (~10K) $100/mo $300/mo $500/mo
tokenmax (~20K) $200/mo $600/mo $1,000/mo

Scales linearly: multiply by 10 for 100K/mo (heavy power user / multi-user fleet); divide by 10 for 1K/mo (light usage). Natural pairings span ~4x. Mismatches (tokenmax+Haiku, conservative+Opus) waste capacity differently — too-big payload overwhelms a cheap model; too-small payload starves an expensive one.

tokenmax adds ~$1.50 per 1K queries in Haiku expansion calls on top of the matrix ($15/mo @ 10K). Cache hits cut all numbers ~50%. The cost picker copy in gbrain init carries the same matrix verbatim — update both when refreshing.

Per-query math vs real-world spend. The matrix above is what an isolated benchmark would measure. Real agent loops with disciplined Anthropic prompt caching see 50-80% discount on top (cache hits skip downstream entirely). The realistic-scale anchor in docs/eval/SEARCH_MODE_METHODOLOGY.md walks the natural pairings at single-power-user volume (~860 turns/mo): tokenmax+Opus ~$700/mo, balanced+Sonnet ~$430/mo, conservative+Haiku ~$170/mo. Setups WITHOUT cache-aware prompt layout (frequent prefix churn) see the per-query matrix dominate — mode + model choice matters more there.

Resolution chain (matches the v0.31.12 model-tier pattern at src/core/model-config.ts:resolveModel):

per-call SearchOpts → per-key config (search.cache.enabled, …) →
  MODE_BUNDLES[search.mode] → MODE_BUNDLES.balanced (fallback)

Mode resolution lives in bare hybridSearch (NOT just the cached wrapper) per [CDX-5+6] in ~/.claude/plans/lets-take-a-look-validated-parrot.md — so gbrain eval replay and gbrain eval longmemeval test the same mode-affected behavior as the production query op.

Cache-key contamination hotfix [CDX-4]: migration v56 added a knobs_hash column to query_cache. The lookup filter is now WHERE source_id = $ AND knobs_hash = $ AND embedding similarity < $ so a tokenmax write (expansion=on, limit=50) can't be served to a conservative read.

v0.36.3.0 knobs_hash v=2 → v=3. The hash now folds the active embedding column name + provider into the cache key, so a query routed through embedding_voyage (1024d Voyage) can't be served a cache row written against embedding (1536d OpenAI). Existing v=2 rows become unreachable on first re-query (one-time miss spike on upgrade); mode.ts:KNOBS_HASH_VERSION is the single source of truth.

Three CLI surfaces:

gbrain search modes              # what is running, with per-knob attribution
gbrain search modes --reset      # clear search.* overrides (mode bundle wins)
gbrain search stats [--days N]   # cache hit rate, intent mix, budget drops
gbrain search tune [--apply]     # data-driven recommendations

The install picker fires inside gbrain init AFTER engine.initSchema() (non-TTY auto-selects). The upgrade banner fires once via runPostUpgrade in src/commands/upgrade.ts, gated by search.mode_upgrade_notice_shown.

Eval discipline (v0.32.3)

Every metric printed by any gbrain eval * or gbrain search stats command resolves through src/core/eval/metric-glossary.ts so industry terms (P@k, nDCG@k, MRR, Jaccard@k) carry a plain-English line in human output and a _meta.metric_glossary block in JSON output (one block per response per [CDX-25], NOT sibling _gloss fields).

The full methodology — datasets, sample selection, pre-registered expectations, threats to validity, paired-bootstrap + Bonferroni p-value discipline [CDX-14] — lives in docs/eval/SEARCH_MODE_METHODOLOGY.md. Auto-regenerated docs/eval/METRIC_GLOSSARY.md is CI-guarded against drift (scripts/check-eval-glossary-fresh.sh).

Per-run records land at <repo>/.gbrain-evals/eval-results.jsonl per [CDX-23]. The user's personal ~/.gbrain brain is NEVER touched — audit trail lives in the source repo's git history.

Skills

Read the skill files in skills/ before doing brain operations. GBrain ships 29 skills organized by skills/RESOLVER.md (AGENTS.md is also accepted as of v0.19):

Original 8 (conformance-migrated): ingest (thin router), query, maintain, enrich, briefing, migrate, setup, publish.

Brain skills (ported from an upstream agent fork): signal-detector, brain-ops, idea-ingest, media-ingest, meeting-ingestion, citation-fixer, repo-architecture, skill-creator, daily-task-manager.

Operational + identity: daily-task-prep, cross-modal-review, cron-scheduler, reports, testing, soul-audit, webhook-transforms, data-research, minion-orchestrator. As of v0.20.4, minion-orchestrator is the single unified skill for both lanes of background work (shell jobs via gbrain jobs submit shell, LLM subagents via gbrain agent run) ... the prior gbrain-jobs skill was merged in, Preconditions are shared, and trigger routing is narrowed to what the skill actually covers.

Skillify loop (v0.19): skillify (the markdown orchestration), skillpack-check (agent-readable health report).

Routing-table compression (v0.32.3.0): skills/functional-area-resolver/ — two-layer dispatch pattern for shrinking large AGENTS.md / RESOLVER.md files (>=12KB) without losing routing accuracy. Replaces one row per skill with one entry per functional area, where each area declares its sub-skills in a (dispatcher for: ...) clause. The static-prompt analog of hierarchical agent routing (AnyTool arXiv:2402.04253, RAG-MCP arXiv:2505.03275, Anthropic Agent Skills progressive disclosure). Empirically validated across Opus 4.7 / Sonnet 4.6 / Haiku 4.5: +13 to +17pp over the verbose baseline at 48% the size (25KB → 13KB on a real fork). The (dispatcher for: ...) clause is the load-bearing signal — strip it and lenient accuracy collapses to 41.7% on Sonnet (the resolver-of-resolvers ablation case). A/B eval surface lives at evals/functional-area-resolver/ (outside skills/ deliberately so the skillpack bundler doesn't ship eval infrastructure to downstream installs): gateway-routed TypeScript harness, 20 training + 5 held-out fixtures, strict + lenient scoring, three committed cross-model receipts in baseline-runs/. Receipt header binds (model, prompt_template_hash, fixtures_hash, harness_sha, ts) so future contributors can verify reproduction. Companion rescore.mjs re-scores existing JSONL with lenient tolerance for zero API cost. Reproduce with cd evals/functional-area-resolver && node harness.mjs --model {opus|sonnet|haiku} (~$0.301.70 per model). Nine v0.33.x follow-up TODOs filed for held-out corpus growth, cross-vendor verification, hierarchical area-of-areas, embedding-based pre-router, and the run-1 vs run-2 prompt-design ablation methodology.

Operational health (v0.19.1): smoke-test (8 post-restart health checks with auto-fix for Bun, CLI, DB, worker, Zod CJS, gateway, API key, brain repo; user-extensible via ~/.gbrain/smoke-tests.d/*.sh).

Conventions: skills/conventions/ has cross-cutting rules (quality, brain-first, model-routing, test-before-bulk, cross-modal). skills/_brain-filing-rules.md and skills/_output-rules.md are shared references.

Bulk-action progress reporting

All bulk commands (doctor, embed, import, export, sync, extract, migrate, repair-jsonb, orphans, check-backlinks, lint, integrity auto, eval, files sync, and apply-migrations) stream progress through the shared reporter at src/core/progress.ts. Agents get heartbeats within 1 second of every iteration regardless of how slow the underlying work is.

Rules:

  • Progress always writes to stderr. Stdout stays clean for data output (--json payloads, final summaries, JSON action events from extract).
  • Non-TTY default: plain one-line-per-event human text. JSON requires the explicit --progress-json flag.
  • Global flags (--quiet, --progress-json, --progress-interval=<ms>) are parsed by src/core/cli-options.ts BEFORE command dispatch.
  • Phase names are machine-stable snake_case.dot.path (e.g. doctor.db_checks, sync.imports). Documented in docs/progress-events.md; additive changes only.
  • scripts/check-progress-to-stdout.sh is a CI guard that fails the build if any new code writes \r progress to stdout. Wired into bun run test.
  • Minion handlers pass job.updateProgress as the onProgress callback to core functions (DB-backed primary progress channel); stderr from jobs work stays coarse for daemon liveness only.

When wiring a new bulk command: import { createProgress } from '../core/progress.ts' and import { getCliOptions, cliOptsToProgressOptions } from '../core/cli-options.ts'. Create a reporter with createProgress(cliOptsToProgressOptions(getCliOptions())), start(phase, total?) before the loop, tick() inside it, finish() after. For single long-running queries, use startHeartbeat(reporter, note) with a try/finally to guarantee cleanup. Never call process.stdout.write('\r...') in bulk paths, the CI guard will fail the build.

Capturing test output (NEVER pipe through tail / head)

Iron rule: when running bun test, bun run test:e2e, bun run typecheck, or any other test/check command, redirect to a file FIRST, then tail the file separately:

# RIGHT — full output preserved, real exit code visible
bun test > /tmp/ship_units.txt 2>&1
echo "EXIT=$?"
tail -50 /tmp/ship_units.txt
grep -E '(fail\)|✗|error:' /tmp/ship_units.txt | head -30
# WRONG — exit code is `tail`'s (always 0), failures truncated, ship gates fail open
bun test 2>&1 | tail -10

The pipe form silently breaks /ship Step T1 (test failure ownership triage) and the test verification gate (Step 16) because:

  • $? after a pipe is the LAST command's exit code (tail → 0), not bun's
  • bun prints failure details before the summary line, so tail -N drops them
  • Step T1 needs the full failure list to classify in-branch vs pre-existing

This bit us during v0.26.2 ship: bun test 2>&1 | tail -10 reported "3911 pass / 23 fail" but no failure details survived, forcing a 23-minute re-run to triage.

Apply the same pattern to any long-running command whose exit code matters: bun run typecheck, bun run ci:local, migration runs, eval suites, etc. For background tasks (run_in_background: true), the harness captures the exit file separately — use it via the bg task's <id>.exit file, not the streamed output.

Build

bun build --compile --outfile bin/gbrain src/cli.ts

Version locations (single source of truth: VERSION file)

Every release advances the version in five files at once. Keep these in sync. /ship enforces this via Step 12's idempotency check (VERSION vs package.json drift), but the canonical list lives here so future runs and the auto-update agent know where to look.

Version format is mandatory: MAJOR.MINOR.PATCH.MICRO (four numeric segments, dot-separated, no leading v). Every new release MUST use the 4-segment form. The .MICRO slot is the dot-suffix follow-up channel: when a release ships its commit subject ahead of its VERSION bump (e.g. PR #795 landing as v0.31.4 without bumping the file), the corrective ship lands as 0.31.4.1 rather than churning the patch number to 0.31.5. Suffixes like -fixwave are still allowed as needed (0.31.1.1-fixwave), but the four numeric segments are required first. Historical 3-segment versions (0.31.3, 0.22.1) remain valid in git log and migration filenames (skills/migrations/v0.21.0.md); do NOT rewrite them. Going forward only.

Required (every release must update all five):

File What lives there Format
VERSION The single source of truth. Read first by /ship, the binary, and CI version-gate. Bare 4-segment string MAJOR.MINOR.PATCH.MICRO (e.g. 0.31.4.1), no leading v.
package.json Bun/npm package version. gbrain --version reads it via the compiled binary's bundled package metadata. CI version-gate cross-checks this against VERSION and fails if they drift. "version": "0.31.4.1"
CHANGELOG.md Top entry header ## [0.31.4.1] - YYYY-MM-DD plus the "To take advantage of v0.31.4.1" block. Standard Keep-a-Changelog header.
TODOS.md Any TODO entries that mention "follow-up from vX.Y.Z.W" use the version of the release that filed them. Update only when filing NEW follow-up TODOs. Inline vX.Y.Z.W references in TODO bodies.
CLAUDE.md The Key Files section's per-file annotations carry vX.Y.Z.W (#NNN) tags noting which release introduced a behavior. Update whenever a wave's annotations get folded in. Inline vX.Y.Z.W (#NNN, contributed by @user) references.

Auto-derived (no manual edit; refreshed by their own commands):

  • bun.lock — root-package version is auto-pinned from package.json. After bumping package.json, run bun install to refresh the lockfile.
  • llms-full.txt / llms.txt — auto-generated documentation bundles. Any CLAUDE.md edit MUST be followed by bun run build:llms in the same commit (or a follow-up commit before push). The committed bundles are checked against fresh generator output by test/build-llms.test.ts, which runs in CI shard 1. If you edited CLAUDE.md and didn't regenerate, CI will fail. This has bitten the wave 3 times — every CLAUDE.md edit gets a bun run build:llms chaser, no exceptions. (The verify gate doesn't run this test; only the full unit suite does. So bun run typecheck clean is NOT enough to know you can push after a CLAUDE.md edit.)

Historical (DO NOT bump on release):

  • skills/migrations/v0.21.0.md — migration files use the version they shipped FROM as their filename. v0.21.0's migration always says v0.21.0.
  • src/commands/migrations/v0_21_0.ts — same: migration code references the schema version it migrates to.
  • test/migrations-v0_21_0.test.ts, test/migration-orchestrator-v0_21_0.test.ts, test/migrate.test.ts — migration tests reference historical migration versions; these are correct as-is and should not move.
  • src/core/db.ts, src/core/migrate.ts, src/core/import-file.ts, src/commands/reindex-code.ts — code comments cite the release that introduced a feature. Once written, these are historical record.
  • README.md — references the latest published feature names by version (e.g. "v0.21.0 Code Cathedral"); update only when the README's marketing copy is intentionally being refreshed, NOT on every micro/patch bump.

The /ship workflow's version idempotency check: Step 12 reads VERSION and package.json, classifies as FRESH / ALREADY_BUMPED / DRIFT_STALE_PKG / DRIFT_UNEXPECTED, and refuses to proceed on DRIFT_UNEXPECTED. This is why the two must move together.

The CI version-gate rejects pushes where VERSION and package.json disagree, OR where VERSION is not strictly greater than master's VERSION. If a queue collision claims your version on master before yours lands, /ship's queue-aware allocator (Step 12) will detect drift and re-bump on the next run.

Mandatory version-consistency audit (run after EVERY merge or commit that touches VERSION, package.json, or CHANGELOG)

The trio MUST agree. Every merge from master will hit conflicts on VERSION + package.json + CHANGELOG.md because master ships its own version bumps. Auto-merge sometimes resolves these silently in unexpected ways. After any merge, branch update, or version-related edit, run this audit. It's three lines and never lies:

echo "VERSION:     $(cat VERSION)"
echo "package.json: $(node -e 'process.stdout.write(require("./package.json").version)')"
grep -E "^## \[" CHANGELOG.md | head -1

All three MUST show the same MAJOR.MINOR.PATCH.MICRO. If any one disagrees, you have not finished the merge. Fix it before pushing or shipping. There is no situation in which "I'll fix it next push" is OK, because:

  • A green local test run with mismatched VERSION/package.json still fails the CI version-gate.
  • A green CHANGELOG entry under the wrong version header silently lies to release-notes consumers.
  • /ship's Step 12 idempotency check classifies a mismatch as DRIFT_UNEXPECTED and HALTS — but only if you remember to run /ship before pushing. Manual git push skips the check.

Merge-conflict recovery procedure (memorize this)

When git merge origin/master reports conflicts on VERSION, package.json, or CHANGELOG.md, resolve in this exact order:

  1. VERSION — overwrite with the wave's version (`echo -n "X.Y.Z.W"

    VERSION`). Highest semver wins; do NOT take master's lower version.

  2. package.json — strip the conflict markers, keep the wave's version line. Sed pattern: sed -i.bak '/^<<<<<<< HEAD$/d; /^=======$/,/^>>>>>>> /d' package.json && rm package.json.bak (assumes ours is above the =======).
  3. CHANGELOG.md — strip ALL three conflict markers; both your entry and master's entry stay. Sed pattern: sed -i.bak '/^<<<<<<< HEAD$/d; /^=======$/d; /^>>>>>>> origin\/master$/d' CHANGELOG.md && rm CHANGELOG.md.bak Then verify your entry is the topmost ## [X.Y.Z.W] and master's newer-than-yours entries (if any) sit below.
  4. Run the 3-line audit above. If it doesn't show your version on all three lines, you missed a marker.
  5. Run bun install to refresh bun.lock against the resolved package.json. Stage and commit if it changed.
  6. Run bun run typecheck before committing the merge.
  7. Only THEN run git commit for the merge.

If the audit shows drift after step 4, do NOT proceed to step 5. Re-run steps 1-3 against the actual file content; you missed a marker or resolved one in the wrong direction.

Anti-pattern to avoid: Resolving via git checkout --ours package.json and git checkout --theirs scripts/test-shard.sh mixed in the same commit. The selective directional resolution is fine, but on VERSION/package.json/CHANGELOG specifically, ALWAYS use the explicit echo > VERSION + sed-strip-markers pattern above. The directional checkout flags have bitten us when the conflict shape was unexpected (e.g. master stripped a section we expected to keep).

Pre-push gate (manual; tighten when you remember to)

Before any git push of a merge commit, run the audit one more time:

echo "VERSION:     $(cat VERSION)"
echo "package.json: $(node -e 'process.stdout.write(require("./package.json").version)')"
grep -E "^## \[" CHANGELOG.md | head -1

If you've been editing the branch via /ship you can rely on Step 12's idempotency check. If you've been editing manually (merge resolution, conflict fix, version bump), the audit is the last line of defense before CI yells at you.

Conductor branch-name = workspace-name (IRON RULE)

Conductor workspaces expect the git branch name to match the workspace directory name. When they disagree, Conductor silently fails to render the PR view + show ship state, leading to "did you actually push?" confusion.

Check this FIRST on every ship and BEFORE creating any PR:

WORKSPACE=$(basename "$PWD")              # e.g. puebla-v4
BRANCH=$(git branch --show-current)        # e.g. garrytan/gstack-requests
case "$BRANCH" in
  */"$WORKSPACE") echo "OK: branch tail matches workspace" ;;
  "$WORKSPACE")   echo "OK: branch == workspace" ;;
  *)              echo "MISMATCH: branch=$BRANCH workspace=$WORKSPACE — RENAME BEFORE SHIPPING" ;;
esac

If MISMATCH (branch is garrytan/foo but workspace is puebla-v4):

# Rename local, push under new name, delete old remote (and old PR if it
# was already created — github auto-closes it when head ref dies).
git branch -m garrytan/<workspace-name>
git push -u origin garrytan/<workspace-name>
git push origin --delete <old-branch-name>
# If a PR existed against the old branch:
#   gh pr comment <old-pr> --body "Superseded by #<new>: branch renamed to match Conductor workspace."
#   gh pr create --base master --title "..." --body "..."  # recreate from renamed branch

Caught the hard way on v0.41.9.0 ship: workspace puebla-v4 but branch garrytan/gstack-requests produced PR #1439 that Conductor wouldn't display. Renamed to garrytan/puebla-v4; recreated as #1440.

The /ship workflow's Step 1 should be augmented to run the mismatch check; until that lands upstream, ALWAYS run the check above before /ship invokes its first push or PR-create step.

Pre-ship requirements

Before shipping (/ship) or reviewing (/review), always run the full test suite. Two equivalent paths:

Path A — local CI gate (recommended, v0.23.1+):

  • bun run ci:local runs the entire stack inside Docker: gitleaks (host), unit tests with DATABASE_URL unset, and all 29 E2E files sequentially against a fresh pgvector container. Stronger than PR CI's 2-file Tier 1 set; closer to what nightly Tier 1 catches. Spins up + tears down postgres automatically via docker-compose.ci.yml. Override the host port with GBRAIN_CI_PG_PORT=5435 bun run ci:local if 5434 collides.
  • bun run ci:local:diff runs only the E2E files matched by the diff selector (scripts/select-e2e.ts), falling back to all 29 on unmapped src/ paths or schema/skills/package.json changes. Fast iteration during a focused branch.

Path B — manual lifecycle (still supported):

  • bun test — unit tests (no database required)
  • Follow the "E2E test DB lifecycle" steps above to spin up the test DB, run bun run test:e2e, then tear it down.

Both must pass. Do not ship with failing E2E tests. Do not skip E2E tests.

Always run typecheck before pushing. bun test (the bun runner) skips TypeScript type checking — it only enforces runtime behavior. Three ways to actually gate on types:

  1. bun run test (npm script in package.json) — includes bun run typecheck plus the four shell pre-checks (check-jsonb-pattern.sh, check-progress-to-stdout.sh, check-trailing-newline.sh, check-wasm-embedded.sh) before the runner. Use this mid-branch.
  2. bun run typechecktsc --noEmit standalone. Fast (~5s on this repo).
  3. bun run ci:local — the full local CI gate from Path A.

The trap is: writing a new test, running bun test test/foo.test.ts, seeing it pass, pushing — and CI's separate typecheck stage rejects an invalid type literal that the runner accepted. Caught one of these shipping the v0.23.2 round-trip E2E (type: 'reflection' is not a member of PageType). Run bun run typecheck once before push, even when only test files changed.

Post-ship requirements (MANDATORY)

After EVERY /ship, you MUST run /document-release. This is NOT optional. Do NOT skip it. Do NOT say "docs look fine" without running it. The skill reads every .md file in the project, cross-references the diff, and updates anything that drifted.

If /ship's Step 8.5 triggers document-release automatically, that counts. But if it gets skipped for ANY reason (timeout, error, oversight), you MUST run it manually before considering the ship complete.

Files that MUST be checked on every ship:

  • README.md — does it reflect new features, commands, or setup steps?
  • CLAUDE.md — does it reflect new files, test files, or architecture changes?
  • CHANGELOG.md — does it cover every commit?
  • TODOS.md — are completed items marked done?
  • docs/ — do any guides need updating?

A ship without updated docs is an incomplete ship. Period.

CHANGELOG + VERSION are branch-scoped

VERSION and CHANGELOG describe what THIS branch adds vs master, not how we got here. Every feature branch that ships gets its own version bump and CHANGELOG entry. The entry is product release notes for users; it is not a log of internal decisions, review rounds, or codex findings.

Write the CHANGELOG entry at /ship time, not during development. Mid-branch iterations, review rounds (CEO/Eng/Codex/DX), and implementation detours belong in the plan file at ~/.claude/plans/, not in the CHANGELOG. One unified entry per branch, covering what the branch added vs the base branch.

Never edit a CHANGELOG entry that already landed on master. If master has v0.18.2 and your branch adds features, bump to the next version (v0.19.0, not editing master's v0.18.2). When merging master into your branch, master may bring new CHANGELOG entries above yours — push your entry above master's latest and verify:

  • Does CHANGELOG have your branch's own entry separate from master's entries?
  • Is VERSION higher than master's VERSION?
  • Is your entry the topmost ## [X.Y.Z] entry?
  • grep "^## \[" CHANGELOG.md shows a contiguous version sequence?

If any answer is no, fix it before continuing.

CHANGELOG is for users, not contributors. Write like product release notes:

  • Lead with what the user can now do that they couldn't before. Sell the capability.
  • Plain language, not implementation details. "You can now..." not "Refactored the..."
  • Never mention internal artifacts: plan file IDs, decision tags (D-CX-#, F-ENG-#), review rounds, codex findings, subcontractor credits. These are invisible to users.
  • Put contributor-facing changes in a separate ### For contributors section at the bottom.
  • Every entry should make someone think "oh nice, I want to try that."

What to omit:

  • "Codex caught X that the CEO review missed" — private process detail.
  • "D-CX-3 split errors/warnings" — tag is meaningless to users; name the feature instead.
  • "Fix-wave PR #N supersedes #M" — supersede chains belong in PR bodies, not release notes.
  • "215 new cases, 3 decisions applied, 7 reviews cleared" — these are planning-mode metrics.

What to keep:

  • The user-facing change: what commands exist now, what flag was added, what behavior fixed.
  • Numbers that mean something to the user: TTHW, commands that timed out before, detection counts.
  • Upgrade instructions: gbrain upgrade + any manual step if needed.
  • Credit to external contributors when a community PR was incorporated.

CHANGELOG voice + release-summary format

IRON RULE: the CHANGELOG describes what the user gets, not how the work happened. Nobody reading release notes cares that codex caught a bug, that the plan went through CEO + eng review, that the migration was originally numbered v68 and renumbered to v79 during master merge, or that two review rounds caught architectural mistakes. The reader cares what gbrain brainstorm does and how to use it. If a fact only exists because of the development process, it does NOT belong in the CHANGELOG.

Specifically forbidden in CHANGELOG entries:

  • Any mention of review processes (CEO review, eng review, codex review, plan-eng-review, outside voice, adversarial review, autoplan, /review).
  • "What we caught and fixed before merging" sections. Bugs found pre-merge are not changes — they're things that didn't ship.
  • Plan file references, plan IDs, plan decision tags (D1, D14, D-CDX-3).
  • Migration version drama ("originally v68", "renumbered to v77", "claimed by parallel waves") — just say "Migration v79 adds X." If the user cares about migration ordering, they read the diff.
  • Round counts, finding counts, decision counts ("25 findings across 2 rounds", "8 architectural decisions", "5/6 expansions accepted").
  • Names of internal collaborators ("codex caught", "the reviewer flagged", "Claude noticed").
  • "Plan + reviews" summary bullets. The plan lives in ~/.claude/plans/; if a future reader wants the backstory they can grep there.
  • Any wording that frames a shipped feature as a recovery from a planning mistake ("the first plan was wrong", "we corrected the approach", "the shipped version supersedes the original design").

Smell test: read the entry as a stranger who has never touched gbrain. If any sentence makes them think "why are you telling me this?", cut it. Every sentence in the release-summary AND in the itemized changes must answer one of three questions: What can I now do? How do I use it? What should I watch for after I upgrade?

Every version entry in CHANGELOG.md MUST start with a release-summary section in the GStack/Garry voice — one viewport's worth of prose + tables that lands like a verdict, not marketing. The itemized changelog (subsections, bullets, files) goes BELOW that summary, separated by a ### Itemized changes header.

The release-summary section gets read by humans, by the auto-update agent, and by anyone deciding whether to upgrade. The itemized list is for agents that need to know exactly what changed.

Release-summary template

Iron rule: lead ELI10, get precise after. The first ~150 words of every entry must be readable by someone who does NOT know gbrain's internals. No file paths, no function names, no internal constants, no acronyms (no "RRF", no "knobsHash", no "MODE_BUNDLES", no "CDX-4"), no jargon that requires reading the codebase to parse. Lead with the user-visible behavior change, in everyday English, like you're explaining it to a smart engineer who has never opened the repo.

THEN, once the reader knows what shipped and why they'd care, drill into the precise details: real file paths, real function names, real config keys, real numbers. The precision part is required (the entry is also the technical record of what changed), but it lives AFTER the plain-English lead, never before it.

The shape:

  1. One-line bold headline. What changed for the user, in human English. No jargon. No internal terms. Example good: "Your search stops boosting weak pages just because they have a lot of links pointing at them." Example bad: "PostFusionOpts gains floorRatio; KNOBS_HASH_VERSION bumped 2→3."
  2. Plain-English opener (~3-5 sentences). Describe the problem this fixes in everyday terms. Pretend the reader has a brain full of meeting notes and people pages and wants to know if this release helps them. Concrete example beats abstract description.
  3. A "How to turn it on" or "How to use it" section with paste-ready commands. Real flags, real config keys. This is where precision starts.
  4. A "What you'd see in a concrete example" or "The X numbers that matter" section with a table. Use everyday-language column headers ("Page", "Match quality", "Has many backlinks?") even when the underlying mechanism is technical. The table teaches what the feature does without requiring the reader to understand how.
  5. A "What's safe to know about" or "Things to watch" section for caveats, side effects, cache invalidation, mid-deploy notes. Still in plain language.
  6. A "What we caught and fixed before merging" section if the work went through review (CEO/eng/codex/outside-voice). Translate review findings into plain English. "We caught a stale-cache bug" beats "knobsHash() did not include floorRatio in the v=2 hash input."
  7. ### Itemized changes (precision lives here). File paths, function names, types, constants, line numbers. This section is for engineers who need to know exactly what moved.

Voice rules (apply throughout):

  • No em dashes (use commas, periods, "...").
  • No AI vocabulary (delve, robust, comprehensive, nuanced, fundamental, etc.) or banned phrases ("here's the kicker", "the bottom line", etc.).
  • Real numbers, real file names, real commands AFTER the ELI10 lead. Not "fast" but "~30s on 30K pages." In the ELI10 lead, "fast enough that you won't notice" or "~30 seconds even on a big brain."
  • Short paragraphs, mix one-sentence punches with 2-3 sentence runs.
  • Connect to user outcomes: "the agent does ~3x less reading" beats "improved precision."
  • Be direct about quality. "Well-designed" or "this is a mess." No dancing.

The smell test: if someone who has never opened gbrain reads the first 150 words and walks away knowing what shipped and whether they care, the entry passes. If they need to grep the codebase to follow along, rewrite the lead.

Canonical examples in this CHANGELOG: v0.35.6.0 (floor-ratio gate, written ELI10-lead-first), v0.34.4.0 (embed stale fix wave). Use those shapes when in doubt. Avoid the shape of entries that lead with internal constants or release mechanics; those exist in older history but should not be the model for new work.

Source material to pull from:

  • CHANGELOG.md previous entry for prior context
  • Latest gbrain-evals/docs/benchmarks/[latest].md for headline numbers (sibling repo)
  • Recent commits (git log <prev-version>..HEAD --oneline) for what shipped
  • Don't make up numbers. If a metric isn't in a benchmark or production data, don't include it. Say "no measurement yet" if asked.

Target length: ~250-350 words for the summary. Should render as one viewport.

"To take advantage of v[version]" block (required, v0.13+)

After the release-summary and BEFORE ### Itemized changes, every ## [X.Y.Z] entry MUST include a human-readable self-repair block under the heading ## To take advantage of v[version].

Why: gbrain upgrade runs gbrain post-upgrade which runs gbrain apply-migrations. This chain has a known weak link — upgrade.ts catches post-upgrade failures as best-effort (so the binary still works). When that chain silently fails, users end up with half-upgraded brains. The self-repair block gives them a paste-ready recovery path; the v0.13+ ~/.gbrain/upgrade-errors.jsonl trail + gbrain doctor integration close the loop.

Template (adapt the verify commands per release):

## To take advantage of v[version]

`gbrain upgrade` should do this automatically. If it didn't, or if `gbrain doctor`
warns about a partial migration:

1. **Run the orchestrator manually:**
   ```bash
   gbrain apply-migrations --yes
  1. Your agent reads skills/migrations/v[version].md the next time you interact with it. [One sentence on whether headless agents need manual action, or whether the orchestrator already handled the mechanical side.]

  2. Verify the outcome:

    [release-specific verify commands, e.g. `gbrain graph ... --depth 2`]
    gbrain stats
    
  3. If any step fails or the numbers look wrong, please file an issue: https://github.com/garrytan/gbrain/issues with:

    • output of gbrain doctor
    • contents of ~/.gbrain/upgrade-errors.jsonl if it exists
    • which step broke

    This feedback loop is how the gbrain maintainers find fragile upgrade paths. Thank you.


**Skip this block** for patches that are pure bug fixes with zero user-facing action
(rare). If the release has a schema migration, data backfill, or new feature the
user needs to verify, the block is required.

The v0.13.0 entry in CHANGELOG.md is the canonical example.

### Itemized changes (the existing rules)

Below the release summary, write `### Itemized changes` and continue with the
detailed subsections (Knowledge Graph Layer, Schema migrations, Security hardening,
Tests, etc.). Same rules as before:

- Lead with what the user can now DO that they couldn't before
- Frame as benefits and capabilities, not files changed or code written
- Make the user think "hell yeah, I want that"
- Bad: "Added GBRAIN_VERIFY.md installation verification runbook"
- Good: "Your agent now verifies the entire GBrain installation end-to-end, catching
  silent sync failures and stale embeddings before they bite you"
- Bad: "Setup skill Phase H and Phase I added"
- Good: "New installs automatically set up live sync so your brain never falls behind"
- **Always credit community contributions.** When a CHANGELOG entry includes work from
  a community PR, name the contributor with `Contributed by @username`. Contributors
  did real work. Thank them publicly every time, no exceptions.

### Reference: v0.12.0 entry as canonical example

The v0.12.0 entry in CHANGELOG.md is the canonical example of the format. Match its
structure for every future version: bold headline, lead paragraph, "numbers that
matter" with BrainBench-style before/after table, "what this means" closer, then
`### Itemized changes` with the detailed sections below.

## Version migrations

Create a migration file at `skills/migrations/v[version].md` when a release
includes changes that existing users need to act on. The auto-update agent
reads these files post-upgrade (Section 17, Step 4) and executes them.

**You need a migration file when:**
- New setup step that existing installs don't have (e.g., v0.5.0 added live sync,
  existing users need to set it up, not just new installs)
- New SKILLPACK section with a MUST ADD setup requirement
- Schema changes that require `gbrain init` or manual SQL
- Changed defaults that affect existing behavior
- Deprecated commands or flags that need replacement
- New verification steps that should run on existing installs
- New cron jobs or background processes that should be registered

**You do NOT need a migration file when:**
- Bug fixes with no behavior changes
- Documentation-only improvements (the agent re-reads docs automatically)
- New optional features that don't affect existing setups
- Performance improvements that are transparent

**The key test:** if an existing user upgrades and does nothing else, will their
brain work worse than before? If yes, migration file. If no, skip it.

Write migration files as agent instructions, not technical notes. Tell the agent
what to do, step by step, with exact commands. See `skills/migrations/v0.5.0.md`
for the pattern.

## Migration is canonical, not advisory

GBrain's job is to deliver a canonical, working setup to every user on upgrade.
Anything that looks like a "host-repo change" — AGENTS.md, cron manifests,
launchctl units, config files outside `~/.gbrain/` — is a GBrain migration
step, not a nudge we leave for the host-repo maintainer. Migrations edit host
files (with backups) to make the canonical setup real. Exceptions: changes
that require human judgment (content edits, renames that break semantics,
host-specific handler registration where shell-exec would be an RCE surface).
Everything mechanical ships in the migration.

**Test:** if shipping a feature requires a sentence that starts with "in
your AGENTS.md, add…" or "in your cron/jobs.json, rewrite…", the migration
orchestrator should be doing that edit, not the user.

**The exception is host-specific code.** For custom Minion handlers
(host-specific integrations like inbox sweeps or third-party API scanners), shipping them as a
data file the worker would exec is an RCE surface. Those get registered in
the host's own repo via the plugin contract (`docs/guides/plugin-handlers.md`);
the migration orchestrator emits a structured TODO to
`~/.gbrain/migrations/pending-host-work.jsonl` + the host agent walks the
TODOs using `skills/migrations/v0.11.0.md` — stays host-agnostic, still
canonical.

## Privacy rule: scrub real names from public docs

**Never reference real people, companies, funds, or private agent names in any
public-facing artifact.** Public artifacts include: `CHANGELOG.md`, `README.md`,
`docs/`, `skills/`, PR titles + bodies, commit messages, and comments in checked-in
code. Query examples, benchmark stories, and migration guides MUST use generic
placeholders.

Why: gbrain runs a personal knowledge brain containing notes on real people and
real companies (YC founders, portfolio companies, funds, investors, meeting
attendees). When a doc copies a query like `gbrain graph diana-hu --depth 2` or
names a specific agent fork like `Wintermute`, that real name gets indexed by
search engines, surfaced in cross-references, and distributed with every release.

**Name mapping** to use in examples:
- Agent forks → `your agent fork`, `a downstream agent`, or `agent-fork`
- Example person → `alice-example`, `charlie-example`, or `a-founder`
- Example company → `acme-example`, `widget-co`, or `a-company`
- Example fund → `fund-a`, `fund-b`, `fund-c`
- Example deal → `acme-seed`, `widget-series-a`
- Example meeting → `meetings/2026-04-03` (generic date is fine)
- Example user → `you` or `the user`, never a proper name

**Specific rule: never say `Wintermute` in any CHANGELOG, README, doc, PR, or
commit message.** When the temptation is to illustrate with the real fork name:
- Reader-facing copy → `your OpenClaw` (covers Wintermute, Hermes, AlphaClaw,
  and any other downstream OpenClaw deployment in one term the reader already
  recognizes).
- First-person / origin-story copy → `Garry's OpenClaw` (honest that this is
  the production deployment driving the feature, without exposing the private
  agent's name).

`Wintermute` may appear in private artifacts (scratch plans under
`~/.gstack/projects/…`, memory files, conversation transcripts, CEO-review
plans) — those aren't distributed. Anything checked into this repo or shipped
in a release must use the OpenClaw phrasing above. Sweeping a stale reference
is a small clean-up PR, not a debate.

**When in doubt, ask yourself:** "Would this query reveal private information
about the user's contacts, investments, or portfolio if it were read by a
stranger?" If yes, replace with generic placeholders.

**Illustrative API examples with household-brand companies** (Stripe, Brex, OpenAI,
GitHub, etc.) are fine — they're public entities, not contacts in anyone's brain.
Do not confuse illustrative API examples with queries that reveal real
relationships.

## Responsible-disclosure rule: don't broadcast attack surface in release notes

**When a release fixes a security gap or a user-impacting bug, describe the fix
functionally. Do not enumerate the attack surface, quantify the exposure window,
or highlight the most sensitive records by name in public-facing artifacts.**

Public-facing artifacts include: `CHANGELOG.md`, `README.md`, `docs/`, PR titles
and bodies, commit messages, GitHub issue titles and comments, release pages,
tweets, blog posts.

**Don't write:**
- "10 tables were publicly readable by the anon key for months, including X, Y, Z"
- "X and Y are the most sensitive ones"
- "N tables exposed. Fix: enable RLS on these specific tables: ..."

**Do write:**
- "Security hardening pass. Fresh installs secure by default. Existing brains
  brought to the same bar automatically on upgrade."
- "If `gbrain doctor` still flags anything after upgrade, the message names each
  table and gives the exact fix."

Why: anyone reading the release page before they've upgraded now has a directed
probe list for unpatched installs. The source code ships the specifics anyway
(`src/schema.sql`, `src/core/migrate.ts`, test fixtures) — reverse engineers can
get them. But the release page is a broadcast channel. Don't hand attackers a
curated list with a banner.

**The test:** if a reader with no prior context could read the release note and
walk away knowing "gbrain at version X has table Y readable by anon key until
they patch," the note is too specific. Rewrite until that's no longer possible.

**What IS fine in public artifacts:**
- The mechanism of the fix ("the check now scans every public table instead of
  a hardcoded allowlist").
- User-facing operator ergonomics (the escape-hatch SQL template, the upgrade
  commands, the breaking-change flag).
- Credit to contributors.
- Generic framing of severity ("security posture tightening pass") without
  quantification.

**What stays in private artifacts (plan files, private memories, internal docs):**
- Specific table names, record counts, exposure duration.
- Which records stand out as highest-risk.
- Detailed before/after tables in the "numbers that matter" format.

If the CEO/Eng review of a plan produces a detailed exposure table, keep it in
the plan file under `~/.claude/plans/` or `~/.gstack/projects/`. Don't copy it
into the CHANGELOG or PR body.

Applies retroactively: if you see a prior CHANGELOG entry naming attack-surface
specifics, scrub it as a small cleanup commit, the same way a stale Wintermute
reference gets swept.

## Schema state tracking

`~/.gbrain/update-state.json` tracks which recommended schema directories the user
adopted, declined, or added custom. The auto-update agent (SKILLPACK Section 17)
reads this during upgrades to suggest new schema additions without re-suggesting
things the user already declined. The setup skill writes the initial state during
Phase C/E. Never modify a user's custom directories or re-suggest declined ones.

## GitHub Actions SHA maintenance

All GitHub Actions in `.github/workflows/` are pinned to commit SHAs. Before shipping
(`/ship`) or reviewing (`/review`), check for stale pins and update them:

```bash
for action in actions/checkout oven-sh/setup-bun actions/upload-artifact actions/download-artifact softprops/action-gh-release gitleaks/gitleaks-action; do
  tag=$(grep -r "$action@" .github/workflows/ | head -1 | grep -o '#.*' | tr -d '# ')
  [ -n "$tag" ] && echo "$action@$tag: $(gh api repos/$action/git/ref/tags/$tag --jq .object.sha 2>/dev/null)"
done

If any SHA differs from what's in the workflow files, update the pin and version comment.

PR title format — version FIRST (IRON RULE)

Every PR title MUST start with the version, then the conventional-commit subject:

vMAJOR.MINOR.PATCH.MICRO <type>(<scope>): <summary> (#issue or wave ref)

Example (correct): v0.42.3.0 feat(search): autocut — score-discontinuity result-sizing (#1663 wave 1)

The version goes at the BEGINNING, never the end. This matches the repo's commit-subject convention (git log shows v0.41.38.0 fix: ..., v0.42.1.0 feat: ...) so the PR list, the merge commit, and the changelog all read version-first. A title with the version parenthesized at the end (feat(search): autocut ... (v0.42.3.0)) is WRONG — fix it with gh pr edit <N> --title "vX.Y.Z.W <type>: <summary>".

This applies to gh pr create and every gh pr edit --title. When /ship (or any flow) sets a PR title, the version is the first token. Same rule for the final commit subject that carries the version bump.

PR descriptions cover the whole branch

Pull request titles and bodies must describe everything in the PR diff against the base branch, not just the most recent commit you made. When you open or update a PR, walk the full commit range with git log --oneline <base>..<head> and write the body to cover all of it. Group by feature area (schema, code, tests, docs) — not chronologically by commit.

This matters because reviewers read the PR body to understand what's shipping. If the body only covers your last commit, they miss everything else and can't review properly. A 7-commit PR with a body that describes commit 7 is worse than no body at all — it actively misleads.

When in doubt, run gh pr view <N> --json commits --jq '[.commits[].messageHeadline]' to see what's actually in the PR before writing the body.

Community PR wave process

Never merge external PRs directly into master. Instead, use the "fix wave" workflow:

  1. Categorize — group PRs by theme (bug fixes, features, infra, docs)
  2. Deduplicate — if two PRs fix the same thing, pick the one that changes fewer lines. Close the other with a note pointing to the winner.
  3. Collector branch — create a feature branch (e.g. garrytan/fix-wave-N), cherry-pick or manually re-implement the best fixes from each PR. Do NOT merge PR branches directly — read the diff, understand the fix, and write it yourself if needed.
  4. Test the wave — verify with bun test && bun run test:e2e (full E2E lifecycle). Every fix in the wave must have test coverage.
  5. Close with context — every closed PR gets a comment explaining why and what (if anything) supersedes it. Contributors did real work; respect that with clear communication and thank them.
  6. Ship as one PR — single PR to master with all attributions preserved via Co-Authored-By: trailers. Include a summary of what merged and what closed.

Community PR guardrails:

  • Always AskUserQuestion before accepting commits that touch voice, tone, or promotional material (README intro, CHANGELOG voice, skill templates).
  • Never auto-merge PRs that remove YC references or "neutralize" the founder perspective.
  • Preserve contributor attribution in commit messages.

Checking out PRs from garrytan-agents

garrytan-agents is the AI-authored PR account and is NOT a collaborator on this repo. Its PRs live in a fork, so GitHub Actions triggered by pull_request events on those PRs do not receive base-repo secrets. Any CI job that needs ANTHROPIC_API_KEY, OPENAI_API_KEY, or similar will fail with empty-env auth errors, regardless of what's set on the base repo. This is a GitHub security default, not a config bug.

When the user says "check out " and the PR is from garrytan-agents (or any other non-collaborator fork), move the branch into the base repo before running CI:

  1. gh pr checkout <N> — pull down the fork's branch. Note the PR number and head branch name (gh pr view <N> --json headRefName --jq .headRefName).
  2. git push origin HEAD:<branch-name> — push the same branch to the base repo (origin points at garrytan/gbrain, not the fork). This is the move that gives CI access to secrets.
  3. gh pr close <N> --comment "moving to base-repo branch for secret access" — close the fork PR so the queue stays clean.
  4. gh pr create --base master --head <branch-name> — open the replacement PR from the base-repo branch. Preserve the original PR's title and body verbatim (gh pr view <N> --json title,body); contributor attribution moves to a Co-Authored-By: trailer if needed.

Why this over alternatives: adding garrytan-agents as a collaborator, or flipping the repo-wide "send secrets to fork PRs" toggle, both broaden secret distribution to every fork PR from that account or any fork. Moving the branch keeps secret scope tight to just the one PR being shipped.

Skill routing

When the user's request matches an available skill, ALWAYS invoke it using the Skill tool as your FIRST action. Do NOT answer directly, do NOT use other tools first. The skill has specialized workflows that produce better results than ad-hoc answers.

NEVER hand-roll ship operations. Do not manually run git commit + push + gh pr create when /ship is available. /ship handles VERSION bump, CHANGELOG, document-release, pre-landing review, test coverage audit, and adversarial review. Manually creating a PR skips all of these. If the user says "commit and ship", "push and ship", "bisect and ship", or any combination that ends with shipping — invoke /ship and let it handle everything including the commits. If the branch name contains a version (e.g. v0.5-live-sync), /ship should use that version for the bump.

Key routing rules:

  • Product ideas, "is this worth building", brainstorming → invoke office-hours
  • Bugs, errors, "why is this broken", 500 errors → invoke investigate
  • Ship, deploy, push, create PR, "commit and ship", "push and ship" → invoke ship
  • QA, test the site, find bugs → invoke qa
  • Code review, check my diff → invoke review
  • Update docs after shipping → invoke document-release
  • Weekly retro → invoke retro
  • Design system, brand → invoke design-consultation
  • Visual audit, design polish → invoke design-review
  • Architecture review → invoke plan-eng-review
  • Save progress, checkpoint, resume → invoke checkpoint
  • Code quality, health check → invoke health