Part of #2946. The hung-COUNT deadline race derived its verdict from a
post-await wall-clock re-check, which races the timer's own drift: on
loaded CI runners setTimeout callbacks fire measurably EARLY relative to
Date.now(), so each padding/boundary adjustment (>=, +1ms) only moved
which wrong status the partial-scan test received ('scanned', then
'partial').
The race's timeout arm now resolves a module-private sentinel; the
sentinel winning IS the deadline verdict (deadlineHit), consulted by the
post-await check without re-reading the clock. A COUNT that resolves
null (failed/absent count) stays distinguishable and does not skip the
scan; a COUNT that resolves slowly without the timer winning is still
caught by the retained wall-clock re-check. The +1ms pad is gone — the
sentinel makes timer drift irrelevant for the hung path.
Verified: partial-scan suite green 8 consecutive runs incl. the new
null-vs-sentinel distinction test.
Claude-Session: https://claude.ai/code/session_01KFWV7zBZFcmD94Vek6xRDT
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The v123 configurable-FTS migration (#2941) logged its completion notice
via console.log. Migrations run lazily inside any command's first DB
connect, so on the nightly heavy run (fresh Postgres service DB) the
line landed as the first line of `gbrain doctor --json` stdout and broke
the fm_wallclock jq parse ("Invalid numeric literal at line 1, column 7",
run 29731426470). runMigrations' contract routes all migration noise to
stderr; move the v123 prints (and the pre-existing v2 slug-rename print)
there.
Also un-vacuous the fm_wallclock harness: its register-source step used
`bun run -e` (bun dumps usage with exit 0 instead of running the code)
and `connect({})` (in-memory), so the source was never registered and
doctor scanned nothing. It now resolves the engine the way the CLI does
and registers the source in the DB doctor actually reads.
Regression test: test/migrate-stdout-clean.test.ts re-runs migrations
from v122 asserting zero stdout writes, plus a source-level guard that
migrate.ts contains no console.log.
Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(sync): --src-subpath + --exclude for monorepo subdir-source support (#753)
Rebased port of PR #774 onto current master. A single git repo can hold N
logical sources at subdirectories: git operations run at the discovered
repo root (git rev-parse --show-toplevel — worktrees and submodules
resolve natively) while file walking, imports, deletes and renames are
scoped to the subpath. Passing the subdirectory directly as the repo path
(auto-discovery) works through the same code path.
Path-containment guards (the point of the feature):
- NAV-1/NAV-2: the realpath-resolved scope must live inside the
realpath-resolved git root — ../-traversal and symlinked subdirs
pointing outside the repo are rejected before any git op.
- NAV-1 TOCTOU: per-file realpath re-validation during the incremental
import drain and rename reimport; symlink-escape files are recorded as
failures (fail-closed — the bookmark cannot advance past an escape).
- NAV-4: an --exclude set that filters out every candidate warns loudly.
Scoped syncs use git-root-relative slugs + source_path in BOTH the full
and incremental paths (runImport gains slugRoot), fixing the original
PR's full/incremental slug divergence in the auto-discovery flow.
--exclude matches scope-relative paths in both paths; exclusion never
deletes previously-imported pages. The full-sync delete reconcile is
scope-restricted and relativizes against the slug base so a healthy
scoped source can't trip the #2828 mass-delete valve. .gitignore
management resolves to the git root at every call site.
Preserves all master-side sync work since the original branch: the #2828
mass-delete safety valve, #1794 resumable checkpoints + pinned targets,
#1950 stall watchdog, #2335 heartbeat bump, #1970 bookmark reachability,
and the git-ls-files walker (#2315/#2462/#2678, whose symlink/cycle
hardening is untouched).
Co-authored-by: Jeremy Knows <jeremy@veefriends.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: listEverCommittedPaths uses gitContextRoot after #753 root-triple refactor
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: Jeremy Knows <jeremy@veefriends.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Squashed superset takeover of the FTS-language trilogy by @rafaelreis-r
(#580 env-var language for query+write side, #581 migration backfill,
#582 gbrain reindex-search-vector), rebased onto current master with the
security review's required fixes applied:
- Migration renumbered v116 -> v123 (master's v116 was already claimed by
code_edges_source_backfill; master is at v122).
- Restored the v120/#1647 search_path hardening: all four CREATE OR
REPLACE trigger-function bodies (migration handler + reindex command)
now carry SET search_path = pg_catalog, public, since CREATE OR REPLACE
resets proconfig and would otherwise strip the hardening on upgrade.
- reindex-search-vector: shared progress reporter (stderr phases
reindex_search_vector.pages/.chunks), id-keyset batched backfill
(5000 rows/UPDATE) instead of single whole-table statements, and
--json no longer bypasses the --yes/TTY confirmation gate.
- Allowlist validation regex + injection tests kept exactly as authored.
- Stale v33/v116 comments swept; docs/guides/multi-language-fts.md
written (README referenced it but no PR added it); llms bundles
regenerated via bun run build:llms.
- Trilogy tests quarantined as *.serial.test.ts (env mutation, per
check-test-isolation).
Verified live on PGLite: fresh init with GBRAIN_FTS_LANGUAGE=portuguese
produces portuguese-stemmed vectors with search_path pinned; the reindex
command retokenizes an english brain to portuguese in place.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Rafael Reis <rafael.reis@contabilizei.com.br>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Takeover of community PR #2387 with the security review's required fixes
applied. Original work by @harrisali0101.
With GBRAIN_RLS_SCOPE_BINDING=1, source-scoped Postgres read methods wrap
their queries in a transaction that binds set_config('app.scopes', $1, true)
(federated sourceIds CSV > scalar sourceId > '*') so operator-managed RLS
policies can filter rows at the SQL layer — defense-in-depth layer 2 under
the mandatory app-layer source filters.
Review fixes on top of the original PR:
- Flag-off is now a TRUE pass-through: no new per-read transaction wrap
(the #1794 PgBouncer pool-exhaustion class). Only the three search
methods keep a transaction when off — exactly the sql.begin() +
SET LOCAL statement_timeout wrap they already had on master — via the
helper's alwaysTransaction option.
- Preserved the PR's latent setseed fix: listCorpusSample pins
setseed() + SELECT to one connection when seeded (alwaysTransaction
gated on opts.seed), so the deterministic path can't split across
pooled connections.
- Updated the two postgres-engine shape tests to pin the new invariant
(search methods route through withScopedReadTransaction with
alwaysTransaction; helper owns the sql.begin(); flag-off path is
callback(this.sql)).
- New behavioral tests (test/postgres-engine-rls-scope.test.ts): flag-off
pass-through, flag-off alwaysTransaction, flag-on set_config emission,
federated > scalar > '*' precedence, and the CSV as a bound parameter
(never interpolated).
- Fixed the helper header comment to match the actual branching behavior.
- Operator docs in docs/ENGINES.md: env var, policy SQL, the
ALTER ROLE ... SET app.scopes='*' default requirement, FORCE ROW LEVEL
SECURITY for owner roles, and the honest caveat about unwrapped paths.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Harris <79081645+harrisali0101@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Five verified-open fixes to the dream/synthesize output family:
- #1586: thread the cycle's resolved sourceId (cycleSourceId) through
runPhaseSynthesize -> SubagentHandlerData.source_id -> subagent tool
OperationContext, and stamp collected refs + summary page with the same
source, so synthesized pages stop landing in 'default'.
- #2415: new config knob dream.synthesize.output_root (default 'wiki',
zero behavior change unless set) drives the synthesize prompt slug
templates, the patterns reflection lookup + prompt, and remaps the
filing-rule allow-list globs. Registered in KNOWN_CONFIG_KEYS.
- #2163: synthesize_concepts writes concept pages through
importFromContent (put_page's parse->chunk->embed pipeline) instead of
bare engine.putPage, so concepts/ pages are chunked + embedded and
reachable by retrieval.
- #2569: stampDreamProvenance persists dream_generated + dream_cycle_date
into pages.frontmatter (JSONB merge via executeRawJsonb) at write time,
so generated pages are DB-queryable and put_page write-through can't
erase the marker.
- #2606: chronicle judge detects stopReason 'length' truncation and
no-JSON-array parse failures as distinct skipped reasons
(judge_truncated / judge_parse_failed) instead of a false terminal
no_events; output cap raised to 4000 and configurable via
chronicle.judge_max_tokens.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Garry Tan <garrytan@gmail.com>
Three verified-open defects, one family: sync silently destroying or
diverging on content it should preserve.
#2404 (P0) — 'ops' was hardcoded in PRUNE_DIR_NAMES (a v0.2.0-era
carve-out), so any path with an ops segment was 'pruned-dir': committed
ops/*.md never imported, and modified ops/* files hit the
unsyncableModified delete loop (whose #1433 guard only spared
'metafile'), silently deleting put-created pages like the bundled
daily-task-manager's canonical ops/tasks on every sync. Fix: remove
'ops' from the prune list (ordinary user content; the vendor/generated
entries stay), and harden the delete loop to also skip 'pruned-dir' —
a page under a pruned dir can only exist via a deliberate put_page.
#2426 (P0) — write-through content stayed DB-only and was deleted by
sync --full. All three compounding bugs fixed:
1. writePageThrough now best-effort commits the artifact (path-limited
git commit) on durability-hardened repos, so the post-commit hook
can push it; result carries committed?: boolean.
2. scripts/brain-commit-push.sh stages+commits BEFORE any pull — the
old fetch+pull-rebase-first order aborted on any dirty tree, so the
helper could never commit a MODIFIED page; brain_push's
rebase-on-reject already handles an advanced remote.
3. The full-sync delete-reconcile partitions stale pages by git
history (listEverCommittedPaths): never-committed source_paths are
DB-only write-through — pages are KEPT and re-exported to the
working tree instead of soft-deleted. Builds on the #2828
mass-delete valve (covers the below-valve cases).
#2607 — the sync --full git ls-files fast path bypassed pruneDir, so a
full pass imported (and resurrected soft-deleted) pages under dot-dirs
and vendored trees that incremental sync excludes. Fix:
isCollectibleForWalker applies the same segment-level pruneDir gate as
classifySync, so full and incremental enumeration agree.
One regression test per defect (all verified failing against master
src): test/sync-ops-pages.serial.test.ts,
test/write-through-commit.serial.test.ts, the #2426 helper-order test
in test/brain-durability-hook.serial.test.ts,
test/sync-reconcile-db-only.serial.test.ts,
test/import-git-fastpath-prune.test.ts.
Fixes#2404Fixes#2426Fixes#2607
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Port of #1837 (mvanhorn) onto current master. The extract_facts cycle
phase unconditionally wipe-and-reinserted every page's fence-owned DB
rows, so re-running a cycle on unchanged content churned rows and — on
the Postgres engine reported in #1781 — accumulated duplicates each run.
The phase now de-dupes extracted facts by the canonical (claim, source)
content key and reconciles the page-scoped DB index: no-op when already
in sync, insert-only for new keys, wipe/reinsert only when stale rows
need cleanup.
Adjustments over the original PR to fit current master:
- preserve #1972's abortSignal threading into the batch embed call
- preserve #1928's excludeSourcePrefixes: ['cli:'] on every wipe, and
exclude cli:-origin conversation facts from the existing-row set so
they neither count as stale (which would force a wipe every cycle)
nor get compared against the fence
Fixes#1781
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Four verified-open issues in the subagent-orchestration family, one PR:
- #1594: dream synthesize subagent job/wait timeouts promoted from
hardcoded 30/35-min constants to config keys
dream.synthesize.subagent_timeout_ms / subagent_wait_timeout_ms.
Approach ported from stale PR #1596 (credit @ai920wisco).
- #2778: add_timeline_entry joins the subagent brain-tool allowlist,
fenced fail-closed by the shared enforceSubagentSlugFence (extracted
from put_page's inline check — same trusted-workspace allow-list /
wiki/agents/<id>/ namespace policy). The per-turn output cap is now
resolveMaxOutputTokens (data.max_tokens → agent.max_output_tokens →
8192, was hardcoded 4096); a max_tokens stop surfaces as
stop_reason 'max_tokens' instead of a silent end_turn, and a
mid-tool-round cap hit injects a truncation note so the model
re-issues the dropped call.
- #2782: patterns phase status now reflects the child outcome —
non-complete outcome with zero writes → fail (PATTERNS_CHILD_<OUTCOME>),
partial writes → warn. Patterns timeouts get the same config-key pair
(dream.patterns.subagent_timeout_ms / subagent_wait_timeout_ms).
- #2113: facts extraction cap is config facts.extraction_max_tokens
(default 4000, was hardcoded 1500); stopReason 'length' is checked,
retried once at 2x the cap, and surfaced on stderr instead of
silently extracting zero facts.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: ai920wisco <ai920wisco@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Absorbs PR #2279's intent (drop the hardcoded ANTHROPIC_API_KEY gate) with
the end-state the gateway world actually wants: the patterns phase now
probes the RESOLVED patterns model through probeChatModel(normalizeModelId)
— the same semantics as think/index.ts and synthesize's makeJudgeClient.
Fixes two misclassifications of the old env gate:
- Non-Anthropic stacks (litellm, deepseek, openrouter, ...) were skipped as
"no upstream" even though the subagent routes them through the gateway
(agent.use_gateway_loop). They now pass the gate; their auth is checked
lazily at dispatch and surfaces in the job outcome.
- Anthropic keys set via `gbrain config set anthropic_api_key` (stdio MCP
launches without shell env) were treated as missing. hasAnthropicKey
inside probeChatModel reads both sources.
Skip reason renames no_api_key → no_provider (carrying the probe's detail).
Both pinning tests updated: the structural test asserts the probe wiring;
the PGLite E2E swaps its env-only helper for the shared hermetic
withoutAnthropicKey (env + config file) so it can't flip to a live LLM call
on a dev machine with a config-file key.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: brettdavies <brettdavies@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Ports the still-unmerged remnant of PR #1891 (#1593 follow-up). Master's
getConfig gained retry-with-reconnect in #1603, but its siblings —
setConfig, unsetConfig, listConfigKeys — still touched `this.sql` bare, so
the first config write/list after a mid-cycle instance-pool teardown threw
the retryable "No database connection" (issue #1678) unhandled instead of
rebuilding the pool.
Adds the connRetry helper from #1891 (same retry+reconnect posture as
batchRetry, but no batch audit JSONL — a config accessor is not a sized
batch), refactors getConfig onto it, and wraps the other three. Writes are
safe to retry: withRetry only retries connection-class failures and both
writes are idempotent (upsert / delete).
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: jalagrange <jalagrange@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Ports the still-unmerged half of PR #2442. OpenAI caches prompt prefixes
automatically, but a stable prompt_cache_key keeps requests that share a
prefix on the same inference engine, lifting the automatic-cache hit rate.
chat() now derives a stable key from the system prompt + sorted tool names
for native-openai models and passes it via providerOptions.openai.
promptCacheKey. Config provider_chat_options still overrides the derived
key; anthropic/google/openai-compatible providers are untouched.
The Anthropic half of #2442 (cache_control "silent no-op") is superseded:
@ai-sdk/anthropic 3.0.74 forwards call-level providerOptions.anthropic.
cacheControl as the Messages API's request-level cache_control (automatic
prefix caching), so master's existing marker is live on current deps.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: CoachRyanNguyen <CoachRyanNguyen@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Port of #1731 (diazMelgarejo) onto current master. gbrain import wrote
~/.gbrain/import-checkpoint.json with the caller's raw dir argument, so
a checkpoint left behind by an interrupted run (e.g. SIGTERM) could carry
"." or a symlinked spelling — an identity that resolves to whatever CWD
the next consumer happens to run from. Downstream tooling that treated
the checkpoint dir as an owned staging boundary could then act on the
wrong directory.
- runImport captures the import target ONCE via resolveImportTargetDir
(resolve + realpathSync) and threads that canonical value through
collection, checkpoint load/save, and resume filtering
- checkpoints are self-describing (schema_version: 1, owner: "gbrain",
kind: "import"); loadCheckpoint tolerates absent metadata (legacy
path-based files) but rejects present-and-wrong metadata and any
relative dir
- checkpoint contract documented in docs/guides/live-sync.md (llms
bundle regenerated)
- test/import-resume.test.ts fixture now realpaths its tmpdir so planted
checkpoints match the canonicalized dir (macOS /var -> /private/var)
Fixes#1728
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Lawrence Melgarejo <Lawrence@cyre.me>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* test: symlink-walker tests use a non-metafile probe (README now skipped by design, #2315)
The import walker deliberately skips README/metafiles since #2315 (closing
#345); the symlink-hardening tests used README.md as their probe file and
went red on the intersection. Probe with notes.md instead; test intent
(cycle hardening + strategy filter) unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: grant security-events write to the OSV caller job — reusable workflow requires it at startup
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
`parseMarkdown` previously walked the lines after the opening `---` and
recorded the first `^#{1,6}\s`-shaped line as a `headingBeforeClose`,
then flagged MISSING_CLOSE when that index came before the actual closing
fence. YAML allows `#` comment lines anywhere inside the document, so a
template that leads with annotation comments inside the fence (e.g. a
`# Research Template` header before the keys) hit a false-positive
MISSING_CLOSE even though the closing `---` was present.
Fix: only walk for the closing `---`. When it is found, content between
the fences is YAML; `#` lines are comments, not headings. When the close
is genuinely missing, surface the first heading-shaped line as a
where-it-went-off-the-rails hint (this path was already correct; we keep
it for the genuine missing-close case).
Two regression tests added to `test/markdown-validation.test.ts`:
- `#` comment lines at the top of a closed frontmatter
- `#` comment lines interleaved with keys
All 68 tests across the four markdown/frontmatter test files stay green.
* fix(extract,ingest,cycle): source-provenance wave — thread source identity through ingest_capture, fs-walk links, and cycle extract (#1522#1747#1503)
Three fixes in the same invariant class (source identity silently dropped
on the write path, collapsing to the 'default' source):
- #1522: the ingest_capture Minion handler validated IngestionEvent
provenance (source_id/source_kind/source_uri) then dropped it on the
importFromContent call. Now threads source_kind/source_uri +
ingested_via='ingest_capture' into the page write, and routes the write
under event.source_id when it names a registered source AND the event
is trusted (fail-closed: untrusted webhook payloads carry a
caller-controlled x-gbrain-source-id header and must not choose their
write source; unregistered emitter ids keep default routing so the
webhook path can't FK-fail).
- #1747: the fs-walk extractors (extractLinksFromDir /
extractTimelineFromDir / extractForSlugs) built batch rows with no
source_id, so addLinksBatch/addTimelineEntriesBatch mapped missing →
'default' and the pages JOIN dropped every row on a non-default source
("Links: created 0 from N pages", no error). ExtractOpts gains
sourceId; the CLI fs path resolves it via resolveSourceId
(--source-id > env > dotfile > registered path > sole-non-default)
and rows are stamped from/to/origin_source_id + timeline source_id.
- #1503: the cycle's extract phase (runPhaseExtract) never passed a
sourceId, so federated-brain dream/autopilot cycles persisted nothing
every night. It now threads cycleSourceId (explicit --source or
resolveSourceForDir(brainDir) — the same seam runPhaseSync uses) into
runExtractCore for both the incremental and full-walk paths.
Regression tests: handler provenance write-through (registered/
unregistered/untrusted routing), fs-walk + CLI resolution + incremental
cycle route landing edges/timeline in the right source (all red on
unfixed master), and a negative control pinning the pre-fix JOIN-drop
shape.
Reuses the ExtractOpts.sourceId threading approach from PR #1719,
rebased onto the current signal-aware signatures and extended to the
cycle + timeline paths.
Fixes#1522Fixes#1747Fixes#1503
Co-authored-by: seungsu <kss530c@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: update cycle source pin for cycleSourceId threading
The #1972 source-pin test asserts the literal runPhaseExtract call site in
cycle.ts. The #1503 fix appended cycleSourceId after opts.signal; signal
threading is unchanged (still the 5th arg, forwarded to runExtractCore).
Pin updated to the new literal — invariant intact.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: seungsu <kss530c@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Recover the worker-owned Postgres pool when promoteDelayed escapes a retryable connection error, preventing the repeated Promotion error: No database connection loop from issue #1491.\n\nAdds a regression test proving reconnect happens before the worker continues to claim work.
Closes#345.
The bulk-import walker isCollectibleForWalker filtered admitted files by
extension only, while incremental sync excludes README/index/log/schema via
isSyncable -> SYNC_SKIP_FILES. A directory import therefore ingested every
directory README as a folder-titled ghost page that trigram-corrupts fuzzy
entity resolution and inflates orphan count. Apply SYNC_SKIP_FILES (basename
guard) at the top of isCollectibleForWalker so both the FS-walk and the
git-fast-path collection routes agree with sync.
Also add RESOLVER.md to SYNC_SKIP_FILES: a structural routing metafile
(docs-aligned with schema.md/index.md/log.md/README.md), not indexable content.
Co-authored-by: ElliotDrel <ElliotDrel@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The instance-pool reconnect() did disconnect() (nulling _sql) BEFORE connect(),
so a connect() failure during a transient Postgres blip left _sql null for the
rest of the process — every subsequent non-retry-wrapped call then fell through
to the never-connected module singleton and threw 'No database connection',
crashing the autopilot worker into a respawn loop. Build-then-swap: snapshot the
live pool, build a fresh one, end the old only once the new validates, restore
on failure. Keeps upstream's reap-detection + pool-recovery audit. Confirmed by
jalagrange on closed PR #1593; this is the Layer-1 fix he left to us.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
D4 invariant ("never advance last_commit on partial", sync.ts comment)
preserved. last_sync_at is a monitoring signal read by doctor
sync_freshness (warn 24h / fail 72h), separate from the import-converged
bookmark. Without this heartbeat write, a cron-driven */15 sync over a
quiet vault pins last_sync_at to the last real commit, so doctor
falsely flags the source as stale for as long as the vault is quiet.
Reproduction (5 lines, no gbrain install required):
1. Setup a fresh obsidian source + commit a single .md file
2. gbrain sync --source obsidian # first_sync, last_sync_at = NOW
3. (do nothing) gbrain sync --source obsidian # up_to_date
4. SELECT last_sync_at FROM sources WHERE id = 'obsidian';
5. Observed: still pinned to step 2. Expected: bumped to step 3.
Fix: in the up_to_date early-return (sync.ts line ~1786), execute a
single `UPDATE sources SET last_sync_at = now() WHERE id = $1` before
returning. The D4-protected writeSyncAnchor path is untouched.
Test: test/sync.test.ts adds a describe block that runs two
consecutive syncs against a quiet vault and asserts last_sync_at
advances while last_commit is unchanged. PGLite + executeRaw pattern
matches the existing #1970 test scaffold.
Workaround: hourly psql touch in WSL crontab documented at
https://github.com/garrytan/gbrain/issues/[link-to-issue]
Co-authored-by: lost9999 <lost9999@users.noreply.github.com>
* fix(cli,config,doctor): CLI/config UX wave — config-get file plane, idempotent archive, honest help + doctor text, prefixed model defaults (#2120#2792#1175#1123#2451)
- config get resolves the file/env plane before the DB plane (runtime
precedence) and reports provenance on stderr; stdout stays a bare value.
- sources archive distinguishes already-archived (friendly no-op, exit 0)
from not-found (clear exit-4 error).
- gbrain --help SOURCES block now lists archive/restore/archived/purge/status
plus a pointer at `sources --help` for the long tail.
- multi_source_drift doctor advice references only real CLI surfaces
(drops the never-built 'sources rehome'; pins delete to GBRAIN_SOURCE=default).
- #2451 (bare model ids in calibration defaults) verified already fixed +
tested on master by #2892 — no change needed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: satisfy test-isolation gate for wave-B tests
check-test-isolation R1 forbids direct process.env mutation in non-serial
unit tests (env leaks across files sharing a shard process). Route the
GBRAIN_HOME / GBRAIN_CHAT_MODEL / GBRAIN_PGLITE_SNAPSHOT overrides through
the canonical withEnv() helper instead.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Rebase-port of #2452 (spinsirr:fix/calibration-profile-scope-and-cli) onto
current master after tonight's merges made the fork branch conflict.
- cli: add 'calibration' to CLI_ONLY so dispatch reaches its existing
handler instead of falling through to "Unknown command" (#2035); honor
--source / GBRAIN_SOURCE in the calibration CLI.
- takes: route takes_list / takes_search / takes_scorecard /
takes_calibration through sourceScopeOpts(ctx) (federated array > scalar
> nothing) and scope engine reads via the take's page.source_id — JOIN
filter for list/search, EXISTS for scorecard/curve — on both engines
(#2200-class).
- bigint: shared takeHitRowToHit coercion in searchTakes /
searchTakesVector (both engines) + bigintToStringReplacer on the cli.ts
output normalizer and the `gbrain call` exit, so int8/BIGSERIAL columns
no longer crash JSON.stringify (#2450); calibration profile id
BIGSERIAL → string.
- calibration: default model ids route through TIER_DEFAULTS
(provider-prefixed) instead of bare model strings; admin calibration
chart endpoints fixed (takes has no page_slug column; month-precision
since_date; Date generated_at; bigint id in drill-down).
The think/gather source-scope slice of the original PR was dropped: it
already landed on master via #2739.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: spinsirr <ID+spinsirr@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
classifyPgliteInitError() routes bare Emscripten aborts to the 'unknown'
verdict, whose hint unconditionally printed "Most common cause: the macOS
26.3 WASM bug (#223)" — even on Windows and Linux (#2195, #1870).
- buildPgliteInitErrorMessage now takes a platform param (default
process.platform): darwin keeps the #223 link as a *possible* cause;
other platforms get the plausible off-macOS causes (lock contention,
damaged data dir) plus `gbrain doctor` / `gbrain reinit-pglite`.
- New stringifyPgliteInitError(): non-Error rejections (Emscripten
aborts can throw plain objects) no longer print "[object Object]".
- Regression tests for both branches + the stringifier in
test/pglite-init-classifier.test.ts.
Canonical for a 7-report class: #2674, #1870, #1195, #1502, #2195,
#939, #391.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The stdio parent-death watchdog was hard-wired to spawnSync('ps'), which
does not exist on Windows. The startup probe therefore failed on every
Windows host, the watchdog was permanently disabled ("watchdog disabled:
ps unavailable"), and an orphaned `gbrain serve` held the PGLite write
lock until reboot. Orphans are especially easy to produce on Windows
because MCP hosts launch the server through a cmd.exe wrapper (.bat) and
killing the wrapper does not kill the bun child.
Windows never re-parents orphans, so the cached process.ppid stays
correct for the process lifetime and the watchdog question inverts from
"did the live PPID change?" to "is the original parent still alive?" --
answered in-process with a signal-0 existence probe (process.kill(ppid, 0),
OpenProcess under the hood). No external binary needed. EPERM counts as
alive. Parent dead reports PID 0, which differs from initialParentPid and
fires the existing shutdown path.
- readLiveParentPid / probeWatchdogAvailable: platform split (ps on
POSIX, signal-0 on win32), exported with a platform test seam so CI on
any OS exercises both branches.
- isPidAlive: shared exported helper.
- Watchdog install guard tightened from `!== 1` to `> 1` so a PID-0
"parent already gone" report cannot install a phantom interval that
compares 0 to 0 forever.
- Disabled-mode log generalized (no longer claims ps is the only
mechanism); existing probe-fail test updated to match.
- New unit suite for the platform defaults (6 tests).
Verified live on Windows 11: serve spawned via a .bat wrapper, wrapper
killed without closing stdin -> bun child exits within ~10s and the
PGLite lock dir is released. Before this change the child survived
indefinitely and every subsequent gbrain invocation timed out waiting
for the lock.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Add provider_chat_options alongside provider_base_urls and thread it through
the gateway config path into chat(). The chat request now deep-merges
provider-scoped options and model-scoped overrides into providerOptions keyed
by recipe id, preserving existing gateway-built options such as Anthropic
cacheControl and leaving absent-config behavior unchanged.
This lets operators disable thinking for small-budget hybrid-reasoning utility
calls without hardcoding that behavior for every use of those models.
The parser/doctor/extractor read the polished page body (compiled_truth +
timeline) instead of the raw turn-by-turn transcript that meeting pages store
in a `raw_transcript` frontmatter sidecar, and the brain's actual raw format
(`Speaker A: ...` / `Speaker B: ...`) had no built-in pattern. Result: scan
returned no_match / 0 messages and conversation-fact extraction produced 0
segments / 0 facts.
- new readConversationBodyForParsing(): prefer the raw_transcript sidecar when
present, fall back to compiled_truth + timeline (src/core/conversation-parser/body.ts)
- wire it into conversation-parser scan, doctor coverage check, and
extract-conversation-facts (drops the old readPageBody helper)
- add a built-in `speaker-letter-no-time` pattern for plain `Speaker A:` lines
Verified: scan now returns phase=regex_match (speaker-letter-no-time), and the
Ben page extracts 61 facts / 7 segments. Tests added; suite green.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
since/until range filters were applied against updated_at, so edited-but-old
entries leaked into time-bounded queries. Filter on the effective date instead.
Closes#1520
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
* feat(extract): recognize inline [Source: ..., YYYY-MM-DD] citations as timeline entries
gbrain's own quality conventions (skills/conventions/quality.md) require a
dated [Source: ..., YYYY-MM-DD] citation on every brain write, so curated
pages are full of dated evidence — but extractTimelineFromContent only
recognized the timeline-bullet and date-header formats. A page whose dates
all live in citations scored zero timeline coverage in brain_score, and
doctor pointed users at a formatting convention their own citations already
satisfied in spirit.
Format 3 files one entry per citation: date and source from the marker,
summary from the annotated line with citation markers stripped. Lines
already captured by Format 1 are skipped so a timeline bullet carrying its
own citation is not double-filed. Bare citations with no surrounding text
are ignored.
Idempotency is unchanged: persistence already dedupes at the DB layer.
* fix(extract): Format 3 citations also in parseTimelineEntries (db-source + ingest path)
The first commit only taught extractTimelineFromContent (fs-source) the
citation format; the db-source extract and the ingest path parse through
parseTimelineEntries in core/link-extraction.ts, which still could not see
citations. Same rules as the fs parser: bullet-captured lines skipped,
bare citations ignored, invalid calendar dates rejected; the citation
source is preserved in the entry detail.
---------
Co-authored-by: pabloglzg <186649799+pabloglzg@users.noreply.github.com>
extractTakesFromPages selected pages by updated_at DESC + LIMIT with no
exclusion of pages that already hold takes. The CLI clamps --max-pages to
1000, so on a corpus larger than one run every re-run rescanned the same
most-recent 1000: the older tail could never be bootstrapped, and each
rescan re-spent Haiku budget producing upsert-identical rows. Seen live on
a 2,311-eligible-page brain — a second run would have covered 0 new pages.
Covered pages are now skipped by default (NOT EXISTS on takes.page_id), so
repeat runs sweep a large corpus in slices; --include-covered restores the
full rescan for refresh use. Usage text documents both plus the 1000 clamp.
Claude-Session: https://claude.ai/code/session_01FQgByq4aqQq2PP8UHCdnfk
Co-authored-by: Paolo Belcastro <p3ob7o@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
extract_atoms minted duplicate atom pages two ways:
A. Trailing-dash twins. The local slugger truncated the title at 60 chars
with no re-strip, so a cut landing on a hyphen left a trailing dash
(`…would-`). The FS-import normalizer (slugifySegment) strips it
(`…would`), so the same atom persisted under two slugs and the
dedup-by-slug check never collapsed them.
B. Cross-day re-mint. The slug used the run date (todayDate()), while the
idempotency guard keys on the whole-file source_hash. Append-only
sources (chat/transcript exports) grow daily, so the file hash changes,
the guard never matches, the source is re-extracted, and each re-mint
lands under a new date prefix → a new slug → no upsert → a duplicate.
Fix: make the atom slug a deterministic function of stable inputs —
`atoms/<source-date>/<stem>-<title-hash>`:
- source date is parsed from the source ref (transcript filename / page
slug), not the run date, so re-extraction converges on the same slug and
putPage upserts instead of duplicating;
- the 6-char title hash keeps two atoms whose titles share the first 60
chars on distinct slugs (no silent clobber of a different atom);
- the stem routes through the canonical slugifySegment and re-strips a
trailing dash, so the two write paths can no longer disagree.
The whole-file source_hash batch check is retained only as a cost fast-path
(skip re-running the model on an unchanged source); correctness no longer
depends on it.
Adds a hermetic PGLite regression test (no DATABASE_URL) asserting the
source-dated prefix, title-hash suffix, trailing-dash strip, and upsert on
re-extraction of a grown append-only transcript.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A lock file can outlive its autopilot process after a crash or forced termination. The previous mtime-only check treated that file as proof that another instance was running, so supervisor restarts could exit repeatedly until the file aged out.
Read the holder PID and probe it with signal 0. Keep the lock whenever that process is alive; take over only dead, malformed, empty, or self-owned locks. EPERM remains conservative and counts as alive, preventing two autopilot instances from running concurrently.
Add hermetic tests for missing, live, dead, malformed, and empty lock states.
Atom-extraction page discovery hardcoded EXTRACTABLE_PAGE_TYPES and ignored the
active pack's `extractable: true` flags, so a type declared extractable in the
manifest (e.g. `note`) never actually extracted. Closes the D2 TODO the code
already flagged (extract-atoms.ts: 'future pack-aware refactor ... pull from the
active pack manifest').
Resolve the allowlist as: legacy hardcoded floor UNION the pack's extractable
types, MINUS synthesis outputs (atom, concept — extracting from these would loop,
since concepts are synthesized from atoms). Mirrors facts/eligibility.ts, which
excludes concept the same way. Back-compat: gbrain-base brains keep every legacy
target via the union; fail-soft falls back to the legacy floor if the pack can't
load.
- pure unionExtractableTypes() policy, unit-tested
- discoverExtractablePages + countExtractAtomsBacklog resolve from the pack
- page-discovery fixture updated (note now extracts; concept stays excluded)
Co-authored-by: Paolo Belcastro <p3ob7o@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`schema use` hardcoded `gbrain-base` (and a fixed bundled list), so other
bundled packs could not be selected by name. Use the shared BUNDLED_PACK_NAMES
set and resolve `<name>.yaml` generically so every bundled pack activates.
Closes#1574
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
claude-sonnet-5 and claude-fable-5 are GA Anthropic models, but neither
was in CANONICAL_PRICING. A brain routing a tier to them (e.g.
models.tier.reasoning = anthropic:claude-sonnet-5) ran with cost
telemetry blind on that tier: canonicalLookup missed, the budget meter
logged BUDGET_METER_NO_PRICING and disabled the gate, and cost views
under-reported spend.
- model-pricing.ts: anthropic:claude-sonnet-5 at $3/$15 and
anthropic:claude-fable-5 at $10/$50. Sonnet 5's launch intro discount
($2/$10 through 2026-08-31) is deliberately not modeled — the table
carries standard rates so estimates stay conservative and the entry
needs no time-bombed edit when the promo lapses.
- takes-quality-eval/pricing.ts: claude-sonnet-5 added to the curated
SUPPORTED_MODELS allowlist (a likely judge override). Fable 5 stays
out — priced for warn-only consumers, not a budgeted-eval panel model.
- model-pricing.test.ts: pin tests for both rows, matching the existing
Opus 4.8/4.7 pattern. Drift guards iterate the table; no changes.
All derived views (ANTHROPIC_PRICING bare view -> budget-tracker,
batch-projection, budget-meter) pick the rows up automatically.
Co-authored-by: Paolo Belcastro <p3ob7o@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
resolveHardExcludes() only ran at DB-query build time (cache miss), so
query_cache rows written by a process without GBRAIN_SEARCH_EXCLUDE could
be served to a process with it (and vice versa), leaking excluded slugs.
hybridSearchCached now resolves the effective hard-exclude list exactly
as the engines' query-build path does and folds it (sorted, append-only
hx= part) into knobsHash via a new KnobsHashContext.hardExcludes field.
KNOBS_HASH_VERSION 11 -> 12: one-time global cache cold-miss on upgrade,
refills within cache.ttl_seconds.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Collector branch superseding the gateway tool-loop duplicate cluster and
adjacent provider fixes. Re-implemented from the best of each PR (deduped by
content, not file-overlap); every fix carries test coverage.
Gateway tool-loop resume (supersedes #1934#2062#2065#2112#2274#2487#2336#2257#2499#2491, test #2063):
- toolLoop now persists the tool-result user turn per round (onToolResultTurn),
so a resumed subagent job reloads a balanced transcript instead of dangling
assistant tool-calls that non-Anthropic providers reject with
AI_MissingToolResultsError.
- runSubagentViaGateway reconciles an already-corrupted transcript on resume:
it heals every dangling assistant tool-call turn (not just the tail) from the
settled subagent_tool_executions rows, re-dispatching idempotent-pending
tools and throwing on non-idempotent, mirroring the legacy Anthropic path.
Terminal early-return for a transcript that already reached end_turn.
- repairToolPairing() is a last-resort normalization at the chat() boundary
(from #2336): back-fills error stubs for any assistant tool-call still
unanswered (partial turns, provider-duplicated/dropped IDs on local models,
length-truncated batches). No-op on balanced input.
- toModelMessages is Date-safe (Postgres timestamptz -> ISO via a JSON round
trip at the SDK boundary, never a ::jsonb cast; degrades bigint/circular to a
string instead of throwing) and drops non-string text blocks reasoning models
emit that AI SDK v6 rejects (#2488).
Adjacent provider fixes:
- Model-aware default max output tokens: thinking-by-default Claude 5 models get
headroom (gateway 32000, think 16000) while everything else stays 4096/4000,
so DeepSeek/OpenAI subagents don't exceed provider caps (#2614#2806).
- DeepSeek: promote reasoning_content into content when content is empty, via a
fail-open recipe fetch shim (#2617).
- OpenRouter: map openrouter_api_key (config + env) into OPENROUTER_API_KEY
through buildGatewayConfig; register agent.use_gateway_loop, zeroentropy and
openrouter keys in KNOWN_CONFIG_KEYS so `config set` accepts them (#2572,
config key from #2112).
Preserves JSONB (no JSON.stringify into ::jsonb), engine parity, source
isolation, and trust-boundary invariants. Verified with the gbrain-pr-test-env
consumer matrix (clawlancer/Postgres, gstack + hivemindos/PGLite) baseline-FAIL
-> candidate-PASS on the same repro, plus bun run verify (31/31) and 439
targeted unit + e2e tests.
Co-authored-by: Sinabina <sinabina@Sinabinas-MacBook-Pro-4.local>
Co-authored-by: thomaskong119 <thomaskong119@hotmail.com>
Co-authored-by: maxpetrusenkoagent <max.petrusenko.agent@gmail.com>
Co-authored-by: brettdavies <brettdavies@users.noreply.github.com>
Co-authored-by: ivandebot <187176982+ivandebot@users.noreply.github.com>
Co-authored-by: Rafael Reis <rafael.reis@contabilizei.com.br>
Co-authored-by: fbal23 <fbal.public@gmail.com>
Co-authored-by: javieraldape <javieraldape@users.noreply.github.com>
Co-authored-by: David Carolan <david@joyrestart.com>
Co-authored-by: Masashi-Ono0611 <masashi.ono.0611@gmail.com>
Co-authored-by: spiky02plateau <155588579+spiky02plateau@users.noreply.github.com>
Co-authored-by: psam-717 <mphilannorbah@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(config): register Life Chronicle keys so the documented enable command works
The v0.42.56.0 release notes say `gbrain config set auto_chronicle true`,
but the key was never added to KNOWN_CONFIG_KEYS — the documented command
fails with 'Unknown config key' and the operator has to discover --force by
reading source. Registers 'auto_chronicle' plus the 'chronicle.' prefix
(chronicle.tz and future knobs). Same registration class as the v0.42.42.0
spend-controls fix. Regression test pins both.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FQgByq4aqQq2PP8UHCdnfk
* fix(config): register takes.bootstrap_enabled too — same unregistered-key class
Hit while enabling the takes bootstrap on a live brain: the onboard
remediation's documented enable key fails 'Unknown config key' exactly like
auto_chronicle did. Registered + pinned by the same regression test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FQgByq4aqQq2PP8UHCdnfk
---------
Co-authored-by: Paolo Belcastro <p3ob7o@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>