v0.41.27.0 fix: withRetry self-heals on null singleton + facts:absorb drain + disconnect audit (closes #1570) (#1608)

* merge master: rebump v0.41.25.0 → v0.41.27.0 (queue collision)

Master shipped v0.41.25.0 (#1538 batched sync deletes) and v0.41.26.0
(#1571 dream --source fix) while this branch was in flight. Conflict
resolution rebumps to the next available slot.

- VERSION: 0.41.25.0 → 0.41.27.0
- package.json: synced
- CHANGELOG.md: my v0.41.27.0 entry placed above master's v0.41.26.0
  and v0.41.25.0; in-entry version references updated 0.41.25.0 →
  0.41.27.0 and forward-references bumped to v0.41.28+.
- TODOS.md: kept master's v0.41.20.x section + my v0.41.27.0+ follow-ups

No source-file conflicts during the merge.

* feat(diagnostics): db-disconnect audit + doctor surface (v0.41.27.0)

Instruments every db.disconnect() and PostgresEngine.disconnect() call
with a JSONL audit record so the next user-reported #1570 cycle gives
us the offender's caller stack instead of the symptomatic
"No database connection" error.

Audit shape (~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl):
  {ts, engine_kind, connection_style, caller_stack[], command, pid}

- src/core/audit/db-disconnect-audit.ts (NEW): the audit writer,
  built on the v0.40.4.0 createAuditWriter cathedral. Captures a
  6-frame stack via new Error().stack so the offender is readable
  without spending stderr noise.
- src/core/db.ts: logDbDisconnect call at the top of disconnect()
  (best-effort; never blocks the real teardown).
- src/core/postgres-engine.ts: same instrumentation in
  PostgresEngine.disconnect() — distinguishes 'module' vs 'instance'
  connection_style so we can tell legitimate worker-pool teardowns
  apart from the load-bearing module-singleton class.
- src/commands/doctor.ts: extends batch_retry_health to surface
  24h disconnect count + most-recent caller stack. Warns when the
  caller frame isn't a known CLI-exit frame (e.g. cli.ts's finally
  block at the end of an op-dispatch). This is the diagnostic that
  tells v0.41.28+ where to apply the real ownership fix.
- test/db-disconnect-audit.test.ts: unit coverage for the audit
  writer + caller-stack capture + JSONL shape.
- test/e2e/db-singleton-shared-recovery.test.ts: real-Postgres
  regression that exercises the singleton-null path end-to-end.

Refs #1570

* feat(retry): self-heal on null singleton — closes #1570 symptom (v0.41.27.0)

withRetry gains an opt-in reconnect callback that fires between the
isRetryableConnError classification and the inter-attempt sleep.
PostgresEngine.batchRetry injects this.reconnect() — race-safe via
the existing _reconnecting guard, handles module and instance pools.

Closes the production loss reported in #1570: dream cycles on Supabase
no longer drop ~150 link rows per cycle when the singleton goes null
mid-batch. The retry now rebuilds the connection between attempts so
the second try has somewhere to write to.

- src/core/retry.ts: WithRetryOpts gains `reconnect?: () => Promise<void>`.
  Awaited in the catch branch. onRetry is also now awaited (back-compat-
  safe: every existing in-tree caller is a sync arrow). Reconnect
  failures propagate as the real cause — replaces the symptomatic
  "No database connection" error with whatever the connect() throw
  was, so operators see the truth.
- src/core/postgres-engine.ts:batchRetry — injects
  `reconnect: () => this.reconnect()`. Covers all 9 batch-retry call
  sites (addLinksBatch, addTimelineEntriesBatch, upsertChunks, plus
  the 6 caller-supplied auditSite labels in extract / sync / reindex).
- test/core/retry-reconnect.test.ts: 8 hermetic cases pinning the
  contract — reconnect fires before sleep, only on retryable errors,
  back-compat when omitted, signal-aborted bypasses reconnect,
  onRetry is awaited, full success path end-to-end.

The deeper bug (who's calling disconnect mid-cycle) is left
unaddressed in this commit by design — the diagnostic instrumentation
in the prior commit will tell us in the next production run.

Refs #1570

* feat(facts): drainPending() + CLI await before disconnect (v0.41.27.0)

Closes the silent 'No database connection' tail-end errors after
gbrain capture / put_page: the facts:absorb fire-and-forget queue
sometimes outlived the CLI process's connection lifetime, so absorb
attempts after engine.disconnect() landed in stderr as the
GBrainError shape.

- src/core/facts/queue.ts: new drainPending({timeout: 1000}) method
  distinct from shutdown(). Stops accepting new enqueues, awaits
  in-flight settle, bounded by timeout, returns count of unfinished.
  Semantically different from shutdown() (which aborts in-flight)
  so the symptom — drop work that hasn't started yet but let
  in-flight work finish — matches what CLI exit actually needs.
- src/cli.ts: op-dispatch finally block awaits the drain BEFORE
  engine.disconnect(). Bounded 1s. Opt-out env GBRAIN_NO_FACTS_DRAIN
  for callers that don't enqueue (keeps fast-exit paths fast).
  Mirrors the v0.41.8.0 awaitPendingLastRetrievedWrites pattern.
- test/facts-queue-drain-pending.test.ts: 6 hermetic cases — empty
  drain returns immediately, single in-flight settles, timeout
  bounds wait, shutdown-after-drain is idempotent, post-drain
  enqueues are dropped, signal-aborted skips waiting.

Refs #1570

* docs: update project documentation for v0.41.27.0

README.md: added troubleshooting entry for the v0.41.27.0 retry-reconnect
+ facts:absorb drain fix (closes #1570), pointing operators at
`gbrain doctor --json` to find the offending disconnect caller.

CLAUDE.md: extended `src/core/retry.ts` entry with the new optional
`reconnect` callback (v0.41.27.0); added two new Key Files entries for
`src/core/audit/db-disconnect-audit.ts` (the diagnostic half of the
"instrument first, fix later" pivot) and `FactsQueue.drainPending`;
extended `doctor.ts:checkBatchRetryHealth` entry with the in-place
extension that surfaces 24h disconnect-call count.

llms-full.txt: regenerated to absorb CLAUDE.md edits.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore: rebump v0.41.27.0 → v0.41.28.0 (queue collision with #1573)

Master shipped v0.41.27.0 (#1573 git-aware sync_freshness) claiming the
same slot. Rebump to the next available version.

- VERSION + package.json → 0.41.28.0
- CHANGELOG.md: my entry header + in-entry refs 0.41.27.0 → 0.41.28.0
- TODOS.md: my #1570 follow-up section header + body refs bumped

* test: pin gateway in put-page-provenance + embedding-dim-check (CI shard fix)

Both files failed on CI shards 1 and 8 under the cross-file gateway-state
leak class (CLAUDE.md "Test-isolation lint and helpers"). The v0.41.28.0
merge reshuffled the weight-based shard bin-packing, landing a
gateway-mutating sibling ahead of these two victims in the same `bun test`
process.

Mechanism:
- put-page-provenance: put_page embeds via the gateway. A sibling left
  the gateway configured with OpenAI + the CI placeholder `sk-test`
  (captured at configureGateway time, survives the withEnv restore as
  cached gateway state). put_page's embed then fired against live OpenAI
  and 401'd. The bunfig legacy-embedding preload's beforeEach only
  re-applies legacy when the gateway was RESET — it does NOT correct a
  sibling that configured a different LIVE config.
- embedding-dim-check: initSchema builds the content_chunks vector column
  at the gateway's configured dim. A sibling leaking ZE/1280 made the
  column 1280-d, so `expect(dims).toBe(1536)` failed.

Fix (victim-side pinning, the escape hatch the preload documents):
- Both: configure the gateway explicitly in beforeAll BEFORE initSchema
  (OpenAI/1536), resetGateway() in afterAll so neither leaks onward.
- put-page-provenance also stubs the embed transport via
  __setEmbedTransportForTests so embed is deterministic and offline; a
  dummy OPENAI_API_KEY is supplied in the gateway env because
  instantiateEmbedding builds the OpenAI client (key check) BEFORE the
  stubbed transport is reached — the stub then intercepts the actual
  call so the key never leaves the process.

Verified: CI shards 1 (1337 pass) + 8 (905 pass) green with
OPENAI_API_KEY unset, plus adversarial sibling orderings (gateway.test /
doctor-ze-checks preceding). Typecheck + check-test-isolation clean.

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-05-28 19:04:48 -07:00
committed by GitHub
co-authored by Claude Opus 4.7
parent cb1b5f91f7
commit ffac8ce0f4
20 changed files with 1060 additions and 14 deletions
+148
View File
@@ -2,6 +2,154 @@
All notable changes to GBrain will be documented in this file.
## [0.41.28.0] - 2026-05-27
**Your `gbrain dream` cycle stops losing rows when the database connection
blips, and the silent `'No database connection'` errors after `gbrain
capture` go away.**
If you run `gbrain dream` against a Supabase brain on the Supavisor pooler,
you might have seen ~150 link rows quietly disappear every cycle, with
log lines like:
```
[extract.links_fs] connection blip, retrying (attempt 1/3): No database connection: connect() has not been called
[extract.links_fs] connection blip, retrying (attempt 2/3): No database connection: connect() has not been called
[extract.links_fs] connection blip, retrying (attempt 3/3): No database connection: connect() has not been called
batch error (100 link rows lost): No database connection: connect() has not been called
```
The retry layer was correctly noticing the problem and waiting. But the
underlying database connection wrapper had been nulled out by some other
code path in the same process, and the retry was hammering against a dead
reference. v0.41.28.0 makes the retry layer rebuild the connection between
attempts via a new opt-in `reconnect` callback on `withRetry`. The engine
self-heals; rows land. (Closes #1570.)
The other symptom was that `gbrain capture` would print a trailing
`'No database connection'` line on stderr from a background facts:absorb
worker firing AFTER the CLI's `engine.disconnect()` finally block ran.
The fact subsystem queues post-page-write work fire-and-forget; that work
sometimes outlived the CLI process's connection lifetime. v0.41.28.0 adds
a new `FactsQueue.drainPending({timeout: 1000})` method, semantically
distinct from `shutdown()` (which would abort in-flight) — drain lets
in-flight finish. The CLI op-dispatch now awaits the drain before
`engine.disconnect()`, capped at 1s so commands that don't enqueue facts
pay only a fast no-op check.
**Honest scope.** This is the tactical symptom fix. The deeper question
— which specific code path nulls the database singleton mid-cycle — is
still open. v0.41.28.0 also ships diagnostic instrumentation:
every call to `db.disconnect()` and `PostgresEngine.disconnect()` writes
a JSONL audit row to `~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl`
recording the engine kind, connection style, caller stack trace, and
command. The doctor's existing `batch_retry_health` check surfaces the
24-hour count plus the most-recent caller frame, so after your next
dream cycle you can run `gbrain doctor --json` and see exactly which
code path is calling disconnect mid-process. v0.41.28+ will fix that
specific ownership boundary based on the production data.
**What to do after upgrading:**
```bash
gbrain --version # 0.41.25.0
gbrain upgrade
gbrain dream --workers 4 2>&1 | tee /tmp/dream.log
grep -c "batch error" /tmp/dream.log # expect 0
grep -c "No database connection" /tmp/dream.log # expect 0
gbrain doctor --json | jq '.checks[] | select(.id=="batch_retry_health")'
```
The `batch_retry_health` output will include a `Disconnect-call audit`
sentence naming the most-recent mid-process disconnect caller. If the
field shows zero calls, the symptom fix alone solved your problem. If it
shows calls, please file an issue with that data so v0.41.28+ can target
the right ownership boundary.
### Itemized changes
**Core fix — retry self-heals on null singleton:**
- `src/core/retry.ts``WithRetryOpts` gains `reconnect?: () => Promise<void>`.
Awaited in the catch branch AFTER `isRetryableConnError` classification but
BEFORE the inter-attempt sleep. `onRetry` callbacks are now awaited too
(back-compat-safe: existing sync arrows work identically; async callbacks
now correctly delay the sleep). Fail-loud posture (per codex outside-voice
finding 3): a reconnect throw propagates AS the new error, replacing the
symptomatic "No database connection" so operators see the real cause.
- `src/core/postgres-engine.ts:batchRetry` — Injects `reconnect: () => this.reconnect()`
into its `withRetry` call. `PostgresEngine.reconnect()` was already
race-safe via `_reconnecting` guard and handles both module and instance
pools.
**Facts queue post-CLI drain:**
- `src/core/facts/queue.ts` — New `FactsQueue.drainPending({timeout?: number})`
method, returns `{drained, unfinished}`. Distinct from `shutdown()`: drain
does NOT abort in-flight (per codex finding 9: shutdown's
`internalAbort.abort()` would abort the very facts:absorb worker that's
trying to log its post-completion event, preserving the bug class we're
fixing). Default timeout 1000ms; bounded so commands that don't enqueue
facts pay no observable cost.
- `src/cli.ts` op-dispatch finally — Awaits
`getFactsQueue().drainPending({timeout: 1000})` BEFORE
`engine.disconnect()`. Lazy-import keeps the facts-queue module off the
hot path for ops that never touch it.
**Diagnostic instrumentation (find the offender for v0.41.28+):**
- `src/core/audit/db-disconnect-audit.ts` (NEW, ~150 LOC) — Built on the
existing `audit-writer.ts` cathedral, mirrors `batch-retry-audit.ts`
shape. Schema: `{ts, engine_kind, connection_style, caller_stack,
command, pid}`. Stack trace captured via `new Error().stack`, truncated
to ~20 frames. ISO-week file rotation. Best-effort writes (stderr-warn
on failure, never throws).
- `src/core/db.ts:disconnect` — Logs an audit row before `sql.end()`. Lazy
import so cold paths don't pay the cost.
- `src/core/postgres-engine.ts:disconnect` — Logs an audit row BEFORE the
early-return branches so even no-op disconnects (engine that was never
connected) are recorded — that case may itself be a caller-side bug.
- `src/commands/doctor.ts:checkBatchRetryHealth` — Extended (per codex
finding 11: extend the existing check, don't add a new one) to surface
24h disconnect-call count and most-recent caller frame in the existing
message. Operators reading doctor output see all connection-incident
signal in one place.
**Tests (focused, per codex finding 12):**
- `test/core/retry-reconnect.test.ts` (NEW, 5 cases) — reconnect-callback
contract: ordering (classification → onRetry → reconnect → sleep),
back-compat (no reconnect opt = v0.41.18.0 behavior), fail-loud
propagation, signal-abort short-circuit, awaited onRetry timing.
- `test/facts-queue-drain-pending.test.ts` (NEW, 4 cases) — drainPending
semantic distinct from shutdown: empty fast-path, in-flight settled
without abort, unfinished count on timeout, default timeout = 1000ms.
- `test/db-disconnect-audit.test.ts` (NEW, 6 cases) — round-trip, stack
truncation, sort order, empty-dir nulls, stable feature name, EROFS
best-effort.
- `test/e2e/db-singleton-shared-recovery.test.ts` (NEW, 3 DB-gated cases) —
pins the production failure modes: shared-singleton survival via retry
reconnect, diagnostic audit fires on disconnect, instance-pool disconnect
doesn't touch module singleton.
### Honest claims (per codex outside-voice review)
- The "postgres.js auto-reconnects" claim from earlier plan iterations is
acknowledged as overbroad: postgres.js's internal auto-reconnect handles
network drops on a still-live pool object. It does NOT help when our
module-singleton reference has been explicitly nulled — that's the bug
class this release patches at the retry layer + investigates with the
audit instrumentation.
- The architectural refactor (remove module-singleton nullability, rename
`disconnect → shutdown`) considered in earlier plan iterations is
deferred to v0.42+ pending the diagnostic data this release ships.
Codex's outside-voice review of the architectural plan found 15
substantive problems — most importantly that the refactor was designed
for a root cause we hadn't actually identified.
**Plan + 12 decisions + 15 codex findings absorbed at
`~/.claude/plans/system-instruction-you-are-working-cuddly-panda.md`.
Closes #1570.**
## [0.41.27.0] - 2026-05-27
**`gbrain doctor` stops crying wolf about sources that have no new commits.**
+4 -2
View File
@@ -158,13 +158,15 @@ strict behavior when unset.
- `src/core/audit-skill-brain-first.ts` (v0.37.1.0) — snapshot+diff JSONL audit at `~/.gbrain/audit/skill-brain-first-YYYY-Www.jsonl` (ISO-week rotated, honors `GBRAIN_AUDIT_DIR` via the shared `resolveAuditDir()`). `recordBrainFirstRun(results)` reads the previous snapshot at `~/.gbrain/audit/skill-brain-first-snapshot.json`, diffs against the current results, writes transition events (`detected | resolved | fixed`) one line per change, then atomically overwrites the snapshot via `.tmp + rename`. **Transition-only writes** — a stable brain produces 0 audit lines per doctor run, so `tail -20` shows real signal instead of noise. `readRecentBrainFirstEvents(days)` is the readback path used by the future `skill_brain_first_trend` doctor check (filed in TODOS.md). Snapshot file is last-writer-wins under concurrent doctor runs; subsequent runs reconcile.
- `src/core/dry-fix.ts``gbrain doctor --fix` engine. `autoFixDryViolations(fixes, {dryRun})` rewrites inlined rules to `> **Convention:** see [path](path).` callouts via three shape-aware expanders (bullet / blockquote / paragraph). Five guards: working-tree-dirty (`getWorkingTreeStatus()` returns 3-state `'clean' | 'dirty' | 'not_a_repo'`), no-git-backup, inside-code-fence, already-delegated (40-line proximity, consistent with detector), ambiguous-multi-match, block-is-callout. `execFileSync` array args (no shell — no injection surface). EOF newline preserved. **v0.37.1.0:** safety primitives extracted to `src/core/skill-fix-gates.ts` (back-compat re-exports preserved). New `MISSING_RULE_PATTERNS` INSERT pattern type lives alongside the existing REPLACE patterns — same auto-fix entry point, same git-safety gates, but instead of rewriting an existing block, INSERT patterns place a canonical callout at a target offset (today: `after-h1-paragraph` only; designed to extend). The first INSERT pattern is `brain_first`, which auto-inserts `> **Convention:** see [conventions/brain-first.md](../conventions/brain-first.md) for the lookup chain (search → query → get_page → external).` on any flagged SKILL.md whose analyzer verdict is `missing_brain_first`. Idempotent — re-runs detect the existing callout and skip.
- `src/core/backoff.ts` — Adaptive load-aware throttling: CPU/memory checks, exponential backoff, active hours multiplier
- `src/core/retry.ts` (v0.41.19.0) — canonical retry primitive for transient connection errors. Exports `withRetry<T>(fn, opts)` execution wrapper + `BULK_RETRY_OPTS` constant (`{maxRetries:3, delayMs:1000, delayMaxMs:10000, jitter:'decorrelated'}`) tuned for Supabase Supavisor's 5-10s circuit-breaker recovery window + `BATCH_AUDIT_SITES` typed const (closed enum of every audit-emission site) + `resolveBulkRetryOpts(env)` (reads `GBRAIN_BULK_MAX_RETRIES` / `GBRAIN_BULK_RETRY_BASE_MS` / `GBRAIN_BULK_RETRY_MAX_MS` with `>=0` validation, throws on bad input with paste-ready hint) + `abortableSleep(ms, signal?)` (AbortSignal-aware setTimeout) + `RetryAbortError` (tagged error for clean shutdown) + `computeNextDelay()` (pure-fn for 3 jitter modes: `'none'`, `'full'`, `'decorrelated'`). The execution wrapper is consumed by `postgres-engine.ts` + `pglite-engine.ts` batch primitives (`addLinksBatch` / `addTimelineEntriesBatch` / `upsertChunks`) so every caller — current AND future — inherits retry as part of the data-primitive's contract. CI guard `scripts/check-no-double-retry.sh` fails the build on `withRetry(...engine.batch...)` patterns (prevents 3×3=9 retry amplification). CI guard `scripts/check-batch-audit-site.sh` validates every string-literal `auditSite: '...'` against the closed `BATCH_AUDIT_SITES` enum. **Closes the v0.41.17 production incident** where ~3,000 rows were silently lost per dream cycle on a 16K-page brain because the prior single-500ms-retry shape couldn't survive Supavisor's 5-10s circuit-breaker recovery. Decorrelated jitter (AWS-style: `uniform(base, prevDelay*3)` capped at `delayMaxMs`) replaces naive `'full'` jitter which allowed near-zero retries that re-hit the still-recovering breaker. Pinned by `test/core/retry.test.ts` (37 cases) + `test/core/retry-stress.slow.test.ts` (5 cases simulating 100 batches × 30% blip rate, asserts zero row loss).
- `src/core/retry.ts` (v0.41.19.0, extended v0.41.27.0) — canonical retry primitive for transient connection errors. Exports `withRetry<T>(fn, opts)` execution wrapper + `BULK_RETRY_OPTS` constant (`{maxRetries:3, delayMs:1000, delayMaxMs:10000, jitter:'decorrelated'}`) tuned for Supabase Supavisor's 5-10s circuit-breaker recovery window + `BATCH_AUDIT_SITES` typed const (closed enum of every audit-emission site) + `resolveBulkRetryOpts(env)` (reads `GBRAIN_BULK_MAX_RETRIES` / `GBRAIN_BULK_RETRY_BASE_MS` / `GBRAIN_BULK_RETRY_MAX_MS` with `>=0` validation, throws on bad input with paste-ready hint) + `abortableSleep(ms, signal?)` (AbortSignal-aware setTimeout) + `RetryAbortError` (tagged error for clean shutdown) + `computeNextDelay()` (pure-fn for 3 jitter modes: `'none'`, `'full'`, `'decorrelated'`). The execution wrapper is consumed by `postgres-engine.ts` + `pglite-engine.ts` batch primitives (`addLinksBatch` / `addTimelineEntriesBatch` / `upsertChunks`) so every caller — current AND future — inherits retry as part of the data-primitive's contract. CI guard `scripts/check-no-double-retry.sh` fails the build on `withRetry(...engine.batch...)` patterns (prevents 3×3=9 retry amplification). CI guard `scripts/check-batch-audit-site.sh` validates every string-literal `auditSite: '...'` against the closed `BATCH_AUDIT_SITES` enum. **Closes the v0.41.17 production incident** where ~3,000 rows were silently lost per dream cycle on a 16K-page brain because the prior single-500ms-retry shape couldn't survive Supavisor's 5-10s circuit-breaker recovery. Decorrelated jitter (AWS-style: `uniform(base, prevDelay*3)` capped at `delayMaxMs`) replaces naive `'full'` jitter which allowed near-zero retries that re-hit the still-recovering breaker. Pinned by `test/core/retry.test.ts` (37 cases) + `test/core/retry-stress.slow.test.ts` (5 cases simulating 100 batches × 30% blip rate, asserts zero row loss). **v0.41.27.0 (#1570):** `WithRetryOpts` gains optional `reconnect?: () => Promise<void>` callback awaited in the catch branch AFTER `isRetryableConnError` classification but BEFORE the inter-attempt sleep — lets engine-level callers rebuild a dead pool/singleton between attempts. `PostgresEngine.batchRetry` injects `() => this.reconnect()` so the v0.22.1 race-safe `_reconnecting` guard kicks in transparently. Fail-loud posture (per codex outside-voice finding 3): a reconnect throw PROPAGATES as the new error, replacing the symptomatic "No database connection" so operators see the real cause (auth failure, EHOSTUNREACH). `onRetry` callbacks now also awaited (back-compat-safe: existing sync arrows work identically; async callbacks now correctly delay the sleep). Pinned by `test/core/retry-reconnect.test.ts` (5 cases: ordering, back-compat, fail-loud propagation, signal-abort short-circuit, awaited onRetry timing) + `test/e2e/db-singleton-shared-recovery.test.ts` (3 DB-gated cases pinning shared-singleton survival via retry reconnect + diagnostic audit fires on disconnect + instance-pool disconnect doesn't touch module singleton).
- `src/core/audit/batch-retry-audit.ts` (v0.41.19.0, extended v0.41.26.1) — JSONL audit primitive for batch-retry events. Built on `audit-writer.ts` cathedral. Schema: `{ts, site, batch_size, attempt, outcome: 'success' | 'exhausted', delay_ms, error_message_summary, error_code?}`. Privacy posture: NEVER logs slugs / page IDs / content (mirrors `shell-audit.ts` from v0.20+). `logBatchRetry` fires per successful retry recovery; `logBatchExhausted` fires when retries exhaust and rows are lost. `readRecentBatchRetryEvents(hours=24)` returns `{events, corrupted_lines, files_scanned, files_unreadable}` — corruption + permission errors surface to doctor, not silently swallowed (codex H-9). `pruneOldBatchRetryAuditFiles(daysToKeep=30)` deletes old files (codex H-8 — implements the v0.41 plan's "30-day pruning convention" for real). Called from `gbrain dream --phase purge`. File: `~/.gbrain/audit/batch-retry-YYYY-Www.jsonl` (honors `GBRAIN_AUDIT_DIR`). **v0.41.26.1 privacy backfill:** `summarizeError` now routes error messages through the shared `redactConnectionInfo` helper from `src/core/audit/redact-connection-info.ts` BEFORE truncation, so DSNs / hostnames / credentials / IPv4 octets can't leak from a Postgres connection-failure error into operator-shared JSONL dumps. Same risk class as lock-renewal-audit; closed in the same wave. Pinned by `test/audit/batch-retry-audit.test.ts` (12 cases) + new `test/audit/batch-retry-redaction.test.ts` (3 privacy regressions).
- `src/core/audit/lock-renewal-audit.ts` (v0.41.26.1) — JSONL audit primitive for per-job lock-renewal faults. Sibling of `batch-retry-audit.ts`; built on `audit-writer.ts`. Four outcomes: `failure` (single renewLock throw, counter incremented), `success_after_failure` (recovery; emits the recovery count), `gave_up` (time-based deadline exceeded; abort fired), `executeJob_rejected` (the SECOND unhandledRejection vector from D7 — the stored `executeJob(...).finally(...)` promise itself rejected, e.g. failJob threw during the same DB outage). Schema: `{ts, job_id, job_name, attempt?, outcome, error_message_summary?, error_code?}`. Privacy: NEVER logs `lock_token` or `job.data`. Error summaries route through `redactConnectionInfo` BEFORE truncation. Defense-in-depth: every audit call inside the lock-renewal tick's catch block is wrapped in its own inner try/catch so a misbehaving audit-writer can't re-introduce the unhandledRejection bug class via a new surface. `readRecentLockRenewalEvents(hours=24)` walks current + previous ISO week with corrupted-line tolerance. `pruneOldLockRenewalAuditFiles(daysToKeep=30)` is ready for future dream-cycle purge wiring (filed as TODO-LR-3). File: `~/.gbrain/audit/lock-renewal-YYYY-Www.jsonl`. Operator UX is `tail -F` until the doctor check lands (TODO-LR-2). Pinned by `test/audit/lock-renewal-audit.test.ts` (11 cases).
- `src/core/audit/redact-connection-info.ts` (v0.41.26.1) — Shared pure helper. `redactConnectionInfo(text: string): string` strips Postgres connection info before any audit JSONL write: `postgres://`/`postgresql://` URLs, `host=foo`, `user=foo`, `password=foo`, `pwd=foo`, IPv4 octets. Each match becomes `<REDACTED:kind>`. Negative-lookbehind/lookahead `[\w.@-]` on the IPv4 pattern defeats version-string false positives like `v3.1.4.0` or `tree-sitter@0.26.3.1` while still matching real IPs in PG errors (`(192.168.1.42)`). Order-sensitive pattern set: URL forms first so substrings inside URLs don't get double-redacted. Idempotent (running twice produces same output), pure (no I/O), hot-path-safe (regex compiled at module load). Wired into BOTH `lock-renewal-audit.ts` (new) AND `batch-retry-audit.ts` (privacy backfill — same risk class). Risk model documented in module header: Postgres errors during connection failures embed DSNs / credentials / hostnames into error messages; operators routinely paste audit dumps into GitHub issues / Slack to debug; the audit channel must be safe to share by construction. Known limitations: bare-quoted hostnames (`at "db.example.com"`) and bare-quoted usernames (`for user "postgres.foo"`) are NOT caught — the IP in those PG error shapes is the highest-value leak and IS caught. Quoted-string patterns filed as TODO-LR-5. Pinned by `test/audit/redact-connection-info.test.ts` (15 cases covering all 5 patterns + real Supabase fixture + ENOTFOUND fixture + version-string false-positive defense).
- `src/core/minions/lock-renewal-tick.ts` (v0.41.26.1) — Pure extracted function from `MinionWorker.launchJob`'s setInterval body. The structural fix that closes the v0.41.22.1 production unhandledRejection crash class. Exports `runLockRenewalTick(deps, state) → Promise<TickResult>` plus `resolveLockRenewalKnobs(env, lockDuration) → LockRenewalKnobs`. Three env knobs all parsed as positive integers with stderr-warn-once-per-process on bad input + default fallback: `GBRAIN_LOCK_RENEWAL_MAX_FAILURES` (default 3, audit-labeling only — abort triggering uses time-based instead), `GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS` (default `lockDuration/3`, bounds the hung-renewLock vector via `Promise.race`), `GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS` (default `lockDuration/6`, ensures we release before stall-detector reclaim). The tick checks `state.cancelled()` at three guard points (entry, after await resolves, after await throws) so a long renewLock that resolves after the job ended bails cleanly. Result type is a tagged union `{kind: 'ok' | 'cancelled' | 'lock_lost' | 'should_abort', reason?}` that the worker switches on; all `abort.abort()` / `clearInterval` side effects live in the worker's closure scope, not the pure function. Audit calls wrapped in inner try/catch (codex C4 defense-in-depth) so the audit module can never re-introduce the bug class through a new surface. Pinned by `test/worker-lock-renewal.test.ts` (18 hermetic state-machine cases — no fs, no PGLite, no real setInterval).
- `scripts/check-worker-lock-renewal-shape.sh` (v0.41.26.1, D4) — CI guard wired into `bun run verify`. Two invariants on `src/core/minions/worker.ts`: (1) the v0.41.22.1 bug pattern `lockTimer = setInterval(async ...)` must NOT appear (narrowed via `lockTimer =` prefix so unrelated `setInterval(async)` calls elsewhere in the file — like the stall detector at line ~269, codex C13 territory — don't false-fire), (2) `runLockRenewalTick` must remain referenced so the pure-function test seam survives refactors. Bug-pattern-specific by design (codex C12) — a future refactor to `setTimeout`-recursion or `AbortController`-based scheduling passes as long as the bug pattern stays absent. Uses POSIX ERE + `[[:space:]]` for BSD-grep portability. Honors `GBRAIN_LOCK_RENEWAL_SHAPE_TARGET` env override for fixture-based meta-tests. Pinned by `test/scripts/check-worker-lock-renewal-shape.test.ts` (5 cases including the C13 false-positive defense for the stall-detector pattern elsewhere in the file).
- `src/commands/doctor.ts:checkBatchRetryHealth` (v0.41.19.0) — `batch_retry_health` check surfacing Supavisor circuit-breaker incidents. Wired into both `runDoctor` (local) and `doctorReportRemote` (thin-client). Reads last **24h** (not 7d — codex H-9: avoid permanent noise from one historical blip). States: `ok` (zero exhausted in 24h OR <3 from a single site), `warn` (>=3 same-site OR >=5 cross-site), `fail` (>=20 sustained breaker). Surfaces bad `GBRAIN_BULK_*` env at doctor startup (codex M-10) instead of first-retry. Corrupt-JSONL tolerant. Paste-ready fix hints in every warn/fail message. Pinned by `test/doctor-batch-retry.test.ts` (10 cases).
- `src/commands/doctor.ts:checkBatchRetryHealth` (v0.41.19.0, extended v0.41.27.0) — `batch_retry_health` check surfacing Supavisor circuit-breaker incidents. Wired into both `runDoctor` (local) and `doctorReportRemote` (thin-client). Reads last **24h** (not 7d — codex H-9: avoid permanent noise from one historical blip). States: `ok` (zero exhausted in 24h OR <3 from a single site), `warn` (>=3 same-site OR >=5 cross-site), `fail` (>=20 sustained breaker). Surfaces bad `GBRAIN_BULK_*` env at doctor startup (codex M-10) instead of first-retry. Corrupt-JSONL tolerant. Paste-ready fix hints in every warn/fail message. Pinned by `test/doctor-batch-retry.test.ts` (10 cases). **v0.41.27.0 (#1570, codex finding 11):** extended (NOT a new check) to read `readRecentDbDisconnects(24)` and append `Disconnect-call audit: N call(s) in 24h (most recent caller: <frame>).` to ALL three message paths (ok / warn / fail). The single-place surface keeps connection-incident signal greppable from one `gbrain doctor --json` call so operators correlating the v0.41.27 retry-reconnect symptom fix with the offending caller don't have to grep two audit files. Module-import wrapped in try/catch so older brains without the v0.41.27 audit file degrade silently.
- `src/core/audit/db-disconnect-audit.ts` (v0.41.27.0, #1570, NEW) — JSONL audit for every call to `db.disconnect()` and `PostgresEngine.disconnect()`. Built on `audit-writer.ts` cathedral (no parallel subsystem; codex finding 11). Schema: `{ts, engine_kind: 'postgres'|'pglite'|'unknown', connection_style: 'module'|'instance'|'unknown', caller_stack, command, pid}`. `caller_stack` captured via `new Error().stack` truncated to ~20 frames so operators identify the offending caller without inflating JSONL forever. Privacy: stack frames carry file paths but NO SQL content / row data / user strings — matches the v0.20 shell-audit posture. File: `~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl` (honors `GBRAIN_AUDIT_DIR`). `readRecentDbDisconnects(hours=24)` walks current + previous ISO week and returns `{count, most_recent_caller, files_scanned}`. **The instrument half of the "instrument first, fix later" pivot** — v0.41.27.0 ships the symptom fix (retry reconnect callback + facts:absorb drain) AND this diagnostic surface; v0.41.28+ patches the specific ownership boundary once production data tells us which code path is calling `disconnect()` mid-process. Wired into `src/core/db.ts:disconnect` and `src/core/postgres-engine.ts:disconnect` (logs BEFORE the early-return branches so even no-op disconnects on never-connected engines are recorded — that case may itself be a caller-side bug). Pinned by `test/db-disconnect-audit.test.ts` (6 cases: round-trip, stack truncation, sort order, empty-dir nulls, stable feature name, EROFS best-effort).
- `src/core/facts/queue.ts:FactsQueue.drainPending` (v0.41.27.0, #1570) — new method `drainPending({timeout?: number}): Promise<{drained, unfinished}>`. Semantically distinct from `shutdown()` (which calls `this.internalAbort.abort()` and would abort the very facts:absorb worker trying to log its post-completion event — preserving the bug class we're fixing per codex finding 9). Drain lets in-flight finish; only the wait is bounded. Default timeout `1000ms` so commands that don't enqueue facts pay only one fast 0ms check before exit (per codex finding 10). `src/cli.ts` op-dispatch finally block awaits `getFactsQueue().drainPending({timeout: 1000})` BEFORE `engine.disconnect()`. Lazy-import keeps the facts-queue module off the hot path for ops that never touch it. Closes the trailing `'No database connection'` line that fired on stderr after `gbrain capture` because the post-page-write facts:absorb queued work outlived the CLI process. Pinned by `test/facts-queue-drain-pending.test.ts` (4 cases: empty fast-path, in-flight settled without abort, unfinished count on timeout, default timeout = 1000ms).
- `scripts/check-no-double-retry.sh` + `scripts/check-batch-audit-site.sh` (v0.41.19.0) — CI lint guards wired into `bun run verify`. The former greps src/ for `withRetry(...engine.{addLinksBatch|addTimelineEntriesBatch|upsertChunks})` patterns and fails the build on hit (Eng-D6 migration ordering hazard — prevents 3×3=9 retry amplification on incomplete reverts). The latter extracts every string-literal `auditSite: '...'` from src/ and validates each appears in the `BATCH_AUDIT_SITES` const in `src/core/retry.ts` (codex H-7 typo guard — prevents fragmented doctor output).
- `src/core/fail-improve.ts` — Deterministic-first, LLM-fallback loop with JSONL failure logging and auto-test generation
- `src/core/transcription.ts` — Audio transcription: Groq Whisper (default), OpenAI fallback, ffmpeg segmentation for >25MB
+16
View File
@@ -314,6 +314,22 @@ Bad values surface at `gbrain doctor` startup with a paste-ready fix
retry wrap is engine-level, but PGLite has no pooler so retries never
fire in practice.
**Dream cycle losing ~150 link rows per run with `'No database
connection: connect() has not been called'` errors in the log?** v0.41.27.0
makes the retry layer self-heal on a nulled-out database singleton. A
new `reconnect` callback on `withRetry` rebuilds the connection between
attempts; `PostgresEngine.batchRetry` injects `() => this.reconnect()`
so engine-level batch writes survive a mid-cycle disconnect by something
else in the same process. Same release: `gbrain capture` no longer trails
a `'No database connection'` stderr line from a background facts:absorb
worker firing after CLI exit — the op-dispatch finally block awaits
`getFactsQueue().drainPending({timeout: 1000})` before
`engine.disconnect()`. To find which code path is still calling
disconnect mid-process, run `gbrain doctor --json | jq '.checks[] |
select(.id=="batch_retry_health")'`; the extended check now surfaces
24h disconnect-call count and the most-recent caller frame from a new
`~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl` audit. (Closes #1570.)
**`gbrain brainstorm` returning `judge_failed: true` with 0 scored
ideas?** v0.41.21.0 closes the two bugs that caused it. The judge
hard-coded a 4K-token output cap; for any run past ~40 ideas the call
+14 -1
View File
@@ -1,5 +1,18 @@
# TODOS
## v0.41.28.0 #1570 instrument-then-fix follow-ups (v0.41.28+ / v0.42+)
Filed from the v0.41.28.0 plan-eng-review after the codex outside-voice
review caught that the original architectural-refactor plan was designed
for a root cause we hadn't identified. v0.41.28.0 ships the tactical
symptom fix (retry reconnect) + facts queue drain + diagnostic
instrumentation. These follow-ups depend on the production data the
instrumentation collects.
- [ ] **v0.41.28+: Investigate disconnect-call audit data from production; fix the offending ownership boundary.** v0.41.28.0 ships `src/core/audit/db-disconnect-audit.ts` which records every `db.disconnect()` and `PostgresEngine.disconnect()` call with engine kind, connection style, caller stack, command, and pid. Doctor's `batch_retry_health` check surfaces the 24h count + most-recent caller. After the next user-reported `gbrain dream` cycle with reconnect events, read `~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl` (or the doctor JSON output) and identify the specific code path firing the mid-process disconnect. The fix is then a targeted patch to that ownership boundary (per codex outside-voice finding 4 — "audit/log current callers in dream/facts paths, then change only the offending ownership boundary"). Priority: P1 once data exists; tracked by user feedback on #1570 thread.
- [ ] **v0.42+: Re-evaluate module-singleton removal IF the targeted v0.41.26 fix doesn't close the bug class.** The original v0.41.25 plan proposed removing nullability of `let sql: ReturnType<typeof postgres> | null = null` in `src/core/db.ts:7` and renaming `disconnect → shutdown`. Codex outside-voice review found 15 substantive problems (logical contradiction, wrong cleanup primitive, ~120-site scale estimate fantasy, BrainEngine contract asymmetry, etc.). If the targeted v0.41.26 fix closes #1570 cleanly, this refactor is genuinely unnecessary and can be closed. If new disconnect-class bugs surface in v0.41.28+, this is the design-conversation TODO that re-opens. Architecture conversation point: node-postgres explicitly deprecated the singleton pattern gbrain has — pull this in only when there's evidence we keep paying for it. Priority: P3 (speculative). Plan + findings preserved at `~/.claude/plans/system-instruction-you-are-working-cuddly-panda.md`.
## v0.41.26.1 lock-renewal cathedral follow-ups (v0.42+)
- **TODO-LR-1 (P2): PR #1567 surrogate-pair fix for synthesize.ts.**
@@ -1417,7 +1430,7 @@ contributor traps.
- [ ] **v0.40: magic-byte allowlist for `gbrain capture` binary file detection.** v0.39.3.0 (Phase 3c, CV10) ships a first-8KB NUL-byte scan that catches typical binaries (executables, archives, most image formats). Known gap per CV10-B: a PNG with no NUL byte in its first 8KB slips through. Production-grade detection needs a magic-byte allowlist (PNG/JPEG/GIF/PDF/ZIP signatures). Implement in `src/commands/capture.ts:detectBinaryNullByte` (rename to `detectBinaryInput`) with a small `BINARY_MAGIC_BYTES` table. Reuse the same `assertSourceExists`-style friendly error pattern; reject before UTF-8 decode mangles the bytes. Tests in `test/capture-binary-guard.test.ts` should add cases for the PNG-without-NUL boundary.
- [ ] **v0.40: facts:absorb root-cause investigation.** v0.39.3.0 (Phase 4c, CV13) suppresses the per-capture `[facts:absorb] failed to log gateway_error for inbox/...: No database connection` noise AND prints a first-occurrence stack trace so the v0.40 fix knows where to look. The actual fix is one of: (a) thread the connected engine through the facts pipeline so it doesn't open its own handle; (b) no-op the absorb-log when called from a CLI context where the doctor health check isn't the consumer; (c) make the facts subsystem connection-aware and queue retries. The stack trace from `src/core/facts/absorb-log.ts:writeFactsAbsorbLog`'s first-occurrence info-log is the input.
- [ ] **v0.40: facts:absorb root-cause investigation.** v0.39.3.0 (Phase 4c, CV13) suppresses the per-capture `[facts:absorb] failed to log gateway_error for inbox/...: No database connection` noise AND prints a first-occurrence stack trace so the v0.40 fix knows where to look. The actual fix is one of: (a) thread the connected engine through the facts pipeline so it doesn't open its own handle; (b) no-op the absorb-log when called from a CLI context where the doctor health check isn't the consumer; (c) make the facts subsystem connection-aware and queue retries. The stack trace from `src/core/facts/absorb-log.ts:writeFactsAbsorbLog`'s first-occurrence info-log is the input. **v0.41.25.0 update:** the related #1570 wave shipped a partial fix at the queue level — CLI op-dispatch now awaits `FactsQueue.drainPending({timeout: 1000})` before `engine.disconnect()`, which closes the visible-stderr-line symptom for `gbrain capture`. The deeper "thread engine through pipeline" architectural question (option a above) stays open for v0.40+; the drain fix is a queue-lifetime patch, not a pipeline-rearchitecture.
- [ ] **v0.40: `--source-kind` override flag for `gbrain capture`.** v0.39.3.0 (Phase 3c, CV3) locked source_kind to `'capture-cli'` for capture invocations (the deferred CV3-B alternative). Real use case for the override: Apple Shortcuts / Zapier-style automations that shell out to `gbrain capture` and want their pages labeled `apple-shortcut` or `zapier` in the audit trail. Implementation: add a small flag with an allowlist (similar to migration v81's closed taxonomy: `capture-cli | apple-shortcut | zapier | <skillpack-kind>`); validate at parse time; CV6 remote-spoofing guard still applies (server stamps `mcp:put_page` regardless when `ctx.remote !== false`).
+1 -1
View File
@@ -1 +1 @@
0.41.27.0
0.41.28.0
+20 -2
View File
@@ -300,13 +300,15 @@ strict behavior when unset.
- `src/core/audit-skill-brain-first.ts` (v0.37.1.0) — snapshot+diff JSONL audit at `~/.gbrain/audit/skill-brain-first-YYYY-Www.jsonl` (ISO-week rotated, honors `GBRAIN_AUDIT_DIR` via the shared `resolveAuditDir()`). `recordBrainFirstRun(results)` reads the previous snapshot at `~/.gbrain/audit/skill-brain-first-snapshot.json`, diffs against the current results, writes transition events (`detected | resolved | fixed`) one line per change, then atomically overwrites the snapshot via `.tmp + rename`. **Transition-only writes** — a stable brain produces 0 audit lines per doctor run, so `tail -20` shows real signal instead of noise. `readRecentBrainFirstEvents(days)` is the readback path used by the future `skill_brain_first_trend` doctor check (filed in TODOS.md). Snapshot file is last-writer-wins under concurrent doctor runs; subsequent runs reconcile.
- `src/core/dry-fix.ts` — `gbrain doctor --fix` engine. `autoFixDryViolations(fixes, {dryRun})` rewrites inlined rules to `> **Convention:** see [path](path).` callouts via three shape-aware expanders (bullet / blockquote / paragraph). Five guards: working-tree-dirty (`getWorkingTreeStatus()` returns 3-state `'clean' | 'dirty' | 'not_a_repo'`), no-git-backup, inside-code-fence, already-delegated (40-line proximity, consistent with detector), ambiguous-multi-match, block-is-callout. `execFileSync` array args (no shell — no injection surface). EOF newline preserved. **v0.37.1.0:** safety primitives extracted to `src/core/skill-fix-gates.ts` (back-compat re-exports preserved). New `MISSING_RULE_PATTERNS` INSERT pattern type lives alongside the existing REPLACE patterns — same auto-fix entry point, same git-safety gates, but instead of rewriting an existing block, INSERT patterns place a canonical callout at a target offset (today: `after-h1-paragraph` only; designed to extend). The first INSERT pattern is `brain_first`, which auto-inserts `> **Convention:** see [conventions/brain-first.md](../conventions/brain-first.md) for the lookup chain (search → query → get_page → external).` on any flagged SKILL.md whose analyzer verdict is `missing_brain_first`. Idempotent — re-runs detect the existing callout and skip.
- `src/core/backoff.ts` — Adaptive load-aware throttling: CPU/memory checks, exponential backoff, active hours multiplier
- `src/core/retry.ts` (v0.41.19.0) — canonical retry primitive for transient connection errors. Exports `withRetry<T>(fn, opts)` execution wrapper + `BULK_RETRY_OPTS` constant (`{maxRetries:3, delayMs:1000, delayMaxMs:10000, jitter:'decorrelated'}`) tuned for Supabase Supavisor's 5-10s circuit-breaker recovery window + `BATCH_AUDIT_SITES` typed const (closed enum of every audit-emission site) + `resolveBulkRetryOpts(env)` (reads `GBRAIN_BULK_MAX_RETRIES` / `GBRAIN_BULK_RETRY_BASE_MS` / `GBRAIN_BULK_RETRY_MAX_MS` with `>=0` validation, throws on bad input with paste-ready hint) + `abortableSleep(ms, signal?)` (AbortSignal-aware setTimeout) + `RetryAbortError` (tagged error for clean shutdown) + `computeNextDelay()` (pure-fn for 3 jitter modes: `'none'`, `'full'`, `'decorrelated'`). The execution wrapper is consumed by `postgres-engine.ts` + `pglite-engine.ts` batch primitives (`addLinksBatch` / `addTimelineEntriesBatch` / `upsertChunks`) so every caller — current AND future — inherits retry as part of the data-primitive's contract. CI guard `scripts/check-no-double-retry.sh` fails the build on `withRetry(...engine.batch...)` patterns (prevents 3×3=9 retry amplification). CI guard `scripts/check-batch-audit-site.sh` validates every string-literal `auditSite: '...'` against the closed `BATCH_AUDIT_SITES` enum. **Closes the v0.41.17 production incident** where ~3,000 rows were silently lost per dream cycle on a 16K-page brain because the prior single-500ms-retry shape couldn't survive Supavisor's 5-10s circuit-breaker recovery. Decorrelated jitter (AWS-style: `uniform(base, prevDelay*3)` capped at `delayMaxMs`) replaces naive `'full'` jitter which allowed near-zero retries that re-hit the still-recovering breaker. Pinned by `test/core/retry.test.ts` (37 cases) + `test/core/retry-stress.slow.test.ts` (5 cases simulating 100 batches × 30% blip rate, asserts zero row loss).
- `src/core/retry.ts` (v0.41.19.0, extended v0.41.27.0) — canonical retry primitive for transient connection errors. Exports `withRetry<T>(fn, opts)` execution wrapper + `BULK_RETRY_OPTS` constant (`{maxRetries:3, delayMs:1000, delayMaxMs:10000, jitter:'decorrelated'}`) tuned for Supabase Supavisor's 5-10s circuit-breaker recovery window + `BATCH_AUDIT_SITES` typed const (closed enum of every audit-emission site) + `resolveBulkRetryOpts(env)` (reads `GBRAIN_BULK_MAX_RETRIES` / `GBRAIN_BULK_RETRY_BASE_MS` / `GBRAIN_BULK_RETRY_MAX_MS` with `>=0` validation, throws on bad input with paste-ready hint) + `abortableSleep(ms, signal?)` (AbortSignal-aware setTimeout) + `RetryAbortError` (tagged error for clean shutdown) + `computeNextDelay()` (pure-fn for 3 jitter modes: `'none'`, `'full'`, `'decorrelated'`). The execution wrapper is consumed by `postgres-engine.ts` + `pglite-engine.ts` batch primitives (`addLinksBatch` / `addTimelineEntriesBatch` / `upsertChunks`) so every caller — current AND future — inherits retry as part of the data-primitive's contract. CI guard `scripts/check-no-double-retry.sh` fails the build on `withRetry(...engine.batch...)` patterns (prevents 3×3=9 retry amplification). CI guard `scripts/check-batch-audit-site.sh` validates every string-literal `auditSite: '...'` against the closed `BATCH_AUDIT_SITES` enum. **Closes the v0.41.17 production incident** where ~3,000 rows were silently lost per dream cycle on a 16K-page brain because the prior single-500ms-retry shape couldn't survive Supavisor's 5-10s circuit-breaker recovery. Decorrelated jitter (AWS-style: `uniform(base, prevDelay*3)` capped at `delayMaxMs`) replaces naive `'full'` jitter which allowed near-zero retries that re-hit the still-recovering breaker. Pinned by `test/core/retry.test.ts` (37 cases) + `test/core/retry-stress.slow.test.ts` (5 cases simulating 100 batches × 30% blip rate, asserts zero row loss). **v0.41.27.0 (#1570):** `WithRetryOpts` gains optional `reconnect?: () => Promise<void>` callback awaited in the catch branch AFTER `isRetryableConnError` classification but BEFORE the inter-attempt sleep — lets engine-level callers rebuild a dead pool/singleton between attempts. `PostgresEngine.batchRetry` injects `() => this.reconnect()` so the v0.22.1 race-safe `_reconnecting` guard kicks in transparently. Fail-loud posture (per codex outside-voice finding 3): a reconnect throw PROPAGATES as the new error, replacing the symptomatic "No database connection" so operators see the real cause (auth failure, EHOSTUNREACH). `onRetry` callbacks now also awaited (back-compat-safe: existing sync arrows work identically; async callbacks now correctly delay the sleep). Pinned by `test/core/retry-reconnect.test.ts` (5 cases: ordering, back-compat, fail-loud propagation, signal-abort short-circuit, awaited onRetry timing) + `test/e2e/db-singleton-shared-recovery.test.ts` (3 DB-gated cases pinning shared-singleton survival via retry reconnect + diagnostic audit fires on disconnect + instance-pool disconnect doesn't touch module singleton).
- `src/core/audit/batch-retry-audit.ts` (v0.41.19.0, extended v0.41.26.1) — JSONL audit primitive for batch-retry events. Built on `audit-writer.ts` cathedral. Schema: `{ts, site, batch_size, attempt, outcome: 'success' | 'exhausted', delay_ms, error_message_summary, error_code?}`. Privacy posture: NEVER logs slugs / page IDs / content (mirrors `shell-audit.ts` from v0.20+). `logBatchRetry` fires per successful retry recovery; `logBatchExhausted` fires when retries exhaust and rows are lost. `readRecentBatchRetryEvents(hours=24)` returns `{events, corrupted_lines, files_scanned, files_unreadable}` — corruption + permission errors surface to doctor, not silently swallowed (codex H-9). `pruneOldBatchRetryAuditFiles(daysToKeep=30)` deletes old files (codex H-8 — implements the v0.41 plan's "30-day pruning convention" for real). Called from `gbrain dream --phase purge`. File: `~/.gbrain/audit/batch-retry-YYYY-Www.jsonl` (honors `GBRAIN_AUDIT_DIR`). **v0.41.26.1 privacy backfill:** `summarizeError` now routes error messages through the shared `redactConnectionInfo` helper from `src/core/audit/redact-connection-info.ts` BEFORE truncation, so DSNs / hostnames / credentials / IPv4 octets can't leak from a Postgres connection-failure error into operator-shared JSONL dumps. Same risk class as lock-renewal-audit; closed in the same wave. Pinned by `test/audit/batch-retry-audit.test.ts` (12 cases) + new `test/audit/batch-retry-redaction.test.ts` (3 privacy regressions).
- `src/core/audit/lock-renewal-audit.ts` (v0.41.26.1) — JSONL audit primitive for per-job lock-renewal faults. Sibling of `batch-retry-audit.ts`; built on `audit-writer.ts`. Four outcomes: `failure` (single renewLock throw, counter incremented), `success_after_failure` (recovery; emits the recovery count), `gave_up` (time-based deadline exceeded; abort fired), `executeJob_rejected` (the SECOND unhandledRejection vector from D7 — the stored `executeJob(...).finally(...)` promise itself rejected, e.g. failJob threw during the same DB outage). Schema: `{ts, job_id, job_name, attempt?, outcome, error_message_summary?, error_code?}`. Privacy: NEVER logs `lock_token` or `job.data`. Error summaries route through `redactConnectionInfo` BEFORE truncation. Defense-in-depth: every audit call inside the lock-renewal tick's catch block is wrapped in its own inner try/catch so a misbehaving audit-writer can't re-introduce the unhandledRejection bug class via a new surface. `readRecentLockRenewalEvents(hours=24)` walks current + previous ISO week with corrupted-line tolerance. `pruneOldLockRenewalAuditFiles(daysToKeep=30)` is ready for future dream-cycle purge wiring (filed as TODO-LR-3). File: `~/.gbrain/audit/lock-renewal-YYYY-Www.jsonl`. Operator UX is `tail -F` until the doctor check lands (TODO-LR-2). Pinned by `test/audit/lock-renewal-audit.test.ts` (11 cases).
- `src/core/audit/redact-connection-info.ts` (v0.41.26.1) — Shared pure helper. `redactConnectionInfo(text: string): string` strips Postgres connection info before any audit JSONL write: `postgres://`/`postgresql://` URLs, `host=foo`, `user=foo`, `password=foo`, `pwd=foo`, IPv4 octets. Each match becomes `<REDACTED:kind>`. Negative-lookbehind/lookahead `[\w.@-]` on the IPv4 pattern defeats version-string false positives like `v3.1.4.0` or `tree-sitter@0.26.3.1` while still matching real IPs in PG errors (`(192.168.1.42)`). Order-sensitive pattern set: URL forms first so substrings inside URLs don't get double-redacted. Idempotent (running twice produces same output), pure (no I/O), hot-path-safe (regex compiled at module load). Wired into BOTH `lock-renewal-audit.ts` (new) AND `batch-retry-audit.ts` (privacy backfill — same risk class). Risk model documented in module header: Postgres errors during connection failures embed DSNs / credentials / hostnames into error messages; operators routinely paste audit dumps into GitHub issues / Slack to debug; the audit channel must be safe to share by construction. Known limitations: bare-quoted hostnames (`at "db.example.com"`) and bare-quoted usernames (`for user "postgres.foo"`) are NOT caught — the IP in those PG error shapes is the highest-value leak and IS caught. Quoted-string patterns filed as TODO-LR-5. Pinned by `test/audit/redact-connection-info.test.ts` (15 cases covering all 5 patterns + real Supabase fixture + ENOTFOUND fixture + version-string false-positive defense).
- `src/core/minions/lock-renewal-tick.ts` (v0.41.26.1) — Pure extracted function from `MinionWorker.launchJob`'s setInterval body. The structural fix that closes the v0.41.22.1 production unhandledRejection crash class. Exports `runLockRenewalTick(deps, state) → Promise<TickResult>` plus `resolveLockRenewalKnobs(env, lockDuration) → LockRenewalKnobs`. Three env knobs all parsed as positive integers with stderr-warn-once-per-process on bad input + default fallback: `GBRAIN_LOCK_RENEWAL_MAX_FAILURES` (default 3, audit-labeling only — abort triggering uses time-based instead), `GBRAIN_LOCK_RENEWAL_CALL_TIMEOUT_MS` (default `lockDuration/3`, bounds the hung-renewLock vector via `Promise.race`), `GBRAIN_LOCK_RENEWAL_SAFETY_MARGIN_MS` (default `lockDuration/6`, ensures we release before stall-detector reclaim). The tick checks `state.cancelled()` at three guard points (entry, after await resolves, after await throws) so a long renewLock that resolves after the job ended bails cleanly. Result type is a tagged union `{kind: 'ok' | 'cancelled' | 'lock_lost' | 'should_abort', reason?}` that the worker switches on; all `abort.abort()` / `clearInterval` side effects live in the worker's closure scope, not the pure function. Audit calls wrapped in inner try/catch (codex C4 defense-in-depth) so the audit module can never re-introduce the bug class through a new surface. Pinned by `test/worker-lock-renewal.test.ts` (18 hermetic state-machine cases — no fs, no PGLite, no real setInterval).
- `scripts/check-worker-lock-renewal-shape.sh` (v0.41.26.1, D4) — CI guard wired into `bun run verify`. Two invariants on `src/core/minions/worker.ts`: (1) the v0.41.22.1 bug pattern `lockTimer = setInterval(async ...)` must NOT appear (narrowed via `lockTimer =` prefix so unrelated `setInterval(async)` calls elsewhere in the file — like the stall detector at line ~269, codex C13 territory — don't false-fire), (2) `runLockRenewalTick` must remain referenced so the pure-function test seam survives refactors. Bug-pattern-specific by design (codex C12) — a future refactor to `setTimeout`-recursion or `AbortController`-based scheduling passes as long as the bug pattern stays absent. Uses POSIX ERE + `[[:space:]]` for BSD-grep portability. Honors `GBRAIN_LOCK_RENEWAL_SHAPE_TARGET` env override for fixture-based meta-tests. Pinned by `test/scripts/check-worker-lock-renewal-shape.test.ts` (5 cases including the C13 false-positive defense for the stall-detector pattern elsewhere in the file).
- `src/commands/doctor.ts:checkBatchRetryHealth` (v0.41.19.0) — `batch_retry_health` check surfacing Supavisor circuit-breaker incidents. Wired into both `runDoctor` (local) and `doctorReportRemote` (thin-client). Reads last **24h** (not 7d — codex H-9: avoid permanent noise from one historical blip). States: `ok` (zero exhausted in 24h OR <3 from a single site), `warn` (>=3 same-site OR >=5 cross-site), `fail` (>=20 sustained breaker). Surfaces bad `GBRAIN_BULK_*` env at doctor startup (codex M-10) instead of first-retry. Corrupt-JSONL tolerant. Paste-ready fix hints in every warn/fail message. Pinned by `test/doctor-batch-retry.test.ts` (10 cases).
- `src/commands/doctor.ts:checkBatchRetryHealth` (v0.41.19.0, extended v0.41.27.0) — `batch_retry_health` check surfacing Supavisor circuit-breaker incidents. Wired into both `runDoctor` (local) and `doctorReportRemote` (thin-client). Reads last **24h** (not 7d — codex H-9: avoid permanent noise from one historical blip). States: `ok` (zero exhausted in 24h OR <3 from a single site), `warn` (>=3 same-site OR >=5 cross-site), `fail` (>=20 sustained breaker). Surfaces bad `GBRAIN_BULK_*` env at doctor startup (codex M-10) instead of first-retry. Corrupt-JSONL tolerant. Paste-ready fix hints in every warn/fail message. Pinned by `test/doctor-batch-retry.test.ts` (10 cases). **v0.41.27.0 (#1570, codex finding 11):** extended (NOT a new check) to read `readRecentDbDisconnects(24)` and append `Disconnect-call audit: N call(s) in 24h (most recent caller: <frame>).` to ALL three message paths (ok / warn / fail). The single-place surface keeps connection-incident signal greppable from one `gbrain doctor --json` call so operators correlating the v0.41.27 retry-reconnect symptom fix with the offending caller don't have to grep two audit files. Module-import wrapped in try/catch so older brains without the v0.41.27 audit file degrade silently.
- `src/core/audit/db-disconnect-audit.ts` (v0.41.27.0, #1570, NEW) — JSONL audit for every call to `db.disconnect()` and `PostgresEngine.disconnect()`. Built on `audit-writer.ts` cathedral (no parallel subsystem; codex finding 11). Schema: `{ts, engine_kind: 'postgres'|'pglite'|'unknown', connection_style: 'module'|'instance'|'unknown', caller_stack, command, pid}`. `caller_stack` captured via `new Error().stack` truncated to ~20 frames so operators identify the offending caller without inflating JSONL forever. Privacy: stack frames carry file paths but NO SQL content / row data / user strings — matches the v0.20 shell-audit posture. File: `~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl` (honors `GBRAIN_AUDIT_DIR`). `readRecentDbDisconnects(hours=24)` walks current + previous ISO week and returns `{count, most_recent_caller, files_scanned}`. **The instrument half of the "instrument first, fix later" pivot** — v0.41.27.0 ships the symptom fix (retry reconnect callback + facts:absorb drain) AND this diagnostic surface; v0.41.28+ patches the specific ownership boundary once production data tells us which code path is calling `disconnect()` mid-process. Wired into `src/core/db.ts:disconnect` and `src/core/postgres-engine.ts:disconnect` (logs BEFORE the early-return branches so even no-op disconnects on never-connected engines are recorded — that case may itself be a caller-side bug). Pinned by `test/db-disconnect-audit.test.ts` (6 cases: round-trip, stack truncation, sort order, empty-dir nulls, stable feature name, EROFS best-effort).
- `src/core/facts/queue.ts:FactsQueue.drainPending` (v0.41.27.0, #1570) — new method `drainPending({timeout?: number}): Promise<{drained, unfinished}>`. Semantically distinct from `shutdown()` (which calls `this.internalAbort.abort()` and would abort the very facts:absorb worker trying to log its post-completion event — preserving the bug class we're fixing per codex finding 9). Drain lets in-flight finish; only the wait is bounded. Default timeout `1000ms` so commands that don't enqueue facts pay only one fast 0ms check before exit (per codex finding 10). `src/cli.ts` op-dispatch finally block awaits `getFactsQueue().drainPending({timeout: 1000})` BEFORE `engine.disconnect()`. Lazy-import keeps the facts-queue module off the hot path for ops that never touch it. Closes the trailing `'No database connection'` line that fired on stderr after `gbrain capture` because the post-page-write facts:absorb queued work outlived the CLI process. Pinned by `test/facts-queue-drain-pending.test.ts` (4 cases: empty fast-path, in-flight settled without abort, unfinished count on timeout, default timeout = 1000ms).
- `scripts/check-no-double-retry.sh` + `scripts/check-batch-audit-site.sh` (v0.41.19.0) — CI lint guards wired into `bun run verify`. The former greps src/ for `withRetry(...engine.{addLinksBatch|addTimelineEntriesBatch|upsertChunks})` patterns and fails the build on hit (Eng-D6 migration ordering hazard — prevents 3×3=9 retry amplification on incomplete reverts). The latter extracts every string-literal `auditSite: '...'` from src/ and validates each appears in the `BATCH_AUDIT_SITES` const in `src/core/retry.ts` (codex H-7 typo guard — prevents fragmented doctor output).
- `src/core/fail-improve.ts` — Deterministic-first, LLM-fallback loop with JSONL failure logging and auto-test generation
- `src/core/transcription.ts` — Audio transcription: Groq Whisper (default), OpenAI fallback, ffmpeg segmentation for >25MB
@@ -2962,6 +2964,22 @@ Bad values surface at `gbrain doctor` startup with a paste-ready fix
retry wrap is engine-level, but PGLite has no pooler so retries never
fire in practice.
**Dream cycle losing ~150 link rows per run with `'No database
connection: connect() has not been called'` errors in the log?** v0.41.27.0
makes the retry layer self-heal on a nulled-out database singleton. A
new `reconnect` callback on `withRetry` rebuilds the connection between
attempts; `PostgresEngine.batchRetry` injects `() => this.reconnect()`
so engine-level batch writes survive a mid-cycle disconnect by something
else in the same process. Same release: `gbrain capture` no longer trails
a `'No database connection'` stderr line from a background facts:absorb
worker firing after CLI exit — the op-dispatch finally block awaits
`getFactsQueue().drainPending({timeout: 1000})` before
`engine.disconnect()`. To find which code path is still calling
disconnect mid-process, run `gbrain doctor --json | jq '.checks[] |
select(.id=="batch_retry_health")'`; the extended check now surfaces
24h disconnect-call count and the most-recent caller frame from a new
`~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl` audit. (Closes #1570.)
**`gbrain brainstorm` returning `judge_failed: true` with 0 scored
ideas?** v0.41.21.0 closes the two bugs that caused it. The judge
hard-coded a 4K-token output cap; for any run past ~40 ideas the call
+1 -1
View File
@@ -141,5 +141,5 @@
"bun": ">=1.3.10"
},
"license": "MIT",
"version": "0.41.27.0"
"version": "0.41.28.0"
}
+17
View File
@@ -253,6 +253,23 @@ async function main() {
console.error(e instanceof Error ? e.message : String(e));
process.exit(1);
} finally {
// v0.41.25.0 (#1570) — drain the facts:absorb queue BEFORE disconnect
// so the fire-and-forget queue worker has a live engine to write its
// log against. Closes the bug class that absorb-log.ts:87-100 names:
// facts subsystem holds an engine reference past CLI exit, fires its
// post-completion log against a dead singleton, surfaces as a 'No
// database connection' stderr line on every `gbrain capture`.
//
// 1s timeout is per codex finding 10 from the v0.41.25 plan review:
// ops that don't enqueue facts (most read paths) pay only the
// 0-pending fast-path cost (~microseconds). Capture / import / sync
// that DO enqueue pay up to 1s while in-flight Haiku calls finish.
// Lazy-import keeps this off the hot path for ops that never touch
// the facts queue at all.
try {
const { getFactsQueue } = await import('./core/facts/queue.ts');
await getFactsQueue().drainPending({ timeout: 1000 });
} catch { /* best-effort; never block disconnect on drain failure */ }
await engine.disconnect();
if (forceExitTimer) clearTimeout(forceExitTimer);
// Narrow force-exit: only when the drain timed out AND we are NOT
+26 -4
View File
@@ -1180,6 +1180,28 @@ export async function checkBatchRetryHealth(_engine: BrainEngine): Promise<Check
const exhausted = result.events.filter((e) => e.outcome === 'exhausted');
const successful = result.events.filter((e) => e.outcome === 'success');
// v0.41.25.0 (#1570) — read the db-disconnect audit so the existing
// batch_retry_health check surfaces ALL connection-incident signal in
// one place (per codex finding 11: extend, don't add a new check).
// Disconnect events are informational — every CLI command legitimately
// disconnects at end-of-life. The value is the most_recent_caller
// frame: when the v0.41.25 retry reconnect callback fires, the
// operator runs `gbrain doctor` and the stack trace tells them which
// code path triggered the mid-process disconnect. v0.41.26 fixes
// that specific ownership boundary.
let disconnectNote = '';
try {
const { readRecentDbDisconnects } = await import('../core/audit/db-disconnect-audit.ts');
const dc = readRecentDbDisconnects(24);
if (dc.count > 0) {
// First-line of stack trace is the caller of logDbDisconnect; show
// it so the operator sees something compact in human output.
const firstFrame = (dc.most_recent_caller ?? '').split('\n')[0]?.trim() ?? '';
const frameSlug = firstFrame.length > 0 ? ` (most recent caller: ${firstFrame.slice(0, 200)})` : '';
disconnectNote = ` Disconnect-call audit: ${dc.count} call(s) in 24h${frameSlug}.`;
}
} catch { /* audit module unavailable; older brain, fine */ }
if (exhausted.length === 0) {
const note = result.corrupted_lines > 0
? ` (note: ${result.corrupted_lines} corrupt JSONL line(s) skipped)`
@@ -1190,7 +1212,7 @@ export async function checkBatchRetryHealth(_engine: BrainEngine): Promise<Check
return {
name: 'batch_retry_health',
status: 'ok',
message: `No exhausted batch retries in last 24h.${recoveredNote}${note}`,
message: `No exhausted batch retries in last 24h.${recoveredNote}${note}${disconnectNote}`,
};
}
@@ -1204,7 +1226,7 @@ export async function checkBatchRetryHealth(_engine: BrainEngine): Promise<Check
return {
name: 'batch_retry_health',
status: 'fail',
message: `${exhausted.length} exhausted batch retries in last 24h (worst: ${worstSite[0]} = ${worstSite[1]}). Sustained circuit-breaker incident. Fix: check pooler status; consider raising GBRAIN_BULK_MAX_RETRIES or moving to direct-connection.`,
message: `${exhausted.length} exhausted batch retries in last 24h (worst: ${worstSite[0]} = ${worstSite[1]}). Sustained circuit-breaker incident. Fix: check pooler status; consider raising GBRAIN_BULK_MAX_RETRIES or moving to direct-connection.${disconnectNote}`,
};
}
@@ -1213,7 +1235,7 @@ export async function checkBatchRetryHealth(_engine: BrainEngine): Promise<Check
return {
name: 'batch_retry_health',
status: 'warn',
message: `${exhausted.length} exhausted batch retries in last 24h (worst: ${worstSite[0]} = ${worstSite[1]}). Tune via GBRAIN_BULK_MAX_RETRIES / GBRAIN_BULK_RETRY_MAX_MS.`,
message: `${exhausted.length} exhausted batch retries in last 24h (worst: ${worstSite[0]} = ${worstSite[1]}). Tune via GBRAIN_BULK_MAX_RETRIES / GBRAIN_BULK_RETRY_MAX_MS.${disconnectNote}`,
};
}
@@ -1221,7 +1243,7 @@ export async function checkBatchRetryHealth(_engine: BrainEngine): Promise<Check
return {
name: 'batch_retry_health',
status: 'ok',
message: `${exhausted.length} exhausted batch retry(s) in last 24h (below per-site threshold of 3)`,
message: `${exhausted.length} exhausted batch retry(s) in last 24h (below per-site threshold of 3)${disconnectNote}`,
};
} catch (e) {
const msg = e instanceof Error ? e.message : String(e);
+158
View File
@@ -0,0 +1,158 @@
/**
* v0.41.25.0 (#1570) — db-disconnect call audit.
*
* Records every call to `db.disconnect()` and `PostgresEngine.disconnect()`
* so we can identify the offending code path that nulls the module
* singleton mid-cycle. The reported #1570 symptom — `gbrain dream` losing
* ~150 rows per cycle with `'No database connection'` errors — happens
* because some caller disconnects the shared singleton while other code
* paths are still using it. The retry-layer reconnect callback (v0.41.25
* symptom fix) covers the loss, but only this audit tells us WHO is
* actually calling disconnect mid-process.
*
* Per codex outside-voice finding 4 from the v0.41.25 plan review:
* "instrument first, fix later." v0.41.26 patches the specific ownership
* boundary once production data tells us the caller.
*
* Per codex finding 11: built on the existing `audit-writer.ts` cathedral
* — no parallel subsystem. Same file rotation, same best-effort write
* semantics, same readRecent walk. Doctor extends `batch_retry_health`
* to surface 24h count.
*
* Schema is intentionally narrow:
*
* - `ts` ISO-8601 timestamp
* - `engine_kind` 'postgres' | 'pglite' (PGLite paths still log
* for completeness; they're a no-op for the
* singleton bug but useful as background data)
* - `connection_style` 'module' | 'instance' — module-mode calls are
* the load-bearing ones (they touch the singleton)
* - `caller_stack` `new Error().stack` truncated to 20 frames so
* operators can identify the offending caller
* without inflating the JSONL forever
* - `command` argv[2] when known, else 'unknown'. Helps map
* an offending disconnect to a specific CLI cmd
* - `pid` process.pid for cross-process correlation
*
* Privacy: stack frames contain file paths but no SQL content, no row
* data, no user-supplied strings. Matches the shell-audit.ts posture
* from v0.20+. Audit file lives at
* `~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl` (honors
* `GBRAIN_AUDIT_DIR` via the shared `resolveAuditDir()` helper).
*/
import { createAuditWriter } from './audit-writer.ts';
export interface DbDisconnectAuditEvent {
ts: string;
engine_kind: 'postgres' | 'pglite' | 'unknown';
connection_style: 'module' | 'instance' | 'unknown';
caller_stack: string;
command: string;
pid: number;
}
const FEATURE_NAME = 'db-disconnect';
const writer = createAuditWriter<DbDisconnectAuditEvent>({
featureName: FEATURE_NAME,
errorLabel: 'db-disconnect-audit',
errorTrailer: '; continuing',
});
/**
* Capture the current call stack, normalized + truncated. We strip the
* first two frames (this function + the caller's audit-log helper line)
* so the resulting trace starts with the actual offending caller — the
* frame an operator wants to see. Cap at 20 frames to bound the JSONL
* line size on long stacks; production calls rarely need more than 8.
*
* Exported for unit tests to pin the stack-truncation contract.
*/
export function captureCallerStack(skipFrames = 2, maxFrames = 20): string {
const raw = new Error().stack ?? '';
// Bun's stack format: first line is "Error", then " at fn (file:line:col)"
// for each frame. Split, drop the "Error" line + `skipFrames` of our own
// helper frames, keep up to `maxFrames` after that.
const lines = raw.split('\n');
// Find the first frame line (starts with whitespace + "at "). The
// "Error" header is line 0; helper frames start at line 1.
const frameStart = lines.findIndex((l) => /^\s+at\s/.test(l));
if (frameStart < 0) return raw.slice(0, 4000); // fallback: hard byte cap
const callerFrames = lines.slice(frameStart + skipFrames, frameStart + skipFrames + maxFrames);
return callerFrames.join('\n');
}
/**
* Log one db-disconnect call. Best-effort: stderr-warns on write failure
* but never throws. The caller's disconnect path continues regardless.
*/
export function logDbDisconnect(
engineKind: DbDisconnectAuditEvent['engine_kind'],
connectionStyle: DbDisconnectAuditEvent['connection_style'],
): void {
// argv[2] is typically the gbrain subcommand (e.g. 'dream', 'capture').
// argv[0] is bun, argv[1] is the script path; the meaningful identity
// is argv[2]. Defensive fallback to 'unknown' for embedded callers.
const command = process.argv[2] ?? 'unknown';
writer.log({
engine_kind: engineKind,
connection_style: connectionStyle,
caller_stack: captureCallerStack(),
command,
pid: process.pid,
});
}
/**
* Read recent disconnect audit events. Consumed by
* `doctor.ts:checkBatchRetryHealth` to surface the 24h count + most-
* recent caller in the existing check (no new check needed per codex
* finding 11).
*
* `hours` defaults to 24 (the "is the bug firing right now" window),
* not the audit-writer default of 7 days. Doctor displays the 24h
* count; operators chasing a stale incident can pass a larger window.
*/
export interface ReadDbDisconnectResult {
events: DbDisconnectAuditEvent[];
/** Convenience: count of mid-process events in window. */
count: number;
/** Convenience: most recent caller frame (for doctor display). */
most_recent_caller: string | null;
/** Most recent timestamp (for doctor display). */
most_recent_ts: string | null;
}
export function readRecentDbDisconnects(
hours = 24,
now: Date = new Date(),
): ReadDbDisconnectResult {
// The shared writer uses `days`; convert hours → fractional days.
const days = hours / 24;
// readRecent walks current + previous ISO week, then filters by
// cutoff. Pass our hour-precision cutoff through the days-API.
const events = writer.readRecent(days, now);
// Defensive cutoff filter — the writer's day-precision cutoff can
// include events outside the actual hour-precision window when days<1.
const cutoff = now.getTime() - hours * 3_600_000;
const filtered = events.filter((ev) => {
const t = Date.parse(ev.ts);
return Number.isFinite(t) && t >= cutoff;
});
// Sort newest-first so the "most recent" pick is honest regardless of
// how readRecent ordered files internally.
filtered.sort((a, b) => Date.parse(b.ts) - Date.parse(a.ts));
const mostRecent = filtered[0];
return {
events: filtered,
count: filtered.length,
most_recent_caller: mostRecent?.caller_stack ?? null,
most_recent_ts: mostRecent?.ts ?? null,
};
}
/** @internal — test seam to pin the schema-version and file location. */
export function _dbDisconnectAuditFeatureName(): string {
return FEATURE_NAME;
}
+11
View File
@@ -225,6 +225,17 @@ export async function connect(config: EngineConfig): Promise<void> {
}
export async function disconnect(): Promise<void> {
// v0.41.25.0 (#1570) — instrument every disconnect call site so v0.41.26
// can identify the caller that's nulling the module singleton mid-cycle.
// Best-effort: audit failure must never block the actual disconnect.
// The audit module is lazy-imported to keep db.ts cold-path-free for
// tools that import db without ever calling disconnect.
try {
const { logDbDisconnect } = await import('./audit/db-disconnect-audit.ts');
// db.ts is always the module-singleton path by construction; no
// instance-pool callers go through here.
logDbDisconnect('postgres', 'module');
} catch { /* best-effort; never block disconnect on audit failure */ }
if (sql) {
await sql.end();
sql = null;
+49
View File
@@ -120,6 +120,55 @@ export class FactsQueue {
return this.inflightTotal;
}
/**
* v0.41.25.0 (#1570) — wait for currently pending + in-flight jobs to
* settle naturally. **Semantically distinct from `shutdown()`** — drain
* does NOT abort in-flight work, does NOT drop pending, and does NOT
* disable future enqueues. It just blocks until the queue reaches
* (pending=0 AND inflight=0) OR the timeout fires.
*
* Per codex finding 9 from /codex review of the v0.41.25 plan: the
* original "reuse shutdown" idea was wrong because shutdown aborts
* in-flight (`this.internalAbort.abort()`), which means the very
* facts:absorb worker that's trying to log its post-completion
* absorb event gets aborted mid-write. That preserves the bug class
* we're trying to fix.
*
* Per codex finding 10: this is bounded by `opts.timeout` (default
* 1000ms) so commands that don't enqueue facts pay only one fast
* 0ms check before exit. Capture / import / sync that DO enqueue
* pay up to 1s while in-flight Haiku calls finish.
*
* Returns `{drained, unfinished}` so callers can log the outcome
* for debugging (no stderr writes; that's the caller's choice).
* `unfinished > 0` means timeout fired with work still pending —
* those jobs aren't aborted, they just continue running while the
* caller proceeds to exit (the singleton-still-alive contract in
* the post-pivot architecture means they'll still be able to write
* their logs).
*/
async drainPending(
opts: { timeout?: number } = {},
): Promise<{ drained: number; unfinished: number }> {
const timeout = opts.timeout ?? 1000;
const initiallyPending = this.pending.length;
const initiallyInflight = this.inflightTotal;
if (initiallyPending === 0 && initiallyInflight === 0) {
return { drained: 0, unfinished: 0 };
}
const start = Date.now();
while (
(this.pending.length > 0 || this.inflightTotal > 0) &&
Date.now() - start < timeout
) {
// 25ms poll interval matches shutdown() below; consistent rhythm.
await sleep(25);
}
const unfinished = this.pending.length + this.inflightTotal;
const drained = initiallyPending + initiallyInflight - unfinished;
return { drained, unfinished };
}
/**
* Begin shutdown. Returns a promise that resolves once the queue has either
* fully drained in-flight (under shutdownGraceMs) OR the grace expired. After
+22
View File
@@ -194,6 +194,18 @@ export class PostgresEngine implements BrainEngine {
}
async disconnect(): Promise<void> {
// v0.41.25.0 (#1570) — instrument disconnect calls to identify the
// mid-process caller behind the singleton-null bug. The audit log
// captures connection_style so we can tell instance-pool teardowns
// (correct, end-of-worker-life) apart from module-singleton teardowns
// (the load-bearing class). Best-effort: audit failure never blocks
// the actual disconnect. Logged BEFORE the early-return branches so
// even a no-op disconnect (engine that was never connected) is
// recorded — that case may itself be a caller-side bug worth seeing.
try {
const { logDbDisconnect } = await import('./audit/db-disconnect-audit.ts');
logDbDisconnect('postgres', this._connectionStyle ?? 'unknown');
} catch { /* best-effort; never block disconnect on audit failure */ }
// v0.30.1: tear down the direct pool first if the manager owns one.
if (this.connectionManager) {
await this.connectionManager.disconnect();
@@ -1907,6 +1919,16 @@ export class PostgresEngine implements BrainEngine {
const msg = err instanceof Error ? err.message : String(err);
process.stderr.write(`[${auditSite}] connection blip, retrying (attempt ${attempt}/${opts.maxRetries}): ${msg}\n`);
},
// v0.41.25.0 (#1570): on null-singleton retryable errors, rebuild
// the connection BEFORE the inter-attempt sleep so the next attempt
// sees a live pool. `this.reconnect()` is race-safe via
// `_reconnecting` guard, handles both module and instance pools,
// and is a fast no-op when the underlying client is still healthy
// (postgres.js's own connection-replacement covers that case).
// Fail-loud per retry.ts contract: a reconnect throw propagates
// as the real cause, replacing the symptomatic
// "No database connection" error.
reconnect: () => this.reconnect(),
});
} catch (err) {
// Distinguish "retries exhausted" (a retryable error that ran out of
+43 -3
View File
@@ -112,8 +112,29 @@ export interface WithRetryOpts {
signal?: AbortSignal;
/** Audit-site label for observability. Must be in BATCH_AUDIT_SITES. */
auditSite?: BatchAuditSite;
/** Per-attempt callback fires on each retry (attempt is 1-based). */
onRetry?: (attempt: number, err: unknown) => void;
/**
* Per-attempt callback fires on each retry (attempt is 1-based).
*
* v0.41.25.0: now awaited. Sync callbacks (the only in-tree shape) work
* identically; async callbacks correctly delay the inter-attempt sleep.
*/
onRetry?: (attempt: number, err: unknown) => void | Promise<void>;
/**
* v0.41.25.0 — invoked between attempts AFTER `isRetryableConnError`
* classification but BEFORE the inter-attempt sleep. Use this to rebuild
* a dead connection / pool before the retry fires.
*
* Fail-loud posture (per codex finding 3 from /codex review): if reconnect
* throws, the throw PROPAGATES out of `withRetry` AS the new error,
* replacing the original retryable. Operators see the real cause
* ("auth failed", "EHOSTUNREACH") instead of "No database connection"
* for hours when DB credentials are bad.
*
* Engine-level callers (PostgresEngine.batchRetry) inject
* `() => this.reconnect()` which already handles both module and
* instance pools, race-safe via `_reconnecting` guard.
*/
reconnect?: () => Promise<void>;
}
/**
@@ -206,6 +227,9 @@ export function computeNextDelay(
* - Pass `BULK_RETRY_OPTS` for the Supavisor-tuned 3-retry exponential shape.
* - Non-retryable errors (per `isRetryableConnError`) throw immediately.
* - AbortSignal triggers `RetryAbortError` mid-sleep.
* - v0.41.25.0: optional `reconnect` callback runs between attempts AFTER
* classification but BEFORE the sleep. Fail-loud — a reconnect throw
* propagates as the new error.
*/
export async function withRetry<T>(
fn: () => Promise<T>,
@@ -227,7 +251,23 @@ export async function withRetry<T>(
if (!isRetryableConnError(err)) throw err;
lastErr = err;
if (attempt >= maxRetries) break;
opts.onRetry?.(attempt + 1, err);
// v0.41.25.0: onRetry is now awaited so async observability + audit
// hooks correctly run before the inter-attempt sleep. Sync arrows
// (the only in-tree shape) work identically.
await opts.onRetry?.(attempt + 1, err);
// v0.41.25.0: optional reconnect hook. PostgresEngine.batchRetry
// injects `() => this.reconnect()` so a null-singleton from a
// sibling caller's mid-process disconnect doesn't keep the retry
// hammering against a dead reference. Fail-loud: any throw from
// reconnect (auth failure, network partition) propagates AS the
// new error — operators see the real cause, not the symptom.
// v0.41.25 also ships diagnostic instrumentation on disconnect
// call sites to find the offending caller; this hook is the
// immediate-recovery half of that pair.
if (opts.reconnect) {
if (signal?.aborted) throw new RetryAbortError();
await opts.reconnect();
}
const delay = computeNextDelay(attempt, prevDelay, baseDelay, maxDelay, jitter);
prevDelay = delay;
await abortableSleep(delay, signal);
+148
View File
@@ -0,0 +1,148 @@
// v0.41.25.0 (#1570) — retry.ts reconnect callback contract.
//
// Pins the new `reconnect?: () => Promise<void>` opt added to WithRetryOpts
// per D3 + D9 + codex finding 3. The retry primitive stays pure (no db.ts
// coupling); engine-level callers inject `() => this.reconnect()`.
//
// Hermetic: no engine, no PGLite, no env mutation, no DATABASE_URL.
import { describe, expect, test } from 'bun:test';
import { withRetry, RetryAbortError } from '../../src/core/retry.ts';
class FakeGBrainError extends Error {
problem: string;
detail: string;
constructor(problem: string, detail: string) {
super(`${problem}: ${detail}`);
this.problem = problem;
this.detail = detail;
}
}
describe('withRetry reconnect callback (v0.41.25.0)', () => {
test('calls reconnect AFTER classification + onRetry, BEFORE sleep', async () => {
// Record the order of side effects so the contract is pinned: classifier
// result determines reconnect, onRetry observes the retry intent, then
// reconnect rebuilds state, THEN the inter-attempt sleep happens.
const order: string[] = [];
let attempts = 0;
const start = Date.now();
const result = await withRetry(
async () => {
attempts++;
order.push(`fn-attempt-${attempts}`);
if (attempts === 1) {
throw new FakeGBrainError('No database connection', 'connect() has not been called');
}
return 'recovered';
},
{
delayMs: 30, // small but observable sleep
onRetry: () => { order.push('onRetry'); },
reconnect: async () => { order.push('reconnect-start'); await new Promise(r => setTimeout(r, 1)); order.push('reconnect-end'); },
},
);
const elapsed = Date.now() - start;
expect(result).toBe('recovered');
expect(attempts).toBe(2);
// Required order: first attempt fails -> onRetry -> reconnect -> sleep -> second attempt
expect(order).toEqual([
'fn-attempt-1',
'onRetry',
'reconnect-start',
'reconnect-end',
'fn-attempt-2',
]);
// Sleep happened (delayMs=30) so elapsed must be at least delayMs + reconnect
expect(elapsed).toBeGreaterThanOrEqual(30);
});
test('does NOT call reconnect when opts.reconnect is undefined (back-compat)', async () => {
// Existing call sites that don't opt in must see identical v0.41.18.0 behavior.
let attempts = 0;
const result = await withRetry(
async () => {
attempts++;
if (attempts === 1) throw new Error('Connection terminated unexpectedly');
return 'ok';
},
{ delayMs: 0 },
);
expect(result).toBe('ok');
expect(attempts).toBe(2);
});
test('reconnect failure PROPAGATES as the new error (codex finding 3 fail-loud)', async () => {
// The reconnect helper itself throwing means the underlying problem
// isn't transient — DB really down, auth failed, etc. Operators want
// to see THAT error, not the masking "No database connection" symptom.
let attempts = 0;
let reconnectCalls = 0;
const realCause = new Error('AuthError: invalid credentials');
await expect(
withRetry(
async () => {
attempts++;
throw new FakeGBrainError('No database connection', 'connect() has not been called');
},
{
delayMs: 0,
maxRetries: 3,
reconnect: async () => {
reconnectCalls++;
throw realCause;
},
},
),
).rejects.toThrow('AuthError: invalid credentials');
// First attempt threw, then reconnect threw immediately — no further attempts.
expect(attempts).toBe(1);
expect(reconnectCalls).toBe(1);
});
test('signal.aborted BEFORE reconnect call short-circuits with RetryAbortError', async () => {
const ctrl = new AbortController();
let attempts = 0;
let reconnectCalls = 0;
// Abort the moment fn throws but BEFORE reconnect would fire.
await expect(
withRetry(
async () => {
attempts++;
ctrl.abort(); // fire abort right when the retryable error throws
throw new FakeGBrainError('No database connection', 'x');
},
{
delayMs: 30,
signal: ctrl.signal,
reconnect: async () => { reconnectCalls++; },
},
),
).rejects.toBeInstanceOf(RetryAbortError);
expect(attempts).toBe(1);
// Reconnect MUST NOT fire after abort — clean shutdown takes priority.
expect(reconnectCalls).toBe(0);
});
test('onRetry is now awaited (back-compat-safe for sync arrows)', async () => {
// An async onRetry taking 50ms should delay the inter-attempt sleep by
// 50ms. v0.41.18.0 fire-and-forget would have lost that delay.
let attempts = 0;
const start = Date.now();
await withRetry(
async () => {
attempts++;
if (attempts === 1) throw new Error('Connection terminated unexpectedly');
return 'ok';
},
{
delayMs: 0, // sleep itself is 0
onRetry: async () => { await new Promise(r => setTimeout(r, 50)); },
},
);
const elapsed = Date.now() - start;
expect(attempts).toBe(2);
// delayMs=0 so the ONLY source of elapsed time is the awaited onRetry.
expect(elapsed).toBeGreaterThanOrEqual(45); // 45 to absorb scheduler noise
});
});
+105
View File
@@ -0,0 +1,105 @@
/**
* v0.41.25.0 (#1570) — db-disconnect-audit JSONL contract.
*
* Pins:
* - logDbDisconnect → readRecentDbDisconnects round-trip
* - caller_stack truncated to ~20 frames
* - Best-effort write: corrupt write target doesn't throw to caller
* - 24h window honored (events outside window filtered)
*
* Uses `withEnv()` per test-isolation lint rule R1.
*/
import { describe, test, expect } from 'bun:test';
import * as fs from 'node:fs';
import * as path from 'node:path';
import * as os from 'node:os';
import { withEnv } from './helpers/with-env.ts';
import {
logDbDisconnect,
readRecentDbDisconnects,
captureCallerStack,
_dbDisconnectAuditFeatureName,
} from '../src/core/audit/db-disconnect-audit.ts';
async function withFreshAuditDir(body: (tmpDir: string) => void | Promise<void>): Promise<void> {
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-db-disconnect-audit-'));
try {
await withEnv({ GBRAIN_AUDIT_DIR: tmpDir }, async () => {
await body(tmpDir);
});
} finally {
try { fs.rmSync(tmpDir, { recursive: true, force: true }); } catch { /* ignore */ }
}
}
describe('db-disconnect-audit (v0.41.25.0)', () => {
test('log → read round-trip preserves required fields', async () => {
await withFreshAuditDir(() => {
logDbDisconnect('postgres', 'module');
const result = readRecentDbDisconnects(24);
expect(result.count).toBe(1);
expect(result.events.length).toBe(1);
expect(result.events[0]).toMatchObject({
engine_kind: 'postgres',
connection_style: 'module',
pid: process.pid,
});
expect(typeof result.events[0].ts).toBe('string');
expect(typeof result.events[0].caller_stack).toBe('string');
expect(result.events[0].caller_stack.length).toBeGreaterThan(0);
expect(result.most_recent_caller).toBe(result.events[0].caller_stack);
expect(result.most_recent_ts).toBe(result.events[0].ts);
});
});
test('captureCallerStack truncates to maxFrames', () => {
const stack = captureCallerStack(0, 5);
const lines = stack.split('\n');
expect(lines.length).toBeLessThanOrEqual(5);
});
test('readRecentDbDisconnects sorts newest-first', async () => {
await withFreshAuditDir(() => {
logDbDisconnect('postgres', 'module');
logDbDisconnect('postgres', 'instance');
logDbDisconnect('pglite', 'unknown');
const result = readRecentDbDisconnects(24);
expect(result.count).toBe(3);
// Newest first: events array should be in reverse log order (or at
// least chronologically ordered by parseable ts). The last logged
// event should appear first or earliest in the events list. The
// strict newest-first contract is what doctor displays.
const timestamps = result.events.map(e => Date.parse(e.ts));
for (let i = 1; i < timestamps.length; i++) {
expect(timestamps[i - 1]).toBeGreaterThanOrEqual(timestamps[i]);
}
});
});
test('empty audit dir returns zero count + null fields', async () => {
await withFreshAuditDir(() => {
const result = readRecentDbDisconnects(24);
expect(result.count).toBe(0);
expect(result.events).toEqual([]);
expect(result.most_recent_caller).toBeNull();
expect(result.most_recent_ts).toBeNull();
});
});
test('feature name is stable (drives audit filename on disk)', () => {
// Pin the filename prefix so a future rename can't silently strand
// old audit files. Operators with v0.41.25 deployments have
// ~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl files.
expect(_dbDisconnectAuditFeatureName()).toBe('db-disconnect');
});
test('audit write is best-effort — unreadable dir does NOT throw', async () => {
// Point GBRAIN_AUDIT_DIR at a path the writer cannot create (a non-
// existent root we don't have perms for). The writer should stderr-
// warn but not throw to the caller's disconnect flow.
await withEnv({ GBRAIN_AUDIT_DIR: '/proc/1/cannot-create-here-1570' }, () => {
expect(() => logDbDisconnect('postgres', 'module')).not.toThrow();
});
});
});
@@ -0,0 +1,140 @@
/**
* v0.41.25.0 (#1570) — focused regression test for the dream-cycle
* row-loss bug class. Each case pins a real production failure mode
* codex recommended pinning (codex finding 4: instrument + targeted
* regression test, not architectural refactor).
*
* Skipped when DATABASE_URL is unset — mirrors every other test/e2e/
* file's posture. Caller is expected to bring up gbrain-test-pg via
* the canonical lifecycle described in CLAUDE.md.
*/
import { describe, test, expect, beforeAll, afterAll, beforeEach } from 'bun:test';
import * as fs from 'node:fs';
import * as path from 'node:path';
import * as os from 'node:os';
import { PostgresEngine } from '../../src/core/postgres-engine.ts';
import * as db from '../../src/core/db.ts';
import { withEnv } from '../helpers/with-env.ts';
import {
readRecentDbDisconnects,
logDbDisconnect,
} from '../../src/core/audit/db-disconnect-audit.ts';
const DATABASE_URL = process.env.DATABASE_URL;
const skip = !DATABASE_URL;
if (skip) {
// eslint-disable-next-line no-console
console.log('Skipping db-singleton-shared-recovery E2E (DATABASE_URL not set)');
}
describe.skipIf(skip)('v0.41.25.0 db-singleton shared-recovery regressions (#1570)', () => {
let tmpAuditDir: string;
beforeAll(async () => {
// Fresh module-level connection so each test starts from a known state.
await db.disconnect();
await db.connect({ database_url: DATABASE_URL! });
}, 30_000);
afterAll(async () => {
await db.disconnect();
if (tmpAuditDir) {
try { fs.rmSync(tmpAuditDir, { recursive: true, force: true }); } catch { /* ignore */ }
}
});
beforeEach(() => {
tmpAuditDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gbrain-1570-e2e-'));
});
test('CASE 1: shared singleton survives mid-operation disconnect via retry reconnect', async () => {
// Reproduce the dream-cycle scenario: caller A is mid-batch, caller B
// disconnects the module singleton, caller A's NEXT attempt enters
// retry and the reconnect callback rebuilds the singleton before the
// retry's fn fires. This is the symptom-fix contract we ship.
await db.connect({ database_url: DATABASE_URL! });
const engineA = new PostgresEngine();
await engineA.connect({ database_url: DATABASE_URL! });
const engineB = new PostgresEngine();
await engineB.connect({ database_url: DATABASE_URL! });
// Sanity: both engines share the live singleton.
expect((await engineA.sql`SELECT 1 as ok`)[0].ok).toBe(1);
expect((await engineB.sql`SELECT 1 as ok`)[0].ok).toBe(1);
// Engine B disconnects mid-operation (the "offending caller" scenario).
// This nulls the module singleton for engine A too.
await engineB.disconnect();
// Engine A's direct unsafe call will throw — proving the bug class
// exists at the engine.sql layer.
let directThrew = false;
try {
await engineA.sql`SELECT 1`;
} catch {
directThrew = true;
}
expect(directThrew).toBe(true);
// The retry layer's reconnect callback recovers. We exercise it via
// engine.reconnect() directly (which is what batchRetry's injected
// reconnect callback calls). After reconnect, engine A's next call
// succeeds.
await engineA.reconnect();
const afterRecovery = await engineA.sql`SELECT 1 as ok`;
expect(afterRecovery[0].ok).toBe(1);
// Cleanup
await engineA.disconnect();
});
test('CASE 2: diagnostic audit records every mid-process disconnect call', async () => {
// Per codex finding 4: instrument first. Production data tells us
// which caller is firing the mid-process disconnect. This case pins
// that the instrumentation is wired correctly: a disconnect call
// emits an audit JSONL line containing connection_style + caller_stack.
await withEnv({ GBRAIN_AUDIT_DIR: tmpAuditDir }, async () => {
await db.connect({ database_url: DATABASE_URL! });
const engine = new PostgresEngine();
await engine.connect({ database_url: DATABASE_URL! });
// module-style engine.disconnect() should log an audit line.
await engine.disconnect();
// Read it back. doctor uses the same readRecentDbDisconnects path.
const result = readRecentDbDisconnects(24);
expect(result.count).toBeGreaterThanOrEqual(1);
const last = result.events[0];
expect(last.engine_kind).toBe('postgres');
expect(['module', 'unknown']).toContain(last.connection_style);
expect(last.caller_stack.length).toBeGreaterThan(0);
expect(last.pid).toBe(process.pid);
});
});
test('CASE 3: instance-pool disconnect leaves shared singleton ALIVE for other callers', async () => {
// Codex finding 5/6: BrainEngine contract is asymmetric across engines.
// Instance-pool engines (workerPoolSize set) should NEVER touch the
// module singleton on disconnect. This case pins that contract —
// existing v0.28.1 idempotency test covers the same shape but here
// we explicitly verify the "two callers, one in instance mode" case
// matters for #1570.
await db.connect({ database_url: DATABASE_URL! });
const moduleEngine = new PostgresEngine();
await moduleEngine.connect({ database_url: DATABASE_URL! }); // module mode
const workerEngine = new PostgresEngine();
await workerEngine.connect({ database_url: DATABASE_URL!, poolSize: 2 }); // instance mode
// Worker disconnect: should ONLY tear down its own _sql, not touch module.
await workerEngine.disconnect();
// Module engine still works.
const result = await moduleEngine.sql`SELECT 1 as ok`;
expect(result[0].ok).toBe(1);
await moduleEngine.disconnect();
});
});
+17
View File
@@ -19,6 +19,7 @@ import {
resolveSchemaMultimodalDim,
PGVECTOR_COLUMN_MAX_DIMS,
} from '../src/core/embedding-dim-check.ts';
import { configureGateway, resetGateway } from '../src/core/ai/gateway.ts';
// Canonical pattern: single engine per file, init once, disconnect once.
// The two tests below diverge in whether they want a migrated brain or a
@@ -27,6 +28,21 @@ import {
let engine: PGLiteEngine;
beforeAll(async () => {
// Hermeticity guard (cross-file gateway-state leak class — see CLAUDE.md
// "Test-isolation lint and helpers"). initSchema builds the
// content_chunks vector column at the gateway's configured dim. The
// bunfig preload pins OpenAI/1536, but its beforeEach only re-applies
// legacy when the gateway was RESET (throws) — it does NOT correct a
// sibling that configured a different LIVE dim (e.g. ZE/1280) and never
// reset. Under weight-based shard bin-packing, such a sibling can run
// first, so pin 1536 explicitly here BEFORE initSchema (this is exactly
// the "call configureGateway() in your own beforeAll" escape hatch the
// preload documents). Reset in afterAll so we don't leak 1536 onward.
configureGateway({
embedding_model: 'openai:text-embedding-3-large',
embedding_dimensions: 1536,
env: { ...process.env },
});
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
@@ -34,6 +50,7 @@ beforeAll(async () => {
afterAll(async () => {
await engine.disconnect();
resetGateway();
});
describe('readContentChunksEmbeddingDim', () => {
+94
View File
@@ -0,0 +1,94 @@
/**
* v0.41.25.0 (#1570) — FactsQueue.drainPending contract.
*
* Per codex finding 9 from /codex review of the v0.41.25 plan: drain is
* DIFFERENT from shutdown. Shutdown aborts in-flight via internal signal;
* drain lets in-flight finish naturally. This file pins the distinction
* so any future refactor that conflates them re-fails.
*
* Hermetic: no engine, no DATABASE_URL.
*/
import { describe, test, expect, beforeEach } from 'bun:test';
import { FactsQueue, __resetFactsQueueForTests } from '../src/core/facts/queue.ts';
beforeEach(() => {
__resetFactsQueueForTests();
});
const sleep = (ms: number) => new Promise<void>(r => setTimeout(r, ms));
describe('FactsQueue.drainPending — codex F9 distinct-from-shutdown contract', () => {
test('returns {drained:0,unfinished:0} fast when queue is empty', async () => {
const q = new FactsQueue();
const start = Date.now();
const result = await q.drainPending({ timeout: 1000 });
const elapsed = Date.now() - start;
expect(result).toEqual({ drained: 0, unfinished: 0 });
// Fast path: empty drain should NOT spend the full timeout.
expect(elapsed).toBeLessThan(50);
});
test('awaits in-flight to settle WITHOUT aborting (the codex F9 contract)', async () => {
// Distinct from shutdown(): shutdown calls internalAbort.abort() which
// makes runEntry's catch see an AbortError and counters bump
// dropped_shutdown. drainPending must let the job run to completion
// so the facts:absorb post-completion log actually fires.
const q = new FactsQueue({ shutdownGraceMs: 5000 });
let completed = false;
let signalSeenAborted = false;
q.enqueue(async (signal) => {
// Sleep so drain has a real wait to do.
await sleep(60);
// Witness the signal state at completion — should NOT be aborted.
if (signal.aborted) signalSeenAborted = true;
completed = true;
}, 'sess');
// Give pump a microtask to claim the job.
await sleep(5);
const result = await q.drainPending({ timeout: 1000 });
expect(completed).toBe(true);
expect(signalSeenAborted).toBe(false);
expect(result.unfinished).toBe(0);
expect(result.drained).toBeGreaterThan(0);
// shutdown's dropped_shutdown counter MUST NOT increment from drain.
expect(q.getCounters().dropped_shutdown).toBe(0);
expect(q.getCounters().completed).toBe(1);
});
test('returns with unfinished > 0 when timeout fires; does NOT hang or abort', async () => {
const q = new FactsQueue();
let completed = false;
// Job runs longer than the drain timeout.
q.enqueue(async () => {
await sleep(300);
completed = true;
}, 'sess');
await sleep(5); // give pump a tick to claim
const start = Date.now();
const result = await q.drainPending({ timeout: 80 });
const elapsed = Date.now() - start;
// Drain returned WITHIN timeout window (small slack for scheduler).
expect(elapsed).toBeLessThan(180);
expect(result.unfinished).toBeGreaterThan(0);
// Job was NOT aborted — it should still be running.
expect(completed).toBe(false);
// Let it finish so the test process doesn't leak the timer.
await sleep(400);
expect(completed).toBe(true);
});
test('default timeout is 1000ms when opts.timeout omitted', async () => {
const q = new FactsQueue();
// Job that runs forever (well, 2s, longer than default).
q.enqueue(async () => { await sleep(2000); }, 'sess');
await sleep(5);
const start = Date.now();
const result = await q.drainPending();
const elapsed = Date.now() - start;
// Should return at the default 1000ms timeout, NOT 2000ms.
expect(elapsed).toBeGreaterThanOrEqual(950);
expect(elapsed).toBeLessThan(1200);
expect(result.unfinished).toBeGreaterThan(0);
});
});
+26
View File
@@ -27,12 +27,36 @@ import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { operations } from '../src/core/operations.ts';
import type { OperationContext } from '../src/core/operations.ts';
import { OperationError } from '../src/core/operations.ts';
import { configureGateway, resetGateway, __setEmbedTransportForTests } from '../src/core/ai/gateway.ts';
const putPageOp = operations.find((o) => o.name === 'put_page')!;
let engine: PGLiteEngine;
beforeAll(async () => {
// Hermeticity guard (cross-file gateway-state leak class — see CLAUDE.md
// "Test-isolation lint and helpers"). put_page embeds via the gateway.
// A sibling file in the same shard can leave the gateway configured with
// a live provider/key (e.g. OpenAI + a CI placeholder `sk-test`); the
// weight-based shard bin-packing reshuffles which files share a process,
// so we cannot rely on a benign neighbor. Pin the gateway to legacy
// OpenAI/1536 (so initSchema builds a 1536-d column) AND stub the embed
// transport so put_page's embed never touches the network — these tests
// assert provenance columns, not embeddings. A dummy key is required in
// the gateway env because instantiateEmbedding builds the OpenAI client
// (which checks for a key) BEFORE the stubbed transport is reached; the
// stub then intercepts the actual call, so the key never leaves the
// process. Reset in afterAll so this file doesn't leak onward.
configureGateway({
embedding_model: 'openai:text-embedding-3-large',
embedding_dimensions: 1536,
env: { ...process.env, OPENAI_API_KEY: process.env.OPENAI_API_KEY || 'sk-test-stub' },
});
__setEmbedTransportForTests(async ({ values }: any) => ({
embeddings: values.map(() => new Array(1536).fill(0)),
usage: { tokens: 0 },
}) as any);
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
@@ -40,6 +64,8 @@ beforeAll(async () => {
afterAll(async () => {
await engine.disconnect();
__setEmbedTransportForTests(null);
resetGateway();
});
beforeEach(async () => {