Files
gbrain/test/migrate.test.ts
T
7267462311 v0.31.4 feat: takes v2 — lessons from 100K-take production extraction (#795)
* feat: takes v2 — lessons from 100K-take production extraction

Consolidates everything learned from the first full takes extraction run
(28,256 pages, 100,720 takes, $361 on Azure GPT-5.5) and subsequent
cross-modal eval (GPT-5.5 + Opus 4.6, scored 6.8/10 overall).

## Fixes

**fix(cli): add recall and forget to CLI_ONLY set**
v0.31 added these commands to handleCliOnly() but forgot the gate set.
Both fell through to cliOps.get() → 'Unknown command'.

**feat(synthesize): auto-enable when corpus dir is configured**
Setting session_corpus_dir is now sufficient — enabled defaults to true
when a corpus dir is set. Explicit enabled=false still wins. Eliminates
the footgun where users configure a corpus dir and nothing happens.

**feat(engine): round takes weights to 0.05 increments**
Cross-modal eval found false precision (0.74, 0.82) implies calibration
accuracy that doesn't exist. Both postgres and pglite engines now round
on insert. 1.0 and 0.0 are preserved exactly.

## Documentation

**docs: takes-vs-facts architectural distinction**
New doc explaining the two epistemological layers, why they must never be
conflated, how the dream cycle consolidate phase bridges them, and
production extraction data (model selection, eval dimensions, key
learnings for extraction prompts).

**docs(takes-fence): clarify holder semantics with eval examples**
Holder = who HOLDS the belief, NOT who it's ABOUT. Expanded JSDoc with
concrete right/wrong examples from the cross-modal eval. Additional
rules: amplification ≠ endorsement, self-reported ≠ verified, founder
describing company → people/founder not companies/slug.

## Tests (17 new, all passing)

- 5 synthesize-enabled-default tests
- 6 takes-holder-semantics tests
- 6 takes-weight-rounding tests

## Cross-Modal Eval Context

| Dimension         | GPT-5.5 | Opus 4.6 | Avg  |
|-------------------|---------|----------|------|
| Accuracy          | 7       | 8        | 7.5  |
| Attribution       | 6       | 7        | 6.5  |
| Weight calibration| 7       | 7        | 7.0  |
| Kind classification| 6      | 7        | 6.5  |
| Signal density    | 7       | 6        | 6.5  |

Top improvements addressed in this PR:
1. Holder vs subject confusion (docs + tests)
2. Weight false precision (runtime enforcement)
3. Takes ≠ facts distinction (architectural doc)
4. Synthesis auto-enable (runtime fix)
5. recall/forget CLI routing (bug fix)

* docs(filing-rules): anchor takes attribution rules (EXP-3)

Adds a "Takes attribution" section to skills/_brain-filing-rules.md
distilling the 6 rules from docs/takes-vs-facts.md into a terse
contract that downstream agents (OpenClaw, Wintermute) can read as
their canonical filing surface.

Documentation only — no in-repo runtime consumer (synthesize.ts reads
the .json file, not the .md). EXP-4 lands the runtime parser-level
holder validation.

Codex review #9: relabels EXP-3 as documentation, not quality work.
The runtime check is EXP-4.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(takes): weight backfill v46 + NaN hardening at 4 sites (EXP-1, Hardening)

Migration v46 (takes_weight_round_to_grid): backfills pre-v0.32 takes.weight
to the 0.05 grid the engine layer (PR #795) enforces on insert. Cross-modal
eval over 100K production takes flagged 0.74, 0.82-style values as false
precision; this brings existing data to the same grid that all new writes
already use.

Tolerance-based comparison (abs > 0.001) avoids the float32-noise re-touch
loop that the naive `weight <> ROUND(...)` form would create — REAL/NUMERIC
comparison promotes weight to DOUBLE PRECISION first, surfacing ~1e-7
representation noise as inequality. The 0.05 grid is 5e-2, so any genuine
off-grid value clears the 1e-3 threshold cleanly.

`transaction: false` (codex review #2 correction): not for mid-statement
resume (a single SQL statement either completes or rolls back). What it
actually buys is freeing the migration runner from holding a long
transaction so other gbrain processes can interleave.

NaN hardening (codex review #8): extracts `normalizeWeightForStorage()` to
takes-fence.ts as a single source of truth used by all 4 takes write sites:
  - pglite-engine.ts addTakesBatch
  - pglite-engine.ts updateTake (was missed in original PR — only clamped,
    didn't round; now rounds AND guards NaN)
  - postgres-engine.ts addTakesBatch
  - postgres-engine.ts updateTake (same fix)

The helper guards `!Number.isFinite()` BEFORE the [0,1] range check (NaN
comparisons are always false, so NaN survived the prior clamp and reached
Math.round(NaN * 20) / 20 = NaN, written through to the DB).

Tests:
- test/migrations-v46-takes-weight-backfill.test.ts: behavioral PGLite test
  (rounding fixture + Codex #2 re-run idempotency + on-grid preservation).
- test/takes-weight-rounding.test.ts: imports the real helper, adds NaN /
  Infinity / -Infinity / null / undefined / updateTake-shape coverage.
- test/migrate.test.ts: structural assertions for v46 SQL shape.

All 52 tests pass; typecheck clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(doctor): takes_weight_grid check + pure helper extraction (EXP-2)

Adds doctor's `takes_weight_grid` slice — the post-migration drift detector
for the 0.05 weight grid v0.31 enforces on insert and v46 backfilled.

Codex review #7 corrected the original plan's "extend test/doctor.test.ts
with 3 cases" estimate. runDoctor() is a side-effectful command with
process.exit branches, and the existing tests are mostly source-structure
assertions. The fix: extract `takesWeightGridCheck(engine: BrainEngine)`
as a pure exported function. runDoctor calls it. Tests target the helper
directly with stubbed engines for the missing-table branch and against
real PGLite for the 4 ratio bands.

Branches:
  - 0 takes total → ok ("No takes yet")
  - off_grid / total > 10% → fail (with apply-migrations fix hint)
  - 1% < off_grid / total ≤ 10% → warn (same fix hint)
  - else → ok
  - takes table missing (pre-v37) → warn, graceful skip

Tolerance comparison matches migration v46 (abs > 1e-3) so float32 noise
doesn't make a healthy brain look broken.

Tests (test/doctor.test.ts):
  - takesWeightGridCheck export shape
  - 0-takes branch (avoids divide-by-zero)
  - 100% on-grid via engine.addTakesBatch (which now normalizes)
  - 8/10 off-grid → fail
  - 5/100 off-grid → warn
  - missing-table branch via stub engine

All 21 doctor tests pass; typecheck clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(takes): holder runtime validation + producer seam (EXP-4)

Adds parser-level holder grammar enforcement so cross-modal eval's #1
attribution error (holder/subject confusion, scored 6.5/10 across 100K
production takes) shows up as a sync-failure record an operator can see.

Changes:

- src/core/sync.ts: exports SLUG_SEGMENT_PATTERN, the actual character
  class slugifySegment() produces ([a-z0-9._-]). Codex review #3 — the
  initial plan's stricter regex would have warned on legitimate slugs
  like `companies/acme.io` and `people/foo_bar`. HOLDER_REGEX now wraps
  this shared pattern instead of inventing a parallel grammar.

- src/core/takes-fence.ts: HOLDER_REGEX + isValidHolder() helper.
  parseTakesFence() emits TAKES_HOLDER_INVALID warnings for non-matching
  holders. Row preserved (markdown source-of-truth contract).

  Catches the eval's failure modes — `Garry`, `people/Garry-Tan`,
  `world/garry-tan`, `users/garry`, whitespace-only — while keeping
  `companies/acme.io`, `people/foo_bar`, `notes/v1.0.0`-style dotted
  slugs valid. Bare-slug form (`garry`, `alice`) accepted as v0.32 legacy
  compat — production brains shipped with bare-slug holders before the
  namespaced JSDoc landed in PR #795. Reserved for v0.33 promotion.

- src/core/cycle/extract-takes.ts (codex review #4 producer seam): adds
  `failedFiles: Array<{path, error}>` to ExtractTakesResult. Both fs
  and db extraction paths populate it from TAKES_HOLDER_INVALID warnings
  so the migration orchestrator can hand it to recordSyncFailures().
  Without this seam, extending classifyErrorCode would do nothing
  (the regex would have nothing to classify).

- src/commands/migrations/v0_28_0.ts: phaseBBackfill calls
  recordSyncFailures(result.failedFiles, 'migration:v0.28.0-backfill')
  after extractTakes completes. Best-effort — persistence failure
  doesn't fail the backfill phase. Doctor's `sync_failures` check now
  shows TAKES_HOLDER_INVALID=N breakdown after upgrade.

- src/core/sync.ts:classifyErrorCode: extends with TAKES_HOLDER_INVALID
  + TAKES_TABLE_MALFORMED / TAKES_ROW_NUM_COLLISION / TAKES_FENCE_UNBALANCED
  bucket. Previously these warnings bucketed to UNKNOWN.

Tests (test/takes-holder-validation.test.ts — 26 cases):
- Canonical forms (world / brain / people-namespace / companies-namespace)
- Codex #3 dotted-slug + underscore-slug positives
- Legacy bare-slug compat positives
- Eval-flagged error mode rejections (uppercase, mixed case, world/<slug>,
  unrecognized prefix, whitespace, embedded slash)
- HOLDER_REGEX anchoring guard
- SLUG_SEGMENT_PATTERN export shape + drift guard against the wrapping regex
- parseTakesFence end-to-end emission contract
- classifyErrorCode regex coverage

127 tests pass across affected files; typecheck clean. No existing fixtures
broken (legacy bare-slug compat preserves old `garry`-style holders during
the v0.32 transition window).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(eval): gbrain eval takes-quality CLI — DB-authoritative + 4-mode (EXP-5)

Reproducible cross-modal quality eval for the takes layer. Three frontier
models score a sample against the 5-dim rubric, the runner aggregates to
PASS/FAIL/INCONCLUSIVE, the receipt persists to eval_takes_quality_runs.
Trend mode segregates by rubric_version; regress mode is a CI gate that
exits 1 when any dim regresses past --threshold.

Subcommands:
  run     [--limit N --cycles N --budget-usd N --slug-prefix P --models a,b,c]
  replay  <receipt-path> [--json]                 # NO BRAIN required
  trend   [--limit N --rubric-version V --json]
  regress --against <receipt> [--threshold T --json]

Codex review integrations (D7 — all 10 findings landed):

  #1 json-repair shim re-exports BOTH parseModelJSON AND the
     ParsedScore + ParsedModelResult types. The original plan only
     re-exported the function, which would have compile-broken
     cross-modal-eval/aggregate.ts:19's type import.

  #3 Receipt name binds (corpus_sha8, prompt_sha8, models_sha8,
     rubric_sha8) so a future rubric tweak segregates trend rows
     instead of silently corrupting the quality-over-time graph.
     RUBRIC_VERSION + rubric_sha8 are persisted in every receipt.

  #4 Pricing fail-closed: any model not in pricing.ts produces an
     actionable PricingNotFoundError before any HTTP call fires.
     Same drift problem as cross-modal-eval/runner.ts:estimateCost(),
     but explicit instead of silent zero.

  #5 Aggregate requires ALL 5 declared rubric dimensions per model.
     Cross-modal-eval v1's union-of-whatever-parsed pattern allowed a
     model to omit a dim and still PASS — that's a regression-gate
     hole. Now: missing-dim drops the contribution, treated identically
     to a parse failure. Empty-scores PASS regression guard preserved.

  #6 DB-authoritative receipt persistence. Original two-phase plan had
     a split-brain reconciliation gap (disk-success/DB-fail vanishes
     from trend; DB-success/disk-fail unreplayable). Now DB row is the
     source of truth (carries full receipt JSON in a JSONB column);
     disk artifact is best-effort. replay reads disk first; loadReceiptFromDb
     reconstructs from DB when the disk file is missing.

  #10 Brain-routing: replay is the only sub-subcommand that doesn't
      need a brain. cli.ts no-DB bypass routes "eval takes-quality replay"
      directly to runReplayNoBrain, which exits 0/1/2 cleanly without
      ever touching the engine. Other modes go through connectEngine.

Files added:
  src/core/eval-shared/json-repair.ts (hoisted from cross-modal-eval)
  src/core/takes-quality-eval/{rubric,pricing,aggregate,receipt-name,
                                receipt-write,receipt,replay,regress,trend,runner}.ts
  src/commands/eval-takes-quality.ts
  docs/eval-takes-quality.md (stable schema_version: 1 contract)
  10 test files (83 cases — aggregate / receipt-name / shim / pricing /
                 rubric / receipt-write / replay / trend / regress / cli)

Files modified:
  src/cli.ts: replay no-DB bypass + engine-required dispatch
  src/core/cross-modal-eval/json-repair.ts → re-export shim
  src/core/migrate.ts: append v47 (eval_takes_quality_runs table)
  src/core/pglite-schema.ts + src/schema.sql: mirror the v47 table for
    fresh-install path. RLS toggled on the new table.
  src/core/schema-embedded.ts: regenerated via build:schema
  test/migrate.test.ts: 6 structural cases for v47

186 tests pass; typecheck clean. Replay verified working end-to-end
(reads receipt JSON file without DATABASE_URL, exits with the verdict
code, prints actionable error on missing file).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(eval): fill EXP-5 unit-test gaps + test-isolation lint fix

Three additions identified during the test-gap audit:

  1. test/eval-takes-quality-boundaries.test.ts (4 cases):
     - empty corpus → "no takes to evaluate" (pre-LLM)
     - source=fs reserved for v0.33 → clear refusal
     - --budget-usd + unknown model → PricingNotFoundError BEFORE any
       network call (codex review #4 fail-closed contract)
     - --budget-usd null + unknown model → no pre-flight pricing error
       (proves pricing pre-flight gates ONLY when budget is set)

  2. test/eval-takes-quality-runner.serial.test.ts (7 cases):
     End-to-end runner integration with mock.module-stubbed gateway.chat.
     Quarantined as *.serial.test.ts because mock.module leaks across
     files in the same shard process (R2 in check-test-isolation.sh).
     Covers:
       - 3 PASS scores → verdict=pass with all dim scores in receipt
       - all model errors → INCONCLUSIVE
       - 1 success + 2 errors → INCONCLUSIVE (need >=2 contributing)
       - 3 successes with low scores → FAIL
       - budget cap fires before cycle 1 (no chat() ever called)
       - budget cap allows cycle when projection fits

  3. test/eval-takes-quality-receipt-write.test.ts: refactored to use
     withEnv() helper for GBRAIN_HOME mutation instead of direct
     process.env writes. The original beforeAll mutation tripped the
     check-test-isolation.sh R1 lint. withEnv() saves/restores via
     try/finally per-test so other shard files don't see the override.

Verification:
  bun run test       → 4977 pass / 0 fail
  bun run test:serial → 179 pass / 0 fail
  bun run verify     → clean (typecheck + 9 pre-checks pass)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(eval): real-Postgres E2E for eval_takes_quality_runs (EXP-5)

Pure-PGLite tests already cover the receipt-write contract; this E2E
verifies the same code path against actual Postgres so the postgres.js
JSONB encoding and the v47 migration apply cleanly under production
conditions.

Coverage (8 cases):
  - migration v47 created the table with all expected columns
  - writeReceiptToDb persists full receipt_json on Postgres
  - 4-sha UNIQUE constraint enforces ON CONFLICT DO NOTHING idempotency
    (3 inserts → 1 row)
  - rubric_version segregation: distinct rubric_sha8 → distinct row
    (codex review #3 — rubric epoch separation)
  - loadTrend reads in DESC order on Postgres
  - loadReceiptFromDb reconstructs receipt JSON via the JSONB column
  - writeReceipt (combined) succeeds with disk artifact + DB row
  - trend SELECT plan executes (planner picks index on larger tables)

Skips gracefully when DATABASE_URL is unset (existing hasDatabase()
helper). Uses the canonical setupDB/teardownDB from test/e2e/helpers.ts.
GBRAIN_HOME mutation is wrapped in withEnv() per the v0.32.0 test-isolation
lint contract.

Verification:
  bash scripts/run-e2e.sh → 71 files / 499 tests / 0 fail (full E2E suite)
  bun test test/e2e/eval-takes-quality.test.ts → 8 / 8 pass standalone

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: fill v0.32 unit + E2E gap audit (3 new files, 36 cases)

Audit of shipped v0.32 code surfaced 4 wiring gaps that the per-EXP unit
tests didn't cover. Adding direct integration tests for each so a future
refactor can't accidentally bypass the helper or unwire the producer seam.

test/extract-takes-holder-producer-seam.test.ts (7 cases) — codex review
#4 producer seam. Verifies extractTakesFromDb populates ExtractTakesResult.
failedFiles[] when parseTakesFence emits TAKES_HOLDER_INVALID warnings,
and that the entry shape is recordSyncFailures-compatible. Without this
test, the v0_28_0 migration's recordSyncFailures call would have silently
fed it nothing if a refactor accidentally dropped the failedFiles append.
Covers: valid holder (no entry), invalid uppercase, world/<slug>, mixed
valid+invalid, legacy bare-slug compat, malformed-table-only (no leak),
recordSyncFailures shape compatibility.

test/engine-weight-rounding-integration.test.ts (15 cases) — codex review
#8 integration coverage. Helper is unit-tested; this proves both engines'
addTakesBatch + updateTake paths actually call it. PGLite-side coverage
mirrors the test/e2e/takes-weight-rounding-postgres.test.ts E2E for real
Postgres. Covers: 0.74→0.75, 0.82→0.80, on-grid identity, NaN→0.5,
Infinity→0.5, clamp high/low, undefined default, mixed batch order,
updateTake rounds (was unhardened pre-v0.32), updateTake NaN, updateTake
preserves prior weight when undefined.

test/e2e/takes-weight-rounding-postgres.test.ts (6 cases, 14 expects) —
real-Postgres write-path coverage. Specifically tests the postgres.js
unnest() bind path that PGLite doesn't exercise:
  - addTakesBatch rounds via the unnest() bind shape
  - addTakesBatch handles NaN at the postgres.js array marshaling layer
  - 10-row mixed batch (4 off-grid) rounds each independently
  - updateTake rounds on real Postgres
  - updateTake handles NaN
  - migration v48 tolerance matches engine-write tolerance (round-trip
    proof — engine-rounded value is invisible to v48's WHERE clause)

Verification:
  bun run test       → 5166 pass / 0 fail (parallel unit, 128s)
  bun run test:serial → 190 pass / 0 fail
  bun run test:e2e   → 71 / 74 files; 3 pre-existing env-inheritance
                       failures (serve-http-oauth, sources-remote-mcp,
                       thin-client — confirmed identical on master in
                       this environment, documented in CLAUDE.md)
  bun run verify     → clean

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): connect engine in withConfiguredSql; unbreak 3 OAuth E2E suites

Real production bug, not just a test-environment issue.
withConfiguredSql in src/commands/auth.ts created a PostgresEngine via
createEngine() but never called engine.connect(). The PostgresEngine.sql
getter falls back to db.getConnection() (the module-level singleton) when
its instance _sql is unset — and db.connect() wasn't called either.

So every `gbrain auth` subcommand (create, list, revoke, register-client,
revoke-client) crashed with the misleading "No database connection:
connect() has not been called" error on real Postgres. Anyone with a
Postgres-backed brain hit this. The error pointed at gbrain init which
made the regression invisible — users assumed they hadn't initialized.

Verified by running `gbrain auth register-client` directly:
  Before: "Error: No database connection: connect() has not been called."
  After:  "OAuth client registered: ..." with credentials printed.

This fix unblocked all 3 previously-failing E2E suites (which all use
register-client in beforeAll):
  serve-http-oauth.test.ts:    0/28 → 28/28 pass
  sources-remote-mcp.test.ts:  0/14 → 14/14 pass
  thin-client.test.ts:         0/7  →  6/7 pass + 1 documented skip

Two surgical test-side fixes also landed:

1. test/e2e/thin-client.test.ts:182 — assertion typo. Test expected
   r.stderr to contain "thin client" (space). Actual refusal message
   says "(thin-client of <url>)" with hyphen. Loosened to /thin[- ]client/
   so a future format tweak doesn't false-fail.

2. test/e2e/thin-client.test.ts:239 — skipped "remote ping triggers
   autopilot-cycle" with a clear TODO. Test asks the wrong question
   against the existing fixture: `gbrain serve --http` deliberately
   does NOT start a job worker (workers run via separate `gbrain jobs
   work` process), so the submitted autopilot-cycle job sits in
   `waiting` forever. Test was supposed to fall back to the self-imposed
   `--timeout`, but `gbrain remote ping --timeout` doesn't honor the cap
   when callRemoteTool hangs (loop only checks elapsed time between
   iterations; a single in-flight callTool with no AbortSignal blocks
   forever). Two real follow-ups would unblock: thread an AbortSignal
   through callRemoteTool's MCP callTool path, OR start a `gbrain jobs
   work` subprocess in beforeAll. Either is its own PR. Wire path
   coverage isn't lost — exercised by every other test in this file
   plus the entire serve-http-oauth.test.ts suite.

Verification:
  bun test test/e2e/serve-http-oauth.test.ts test/e2e/sources-remote-mcp.test.ts test/e2e/thin-client.test.ts
    → 47 pass / 1 skip / 0 fail in 8.4s
  bun run verify → clean

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: garrytan-agents <garrytan-agents@users.noreply.github.com>
Co-authored-by: Garry Tan <garrytan@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 06:34:40 -07:00

1391 lines
66 KiB
TypeScript

import { describe, test, expect, beforeAll, afterAll, spyOn } from 'bun:test';
import { LATEST_VERSION, runMigrations, MIGRATIONS, getIdleBlockers, hasPendingMigrations } from '../src/core/migrate.ts';
import type { IdleBlocker } from '../src/core/migrate.ts';
import type { BrainEngine } from '../src/core/engine.ts';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { readFileSync } from 'fs';
import { resolve } from 'path';
describe('migrate', () => {
test('LATEST_VERSION is a number >= 1', () => {
expect(typeof LATEST_VERSION).toBe('number');
expect(LATEST_VERSION).toBeGreaterThanOrEqual(1);
});
test('runMigrations is exported and callable', async () => {
expect(typeof runMigrations).toBe('function');
});
// Integration tests for actual migration execution require DATABASE_URL
// and are covered in the E2E suite (test/e2e/mechanical.test.ts)
});
// v0.28.5 — A1: cheap probe used by `connectEngine` to gate `initSchema()`
// so already-migrated brains don't pay the schema-replay cost on every
// short-lived CLI invocation. Closes #651 in cooperation with X1's
// post-upgrade auto-apply, without #652's perf regression.
describe('hasPendingMigrations', () => {
test('returns false on a fully-migrated brain (version === LATEST)', async () => {
const engine = new PGLiteEngine();
await engine.connect({});
try {
await engine.initSchema(); // applies all migrations through LATEST_VERSION
expect(await hasPendingMigrations(engine)).toBe(false);
} finally {
await engine.disconnect();
}
}, 30000);
test('returns true when version config is behind LATEST_VERSION', async () => {
const engine = new PGLiteEngine();
await engine.connect({});
try {
await engine.initSchema();
// Simulate an older brain by rewinding the version row.
await engine.setConfig('version', '1');
expect(await hasPendingMigrations(engine)).toBe(true);
} finally {
await engine.disconnect();
}
}, 30000);
test('returns true when version config is missing entirely (defensive default)', async () => {
const engine = new PGLiteEngine();
await engine.connect({});
try {
// Don't call initSchema. Probe against an empty PGlite — getConfig should
// either return null (treated as version=1) or throw on missing config
// table; either way the probe must say "yes pending."
expect(await hasPendingMigrations(engine)).toBe(true);
} finally {
await engine.disconnect();
}
}, 30000);
});
// ─────────────────────────────────────────────────────────────────
// v0.18.0 — v16 sources_table_additive (Step 1, Lane A)
// ─────────────────────────────────────────────────────────────────
// v16 is the ADDITIVE-ONLY migration: it installs the sources primitive
// without breaking the engine's existing ON CONFLICT (slug) upserts.
// The breaking schema changes (pages.source_id NOT NULL, composite
// UNIQUE, files.page_slug → page_id, file_migration_ledger,
// links.resolution_type) land in v17 alongside the engine API rewrite
// so the engine can execute the new ON CONFLICT (source_id, slug)
// atomically with the schema change.
// ─────────────────────────────────────────────────────────────────
describe('migrate v20 — sources_table_additive', () => {
const v20 = MIGRATIONS.find(m => m.version === 20);
test('v20 exists', () => {
expect(v20).toBeDefined();
expect(v20!.name).toBe('sources_table_additive');
});
test('v20 creates sources table', () => {
expect(v20!.sql).toContain('CREATE TABLE IF NOT EXISTS sources');
expect(v20!.sql).toContain('id TEXT PRIMARY KEY');
expect(v20!.sql).toContain('name TEXT NOT NULL UNIQUE');
expect(v20!.sql).toContain('config JSONB NOT NULL');
});
test("v20 seeds 'default' source inheriting sync config", () => {
expect(v20!.sql).toContain("INSERT INTO sources (id, name, local_path, last_commit, config)");
expect(v20!.sql).toContain("'default'");
// The default source pulls from existing config so post-upgrade
// identity is preserved.
expect(v20!.sql).toContain("SELECT value FROM config WHERE key = 'sync.repo_path'");
expect(v20!.sql).toContain("SELECT value FROM config WHERE key = 'sync.last_commit'");
});
test('v20 default source is federated=true (backward-compat)', () => {
// federated=true ensures pre-v0.17 brains keep single-namespace
// search semantics — every page appears in unqualified search.
expect(v20!.sql).toContain('"federated": true');
});
test('v20 is idempotent on re-run', () => {
// CREATE TABLE IF NOT EXISTS + NOT EXISTS subquery on INSERT.
expect(v20!.sql).toContain('CREATE TABLE IF NOT EXISTS sources');
expect(v20!.sql).toContain('WHERE NOT EXISTS (SELECT 1 FROM sources WHERE id = ');
});
test('v20 does NOT touch pages / ingest_log / files / links', () => {
// Step 1 is additive-only. Breaking changes deferred to v17 so they
// land with the engine rewrite (Step 2). Guard against anyone
// accidentally re-expanding v16's scope.
expect(v20!.sql).not.toContain('ALTER TABLE pages');
expect(v20!.sql).not.toContain('ALTER TABLE ingest_log');
expect(v20!.sql).not.toContain('ALTER TABLE files');
expect(v20!.sql).not.toContain('ALTER TABLE links');
expect(v20!.handler).toBeUndefined();
});
});
// ─────────────────────────────────────────────────────────────────
// v0.18.0 — v17 pages_source_id_composite_unique (Step 2, Lane B)
// ─────────────────────────────────────────────────────────────────
// ─────────────────────────────────────────────────────────────────
// v0.26.3 — v33 admin_dashboard_columns_v0_26_3
// ─────────────────────────────────────────────────────────────────
// SQL-shape guard: PR #586 referenced 5 columns + a new index that didn't
// exist in any prior migration. Without v33, /admin/api/agents 503s and
// the request-log INSERT silently swallows column-doesn't-exist errors.
// This test pins the column set so a future refactor can't silently drop
// part of the migration without the test failing.
describe('migrate v33 — admin_dashboard_columns_v0_26_3', () => {
const v33 = MIGRATIONS.find(m => m.version === 33);
test('v33 exists with the expected name', () => {
expect(v33).toBeDefined();
expect(v33!.name).toBe('admin_dashboard_columns_v0_26_3');
});
test('v33 adds all 5 columns referenced by serve-http.ts and oauth-provider.ts', () => {
const sql = v33!.sql;
expect(sql).toContain('ALTER TABLE oauth_clients');
expect(sql).toContain('ADD COLUMN IF NOT EXISTS token_ttl INTEGER');
expect(sql).toContain('ADD COLUMN IF NOT EXISTS deleted_at TIMESTAMPTZ');
expect(sql).toContain('ALTER TABLE mcp_request_log');
expect(sql).toContain('ADD COLUMN IF NOT EXISTS agent_name TEXT');
expect(sql).toContain('ADD COLUMN IF NOT EXISTS params JSONB');
expect(sql).toContain('ADD COLUMN IF NOT EXISTS error_message TEXT');
});
test('v33 backfills mcp_request_log.agent_name from oauth_clients + access_tokens', () => {
const sql = v33!.sql;
expect(sql).toContain('UPDATE mcp_request_log');
expect(sql).toContain('SET agent_name = COALESCE(');
expect(sql).toContain('FROM oauth_clients WHERE client_id = m.token_name');
expect(sql).toContain('FROM access_tokens WHERE name = m.token_name');
expect(sql).toContain('WHERE agent_name IS NULL');
});
test('v33 creates idx_mcp_log_agent_time for the new agent filter', () => {
expect(v33!.sql).toContain('idx_mcp_log_agent_time');
expect(v33!.sql).toContain('mcp_request_log(agent_name, created_at DESC)');
});
test('v33 uses ADD COLUMN IF NOT EXISTS so re-runs are idempotent', () => {
// All ALTER lines must be IF NOT EXISTS — re-running migrations on a
// brain that already has v33 columns must be a no-op, not a duplicate
// column error.
const sql = v33!.sql;
const addColumnLines = sql.match(/ADD COLUMN[^,;]+/gi) || [];
expect(addColumnLines.length).toBeGreaterThanOrEqual(5);
for (const line of addColumnLines) {
expect(line).toContain('IF NOT EXISTS');
}
});
});
// ============================================================
// v0.27 — v35 subagent_provider_neutral_persistence_v0_27
// ============================================================
// Codex F-OV-1 / D11. The subagent_messages and subagent_tool_executions
// tables stored Anthropic-shaped tool_use / tool_result blocks as JSONB.
// When a worker resumes mid-loop and the live model is OpenAI/DeepSeek/etc,
// the persisted shape is the runtime contract — translation at read time
// is lossy.
//
// Fix: schema_version + provider_id columns. v=1 = legacy Anthropic shape,
// v=2 = provider-neutral ChatBlock format (commit 2). subagent.ts (commit
// 2) writes v=2 going forward.
//
// Renumbered v34→v35→v36 across master merges: master's v34
// (destructive_guard_columns) and v35 (auto_rls_event_trigger) landed first.
describe('migrate v36 — subagent_provider_neutral_persistence_v0_27', () => {
const v36 = MIGRATIONS.find(m => m.version === 36);
test('v36 exists with the expected name', () => {
expect(v36).toBeDefined();
expect(v36!.name).toBe('subagent_provider_neutral_persistence_v0_27');
});
test('v36 adds schema_version + provider_id to both subagent tables', () => {
const sql = v36!.sql;
expect(sql).toContain('ALTER TABLE subagent_messages');
expect(sql).toContain('ALTER TABLE subagent_tool_executions');
// schema_version present in both tables
const schemaVersionMatches = sql.match(/ADD COLUMN IF NOT EXISTS schema_version INTEGER NOT NULL DEFAULT 1/g) || [];
expect(schemaVersionMatches.length).toBe(2);
// provider_id present in both tables
const providerIdMatches = sql.match(/ADD COLUMN IF NOT EXISTS provider_id TEXT/g) || [];
expect(providerIdMatches.length).toBe(2);
});
test('v36 keeps DEFAULT 1 so existing rows are taggable as legacy Anthropic shape', () => {
// Existing rows backfill to schema_version=1 (legacy) automatically via
// DEFAULT. No explicit UPDATE needed; subagent.ts read path checks the
// version and dispatches the right mapper.
expect(v36!.sql).toContain('DEFAULT 1');
});
test('v36 creates idx_subagent_messages_provider for cost rollups', () => {
expect(v36!.sql).toContain('idx_subagent_messages_provider');
expect(v36!.sql).toContain('subagent_messages (job_id, provider_id)');
});
test('v36 ALTERs are idempotent (ADD COLUMN IF NOT EXISTS)', () => {
const sql = v36!.sql;
const addColumnLines = sql.match(/ADD COLUMN[^,;]+/gi) || [];
expect(addColumnLines.length).toBe(4);
for (const line of addColumnLines) {
expect(line).toContain('IF NOT EXISTS');
}
// Index creation must also be idempotent.
expect(sql).toContain('CREATE INDEX IF NOT EXISTS');
});
test('PGLite fresh-install schema reflects v36 columns', async () => {
const { PGLITE_SCHEMA_SQL } = await import('../src/core/pglite-schema.ts');
expect(PGLITE_SCHEMA_SQL).toContain('schema_version INTEGER NOT NULL DEFAULT 1');
expect(PGLITE_SCHEMA_SQL).toContain('provider_id TEXT');
expect(PGLITE_SCHEMA_SQL).toContain('idx_subagent_messages_provider');
});
test('embedded schema (src/core/schema-embedded.ts) reflects v36 columns', async () => {
const { SCHEMA_SQL } = await import('../src/core/schema-embedded.ts');
expect(SCHEMA_SQL).toContain('schema_version');
expect(SCHEMA_SQL).toContain('provider_id');
expect(SCHEMA_SQL).toContain('idx_subagent_messages_provider');
});
});
describe('migrate v21 — pages_source_id_composite_unique', () => {
const v21 = MIGRATIONS.find(m => m.version === 21);
test('v21 exists and is paired with Step 2 engine rewrite', () => {
expect(v21).toBeDefined();
expect(v21!.name).toBe('pages_source_id_composite_unique');
});
// Post-codex restructure: v21 is engine-split.
// Postgres path = additive only (source_id + index). The UNIQUE swap
// and files_page_slug_fkey drop moved into v23's atomic transaction.
// PGLite path = full (add + unique swap) because PGLite has no
// concurrent writers so the integrity window doesn't apply.
test('v21 uses sqlFor for engine-specific paths (post-codex)', () => {
expect(v21!.sql).toBe('');
expect(v21!.sqlFor).toBeDefined();
expect(v21!.sqlFor!.postgres).toBeDefined();
expect(v21!.sqlFor!.pglite).toBeDefined();
});
test('v21 Postgres path: additive only (source_id + index)', () => {
const pg = v21!.sqlFor!.postgres!;
expect(pg).toContain('ALTER TABLE pages ADD COLUMN IF NOT EXISTS source_id TEXT');
// DEFAULT 'default' closes the race where an INSERT between ADD COLUMN
// and SET NOT NULL could leave source_id NULL (Codex second-pass review).
expect(pg).toContain("NOT NULL DEFAULT 'default' REFERENCES sources(id)");
expect(pg).toContain('CREATE INDEX IF NOT EXISTS idx_pages_source_id');
// The UNIQUE swap and files FK drop must NOT be in the Postgres path.
// They moved into v23's atomic transaction to close the partial-state
// window codex identified.
expect(pg).not.toContain('pages_slug_key');
expect(pg).not.toContain('files_page_slug_fkey');
});
test('v21 PGLite path: additive + UNIQUE swap (no integrity window)', () => {
const pgl = v21!.sqlFor!.pglite!;
expect(pgl).toContain('ALTER TABLE pages ADD COLUMN IF NOT EXISTS source_id TEXT');
expect(pgl).toContain('CREATE INDEX IF NOT EXISTS idx_pages_source_id');
// PGLite swaps the unique here (no files table means no FK to drop).
expect(pgl).toContain('ALTER TABLE pages DROP CONSTRAINT IF EXISTS pages_slug_key');
expect(pgl).toContain('pages_source_slug_key');
expect(pgl).toContain('UNIQUE (source_id, slug)');
// PGLite path doesn't touch files (doesn't exist on PGLite).
expect(pgl).not.toContain('files_page_slug_fkey');
});
});
// ─────────────────────────────────────────────────────────────────
// v0.18.0 — v19 files_source_id_page_id_ledger (Step 7, Lane E)
// ─────────────────────────────────────────────────────────────────
describe('migrate v23 — files_source_id_page_id_ledger', () => {
const v23 = MIGRATIONS.find(m => m.version === 23);
test('v23 exists as handler-only (Postgres files table, PGLite no-op)', () => {
expect(v23).toBeDefined();
expect(v23!.name).toBe('files_source_id_page_id_ledger');
expect(v23!.sql).toBe('');
expect(v23!.handler).toBeDefined();
});
test('v23 handler gates on engine.kind for PGLite (no files table)', () => {
expect(v23!.handler!.toString()).toMatch(/engine\.kind\s*===\s*["']pglite["']/);
});
test('v23 adds files.source_id + files.page_id + ledger creation', () => {
const body = v23!.handler!.toString();
expect(body).toContain('ALTER TABLE files ADD COLUMN IF NOT EXISTS source_id');
expect(body).toContain('ALTER TABLE files ADD COLUMN IF NOT EXISTS page_id');
expect(body).toContain('CREATE TABLE IF NOT EXISTS file_migration_ledger');
});
test('v23 is atomic: wraps all work in engine.transaction (integrity-window fix)', () => {
const body = v23!.handler!.toString();
// Codex caught: if files_page_slug_fkey is dropped in v21 but the
// replacement files.page_id is only added in v23, a process-death
// between v21 and v23 leaves files permanently unconstrained.
// Fix: move BOTH the FK drop AND the pages UNIQUE swap into v23,
// wrap everything in engine.transaction so it commits atomically.
expect(body).toContain('engine.transaction');
expect(body).toContain('files_page_slug_fkey');
expect(body).toContain('pages_slug_key');
expect(body).toContain('pages_source_slug_key');
});
test('v23 backfills files.page_id scoped to default source (Codex fix)', () => {
const body = v23!.handler!.toString();
// Without source_id='default' scope, the JOIN could hit the wrong
// page after new sources with duplicate slugs are added.
expect(body).toContain('UPDATE files f');
expect(body).toContain("p.source_id = 'default'");
});
test('v23 ledger PK is file_id (Codex: two sources can share old path)', () => {
const body = v23!.handler!.toString();
expect(body).toContain('file_id INTEGER PRIMARY KEY');
// State machine values all present.
for (const state of ['pending', 'copy_done', 'db_updated', 'complete', 'failed']) {
expect(body).toContain(`'${state}'`);
}
});
});
describe('migrate — ordering guarantee (v15 must NOT be skipped by v16)', () => {
test('runMigrations sorts by version ascending', async () => {
// Regression: if v16 preceded v15 in the MIGRATIONS array, the iterator
// would setConfig(version, 16) first, then skip v15 on the next pass.
// runMigrations applies a defensive sort so array order doesn't matter.
// This test asserts v15 exists (if we broke the sort, v15 would still
// exist in MIGRATIONS but would never apply at runtime).
const v15 = MIGRATIONS.find(m => m.version === 15);
const v20 = MIGRATIONS.find(m => m.version === 20);
expect(v15).toBeDefined();
expect(v20).toBeDefined();
// Sanity: versions are distinct and progress.
const versions = MIGRATIONS.map(m => m.version);
const uniq = new Set(versions);
expect(uniq.size).toBe(versions.length);
});
});
// ─────────────────────────────────────────────────────────────────
// v0.18.1 RLS hardening — structural guard for migration v24
// ─────────────────────────────────────────────────────────────────
//
// The base schema shipped 8 gbrain-managed public tables without RLS
// enabled (access_tokens, mcp_request_log, minion_inbox,
// minion_attachments, subagent_messages, subagent_tool_executions,
// subagent_rate_leases, gbrain_cycle_locks). Migration v12 created
// two more (budget_ledger, budget_reservations) without RLS.
// Migration v24 backfills the ENABLE RLS statements for existing
// brains. This test guards against regressions where the migration
// gets truncated or the wrong tables get enabled.
describe('migration v24 — rls_backfill_missing_tables', () => {
const RLS_BACKFILL_TABLES = [
'access_tokens',
'mcp_request_log',
'minion_inbox',
'minion_attachments',
'subagent_messages',
'subagent_tool_executions',
'subagent_rate_leases',
'gbrain_cycle_locks',
'budget_ledger',
'budget_reservations',
];
test('exists with the expected name', () => {
const v24 = MIGRATIONS.find(m => m.version === 24);
expect(v24).toBeDefined();
expect(v24?.name).toBe('rls_backfill_missing_tables');
});
test('enables RLS on all 10 backfill tables', () => {
const v24 = MIGRATIONS.find(m => m.version === 24);
expect(v24).toBeDefined();
const sql = v24!.sql || '';
for (const tbl of RLS_BACKFILL_TABLES) {
expect(sql).toContain(`ALTER TABLE ${tbl} ENABLE ROW LEVEL SECURITY`);
}
});
test('is gated on BYPASSRLS so it never locks a non-bypass session out of its data', () => {
const v24 = MIGRATIONS.find(m => m.version === 24);
const sql = v24!.sql || '';
expect(sql).toContain('rolbypassrls');
// The gate can be either IF has_bypass / early-raise pattern.
expect(sql).toMatch(/IF (NOT )?has_bypass/);
});
// Self-healing guard: the budget_* tables are migration-only (v12). If an
// operator manually dropped them, or if a brain was somehow pinned to a
// pre-v12 version when those tables didn't exist, a bare `ALTER TABLE
// budget_ledger ...` would fail with 42P01 and abort v24. Wrapping those
// two ALTERs in an `IF EXISTS (information_schema.tables ...)` check lets
// the migration skip them silently instead of erroring out. The other 8
// tables are created by schema.sql on every initSchema and don't need
// the guard — bare ALTER is fine.
test('guards budget_ledger + budget_reservations with information_schema.tables IF EXISTS', () => {
const v24 = MIGRATIONS.find(m => m.version === 24);
const sql = v24!.sql || '';
// Both budget tables must be wrapped in an existence check.
expect(sql).toMatch(
/IF EXISTS \(SELECT 1 FROM information_schema\.tables[\s\S]{0,200}table_name = 'budget_ledger'\)[\s\S]{0,200}ALTER TABLE budget_ledger ENABLE ROW LEVEL SECURITY/,
);
expect(sql).toMatch(
/IF EXISTS \(SELECT 1 FROM information_schema\.tables[\s\S]{0,200}table_name = 'budget_reservations'\)[\s\S]{0,200}ALTER TABLE budget_reservations ENABLE ROW LEVEL SECURITY/,
);
});
// Codex found: if v24 RAISE WARNINGs instead of raising on non-BYPASSRLS,
// the migration runner still bumps schema_version to 24, permanently
// skipping the backfill on future runs even after the role is fixed.
// The fix is to raise loudly so the transaction aborts, version stays
// at 23, and the next initSchema call retries after role reassignment.
test('fails loudly on non-BYPASSRLS roles instead of silently bumping version', () => {
const v24 = MIGRATIONS.find(m => m.version === 24);
const sql = v24!.sql || '';
expect(sql).toMatch(/RAISE EXCEPTION[^;]*BYPASSRLS/);
expect(sql).not.toMatch(/RAISE WARNING[^;]*BYPASSRLS/);
});
test('LATEST_VERSION has caught up to 24', () => {
expect(LATEST_VERSION).toBeGreaterThanOrEqual(24);
});
// PGLite has no RLS engine and is intrinsically single-tenant. The 8 RLS
// backfill ALTER statements target tables that may not exist on PGLite
// (subagent_*, minion_inbox aren't always present in pglite-schema.ts).
// sqlFor.pglite='' makes v24 a no-op on PGLite while still bumping the
// version counter. Engine.kind discrimination in runMigrations selects
// sqlFor[engine.kind] over m.sql. Issue #395.
test('uses a PGLite no-op override so local brains skip Postgres-only RLS ALTER TABLEs', () => {
const v24 = MIGRATIONS.find(m => m.version === 24);
expect(v24?.sqlFor?.pglite).toBe('');
});
});
// ─────────────────────────────────────────────────────────────────
// v0.26.7 — migration v35 structural guards (auto-RLS event trigger)
// ─────────────────────────────────────────────────────────────────
//
// The PR review caught that the original v35 had three correctness issues:
// - FORCE ROW LEVEL SECURITY locked out non-BYPASSRLS table owners.
// - Trigger fired on Supabase-managed schemas (auth/storage/realtime/...).
// - EXCEPTION WHEN OTHERS would silently swallow per-table failures and
// replace a transactional rollback (loud) with a permissive default (quiet).
// These tests pin the corrected shape so a future revert can't reintroduce
// the original bugs.
describe('migration v35 — auto_rls_event_trigger structural guards', () => {
test('exists with the expected name and SQL shape', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
expect(v35).toBeDefined();
expect(v35?.name).toBe('auto_rls_event_trigger');
expect((v35?.sqlFor as any)?.postgres?.length).toBeGreaterThan(0);
});
test('uses a PGLite no-op override (no event trigger support on PGLite)', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
expect(v35?.sqlFor?.pglite).toBe('');
});
test('does NOT issue FORCE ROW LEVEL SECURITY (D1: ENABLE only)', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
const sql = ((v35?.sqlFor as any)?.postgres ?? '') as string;
expect(sql).not.toMatch(/FORCE\s+ROW\s+LEVEL\s+SECURITY/i);
expect(sql).toMatch(/ENABLE\s+ROW\s+LEVEL\s+SECURITY/i);
});
test('trigger function is scoped to schema_name = public (D2)', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
const sql = ((v35?.sqlFor as any)?.postgres ?? '') as string;
expect(sql).toMatch(/schema_name\s*=\s*'public'/);
});
test('WHEN TAG covers CREATE TABLE, CREATE TABLE AS, and SELECT INTO (D6)', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
const sql = ((v35?.sqlFor as any)?.postgres ?? '') as string;
expect(sql).toMatch(/WHEN\s+TAG\s+IN\s*\([^)]*'CREATE TABLE'[^)]*\)/i);
expect(sql).toMatch(/'CREATE TABLE AS'/);
expect(sql).toMatch(/'SELECT INTO'/);
});
test('does NOT contain EXCEPTION WHEN OTHERS inside the trigger function (D5 reversed)', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
const sql = ((v35?.sqlFor as any)?.postgres ?? '') as string;
// ddl_command_end fires inside the DDL transaction, so a failed ALTER
// aborts the offending CREATE TABLE — that's the security guarantee.
// Wrapping in EXCEPTION WHEN OTHERS would convert that loud rollback
// into a silent permissive default. Pin the absence.
expect(sql.toUpperCase()).not.toContain('EXCEPTION WHEN OTHERS');
});
test('backfill block uses %I.%I identifier quoting (codex correction)', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
const sql = ((v35?.sqlFor as any)?.postgres ?? '') as string;
// The backfill iterates pg_class and ALTERs each non-exempt RLS-off public
// table. Mixed-case identifiers require %I quoting; raw concat would break.
expect(sql).toMatch(/format\(\s*'ALTER TABLE %I\.%I/);
});
test('backfill exemption regex matches the doctor.ts contract', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
const sql = ((v35?.sqlFor as any)?.postgres ?? '') as string;
// doctor.ts:418 EXEMPT_RE = /^GBRAIN:RLS_EXEMPT\s+reason=\S.{3,}/
// The plpgsql side must use the same pattern (via ~) so the two surfaces
// honor identical exemptions.
expect(sql).toMatch(/'\^GBRAIN:RLS_EXEMPT\\s\+reason=\\S\.\{3,\}'/);
});
test('backfill is gated on rolbypassrls (matches v24 posture)', () => {
const v35 = MIGRATIONS.find(m => m.version === 35);
const sql = ((v35?.sqlFor as any)?.postgres ?? '') as string;
expect(sql).toMatch(/rolbypassrls/);
expect(sql).toMatch(/RAISE\s+EXCEPTION/i);
});
});
// ─────────────────────────────────────────────────────────────────
// REGRESSION TESTS — migrations v8 + v9 perf on duplicate-heavy tables
// ─────────────────────────────────────────────────────────────────
//
// Garry's production brain hit Supabase Management API's 60s ceiling because
// the DELETE...USING self-join in migrations v8 + v9 was O(n²) without an
// index on the dedup columns. The fix pre-creates a btree helper index
// before the DELETE, then drops it. These tests guard against any future
// change that re-introduces the missing helper index.
//
// Two-layer guard:
// 1. Structural — assert the migration SQL literally contains the helper
// CREATE INDEX + DROP INDEX (deterministic, fast, catches the regression
// even at 0-row scale where wall-clock can't distinguish O(n²) from O(1)).
// 2. Behavioral — populate 1000 duplicates and assert the migration completes
// under the wall-clock cap. Sanity check at small scale; the structural
// assertion is the real guard.
describe('migrations v8 + v9 — structural guard for helper-index fix', () => {
test('migration v8 SQL contains idx_links_dedup_helper CREATE+DROP around the DELETE', () => {
const v8 = MIGRATIONS.find(m => m.version === 8);
expect(v8).toBeDefined();
const sql = v8!.sql;
// The fix must: (a) create the helper btree, (b) DELETE...USING, (c) drop the helper, (d) add the unique constraint.
// If anyone reorders or removes the helper-index lines, this fails.
expect(sql).toContain('CREATE INDEX IF NOT EXISTS idx_links_dedup_helper');
expect(sql).toContain('ON links(from_page_id, to_page_id, link_type)');
expect(sql).toContain('DROP INDEX IF EXISTS idx_links_dedup_helper');
expect(sql).toContain('DELETE FROM links a USING links b');
expect(sql).toContain('ALTER TABLE links ADD CONSTRAINT links_from_to_type_unique');
// Order matters: CREATE INDEX before DELETE, DROP INDEX after DELETE, before ADD CONSTRAINT.
const createIdx = sql.indexOf('CREATE INDEX IF NOT EXISTS idx_links_dedup_helper');
const deleteUsing = sql.indexOf('DELETE FROM links a USING links b');
const dropIdx = sql.indexOf('DROP INDEX IF EXISTS idx_links_dedup_helper');
const addConstraint = sql.indexOf('ALTER TABLE links ADD CONSTRAINT links_from_to_type_unique');
expect(createIdx).toBeLessThan(deleteUsing);
expect(deleteUsing).toBeLessThan(dropIdx);
expect(dropIdx).toBeLessThan(addConstraint);
});
test('migration v9 SQL contains idx_timeline_dedup_helper CREATE+DROP around the DELETE', () => {
const v9 = MIGRATIONS.find(m => m.version === 9);
expect(v9).toBeDefined();
const sql = v9!.sql;
expect(sql).toContain('CREATE INDEX IF NOT EXISTS idx_timeline_dedup_helper');
expect(sql).toContain('ON timeline_entries(page_id, date, summary)');
expect(sql).toContain('DROP INDEX IF EXISTS idx_timeline_dedup_helper');
expect(sql).toContain('DELETE FROM timeline_entries a USING timeline_entries b');
expect(sql).toContain('CREATE UNIQUE INDEX IF NOT EXISTS idx_timeline_dedup');
const createHelper = sql.indexOf('CREATE INDEX IF NOT EXISTS idx_timeline_dedup_helper');
const deleteUsing = sql.indexOf('DELETE FROM timeline_entries a USING timeline_entries b');
const dropHelper = sql.indexOf('DROP INDEX IF EXISTS idx_timeline_dedup_helper');
const createUnique = sql.indexOf('CREATE UNIQUE INDEX IF NOT EXISTS idx_timeline_dedup');
expect(createHelper).toBeLessThan(deleteUsing);
expect(deleteUsing).toBeLessThan(dropHelper);
expect(dropHelper).toBeLessThan(createUnique);
});
});
// v0.14.1 — fix wave structural assertions (migrations renumbered from v12/v13 to
// v14/v15 after master merged budget_ledger (v12) + minion_quiet_hours_stagger (v13)).
describe('migrate v14 — pages_updated_at_index (handler-based, engine-aware)', () => {
const v14 = MIGRATIONS.find(m => m.version === 14);
test('v14 exists and uses a handler (not pure SQL) for engine-aware branching', () => {
expect(v14).toBeDefined();
expect(v14!.name).toBe('pages_updated_at_index');
expect(typeof v14!.handler).toBe('function');
expect(v14!.sql).toBe('');
});
test('v14 handler source contains CONCURRENTLY + invalid-index cleanup for Postgres branch', async () => {
const { readFileSync } = await import('fs');
const src = readFileSync('src/core/migrate.ts', 'utf-8');
const v14Start = src.indexOf("name: 'pages_updated_at_index'");
expect(v14Start).toBeGreaterThan(-1);
const v14Block = src.slice(v14Start, v14Start + 3000);
expect(v14Block).toContain('pg_index');
expect(v14Block).toContain('indisvalid');
expect(v14Block).toContain('DROP INDEX CONCURRENTLY IF EXISTS idx_pages_updated_at_desc');
expect(v14Block).toContain('CREATE INDEX CONCURRENTLY IF NOT EXISTS idx_pages_updated_at_desc');
// Order within the handler body: DROP IF EXISTS must precede CREATE IF NOT EXISTS,
// so a failed prior CONCURRENTLY build is cleaned before re-create. Anchor on the
// explicit "IF EXISTS" / "IF NOT EXISTS" phrases so the header doc-comment
// (which mentions both unqualified) doesn't fool the ordering assertion.
const dropIdx = v14Block.indexOf('DROP INDEX CONCURRENTLY IF EXISTS');
const createIdx = v14Block.indexOf('CREATE INDEX CONCURRENTLY IF NOT EXISTS');
expect(dropIdx).toBeLessThan(createIdx);
expect(v14Block).toContain('engine.kind');
});
});
describe('migrate v15 — minion_jobs_max_stalled_default_5', () => {
const v15 = MIGRATIONS.find(m => m.version === 15);
test('v15 exists and alters max_stalled default to 5', () => {
expect(v15).toBeDefined();
expect(v15!.name).toBe('minion_jobs_max_stalled_default_5');
expect(v15!.sql).toContain('ALTER TABLE minion_jobs ALTER COLUMN max_stalled SET DEFAULT 5');
});
test('v15 backfill UPDATE targets the correct non-terminal statuses', () => {
const sql = v15!.sql;
expect(sql).toContain(`'waiting'`);
expect(sql).toContain(`'active'`);
expect(sql).toContain(`'delayed'`);
expect(sql).toContain(`'waiting-children'`);
expect(sql).toContain(`'paused'`);
expect(sql).not.toContain(`'completed'`);
expect(sql).not.toContain(`'dead'`);
expect(sql).not.toContain(`'cancelled'`);
expect(sql).not.toContain(`'claimed'`);
expect(sql).not.toContain(`'running'`);
expect(sql).not.toContain(`'stalled'`);
});
test('v15 UPDATE clause has the < 5 guard so idempotent re-runs are no-ops', () => {
expect(v15!.sql).toContain('max_stalled < 5');
});
});
describe('migrate — runner behavioral (v14 handler + v15 backfill)', () => {
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
test('v14 created idx_pages_updated_at_desc on PGLite via handler branch', async () => {
const rows = await (engine as any).db.query(
`SELECT indexname FROM pg_indexes WHERE indexname = 'idx_pages_updated_at_desc'`
);
expect(rows.rows.length).toBe(1);
});
test('v15 backfilled any max_stalled=1 rows (smoke: schema default is 5)', async () => {
await (engine as any).db.exec(
`INSERT INTO minion_jobs (name, queue, status, max_stalled) VALUES ('test', 'default', 'waiting', 1)`
);
await (engine as any).db.exec(
`UPDATE minion_jobs SET max_stalled = 5
WHERE status IN ('waiting','active','delayed','waiting-children','paused')
AND max_stalled < 5`
);
const rows = await (engine as any).db.query(
`SELECT max_stalled FROM minion_jobs WHERE name = 'test'`
);
expect((rows.rows[0] as any).max_stalled).toBe(5);
await (engine as any).db.exec(
`UPDATE minion_jobs SET max_stalled = 5
WHERE status IN ('waiting','active','delayed','waiting-children','paused')
AND max_stalled < 5`
);
const rows2 = await (engine as any).db.query(
`SELECT max_stalled FROM minion_jobs WHERE name = 'test'`
);
expect((rows2.rows[0] as any).max_stalled).toBe(5);
});
});
describe('migrate: v8 (links_dedup) regression — must be fast on 1K duplicate rows', () => {
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
test('1000 duplicate links dedup completes in <90s and leaves table deduped', async () => {
// Set up: drop BOTH the old (v8) and new (v11) unique constraints so
// duplicates can be inserted, then reset version so v8 + v11 re-run.
// v11 replaces the v8 constraint name; we drop whichever is present.
const db = (engine as any).db;
await db.exec(`ALTER TABLE links DROP CONSTRAINT IF EXISTS links_from_to_type_unique`);
await db.exec(`ALTER TABLE links DROP CONSTRAINT IF EXISTS links_from_to_type_source_origin_unique`);
// Two pages so the FK is satisfied
await engine.putPage('p/from', { type: 'concept', title: 'F', compiled_truth: '', timeline: '' });
await engine.putPage('p/to', { type: 'concept', title: 'T', compiled_truth: '', timeline: '' });
const fromId = (await db.query(`SELECT id FROM pages WHERE slug = 'p/from'`)).rows[0].id;
const toId = (await db.query(`SELECT id FROM pages WHERE slug = 'p/to'`)).rows[0].id;
// Insert 1000 duplicates of the same (from, to, type) row
for (let i = 0; i < 1000; i++) {
await db.query(
`INSERT INTO links (from_page_id, to_page_id, link_type, context) VALUES ($1, $2, $3, $4)`,
[fromId, toId, 'mention', `dup-${i}`]
);
}
const beforeCount = (await db.query(`SELECT COUNT(*)::int AS c FROM links`)).rows[0].c;
expect(beforeCount).toBe(1000);
// Reset version to 7 so v8 + v9 + v10 + v11 re-run
await engine.setConfig('version', '7');
// Run migrations and assert wall-clock + correctness.
//
// Budget note: 90s, not 5s. The 5s budget guarded the original O(n²) v8
// regression in isolation when the chain only had ~8 migrations to run.
// Cathedral II (v0.21.0) added v27 + v28 (TSVECTOR column + GIN index +
// plpgsql trigger compile + 2 new tables w/ FK CASCADE), pushing the
// full v7→v28 chain to ~30-40s on PGLite WASM. The O(n²) regression
// would still take MINUTES on 1K duplicate rows (the original incident
// was multi-minute), so 90s preserves the gate intent while
// accommodating the longer schema chain.
const start = Date.now();
await runMigrations(engine);
const elapsedMs = Date.now() - start;
expect(elapsedMs).toBeLessThan(90_000);
const afterCount = (await db.query(`SELECT COUNT(*)::int AS c FROM links`)).rows[0].c;
expect(afterCount).toBe(1); // deduped to one row
// v11 replaces v8's constraint name. Assert the current (v11) constraint
// exists and the legacy v8 name is gone.
const constraints = (await db.query(`
SELECT conname FROM pg_constraint
WHERE conrelid = 'links'::regclass AND contype = 'u'
`)).rows;
expect(constraints.some((c: { conname: string }) => c.conname === 'links_from_to_type_source_origin_unique')).toBe(true);
expect(constraints.some((c: { conname: string }) => c.conname === 'links_from_to_type_unique')).toBe(false);
// Helper index was dropped after dedup
const helperIdx = (await db.query(`
SELECT indexname FROM pg_indexes
WHERE tablename = 'links' AND indexname = 'idx_links_dedup_helper'
`)).rows;
expect(helperIdx.length).toBe(0);
});
});
describe('migrate: v9 (timeline_dedup_index) regression — must be fast on 1K duplicate rows', () => {
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
test('1000 duplicate timeline entries dedup completes in <90s and leaves table deduped', async () => {
const db = (engine as any).db;
await db.exec(`DROP INDEX IF EXISTS idx_timeline_dedup`);
await engine.putPage('p/timeline', { type: 'concept', title: 'TL', compiled_truth: '', timeline: '' });
const pageId = (await db.query(`SELECT id FROM pages WHERE slug = 'p/timeline'`)).rows[0].id;
// Insert 1000 duplicates of the same (page_id, date, summary) row
for (let i = 0; i < 1000; i++) {
await db.query(
`INSERT INTO timeline_entries (page_id, date, source, summary, detail) VALUES ($1, $2::date, $3, $4, $5)`,
[pageId, '2024-01-15', `src-${i}`, 'Founded NovaMind', `detail-${i}`]
);
}
const beforeCount = (await db.query(`SELECT COUNT(*)::int AS c FROM timeline_entries`)).rows[0].c;
expect(beforeCount).toBe(1000);
await engine.setConfig('version', '7');
// Same 90s budget as the v8 link-dedup test for the same reason — see
// its "Budget note" comment. The 5s budget was for v9 in isolation;
// post-Cathedral II the chain runs through v28's TSVECTOR + GIN setup.
const start = Date.now();
await runMigrations(engine);
const elapsedMs = Date.now() - start;
expect(elapsedMs).toBeLessThan(90_000);
const afterCount = (await db.query(`SELECT COUNT(*)::int AS c FROM timeline_entries`)).rows[0].c;
expect(afterCount).toBe(1);
const uniqueIdx = (await db.query(`
SELECT indexname FROM pg_indexes
WHERE tablename = 'timeline_entries' AND indexname = 'idx_timeline_dedup'
`)).rows;
expect(uniqueIdx.length).toBe(1);
const helperIdx = (await db.query(`
SELECT indexname FROM pg_indexes
WHERE tablename = 'timeline_entries' AND indexname = 'idx_timeline_dedup_helper'
`)).rows;
expect(helperIdx.length).toBe(0);
});
});
// ─────────────────────────────────────────────────────────────────
// resolvePoolSize — GBRAIN_POOL_SIZE env override
// ─────────────────────────────────────────────────────────────────
//
// Guards the Bug 2 fix: users on constrained poolers (Supabase port 6543)
// must be able to cap the pool size via GBRAIN_POOL_SIZE. The default
// (10) is unchanged when the env var is unset.
describe('resolvePoolSize — env var + explicit override', () => {
const { resolvePoolSize } = require('../src/core/db.ts');
const original = process.env.GBRAIN_POOL_SIZE;
afterAll(() => {
if (original === undefined) delete process.env.GBRAIN_POOL_SIZE;
else process.env.GBRAIN_POOL_SIZE = original;
});
test('returns 10 default when unset and no explicit override', () => {
delete process.env.GBRAIN_POOL_SIZE;
expect(resolvePoolSize()).toBe(10);
});
test('reads GBRAIN_POOL_SIZE as an integer', () => {
process.env.GBRAIN_POOL_SIZE = '2';
expect(resolvePoolSize()).toBe(2);
process.env.GBRAIN_POOL_SIZE = '5';
expect(resolvePoolSize()).toBe(5);
});
test('ignores invalid GBRAIN_POOL_SIZE values', () => {
process.env.GBRAIN_POOL_SIZE = 'not-a-number';
expect(resolvePoolSize()).toBe(10);
process.env.GBRAIN_POOL_SIZE = '0';
expect(resolvePoolSize()).toBe(10);
process.env.GBRAIN_POOL_SIZE = '-1';
expect(resolvePoolSize()).toBe(10);
});
test('explicit argument wins over env + default', () => {
delete process.env.GBRAIN_POOL_SIZE;
expect(resolvePoolSize(3)).toBe(3);
process.env.GBRAIN_POOL_SIZE = '7';
expect(resolvePoolSize(3)).toBe(3);
});
});
// ─────────────────────────────────────────────────────────────────
// PR #356 regression guards — migration hardening
// ─────────────────────────────────────────────────────────────────
//
// These tests guard the codex + eng review findings folded into PR #356.
// If anyone refactors away the fixes, these catch it.
describe('PR #356 — LATEST_VERSION is max(versions), not array[-1]', () => {
test('LATEST_VERSION equals Math.max of all migration versions', () => {
// The bug it closes: MIGRATIONS is NOT stored in ascending order.
// array[-1] returned v16 when the true max was v23 — every Postgres
// user was told "up to date at v16" while 7 migrations were behind.
// This regression guard catches any refactor back to array[-1].
const expectedMax = Math.max(...MIGRATIONS.map(m => m.version));
expect(LATEST_VERSION).toBe(expectedMax);
});
test('Math.max is robust to any array order (structural check)', () => {
// The array ordering is not a guarantee we maintain. v0.18.0's v21/v22/v23
// sat out-of-order in the middle of the array (release-order reasons);
// v0.18.1's v24 was appended sensibly. Both need to work. The invariant
// is: LATEST_VERSION equals max across any ordering. Scramble and verify.
const scrambled = [...MIGRATIONS].sort(() => Math.random() - 0.5);
const scrambledMax = Math.max(...scrambled.map(m => m.version));
expect(scrambledMax).toBe(LATEST_VERSION);
// Guard against regression to array[-1]: the production source must use
// Math.max, never indexed access to the last element.
const src = readFileSync(resolve('src/core/migrate.ts'), 'utf-8');
expect(src).toMatch(/LATEST_VERSION\s*=\s*MIGRATIONS\.length[\s\S]{0,200}Math\.max/);
expect(src).not.toMatch(/MIGRATIONS\[MIGRATIONS\.length\s*-\s*1\]\.version/);
});
});
describe('PR #356 — getIdleBlockers pg_stat_activity shape', () => {
// Minimal mock of BrainEngine — we only need kind + executeRaw.
function mockEngine(kind: 'postgres' | 'pglite', rows: IdleBlocker[] | Error): BrainEngine {
return {
kind,
async executeRaw<T>(_sql: string, _params?: unknown[]): Promise<T[]> {
if (rows instanceof Error) throw rows;
return rows as unknown as T[];
},
} as unknown as BrainEngine;
}
test('returns [] on PGLite (no pool, no idle-in-tx concept)', async () => {
const engine = mockEngine('pglite', [{ pid: 1, state: 'idle in transaction', query_start: 'x', query: 'y' }]);
const blockers = await getIdleBlockers(engine);
expect(blockers).toEqual([]);
});
test('returns rows from pg_stat_activity on Postgres', async () => {
const fixture: IdleBlocker[] = [
{ pid: 12345, state: 'idle in transaction', query_start: '2026-04-22 06:00:00+00', query: 'BEGIN; SELECT * FROM pages' },
];
const engine = mockEngine('postgres', fixture);
const blockers = await getIdleBlockers(engine);
expect(blockers).toEqual(fixture);
});
test('returns [] (not throw) when pg_stat_activity query fails', async () => {
// Some managed Postgres tenants restrict pg_stat_activity. The helper
// should degrade gracefully: doctor --locks prints "no blockers" and
// migration pre-flight skips the warning.
const engine = mockEngine('postgres', new Error('permission denied'));
const blockers = await getIdleBlockers(engine);
expect(blockers).toEqual([]);
});
});
describe('PR #356 — 57014 catch path emits actionable 4-part diagnostic', () => {
test('runMigrations surfaces SQLSTATE 57014 with fix + verify steps', async () => {
// Mock an engine whose runMigration throws a code-57014 error
// once; the catch branch should log the 4-part structure AND
// rethrow preserving err.code so callers can re-branch.
//
// v0.30.1: retry wrapper now retries 3x on 57014. We set
// GBRAIN_MIGRATE_BACKOFF_MS=0 in test env to skip the 5s/15s wait
// so the test still completes within its budget. The final throw
// is a MigrationRetryExhausted whose message names the (mocked,
// empty) blocker set; the legacy err.code preservation is no longer
// primary surface — callers handle MigrationRetryExhausted explicitly.
const original = process.env.GBRAIN_MIGRATE_BACKOFF_MS;
process.env.GBRAIN_MIGRATE_BACKOFF_MS = '0';
const err = Object.assign(new Error('canceling statement due to statement timeout'), { code: '57014' });
let caughtCode: string | undefined;
let caughtName: string | undefined;
// getConfig returns '15' so pending starts with v16 (has sql content
// in the MIGRATIONS array). The first migration's SQL execution
// hits the 57014-throwing mock and fires the diagnostic branch.
const engine = {
kind: 'postgres' as const,
async getConfig(_k: string) { return '15'; },
async setConfig() {},
async executeRaw() { return []; },
async transaction<T>(fn: (e: BrainEngine) => Promise<T>): Promise<T> { return fn(engine as unknown as BrainEngine); },
async withReservedConnection() { throw new Error('unreached'); },
async runMigration() { throw err; },
} as unknown as BrainEngine;
const errSpy = spyOn(console, 'error').mockImplementation(() => {});
try {
await runMigrations(engine);
} catch (e: unknown) {
caughtCode = (e as { code?: string }).code;
caughtName = (e as { name?: string }).name;
}
if (original === undefined) delete process.env.GBRAIN_MIGRATE_BACKOFF_MS;
else process.env.GBRAIN_MIGRATE_BACKOFF_MS = original;
// v0.30.1: the throw is now a MigrationRetryExhausted (retry wrapper
// wraps the original err after 3 attempts). The original 57014 code
// is preserved on the `lastError` member of the envelope.
expect(caughtName).toBe('MigrationRetryExhausted');
// Defensive: legacy callers checking .code still work via `lastError`.
void caughtCode;
// Assert the diagnostic lines hit stderr with the agent-driven shape.
// v0.30.1: the header reads "exhausted retries" instead of
// "hit statement_timeout (SQLSTATE 57014)" because the retry wrapper
// wrapped the underlying timeout. The Cause/Fix/Verify body still fires
// when no blockers were detected (empty pg_stat_activity in the mock).
const msgs = errSpy.mock.calls.map(c => String(c[0]));
const joined = msgs.join('\n');
expect(joined).toContain('exhausted retries');
expect(joined).toContain('gbrain doctor --locks');
expect(joined).toContain('gbrain apply-migrations --yes');
expect(joined).toContain('Verify:');
expect(joined).toContain('gbrain doctor');
errSpy.mockRestore();
});
});
describe('PR #356 — apply-migrations pre-flight schema-version warning', () => {
test('source contains the pre-flight check branch before plan execution', () => {
// Structural check: the pre-flight block compares the engine's
// reported schema version against LATEST_VERSION and warns if
// behind. If someone removes this branch, users who run
// apply-migrations expecting it to handle schema migrations get
// the silent-gaslight experience from the field report.
const source = readFileSync(resolve('src/commands/apply-migrations.ts'), 'utf-8');
expect(source).toContain('LATEST_VERSION');
expect(source).toContain('Schema version');
expect(source).toContain('is behind latest');
});
});
describe('PR #356 + #363 — session timeouts applied via startup parameters', () => {
test('structural: setSessionDefaults exists for back-compat; resolveSessionTimeouts is the source of truth', () => {
// PR #356 introduced setSessionDefaults (post-pool SET).
// PR #363 superseded it with resolveSessionTimeouts (startup parameters,
// PgBouncer-transaction-mode-safe). The setSessionDefaults function is
// kept as a no-op shim for back-compat with existing call sites.
const dbSrc = readFileSync(resolve('src/core/db.ts'), 'utf-8');
const pgSrc = readFileSync(resolve('src/core/postgres-engine.ts'), 'utf-8');
// Helper still exists for back-compat
expect(dbSrc).toContain('export async function setSessionDefaults');
// The new source-of-truth function exists
expect(dbSrc).toContain('export function resolveSessionTimeouts');
expect(dbSrc).toContain('idle_in_transaction_session_timeout');
// Both connect paths call resolveSessionTimeouts() and feed it through
// postgres.js's connection option (startup parameters)
expect(dbSrc).toContain('resolveSessionTimeouts()');
expect(pgSrc).toContain('resolveSessionTimeouts()');
// setSessionDefaults still callable (no-op) so existing call sites
// don't break, but the SET command itself is gone — the work has
// already happened at connection startup time.
expect(pgSrc).toContain('db.setSessionDefaults');
// Critically: no SET idle_in_transaction in source — startup parameters
// are the durable mechanism for PgBouncer transaction mode.
const setMatches = dbSrc.match(/SET idle_in_transaction_session_timeout/g) || [];
expect(setMatches.length).toBe(0);
});
});
describe('PR #356 — non-transactional DDL runs via reserved connection', () => {
test('runMigrationSQL uses withReservedConnection for transaction:false branch', () => {
// The else-branch of runMigrationSQL (CREATE INDEX CONCURRENTLY etc.)
// must go through engine.withReservedConnection + SET statement_timeout,
// NOT engine.runMigration on the shared pool. Codex caught that the
// prior code left CONCURRENTLY DDL exposed to Supabase's 2-min timeout
// with no session-level override.
//
// v0.30.1: anchor on the exact function signature (open paren) so we
// don't match the new `runMigrationSQLWithRetry` wrapper that lives
// immediately above. The wrapper calls runMigrationSQL inside its retry
// body, so it must come BEFORE in the source — which is why a prefix
// match would catch the wrong function.
const source = readFileSync(resolve('src/core/migrate.ts'), 'utf-8');
const runFnIdx = source.indexOf('async function runMigrationSQL(');
expect(runFnIdx).toBeGreaterThan(-1);
const fnBody = source.slice(runFnIdx, runFnIdx + 2500);
expect(fnBody).toContain('withReservedConnection');
expect(fnBody).toContain("SET statement_timeout = '600000'");
});
});
describe('migration v31 — eval_capture_tables', () => {
test('exists with the expected name and is engine-specific (sqlFor)', () => {
const v31 = MIGRATIONS.find(m => m.version === 31);
expect(v31).toBeDefined();
expect(v31?.name).toBe('eval_capture_tables');
expect(v31?.sqlFor?.postgres).toBeDefined();
expect(v31?.sqlFor?.pglite).toBeDefined();
expect(v31?.sql).toBe('');
});
test('creates both eval_candidates and eval_capture_failures on both engines', () => {
const v31 = MIGRATIONS.find(m => m.version === 31)!;
for (const variant of ['postgres', 'pglite'] as const) {
const sql = v31.sqlFor![variant]!;
expect(sql).toContain('CREATE TABLE IF NOT EXISTS eval_candidates');
expect(sql).toContain('CREATE TABLE IF NOT EXISTS eval_capture_failures');
}
});
test('enforces CHECK length(query) <= 51200', () => {
const v31 = MIGRATIONS.find(m => m.version === 31)!;
for (const variant of ['postgres', 'pglite'] as const) {
expect(v31.sqlFor![variant]!).toContain('CHECK (length(query) <= 51200)');
}
});
test('enforces tool_name enum + reason enum', () => {
const v31 = MIGRATIONS.find(m => m.version === 31)!;
for (const variant of ['postgres', 'pglite'] as const) {
const sql = v31.sqlFor![variant]!;
expect(sql).toContain(`tool_name IN ('query', 'search')`);
expect(sql).toContain(`reason IN ('db_down', 'rls_reject', 'check_violation', 'scrubber_exception', 'other')`);
}
});
test('creates DESC indexes on both tables', () => {
const v31 = MIGRATIONS.find(m => m.version === 31)!;
for (const variant of ['postgres', 'pglite'] as const) {
const sql = v31.sqlFor![variant]!;
expect(sql).toContain('idx_eval_candidates_created_at');
expect(sql).toContain('idx_eval_capture_failures_ts');
expect(sql).toContain('created_at DESC');
expect(sql).toContain('ts DESC');
}
});
test('Postgres variant gates RLS on BYPASSRLS and fails loudly', () => {
const pgSql = MIGRATIONS.find(m => m.version === 31)!.sqlFor!.postgres!;
expect(pgSql).toContain('rolbypassrls');
expect(pgSql).toMatch(/IF NOT has_bypass/);
expect(pgSql).toMatch(/RAISE EXCEPTION[^;]*BYPASSRLS/);
expect(pgSql).toContain('ALTER TABLE eval_candidates ENABLE ROW LEVEL SECURITY');
expect(pgSql).toContain('ALTER TABLE eval_capture_failures ENABLE ROW LEVEL SECURITY');
});
test('PGLite variant has no RLS / no BYPASSRLS gate', () => {
const pgliteSql = MIGRATIONS.find(m => m.version === 31)!.sqlFor!.pglite!;
expect(pgliteSql).not.toContain('rolbypassrls');
expect(pgliteSql).not.toContain('ENABLE ROW LEVEL SECURITY');
});
test('LATEST_VERSION caught up to 31', () => {
expect(LATEST_VERSION).toBeGreaterThanOrEqual(31);
});
});
describe('migration v40 — pages_emotional_weight (v0.29)', () => {
// v0.29 ships off master. Master is at v39 (multimodal_dual_column_v0_27_1);
// v0.29 lands at v40. Idempotent ADD COLUMN IF NOT EXISTS, so brains that
// applied this at any prior number on a feature branch see v40 as new and
// run cleanly.
test('exists with the expected name', () => {
const v40 = MIGRATIONS.find(m => m.version === 40);
expect(v40).toBeDefined();
expect(v40?.name).toBe('pages_emotional_weight');
});
test('adds emotional_weight REAL NOT NULL DEFAULT 0.0 to pages', () => {
const v40 = MIGRATIONS.find(m => m.version === 40);
const sql = v40!.sql || '';
expect(sql).toContain('ALTER TABLE pages');
expect(sql).toContain('ADD COLUMN IF NOT EXISTS emotional_weight');
expect(sql).toContain('REAL');
expect(sql).toContain('NOT NULL DEFAULT 0.0');
});
test('does NOT create an idx_pages_emotional_weight index (eng review D6)', () => {
// Salience query orders by computed score, not raw weight; the index
// would never be used. Adding it later requires a separate migration.
const v40 = MIGRATIONS.find(m => m.version === 40);
const sql = v40!.sql || '';
expect(sql).not.toContain('idx_pages_emotional_weight');
expect(sql).not.toContain('CREATE INDEX');
});
test('LATEST_VERSION caught up to 40', () => {
expect(LATEST_VERSION).toBeGreaterThanOrEqual(40);
});
});
describe('migration v48 — takes_weight_round_to_grid (v0.32)', () => {
// v0.32 — Takes v2 wave. Renumbered from v46 → v48 after merging master's
// v0.31.3 wave (which claimed v46 with mcp_request_log_params_jsonb_normalize).
// Backfill the pre-v0.32 weight column to the 0.05 grid the engine layer
// (PR #795) enforces on insert. Cross-modal eval over 100K production
// takes flagged 0.74, 0.82-style values as false precision; this migration
// brings existing data to the grid that all new writes already match.
test('exists with the expected name', () => {
const v48 = MIGRATIONS.find(m => m.version === 48);
expect(v48).toBeDefined();
expect(v48?.name).toBe('takes_weight_round_to_grid');
});
test('uses transaction:false (codex review #2 — non-blocking, idempotent via WHERE)', () => {
// The original plan called this "mid-statement resume" — that was wrong.
// What transaction:false actually buys is freeing the migration runner
// from a long transaction so other gbrain processes can interleave.
const v48 = MIGRATIONS.find(m => m.version === 48);
expect(v48?.transaction).toBe(false);
});
test('UPDATE rounds weight to 0.05 grid', () => {
const v48 = MIGRATIONS.find(m => m.version === 48);
const sql = v48!.sql || '';
expect(sql).toContain('UPDATE takes');
expect(sql).toContain('ROUND(weight::numeric * 20) / 20');
});
test('WHERE uses tolerance comparison (REAL float32 noise vs 0.05 grid)', () => {
// Codex #2 idempotency correction + REAL/float32 implementation note:
// a naive `weight <> ROUND(...)` form fires every time because mixed
// REAL/NUMERIC comparison promotes weight to DOUBLE PRECISION first,
// surfacing ~1e-7 representation noise as inequality. The tolerance
// form (abs(...) > 0.001) catches genuinely off-grid values (the 0.05
// grid is 5e-2, far above 1e-3) while ignoring float32 round-trip noise.
const v48 = MIGRATIONS.find(m => m.version === 48);
const sql = v48!.sql || '';
expect(sql).toContain('WHERE');
expect(sql).toContain('abs(weight::numeric');
expect(sql).toContain('> 0.001');
});
test('IS NOT NULL guard (insurance against stale schema)', () => {
const v48 = MIGRATIONS.find(m => m.version === 48);
const sql = v48!.sql || '';
expect(sql).toContain('weight IS NOT NULL');
});
test('LATEST_VERSION caught up to 48', () => {
expect(LATEST_VERSION).toBeGreaterThanOrEqual(48);
});
});
describe('migration v49 — eval_takes_quality_runs (v0.32)', () => {
// v0.32 EXP-5 — Renumbered from v47 → v49 after merging master's v0.31.3 wave.
// DB-authoritative receipts table for `gbrain eval takes-quality`.
// Codex review #6 corrected the original two-phase split-brain plan: DB row
// is the source of truth (carries full receipt JSON), disk artifact is
// best-effort. The 4-sha unique key (corpus, prompt, models, rubric) makes
// re-running identical evals an `INSERT ... ON CONFLICT DO NOTHING` no-op.
test('exists with the expected name', () => {
const v49 = MIGRATIONS.find(m => m.version === 49);
expect(v49).toBeDefined();
expect(v49?.name).toBe('eval_takes_quality_runs');
});
test('creates the table with all 4 receipt sha columns + receipt_json JSONB', () => {
const v49 = MIGRATIONS.find(m => m.version === 49);
const sql = v49!.sql || '';
expect(sql).toContain('CREATE TABLE IF NOT EXISTS eval_takes_quality_runs');
expect(sql).toContain('receipt_sha8_corpus');
expect(sql).toContain('receipt_sha8_prompt');
expect(sql).toContain('receipt_sha8_models');
expect(sql).toContain('receipt_sha8_rubric');
expect(sql).toContain('receipt_json JSONB');
});
test('has 4-sha UNIQUE constraint (idempotent re-runs)', () => {
const v49 = MIGRATIONS.find(m => m.version === 49);
const sql = v49!.sql || '';
expect(sql).toContain('UNIQUE (receipt_sha8_corpus, receipt_sha8_prompt, receipt_sha8_models, receipt_sha8_rubric)');
});
test('verdict column has CHECK constraint for the 3 verdict values', () => {
const v49 = MIGRATIONS.find(m => m.version === 49);
const sql = v49!.sql || '';
expect(sql).toContain("CHECK (verdict IN ('pass','fail','inconclusive'))");
});
test('trend index orders by (rubric_version, created_at DESC)', () => {
// Codex review #3 — trend mode segregates by rubric_version + reads
// ordered DESC. Index shape must match the query shape exactly.
const v49 = MIGRATIONS.find(m => m.version === 49);
const sql = v49!.sql || '';
expect(sql).toContain('eval_takes_quality_runs_trend_idx');
expect(sql).toContain('(rubric_version, created_at DESC)');
});
test('LATEST_VERSION caught up to 49', () => {
expect(LATEST_VERSION).toBeGreaterThanOrEqual(49);
});
});
// ─────────────────────────────────────────────────────────────────
// PR #363 regression guards — session timeouts via startup parameters
// resolveSessionTimeouts — GBRAIN_*_TIMEOUT env overrides
// ─────────────────────────────────────────────────────────────────
//
// Guards: orphan pgbouncer backends that hold table locks for hours when
// the postgres.js client disconnects mid-transaction. Session-level
// statement_timeout + idle_in_transaction_session_timeout delivered as
// startup parameters kill those backends on the server side.
describe('resolveSessionTimeouts — env var overrides', () => {
const { resolveSessionTimeouts } = require('../src/core/db.ts');
const origStatement = process.env.GBRAIN_STATEMENT_TIMEOUT;
const origIdleTx = process.env.GBRAIN_IDLE_TX_TIMEOUT;
const origCheck = process.env.GBRAIN_CLIENT_CHECK_INTERVAL;
afterAll(() => {
const restore = (key: string, val: string | undefined) => {
if (val === undefined) delete process.env[key];
else process.env[key] = val;
};
restore('GBRAIN_STATEMENT_TIMEOUT', origStatement);
restore('GBRAIN_IDLE_TX_TIMEOUT', origIdleTx);
restore('GBRAIN_CLIENT_CHECK_INTERVAL', origCheck);
});
const resetEnv = () => {
delete process.env.GBRAIN_STATEMENT_TIMEOUT;
delete process.env.GBRAIN_IDLE_TX_TIMEOUT;
delete process.env.GBRAIN_CLIENT_CHECK_INTERVAL;
};
test('returns statement_timeout + idle_in_transaction defaults when unset', () => {
resetEnv();
const t = resolveSessionTimeouts();
expect(t.statement_timeout).toBe('5min');
// Default bumped from #363's original 2min to 5min on merge with v0.21.0's
// setSessionDefaults posture, to avoid regressing long embed/CREATE INDEX
// passes that have legitimate idle gaps.
expect(t.idle_in_transaction_session_timeout).toBe('5min');
// client_connection_check_interval is opt-in only (Postgres 14+)
expect(t.client_connection_check_interval).toBeUndefined();
});
test('env vars override the defaults', () => {
resetEnv();
process.env.GBRAIN_STATEMENT_TIMEOUT = '10min';
process.env.GBRAIN_IDLE_TX_TIMEOUT = '30s';
process.env.GBRAIN_CLIENT_CHECK_INTERVAL = '15s';
const t = resolveSessionTimeouts();
expect(t.statement_timeout).toBe('10min');
expect(t.idle_in_transaction_session_timeout).toBe('30s');
expect(t.client_connection_check_interval).toBe('15s');
});
test("'0' disables a specific GUC", () => {
resetEnv();
process.env.GBRAIN_STATEMENT_TIMEOUT = '0';
const t = resolveSessionTimeouts();
expect(t.statement_timeout).toBeUndefined();
expect(t.idle_in_transaction_session_timeout).toBe('5min');
});
test("'off' disables a specific GUC", () => {
resetEnv();
process.env.GBRAIN_IDLE_TX_TIMEOUT = 'off';
const t = resolveSessionTimeouts();
expect(t.statement_timeout).toBe('5min');
expect(t.idle_in_transaction_session_timeout).toBeUndefined();
});
test('all three can be disabled independently', () => {
resetEnv();
process.env.GBRAIN_STATEMENT_TIMEOUT = '0';
process.env.GBRAIN_IDLE_TX_TIMEOUT = 'off';
const t = resolveSessionTimeouts();
expect(Object.keys(t)).toHaveLength(0);
});
});