Files
gbrain/test/brain-writer-partial-scan.test.ts
T
41ab138462 v0.40.8.0 test: e2e + unit gap coverage + master flake root-cause fixes (#1313)
* fix(tests): root-cause two master test-infra flakes

gateway.test.ts: add afterAll(resetGateway) hook. The file's tests use
beforeEach(resetGateway) for per-test isolation, but the FINAL test was
leaving configureGateway({env: {OPENAI_API_KEY: 'openai-fake'}}) in the
gateway module state. Sibling files in the same bun shard (e.g.
test/ingestion/ingest-capture.test.ts) then triggered embed() against
the real OpenAI endpoint with the leaked fake key, wedging the shard
with 'Incorrect API key provided: openai-fake'.

header-transport.test.ts → .serial.test.ts: the file mutates RECIPES
(gateway's module-scoped recipe map) plus configureGateway, then asserts
fakeChatFetch was invoked. Under bun's intra-shard parallelism, sibling
files like test/ai/rerank.test.ts race the same state — chat would see
result.text === '[]' instead of 'ok' because another test called
resetGateway between this test's configureGateway and chat. Quarantining
as .serial.test.ts moves the file into the post-parallel serial pass
at --max-concurrency=1 per repo convention.

* refactor(doctor): extract buildChecks seam + behavioral coverage

src/commands/doctor.ts: extract buildChecks(engine, args, dbSource):
Promise<Check[]> from runDoctor. The check-building logic moves into
the new exported function; existing exported computeDoctorReport(checks)
at line 78 stays untouched. runDoctor becomes a thin wrapper:
buildChecks → computeDoctorReport → render + process.exit. All 10
process.exit sites stay in place. The two early-return paths drop
their inline outputResults+process.exit calls and return the partial
check list; the wrapper still produces identical observable output.

test/doctor-behavioral.test.ts (13 cases): pure pure-aggregation cases
pin computeDoctorReport math (3 fails → -60 points, score clamped at 0,
mixed outcome → unhealthy with fail dominating). Orchestrator cases
assert --fast flag honors the skip set, --json doesn't alter the list,
no-engine path returns partial without process.exit, and the snapshot
of load-bearing check names catches accidental drop-outs during
future refactors.

test/doctor-cli-smoke.serial.test.ts (1 case): subprocess smoke spawning
'bun run src/cli.ts doctor --json' against a fresh PGLite tempdir brain.
Catches render-path bugs that buildChecks-only tests miss — the class
the v0.38.2.0 partial-scan wave exposed. Quarantined as .serial because
PGLite write-locks don't play well with parallel runners; skippable via
GBRAIN_SKIP_SUBPROCESS_TESTS=1.

* feat(operations): trust-boundary contract test + filter-bypass shell guard

test/operations-trust-boundary.test.ts (14 cases): hybrid design per
plan D7. Pure assertions over all 74 ops cover the drift-detection
win (every op has a scope; every mutating op has a non-read scope;
hasScope(['read'], op.scope) correctly rejects 'admin' or 'write').
Plus the canonical filter contract: every localOnly: true op is
excluded from operations.filter(op => !op.localOnly). Plus targeted
handler-invocation cases for the two historically-broken HTTP-callable
classes: submit_job(name='shell', ctx.remote=true) MUST reject (F7b
HTTP MCP shell-job RCE class), and search_by_image(image_path,
ctx.remote=true) MUST reject (D18 P0 image-leak class). file_upload
and sync_brain are deliberately omitted from handler-invocation tests
because they're localOnly — calling their handlers directly tests an
impossible production path (codex CMT-3). All 7 localOnly ops are
snapshot-pinned by name to catch future flag-flips.

scripts/check-operations-filter-bypass.sh: greps src/ for any module
that imports the 'operations' value from core/operations.ts outside
the canonical filter site. Three import shapes detected: destructured,
aliased ('as ops'), namespace ('import * as'). Explicit allow-list of
10 known-safe importers with one-line rationale per entry. Plus a
filter-presence check on serve-http.ts that fails if the canonical
filter expression is refactored out. Codex /ship adversarial review
caught the original narrow regex missed aliased + namespace bypasses;
the expanded regex closes that class. Type-only imports of sibling
exports (sourceScopeOpts, OperationContext) are not flagged.

package.json: wires check:operations-filter-bypass into the verify
chain alongside check:jsonb and check:progress.

* refactor(cycle): export runPhaseLint+runPhaseBacklinks; add wrapper tests

src/core/cycle.ts: adds 'export' keyword to two existing phase
functions so behavioral tests can drive them without going through
runCycle's full setup cost. No body changes; no behavior change.
Documented as internal helpers exposed for test-only consumption —
downstream code should NOT take a dependency on them; existing
plan-eng-review D9 explicitly accepted the API-widening tax for
testability.

test/cycle-legacy-phases.test.ts (11 cases): combined file with two
describe blocks per plan D5 (DRY — shared setup, future phase
wrappers land as additional describes). Narrowed to result-mapping
+ error envelope per codex CMT-1 (legacy phases don't extend
BaseCyclePhase and don't take a progress reporter or AbortSignal
directly, so the contract surface is counter → status enum + try/catch
envelope). Cases: clean run → status='ok', partial fix → status='warn'
with dryRun in details, dry-run path doesn't write, throw-from-lib
→ status='fail' with envelope populated (no exception escape).
Verified runLintCore and runBacklinksCore both throw on missing dir,
so the throw-from-lib cases are deterministic.

* chore: bump version and changelog (v0.40.4.1)

E2E + unit test gap coverage wave. Closes 4 audit gaps with new
behavioral coverage (doctor orchestrator + subprocess smoke, operations
trust-boundary contract + filter-bypass guard, cycle phase wrapper
result-mapping). Verifies 3 audit gaps were already covered (ingestion
dedup/daemon/skillpack-load, phantom-redirect, ingestion test-harness).
Root-causes 2 pre-existing master flakes (gateway state leak, header-
transport cross-shard race). Files 5 follow-up TODOs from codex
adversarial-review findings.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: note v0.40.4.1 doctor.buildChecks + cycle phase exports in CLAUDE.md

Adds three Key-files entries pinning the v0.40.4.1 test-wave additions:
- doctor.ts extension: buildChecks seam + behavioral tests (13+1 cases)
- cycle.ts extension: runPhaseLint + runPhaseBacklinks exports (11 cases)
- operations-trust-boundary contract + check-operations-filter-bypass.sh

Regenerates llms-full.txt to match (CLAUDE.md edits require build:llms per
project rule, otherwise test/build-llms.test.ts fails in CI shard 1).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(tests): reset gateway in put_page write-through tests to skip embed in CI

CI failure mode: 9 tests in test/ingestion/put-page-write-through.test.ts
failed with `AIConfigError: [embed(zeroentropyai:zembed-1)] Unauthorized`
because put_page's handler at src/core/operations.ts:622 computes
`noEmbed = !isAvailable('embedding')`. When the gateway state has been
configured by a sibling test (or by the cli.ts module-load path reading
.env.testing) with a fake/stale ZEROENTROPY_API_KEY, isAvailable returns
true → put_page tries to embed → the real ZeroEntropy API returns 401.

Local dev passes because real ZE keys are present; CI doesn't have them.

Fix: call resetGateway() in beforeEach so isAvailable('embedding')
returns false → put_page's noEmbed path activates → no network call.
Also reset in afterAll to avoid leaking the cleared state to sibling
files in the same bun shard (the v0.40.4.1 gateway state-leak class
that motivated the earlier gateway.test.ts fix).

The test exercises write-through behavior, not embedding. No need for
a real or fake embed transport — bypass entirely.

* fix(tests): widen brain-writer partial-scan deadlines to absorb CI timing variance

CI failure: scanBrainSources partial-scan state > "hanging COUNT does not
exceed deadline — Promise.race timeout fires" failed once on a GitHub
Actions runner. The test asserted a 100ms deadline budget with a 500ms
bound; observed test duration was 187ms on CI (passes locally 20/20
runs at the original budget).

Root cause: Node.js timer drift under shard parallelism. The deadline
check at src/core/brain-writer.ts:503 uses strict `Date.now() > deadline`,
so when the setTimeout in Promise.race fires exactly at the boundary
(e.g. start+100ms when deadline is start+100ms), the post-await check
sees equality and skips the markRemainingSkipped branch. The test
also asserts elapsed < 500ms; CI overhead can push elapsed past that
bound when setTimeout drifts.

Fix: widen the deadline budget on both deadline-race tests proportionally
(keeps the same 2x ratio that proves "query exceeds deadline"). No
src/ changes — this is purely a test robustness widening.

  - "hanging COUNT" test: 100ms → 500ms deadline, 500ms → 2500ms bound
  - "slow COUNT" test: 50ms → 250ms deadline, 100ms → 500ms query delay

Verified locally: 20/20 stress runs at the widened budgets, no fails.

* fix(brain-writer): deadline check is >= not > (closes CI flake at boundary)

CI failure recurred: same "hanging COUNT does not exceed deadline" test
failed again at 588ms (past my previous 500ms deadline + 2500ms bound
widening). The root cause isn't test timing — it's an off-by-one in
the source.

src/core/brain-writer.ts had two deadline checks using strict `>`:
- line 445 (between-source abort)
- line 503 (post-COUNT-await re-check)

The Promise.race setTimeout resolves null at exactly `remainingMs` from
now, so post-await Date.now() OFTEN equals the deadline within
integer-ms precision. With `>`, the check skipped → scanOneSource ran
on the source whose budget had just been eaten → that source got
status='scanned' instead of 'skipped'. The test's `expect(firstSource
.status).toBe('skipped')` failed.

Fix: both checks now use `>=`. When Date.now() equals deadline exactly,
the budget IS exhausted — proceeding would let the next source eat its
own budget on top of what's already spent. Matches the boundary the
Promise.race's remainingMs <= 0 immediate-null path uses (line 481).

This is the real fix for the v0.40.x CI flakes; my earlier test-budget
widening papered over the symptom without closing the boundary. Kept
the wider 500ms deadline for headroom but added a comment pointing at
the operator fix as the load-bearing change.

Verified: 20/20 stress runs green locally after the operator fix.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-23 18:43:10 -07:00

210 lines
8.8 KiB
TypeScript

/**
* v0.38.2.0 — partial-scan state tests for scanBrainSources.
*
* Codex outside-voice C1 caught that AbortSignal.timeout cannot interrupt
* the sync walker (event loop blocked by readdirSync / readFileSync). The
* load-bearing interruption mechanism is `deadline?: number` checked
* inside scanOneSource's visit closure before parsing each file.
*
* These tests use `deadline: Date.now() - 1` (already-expired) to force
* partial state deterministically — NOT AbortSignal, which doesn't fire
* in the sync loop and would make this test flake or never trigger.
*
* They also cover codex C2 (`ok` after abort must be false even on clean
* prefix), C4 (`files_scanned` numerator surfaced), and the
* `aborted_at_source` field that lets doctor name the partial source.
*/
import { describe, expect, test, beforeAll, afterAll, beforeEach } from 'bun:test';
import { mkdtempSync, rmSync, writeFileSync, mkdirSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';
import { scanBrainSources } from '../src/core/brain-writer.ts';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';
let engine: PGLiteEngine;
let sourceA: string;
let sourceB: string;
let sourceC: string;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
// Three source dirs, each with a few markdown files.
sourceA = mkdtempSync(join(tmpdir(), 'partial-scan-a-'));
sourceB = mkdtempSync(join(tmpdir(), 'partial-scan-b-'));
sourceC = mkdtempSync(join(tmpdir(), 'partial-scan-c-'));
for (const dir of [sourceA, sourceB, sourceC]) {
mkdirSync(join(dir, 'people'), { recursive: true });
for (let i = 0; i < 5; i++) {
writeFileSync(
join(dir, 'people', `p${i}.md`),
`---\ntitle: Person ${i}\n---\n\nbody\n`,
);
}
}
});
afterAll(async () => {
await engine.disconnect();
for (const d of [sourceA, sourceB, sourceC]) {
rmSync(d, { recursive: true, force: true });
}
});
beforeEach(async () => {
await resetPgliteState(engine);
// Register all three sources for each test.
await engine.executeRaw(
`INSERT INTO sources (id, name, local_path) VALUES ('src-a', 'A', $1), ('src-b', 'B', $2), ('src-c', 'C', $3)`,
[sourceA, sourceB, sourceC],
);
});
describe('scanBrainSources partial-scan state', () => {
test('no deadline + no abort: every source scanned, partial=false, ok reflects grandTotal', async () => {
const report = await scanBrainSources(engine);
expect(report.partial).toBe(false);
expect(report.aborted_at_source).toBe(null);
expect(report.per_source.length).toBe(3);
for (const src of report.per_source) {
expect(src.status).toBe('scanned');
expect(src.files_scanned).toBe(5);
}
expect(report.total).toBe(0);
expect(report.ok).toBe(true);
});
test('deadline expired before any source starts: all three skipped', async () => {
const report = await scanBrainSources(engine, {
deadline: Date.now() - 1, // already expired
});
expect(report.partial).toBe(true);
expect(report.per_source.length).toBe(3);
for (const src of report.per_source) {
expect(src.status).toBe('skipped');
expect(src.files_scanned).toBe(0);
}
// ok must be false even though zero errors were found — partial state
// means the clean count can't speak for unscanned files (codex C2).
expect(report.ok).toBe(false);
});
test('after-abort ok field is false even on clean prefix (codex C2 regression guard)', async () => {
// Force the abort path: deadline already expired. Even though no
// errors found (because no files scanned), `ok` must reflect the
// partial-scan reality.
const report = await scanBrainSources(engine, {
deadline: Date.now() - 1,
});
expect(report.total).toBe(0);
expect(report.partial).toBe(true);
expect(report.ok).toBe(false);
});
test('files_scanned numerator populated on completed sources (codex C4 regression guard)', async () => {
const report = await scanBrainSources(engine);
for (const src of report.per_source) {
// Each source has 5 .md files under people/; all syncable.
expect(src.files_scanned).toBe(5);
}
});
test('dbPageCountForSource hook plumbed onto db_page_count; failure degrades to null', async () => {
let calls = 0;
const report = await scanBrainSources(engine, {
dbPageCountForSource: async (sourceId) => {
calls++;
if (sourceId === 'src-b') throw new Error('synthetic query failure');
return sourceId === 'src-a' ? 42 : 99;
},
});
expect(calls).toBe(3);
const a = report.per_source.find(r => r.source_id === 'src-a')!;
const b = report.per_source.find(r => r.source_id === 'src-b')!;
const c = report.per_source.find(r => r.source_id === 'src-c')!;
expect(a.db_page_count).toBe(42);
// Throw → null, no crash, scan continues.
expect(b.db_page_count).toBe(null);
expect(c.db_page_count).toBe(99);
// files_scanned numerator still populated regardless of denominator outcome.
expect(b.files_scanned).toBe(5);
});
// Codex adversarial #3 regression: when the outer-loop deadline check fires
// BEFORE any source starts, aborted_at_source MUST stamp the first
// would-have-been-scanned source so the doctor message can name it.
test('aborted_at_source stamped when deadline fires before any source starts', async () => {
const report = await scanBrainSources(engine, {
deadline: Date.now() - 1,
});
expect(report.partial).toBe(true);
// First source in deterministic order (sources ORDER BY id) is 'src-a'.
expect(report.aborted_at_source).toBe('src-a');
// Every source skipped, no scans started.
expect(report.per_source.every(r => r.status === 'skipped')).toBe(true);
});
// Codex adversarial #2 regression: a slow dbPageCountForSource that exceeds
// the deadline must NOT result in scanOneSource running and reporting
// status='partial' with files_scanned=0 (misleading — nothing was scanned).
// The post-await deadline re-check should mark the source as 'skipped'.
test('slow COUNT that exceeds deadline marks source skipped, not partial', async () => {
const start = Date.now();
// CI timing fix (same class as the "hanging COUNT" test below): widened
// 50→250ms deadline + 100→500ms simulated query delay. Keeps the 2x
// ratio that proves "query exceeds deadline" while giving CI runners
// enough headroom to schedule both setTimeout callbacks on-budget.
const report = await scanBrainSources(engine, {
deadline: start + 250,
dbPageCountForSource: async () => {
// Simulate a hung query: take 500ms (past the deadline).
await new Promise(resolve => setTimeout(resolve, 500));
return 42;
},
});
// The first source should be skipped (post-await deadline re-check fires),
// NOT marked partial with files_scanned=0.
const firstSource = report.per_source.find(r => r.source_id === 'src-a')!;
expect(firstSource.status).toBe('skipped');
expect(firstSource.files_scanned).toBe(0);
expect(report.partial).toBe(true);
expect(report.aborted_at_source).toBe('src-a');
});
// Codex adversarial #4 regression: even when dbPageCountForSource itself
// would hang indefinitely, the Promise.race against the deadline must
// resolve null and the scan must abort cleanly.
//
// Boundary fix (CI flake): brain-writer.ts post-await deadline check uses
// `>=` not `>`. The Promise.race setTimeout resolves at exactly
// `remainingMs` from now, so post-await Date.now() often equals deadline
// within integer-ms precision. Strict `>` missed those landings on CI
// runners and let scanOneSource run anyway, marking src-a as 'scanned'
// instead of 'skipped'. Test budget is generous enough to absorb runner
// timer drift; the real fix is the operator change in brain-writer.ts.
test('hanging COUNT does not exceed deadline — Promise.race timeout fires', async () => {
const start = Date.now();
const DEADLINE_MS = 500;
const BOUND_MS = 2500;
const report = await scanBrainSources(engine, {
deadline: start + DEADLINE_MS,
dbPageCountForSource: () => {
// Never resolves — would hang forever without the deadline race.
return new Promise<number | null>(() => {});
},
});
const elapsed = Date.now() - start;
expect(elapsed).toBeLessThan(BOUND_MS);
expect(report.partial).toBe(true);
const firstSource = report.per_source.find(r => r.source_id === 'src-a')!;
expect(firstSource.status).toBe('skipped');
// Skipped sources never get db_page_count set — they weren't attempted.
// (Either null from the race or undefined from never reaching the DB
// path; both express "no denominator available" honestly.)
expect(firstSource.db_page_count == null).toBe(true);
});
});