Files
gbrain/test/facts-backstop.test.ts
T
0c6fcab555 v0.35.4.0 fix(doctor,entities): supervisor crash classification + bare-name resolver + 58x perf + stub guard observability (#1085)
* fix(doctor,entities): supervisor crash classification + bare-name resolver + stub guard

- doctor.ts/jobs.ts: classify worker exits with code !== 0 as real crashes
  vs code === 0 clean restarts (separate counter); fixes false-positive
  WARN on healthy supervisors
- entities/resolve.ts: prefix-expansion step between fuzzy match and
  slugify fallback catches bare first names that score too low on pg_trgm;
  picks highest-connection candidate as tiebreaker
- facts/fence-write.ts: stub-creation guard refuses to spawn unprefixed
  entity pages at brain root
- facts/backstop.ts: routes stubGuardBlocked facts to engine.insertFact
  so the fact still persists even when no markdown file is created
- docs/issues/doctor-auto-heal-and-scoring.md: spec for follow-up doctor
  health-score improvements
- .gitignore: guard reports/network-intelligence/ (private brain exports)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(privacy): scrub real names from entity-resolve test fixtures and JSDoc

Replace YC partner names with placeholders per CLAUDE.md privacy rule:
alice-example, bob-example, charlie-example, dave-example. Stripe and
Stripe Atlas retained (allowed household brands; exercises the two-word
company-prefix case).

Test semantics preserved:
- Alice / Dave: single-match cases
- Bob / Charlie: multi-match tiebreaker cases (winner has more chunks)

All 13 entity-resolve cases pass with the scrubbed fixtures.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(supervisor): extract classifyWorkerExit() helper (DRY)

Three call sites were inline-classifying worker exits: supervisor's
restart policy (child-worker-supervisor.ts:291), doctor's supervisor
check (doctor.ts:1016), and jobs supervisor status (jobs.ts:806). Same
rule, three copies — drift risk if one is updated without the others.

Extract to src/core/minions/exit-classification.ts as a pure function.
Signature consumes audit-JSON shape ({ code: number | null }) so doctor
and jobs (which read serialized events from JSONL) and supervisor (which
reads Node's exit callback) call the same function. Helper's classification
rule: code === 0 → clean_exit, everything else (non-zero, null, undefined,
missing) → crash. Default-to-crash prevents corrupted rows from silently
demoting into the clean-restart bucket.

5 hermetic unit tests (test/exit-classification.test.ts) pin all edge cases.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(facts): audit + sunset comment for stub-guard fires

Wire telemetry into the v0.34.5 stub-guard at fence-write.ts:190. Every
guard fire now appends a JSONL line to
~/.gbrain/audit/stub-guard-YYYY-Www.jsonl with {ts, slug, source_id,
fact_count}. Operator visibility for the sunset criterion: when the new
audit log reads <5 hits/week for 3 consecutive weeks on production
brains, the prefix-expansion in resolveEntitySlug is sufficient and the
guard can be removed in v0.36.

Reader (readRecentStubGuardEvents) deliberately diverges from
supervisor-audit.ts:readSupervisorEvents — it reads BOTH the current AND
previous ISO-week file before filtering by ts. supervisor-audit's reader
only reads the current week, which loses 24h-window correctness across
Monday 00:00 UTC (a Sunday 23:55 event lives in last week's file). The
2-file read costs nothing and makes the window actually 24h.

9 hermetic unit tests pin filename math, the writer's
swallows-errors contract, the cross-week-boundary read, sort order,
missing-file behavior, and malformed-row tolerance. The cross-week test
is the regression guard: if a future refactor copies the supervisor's
single-file pattern, that test fails.

Follow-up TODO (not in this PR): fix readSupervisorEvents to use the
same 2-file pattern. The new stub-guard reader becomes the canonical
template to copy back.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(doctor): stub_guard_24h check surfaces resolver gaps

Adds a new doctor check that reads ~/.gbrain/audit/stub-guard-YYYY-Www.jsonl
(via the dual-week-aware reader from T8) and surfaces the 24h fire count.
WARN at >10 fires — at that rate the prefix-expansion in resolveEntitySlug
is probably missing a case (typo prefix, alias, non-Latin script) and
operators should grep the audit log for the offending slugs. Below the
threshold but non-zero shows as OK with a count, so operators can watch
the v0.36 sunset criterion (<5/week for 3 weeks → guard can be removed).
Zero hits emits no check, keeping the doctor output clean on healthy
brains.

5 source-grep regression tests pin the contract: check name, WARN
threshold, fix hint mentions the audit log + the resolver function name,
reader is the dual-week-aware variant (NOT the supervisor-audit single-
week pattern), and zero-hits stays silent.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(facts): pin stub-guard contract at writeFactsToFence + backstop layers

- fence-write.test.ts: 3 new cases for the v0.34.5 stub guard. Bare slugs
  return {inserted: 0, stubGuardBlocked: true, ids: []} and create no
  file/.tmp at brain root. Prefixed slugs bypass the guard (regression
  guard against accidentally inverting the slug.includes('/') check).
  Empty facts array short-circuits before the guard fires.
- facts-backstop.test.ts: 1 new case for the end-to-end routing. A
  bare-name LLM extraction resolves through to a bare slug, hits the
  guard, and lands in the facts table via engine.insertFact (DB-only).
  No phantom .md file; entity_slug stores the bare slug;
  source_markdown_slug is null. This is the routing contract Codex
  flagged as a "split-brain" data shape — the test pins the by-design
  behavior so a future refactor can't silently drop these facts.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(supervisor): pin classifyWorkerExit consumer wire-up + regressions

12 new cases on top of the 5 helper unit tests:
- doctor.ts / jobs.ts / child-worker-supervisor.ts each import the helper
- All three call classifyWorkerExit at least once
- doctor.ts and jobs.ts no longer carry the pre-T7 inline filter
- supervisor uses the helper result to choose the clean_exit branch
- audit-event shape round-trip: code=0 → clean_exit, code=1 → crash,
  code=null+SIGKILL → crash (catches future shape changes)

The regression guards (3) and the wire-up checks (6) close the gap that
motivated T7 in the first place: if a future change accidentally re-inlines
the filter or shifts the audit event shape, the test fails before
production sees the silent divergence.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(entities): correlated subqueries scoped to slug-LIKE candidates

Replace the derived-table JOIN shape in tryPrefixExpansion with
correlated subqueries. The pre-fix SQL did

  LEFT JOIN (SELECT to_page_id, COUNT(*) FROM links GROUP BY to_page_id) li ON ...

which forced the planner to aggregate the entire links + content_chunks
tables on every prefix-expansion call — O(N) per call where N is total
links/chunks in the brain. On a 100K-link / 50K-chunk brain that's slow
enough to bottleneck fact-extraction.

New shape uses correlated subqueries:

  (SELECT COUNT(*) FROM links WHERE to_page_id = p.id)
    + (SELECT COUNT(*) FROM links WHERE from_page_id = p.id)
    + (SELECT COUNT(*) FROM content_chunks WHERE page_id = p.id)

The slug LIKE filter is already selective (typical brain has 0-5 pages
per prefix), so the three subqueries run N≈3 times per matched row
against the existing indexes on links.to_page_id, links.from_page_id,
and content_chunks.page_id. Behavior preserved: 13/13 entity-resolve
tests pass (single-match + multi-match tiebreaker + edge cases).

Codex's outside-voice review caught the dead-end design that an earlier
draft of this plan proposed (a CTE with `LIMIT 50` candidate cap — would
have excluded correct high-connection candidates if their slug sorted
late). Correlated subqueries without a candidate cap are the cleaner
shape that lets the LIKE filter do the bounding work.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(entities): perf regression guard for prefix-expansion (58x speedup)

Hermetic PGLite benchmark with 5K pages + 50K links + 25K chunks. Runs
the pre-T12 derived-table shape and the new correlated-subquery shape
side-by-side against the same fixture, asserts NEW >= 5x faster than OLD.
Baseline-ratio, not absolute wall-clock — different machines / Bun
versions / CI load can shift absolute timings by 10x without indicating
a real regression, but the SHAPE difference between "aggregate the full
tables" and "correlated subquery per candidate" is what we care about.

Measured: old_median=18.16ms, new_median=0.31ms, speedup=58.22x.
The 5x assertion has plenty of headroom.

The OLD SQL is embedded verbatim as the regression baseline. If a future
refactor re-introduces full-table aggregation (LEFT JOIN against
SELECT...GROUP BY over the whole links or content_chunks table), the
test fails. PGLite-only — Postgres planner can shape derived-table
JOINs differently enough that the 5x ratio could be noise on a 5K-page
fixture. The structural correctness of the rewrite is the same on both;
this is purely a planner-shape regression guard.

.slow.test.ts suffix keeps it out of the fast loop (run via
`bun run test:slow`).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.35.2.0)

Wave content:
- Privacy scrub: PII rebuilt out of branch history; real names → placeholders
- Bug fix: doctor + jobs no longer count clean worker exits as crashes
- Bug fix: entity resolver prefix-expansion catches bare first names
- DRY refactor: classifyWorkerExit() helper (one rule, 3 call sites)
- Observability: stub_guard_24h doctor check + ISO-week audit log
- Perf: 58x speedup on tryPrefixExpansion query shape

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: rebump v0.35.2.0 → v0.35.4.0 + scrub TODOS.md privacy violation

VERSION/package.json/CHANGELOG header rebumped to v0.35.4.0 per user
request (queue allocation). TODOS.md rephrased to not literally name
the banned private-agent string — that was the CI failure root cause
on the v0.35.2.0 push. CHANGELOG.md is on check-privacy.sh's allow-list
(meta-documentation exception); TODOS.md is not.

CI re-runs against this commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 08:51:33 -07:00

304 lines
12 KiB
TypeScript

/**
* v0.31.2 — runFactsBackstop pipeline tests.
*
* Pins the helper's full contract: eligibility/kill-switch gates,
* mode dispatch (queue vs inline), notability filter, dedup fast-path,
* abort propagation, and skipped envelope shapes.
*
* Real PGLite engine (in-memory, no DATABASE_URL). LLM is stubbed via
* __setChatTransportForTests so tests are deterministic + fast.
*/
import { describe, test, expect, beforeAll, afterAll, afterEach } from 'bun:test';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { runFactsBackstop } from '../src/core/facts/backstop.ts';
import type { FactsBackstopCtx } from '../src/core/facts/backstop.ts';
import {
__setChatTransportForTests,
resetGateway,
type ChatResult,
} from '../src/core/ai/gateway.ts';
import { __resetFactsQueueForTests } from '../src/core/facts/queue.ts';
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
afterEach(() => {
__setChatTransportForTests(null);
resetGateway();
__resetFactsQueueForTests();
});
const LONG_BODY = 'this is a real meeting note longer than 80 chars '.repeat(3);
function chatStub(facts: Array<{ fact: string; kind: string; notability: 'high' | 'medium' | 'low'; entity?: string | null }>) {
__setChatTransportForTests(async (): Promise<ChatResult> => ({
text: JSON.stringify({
facts: facts.map(f => ({
fact: f.fact,
kind: f.kind,
entity: f.entity ?? null,
confidence: 1.0,
notability: f.notability,
})),
}),
blocks: [],
stopReason: 'end',
usage: { input_tokens: 0, output_tokens: 0, cache_read_tokens: 0, cache_creation_tokens: 0 },
model: 'test:stub',
providerId: 'test',
}));
}
function makeCtx(overrides: Partial<FactsBackstopCtx> = {}): FactsBackstopCtx {
return {
engine,
sourceId: 'default',
sessionId: null,
source: 'mcp:put_page',
...overrides,
};
}
const meetingPage = (slug = 'meetings/test-' + Math.random().toString(36).slice(2, 9)) => ({
slug,
type: 'meeting' as const,
compiled_truth: LONG_BODY,
frontmatter: {} as Record<string, unknown>,
});
describe('runFactsBackstop — eligibility + kill-switch gates', () => {
test('skips with extraction_disabled when kill-switch off', async () => {
await engine.setConfig('facts.extraction_enabled', 'false');
const r = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'inline' }));
expect(r.mode).toBe('inline');
if (r.mode === 'inline') expect(r.skipped).toBe('extraction_disabled');
await engine.setConfig('facts.extraction_enabled', 'true');
});
test('skips with eligibility_failed:<reason> when subagent namespace', async () => {
chatStub([]);
const r = await runFactsBackstop(
{ ...meetingPage('wiki/agents/scratch/foo'), type: 'meeting' },
makeCtx({ mode: 'inline' }),
);
expect(r.mode).toBe('inline');
if (r.mode === 'inline') expect(r.skipped).toBe('eligibility_failed:subagent_namespace');
});
test('skips with eligibility_failed:dream_generated when frontmatter set', async () => {
chatStub([]);
const page = meetingPage();
page.frontmatter = { dream_generated: true };
const r = await runFactsBackstop(page, makeCtx({ mode: 'inline' }));
expect(r.mode).toBe('inline');
if (r.mode === 'inline') expect(r.skipped).toBe('eligibility_failed:dream_generated');
});
});
describe('runFactsBackstop — mode: inline', () => {
test('inserts the LLM-extracted facts and returns counts', async () => {
chatStub([
{ fact: 'mode-inline-event-1', kind: 'event', notability: 'high', entity: 'people/alice-example' },
{ fact: 'mode-inline-event-2', kind: 'event', notability: 'medium', entity: 'people/alice-example' },
]);
const r = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'inline' }));
expect(r.mode).toBe('inline');
if (r.mode === 'inline') {
expect(r.inserted).toBe(2);
expect(r.duplicate).toBe(0);
expect(r.fact_ids.length).toBe(2);
}
});
test('notabilityFilter=high-only drops MEDIUM + LOW from the insert path', async () => {
chatStub([
{ fact: 'high-only-1', kind: 'event', notability: 'high', entity: 'people/bob-test' },
{ fact: 'high-only-2-skip', kind: 'event', notability: 'medium', entity: 'people/bob-test' },
{ fact: 'high-only-3-skip', kind: 'event', notability: 'low', entity: 'people/bob-test' },
]);
const r = await runFactsBackstop(
meetingPage(),
makeCtx({ mode: 'inline', notabilityFilter: 'high-only' }),
);
expect(r.mode).toBe('inline');
if (r.mode === 'inline') {
expect(r.inserted).toBe(1);
expect(r.fact_ids.length).toBe(1);
}
});
test('source string lands on the inserted row', async () => {
const sessionId = 'source-pin-session-' + Math.random().toString(36).slice(2, 9);
chatStub([{ fact: 'source-pin-fact', kind: 'fact', notability: 'medium', entity: null }]);
const r = await runFactsBackstop(
meetingPage(),
makeCtx({ mode: 'inline', source: 'sync:import', sessionId }),
);
expect(r.mode).toBe('inline');
if (r.mode === 'inline' && r.fact_ids.length > 0) {
// Query by source_session (deterministic, no resolveEntitySlug rewrite).
const rows = await engine.listFactsBySession('default', sessionId);
const ours = rows.find(x => x.id === r.fact_ids[0]);
expect(ours?.source).toBe('sync:import');
expect(ours?.source_session).toBe(sessionId);
}
});
test('empty extraction → zero counts', async () => {
chatStub([]);
const r = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'inline' }));
expect(r.mode).toBe('inline');
if (r.mode === 'inline') {
expect(r.inserted).toBe(0);
expect(r.duplicate).toBe(0);
expect(r.fact_ids.length).toBe(0);
}
});
test('aborted before LLM call → zero counts, no throw', async () => {
chatStub([]);
const ac = new AbortController();
ac.abort();
const r = await runFactsBackstop(
meetingPage(),
makeCtx({ mode: 'inline', abortSignal: ac.signal }),
);
expect(r.mode).toBe('inline');
if (r.mode === 'inline') expect(r.inserted).toBe(0);
});
});
describe('runFactsBackstop — mode: queue', () => {
test('default mode is queue; returns enqueued: true', async () => {
chatStub([{ fact: 'queue-default-1', kind: 'event', notability: 'high', entity: 'people/queue-test' }]);
const r = await runFactsBackstop(meetingPage(), makeCtx()); // no mode → default 'queue'
expect(r.mode).toBe('queue');
if (r.mode === 'queue') {
expect(r.enqueued).toBe(true);
expect(r.queueDepth).toBeGreaterThanOrEqual(0);
}
});
test('explicit mode=queue returns immediately with enqueued: true', async () => {
chatStub([{ fact: 'queue-explicit', kind: 'event', notability: 'high', entity: 'people/queue-explicit' }]);
const r = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'queue' }));
expect(r.mode).toBe('queue');
if (r.mode === 'queue') expect(r.enqueued).toBe(true);
});
test('queue mode with extraction_disabled returns enqueued: false + skipped reason', async () => {
await engine.setConfig('facts.extraction_enabled', 'false');
const r = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'queue' }));
expect(r.mode).toBe('queue');
if (r.mode === 'queue') {
expect(r.enqueued).toBe(false);
expect(r.skipped).toBe('extraction_disabled');
}
await engine.setConfig('facts.extraction_enabled', 'true');
});
});
describe('runFactsBackstop — dedup fast-path', () => {
test('two near-identical inserts: second comes back as duplicate', async () => {
// Insert first via inline mode; we'll re-fire with the same fact text
// and rely on cosine ≥ 0.95 when the embedding matches. The B1 smoke
// path's stubbed transport means embedding stays null (gateway not
// configured), so the dedup path needs candidates with embeddings.
// Skip the embedding-match assertion here and pin the no-dedup path:
chatStub([
{ fact: 'distinct-fact-A', kind: 'event', notability: 'high', entity: 'people/dedup-a' },
]);
const r1 = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'inline' }));
expect(r1.mode).toBe('inline');
if (r1.mode === 'inline') expect(r1.inserted).toBe(1);
// Without embeddings present, dedup short-circuits; the second call
// inserts a new row (insertFact does no further dedup unless caller
// passes supersedeId). That's the contract for queue+inline backstop.
chatStub([
{ fact: 'distinct-fact-A', kind: 'event', notability: 'high', entity: 'people/dedup-a' },
]);
const r2 = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'inline' }));
expect(r2.mode).toBe('inline');
if (r2.mode === 'inline') {
// Without embeddings the dedup fast-path can't fire; second insert lands.
// Real production has embeddings via gateway — covered by E2E in a
// future test that points at a configured chat+embed gateway.
expect(r2.inserted + r2.duplicate).toBe(1);
}
});
});
describe('runFactsBackstop — stub guard routing (v0.34.5)', () => {
test('bare-name entity routes to legacy DB-only path (no phantom page)', async () => {
// Set up: configure default source with a real local_path so the
// backstop reaches Phase 5 (fence write) instead of Phase 4 (legacy).
// This is the scenario where the stub guard actually fires.
const { mkdtempSync, rmSync, existsSync } = await import('node:fs');
const { tmpdir } = await import('node:os');
const { join } = await import('node:path');
const brainDir = mkdtempSync(join(tmpdir(), 'backstop-stub-guard-'));
try {
// Point the default source at the tempdir so localPath is non-null.
// eslint-disable-next-line @typescript-eslint/no-explicit-any
await (engine as any).db.query(
`UPDATE sources SET local_path = $1 WHERE id = 'default'`,
[brainDir],
);
// Stub the chat to return a fact with a bare-name entity. The
// resolver will:
// 1. exact slug match → miss (no 'noresolvable' row)
// 2. fuzzy match → miss (no title contains noresolvable)
// 3. prefix expansion → miss (no people/noresolvable-* rows)
// 4. slugify fallback → 'noresolvable' (bare)
// The bare slug then trips the stub guard in writeFactsToFence,
// which returns stubGuardBlocked: true, and backstop routes the
// fact to engine.insertFact (DB-only).
chatStub([
{ fact: 'said hello at the meeting', kind: 'event', notability: 'high', entity: 'noresolvable' },
]);
const r = await runFactsBackstop(meetingPage(), makeCtx({ mode: 'inline' }));
expect(r.mode).toBe('inline');
if (r.mode === 'inline') {
// The fact MUST be persisted via the DB-only fallback, not dropped.
expect(r.inserted).toBe(1);
expect(r.fact_ids.length).toBe(1);
// No phantom file at the brain root (this is the whole point of the guard).
expect(existsSync(join(brainDir, 'noresolvable.md'))).toBe(false);
// The fact is in the DB with the bare entity_slug. Query directly to
// confirm — the routing is the contract under test.
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const rows = await (engine as any).db.query(
`SELECT entity_slug, fact, source_markdown_slug FROM facts WHERE id = $1`,
[r.fact_ids[0]],
);
expect(rows.rows[0].entity_slug).toBe('noresolvable');
expect(rows.rows[0].fact).toBe('said hello at the meeting');
// source_markdown_slug is the fence-tracking column; under DB-only
// fallback it stays null (no .md file backs the row).
expect(rows.rows[0].source_markdown_slug).toBeNull();
}
} finally {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
await (engine as any).db.query(`UPDATE sources SET local_path = NULL WHERE id = 'default'`);
rmSync(brainDir, { recursive: true, force: true });
}
});
});