Files
gbrain/test/query-cache-gate.test.ts
T
9bf96db807 v0.42.51.0 fix(sync): contention-free clock + checkpoint integrity + honest sync freshness (#2255)
* fix(sync): contention-free page-generation clock — sequence swap

The page-generation clock backed the query-cache Layer-1 bookmark via a
FOR EACH STATEMENT trigger running `UPDATE page_generation_clock SET
value=value+1 WHERE id=1`. That took a transaction-length RowExclusiveLock
on one tuple, so every concurrent page writer serialized on the prior
writer's COMMIT — sync ran at ~0.8 cores regardless of worker count.

Swap to a SEQUENCE bumped by nextval() (a microsecond LWLock, never a row
lock). The clock's only contract is monotonic advancement on any page
INSERT/UPDATE/DELETE; last_value is non-transactional, so rolled-back or
concurrent-uncommitted writers only OVER-invalidate the cache (lose a hit),
never serve stale.

- migration v118: CREATE SEQUENCE + load-bearing 2-arg setval (is_called=
  true, floor 1, seeded >= old clock and MAX(generation)) + repoint the
  trigger function body + DELETE query_cache so no old-clock bookmark
  survives the swap. v107 left immutable.
- query-cache-gate.ts: 3 readers -> SELECT last_value FROM page_generation_clock_seq.
- schema.sql + pglite-schema.ts (+ regenerated schema-embedded.ts) ship the
  sequence on fresh install; table + trigger names retained.
- tests: clockValue reads last_value; mechanism proof (trigger fn uses
  nextval not the row UPDATE); rollback-advances-clock safety pin; real
  PGLite sequence round-trip (is_called gotcha); shape test requires _seq.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sync): op_checkpoints array-shape guard — CHECK + repair + defensive loader

completed_keys is JSONB and the checkpoint loader runs
jsonb_array_elements_text over it. A non-array (scalar) value makes that
throw "cannot extract elements from a scalar", which takes down the whole
UNION load — including the valid op_checkpoint_paths child rows — and loses
all checkpoint progress for that key. No current writer produces a scalar,
but an older binary / external script / future bug could.

Make the corruption class structurally impossible and self-healing:

- migration v119: LOCK TABLE (so an out-of-band scalar can't land between
  repair and constrain; no-op on single-connection PGLite), repair any
  pre-existing scalar to '[]' (op_checkpoint_paths child rows are the
  append-only source of truth, so the reset loses nothing), then add the
  named CHECK (jsonb_typeof(completed_keys) = 'array') via a pg_constraint
  IF NOT EXISTS guard. A DB-enforced always-on guard — the correct pattern
  vs a migration verify-hook, which never runs on already-stamped brains.
- schema.sql + pglite-schema.ts (+ regenerated schema-embedded.ts) ship the
  same NAMED inline CHECK so fresh installs match migrated brains and v119
  skips the duplicate.
- op-checkpoint.ts loader: gate the legacy arm on jsonb_typeof = 'array' so
  a scalar parent is skipped (children still load) instead of throwing the
  whole union, and log a specific corruption warning when one is seen.
- tests: CHECK rejects a scalar (exactly one constraint, no blob+migration
  dupe); loader survives a scalar parent and returns the children; v119
  repair converts a scalar to '[]'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(doctor): report actively-running sync via live lock, not stale freshness

A slow source that makes partial progress every cycle but never fully
completes used to read as permanently "stale" / "never synced" because
last_sync_at only advances on a full successful sync. The naive fix
(treat recent checkpoint banking as "in progress") is unsafe: a blocked
sync banks the good files then writes no anchor, so banking can't tell
in-progress from wedged.

Use the only honest signal: a LIVE, non-expired per-source sync lock
(inspectLock + syncLockId against gbrain_cycle_locks). Every non-skipLock
sync holds it and refreshes it; a blocked/failed sync's process has exited
(no lock row) and a wedged holder stops refreshing (TTL lapses), so either
correctly falls through to the stale path and is NEVER masked. An
actively-syncing source (including a never-synced source doing its first
sync) counts as synced_recently, preserving the pinned 3-bucket invariant.
The lock lookup reuses doctor's existing dynamic db-lock import and swallows
any throw (stub engine, pre-lock-table brain) to false, so it can only ADD
an in-progress verdict, never suppress a real stale one.

Tests (real PGLiteEngine + real lock rows): stale+no-lock -> fail;
stale+live-lock -> ok; never-synced+live-lock -> ok; never-synced+no-lock
-> fail; expired-TTL lock -> fail (wedged not masked); blocked source with
banked checkpoint rows but no lock -> still fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sync): honest --force-break-lock diagnostic when no lock is held

--force-break-lock used to emit the same terse "Lock ... is not held
(nothing to break)" line and exit 0 even when a sync was genuinely wedged,
sending the operator down a dead end — the wedge was not a held lock. Keep
rc=0 (breaking a non-existent lock is idempotently successful; flipping the
exit code would break automation), but under --force say plainly that
nothing was broken and point at the real next step (gbrain sync / gbrain
doctor) plus a `wedge_hint` field in --json output. The non-force path is
byte-for-byte unchanged.

runBreakLock is exported for the test. Tests: force+no-lock -> wedge_hint
JSON + human hint, rc 0; non-force+no-lock -> unchanged terse line, no hint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(doctor): surface the in-progress sync holder in the freshness message

Plan-completion follow-up to the BUG 4 live-lock signal: when a source is
actively syncing, name the holder (pid + host) in the check message instead
of silently folding it into synced_recently. The note is appended only when
something is in progress, so steady-state messages stay byte-for-byte
unchanged (the pinned exact-message + 3-bucket-invariant tests still pass).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sync): pre-landing review fixes — monotonic clock seed, scoped CHECK guard

Adversarial (codex) review of the implementation diff caught three:

- P1 (correctness): the fresh-schema setval was not monotonic. initSchema
  replays the schema blob, and the unconditional setval(MAX(generation))
  could move page_generation_clock_seq.last_value BACKWARD on an
  already-upgraded brain, letting a stored query_cache bookmark serve stale
  rows. Seed via GREATEST over the sequence's OWN last_value (+ old table
  value + MAX(generation)) in all 3 fresh schemas and migration v118, so a
  replay is idempotent — mirrors the old table's ON CONFLICT DO NOTHING.
  Pinned by a new monotonic regression test.
- P2: v119's CHECK-exists guard keyed on conname only (not globally unique).
  Scope it to conrelid = 'op_checkpoints'::regclass.
- P3: in-progress note ran into the prior sentence in fail/warn doctor
  messages; separate it with '. '.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: make Anthropic/ZE no-key tests hermetic against a dev config key

These "no key" tests cleared only ANTHROPIC_API_KEY / ZEROENTROPY_API_KEY
from the env, but hasAnthropicKey() and checkZeEmbeddingHealth() also read
the key from ~/.gbrain/config.json. On a dev machine whose real config holds
a key, the no-key assertions flipped and the tests failed locally (they
passed only in key-less CI). Add a shared with-env emptyHome() helper and
point GBRAIN_HOME at an empty dir in every no-key path so loadConfig finds
nothing — matching the already-hermetic anthropic-key / gateway-probe tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.44.1.0)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(key-files): sync doctor + op-checkpoint entries to v0.44.1.0 truth

checkSyncFreshness now reports an actively-running sync via the live
per-source lock (names holder pid+host, counts as synced_recently) instead
of flagging it stale; loadOpCheckpoint gates the legacy union arm on
jsonb_typeof = 'array' so a scalar parent can't take down the whole load,
and migration v119's CHECK constraint makes the corruption class
structurally impossible. Reference docs describe current behavior only —
both entries updated in place, no release-clause appends. Guard + llms
freshness test green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: re-version to v0.42.51.0 (natural next-off-master)

Maintainer override of the queue allocator's leap to 0.44.1.0 (it jumped past
in-flight sibling PR claims at 0.42.50/0.43.0/0.44.0). Take the natural next
slot in the 0.42.x line above the immediate sibling claim (0.42.50.0); a
merge re-bump resolves any collision if a cathedral PR lands first.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(e2e): bound + retry the OpenClaw install so a transient npm hang can't burn the Tier 2 budget

The Tier 2 (LLM Skills) job failed at 30m16s — the `npm install -g
openclaw@2026.4.9` step hung on a transient npm/registry stall (orphan
`npm install openclaw` was still running at cancel time) and consumed the
entire 30m job budget that v0.42.50.0 (#2254) introduced. The install
normally finishes in under a minute (Tier 2 is ~4m end to end on master),
so this is flaky-install infra, not a test failure.

Wrap the install in `timeout 120` + a 3-attempt retry loop with an 8-minute
step backstop: a hung attempt is killed in 2 min and retried instead of
eating the whole job. Same bound-the-hang philosophy as #2254's job timeouts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 14:02:47 -07:00

285 lines
11 KiB
TypeScript

/**
* v0.40.3.0 — query-cache-gate.ts (cache invalidation gate)
*
* Pure unit tests for the two helpers. Both helpers are pure functions
* (the validator) or take an engine argument (the snapshot builder); the
* latter uses a PGLite engine to test the real SQL fragment.
*
* Coverage:
* - buildPageGenerationsSnapshot: empty pageIds, populated pageIds,
* bookmark capture, missing pages excluded (LEFT JOIN equivalent).
* - validateCacheRowAgainstPages: vacuously valid for legacy/empty,
* bookmark short-circuit, single-page bump invalidates, deleted page
* invalidates, multi-page partial bump invalidates, integer-JSONB
* shape regression (string-typed values).
* - CACHE_GATE_WHERE_CLAUSE: source-text shape regression (grep guards).
*/
import { afterAll, beforeAll, beforeEach, describe, expect, test } from 'bun:test';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';
import {
buildPageGenerationsSnapshot,
validateCacheRowAgainstPages,
CACHE_GATE_WHERE_CLAUSE,
type PageGenerationsSnapshot,
} from '../src/core/search/query-cache-gate.ts';
describe('validateCacheRowAgainstPages (pure validator)', () => {
test('v0.41.19.0 D20/CDX-6 inversion: empty snapshot invalidates when bookmark fires', () => {
const snapshot: PageGenerationsSnapshot = {
page_generations: {},
max_generation_at_store: 0,
};
// Pre-v0.41.19.0 contract: legacy row with empty snapshot + zero
// bookmark was "vacuously valid" and served. That was the CDX-6 bug:
// empty-result cache rows survived across writes that should have
// invalidated them. Post-v0.41.19.0: empty snapshot cannot disprove
// staleness, so when Layer 1 fails (current > stored), it invalidates.
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 999, // Clock advanced since store
page_generations: {},
});
expect(ok).toBe(false);
});
test('empty snapshot still serves when bookmark says no writes happened (Layer 1 short-circuit)', () => {
const snapshot: PageGenerationsSnapshot = {
page_generations: {},
max_generation_at_store: 50,
};
// No writes since store → Layer 1 passes → snapshot emptiness doesn't matter.
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 50,
page_generations: {},
});
expect(ok).toBe(true);
});
test('bookmark short-circuit: MAX <= stored → valid without per-page work', () => {
const snapshot: PageGenerationsSnapshot = {
page_generations: { '1': 5, '2': 7 },
max_generation_at_store: 10,
};
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 10, // No writes since store
page_generations: { '1': 5, '2': 7 },
});
expect(ok).toBe(true);
});
test('single-page bumped invalidates', () => {
const snapshot: PageGenerationsSnapshot = {
page_generations: { '1': 5, '2': 7 },
max_generation_at_store: 7,
};
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 8, // Bookmark fires
page_generations: { '1': 5, '2': 8 }, // Page 2 bumped
});
expect(ok).toBe(false);
});
test('deleted page invalidates', () => {
const snapshot: PageGenerationsSnapshot = {
page_generations: { '1': 5, '2': 7 },
max_generation_at_store: 7,
};
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 8,
page_generations: { '1': 5, '2': undefined }, // Page 2 deleted
});
expect(ok).toBe(false);
});
test('multi-page partial bump invalidates', () => {
const snapshot: PageGenerationsSnapshot = {
page_generations: { '1': 5, '2': 7, '3': 9 },
max_generation_at_store: 9,
};
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 10,
page_generations: { '1': 5, '2': 7, '3': 10 }, // Only page 3 bumped
});
expect(ok).toBe(false);
});
test('codex D11 critical case (NON-empty snapshot): new page after store → Layer 1 fires → snapshot intact → row serves', () => {
// A brand-new page makes the clock advance but the cache row's
// page_generations snapshot doesn't reference it. The bookmark
// detects the corpus changed. Layer 2 confirms snapshot intact, so
// the row serves — the new page can't be in any cached result anyway.
// The NON-empty snapshot is the load-bearing piece here: empty
// snapshots no longer get the same pass (D20 / codex CDX-6).
const snapshot: PageGenerationsSnapshot = {
page_generations: { '1': 5, '2': 7 },
max_generation_at_store: 7,
};
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 8, // Page 3 was created
page_generations: { '1': 5, '2': 7 }, // Pages in snapshot unchanged
});
expect(ok).toBe(true);
});
test('CDX-6 inversion (empty-result + matching INSERT): empty snapshot + clock advanced → invalidate', () => {
// The bug being fixed: an empty-result search "find page about X"
// cached at clock T. Subsequently INSERT a matching page → clock T+1.
// Pre-v0.41.19.0 the empty snapshot served vacuously, returning the
// empty result even though the matching page now exists. Post-fix:
// invalidates so the next lookup re-queries.
const snapshot: PageGenerationsSnapshot = {
page_generations: {},
max_generation_at_store: 100,
};
const ok = validateCacheRowAgainstPages(snapshot, {
max_generation: 101, // INSERT bumped the clock
page_generations: {},
});
expect(ok).toBe(false);
});
});
describe('buildPageGenerationsSnapshot (PGLite-backed)', () => {
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await resetPgliteState(engine);
});
test('empty pageIds: snapshot has empty page_generations, MAX bookmark only', async () => {
// Seed a page so MAX > 0.
await engine.putPage('test/p1', {
type: 'note',
title: 'p1',
compiled_truth: 'body',
timeline: '',
frontmatter: {},
});
const snap = await buildPageGenerationsSnapshot(engine, []);
expect(snap.page_generations).toEqual({});
expect(snap.max_generation_at_store).toBeGreaterThan(0);
});
test('populated pageIds: snapshot captures generation per page + MAX', async () => {
const p1 = await engine.putPage('test/p1', {
type: 'note',
title: 'p1',
compiled_truth: 'body1',
timeline: '',
frontmatter: {},
});
const p2 = await engine.putPage('test/p2', {
type: 'note',
title: 'p2',
compiled_truth: 'body2',
timeline: '',
frontmatter: {},
});
const snap = await buildPageGenerationsSnapshot(engine, [p1.id, p2.id]);
expect(Object.keys(snap.page_generations).length).toBe(2);
expect(snap.page_generations[String(p1.id)]).toBeGreaterThan(0);
expect(snap.page_generations[String(p2.id)]).toBeGreaterThan(0);
expect(snap.max_generation_at_store).toBeGreaterThanOrEqual(
Math.max(snap.page_generations[String(p1.id)], snap.page_generations[String(p2.id)]),
);
});
test('integer-typed JSONB shape regression: stored values are numbers, not strings', async () => {
const p1 = await engine.putPage('test/p1', {
type: 'note',
title: 'p1',
compiled_truth: 'body1',
timeline: '',
frontmatter: {},
});
const snap = await buildPageGenerationsSnapshot(engine, [p1.id]);
const v = snap.page_generations[String(p1.id)];
expect(typeof v).toBe('number');
expect(Number.isInteger(v)).toBe(true);
});
test('non-existent pageIds: skipped (LEFT JOIN equivalent — no entry in page_generations)', async () => {
await engine.putPage('test/p1', {
type: 'note',
title: 'p1',
compiled_truth: 'body',
timeline: '',
frontmatter: {},
});
// Pass two IDs: one valid, one that doesn't exist.
const rows = await engine.executeRaw<{ id: number }>(`SELECT id FROM pages LIMIT 1`);
const p1Id = rows[0].id;
const snap = await buildPageGenerationsSnapshot(engine, [p1Id, 999999]);
expect(snap.page_generations[String(p1Id)]).toBeGreaterThan(0);
expect(snap.page_generations['999999']).toBeUndefined();
});
test('after content UPDATE: generation bumps and snapshot reflects new value', async () => {
const p1 = await engine.putPage('test/p1', {
type: 'note',
title: 'p1',
compiled_truth: 'body-v1',
timeline: '',
frontmatter: {},
});
const before = await buildPageGenerationsSnapshot(engine, [p1.id]);
const beforeGen = before.page_generations[String(p1.id)];
// Update compiled_truth (content allow-list column → trigger bumps).
await engine.putPage('test/p1', {
type: 'note',
title: 'p1',
compiled_truth: 'body-v2',
timeline: '',
frontmatter: {},
});
const after = await buildPageGenerationsSnapshot(engine, [p1.id]);
const afterGen = after.page_generations[String(p1.id)];
expect(afterGen).toBeGreaterThan(beforeGen);
});
});
describe('CACHE_GATE_WHERE_CLAUSE (SQL shape regression)', () => {
test('v0.42.x: Layer 1 reads page_generation_clock_seq.last_value (not the locked row, not MAX)', () => {
expect(CACHE_GATE_WHERE_CLAUSE).toContain('page_generation_clock_seq');
expect(CACHE_GATE_WHERE_CLAUSE).toContain('last_value');
expect(CACHE_GATE_WHERE_CLAUSE).toContain('qc.max_generation_at_store');
// Negative regression guard: the old MAX(generation) read shape MUST
// be gone (codex CDX-1/CDX-2: it silently served stale on
// UPDATE-to-non-max and DELETE).
expect(CACHE_GATE_WHERE_CLAUSE).not.toContain('MAX(generation) FROM pages');
// The locked single-row read (the BUG 1 contention source) MUST be gone.
expect(CACHE_GATE_WHERE_CLAUSE).not.toContain('value FROM page_generation_clock WHERE id');
});
test('contains Layer 2 per-page snapshot (jsonb_each + LEFT JOIN)', () => {
expect(CACHE_GATE_WHERE_CLAUSE).toContain('jsonb_each(qc.page_generations)');
expect(CACHE_GATE_WHERE_CLAUSE).toContain('LEFT JOIN pages');
});
test('v0.41.19.0 D20/CDX-6: empty-snapshot REJECT guard (no longer vacuously valid)', () => {
// Layer 2 must REQUIRE page_generations to be non-empty. Pre-fix
// shape was `qc.page_generations = '{}'::jsonb OR NOT EXISTS(...)`
// which let empty snapshots survive any clock bump.
expect(CACHE_GATE_WHERE_CLAUSE).toContain(`qc.page_generations <> '{}'::jsonb`);
expect(CACHE_GATE_WHERE_CLAUSE).not.toMatch(/qc\.page_generations = '\{\}'::jsonb\s*OR/);
});
test('per-page mismatch path checks both deletion (NULL) and bump (!=)', () => {
expect(CACHE_GATE_WHERE_CLAUSE).toContain('p.id IS NULL');
expect(CACHE_GATE_WHERE_CLAUSE).toContain('p.generation <>');
});
});