mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* feat(sync): migration v98 last_refreshed_at + deleteLockRowIfStale helper Schema foundation for v0.41.15.0's `gbrain sync --break-lock --max-age <s>` flag. Adds `gbrain_cycle_locks.last_refreshed_at TIMESTAMPTZ` as the heartbeat signal that distinguishes wedged-but-alive lock holders from healthy long-running syncs that are actively refreshing. Why last_refreshed_at not acquired_at: `withRefreshingLock` already bumps `ttl_expires_at` every ~5 min while work runs, but leaves `acquired_at` at the original timestamp. A 35-min media-corpus sync that's healthy has `acquired_at` 35 min ago but `last_refreshed_at` 30 seconds ago. Using acquired_at for --max-age would steal healthy locks; last_refreshed_at correctly identifies only holders whose JS interval has stopped firing. D-V4-1 rollout safety: migration v98 backfills `last_refreshed_at = NOW()` (NOT `= acquired_at`) so pre-upgrade holders running the old binary get a 30-min protection window. After that window all pre-upgrade syncs are either complete (lock released) OR genuinely wedged (--max-age does the right thing). Documented as a known caveat in CHANGELOG. D-V4-mech-4 SQL cast: deleteLockRowIfStale uses `$N * INTERVAL '1 second'` not `$N::interval` (Postgres does not cast integer to interval the latter way). Atomic DELETE keyed on (id, holder_pid, last_refreshed_at < NOW() - $N * INTERVAL '1 second') RETURNING id, last_refreshed_at — no TOCTOU between inspect + delete. D-V4-mech-3 schema-snapshot parity: column added to all 3 snapshots so fresh init paths (pglite-schema.ts, schema.sql) initialize correctly without depending on the migration runner. schema-embedded.ts regenerated via `bun run build:schema`. Pinned by 13 PGLite cases in test/sync-break-lock-all.test.ts: tryAcquireDbLock writes on INSERT, withRefreshingLock refresh bumps both columns, inspectLock surfaces the new field, deleteLockRowIfStale refuses fresh / breaks stale / safe on holder_pid mismatch / refuses NULL (pre-v98). R1 + R6 regression invariants from the v4 plan. Closes #1472 (RFC from @garrytan-agents) — schema foundation only; performSync abort threading + CLI flags + consumer threading land in follow-up commits in this PR. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(sync): --timeout + --max-age + partial status + per-source AbortController The CLI surface for v0.41.15.0. Wires `gbrain sync --timeout <s>` (graceful self-termination) and `gbrain sync --break-lock --all --max-age <s>` (cron-self-heal) end-to-end through `performSync`, `runOne`, `runBreakLock`, and all `SyncResult.status` consumers. Surface 1: `gbrain sync --timeout <s>` - New `SyncOpts.signal?: AbortSignal` threads through `performSync` → `withRefreshingLock` work callback → `performSyncInner`. - D-V3-1 honest scope: abort checks fire ONLY in pre-bookmark phases (pull, delete, rename, import). Extract + embed run to completion if reached. The `last_commit` bookmark write at sync.ts:1261 is the invariant boundary — partial CANNOT advance the bookmark because the abort checkpoints sit strictly before that write. - D-V3-2 per-iteration: abort check at top of every loop iteration (delete, rename, serial import, each parallel worker's while loop) matches the per-file granularity the existing loops already have. - D-V3-3 per-source AbortController: `--timeout --all` creates ONE controller inside runOne per source so each gets its own budget; NOT a shared global controller (which would starve later sources). try/finally + timer.unref() guarantees cleanup on throw. - D-V4-mech-7 pull error.cause: pullRepo wraps execFileSync errors in GitOperationError. The catch inspects e.cause.code === 'ETIMEDOUT' and e.cause.signal === 'SIGTERM' (NOT the top-level error) to distinguish timeout (partial reason='pull_timeout') from ordinary pull failure (existing warn-and-continue, R2 invariant preserved). Surface 2: `gbrain sync --break-lock [--all] [--max-age <s>]` - Drops the --all refusal at sync.ts:1610. When combined with --all, runBreakLock iterates every active source and prints per-source verdict. - --max-age routes through the new deleteLockRowIfStale helper from db-lock.ts (atomic age-gated DELETE; no TOCTOU). Healthy refreshing holders survive by construction; only wedged-but-alive holders trip. D-V3-5 partial-status consumer threading (conservative posture matching blocked_by_failures): - printSyncResult: new `case 'partial':` arm reports filesImported + reason; tells operator to re-run to continue. - manageGitignore (both single-source and parallel runOne sites, plus watch mode): excludes partial from the gate. A partial sync's db_only path set isn't fully reconciled. - Auto-embed-backfill enqueue inside runOne: excludes partial. The next clean sync will re-walk and re-decide. CLI flag parsing (T16): - parseDurationSeconds in sync-concurrency.ts: accepts 60s/10m/1h/bare int; rejects 0/negatives/decimals/garbage. Names the failing flag in the error message. - --timeout requires --source OR --all (validation rejects bare `gbrain sync --timeout`). - --max-age requires --break-lock; mutually exclusive with --force-break-lock. Coverage: - 15 unit cases (test/sync-timeout.test.ts) pin parseDurationSeconds + SyncResult union additivity. - 2 E2E cases (test/e2e/sync-parallel.test.ts) pin the abort-mid-import contract against real Postgres: status='partial', last_commit unchanged, filesImported bounded. Closes #1472 (RFC from @garrytan-agents) — CLI surface; schema foundation landed in the previous commit. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(heavy): sync_timeout_rescue.sh reproducer for the cron-cascade 10K-page seed × 4 sources × deliberately tight --timeout × 3 sequential cron emulations. Asserts every source reaches `last_commit === HEAD` within 3 waves. Proves the v0.41.15.0 fix breaks the cascade the PR #1472 RFC documented. Workload (tests/heavy/_sync_timeout_rescue_workload.ts) is PGLite-only because the PGLite engine forces serial sync internally (parallelEligible excludes it). The parallel-fan-out + per-source AbortController case lives in test/e2e/sync-parallel.test.ts against real Postgres. This heavy test pins the contract that matters for cron: aborts → partial returns → next wave content_hash-short-circuits + makes new progress. Smoke-tested locally at PAGES=50 WAVES=2 TIMEOUT_SECONDS=2: every source converges within 2 waves. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(v0.41.15.0): CHANGELOG + README + TODOS + version bump Bumps VERSION + package.json to 0.41.15.0 (next slot after master's v0.41.14.0). CHANGELOG entry leads ELI10 per gstack voice rules and documents the 3 intentional honest gaps: 1. --timeout covers pull + delete + rename + import only; extract + embed run to completion (D-V3-1 honest scope). 2. First 30 min after migration v98, --max-age cannot identify wedged pre-upgrade holders (D-V4-1 rollout trade-off). 3. Full-sync triggers (first sync, --full, chunker-version rewalk) don't respect --timeout yet (deferred to v0.42+). README troubleshooting section: paste-ready cron pattern with shell timeout(1) for OS-level process isolation + gbrain's --timeout for graceful self-termination half-a-minute earlier. TODOS.md: v0.42+ entries for subprocess fan-out (revisit if shell timeout(1) proves insufficient), full-sync --timeout coverage via AbortSignal in runImport, and runFactsBackstop microtask-queue process-alive caveat. llms-full.txt regenerated via `bun run build:llms`. Closes #1472 (RFC from @garrytan-agents). Credit to @garrytan-agents in the CHANGELOG for surfacing the production cron-failure data that motivated the work. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(test-isolation): rewrite JSDoc to not match mock.module() lint regex scripts/check-test-isolation.sh greps for the literal string `mock.module(` to flag top-level module mocks (R2 rule — top-level mocks leak across files in the shard process). The regex doesn't know about comments, so my two new test files tripped the lint with JSDoc lines literally describing the rule: test/sync-timeout.test.ts:11 "* `mock.module()` (R2). Engine ..." test/sync-break-lock-all.test.ts:15 "* mock.module(), no process.env ..." Both files had ZERO actual mock.module() calls — only the comment text matched. Rewrote both JSDocs to refer to "top-level module mocks" instead of the literal token. Same meaning; doesn't trip the regex. `bun run check:test-isolation` now passes (714 non-serial unit files scanned). `bun run verify` clean (22/22 checks pass). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
284 lines
11 KiB
TypeScript
284 lines
11 KiB
TypeScript
/**
|
||
* E2E test for parallel sync against real Postgres.
|
||
*
|
||
* T2 — happy path: 60-file sync at concurrency=4 against PostgresEngine
|
||
* actually constructs N worker engines, imports correctly, and does
|
||
* not leak connections (probe pg_stat_activity before/after).
|
||
* P4 — benchmark: serial vs concurrency=4 timing on the same fixture so
|
||
* the v0.22.13 CHANGELOG can quote a real number instead of "~4×".
|
||
*
|
||
* Gated on DATABASE_URL. Run via:
|
||
* docker run -d --name gbrain-test-pg -e POSTGRES_USER=postgres \
|
||
* -e POSTGRES_PASSWORD=postgres -e POSTGRES_DB=gbrain_test \
|
||
* -p 5435:5432 pgvector/pgvector:pg16
|
||
* DATABASE_URL=postgresql://postgres:postgres@localhost:5435/gbrain_test \
|
||
* bun test test/e2e/sync-parallel.test.ts
|
||
*/
|
||
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
||
import { mkdtempSync, rmSync, writeFileSync, mkdirSync } from 'fs';
|
||
import { join } from 'path';
|
||
import { tmpdir } from 'os';
|
||
import { execSync } from 'child_process';
|
||
import { hasDatabase, setupDB, teardownDB, getEngine, getConn } from './helpers.ts';
|
||
|
||
const skip = !hasDatabase();
|
||
const describeE2E = skip ? describe.skip : describe;
|
||
|
||
if (skip) {
|
||
console.log('Skipping E2E sync-parallel tests (DATABASE_URL not set)');
|
||
}
|
||
|
||
function seedRepo(repoPath: string, fileCount: number): string {
|
||
execSync('git init', { cwd: repoPath, stdio: 'pipe' });
|
||
execSync('git config user.email "test@test.com"', { cwd: repoPath, stdio: 'pipe' });
|
||
execSync('git config user.name "Test"', { cwd: repoPath, stdio: 'pipe' });
|
||
mkdirSync(join(repoPath, 'people'), { recursive: true });
|
||
for (let i = 0; i < fileCount; i++) {
|
||
writeFileSync(join(repoPath, `people/p${i}.md`), [
|
||
'---',
|
||
'type: person',
|
||
`title: Person ${i}`,
|
||
'---',
|
||
'',
|
||
`Person ${i} body — some text long enough to chunk.`,
|
||
`Iteration index ${i}, generated by sync-parallel E2E.`,
|
||
].join('\n'));
|
||
}
|
||
execSync('git add -A && git commit -m "initial"', { cwd: repoPath, stdio: 'pipe' });
|
||
return execSync('git rev-parse HEAD', { cwd: repoPath, encoding: 'utf-8' }).trim();
|
||
}
|
||
|
||
async function activeConnections(): Promise<number> {
|
||
const conn = getConn();
|
||
const rows = await conn.unsafe(`
|
||
SELECT count(*) AS n FROM pg_stat_activity
|
||
WHERE datname = current_database()
|
||
AND state IS NOT NULL
|
||
`) as Array<{ n: string }>;
|
||
return parseInt(rows[0]?.n ?? '0', 10);
|
||
}
|
||
|
||
describeE2E('E2E sync-parallel: T2 happy path + leak probe', () => {
|
||
let repoPath: string;
|
||
|
||
beforeAll(async () => {
|
||
await setupDB();
|
||
}, 30_000);
|
||
|
||
afterAll(async () => {
|
||
if (repoPath) rmSync(repoPath, { recursive: true, force: true });
|
||
await teardownDB();
|
||
});
|
||
|
||
test('60-file Postgres sync at concurrency=4 imports all + no connection leak', async () => {
|
||
repoPath = mkdtempSync(join(tmpdir(), 'gbrain-e2e-par-'));
|
||
seedRepo(repoPath, 60);
|
||
|
||
const before = await activeConnections();
|
||
|
||
const { performSync } = await import('../../src/commands/sync.ts');
|
||
const engine = getEngine();
|
||
const result = await performSync(engine, {
|
||
repoPath,
|
||
noPull: true,
|
||
noEmbed: true,
|
||
concurrency: 4,
|
||
});
|
||
|
||
// First sync routes through performFullSync (delegates to runImport which
|
||
// also accepts --workers); status is 'first_sync'.
|
||
expect(result.status).toBe('first_sync');
|
||
|
||
const after = await activeConnections();
|
||
|
||
// Allow some slack — the helper engine + sync's normal pool stay open.
|
||
// Worker engines (4 × 2 = 8 connections) MUST have closed; if they
|
||
// hadn't, after - before would be at least 8.
|
||
expect(after - before).toBeLessThan(4);
|
||
|
||
// Verify pages are actually in the DB (via raw SQL — engine API also works).
|
||
const conn = getConn();
|
||
const pageRows = await conn.unsafe(
|
||
`SELECT count(*) AS n FROM pages WHERE slug LIKE 'people/p%'`,
|
||
) as Array<{ n: string }>;
|
||
const count = parseInt(pageRows[0]?.n ?? '0', 10);
|
||
expect(count).toBe(60);
|
||
}, 60_000);
|
||
});
|
||
|
||
describeE2E('E2E sync-parallel: P4 benchmark serial vs concurrency=4', () => {
|
||
let repoSerial: string;
|
||
let repoParallel: string;
|
||
|
||
beforeAll(async () => {
|
||
await setupDB();
|
||
}, 30_000);
|
||
|
||
afterAll(async () => {
|
||
if (repoSerial) rmSync(repoSerial, { recursive: true, force: true });
|
||
if (repoParallel) rmSync(repoParallel, { recursive: true, force: true });
|
||
await teardownDB();
|
||
});
|
||
|
||
test('120-file benchmark: report serial and parallel wall-clock', async () => {
|
||
// Two separate repos so neither sync's chunks bleed into the other.
|
||
repoSerial = mkdtempSync(join(tmpdir(), 'gbrain-bench-serial-'));
|
||
repoParallel = mkdtempSync(join(tmpdir(), 'gbrain-bench-parallel-'));
|
||
seedRepo(repoSerial, 120);
|
||
seedRepo(repoParallel, 120);
|
||
|
||
const { performSync } = await import('../../src/commands/sync.ts');
|
||
const engine = getEngine();
|
||
|
||
// Truncate between runs to keep the benchmark honest.
|
||
const conn = getConn();
|
||
|
||
const t1 = Date.now();
|
||
await performSync(engine, {
|
||
repoPath: repoSerial,
|
||
noPull: true,
|
||
noEmbed: true,
|
||
concurrency: 1,
|
||
});
|
||
const serialMs = Date.now() - t1;
|
||
|
||
// Wipe pages before second run so neither one is "incremental".
|
||
await conn.unsafe(`TRUNCATE pages CASCADE`);
|
||
await conn.unsafe(`TRUNCATE config CASCADE`);
|
||
|
||
const t2 = Date.now();
|
||
await performSync(engine, {
|
||
repoPath: repoParallel,
|
||
noPull: true,
|
||
noEmbed: true,
|
||
concurrency: 4,
|
||
});
|
||
const parallelMs = Date.now() - t2;
|
||
|
||
const speedup = (serialMs / parallelMs).toFixed(2);
|
||
// Emit as a single line stdout consumers can grep for.
|
||
console.log(`SYNC_PARALLEL_BENCH 120 files | serial=${serialMs}ms | parallel(4)=${parallelMs}ms | speedup=${speedup}x`);
|
||
|
||
// Soft assertion: parallel must not be slower than serial. The actual
|
||
// speedup ratio depends heavily on Postgres latency profile and is what
|
||
// the CHANGELOG quotes — don't gate the test on a specific multiplier.
|
||
expect(parallelMs).toBeLessThanOrEqual(serialMs * 1.5); // +50% slack for noisy CI
|
||
}, 120_000);
|
||
});
|
||
|
||
describeE2E('E2E sync-parallel: T18 --timeout returns partial; last_commit unchanged', () => {
|
||
let repoPath: string;
|
||
|
||
beforeAll(async () => {
|
||
await setupDB();
|
||
}, 30_000);
|
||
|
||
afterAll(async () => {
|
||
if (repoPath) rmSync(repoPath, { recursive: true, force: true });
|
||
await teardownDB();
|
||
});
|
||
|
||
test('signal aborted mid-import returns partial and does not advance last_commit', async () => {
|
||
// v0.41.13.0 (T18 / D-V4-mech-10): real-Postgres E2E for the
|
||
// --timeout partial-status contract. PGLite tests cover the
|
||
// single-source AbortSignal threading in test/sync-break-lock-all.test.ts;
|
||
// this case verifies the same contract on the actual Postgres engine
|
||
// because parallelEligible excludes PGLite from the worker fan-out
|
||
// and the bookmark-write semantic uses real Postgres timestamp + index
|
||
// behavior.
|
||
repoPath = mkdtempSync(join(tmpdir(), 'gbrain-e2e-timeout-'));
|
||
seedRepo(repoPath, 200);
|
||
|
||
const { performSync } = await import('../../src/commands/sync.ts');
|
||
const engine = getEngine();
|
||
|
||
// Register a source so per-source last_commit lives in `sources`.
|
||
const conn = getConn();
|
||
await conn.unsafe(
|
||
`INSERT INTO sources (id, name, local_path) VALUES ($1, $2, $3)
|
||
ON CONFLICT (id) DO UPDATE SET local_path = EXCLUDED.local_path`,
|
||
['e2e-timeout-source', 'e2e-timeout-source', repoPath],
|
||
);
|
||
|
||
// Fire abort immediately. With a 200-file diff, performSync's per-file
|
||
// abort check at the top of the import loop fires before file 1 starts,
|
||
// so files_imported should be 0 and last_commit should stay null.
|
||
const controller = new AbortController();
|
||
controller.abort();
|
||
|
||
const result = await performSync(engine, {
|
||
repoPath,
|
||
sourceId: 'e2e-timeout-source',
|
||
noPull: true,
|
||
noEmbed: true,
|
||
noExtract: true,
|
||
concurrency: 1,
|
||
signal: controller.signal,
|
||
});
|
||
|
||
expect(result.status).toBe('partial');
|
||
expect(result.reason).toBeDefined();
|
||
// last_commit must NOT have advanced (D-V3-1 invariant — partial
|
||
// fires strictly before the writeSyncAnchor call).
|
||
const rows = await conn.unsafe(
|
||
`SELECT last_commit FROM sources WHERE id = $1`,
|
||
['e2e-timeout-source'],
|
||
) as Array<{ last_commit: string | null }>;
|
||
expect(rows[0]?.last_commit).toBeNull();
|
||
}, 60_000);
|
||
|
||
test('signal aborted after a few imports leaves last_commit unchanged and reports partial files_imported', async () => {
|
||
// Rebuild a fresh repo for this test; the prior describe path uses
|
||
// its own repoPath variable.
|
||
const repo2 = mkdtempSync(join(tmpdir(), 'gbrain-e2e-timeout-partial-'));
|
||
try {
|
||
seedRepo(repo2, 50);
|
||
const { performSync } = await import('../../src/commands/sync.ts');
|
||
const engine = getEngine();
|
||
const conn = getConn();
|
||
await conn.unsafe(
|
||
`INSERT INTO sources (id, name, local_path) VALUES ($1, $2, $3)
|
||
ON CONFLICT (id) DO UPDATE SET local_path = EXCLUDED.local_path`,
|
||
['e2e-timeout-partial', 'e2e-timeout-partial', repo2],
|
||
);
|
||
|
||
// Schedule abort 250ms in. On a 50-file repo with real Postgres
|
||
// round-trips per import, some files persist before abort fires.
|
||
// We assert that:
|
||
// - status is partial OR first_sync (race-tolerant — if Postgres
|
||
// is fast enough that all 50 imports finish in <250ms, the run
|
||
// completes successfully which is also a valid outcome)
|
||
// - if partial: filesImported is bounded between 1 and 49
|
||
// - if partial: last_commit is null (never advanced past partial)
|
||
const controller = new AbortController();
|
||
setTimeout(() => controller.abort(), 250).unref();
|
||
|
||
const result = await performSync(engine, {
|
||
repoPath: repo2,
|
||
sourceId: 'e2e-timeout-partial',
|
||
noPull: true,
|
||
noEmbed: true,
|
||
noExtract: true,
|
||
concurrency: 1,
|
||
signal: controller.signal,
|
||
});
|
||
|
||
if (result.status === 'partial') {
|
||
expect(result.filesImported).toBeGreaterThanOrEqual(0);
|
||
expect(result.filesImported).toBeLessThanOrEqual(50);
|
||
const rows = await conn.unsafe(
|
||
`SELECT last_commit FROM sources WHERE id = $1`,
|
||
['e2e-timeout-partial'],
|
||
) as Array<{ last_commit: string | null }>;
|
||
expect(rows[0]?.last_commit).toBeNull();
|
||
} else {
|
||
// first_sync or synced — sub-250ms full run; not a contract violation.
|
||
// The point of the test is that IF partial happens, the invariants hold.
|
||
expect(['first_sync', 'synced']).toContain(result.status);
|
||
}
|
||
} finally {
|
||
rmSync(repo2, { recursive: true, force: true });
|
||
}
|
||
}, 60_000);
|
||
});
|