mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
814258dda67945ffec9457a1e73980e947b7e462
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9bf96db807 |
v0.42.51.0 fix(sync): contention-free clock + checkpoint integrity + honest sync freshness (#2255)
* fix(sync): contention-free page-generation clock — sequence swap The page-generation clock backed the query-cache Layer-1 bookmark via a FOR EACH STATEMENT trigger running `UPDATE page_generation_clock SET value=value+1 WHERE id=1`. That took a transaction-length RowExclusiveLock on one tuple, so every concurrent page writer serialized on the prior writer's COMMIT — sync ran at ~0.8 cores regardless of worker count. Swap to a SEQUENCE bumped by nextval() (a microsecond LWLock, never a row lock). The clock's only contract is monotonic advancement on any page INSERT/UPDATE/DELETE; last_value is non-transactional, so rolled-back or concurrent-uncommitted writers only OVER-invalidate the cache (lose a hit), never serve stale. - migration v118: CREATE SEQUENCE + load-bearing 2-arg setval (is_called= true, floor 1, seeded >= old clock and MAX(generation)) + repoint the trigger function body + DELETE query_cache so no old-clock bookmark survives the swap. v107 left immutable. - query-cache-gate.ts: 3 readers -> SELECT last_value FROM page_generation_clock_seq. - schema.sql + pglite-schema.ts (+ regenerated schema-embedded.ts) ship the sequence on fresh install; table + trigger names retained. - tests: clockValue reads last_value; mechanism proof (trigger fn uses nextval not the row UPDATE); rollback-advances-clock safety pin; real PGLite sequence round-trip (is_called gotcha); shape test requires _seq. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sync): op_checkpoints array-shape guard — CHECK + repair + defensive loader completed_keys is JSONB and the checkpoint loader runs jsonb_array_elements_text over it. A non-array (scalar) value makes that throw "cannot extract elements from a scalar", which takes down the whole UNION load — including the valid op_checkpoint_paths child rows — and loses all checkpoint progress for that key. No current writer produces a scalar, but an older binary / external script / future bug could. Make the corruption class structurally impossible and self-healing: - migration v119: LOCK TABLE (so an out-of-band scalar can't land between repair and constrain; no-op on single-connection PGLite), repair any pre-existing scalar to '[]' (op_checkpoint_paths child rows are the append-only source of truth, so the reset loses nothing), then add the named CHECK (jsonb_typeof(completed_keys) = 'array') via a pg_constraint IF NOT EXISTS guard. A DB-enforced always-on guard — the correct pattern vs a migration verify-hook, which never runs on already-stamped brains. - schema.sql + pglite-schema.ts (+ regenerated schema-embedded.ts) ship the same NAMED inline CHECK so fresh installs match migrated brains and v119 skips the duplicate. - op-checkpoint.ts loader: gate the legacy arm on jsonb_typeof = 'array' so a scalar parent is skipped (children still load) instead of throwing the whole union, and log a specific corruption warning when one is seen. - tests: CHECK rejects a scalar (exactly one constraint, no blob+migration dupe); loader survives a scalar parent and returns the children; v119 repair converts a scalar to '[]'. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(doctor): report actively-running sync via live lock, not stale freshness A slow source that makes partial progress every cycle but never fully completes used to read as permanently "stale" / "never synced" because last_sync_at only advances on a full successful sync. The naive fix (treat recent checkpoint banking as "in progress") is unsafe: a blocked sync banks the good files then writes no anchor, so banking can't tell in-progress from wedged. Use the only honest signal: a LIVE, non-expired per-source sync lock (inspectLock + syncLockId against gbrain_cycle_locks). Every non-skipLock sync holds it and refreshes it; a blocked/failed sync's process has exited (no lock row) and a wedged holder stops refreshing (TTL lapses), so either correctly falls through to the stale path and is NEVER masked. An actively-syncing source (including a never-synced source doing its first sync) counts as synced_recently, preserving the pinned 3-bucket invariant. The lock lookup reuses doctor's existing dynamic db-lock import and swallows any throw (stub engine, pre-lock-table brain) to false, so it can only ADD an in-progress verdict, never suppress a real stale one. Tests (real PGLiteEngine + real lock rows): stale+no-lock -> fail; stale+live-lock -> ok; never-synced+live-lock -> ok; never-synced+no-lock -> fail; expired-TTL lock -> fail (wedged not masked); blocked source with banked checkpoint rows but no lock -> still fail. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sync): honest --force-break-lock diagnostic when no lock is held --force-break-lock used to emit the same terse "Lock ... is not held (nothing to break)" line and exit 0 even when a sync was genuinely wedged, sending the operator down a dead end — the wedge was not a held lock. Keep rc=0 (breaking a non-existent lock is idempotently successful; flipping the exit code would break automation), but under --force say plainly that nothing was broken and point at the real next step (gbrain sync / gbrain doctor) plus a `wedge_hint` field in --json output. The non-force path is byte-for-byte unchanged. runBreakLock is exported for the test. Tests: force+no-lock -> wedge_hint JSON + human hint, rc 0; non-force+no-lock -> unchanged terse line, no hint. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(doctor): surface the in-progress sync holder in the freshness message Plan-completion follow-up to the BUG 4 live-lock signal: when a source is actively syncing, name the holder (pid + host) in the check message instead of silently folding it into synced_recently. The note is appended only when something is in progress, so steady-state messages stay byte-for-byte unchanged (the pinned exact-message + 3-bucket-invariant tests still pass). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(sync): pre-landing review fixes — monotonic clock seed, scoped CHECK guard Adversarial (codex) review of the implementation diff caught three: - P1 (correctness): the fresh-schema setval was not monotonic. initSchema replays the schema blob, and the unconditional setval(MAX(generation)) could move page_generation_clock_seq.last_value BACKWARD on an already-upgraded brain, letting a stored query_cache bookmark serve stale rows. Seed via GREATEST over the sequence's OWN last_value (+ old table value + MAX(generation)) in all 3 fresh schemas and migration v118, so a replay is idempotent — mirrors the old table's ON CONFLICT DO NOTHING. Pinned by a new monotonic regression test. - P2: v119's CHECK-exists guard keyed on conname only (not globally unique). Scope it to conrelid = 'op_checkpoints'::regclass. - P3: in-progress note ran into the prior sentence in fail/warn doctor messages; separate it with '. '. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: make Anthropic/ZE no-key tests hermetic against a dev config key These "no key" tests cleared only ANTHROPIC_API_KEY / ZEROENTROPY_API_KEY from the env, but hasAnthropicKey() and checkZeEmbeddingHealth() also read the key from ~/.gbrain/config.json. On a dev machine whose real config holds a key, the no-key assertions flipped and the tests failed locally (they passed only in key-less CI). Add a shared with-env emptyHome() helper and point GBRAIN_HOME at an empty dir in every no-key path so loadConfig finds nothing — matching the already-hermetic anthropic-key / gateway-probe tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: bump version and changelog (v0.44.1.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(key-files): sync doctor + op-checkpoint entries to v0.44.1.0 truth checkSyncFreshness now reports an actively-running sync via the live per-source lock (names holder pid+host, counts as synced_recently) instead of flagging it stale; loadOpCheckpoint gates the legacy union arm on jsonb_typeof = 'array' so a scalar parent can't take down the whole load, and migration v119's CHECK constraint makes the corruption class structurally impossible. Reference docs describe current behavior only — both entries updated in place, no release-clause appends. Guard + llms freshness test green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: re-version to v0.42.51.0 (natural next-off-master) Maintainer override of the queue allocator's leap to 0.44.1.0 (it jumped past in-flight sibling PR claims at 0.42.50/0.43.0/0.44.0). Take the natural next slot in the 0.42.x line above the immediate sibling claim (0.42.50.0); a merge re-bump resolves any collision if a cathedral PR lands first. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(e2e): bound + retry the OpenClaw install so a transient npm hang can't burn the Tier 2 budget The Tier 2 (LLM Skills) job failed at 30m16s — the `npm install -g openclaw@2026.4.9` step hung on a transient npm/registry stall (orphan `npm install openclaw` was still running at cancel time) and consumed the entire 30m job budget that v0.42.50.0 (#2254) introduced. The install normally finishes in under a minute (Tier 2 is ~4m end to end on master), so this is flaky-install infra, not a test failure. Wrap the install in `timeout 120` + a 3-attempt retry loop with an 8-minute step backstop: a hung attempt is killed in 2 min and retried instead of eating the whole job. Same bound-the-hang philosophy as #2254's job timeouts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cd8efee0ea |
v0.41.15.0 feat(sync): --timeout + --max-age + partial status (closes #1472 RFC) (#1506)
* feat(sync): migration v98 last_refreshed_at + deleteLockRowIfStale helper Schema foundation for v0.41.15.0's `gbrain sync --break-lock --max-age <s>` flag. Adds `gbrain_cycle_locks.last_refreshed_at TIMESTAMPTZ` as the heartbeat signal that distinguishes wedged-but-alive lock holders from healthy long-running syncs that are actively refreshing. Why last_refreshed_at not acquired_at: `withRefreshingLock` already bumps `ttl_expires_at` every ~5 min while work runs, but leaves `acquired_at` at the original timestamp. A 35-min media-corpus sync that's healthy has `acquired_at` 35 min ago but `last_refreshed_at` 30 seconds ago. Using acquired_at for --max-age would steal healthy locks; last_refreshed_at correctly identifies only holders whose JS interval has stopped firing. D-V4-1 rollout safety: migration v98 backfills `last_refreshed_at = NOW()` (NOT `= acquired_at`) so pre-upgrade holders running the old binary get a 30-min protection window. After that window all pre-upgrade syncs are either complete (lock released) OR genuinely wedged (--max-age does the right thing). Documented as a known caveat in CHANGELOG. D-V4-mech-4 SQL cast: deleteLockRowIfStale uses `$N * INTERVAL '1 second'` not `$N::interval` (Postgres does not cast integer to interval the latter way). Atomic DELETE keyed on (id, holder_pid, last_refreshed_at < NOW() - $N * INTERVAL '1 second') RETURNING id, last_refreshed_at — no TOCTOU between inspect + delete. D-V4-mech-3 schema-snapshot parity: column added to all 3 snapshots so fresh init paths (pglite-schema.ts, schema.sql) initialize correctly without depending on the migration runner. schema-embedded.ts regenerated via `bun run build:schema`. Pinned by 13 PGLite cases in test/sync-break-lock-all.test.ts: tryAcquireDbLock writes on INSERT, withRefreshingLock refresh bumps both columns, inspectLock surfaces the new field, deleteLockRowIfStale refuses fresh / breaks stale / safe on holder_pid mismatch / refuses NULL (pre-v98). R1 + R6 regression invariants from the v4 plan. Closes #1472 (RFC from @garrytan-agents) — schema foundation only; performSync abort threading + CLI flags + consumer threading land in follow-up commits in this PR. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(sync): --timeout + --max-age + partial status + per-source AbortController The CLI surface for v0.41.15.0. Wires `gbrain sync --timeout <s>` (graceful self-termination) and `gbrain sync --break-lock --all --max-age <s>` (cron-self-heal) end-to-end through `performSync`, `runOne`, `runBreakLock`, and all `SyncResult.status` consumers. Surface 1: `gbrain sync --timeout <s>` - New `SyncOpts.signal?: AbortSignal` threads through `performSync` → `withRefreshingLock` work callback → `performSyncInner`. - D-V3-1 honest scope: abort checks fire ONLY in pre-bookmark phases (pull, delete, rename, import). Extract + embed run to completion if reached. The `last_commit` bookmark write at sync.ts:1261 is the invariant boundary — partial CANNOT advance the bookmark because the abort checkpoints sit strictly before that write. - D-V3-2 per-iteration: abort check at top of every loop iteration (delete, rename, serial import, each parallel worker's while loop) matches the per-file granularity the existing loops already have. - D-V3-3 per-source AbortController: `--timeout --all` creates ONE controller inside runOne per source so each gets its own budget; NOT a shared global controller (which would starve later sources). try/finally + timer.unref() guarantees cleanup on throw. - D-V4-mech-7 pull error.cause: pullRepo wraps execFileSync errors in GitOperationError. The catch inspects e.cause.code === 'ETIMEDOUT' and e.cause.signal === 'SIGTERM' (NOT the top-level error) to distinguish timeout (partial reason='pull_timeout') from ordinary pull failure (existing warn-and-continue, R2 invariant preserved). Surface 2: `gbrain sync --break-lock [--all] [--max-age <s>]` - Drops the --all refusal at sync.ts:1610. When combined with --all, runBreakLock iterates every active source and prints per-source verdict. - --max-age routes through the new deleteLockRowIfStale helper from db-lock.ts (atomic age-gated DELETE; no TOCTOU). Healthy refreshing holders survive by construction; only wedged-but-alive holders trip. D-V3-5 partial-status consumer threading (conservative posture matching blocked_by_failures): - printSyncResult: new `case 'partial':` arm reports filesImported + reason; tells operator to re-run to continue. - manageGitignore (both single-source and parallel runOne sites, plus watch mode): excludes partial from the gate. A partial sync's db_only path set isn't fully reconciled. - Auto-embed-backfill enqueue inside runOne: excludes partial. The next clean sync will re-walk and re-decide. CLI flag parsing (T16): - parseDurationSeconds in sync-concurrency.ts: accepts 60s/10m/1h/bare int; rejects 0/negatives/decimals/garbage. Names the failing flag in the error message. - --timeout requires --source OR --all (validation rejects bare `gbrain sync --timeout`). - --max-age requires --break-lock; mutually exclusive with --force-break-lock. Coverage: - 15 unit cases (test/sync-timeout.test.ts) pin parseDurationSeconds + SyncResult union additivity. - 2 E2E cases (test/e2e/sync-parallel.test.ts) pin the abort-mid-import contract against real Postgres: status='partial', last_commit unchanged, filesImported bounded. Closes #1472 (RFC from @garrytan-agents) — CLI surface; schema foundation landed in the previous commit. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(heavy): sync_timeout_rescue.sh reproducer for the cron-cascade 10K-page seed × 4 sources × deliberately tight --timeout × 3 sequential cron emulations. Asserts every source reaches `last_commit === HEAD` within 3 waves. Proves the v0.41.15.0 fix breaks the cascade the PR #1472 RFC documented. Workload (tests/heavy/_sync_timeout_rescue_workload.ts) is PGLite-only because the PGLite engine forces serial sync internally (parallelEligible excludes it). The parallel-fan-out + per-source AbortController case lives in test/e2e/sync-parallel.test.ts against real Postgres. This heavy test pins the contract that matters for cron: aborts → partial returns → next wave content_hash-short-circuits + makes new progress. Smoke-tested locally at PAGES=50 WAVES=2 TIMEOUT_SECONDS=2: every source converges within 2 waves. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(v0.41.15.0): CHANGELOG + README + TODOS + version bump Bumps VERSION + package.json to 0.41.15.0 (next slot after master's v0.41.14.0). CHANGELOG entry leads ELI10 per gstack voice rules and documents the 3 intentional honest gaps: 1. --timeout covers pull + delete + rename + import only; extract + embed run to completion (D-V3-1 honest scope). 2. First 30 min after migration v98, --max-age cannot identify wedged pre-upgrade holders (D-V4-1 rollout trade-off). 3. Full-sync triggers (first sync, --full, chunker-version rewalk) don't respect --timeout yet (deferred to v0.42+). README troubleshooting section: paste-ready cron pattern with shell timeout(1) for OS-level process isolation + gbrain's --timeout for graceful self-termination half-a-minute earlier. TODOS.md: v0.42+ entries for subprocess fan-out (revisit if shell timeout(1) proves insufficient), full-sync --timeout coverage via AbortSignal in runImport, and runFactsBackstop microtask-queue process-alive caveat. llms-full.txt regenerated via `bun run build:llms`. Closes #1472 (RFC from @garrytan-agents). Credit to @garrytan-agents in the CHANGELOG for surfacing the production cron-failure data that motivated the work. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(test-isolation): rewrite JSDoc to not match mock.module() lint regex scripts/check-test-isolation.sh greps for the literal string `mock.module(` to flag top-level module mocks (R2 rule — top-level mocks leak across files in the shard process). The regex doesn't know about comments, so my two new test files tripped the lint with JSDoc lines literally describing the rule: test/sync-timeout.test.ts:11 "* `mock.module()` (R2). Engine ..." test/sync-break-lock-all.test.ts:15 "* mock.module(), no process.env ..." Both files had ZERO actual mock.module() calls — only the comment text matched. Rewrote both JSDocs to refer to "top-level module mocks" instead of the literal token. Same meaning; doesn't trip the regex. `bun run check:test-isolation` now passes (714 non-serial unit files scanned). `bun run verify` clean (22/22 checks pass). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |