mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* fix(minions): self-identifying RSS watchdog + cgroup-aware default + pooler-reap self-heal (#1678) Problem 1: distinct WORKER_EXIT_RSS_WATCHDOG exit code + cause-keyed supervisor breaker (bypasses the stable-run reset that hid the 400x/24h loop) + rss_watchdog audit bucket + 80% soft-warn; cgroup-aware resolveDefaultMaxRssMb replaces the flat 2048 default at every spawn site. Problem 2: CONNECTION_ENDED classified retryable; postgres-engine sql getter throws a retryable error on a reaped instance pool instead of the misleading module-singleton fallthrough; promoteDelayed reconnect-retry; claim recovers on the next poll tick (no double-claim); lock-renewal tick reconnect-once dep. * feat(cycle): surface silent extract_atoms backlog + bounded --drain + fix lint clobbering the shared DB connection (#1678) Problem 3: extract_atoms_backlog doctor check + pack_gated skip marker + shared countExtractAtomsBacklog; `gbrain dream --phase extract_atoms --drain [--window N]` single-hold bounded drain (same cycleLockIdFor, rediscover each batch, reports remaining, exits non-zero while work remains). Also fixes a real production bug found via E2E: the cycle lint phase's resolveLintContentSanity created + disconnected a module-style engine that nulled the shared db singleton mid-cycle, breaking every later phase with "connect() has not been called". Lint now reuses the caller's live engine (cycle + Minion handlers thread it; standalone CLI keeps the create-own path). * chore: bump version and changelog (v0.41.39.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1678): pre-landing review — route transaction/withReservedConnection through the sql getter + drain treats failed count as incomplete Codex adversarial review findings: - #2: transaction(), withReservedConnection(), and one other site bypassed the v0.42.2.0 sql-getter self-heal via `this._sql || db.getConnection()`, so a reaped instance pool fell through to the module singleton there. Route all three through `this.sql` so they throw the retryable instance-pool error and recover consistently (MinionQueue.transaction hits this). - #4: `gbrain dream --drain` treated a null backlog count (query failure) as success via `remaining ?? 0`; now null exits EXIT_DRAIN_INCOMPLETE so automation never believes an unverified backlog drained. - #1 (claim orphan) + #3 (PGLite drain lock) documented as follow-ups in TODOS. * docs: document v0.42.2.0 #1678 modules + behavior in CLAUDE.md Adds Key Files entries for worker-exit-codes.ts, rss-default.ts, and extract-atoms-drain.ts, plus v0.42.2.0 annotations on worker.ts, child-worker-supervisor.ts, lock-renewal-tick.ts, and dream.ts. Regenerated llms-full.txt to match (test/build-llms.test.ts gate). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: re-version v0.42.2.0 → v0.42.5.0 across VERSION/package.json/CHANGELOG/docs/comments Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
71 lines
2.6 KiB
TypeScript
71 lines
2.6 KiB
TypeScript
/**
|
|
* issue #1678 — Minion hot-path lock recovery contract.
|
|
*
|
|
* - promoteDelayed (idempotent) self-heals: a reaped-socket CONNECTION_ENDED
|
|
* triggers a reconnect + retry against a fresh pool.
|
|
* - claim does NOT retry inline (Codex #1): blind-retrying a claim whose
|
|
* UPDATE...RETURNING may have committed could double-claim a job. The error
|
|
* propagates; the worker poll loop reconnects + re-claims on the next tick.
|
|
*
|
|
* Hermetic: a fake BrainEngine whose executeRaw is scripted; no real DB.
|
|
*/
|
|
|
|
import { describe, it, expect } from 'bun:test';
|
|
import { tmpdir } from 'os';
|
|
import { join } from 'path';
|
|
import { MinionQueue } from '../src/core/minions/queue.ts';
|
|
import { withEnv } from './helpers/with-env.ts';
|
|
|
|
function connEndedError(): Error & { code: string } {
|
|
const e = new Error('write CONNECTION_ENDED localhost:6543') as Error & { code: string };
|
|
e.code = 'CONNECTION_ENDED';
|
|
return e;
|
|
}
|
|
|
|
const AUDIT_DIR = join(tmpdir(), `gbrain-queue-lock-retry-${process.pid}-${Date.now()}`);
|
|
// Fast retry + isolated audit dir so the test doesn't sleep ~1s or pollute ~/.gbrain.
|
|
const FAST_ENV = {
|
|
GBRAIN_BULK_RETRY_BASE_MS: '1',
|
|
GBRAIN_BULK_RETRY_MAX_MS: '2',
|
|
GBRAIN_AUDIT_DIR: AUDIT_DIR,
|
|
};
|
|
|
|
describe('MinionQueue lock-path recovery (issue #1678)', () => {
|
|
it('promoteDelayed reconnects + retries on a reaped-socket error', async () => {
|
|
await withEnv(FAST_ENV, async () => {
|
|
let calls = 0;
|
|
let reconnects = 0;
|
|
const engine = {
|
|
kind: 'postgres',
|
|
executeRaw: async () => {
|
|
calls++;
|
|
if (calls === 1) throw connEndedError();
|
|
return [];
|
|
},
|
|
reconnect: async () => { reconnects++; },
|
|
} as unknown as ConstructorParameters<typeof MinionQueue>[0];
|
|
|
|
const q = new MinionQueue(engine);
|
|
const out = await q.promoteDelayed();
|
|
expect(out).toEqual([]);
|
|
expect(calls).toBe(2); // first attempt threw, retry succeeded
|
|
expect(reconnects).toBe(1); // reconnect fired between attempts
|
|
});
|
|
});
|
|
|
|
it('claim does NOT retry inline on a reaped-socket error (Codex #1 double-claim guard)', async () => {
|
|
await withEnv(FAST_ENV, async () => {
|
|
let calls = 0;
|
|
const engine = {
|
|
kind: 'postgres',
|
|
executeRaw: async () => { calls++; throw connEndedError(); },
|
|
reconnect: async () => {},
|
|
} as unknown as ConstructorParameters<typeof MinionQueue>[0];
|
|
|
|
const q = new MinionQueue(engine);
|
|
await expect(q.claim('tok', 1000, 'default', ['sync'])).rejects.toThrow('CONNECTION_ENDED');
|
|
expect(calls).toBe(1); // exactly one attempt — no inline retry
|
|
});
|
|
});
|
|
});
|