mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-31 04:07:52 +00:00
* fix(minions): self-identifying RSS watchdog + cgroup-aware default + pooler-reap self-heal (#1678) Problem 1: distinct WORKER_EXIT_RSS_WATCHDOG exit code + cause-keyed supervisor breaker (bypasses the stable-run reset that hid the 400x/24h loop) + rss_watchdog audit bucket + 80% soft-warn; cgroup-aware resolveDefaultMaxRssMb replaces the flat 2048 default at every spawn site. Problem 2: CONNECTION_ENDED classified retryable; postgres-engine sql getter throws a retryable error on a reaped instance pool instead of the misleading module-singleton fallthrough; promoteDelayed reconnect-retry; claim recovers on the next poll tick (no double-claim); lock-renewal tick reconnect-once dep. * feat(cycle): surface silent extract_atoms backlog + bounded --drain + fix lint clobbering the shared DB connection (#1678) Problem 3: extract_atoms_backlog doctor check + pack_gated skip marker + shared countExtractAtomsBacklog; `gbrain dream --phase extract_atoms --drain [--window N]` single-hold bounded drain (same cycleLockIdFor, rediscover each batch, reports remaining, exits non-zero while work remains). Also fixes a real production bug found via E2E: the cycle lint phase's resolveLintContentSanity created + disconnected a module-style engine that nulled the shared db singleton mid-cycle, breaking every later phase with "connect() has not been called". Lint now reuses the caller's live engine (cycle + Minion handlers thread it; standalone CLI keeps the create-own path). * chore: bump version and changelog (v0.41.39.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#1678): pre-landing review — route transaction/withReservedConnection through the sql getter + drain treats failed count as incomplete Codex adversarial review findings: - #2: transaction(), withReservedConnection(), and one other site bypassed the v0.42.2.0 sql-getter self-heal via `this._sql || db.getConnection()`, so a reaped instance pool fell through to the module singleton there. Route all three through `this.sql` so they throw the retryable instance-pool error and recover consistently (MinionQueue.transaction hits this). - #4: `gbrain dream --drain` treated a null backlog count (query failure) as success via `remaining ?? 0`; now null exits EXIT_DRAIN_INCOMPLETE so automation never believes an unverified backlog drained. - #1 (claim orphan) + #3 (PGLite drain lock) documented as follow-ups in TODOS. * docs: document v0.42.2.0 #1678 modules + behavior in CLAUDE.md Adds Key Files entries for worker-exit-codes.ts, rss-default.ts, and extract-atoms-drain.ts, plus v0.42.2.0 annotations on worker.ts, child-worker-supervisor.ts, lock-renewal-tick.ts, and dream.ts. Regenerated llms-full.txt to match (test/build-llms.test.ts gate). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: re-version v0.42.2.0 → v0.42.5.0 across VERSION/package.json/CHANGELOG/docs/comments Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
70 lines
2.8 KiB
TypeScript
70 lines
2.8 KiB
TypeScript
/**
|
|
* Test fixture: spawns a MinionSupervisor with options parsed from env vars.
|
|
*
|
|
* Used by test/supervisor.test.ts integration tests. Separate file because
|
|
* the supervisor calls `process.exit()` at the end of its lifecycle — tests
|
|
* spawn this runner as a subprocess to observe exit codes and audit events
|
|
* without killing the test runner itself.
|
|
*
|
|
* Env vars (all optional, sensible defaults for tests):
|
|
* SUP_CLI_PATH — worker binary path (default: /bin/sh exit-1 script)
|
|
* SUP_PID_FILE — PID file path (REQUIRED; each test uses a unique one)
|
|
* SUP_MAX_CRASHES — max consecutive crashes (default: 3)
|
|
* SUP_BACKOFF_FLOOR_MS — test-only short backoff (default: 1)
|
|
* SUP_HEALTH_INTERVAL_MS — how often healthCheck fires (default: 999_999 off)
|
|
* SUP_ALLOW_SHELL_JOBS — "1" to set allowShellJobs:true, else false
|
|
* SUP_QUEUE — queue name (default: 'default')
|
|
* SUP_AUDIT_DIR — GBRAIN_AUDIT_DIR override (default: tmpdir/supervisor-test)
|
|
*/
|
|
|
|
import { MinionSupervisor } from '../../src/core/minions/supervisor.ts';
|
|
import { writeSupervisorEvent } from '../../src/core/minions/handlers/supervisor-audit.ts';
|
|
import type { BrainEngine } from '../../src/core/engine.ts';
|
|
|
|
// Mock engine: healthCheck() calls engine.executeRaw; return empty rows so
|
|
// the query path exercises without needing Postgres.
|
|
const mockEngine: Partial<BrainEngine> = {
|
|
kind: 'postgres' as const,
|
|
executeRaw: async () => [],
|
|
} as unknown as BrainEngine;
|
|
|
|
const pidFile = process.env.SUP_PID_FILE;
|
|
if (!pidFile) {
|
|
console.error('SUP_PID_FILE env var is required');
|
|
process.exit(99);
|
|
}
|
|
|
|
const cliPath = process.env.SUP_CLI_PATH ?? '/bin/sh';
|
|
const maxCrashes = parseInt(process.env.SUP_MAX_CRASHES ?? '3', 10);
|
|
const backoffFloor = parseInt(process.env.SUP_BACKOFF_FLOOR_MS ?? '1', 10);
|
|
const healthInterval = parseInt(process.env.SUP_HEALTH_INTERVAL_MS ?? '999999', 10);
|
|
const allowShellJobs = process.env.SUP_ALLOW_SHELL_JOBS === '1';
|
|
const queueName = process.env.SUP_QUEUE ?? 'default';
|
|
// SUP_MAX_RSS: when set, pin an explicit watchdog cap (tests the passthrough
|
|
// path). When unset, MinionSupervisor auto-sizes cgroup-aware (issue #1678).
|
|
const maxRssExplicit = process.env.SUP_MAX_RSS !== undefined
|
|
? parseInt(process.env.SUP_MAX_RSS, 10)
|
|
: undefined;
|
|
|
|
if (process.env.SUP_AUDIT_DIR) {
|
|
process.env.GBRAIN_AUDIT_DIR = process.env.SUP_AUDIT_DIR;
|
|
}
|
|
|
|
const supervisorPid = process.pid;
|
|
|
|
const supervisor = new MinionSupervisor(mockEngine as BrainEngine, {
|
|
concurrency: 1,
|
|
queue: queueName,
|
|
pidFile,
|
|
maxCrashes,
|
|
healthInterval,
|
|
cliPath,
|
|
allowShellJobs,
|
|
json: true,
|
|
_backoffFloorMs: backoffFloor,
|
|
...(maxRssExplicit !== undefined ? { maxRssMb: maxRssExplicit } : {}),
|
|
onEvent: (emission) => writeSupervisorEvent(emission, supervisorPid),
|
|
});
|
|
|
|
await supervisor.start();
|