mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-28 14:59:47 +00:00
* fix: supervisor treats code=0 watchdog exits as crashes
The RSS watchdog triggers gracefulShutdown() which exits with code 0.
The supervisor was counting ALL exits < 5min as crashes, including
clean code=0 exits. After 10 watchdog-triggered restarts (typical with
a 96K-page brain where autopilot inflates RSS), the supervisor gave up
with max_crashes_exceeded.
Fix: code=0 exits reset crashCount to 0 and restart immediately with
no backoff. Only code≠0 exits count toward the crash limit.
Root cause: process.memoryUsage().rss reports 7GB during autopilot
sync on large repos (possibly shared page inflation from git mmap).
The 4096MB threshold triggers on every cycle. This is a separate
issue (RSS measurement accuracy) but the supervisor should handle
clean exits regardless.
* fix: use RssAnon instead of VmRSS for watchdog threshold
process.memoryUsage().rss returns VmRSS which includes file-backed
mmap'd pages. On repos with large git packfiles (96K+ pages), git
operations inflate VmRSS to 7GB+ while actual heap usage is ~100MB.
The kernel reclaims these pages under memory pressure — they're cache.
Replace with /proc/self/status RssAnon + RssShmem which measures only
anonymous pages (heap, stack, anonymous mmap). This is the memory that
actually matters for OOM risk.
Falls back to process.memoryUsage().rss on non-Linux.
Before: watchdog triggers every autopilot cycle (7GB VmRSS > 4GB threshold)
After: watchdog only triggers on real memory growth (~100MB << 4GB threshold)
Related: #1002 (supervisor crash-count fix for the same symptom)
* refactor(minions): extract ChildWorkerSupervisor with D1/D2 amendments
MinionSupervisor and src/commands/autopilot.ts each owned a separate
spawn-and-respawn loop. PR #1003 fixed the supervisor's crash-counter
bug (counting code=0 watchdog drains as crashes) but the autopilot
loop has the same bug class. Worse, the as-shipped #1003 fix reset
crashCount=0 on every code=0 exit, which lost the "flapping worker"
signal in mixed-exit sequences.
Extract the shared spawn loop into ChildWorkerSupervisor so both
consumers compose one tested core. The new class bakes in two
amendments resolved during plan-eng-review:
D1 (lastExitCode track): code=0 exits no longer touch crashCount.
They emit ms:0 backoff and restart immediately, but the counter
survives across them. A worker alternating exit 1 / exit 0 / exit 1
correctly trips max_crashes; a worker drained 100 times by the
watchdog stays at crashCount=0 and runs forever (also correct).
D2 (clean-restart budget): on platforms where the watchdog measures
VmRSS instead of RssAnon (macOS, kernel <4.5, restricted containers),
a perpetually over-threshold worker could clean-exit in a tight loop
with no observability. New `cleanRestartBudget` option (default 10
clean restarts per 60s window) emits a `health_warn` and applies
backoff once exceeded.
The supervisor now delegates spawn/respawn/backoff to the inner
class and maps ChildSupervisorEvent → existing SupervisorEvent
emit() channel so JSONL audit consumers see byte-compatible output.
PID lock, signal handlers, health check, and process.exit on
max-crashes stay in MinionSupervisor (those are standalone-daemon
concerns the autopilot composer doesn't need).
Tests: 6 new ChildWorkerSupervisor cases (D1 classifier, interleaved
exits, stable-run + clean-exit interaction, D2 budget tripping, per-
instance config isolation, event shape regression). Existing supervisor
tests updated to use exit-1 workers where they previously relied on
clean-exit-as-crash semantics; their assertions (env plumbing, PID
lock, audit shape) are unaffected.
Co-Authored-By: Wintermute <wintermute@garrytan.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(autopilot): compose ChildWorkerSupervisor instead of inline spawn loop
src/commands/autopilot.ts:165-197 used to have its own spawn-and-
respawn loop separate from MinionSupervisor's. It hardcoded
maxCrashes=5, fixed 10s backoff, and counted every exit (including
code=0) toward the crash limit. Codex flagged this during plan-eng
review: the parallel implementation had the same bug class fixed
in #1003, just on a different code path. Anyone running
`gbrain autopilot` as a long-running daemon (instead of
`gbrain jobs supervisor`) would hit it.
Replace the inline `startWorker` + `child.on('exit')` block with
a ChildWorkerSupervisor instance. Drops the parallel `crashCount`,
`lastWorkerStartTime`, and `STABLE_RUN_RESET_MS` state. The
ChildWorkerSupervisor's D1 lastExitCode track + D2 clean-restart
budget apply to autopilot for free.
Shutdown now drains via the supervisor's killChild + awaitChildExit
typed surface instead of reaching into `workerProc` directly. The
onMaxCrashesExceeded callback routes through autopilot's existing
shutdown('max_crashes') path so the lockfile gets cleaned up
(pre-refactor, the inline loop called process.exit(1) directly and
bypassed the cleanup).
Regression coverage in test/autopilot-supervisor-wiring.test.ts:
static-shape grep guards for `--max-rss 2048`, `maxCrashes: 5`,
the shutdown-via-callback wiring, and absence of the legacy inline
names (startWorker, workerProc, crashCount, lastWorkerStartTime,
STABLE_RUN_RESET_MS).
Co-Authored-By: Wintermute <wintermute@garrytan.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(worker): parse RssAnon as field-presence + soften OOM docstring
Two follow-ups to the RssAnon watchdog fix (b81c598f), both surfaced
during plan-eng-review by Codex.
M1: getAccurateRss() used `if (anonKb > 0) return ...` to decide
whether to use the /proc/self/status reading or fall back to
process.memoryUsage().rss. That conflated "RssAnon field missing"
(old kernel, non-Linux) with "RssAnon field present but zero" (a
near-empty worker process whose only memory is shmem). The legitimate
shmem-only worker case fell through to VmRSS even though /proc had a
valid reading.
Fix: split the pure parser (parseRssFromProcStatus) into a separate
exported function that checks field presence via regex match, not
value comparison. Returns null only when the field text doesn't
match `^RssAnon:\s+(\d+)` AND `^RssShmem:\s+(\d+)`. Both fields
present + both zero is now a valid reading of 0 bytes.
M2: the docstring claimed RssAnon + RssShmem was "the memory that
actually matters for OOM risk." Codex pushed back: this is correct
for per-process leak detection but NOT a full container-OOM metric,
because cgroup memory pressure includes page cache. Soften to
"non-file-backed resident memory used for per-process leak
detection" and call out the cgroup caveat explicitly.
getAccurateRss now takes an optional readStatus function for
testability. Production callers use the default; tests inject
canned status text to cover the M1 regression and the fallback paths
without mocking the filesystem.
Tests: 11 cases covering parseRssFromProcStatus (normal, M1 regression
with anon=0 + shmem>0, both-zero, missing fields, malformed values,
shmem-only) and getAccurateRss (injected reader, ENOENT fallback,
old-kernel fallback, malformed-value fallback).
Co-Authored-By: Wintermute <wintermute@garrytan.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(minions): awaitChildExit short-circuits when child already exited
Pre-fix, awaitChildExit registered `child.once('exit', ...)` without
checking whether the child had already terminated. If the child drained
between killChild('SIGTERM') and awaitChildExit() — common on fast
SIGTERM responders — Node's 'exit' event had already fired, the late
listener never resolved, and the caller waited out the full timeout.
On the supervisor's clean shutdown path that's a 35-second hang on
every quick child.
Probe `child.exitCode` and `child.signalCode` first; resolve
immediately when either is non-null. Sub-second clean shutdown
restored.
Pre-existing in the legacy supervisor.ts shape (same bug pattern),
but since the refactor consolidates child-process management into one
class, fix the pattern at the new seam.
Regression test in test/child-worker-supervisor.test.ts: run one full
spawn cycle, then call awaitChildExit on the already-finished cycle
and assert it returns in under 200ms (well under any test timeout).
Surfaced during pre-landing /review on the fix wave.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: bump version and changelog (v0.34.3.0)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: update CLAUDE.md key-files entries for v0.34.3.0
Reflects the ChildWorkerSupervisor extraction shipped in this branch:
- Add new entry for src/core/minions/child-worker-supervisor.ts
covering D1 lastExitCode classifier, D2 clean-restart budget, the
awaitChildExit short-circuit, and test pinning at
test/child-worker-supervisor.test.ts
- Update src/core/minions/supervisor.ts entry to note the spawn-loop
extraction into the shared core + the byte-compatible event-shape
mapping that preserves JSONL audit consumers
- Update src/commands/autopilot.ts entry to note the parallel-
supervisor elimination + the shutdown-via-callback wiring
- Update src/core/minions/worker.ts entry with the new RssAnon /
getAccurateRss exports + the M1 field-presence parser fix
Regenerated llms-full.txt to match (per project rule: every CLAUDE.md
edit must be followed by bun run build:llms).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Wintermute <wintermute@garrytan.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
459 lines
18 KiB
TypeScript
459 lines
18 KiB
TypeScript
import { describe, it, expect, afterEach } from 'bun:test';
|
||
import { existsSync, readFileSync, writeFileSync, unlinkSync, chmodSync, mkdirSync, rmSync } from 'fs';
|
||
import { spawn } from 'child_process';
|
||
import { join } from 'path';
|
||
import { tmpdir } from 'os';
|
||
import { readSupervisorEvents, computeSupervisorAuditFilename } from '../src/core/minions/handlers/supervisor-audit.ts';
|
||
import { calculateBackoffMs } from '../src/core/minions/supervisor.ts';
|
||
|
||
const TEST_PID_FILE = '/tmp/gbrain-supervisor-test.pid';
|
||
|
||
afterEach(() => {
|
||
try { unlinkSync(TEST_PID_FILE); } catch { /* noop */ }
|
||
});
|
||
|
||
// ----- Integration test helpers -----
|
||
|
||
interface IntegrationHarness {
|
||
pidFile: string;
|
||
auditDir: string;
|
||
workerScript: string;
|
||
envOutFile: string;
|
||
cleanup: () => void;
|
||
}
|
||
|
||
/** Create per-test temp files + a fake worker shell script. */
|
||
function makeHarness(name: string, workerBody: string): IntegrationHarness {
|
||
const tmpRoot = join(tmpdir(), `gbrain-sup-test-${name}-${process.pid}-${Date.now()}`);
|
||
mkdirSync(tmpRoot, { recursive: true });
|
||
const pidFile = join(tmpRoot, 'supervisor.pid');
|
||
const auditDir = join(tmpRoot, 'audit');
|
||
const workerScript = join(tmpRoot, 'worker.sh');
|
||
const envOutFile = join(tmpRoot, 'env-out.txt');
|
||
|
||
writeFileSync(workerScript, `#!/bin/sh\n${workerBody}\n`, 'utf8');
|
||
chmodSync(workerScript, 0o755);
|
||
|
||
return {
|
||
pidFile,
|
||
auditDir,
|
||
workerScript,
|
||
envOutFile,
|
||
cleanup: () => { try { rmSync(tmpRoot, { recursive: true, force: true }); } catch { /* noop */ } },
|
||
};
|
||
}
|
||
|
||
/**
|
||
* Spawn the supervisor runner as a subprocess. Returns a handle with the
|
||
* child, a promise resolving to exit code + signal, and a kill helper.
|
||
*/
|
||
function spawnSupervisor(h: IntegrationHarness, overrides: Record<string, string> = {}) {
|
||
const env: Record<string, string> = {
|
||
...(process.env as Record<string, string>),
|
||
SUP_PID_FILE: h.pidFile,
|
||
SUP_CLI_PATH: h.workerScript,
|
||
SUP_AUDIT_DIR: h.auditDir,
|
||
SUP_BACKOFF_FLOOR_MS: '5',
|
||
SUP_MAX_CRASHES: '3',
|
||
SUP_HEALTH_INTERVAL_MS: '999999', // effectively off
|
||
...overrides,
|
||
};
|
||
|
||
const child = spawn('bun', [join(import.meta.dir, 'fixtures/supervisor-runner.ts')], {
|
||
env,
|
||
stdio: ['ignore', 'pipe', 'pipe'],
|
||
});
|
||
|
||
let stdout = '';
|
||
let stderr = '';
|
||
child.stdout?.on('data', (d) => { stdout += d.toString(); });
|
||
child.stderr?.on('data', (d) => { stderr += d.toString(); });
|
||
|
||
const exited = new Promise<{ code: number | null; signal: NodeJS.Signals | null }>((resolve) => {
|
||
child.on('exit', (code, signal) => resolve({ code, signal }));
|
||
});
|
||
|
||
return {
|
||
child,
|
||
exited,
|
||
getStdout: () => stdout,
|
||
getStderr: () => stderr,
|
||
};
|
||
}
|
||
|
||
/** Read the audit JSONL for the current week. */
|
||
function readAudit(auditDir: string) {
|
||
const origEnv = process.env.GBRAIN_AUDIT_DIR;
|
||
process.env.GBRAIN_AUDIT_DIR = auditDir;
|
||
try {
|
||
return readSupervisorEvents();
|
||
} finally {
|
||
if (origEnv === undefined) delete process.env.GBRAIN_AUDIT_DIR;
|
||
else process.env.GBRAIN_AUDIT_DIR = origEnv;
|
||
}
|
||
}
|
||
|
||
/** Poll until predicate returns true or deadline elapses. */
|
||
async function waitFor(pred: () => boolean, timeoutMs: number, tickMs = 20): Promise<boolean> {
|
||
const deadline = Date.now() + timeoutMs;
|
||
while (Date.now() < deadline) {
|
||
if (pred()) return true;
|
||
await new Promise(r => setTimeout(r, tickMs));
|
||
}
|
||
return pred();
|
||
}
|
||
|
||
describe('MinionSupervisor', () => {
|
||
describe('calculateBackoffMs', () => {
|
||
it('returns ~1s for first crash', () => {
|
||
const backoff = calculateBackoffMs(0);
|
||
expect(backoff).toBeGreaterThanOrEqual(1000);
|
||
expect(backoff).toBeLessThan(1200); // 1000 + 10% jitter max
|
||
});
|
||
|
||
it('doubles with each crash', () => {
|
||
const b0 = calculateBackoffMs(0);
|
||
const b1 = calculateBackoffMs(1);
|
||
const b2 = calculateBackoffMs(2);
|
||
// Approximate: b1 should be ~2x b0, b2 ~2x b1 (within jitter)
|
||
expect(b1).toBeGreaterThan(1800);
|
||
expect(b2).toBeGreaterThan(3600);
|
||
});
|
||
|
||
it('caps at 60s', () => {
|
||
const backoff = calculateBackoffMs(20); // 2^20 * 1000 would be huge
|
||
expect(backoff).toBeLessThanOrEqual(66_000); // 60s + 10% jitter
|
||
});
|
||
|
||
it('includes jitter (not perfectly deterministic)', () => {
|
||
const values = new Set<number>();
|
||
for (let i = 0; i < 10; i++) {
|
||
values.add(Math.round(calculateBackoffMs(3)));
|
||
}
|
||
// With 10% jitter, we should get some variation
|
||
expect(values.size).toBeGreaterThan(1);
|
||
});
|
||
});
|
||
|
||
describe('PID file management', () => {
|
||
it('detects stale PID files', () => {
|
||
// Write a PID file with a non-existent PID
|
||
writeFileSync(TEST_PID_FILE, '999999999');
|
||
expect(existsSync(TEST_PID_FILE)).toBe(true);
|
||
|
||
// A real supervisor would detect this as stale and overwrite
|
||
const existingPid = parseInt(readFileSync(TEST_PID_FILE, 'utf8').trim(), 10);
|
||
let isAlive = false;
|
||
try {
|
||
process.kill(existingPid, 0);
|
||
isAlive = true;
|
||
} catch {
|
||
isAlive = false;
|
||
}
|
||
expect(isAlive).toBe(false);
|
||
});
|
||
|
||
it('detects live PID files (current process)', () => {
|
||
// Write our own PID
|
||
writeFileSync(TEST_PID_FILE, String(process.pid));
|
||
|
||
const existingPid = parseInt(readFileSync(TEST_PID_FILE, 'utf8').trim(), 10);
|
||
let isAlive = false;
|
||
try {
|
||
process.kill(existingPid, 0);
|
||
isAlive = true;
|
||
} catch {
|
||
isAlive = false;
|
||
}
|
||
expect(isAlive).toBe(true);
|
||
expect(existingPid).toBe(process.pid);
|
||
});
|
||
});
|
||
|
||
describe('crash count tracking', () => {
|
||
it('backoff escalates with crash count', () => {
|
||
const backoffs = [];
|
||
for (let i = 0; i < 7; i++) {
|
||
backoffs.push(calculateBackoffMs(i));
|
||
}
|
||
// Each should be roughly 2x the previous (within jitter)
|
||
for (let i = 1; i < 6; i++) {
|
||
// The base doubles, so even with jitter the next should be > 1.5x previous
|
||
expect(backoffs[i]).toBeGreaterThan(backoffs[i - 1] * 1.5);
|
||
}
|
||
});
|
||
});
|
||
|
||
// --------------------------------------------------------------
|
||
// Integration tests: real spawn(), real signals, real audit file.
|
||
// Each test uses a unique tmpdir harness so they can run in parallel
|
||
// without colliding. `_backoffFloorMs: 5` (set via SUP_BACKOFF_FLOOR_MS)
|
||
// keeps the whole suite under a few seconds.
|
||
// --------------------------------------------------------------
|
||
|
||
describe('integration: crash → restart → max-crashes lifecycle', () => {
|
||
it('respawns the worker after a crash and eventually exits with max-crashes code=1', async () => {
|
||
// Worker always exits with code 1; supervisor should respawn it 3 times,
|
||
// hit max-crashes, then exit via shutdown() with code 1.
|
||
const h = makeHarness('max-crashes', 'exit 1');
|
||
try {
|
||
const sup = spawnSupervisor(h, { SUP_MAX_CRASHES: '3' });
|
||
const { code } = await sup.exited;
|
||
|
||
expect(code).toBe(1);
|
||
|
||
// PID file cleaned up on exit (synchronous process.on('exit') handler).
|
||
expect(existsSync(h.pidFile)).toBe(false);
|
||
|
||
// Audit file should contain started + 3x worker_spawned/worker_exited +
|
||
// max_crashes_exceeded + shutting_down + stopped.
|
||
const events = readAudit(h.auditDir);
|
||
const eventTypes = events.map(e => e.event);
|
||
expect(eventTypes).toContain('started');
|
||
expect(eventTypes.filter(t => t === 'worker_spawned').length).toBeGreaterThanOrEqual(3);
|
||
expect(eventTypes.filter(t => t === 'worker_exited').length).toBeGreaterThanOrEqual(3);
|
||
expect(eventTypes).toContain('max_crashes_exceeded');
|
||
expect(eventTypes).toContain('shutting_down');
|
||
expect(eventTypes).toContain('stopped');
|
||
|
||
// The stopped event should carry exit_code=1 and reason=max_crashes.
|
||
const stoppedEvt = events.filter(e => e.event === 'stopped').pop();
|
||
expect((stoppedEvt as Record<string, unknown>).exit_code).toBe(1);
|
||
expect((stoppedEvt as Record<string, unknown>).reason).toBe('max_crashes');
|
||
} finally {
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: graceful SIGTERM during backoff', () => {
|
||
it('receives SIGTERM while sleeping between crashes and exits 0 cleanly', async () => {
|
||
// Worker always exits with code 1; supervisor has a high max-crashes
|
||
// and a long-enough backoff floor that we can reliably catch it mid-sleep.
|
||
const h = makeHarness('sigterm-backoff', 'exit 1');
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
SUP_MAX_CRASHES: '100',
|
||
SUP_BACKOFF_FLOOR_MS: '800', // 800ms between restarts — enough to catch
|
||
});
|
||
|
||
// Wait until the supervisor has written the PID file AND survived at
|
||
// least one worker_exited (so it's definitely in the backoff sleep).
|
||
const ready = await waitFor(() => {
|
||
if (!existsSync(h.pidFile)) return false;
|
||
const events = readAudit(h.auditDir);
|
||
return events.some(e => e.event === 'worker_exited');
|
||
}, 3000);
|
||
expect(ready).toBe(true);
|
||
|
||
// Now SIGTERM the supervisor. It must exit cleanly within 200ms
|
||
// (short-circuits the 800ms backoff sleep via the stopping flag).
|
||
const sigSentAt = Date.now();
|
||
sup.child.kill('SIGTERM');
|
||
|
||
const { code, signal } = await sup.exited;
|
||
const elapsed = Date.now() - sigSentAt;
|
||
|
||
// Exit code 0 = clean; signal=null means we exited via process.exit, not got killed.
|
||
expect(code).toBe(0);
|
||
expect(signal).toBe(null);
|
||
// Graceful, not hung: exit within 5s (process.exit() through shutdown()
|
||
// should be near-instant; generous bound to tolerate CI slowness).
|
||
expect(elapsed).toBeLessThan(5000);
|
||
|
||
const events = readAudit(h.auditDir);
|
||
const eventTypes = events.map(e => e.event);
|
||
expect(eventTypes).toContain('shutting_down');
|
||
expect(eventTypes).toContain('stopped');
|
||
|
||
const shuttingEvt = events.filter(e => e.event === 'shutting_down').pop();
|
||
expect((shuttingEvt as Record<string, unknown>).reason).toBe('SIGTERM');
|
||
|
||
// PID file cleaned up.
|
||
expect(existsSync(h.pidFile)).toBe(false);
|
||
} finally {
|
||
h.cleanup();
|
||
}
|
||
}, 20_000);
|
||
});
|
||
|
||
describe('integration: env-var inheritance regression (codex #9 / eng #8)', () => {
|
||
it('strips inherited GBRAIN_ALLOW_SHELL_JOBS when allowShellJobs=false, even if parent has it set', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-env-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
// Worker writes env to OUT_FILE then exits 1. exit=1 is required (not
|
||
// exit=0) because post-D1/D2 (v0.33) clean exits don't count toward
|
||
// crashCount — the supervisor would respawn forever. The test's
|
||
// assertion is on the OUT_FILE contents (env plumbing), not the
|
||
// exit code, so any non-zero code that trips SUP_MAX_CRASHES=1 works.
|
||
const h = makeHarness('env-strip-outfile', `printf '%s\\n' "\${GBRAIN_ALLOW_SHELL_JOBS-UNSET}" > "$OUT_FILE" ; exit 1`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
GBRAIN_ALLOW_SHELL_JOBS: '1', // parent has it
|
||
SUP_ALLOW_SHELL_JOBS: '0', // supervisor says NO
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
// Worker should have written "UNSET" (parent env var stripped from child).
|
||
expect(existsSync(outFile)).toBe(true);
|
||
const childSawEnv = readFileSync(outFile, 'utf8').trim();
|
||
expect(childSawEnv).toBe('UNSET');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
|
||
it('DOES pass GBRAIN_ALLOW_SHELL_JOBS to child when allowShellJobs is true', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-env-ok-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
// Worker exits 1 (not 0) so SUP_MAX_CRASHES=1 actually trips. See
|
||
// the comment on the env-strip test above for the v0.33 rationale.
|
||
const h = makeHarness('env-pass-on-opt-in', `printf '%s\\n' "\${GBRAIN_ALLOW_SHELL_JOBS-UNSET}" > "$OUT_FILE" ; exit 1`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
SUP_ALLOW_SHELL_JOBS: '1',
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
expect(existsSync(outFile)).toBe(true);
|
||
expect(readFileSync(outFile, 'utf8').trim()).toBe('1');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: GBRAIN_SUPERVISED env var (v0.22.14)', () => {
|
||
it('sets GBRAIN_SUPERVISED=1 on spawned worker child', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-supervised-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
// exit 1 required post-D1/D2 to trip SUP_MAX_CRASHES=1; clean exits
|
||
// no longer count toward the crash limit.
|
||
const h = makeHarness('supervised-env', `printf '%s\n' "\${GBRAIN_SUPERVISED-UNSET}" > "$OUT_FILE" ; exit 1`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
expect(existsSync(outFile)).toBe(true);
|
||
const childSawEnv = readFileSync(outFile, 'utf8').trim();
|
||
expect(childSawEnv).toBe('1');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('regression (R3): healthInterval=0 disables timer (v0.22.14)', () => {
|
||
// Pre-fix: supervisor unconditionally called setInterval(callback, 0),
|
||
// which schedules a tight loop on the next event-loop tick. The
|
||
// operator-facing CLI claim "Use 0 to disable" was a lie — passing 0
|
||
// produced a DB-probe loop that hammered Postgres.
|
||
//
|
||
// Post-fix: setInterval is gated on healthInterval > 0. With 0, the
|
||
// supervisor runs its supervise loop normally with the health timer
|
||
// entirely absent.
|
||
//
|
||
// Assertion strategy: spawn the supervisor with SUP_HEALTH_INTERVAL_MS=0,
|
||
// a fast worker that exits cleanly, and SUP_MAX_CRASHES=1. A working fix
|
||
// should produce a single worker spawn → exit → supervisor shutdown
|
||
// sequence. If the tight-loop bug returned, the supervisor would still
|
||
// exit (max-crashes path) but the audit trail would show the tell-tale
|
||
// signature of an extremely high health-check call rate during the brief
|
||
// window before max-crashes fires. We assert the basic completion path
|
||
// and let CI's wall-clock detect any pathological CPU spike.
|
||
it('completes a normal supervise lifecycle with healthInterval=0', async () => {
|
||
// exit 1 (not exit 0) because post-D1/D2 (v0.33) clean exits don't
|
||
// count toward max_crashes — a code=0 worker would respawn forever.
|
||
// The test's purpose is regression coverage that healthInterval=0
|
||
// disables the timer; the exit code doesn't matter to that assertion.
|
||
const h = makeHarness('health-interval-zero', 'exit 1');
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
SUP_HEALTH_INTERVAL_MS: '0',
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
const start = Date.now();
|
||
const { code } = await sup.exited;
|
||
const elapsedMs = Date.now() - start;
|
||
|
||
// Clean exit (max-crashes path returns 1; this is fine — we just
|
||
// want to confirm the supervisor reached its terminal state without
|
||
// hanging or runaway looping).
|
||
expect(code).toBe(1);
|
||
|
||
// Sanity: a tight loop on setInterval(0) plus the spawn-respawn
|
||
// loop would still terminate at max-crashes, but it would be
|
||
// measurably slower than a clean run because the event loop is
|
||
// saturated with health-check callbacks. Cap the upper bound at
|
||
// 10s — clean runs typically finish in 1–2s.
|
||
expect(elapsedMs).toBeLessThan(10_000);
|
||
} finally {
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: --max-rss spawn args (v0.21)', () => {
|
||
it('passes --max-rss 2048 to spawned worker by default', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-maxrss-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
// Worker logs its argv to OUT_FILE so the test can assert --max-rss 2048
|
||
// landed there. spawnOnce in supervisor.ts builds:
|
||
// ['jobs', 'work', '--concurrency', '1', '--queue', 'default', '--max-rss', '2048']
|
||
// exit 1 required post-D1/D2: code=0 workers respawn forever.
|
||
const h = makeHarness('maxrss-default', `printf '%s\\n' "$*" > "$OUT_FILE" ; exit 1`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
expect(existsSync(outFile)).toBe(true);
|
||
const argv = readFileSync(outFile, 'utf8').trim();
|
||
expect(argv).toContain('--max-rss 2048');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: audit file rotation + helper', () => {
|
||
it('computeSupervisorAuditFilename returns supervisor-YYYY-Www.jsonl format', () => {
|
||
const jan15_2026 = new Date(Date.UTC(2026, 0, 15)); // Thu
|
||
expect(computeSupervisorAuditFilename(jan15_2026)).toMatch(/^supervisor-2026-W\d\d\.jsonl$/);
|
||
});
|
||
|
||
it('year-boundary ISO week: 2027-01-01 reports as 2026-W53', () => {
|
||
const jan1_2027 = new Date(Date.UTC(2027, 0, 1));
|
||
// ISO week: 2027-01-01 is Friday of W53 of 2026
|
||
expect(computeSupervisorAuditFilename(jan1_2027)).toBe('supervisor-2026-W53.jsonl');
|
||
});
|
||
});
|
||
});
|