mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-28 06:23:01 +00:00
* feat(minions): add self-health-monitoring to bare worker mode
Bare `gbrain jobs work` (without supervisor) previously had zero health
monitoring. If the Postgres connection dropped or the worker's event loop
deadlocked, the process stayed alive doing nothing — jobs piled up while
external process managers (systemd, Docker, cron) thought it was healthy.
Changes:
1. **Self-health-check timer** (worker.ts): Runs every 60s when not under
a supervisor. Two probes:
- DB liveness: `SELECT 1` — 3 consecutive failures → exit(1)
- Stall detection: waiting jobs + 0 in-flight + no completions for 5m
→ warning; 10m → exit(1)
2. **GBRAIN_SUPERVISED env var** (supervisor.ts): Supervisor sets this on
its child worker to prevent duplicate health checks. The supervisor
already has its own health monitoring.
3. **RSS watchdog default** (jobs.ts): Bare workers now default to
`--max-rss 2048` (matching supervisor default). Opt out: `--max-rss 0`.
4. **--health-interval flag** (jobs.ts): Configurable health check period.
`--health-interval 0` disables. Default: 60000ms.
5. **parseMaxRssFlag returns undefined** when flag is absent (vs 0), so
callers can distinguish 'not set' from 'explicitly disabled'.
The design ensures bare workers get supervisor-grade monitoring while
remaining compatible with any external process manager — the worker just
exits with code 1 on detected failure, letting the PM handle the restart.
Tests: 4 new tests (3 worker health, 1 supervisor env var). All 178 pass.
* fix(minions): harden bare-worker self-health-check after multi-round review
Layered fixes from 5 rounds of plan-eng-review + codex outside voice on top of
the original PR #503 (feat: bare-worker self-health-monitoring). Every change
below is in service of "fail-stop into the operator's process manager" without
introducing new ways the library can kill its caller.
worker.ts:
- MinionWorker now extends EventEmitter; emits `'unhealthy'` event with structured
reason payload (`db_dead` | `stalled`). CLI subscribes; library no longer calls
process.exit directly.
- emitUnhealthy() falls back to process.exit(1) when listenerCount('unhealthy') === 0
so direct API consumers without a listener inherit the pre-refactor fail-stop
default. Inline paths opt out via healthCheckInterval=0.
- Stall detection: count(*) query now filters by registered handler names
(`AND name = ANY($2::text[])`) so workers with handlers for {embed,sync} don't
false-positive when waiting jobs of unhandled names accumulate.
- Stall exit threshold measured from lastCompletionTime (not from warn-since), so
defaults of 5min warn / 10min exit fire at idle=10min total — matching the
documented contract.
- Recursive setTimeout pattern with running flag replaces setInterval, eliminating
callback overlap on slow DB probes.
- DB liveness probe wrapped in Promise.race against AbortController-driven
timeout (default 10s) so a hung executeRaw can't wedge the recursive chain
forever. Hung probes count as failures and feed dbFailExitAfter.
- Constructor validates stallExitAfterMs > stallWarnAfterMs and throws loudly
on misconfiguration. Internal timer-installation invariants documented inline.
- GBRAIN_SUPERVISED env-var check tightened from `!!process.env.X` to `=== '1'`.
types.ts:
- Added 5 new MinionWorkerOpts fields with documented contracts:
healthCheckInterval, stallWarnAfterMs, stallExitAfterMs, dbFailExitAfter,
dbProbeTimeoutMs.
- Exported `UnhealthyReason` discriminated union for the 'unhealthy' event payload.
supervisor.ts:
- GBRAIN_SUPERVISED=1 injected on the spawned worker child's env so the child's
self-health timer is skipped (no double-monitoring).
- setInterval(callback, healthInterval) gated behind `> 0`, so the
`--health-interval 0` documented disable contract actually disables instead
of producing a tight DB-hammer loop.
jobs.ts:
- `gbrain jobs work` subscribes to 'unhealthy' and calls process.exit(1) at the
CLI layer. Default --max-rss bumped from 0 to 2048 (matches supervisor default;
catches memory-leak stalls that previously went undetected).
- New --health-interval flag with aggressive validation (NaN/negative/sub-1000ms
rejected; parity with --max-rss) on both `jobs work` and `jobs supervisor`.
- `jobs submit --follow` and `jobs smoke` now pass healthCheckInterval=0 to
disable the self-health timer entirely. These are inline/one-shot flows with
no PM to restart them; the no-listener emitUnhealthy fallback could otherwise
trip on a DB blip and kill the user's CLI session.
- parseMaxRssFlag returns `number | undefined` (was `number`) so callers can
distinguish absent (use the default) from explicit-disable (--max-rss 0).
doctor.ts:
- New queue_health subcheck reports RSS-watchdog kills in the last 24h.
Detects via exact-match `error_text = 'aborted: watchdog'` (the worker's
failJob signature when gracefulShutdown('watchdog') aborts in-flight jobs)
scoped to status IN ('dead','failed'). Tight match avoids over-counting parent
jobs that propagate child failures via on_child_fail='fail_parent'.
* test(minions): self-health behavior + regression tests
7 new tests covering the production failure modes that drove the original PR,
plus regressions for fixes landed during multi-round review.
minions.test.ts:
- DB 3-strike → 'unhealthy' event with reason='db_dead' (the production-incident signature)
- DB recovery resets failure counter (no exit on intermittent failures)
- Stall warn-then-exit (clock-driven; idleMs > stallExitAfterMs is the new contract)
- inFlight > 0 blocks stall detection (long-running legitimate jobs don't false-trip)
- Regression for D1 fix: jobs of unregistered handler names don't trigger stall exit;
also captures the SQL via probe engine and asserts the predicate text contains
`name = ANY` so a future refactor that drops the filter is caught at test time.
- Regression for R3 constructor validation: throws when stallExitAfterMs <= stallWarnAfterMs
(covers both `<` and `=` cases); defaults still construct cleanly.
supervisor.test.ts:
- Regression for R3: supervisor with healthInterval=0 completes a normal lifecycle
within 10s. A tight setInterval(0) loop (the bug we fixed) would saturate the
event loop and slow this past the cap.
Tests use a Proxy-based engine helper (makeProbeEngine) that intercepts SELECT 1
and the count(*) WHERE status='waiting' query while passing through everything
else to the real PGLite engine. This isolates health-check semantics from claim
plumbing without mocking the entire engine surface.
* docs(v0.22.14): migration walkthrough + follow-up TODOs
skills/migrations/v0.22.14.md (new):
- Pre-flight per-PM restart-policy table (systemd Restart=always, Docker
restart: always, launchd KeepAlive, cron watchdog, supervisord autorestart).
v0.22.14 makes bare-worker behavior fail-stop — without an external restart
loop the worker exits and stays dead. Migration calls this out loudly so
OpenClaw/Hermes-style downstream agents can verify their PM before upgrade.
- Five new MinionWorkerOpts fields documented with defaults and rationale.
- Worker-side process.exit(1) fallback semantics explained: CLI subscribes to
'unhealthy', but direct API consumers without a listener inherit fail-stop.
- AskUserQuestion-driven flow for the --max-rss 2048 default (raise / opt out /
keep) with concrete edits per PM (systemd unit, cron line, Docker compose,
launchctl plist).
- Verification commands (gbrain jobs stats, gbrain doctor --json, RSS check)
and a triage paragraph for opening an issue if anything fails.
TODOS.md:
- v0.22.15 embed cooperative-abort (P0, daily pain): plumb signal through
runPhaseEmbed → embed.ts → embedBatch; check signal.aborted between OpenAI
batch calls and between slugs. Closes the daily wedge where embed > 600s
timeout dead-letters the job but keeps running, holding gbrain_cycle_locks
until the lock TTL expires. PR #503 catches the symptom (worker stalled);
this captures the cause-side fix that's the real production resolution.
- v0.23+ bare-worker engine reconnect parity: extract supervisor's
reconnect-then-fail pattern (#406) into MinionWorker so transient PgBouncer
blips don't force a full process restart.
- v0.23+ minion_workers heartbeat table for queue_health doctor check (B7
follow-up): replace lock_until proxy with ground-truth worker liveness
signal so doctor stops crying wolf on legitimately idle workers.
* chore: bump version and changelog (v0.22.14)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Wintermute <wintermute@garrytan.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
445 lines
17 KiB
TypeScript
445 lines
17 KiB
TypeScript
import { describe, it, expect, afterEach } from 'bun:test';
|
||
import { existsSync, readFileSync, writeFileSync, unlinkSync, chmodSync, mkdirSync, rmSync } from 'fs';
|
||
import { spawn } from 'child_process';
|
||
import { join } from 'path';
|
||
import { tmpdir } from 'os';
|
||
import { readSupervisorEvents, computeSupervisorAuditFilename } from '../src/core/minions/handlers/supervisor-audit.ts';
|
||
import { calculateBackoffMs } from '../src/core/minions/supervisor.ts';
|
||
|
||
const TEST_PID_FILE = '/tmp/gbrain-supervisor-test.pid';
|
||
|
||
afterEach(() => {
|
||
try { unlinkSync(TEST_PID_FILE); } catch { /* noop */ }
|
||
});
|
||
|
||
// ----- Integration test helpers -----
|
||
|
||
interface IntegrationHarness {
|
||
pidFile: string;
|
||
auditDir: string;
|
||
workerScript: string;
|
||
envOutFile: string;
|
||
cleanup: () => void;
|
||
}
|
||
|
||
/** Create per-test temp files + a fake worker shell script. */
|
||
function makeHarness(name: string, workerBody: string): IntegrationHarness {
|
||
const tmpRoot = join(tmpdir(), `gbrain-sup-test-${name}-${process.pid}-${Date.now()}`);
|
||
mkdirSync(tmpRoot, { recursive: true });
|
||
const pidFile = join(tmpRoot, 'supervisor.pid');
|
||
const auditDir = join(tmpRoot, 'audit');
|
||
const workerScript = join(tmpRoot, 'worker.sh');
|
||
const envOutFile = join(tmpRoot, 'env-out.txt');
|
||
|
||
writeFileSync(workerScript, `#!/bin/sh\n${workerBody}\n`, 'utf8');
|
||
chmodSync(workerScript, 0o755);
|
||
|
||
return {
|
||
pidFile,
|
||
auditDir,
|
||
workerScript,
|
||
envOutFile,
|
||
cleanup: () => { try { rmSync(tmpRoot, { recursive: true, force: true }); } catch { /* noop */ } },
|
||
};
|
||
}
|
||
|
||
/**
|
||
* Spawn the supervisor runner as a subprocess. Returns a handle with the
|
||
* child, a promise resolving to exit code + signal, and a kill helper.
|
||
*/
|
||
function spawnSupervisor(h: IntegrationHarness, overrides: Record<string, string> = {}) {
|
||
const env: Record<string, string> = {
|
||
...(process.env as Record<string, string>),
|
||
SUP_PID_FILE: h.pidFile,
|
||
SUP_CLI_PATH: h.workerScript,
|
||
SUP_AUDIT_DIR: h.auditDir,
|
||
SUP_BACKOFF_FLOOR_MS: '5',
|
||
SUP_MAX_CRASHES: '3',
|
||
SUP_HEALTH_INTERVAL_MS: '999999', // effectively off
|
||
...overrides,
|
||
};
|
||
|
||
const child = spawn('bun', [join(import.meta.dir, 'fixtures/supervisor-runner.ts')], {
|
||
env,
|
||
stdio: ['ignore', 'pipe', 'pipe'],
|
||
});
|
||
|
||
let stdout = '';
|
||
let stderr = '';
|
||
child.stdout?.on('data', (d) => { stdout += d.toString(); });
|
||
child.stderr?.on('data', (d) => { stderr += d.toString(); });
|
||
|
||
const exited = new Promise<{ code: number | null; signal: NodeJS.Signals | null }>((resolve) => {
|
||
child.on('exit', (code, signal) => resolve({ code, signal }));
|
||
});
|
||
|
||
return {
|
||
child,
|
||
exited,
|
||
getStdout: () => stdout,
|
||
getStderr: () => stderr,
|
||
};
|
||
}
|
||
|
||
/** Read the audit JSONL for the current week. */
|
||
function readAudit(auditDir: string) {
|
||
const origEnv = process.env.GBRAIN_AUDIT_DIR;
|
||
process.env.GBRAIN_AUDIT_DIR = auditDir;
|
||
try {
|
||
return readSupervisorEvents();
|
||
} finally {
|
||
if (origEnv === undefined) delete process.env.GBRAIN_AUDIT_DIR;
|
||
else process.env.GBRAIN_AUDIT_DIR = origEnv;
|
||
}
|
||
}
|
||
|
||
/** Poll until predicate returns true or deadline elapses. */
|
||
async function waitFor(pred: () => boolean, timeoutMs: number, tickMs = 20): Promise<boolean> {
|
||
const deadline = Date.now() + timeoutMs;
|
||
while (Date.now() < deadline) {
|
||
if (pred()) return true;
|
||
await new Promise(r => setTimeout(r, tickMs));
|
||
}
|
||
return pred();
|
||
}
|
||
|
||
describe('MinionSupervisor', () => {
|
||
describe('calculateBackoffMs', () => {
|
||
it('returns ~1s for first crash', () => {
|
||
const backoff = calculateBackoffMs(0);
|
||
expect(backoff).toBeGreaterThanOrEqual(1000);
|
||
expect(backoff).toBeLessThan(1200); // 1000 + 10% jitter max
|
||
});
|
||
|
||
it('doubles with each crash', () => {
|
||
const b0 = calculateBackoffMs(0);
|
||
const b1 = calculateBackoffMs(1);
|
||
const b2 = calculateBackoffMs(2);
|
||
// Approximate: b1 should be ~2x b0, b2 ~2x b1 (within jitter)
|
||
expect(b1).toBeGreaterThan(1800);
|
||
expect(b2).toBeGreaterThan(3600);
|
||
});
|
||
|
||
it('caps at 60s', () => {
|
||
const backoff = calculateBackoffMs(20); // 2^20 * 1000 would be huge
|
||
expect(backoff).toBeLessThanOrEqual(66_000); // 60s + 10% jitter
|
||
});
|
||
|
||
it('includes jitter (not perfectly deterministic)', () => {
|
||
const values = new Set<number>();
|
||
for (let i = 0; i < 10; i++) {
|
||
values.add(Math.round(calculateBackoffMs(3)));
|
||
}
|
||
// With 10% jitter, we should get some variation
|
||
expect(values.size).toBeGreaterThan(1);
|
||
});
|
||
});
|
||
|
||
describe('PID file management', () => {
|
||
it('detects stale PID files', () => {
|
||
// Write a PID file with a non-existent PID
|
||
writeFileSync(TEST_PID_FILE, '999999999');
|
||
expect(existsSync(TEST_PID_FILE)).toBe(true);
|
||
|
||
// A real supervisor would detect this as stale and overwrite
|
||
const existingPid = parseInt(readFileSync(TEST_PID_FILE, 'utf8').trim(), 10);
|
||
let isAlive = false;
|
||
try {
|
||
process.kill(existingPid, 0);
|
||
isAlive = true;
|
||
} catch {
|
||
isAlive = false;
|
||
}
|
||
expect(isAlive).toBe(false);
|
||
});
|
||
|
||
it('detects live PID files (current process)', () => {
|
||
// Write our own PID
|
||
writeFileSync(TEST_PID_FILE, String(process.pid));
|
||
|
||
const existingPid = parseInt(readFileSync(TEST_PID_FILE, 'utf8').trim(), 10);
|
||
let isAlive = false;
|
||
try {
|
||
process.kill(existingPid, 0);
|
||
isAlive = true;
|
||
} catch {
|
||
isAlive = false;
|
||
}
|
||
expect(isAlive).toBe(true);
|
||
expect(existingPid).toBe(process.pid);
|
||
});
|
||
});
|
||
|
||
describe('crash count tracking', () => {
|
||
it('backoff escalates with crash count', () => {
|
||
const backoffs = [];
|
||
for (let i = 0; i < 7; i++) {
|
||
backoffs.push(calculateBackoffMs(i));
|
||
}
|
||
// Each should be roughly 2x the previous (within jitter)
|
||
for (let i = 1; i < 6; i++) {
|
||
// The base doubles, so even with jitter the next should be > 1.5x previous
|
||
expect(backoffs[i]).toBeGreaterThan(backoffs[i - 1] * 1.5);
|
||
}
|
||
});
|
||
});
|
||
|
||
// --------------------------------------------------------------
|
||
// Integration tests: real spawn(), real signals, real audit file.
|
||
// Each test uses a unique tmpdir harness so they can run in parallel
|
||
// without colliding. `_backoffFloorMs: 5` (set via SUP_BACKOFF_FLOOR_MS)
|
||
// keeps the whole suite under a few seconds.
|
||
// --------------------------------------------------------------
|
||
|
||
describe('integration: crash → restart → max-crashes lifecycle', () => {
|
||
it('respawns the worker after a crash and eventually exits with max-crashes code=1', async () => {
|
||
// Worker always exits with code 1; supervisor should respawn it 3 times,
|
||
// hit max-crashes, then exit via shutdown() with code 1.
|
||
const h = makeHarness('max-crashes', 'exit 1');
|
||
try {
|
||
const sup = spawnSupervisor(h, { SUP_MAX_CRASHES: '3' });
|
||
const { code } = await sup.exited;
|
||
|
||
expect(code).toBe(1);
|
||
|
||
// PID file cleaned up on exit (synchronous process.on('exit') handler).
|
||
expect(existsSync(h.pidFile)).toBe(false);
|
||
|
||
// Audit file should contain started + 3x worker_spawned/worker_exited +
|
||
// max_crashes_exceeded + shutting_down + stopped.
|
||
const events = readAudit(h.auditDir);
|
||
const eventTypes = events.map(e => e.event);
|
||
expect(eventTypes).toContain('started');
|
||
expect(eventTypes.filter(t => t === 'worker_spawned').length).toBeGreaterThanOrEqual(3);
|
||
expect(eventTypes.filter(t => t === 'worker_exited').length).toBeGreaterThanOrEqual(3);
|
||
expect(eventTypes).toContain('max_crashes_exceeded');
|
||
expect(eventTypes).toContain('shutting_down');
|
||
expect(eventTypes).toContain('stopped');
|
||
|
||
// The stopped event should carry exit_code=1 and reason=max_crashes.
|
||
const stoppedEvt = events.filter(e => e.event === 'stopped').pop();
|
||
expect((stoppedEvt as Record<string, unknown>).exit_code).toBe(1);
|
||
expect((stoppedEvt as Record<string, unknown>).reason).toBe('max_crashes');
|
||
} finally {
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: graceful SIGTERM during backoff', () => {
|
||
it('receives SIGTERM while sleeping between crashes and exits 0 cleanly', async () => {
|
||
// Worker always exits with code 1; supervisor has a high max-crashes
|
||
// and a long-enough backoff floor that we can reliably catch it mid-sleep.
|
||
const h = makeHarness('sigterm-backoff', 'exit 1');
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
SUP_MAX_CRASHES: '100',
|
||
SUP_BACKOFF_FLOOR_MS: '800', // 800ms between restarts — enough to catch
|
||
});
|
||
|
||
// Wait until the supervisor has written the PID file AND survived at
|
||
// least one worker_exited (so it's definitely in the backoff sleep).
|
||
const ready = await waitFor(() => {
|
||
if (!existsSync(h.pidFile)) return false;
|
||
const events = readAudit(h.auditDir);
|
||
return events.some(e => e.event === 'worker_exited');
|
||
}, 3000);
|
||
expect(ready).toBe(true);
|
||
|
||
// Now SIGTERM the supervisor. It must exit cleanly within 200ms
|
||
// (short-circuits the 800ms backoff sleep via the stopping flag).
|
||
const sigSentAt = Date.now();
|
||
sup.child.kill('SIGTERM');
|
||
|
||
const { code, signal } = await sup.exited;
|
||
const elapsed = Date.now() - sigSentAt;
|
||
|
||
// Exit code 0 = clean; signal=null means we exited via process.exit, not got killed.
|
||
expect(code).toBe(0);
|
||
expect(signal).toBe(null);
|
||
// Graceful, not hung: exit within 5s (process.exit() through shutdown()
|
||
// should be near-instant; generous bound to tolerate CI slowness).
|
||
expect(elapsed).toBeLessThan(5000);
|
||
|
||
const events = readAudit(h.auditDir);
|
||
const eventTypes = events.map(e => e.event);
|
||
expect(eventTypes).toContain('shutting_down');
|
||
expect(eventTypes).toContain('stopped');
|
||
|
||
const shuttingEvt = events.filter(e => e.event === 'shutting_down').pop();
|
||
expect((shuttingEvt as Record<string, unknown>).reason).toBe('SIGTERM');
|
||
|
||
// PID file cleaned up.
|
||
expect(existsSync(h.pidFile)).toBe(false);
|
||
} finally {
|
||
h.cleanup();
|
||
}
|
||
}, 20_000);
|
||
});
|
||
|
||
describe('integration: env-var inheritance regression (codex #9 / eng #8)', () => {
|
||
it('strips inherited GBRAIN_ALLOW_SHELL_JOBS when allowShellJobs=false, even if parent has it set', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-env-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
const h = makeHarness('env-strip-outfile', `printf '%s\\n' "\${GBRAIN_ALLOW_SHELL_JOBS-UNSET}" > "$OUT_FILE" ; exit 0`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
GBRAIN_ALLOW_SHELL_JOBS: '1', // parent has it
|
||
SUP_ALLOW_SHELL_JOBS: '0', // supervisor says NO
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
// Worker should have written "UNSET" (parent env var stripped from child).
|
||
expect(existsSync(outFile)).toBe(true);
|
||
const childSawEnv = readFileSync(outFile, 'utf8').trim();
|
||
expect(childSawEnv).toBe('UNSET');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
|
||
it('DOES pass GBRAIN_ALLOW_SHELL_JOBS to child when allowShellJobs is true', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-env-ok-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
const h = makeHarness('env-pass-on-opt-in', `printf '%s\\n' "\${GBRAIN_ALLOW_SHELL_JOBS-UNSET}" > "$OUT_FILE" ; exit 0`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
SUP_ALLOW_SHELL_JOBS: '1',
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
expect(existsSync(outFile)).toBe(true);
|
||
expect(readFileSync(outFile, 'utf8').trim()).toBe('1');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: GBRAIN_SUPERVISED env var (v0.22.14)', () => {
|
||
it('sets GBRAIN_SUPERVISED=1 on spawned worker child', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-supervised-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
const h = makeHarness('supervised-env', `printf '%s\n' "\${GBRAIN_SUPERVISED-UNSET}" > "$OUT_FILE" ; exit 0`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
expect(existsSync(outFile)).toBe(true);
|
||
const childSawEnv = readFileSync(outFile, 'utf8').trim();
|
||
expect(childSawEnv).toBe('1');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('regression (R3): healthInterval=0 disables timer (v0.22.14)', () => {
|
||
// Pre-fix: supervisor unconditionally called setInterval(callback, 0),
|
||
// which schedules a tight loop on the next event-loop tick. The
|
||
// operator-facing CLI claim "Use 0 to disable" was a lie — passing 0
|
||
// produced a DB-probe loop that hammered Postgres.
|
||
//
|
||
// Post-fix: setInterval is gated on healthInterval > 0. With 0, the
|
||
// supervisor runs its supervise loop normally with the health timer
|
||
// entirely absent.
|
||
//
|
||
// Assertion strategy: spawn the supervisor with SUP_HEALTH_INTERVAL_MS=0,
|
||
// a fast worker that exits cleanly, and SUP_MAX_CRASHES=1. A working fix
|
||
// should produce a single worker spawn → exit → supervisor shutdown
|
||
// sequence. If the tight-loop bug returned, the supervisor would still
|
||
// exit (max-crashes path) but the audit trail would show the tell-tale
|
||
// signature of an extremely high health-check call rate during the brief
|
||
// window before max-crashes fires. We assert the basic completion path
|
||
// and let CI's wall-clock detect any pathological CPU spike.
|
||
it('completes a normal supervise lifecycle with healthInterval=0', async () => {
|
||
const h = makeHarness('health-interval-zero', 'exit 0');
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
SUP_HEALTH_INTERVAL_MS: '0',
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
const start = Date.now();
|
||
const { code } = await sup.exited;
|
||
const elapsedMs = Date.now() - start;
|
||
|
||
// Clean exit (max-crashes path returns 1; this is fine — we just
|
||
// want to confirm the supervisor reached its terminal state without
|
||
// hanging or runaway looping).
|
||
expect(code).toBe(1);
|
||
|
||
// Sanity: a tight loop on setInterval(0) plus the spawn-respawn
|
||
// loop would still terminate at max-crashes, but it would be
|
||
// measurably slower than a clean run because the event loop is
|
||
// saturated with health-check callbacks. Cap the upper bound at
|
||
// 10s — clean runs typically finish in 1–2s.
|
||
expect(elapsedMs).toBeLessThan(10_000);
|
||
} finally {
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: --max-rss spawn args (v0.21)', () => {
|
||
it('passes --max-rss 2048 to spawned worker by default', async () => {
|
||
const outFile = join(tmpdir(), `gbrain-sup-maxrss-${process.pid}-${Date.now()}.txt`);
|
||
try { unlinkSync(outFile); } catch { /* may not exist */ }
|
||
|
||
// Worker logs its argv to OUT_FILE so the test can assert --max-rss 2048
|
||
// landed there. spawnOnce in supervisor.ts builds:
|
||
// ['jobs', 'work', '--concurrency', '1', '--queue', 'default', '--max-rss', '2048']
|
||
const h = makeHarness('maxrss-default', `printf '%s\\n' "$*" > "$OUT_FILE" ; exit 0`);
|
||
|
||
try {
|
||
const sup = spawnSupervisor(h, {
|
||
OUT_FILE: outFile,
|
||
SUP_MAX_CRASHES: '1',
|
||
});
|
||
|
||
await sup.exited;
|
||
|
||
expect(existsSync(outFile)).toBe(true);
|
||
const argv = readFileSync(outFile, 'utf8').trim();
|
||
expect(argv).toContain('--max-rss 2048');
|
||
} finally {
|
||
try { unlinkSync(outFile); } catch { /* noop */ }
|
||
h.cleanup();
|
||
}
|
||
}, 15_000);
|
||
});
|
||
|
||
describe('integration: audit file rotation + helper', () => {
|
||
it('computeSupervisorAuditFilename returns supervisor-YYYY-Www.jsonl format', () => {
|
||
const jan15_2026 = new Date(Date.UTC(2026, 0, 15)); // Thu
|
||
expect(computeSupervisorAuditFilename(jan15_2026)).toMatch(/^supervisor-2026-W\d\d\.jsonl$/);
|
||
});
|
||
|
||
it('year-boundary ISO week: 2027-01-01 reports as 2026-W53', () => {
|
||
const jan1_2027 = new Date(Date.UTC(2027, 0, 1));
|
||
// ISO week: 2027-01-01 is Friday of W53 of 2026
|
||
expect(computeSupervisorAuditFilename(jan1_2027)).toBe('supervisor-2026-W53.jsonl');
|
||
});
|
||
});
|
||
});
|