Files
gbrain/test/worker-watchdog-trigger.test.ts
T
766604dea0 v0.42.5.0 fix(minions): RSS watchdog opacity + pooler-reap self-heal + silent lens backlog + cycle lint DB-disconnect (#1678) (#1735)
* fix(minions): self-identifying RSS watchdog + cgroup-aware default + pooler-reap self-heal (#1678)

Problem 1: distinct WORKER_EXIT_RSS_WATCHDOG exit code + cause-keyed supervisor
breaker (bypasses the stable-run reset that hid the 400x/24h loop) + rss_watchdog
audit bucket + 80% soft-warn; cgroup-aware resolveDefaultMaxRssMb replaces the
flat 2048 default at every spawn site.

Problem 2: CONNECTION_ENDED classified retryable; postgres-engine sql getter
throws a retryable error on a reaped instance pool instead of the misleading
module-singleton fallthrough; promoteDelayed reconnect-retry; claim recovers on
the next poll tick (no double-claim); lock-renewal tick reconnect-once dep.

* feat(cycle): surface silent extract_atoms backlog + bounded --drain + fix lint clobbering the shared DB connection (#1678)

Problem 3: extract_atoms_backlog doctor check + pack_gated skip marker +
shared countExtractAtomsBacklog; `gbrain dream --phase extract_atoms --drain
[--window N]` single-hold bounded drain (same cycleLockIdFor, rediscover each
batch, reports remaining, exits non-zero while work remains).

Also fixes a real production bug found via E2E: the cycle lint phase's
resolveLintContentSanity created + disconnected a module-style engine that
nulled the shared db singleton mid-cycle, breaking every later phase with
"connect() has not been called". Lint now reuses the caller's live engine
(cycle + Minion handlers thread it; standalone CLI keeps the create-own path).

* chore: bump version and changelog (v0.41.39.0)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#1678): pre-landing review — route transaction/withReservedConnection through the sql getter + drain treats failed count as incomplete

Codex adversarial review findings:
- #2: transaction(), withReservedConnection(), and one other site bypassed the
  v0.42.2.0 sql-getter self-heal via `this._sql || db.getConnection()`, so a
  reaped instance pool fell through to the module singleton there. Route all
  three through `this.sql` so they throw the retryable instance-pool error and
  recover consistently (MinionQueue.transaction hits this).
- #4: `gbrain dream --drain` treated a null backlog count (query failure) as
  success via `remaining ?? 0`; now null exits EXIT_DRAIN_INCOMPLETE so
  automation never believes an unverified backlog drained.
- #1 (claim orphan) + #3 (PGLite drain lock) documented as follow-ups in TODOS.

* docs: document v0.42.2.0 #1678 modules + behavior in CLAUDE.md

Adds Key Files entries for worker-exit-codes.ts, rss-default.ts, and
extract-atoms-drain.ts, plus v0.42.2.0 annotations on worker.ts,
child-worker-supervisor.ts, lock-renewal-tick.ts, and dream.ts. Regenerated
llms-full.txt to match (test/build-llms.test.ts gate).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: re-version v0.42.2.0 → v0.42.5.0 across VERSION/package.json/CHANGELOG/docs/comments

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 22:01:14 -07:00

81 lines
2.6 KiB
TypeScript

/**
* issue #1678 — worker-side RSS watchdog behavior.
*
* Pins that:
* 1. Crossing the cap sets `rssWatchdogTriggered` (the flag the CLI reads to
* exit with WORKER_EXIT_RSS_WATCHDOG) and drains the worker.
* 2. The 80%-of-cap soft warn fires BEFORE the kill, once per crossing,
* carrying the peak + in-flight job kinds — and does NOT drain.
*
* Uses a real in-memory PGLite engine (canonical block per CLAUDE.md R3+R4)
* + a stubbed `getRss` so the test is deterministic and hermetic.
*/
import { describe, it, expect, beforeAll, afterAll } from 'bun:test';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { MinionWorker } from '../src/core/minions/worker.ts';
const MB = 1024 * 1024;
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
describe('worker RSS watchdog (issue #1678)', () => {
it('crossing the cap sets rssWatchdogTriggered and drains', async () => {
const worker = new MinionWorker(engine, {
queue: 'default',
concurrency: 1,
maxRssMb: 100,
getRss: () => 500 * MB, // 5x the cap
rssCheckInterval: 25,
healthCheckInterval: 0, // no self-health timer in this test
pollInterval: 25,
});
worker.register('noop', async () => {});
// start() resolves on its own: the periodic check trips the watchdog,
// gracefulShutdown sets running=false, the loop exits.
await worker.start();
expect(worker.rssWatchdogTriggered).toBe(true);
});
it('80% soft-warn fires before the kill and does not drain', async () => {
const warns: string[] = [];
const origWarn = console.warn;
console.warn = (...a: unknown[]) => { warns.push(a.join(' ')); };
const worker = new MinionWorker(engine, {
queue: 'default',
concurrency: 1,
maxRssMb: 100,
getRss: () => 85 * MB, // 85% — above soft line, below cap
rssCheckInterval: 25,
healthCheckInterval: 0,
pollInterval: 25,
});
worker.register('noop', async () => {});
const runPromise = worker.start();
// Let a couple of periodic checks fire.
await new Promise((r) => setTimeout(r, 120));
worker.stop();
await runPromise;
console.warn = origWarn;
const softWarn = warns.find((w) => w.includes('approaching cap'));
expect(softWarn).toBeDefined();
expect(softWarn).toContain('85%');
// Soft warn must NOT have drained the worker.
expect(worker.rssWatchdogTriggered).toBe(false);
});
});