mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-30 03:12:32 +00:00
* fix(minions): honest attempt accounting + cooperative abort-honoring + per-handler timeouts (#1737) - Wall-clock and stall dead-letter paths now increment attempts_made (terminal, no retry — wall-clock fires at 2x cumulative timeout; retrying non-idempotent embed/subagent work would duplicate side effects). Surface stalled_counter in jobs get so 'started 3 / stalled 2 / attempts 0' reads true instead of looking like broken accounting. - Thread AbortSignal through embed-backfill/autopilot-cycle -> runPhaseEmbed -> runEmbedCore -> embedAll(Stale)/embedPage, checking it on BOTH --stale and --all paths and between embed batches. A timed-out embed phase now bails within a batch, so the cycle finally releases gbrain_cycle_locks instead of running the full 10-15 min after the job was killed (the daily cycle-wedge). New shared src/core/abort-check.ts (isAborted/throwIfAborted/anySignal). - Per-handler default wall-clock budget (handler-timeouts.ts) stamped at submit for long handlers without an explicit timeout_ms. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(minions): queue-scoped DB supervisor singleton + canonical pidfile + doctor max-rss check (#1849) - Acquire a queue-scoped DB lock (tryAcquireDbLock, keyed on the raw DB identity + queue) on supervisor.start(): a second supervisor on the same (db, queue) fails fast with exit 2 regardless of $HOME/--pid-file. Refresh on a dedicated timer; on refresh failure past the threshold, fail SAFE (exit non-zero) before the TTL could lapse and let a second supervisor take over. Release on shutdown. - Canonical default pidfile keyed on brain id (currentBrainId, config-only, no DB connect) so two brains under one HOME no longer share supervisor.pid. - doctor: new supervisor_singleton check surfaces the effective --max-rss (from the started audit event) and warns when the lock holder's (host,pid) differs from the local pidfile — comparing host+pid, not bare pid. Registered in doctor-categories. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(agent-voice): topic-aware persona context via server-resolved topicId (#1851) Summon Mars/Venus into a specific conversation topic so they boot already knowing the recent thread. A per-topic call link carries only topicId (+ optional display topicName); the server resolves the recent-conversation context from the brain (topics/<topicId>.md) — topic CONTENT is never accepted over the wire (that would be prompt injection + a leak into URLs/referrers/logs). topicId is a strict slug with a path-traversal guard. New '# Topic Context' prompt slot injected after the persona body so identity-first ordering still wins; persona identity unchanged. No topicId -> generic behavior. Contract doc + persona skill docs updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: file #1737 slot-reservation fair-scheduling follow-up TODO (F7) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: bump version and changelog (v0.42.29.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: sync KEY_FILES.md for v0.42.29.0 minions + abort wave (#1737, #1849) Fold the #1737/#1849 behavior into the existing per-file entries and add the two new core files, keeping the doc at current-state truth: - queue.ts: honest attempt accounting on wall-clock + stall dead-letter paths; defaultTimeoutMsFor stamping at submit. - supervisor.ts: queue-scoped DB singleton lock (supervisorLockId, classifySupervisorSingleton, LOCK_LOST, refresh-fail-safe, brain-id pidfile, max_rss_mb audit). - worker-registry.ts: currentDbIdentity(). - New entries: src/core/abort-check.ts, src/core/minions/handler-timeouts.ts. - New doctor.ts extension: supervisor_singleton check. - cycle.ts / embed.ts extensions: AbortSignal threading note. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
83 lines
3.6 KiB
TypeScript
83 lines
3.6 KiB
TypeScript
/**
|
|
* abort-check.ts — one canonical place for cooperative-abort checks (#1737).
|
|
*
|
|
* gbrain has several long-running loops (embed --stale, embed --all, dream
|
|
* cycle phases) that each grew their own `signal?.aborted` check. When a job
|
|
* is killed by the Minions worker (wall-clock timeout, lock loss, SIGTERM) the
|
|
* handler keeps running unless every loop cooperatively checks its signal — and
|
|
* a missed loop is exactly the daily cycle-wedge in #1737: the embed phase ran
|
|
* to completion ignoring the abort, so `gbrain_cycle_locks` stayed held and
|
|
* every later autopilot cycle skipped with `cycle_already_running`.
|
|
*
|
|
* ┌── worker fires job.signal.abort() ──┐
|
|
* │ (wall-clock / lock-loss / SIGTERM) │
|
|
* └──────────────┬─────────────────────┘
|
|
* ▼
|
|
* handler → runPhaseEmbed → runEmbedCore → embedAll(Stale)
|
|
* │ │
|
|
* └─ throwIfAborted(signal) ─────┘ ← bail here, not 15 min later
|
|
* ▼
|
|
* finally releases gbrain_cycle_locks → next cycle runs
|
|
*
|
|
* Two shapes, because the call sites want different control flow:
|
|
* - `isAborted(signal)` — boolean; for loops that `break` cleanly and
|
|
* return partial progress (embed loops).
|
|
* - `throwIfAborted(signal)` — throws an AbortError; for phase boundaries
|
|
* that want to unwind to the cycle's finally.
|
|
*/
|
|
|
|
/** True iff the signal exists and has fired. Null/undefined → never aborted. */
|
|
export function isAborted(signal?: AbortSignal | null): boolean {
|
|
return !!signal?.aborted;
|
|
}
|
|
|
|
/** Error thrown by {@link throwIfAborted}; `name === 'AbortError'`. */
|
|
export class AbortError extends Error {
|
|
constructor(message = 'aborted') {
|
|
super(message);
|
|
this.name = 'AbortError';
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Throw an {@link AbortError} if the signal has fired. The thrown message
|
|
* prefers the signal's own `reason` (the worker sets it to the abort cause —
|
|
* 'wall-clock', 'lock-lost', 'shutdown') so the unwind is self-describing.
|
|
*/
|
|
export function throwIfAborted(signal?: AbortSignal | null, label?: string): void {
|
|
if (!signal?.aborted) return;
|
|
const reason =
|
|
signal.reason instanceof Error
|
|
? signal.reason.message
|
|
: String(signal.reason ?? 'aborted');
|
|
throw new AbortError(label ? `${label}: ${reason}` : reason);
|
|
}
|
|
|
|
/**
|
|
* Compose an external abort signal with an internal one (e.g. a wall-clock
|
|
* budget timer) so a single combined signal fires when EITHER does. Returns
|
|
* the internal signal unchanged when there's no external signal, so callers
|
|
* that never pass one pay nothing. Uses the platform `AbortSignal.any` (Node
|
|
* 20+/Bun) and falls back to a manual relay if it's somehow unavailable.
|
|
*/
|
|
export function anySignal(
|
|
internal: AbortSignal,
|
|
external?: AbortSignal | null,
|
|
): AbortSignal {
|
|
if (!external) return internal;
|
|
if (typeof (AbortSignal as { any?: unknown }).any === 'function') {
|
|
return (AbortSignal as unknown as { any(s: AbortSignal[]): AbortSignal }).any([
|
|
internal,
|
|
external,
|
|
]);
|
|
}
|
|
// Fallback relay (older runtimes): forward whichever fires first.
|
|
const ac = new AbortController();
|
|
const relay = (s: AbortSignal) => ac.abort(s.reason);
|
|
if (internal.aborted) relay(internal);
|
|
else internal.addEventListener('abort', () => relay(internal), { once: true });
|
|
if (external.aborted) relay(external);
|
|
else external.addEventListener('abort', () => relay(external), { once: true });
|
|
return ac.signal;
|
|
}
|