Files
gbrain/src/core/embed-backfill-submit.ts
T
5c49225e4b v0.42.45.0 feat(sync): delta-aware cost estimator — stop wedging the daily cron (#2139) (#2224)
* feat(core): shared computeSyncDelta + spend-posture module (#2139)

sync-delta.ts: ONE implementation of "what changed since last_commit",
consumed by both the sync executor and the inline cost estimator so the
gate's dollar figure can't drift from what the sync imports.

spend-posture.ts: spend.posture config + parseUsdLimit/formatUsdLimit
off-switch parsing (off/unlimited/none → Infinity; undefined at the budget
boundary so ledger rows never serialize null).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sync): delta-aware cost estimator + non-TTY auto-defer + per-source failure acks (#2139)

The inline-embed cost gate was a ~400x phantom: it priced the entire tree
whenever the working tree was dirty (always, on an active brain), then blocked
the daily cron with exit 2. Now:

- performSyncInner + the estimator both route through computeSyncDelta, so the
  estimate mirrors execution (fetch-first delta; dirty-but-caught-up tree → $0).
- shouldBlockSync is posture-aware; non-TTY above floor AUTO-DEFERS embeds to
  capped backfill jobs (exit 0) instead of wedging — single shared
  runInlineCostGate on both --all and single-source paths.
- --full prices delta + stale backlog (full sync sweeps it inline).
- off/unlimited on the cost knobs; tokenmax bypasses the backfill cap (still
  ledgered) but never the cooldown.
- --skip-failed/--retry-failed scoped per source; the D15 parallel refusal is
  lifted (the #1939 ledger is per-source + lock-serialized).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(config): register spend-control keys + validate spend.posture (#2139)

Adds spend.posture + the five previously --force-only spend knobs to
KNOWN_CONFIG_KEYS so `config set` accepts them directly (removes the
archaeology the issue complained about), and rejects invalid spend.posture
values at set time with a paste-ready hint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reindex,enrich,onboard): spend.posture across the remaining cost gates (#2139)

reindex-code: tokenmax makes the cost gate informational; --max-cost accepts
off/unlimited. enrich + onboard --auto: tokenmax lifts the refuse-without-cap
guardrail and runs UNCAPPED (spend still ledgered by BudgetTracker). Explicit
--max-usd always wins over posture.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: cost-gate, delta estimator, spend-posture, off-switch coverage (#2139)

New sync-delta + sync-cost-estimate unit suites; rewritten cost-gate serial
tests (auto-defer instead of exit 2, posture, off-switch, format split,
single-source); parseUsdLimit/posture-aware shouldBlockSync; backfill cap-off
+ tokenmax-bypass + cooldown-still-refuses; config known-key acceptance.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(spend-controls): single spend-control surface + ref-map + follow-up TODOs (#2139)

New docs/operations/spend-controls.md (every gate, key, default, off switch,
posture interaction); CLAUDE.md reference-map row; two P3 follow-up TODOs
(measured chunk-count gating, per-source defer granularity). llms bundles
regenerated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(spend): SSRF-harden estimator fetch + complete off/uncapped across reindex/enrich/onboard (#2139)

Ship-stage codex pre-landing review caught four P1s in the secondary cost gates:

- The delta estimator's fetch-first ran `git fetch` through the plain git()
  helper, bypassing the GIT_SSRF_FLAGS + GIT_TERMINAL_PROMPT=0 hardening that
  real sync uses. Added `fetchRemote()` to git-remote.ts (same flags as
  pullRepo) and route the estimator through it — a cost preview / dry-run can
  no longer hit a remote through a less-protected path.
- `reindex --max-cost off`, `enrich --max-usd off`, `onboard --auto --max-usd
  off` were parsed but didn't actually proceed/uncap. Now: explicit off (and
  spend.posture=tokenmax) proceed past the confirmation/missing-cap refusal AND
  run uncapped. enrich threads an Infinity sentinel mapped to "no BudgetTracker
  ceiling" (never raw Infinity → no null in audit rows); reindex/onboard use
  their native undefined=uncapped path. Spend still ledgered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.42.45.0)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(KEY_FILES): update sync/embedding/git-remote/reindex entries to post-#2139 truth

document-release pass: the cost-gate entries described the pre-#2139 behavior
(full-tree-ceiling estimator, --skip-failed-rejects-under-parallel, exit-2
confirmation gate). Updated to current truth — delta-aware estimator via the
shared computeSyncDelta, per-source failure acks under parallel, non-TTY
auto-defer (no exit 2), posture-aware shouldBlockSync. Added entries for the
two new core modules (sync-delta.ts, spend-posture.ts) + fetchRemote on
git-remote.ts + reindex --max-cost off.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 15:08:47 -07:00

224 lines
8.9 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* Single submission entry point for the `embed-backfill` minion job
* (v0.40 Federated Sync v2 — D19).
*
* Every caller routes through `submitEmbedBackfill`:
* - Parallel sync --all completion (D18)
* - Extended `sync` handler (D22, auto_embed_backfill)
* - POST /webhooks/github
* - `sources federate` / `unfederate` flip hook
* - `gbrain sync trigger`
* - autopilot per-source dispatch (when source is stale and degraded)
*
* Why centralize: D2 added a per-source DB lock at handler entry. That
* protects against double-RUN but not double-SUBMIT — a webhook storm could
* still queue 20 embed-backfill jobs, each one waking, acquiring the lock,
* finding it busy, completing as `already_in_progress`. Net effect: zero
* wasted Voyage spend (D6 budget cap), but high queue churn. Worse, even if
* the lock held, a 30-second idempotency bucket only coalesces SUBMITS within
* the same window. Multi-hour push activity racks up unbounded calls.
*
* D19 layered defenses (composed here):
* 1. Per-source cooldown (default 10min). Refuses submission if the most
* recent embed-backfill for this source finished or is still active
* inside the window.
* 2. Per-source 24h rolling spend cap (default $25). Computed from the
* embed-backfill-tagged rows in the budget audit JSONL. Refuses
* submission when spend has hit the cap.
*
* Both bounds are config-overridable:
* - `embed.backfill_cooldown_min` (default 10)
* - `embed.backfill_max_usd_per_source_24h` (default 25)
*
* Returns a tagged-union status so callers can render the right user signal
* (`gbrain sources status`, webhook response body, sync completion banner).
*/
import type { BrainEngine } from './engine.ts';
import { MinionQueue } from './minions/queue.ts';
import { parseUsdLimit, resolveSpendPosture, type SpendPosture } from './spend-posture.ts';
export const COOLDOWN_CONFIG_KEY = 'embed.backfill_cooldown_min';
export const SPEND_CAP_CONFIG_KEY = 'embed.backfill_max_usd_per_source_24h';
const DEFAULT_COOLDOWN_MIN = 10;
const DEFAULT_SPEND_CAP_USD = 25;
export type SubmitEmbedBackfillStatus =
| 'submitted'
| 'cooldown'
| 'spend_capped';
export interface SubmitEmbedBackfillResult {
status: SubmitEmbedBackfillStatus;
/** Set when status === 'submitted'. */
jobId?: number;
/** Set when status === 'cooldown'. Seconds remaining until cooldown lifts. */
cooldownRemainingSeconds?: number;
/** Set when status === 'spend_capped'. Dollars spent in the 24h window. */
spend24hUsd?: number;
/** Set when status === 'spend_capped'. Active cap. */
spendCapUsd?: number;
/**
* Set true when `spend.posture=tokenmax` waved the job past the 24h spend
* cap (#2139). The spend is still LEDGERED by the per-job BudgetTracker —
* posture removes the ceiling, not the accounting. Cooldown is NOT bypassed
* (it's queue-churn protection, not a spend gate).
*/
spendCapBypassed?: boolean;
}
export interface SubmitEmbedBackfillOpts {
/** Logged into the job's data row for audit. */
reason: string;
/** Override cooldown lookup (tests). */
cooldownMinOverride?: number;
/** Override spend-cap lookup (tests). */
spendCapUsdOverride?: number;
/** Override the 24h spend aggregator (tests). Returns spend in USD. */
spend24hFn?: (engine: BrainEngine, sourceId: string) => Promise<number>;
/** Override `Date.now` (tests). */
nowMs?: number;
/** Job priority. Default 5 (lower than autopilot's 0; above default jobs). */
priority?: number;
/** Override the resolved spend posture (tests). Default: read from config. */
postureOverride?: SpendPosture;
}
/**
* Default 24h-spend aggregator. Reads completed embed-backfill rows in
* `minion_jobs` and sums their token usage × known-price proxy. We
* deliberately do NOT read the budget-tracker audit JSONL — the file
* lives on the worker's local disk and may not be present on the
* machine submitting the job (e.g. a remote MCP serve-http process).
*
* The minion_jobs.tokens_input/output columns ARE populated by the
* subagent flow but NOT by the embed flow (gateway.embed doesn't go
* through the chat-completion roll-up). For v0.40 we use a job-COUNT
* proxy capped at the daily limit: 1 job ≈ default-cap-share. Accepts
* imprecision in exchange for cross-process visibility.
*
* A precise rolling-spend tracker is filed as a v0.41 TODO.
*/
async function defaultSpend24hForSource(
engine: BrainEngine,
sourceId: string,
): Promise<number> {
// Conservative proxy: count jobs that completed (or are running) in the
// 24h window. Each is treated as worth `DEFAULT_SPEND_CAP_USD / 25` ($1)
// toward the cap — i.e. 25 jobs in 24h saturate the default cap.
const rows = await engine.executeRaw<{ n: number }>(
`SELECT COUNT(*)::int AS n
FROM minion_jobs
WHERE name = 'embed-backfill'
AND data->>'sourceId' = $1
AND status IN ('active', 'completed')
AND created_at > NOW() - INTERVAL '24 hours'`,
[sourceId],
);
const jobCount = rows[0]?.n ?? 0;
return jobCount * 1; // $1 / job placeholder; configurable in a later wave.
}
/**
* Look up an integer-valued config key with sane defaults.
* Returns `def` on missing / NaN / non-positive.
*/
async function readIntConfig(
engine: BrainEngine,
key: string,
def: number,
): Promise<number> {
const raw = await engine.getConfig(key);
if (raw === null || raw === undefined) return def;
const n = Number(raw);
return Number.isFinite(n) && n > 0 ? n : def;
}
export async function submitEmbedBackfill(
engine: BrainEngine,
sourceId: string,
opts: SubmitEmbedBackfillOpts,
): Promise<SubmitEmbedBackfillResult> {
const now = opts.nowMs ?? Date.now();
const cooldownMin =
opts.cooldownMinOverride ??
(await readIntConfig(engine, COOLDOWN_CONFIG_KEY, DEFAULT_COOLDOWN_MIN));
// v0.42.42.0 (#2139): spend cap honors `off`/`unlimited`/`none` → Infinity.
// `0` still falls back to the default (off semantics ≠ 0).
const spendCap =
opts.spendCapUsdOverride ??
(raw => parseUsdLimit(raw, DEFAULT_SPEND_CAP_USD))(await engine.getConfig(SPEND_CAP_CONFIG_KEY));
const posture = opts.postureOverride ?? (await resolveSpendPosture(engine));
// ── Source-level cooldown ─────────────────────────────────────
// Block re-submission if (a) an embed-backfill is currently active for this
// source, OR (b) the most-recent embed-backfill finished within the
// cooldown window.
const lastJob = await engine.executeRaw<{
finished_at: Date | null;
status: string;
}>(
`SELECT finished_at, status
FROM minion_jobs
WHERE name = 'embed-backfill'
AND data->>'sourceId' = $1
ORDER BY id DESC LIMIT 1`,
[sourceId],
);
if (lastJob[0]) {
if (lastJob[0].status === 'active' || lastJob[0].status === 'waiting') {
// Active or waiting: no cooldown-remaining number (would be misleading).
return { status: 'cooldown' };
}
if (lastJob[0].finished_at) {
const finishedMs = new Date(lastJob[0].finished_at).getTime();
const ageMs = now - finishedMs;
const cooldownMs = cooldownMin * 60 * 1000;
if (ageMs < cooldownMs) {
return {
status: 'cooldown',
cooldownRemainingSeconds: Math.ceil((cooldownMs - ageMs) / 1000),
};
}
}
}
// ── 24h rolling spend cap ─────────────────────────────────────
// v0.42.42.0 (#2139): `spend.posture=tokenmax` waves past the cap (the
// operator declared cost isn't the constraint). The per-job BudgetTracker
// still ledgers the spend — posture removes the ceiling, not the accounting.
// An `off`/`unlimited` cap (Infinity) is likewise never tripped.
const spend24hFn = opts.spend24hFn ?? defaultSpend24hForSource;
const spend24h = await spend24hFn(engine, sourceId);
const spendCapBypassed = posture === 'tokenmax' && spend24h >= spendCap;
if (spend24h >= spendCap && !spendCapBypassed) {
return {
status: 'spend_capped',
spend24hUsd: spend24h,
spendCapUsd: spendCap,
};
}
// ── Submission ────────────────────────────────────────────────
const queue = new MinionQueue(engine);
const job = await queue.add(
'embed-backfill',
{ sourceId, batchSize: 500, reason: opts.reason },
{
priority: opts.priority ?? 5,
idempotency_key: `embed-backfill:${sourceId}:${bucketize(now, 5 * 60_000)}`,
maxWaiting: 1,
},
);
return spendCapBypassed
? { status: 'submitted', jobId: job.id, spendCapBypassed: true, spend24hUsd: spend24h }
: { status: 'submitted', jobId: job.id };
}
/** Round timestamp down to the nearest `bucketMs` boundary. */
function bucketize(ms: number, bucketMs: number): number {
return Math.floor(ms / bucketMs);
}