Files
gbrain/docs/operations/spend-controls.md
5c49225e4b v0.42.45.0 feat(sync): delta-aware cost estimator — stop wedging the daily cron (#2139) (#2224)
* feat(core): shared computeSyncDelta + spend-posture module (#2139)

sync-delta.ts: ONE implementation of "what changed since last_commit",
consumed by both the sync executor and the inline cost estimator so the
gate's dollar figure can't drift from what the sync imports.

spend-posture.ts: spend.posture config + parseUsdLimit/formatUsdLimit
off-switch parsing (off/unlimited/none → Infinity; undefined at the budget
boundary so ledger rows never serialize null).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sync): delta-aware cost estimator + non-TTY auto-defer + per-source failure acks (#2139)

The inline-embed cost gate was a ~400x phantom: it priced the entire tree
whenever the working tree was dirty (always, on an active brain), then blocked
the daily cron with exit 2. Now:

- performSyncInner + the estimator both route through computeSyncDelta, so the
  estimate mirrors execution (fetch-first delta; dirty-but-caught-up tree → $0).
- shouldBlockSync is posture-aware; non-TTY above floor AUTO-DEFERS embeds to
  capped backfill jobs (exit 0) instead of wedging — single shared
  runInlineCostGate on both --all and single-source paths.
- --full prices delta + stale backlog (full sync sweeps it inline).
- off/unlimited on the cost knobs; tokenmax bypasses the backfill cap (still
  ledgered) but never the cooldown.
- --skip-failed/--retry-failed scoped per source; the D15 parallel refusal is
  lifted (the #1939 ledger is per-source + lock-serialized).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(config): register spend-control keys + validate spend.posture (#2139)

Adds spend.posture + the five previously --force-only spend knobs to
KNOWN_CONFIG_KEYS so `config set` accepts them directly (removes the
archaeology the issue complained about), and rejects invalid spend.posture
values at set time with a paste-ready hint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reindex,enrich,onboard): spend.posture across the remaining cost gates (#2139)

reindex-code: tokenmax makes the cost gate informational; --max-cost accepts
off/unlimited. enrich + onboard --auto: tokenmax lifts the refuse-without-cap
guardrail and runs UNCAPPED (spend still ledgered by BudgetTracker). Explicit
--max-usd always wins over posture.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: cost-gate, delta estimator, spend-posture, off-switch coverage (#2139)

New sync-delta + sync-cost-estimate unit suites; rewritten cost-gate serial
tests (auto-defer instead of exit 2, posture, off-switch, format split,
single-source); parseUsdLimit/posture-aware shouldBlockSync; backfill cap-off
+ tokenmax-bypass + cooldown-still-refuses; config known-key acceptance.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(spend-controls): single spend-control surface + ref-map + follow-up TODOs (#2139)

New docs/operations/spend-controls.md (every gate, key, default, off switch,
posture interaction); CLAUDE.md reference-map row; two P3 follow-up TODOs
(measured chunk-count gating, per-source defer granularity). llms bundles
regenerated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(spend): SSRF-harden estimator fetch + complete off/uncapped across reindex/enrich/onboard (#2139)

Ship-stage codex pre-landing review caught four P1s in the secondary cost gates:

- The delta estimator's fetch-first ran `git fetch` through the plain git()
  helper, bypassing the GIT_SSRF_FLAGS + GIT_TERMINAL_PROMPT=0 hardening that
  real sync uses. Added `fetchRemote()` to git-remote.ts (same flags as
  pullRepo) and route the estimator through it — a cost preview / dry-run can
  no longer hit a remote through a less-protected path.
- `reindex --max-cost off`, `enrich --max-usd off`, `onboard --auto --max-usd
  off` were parsed but didn't actually proceed/uncap. Now: explicit off (and
  spend.posture=tokenmax) proceed past the confirmation/missing-cap refusal AND
  run uncapped. enrich threads an Infinity sentinel mapped to "no BudgetTracker
  ceiling" (never raw Infinity → no null in audit rows); reindex/onboard use
  their native undefined=uncapped path. Spend still ledgered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.42.45.0)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(KEY_FILES): update sync/embedding/git-remote/reindex entries to post-#2139 truth

document-release pass: the cost-gate entries described the pre-#2139 behavior
(full-tree-ceiling estimator, --skip-failed-rejects-under-parallel, exit-2
confirmation gate). Updated to current truth — delta-aware estimator via the
shared computeSyncDelta, per-source failure acks under parallel, non-TTY
auto-defer (no exit 2), posture-aware shouldBlockSync. Added entries for the
two new core modules (sync-delta.ts, spend-posture.ts) + fetchRemote on
git-remote.ts + reindex --max-cost off.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 15:08:47 -07:00

5.6 KiB

Spend controls

GBrain's embedding-spend gates in one place: every gate, its config key, default, whether it blocks or just informs, how to widen or disable it, and how the spend.posture switch governs all of them.

The orienting idea: GBrain itself is rounding error; the spend that matters is downstream embedding. These gates exist so a routine sync or enrich can't run up an unexpected embedding bill, while never wedging an unattended cron.

spend.posture — one switch for "cost is not my constraint"

gbrain config set spend.posture tokenmax   # all cost gates become informational
gbrain config set spend.posture gated      # default — gates enforce
Value Effect
gated (default) Every cost gate enforces its limit as documented below.
tokenmax Every cost gate prints its estimate and proceeds — informational only. Spend is still recorded to the ledger; posture removes the ceiling, not the accounting.

spend.posture is deliberately separate from search.mode=tokenmax (which governs retrieval payload size, not embedding spend). When a gate fires and search.mode=tokenmax but spend.posture is unset, the gate prints a one-line hint pointing at this switch.

Precedence: an explicit per-call cap (--max-usd N, --max-cost N) always wins over posture. tokenmax only governs the default/absent case — it never overrides a number you typed on the command line.

Off switches (off / unlimited / none)

The USD-limit knobs accept off, unlimited, or none (case-insensitive) to mean "no limit" — no more setting sentinel values like 100000.

  • 0 is not "off". On sync.cost_gate_min_usd, 0 means "block on any nonzero spend" (a real choice). On the backfill caps, 0 falls back to the default.
  • Internally "no limit" is the string unlimited in any printed/JSON output and "no cap" inside the budget tracker — never a raw Infinity (which would serialize to null in ledger rows).

The gates

Gate Config key Default Blocks? Off switch tokenmax
Sync inline-embed cost gate sync.cost_gate_min_usd 0.50 TTY prompt / non-TTY auto-defer off (or 0 = block-on-any) informational
Backfill 24h per-source spend cap embed.backfill_max_usd_per_source_24h 25 refuses submission off (0 → default) bypassed (still ledgered)
Backfill per-job budget embed.backfill_max_usd 10 caps the job's tracker off (0 → default) uncapped (still ledgered)
Backfill cooldown embed.backfill_cooldown_min 10 skips re-submission inside window — (latency knob, not spend) not bypassed
reindex-code cost gate — (preview before re-embed) TTY prompt / non-TTY refuse + exit 2 --max-cost off informational
enrich / onboard --auto --max-usd (per-call) refuse without a cap (non-TTY) --max-usd off runs uncapped (still ledgered)

Sync inline-embed cost gate

Fires only when sync embeds inline (federated_v2 off, or --serial without --no-embed). Under federated_v2 + parallel, embedding is deferred to capped backfill jobs and the gate is informational. The estimate prices the delta — the files this sync will actually import (fetched-first, so it sees commits the run is about to pull) — not the whole tree. A busy brain with a dirty working tree but caught-up commits estimates $0, because an attached-HEAD sync imports only the committed diff.

Behavior above the floor:

  • TTY: prompts [y/N].
  • Non-interactive (cron/agent): auto-defers embeds to capped backfill jobs and exits 0 — it never wedges the pipeline. The backlog drains via the jobs worker or gbrain embed --stale. Pass --yes to embed inline instead.

Output format splits on the explicit --json flag: --json emits a structured envelope; otherwise human text. Every gate message carries paste-ready knobs.

--full re-embeds the stale backlog inline (full sync sweeps it), so a --full estimate is delta + stale backlog, labeled as such.

Estimate labels

  • ~N tokens (delta: changed files since last sync) — the precise estimate.
  • <=N tokens (full-tree ceiling for K source(s): <reasons> …) — a conservative over-count used only when a precise delta can't be computed: a first sync, a chunker version drift (forces a full re-chunk), or git being unavailable. Unchanged files still skip via content_hash at execution, so the ceiling over-states real spend.

Notes & limits

  • Pre-pull window: the gate fetches before estimating, so it prices what the run will pull. If a fetch fails (offline), it estimates against local HEAD and labels the result; the bounded residual is priced on the next run.
  • Single-source gbrain sync carries the same gate as sync --all (it previously embedded inline with no preview).
  • Recovery under parallel: --skip-failed / --retry-failed work under parallel sync (the failure ledger is per-source and lock-serialized) — you no longer have to drop to --serial, which is what used to arm the inline gate.

Escape hatches at a glance

# Never gate this brain on cost:
gbrain config set spend.posture tokenmax

# Widen the sync inline floor to $5:
gbrain config set sync.cost_gate_min_usd 5

# Disable the sync inline floor entirely:
gbrain config set sync.cost_gate_min_usd off

# Lift the backfill 24h spend cap:
gbrain config set embed.backfill_max_usd_per_source_24h off

# Run enrich uncapped non-interactively:
gbrain enrich --max-usd off        # or: gbrain config set spend.posture tokenmax