mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-30 19:49:14 +00:00
* feat(core): shared computeSyncDelta + spend-posture module (#2139) sync-delta.ts: ONE implementation of "what changed since last_commit", consumed by both the sync executor and the inline cost estimator so the gate's dollar figure can't drift from what the sync imports. spend-posture.ts: spend.posture config + parseUsdLimit/formatUsdLimit off-switch parsing (off/unlimited/none → Infinity; undefined at the budget boundary so ledger rows never serialize null). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(sync): delta-aware cost estimator + non-TTY auto-defer + per-source failure acks (#2139) The inline-embed cost gate was a ~400x phantom: it priced the entire tree whenever the working tree was dirty (always, on an active brain), then blocked the daily cron with exit 2. Now: - performSyncInner + the estimator both route through computeSyncDelta, so the estimate mirrors execution (fetch-first delta; dirty-but-caught-up tree → $0). - shouldBlockSync is posture-aware; non-TTY above floor AUTO-DEFERS embeds to capped backfill jobs (exit 0) instead of wedging — single shared runInlineCostGate on both --all and single-source paths. - --full prices delta + stale backlog (full sync sweeps it inline). - off/unlimited on the cost knobs; tokenmax bypasses the backfill cap (still ledgered) but never the cooldown. - --skip-failed/--retry-failed scoped per source; the D15 parallel refusal is lifted (the #1939 ledger is per-source + lock-serialized). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(config): register spend-control keys + validate spend.posture (#2139) Adds spend.posture + the five previously --force-only spend knobs to KNOWN_CONFIG_KEYS so `config set` accepts them directly (removes the archaeology the issue complained about), and rejects invalid spend.posture values at set time with a paste-ready hint. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reindex,enrich,onboard): spend.posture across the remaining cost gates (#2139) reindex-code: tokenmax makes the cost gate informational; --max-cost accepts off/unlimited. enrich + onboard --auto: tokenmax lifts the refuse-without-cap guardrail and runs UNCAPPED (spend still ledgered by BudgetTracker). Explicit --max-usd always wins over posture. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: cost-gate, delta estimator, spend-posture, off-switch coverage (#2139) New sync-delta + sync-cost-estimate unit suites; rewritten cost-gate serial tests (auto-defer instead of exit 2, posture, off-switch, format split, single-source); parseUsdLimit/posture-aware shouldBlockSync; backfill cap-off + tokenmax-bypass + cooldown-still-refuses; config known-key acceptance. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(spend-controls): single spend-control surface + ref-map + follow-up TODOs (#2139) New docs/operations/spend-controls.md (every gate, key, default, off switch, posture interaction); CLAUDE.md reference-map row; two P3 follow-up TODOs (measured chunk-count gating, per-source defer granularity). llms bundles regenerated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(spend): SSRF-harden estimator fetch + complete off/uncapped across reindex/enrich/onboard (#2139) Ship-stage codex pre-landing review caught four P1s in the secondary cost gates: - The delta estimator's fetch-first ran `git fetch` through the plain git() helper, bypassing the GIT_SSRF_FLAGS + GIT_TERMINAL_PROMPT=0 hardening that real sync uses. Added `fetchRemote()` to git-remote.ts (same flags as pullRepo) and route the estimator through it — a cost preview / dry-run can no longer hit a remote through a less-protected path. - `reindex --max-cost off`, `enrich --max-usd off`, `onboard --auto --max-usd off` were parsed but didn't actually proceed/uncap. Now: explicit off (and spend.posture=tokenmax) proceed past the confirmation/missing-cap refusal AND run uncapped. enrich threads an Infinity sentinel mapped to "no BudgetTracker ceiling" (never raw Infinity → no null in audit rows); reindex/onboard use their native undefined=uncapped path. Spend still ledgered. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: bump version and changelog (v0.42.45.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(KEY_FILES): update sync/embedding/git-remote/reindex entries to post-#2139 truth document-release pass: the cost-gate entries described the pre-#2139 behavior (full-tree-ceiling estimator, --skip-failed-rejects-under-parallel, exit-2 confirmation gate). Updated to current truth — delta-aware estimator via the shared computeSyncDelta, per-source failure acks under parallel, non-TTY auto-defer (no exit 2), posture-aware shouldBlockSync. Added entries for the two new core modules (sync-delta.ts, spend-posture.ts) + fetchRemote on git-remote.ts + reindex --max-cost off. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
112 lines
5.6 KiB
Markdown
112 lines
5.6 KiB
Markdown
# Spend controls
|
|
|
|
GBrain's embedding-spend gates in one place: every gate, its config key, default,
|
|
whether it blocks or just informs, how to widen or disable it, and how the
|
|
`spend.posture` switch governs all of them.
|
|
|
|
The orienting idea: **GBrain itself is rounding error; the spend that matters is
|
|
downstream embedding.** These gates exist so a routine sync or enrich can't run up
|
|
an unexpected embedding bill, while never wedging an unattended cron.
|
|
|
|
## `spend.posture` — one switch for "cost is not my constraint"
|
|
|
|
```bash
|
|
gbrain config set spend.posture tokenmax # all cost gates become informational
|
|
gbrain config set spend.posture gated # default — gates enforce
|
|
```
|
|
|
|
| Value | Effect |
|
|
|-------|--------|
|
|
| `gated` (default) | Every cost gate enforces its limit as documented below. |
|
|
| `tokenmax` | Every cost gate prints its estimate and **proceeds** — informational only. Spend is still recorded to the ledger; posture removes the *ceiling*, not the *accounting*. |
|
|
|
|
`spend.posture` is deliberately separate from `search.mode=tokenmax` (which governs
|
|
retrieval payload size, not embedding spend). When a gate fires and
|
|
`search.mode=tokenmax` but `spend.posture` is unset, the gate prints a one-line hint
|
|
pointing at this switch.
|
|
|
|
**Precedence:** an explicit per-call cap (`--max-usd N`, `--max-cost N`) always wins
|
|
over posture. `tokenmax` only governs the default/absent case — it never overrides a
|
|
number you typed on the command line.
|
|
|
|
## Off switches (`off` / `unlimited` / `none`)
|
|
|
|
The USD-limit knobs accept `off`, `unlimited`, or `none` (case-insensitive) to mean
|
|
"no limit" — no more setting sentinel values like `100000`.
|
|
|
|
- `0` is **not** "off". On `sync.cost_gate_min_usd`, `0` means "block on any nonzero
|
|
spend" (a real choice). On the backfill caps, `0` falls back to the default.
|
|
- Internally "no limit" is the string `unlimited` in any printed/JSON output and "no
|
|
cap" inside the budget tracker — never a raw `Infinity` (which would serialize to
|
|
`null` in ledger rows).
|
|
|
|
## The gates
|
|
|
|
| Gate | Config key | Default | Blocks? | Off switch | tokenmax |
|
|
|------|-----------|---------|---------|-----------|----------|
|
|
| Sync inline-embed cost gate | `sync.cost_gate_min_usd` | `0.50` | TTY prompt / non-TTY auto-defer | `off` (or `0` = block-on-any) | informational |
|
|
| Backfill 24h per-source spend cap | `embed.backfill_max_usd_per_source_24h` | `25` | refuses submission | `off` (`0` → default) | bypassed (still ledgered) |
|
|
| Backfill per-job budget | `embed.backfill_max_usd` | `10` | caps the job's tracker | `off` (`0` → default) | uncapped (still ledgered) |
|
|
| Backfill cooldown | `embed.backfill_cooldown_min` | `10` | skips re-submission inside window | — (latency knob, not spend) | **not** bypassed |
|
|
| `reindex-code` cost gate | — (preview before re-embed) | — | TTY prompt / non-TTY refuse + exit 2 | `--max-cost off` | informational |
|
|
| `enrich` / `onboard --auto` | `--max-usd` (per-call) | — | refuse without a cap (non-TTY) | `--max-usd off` | runs uncapped (still ledgered) |
|
|
|
|
### Sync inline-embed cost gate
|
|
|
|
Fires only when sync embeds **inline** (federated_v2 off, or `--serial` without
|
|
`--no-embed`). Under federated_v2 + parallel, embedding is deferred to capped backfill
|
|
jobs and the gate is informational. The estimate prices the **delta** — the files this
|
|
sync will actually import (fetched-first, so it sees commits the run is about to pull) —
|
|
not the whole tree. A busy brain with a dirty working tree but caught-up commits
|
|
estimates `$0`, because an attached-HEAD sync imports only the committed diff.
|
|
|
|
Behavior above the floor:
|
|
- **TTY:** prompts `[y/N]`.
|
|
- **Non-interactive (cron/agent):** **auto-defers** embeds to capped backfill jobs and
|
|
exits 0 — it never wedges the pipeline. The backlog drains via the jobs worker or
|
|
`gbrain embed --stale`. Pass `--yes` to embed inline instead.
|
|
|
|
Output format splits on the explicit `--json` flag: `--json` emits a structured
|
|
envelope; otherwise human text. Every gate message carries paste-ready knobs.
|
|
|
|
`--full` re-embeds the stale backlog inline (full sync sweeps it), so a `--full`
|
|
estimate is `delta + stale backlog`, labeled as such.
|
|
|
|
### Estimate labels
|
|
|
|
- `~N tokens (delta: changed files since last sync)` — the precise estimate.
|
|
- `<=N tokens (full-tree ceiling for K source(s): <reasons> …)` — a conservative
|
|
over-count used only when a precise delta can't be computed: a first sync, a chunker
|
|
version drift (forces a full re-chunk), or git being unavailable. Unchanged files
|
|
still skip via `content_hash` at execution, so the ceiling over-states real spend.
|
|
|
|
## Notes & limits
|
|
|
|
- **Pre-pull window:** the gate fetches before estimating, so it prices what the run
|
|
will pull. If a fetch fails (offline), it estimates against local HEAD and labels the
|
|
result; the bounded residual is priced on the next run.
|
|
- **Single-source `gbrain sync`** carries the same gate as `sync --all` (it previously
|
|
embedded inline with no preview).
|
|
- **Recovery under parallel:** `--skip-failed` / `--retry-failed` work under parallel
|
|
sync (the failure ledger is per-source and lock-serialized) — you no longer have to
|
|
drop to `--serial`, which is what used to arm the inline gate.
|
|
|
|
## Escape hatches at a glance
|
|
|
|
```bash
|
|
# Never gate this brain on cost:
|
|
gbrain config set spend.posture tokenmax
|
|
|
|
# Widen the sync inline floor to $5:
|
|
gbrain config set sync.cost_gate_min_usd 5
|
|
|
|
# Disable the sync inline floor entirely:
|
|
gbrain config set sync.cost_gate_min_usd off
|
|
|
|
# Lift the backfill 24h spend cap:
|
|
gbrain config set embed.backfill_max_usd_per_source_24h off
|
|
|
|
# Run enrich uncapped non-interactively:
|
|
gbrain enrich --max-usd off # or: gbrain config set spend.posture tokenmax
|
|
```
|