mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-30 11:22:34 +00:00
* fix: canonical Anthropic model IDs + reverse alias + Opus 4.7 pricing
Replace claude-sonnet-4-6-20250929 with claude-sonnet-4-6 everywhere it
appears as a model ID. Starting with Claude 4.6, Anthropic API IDs are
dateless and pinned — the date suffix was carried forward from Sonnet 4.5
by mistake, producing a phantom ID that 404'd on every call.
Production impact in v0.31.6: isAvailable("chat") returned false in every
code path that loaded the recipe's model list, and extractFactsFromTurn
silently returned []. The headline real-time facts extraction feature
was a no-op on the happy path.
- gateway.ts:46 DEFAULT_CHAT_MODEL -> anthropic:claude-sonnet-4-6
- recipes/anthropic.ts: chat + expansion model lists drop date suffix;
remove wrong-direction alias (claude-sonnet-4-6 -> -20250929);
add reverse alias (-20250929 -> claude-sonnet-4-6) so stale user
configs in models.dream.synthesize etc. keep working
- facts/extract.ts: routes through resolveModel; both fallbacks corrected
- anthropic-pricing.ts: Opus 4.7 corrected $15/$75 -> $5/$25 per
Anthropic docs (the $15/$75 was Opus 4.0 pricing)
- cross-modal-eval/runner.ts: PRICING now reads from ANTHROPIC_PRICING
for Anthropic models instead of duplicating the map (single source of
truth — fixes the drift trap that motivated this whole patch)
Tests: cherry-pick PR #830's test/anthropic-model-ids.test.ts verbatim
(6 recipe-shape guardrails). Update gateway-chat tests to assert reverse
alias resolves correctly. Update budget-meter test for new Opus pricing.
Co-Authored-By: garrytan-agents <garrytan-agents@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: model tier system + recipe-models merge + async reconfigure hook
Add 4-tier model routing (utility/reasoning/deep/subagent) so users can
swap defaults with one config key. Each tier maps to a class of work;
override globally via models.default or per-tier via models.tier.<tier>.
Codex flagged three real architecture issues in the v0.31.12 plan review;
this commit addresses each.
F3 — sync/async timing of configureGateway:
- buildGatewayConfig stays synchronous (pre-engine-connect callers
keep working)
- New reconfigureGatewayWithEngine(engine) async function re-resolves
expansion + chat defaults through resolveModel after engine.connect()
- cli.ts wires the re-stamp into the post-connect path
F4/F5 — softening assertTouchpoint was too broad:
- Earlier plan was to flip native-recipe validation from throw to warn,
affecting gateway.chat AND gateway.expand AND gateway.embed
- Instead: per-gateway-instance recipe-models merge. assertTouchpoint
gets an optional extendedModels Set; when the user opted into a model
via config, it bypasses the throw. Source-code typos still fail fast.
- Existing contract test (test/ai/gateway-chat.test.ts:106) preserved
Tier defaults are TIER_DEFAULTS in model-config.ts. Resolution chain
inserts at step 5 (between models.default and env var). Each existing
resolveModel call site gains a tier: arg — think (deep), cycle/synthesize
(reasoning + utility for verdict), patterns/drift (reasoning), auto-think
(deep), facts/extract (reasoning).
Plus 10 new tests pinning tier precedence, subagent-tier fallback when
models.default is non-Anthropic, and the F6 alias-chain conflict case.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: subagent runtime enforcement for non-Anthropic models (3 layers)
The subagent loop uses Anthropic's Messages API with prompt caching on
system + tools. OpenAI/Google have different shapes. Setting
models.default = openai:gpt-5.5 and routing the subagent there silently
breaks the loop.
Codex F1+F2+F13 in the v0.31.12 plan review pointed out that "warn at
doctor" wasn't enough — handlers/subagent.ts:148 still did
`const model = data.model ?? DEFAULT_MODEL` and called Anthropic directly,
so a job submitted with data.model = openai:gpt-5.5 bypassed any tier
logic and failed at runtime with a confusing provider error.
Three layers of enforcement, defense in depth:
Layer 1 (queue.ts:add) — submit-time guard. When name === 'subagent'
and data.model is set, validate the provider. Non-Anthropic rejects
before the job enters the queue.
Layer 2 (handlers/subagent.ts) — tier-resolution fallback. The handler
routes through resolveModel({ tier: 'subagent' }). If the chain resolves
to a non-Anthropic provider (via models.default or models.tier.subagent),
the resolver warns + falls back to TIER_DEFAULTS.subagent
(claude-sonnet-4-6).
Layer 3 (doctor.ts:checkSubagentProvider) — surfacing layer. Warns when
models.tier.subagent or models.default is explicitly set to a
non-Anthropic provider, with a paste-ready fix command. Lets users see
config drift before submitting a job.
Tests: 3 new cases in test/agent-cli.test.ts asserting the queue-level
guard rejects non-Anthropic data.model. Existing test/subagent-handler
suite still passes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat: gbrain models CLI + doctor probe + silent-no-op regression test
New gbrain models CLI gives the agent and user visibility into routing.
Read mode prints the tier table, current overrides, per-task config,
and aliases with source-of-truth attribution per row. Doctor subcommand
fires a 1-token probe to each configured chat/expansion model and
classifies failures (model_not_found / auth / rate_limit / network /
unknown) so config-time invalid IDs surface without waiting for a
production call that silently degrades.
Per Codex F11 — no specific dollar cost claim in either the help text
or the CHANGELOG (providers have minimum-output billing and prompt-cache
rounding that vary). Probe is opt-in (gbrain doctor --probe-models),
never auto-runs. --skip=<provider> narrows the matrix for cost-sensitive
operators.
Per Codex F7+F8+F15 (the structural regression gap): new
test/facts-extract-silent-no-op.test.ts is THE regression test for the
bug class that motivated v0.31.12. Five cases including the smoking-gun:
when chat IS available, extractFactsFromTurn MUST actually call the chat
transport, not silently return []. Uses the gateway's
__setChatTransportForTests seam so it runs in every shard with no API key.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: bump version and changelog (v0.31.12)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: document v0.31.12 model tier system + gbrain models CLI
Add CLAUDE.md Key Files annotations for the v0.31.12 work:
src/core/model-config.ts (tier system + isAnthropicProvider + TIER_DEFAULTS),
src/core/ai/model-resolver.ts (assertTouchpoint extendedModels arg),
src/core/ai/gateway.ts (reconfigureGatewayWithEngine + extended-models registry),
src/core/minions/queue.ts (subagent submit-time guard, layer 1 of 3),
src/commands/models.ts (new gbrain models CLI + doctor probe),
src/commands/doctor.ts (subagent_provider check, layer 3 of 3),
src/core/ai/recipes/anthropic.ts (canonical model IDs + reverse alias),
src/core/anthropic-pricing.ts (Opus 4.7 corrected to \$5/\$25).
Add CLAUDE.md commands section for gbrain models + gbrain models doctor
+ power-user config recipes. Add README.md command-table rows for the
same. Regenerate llms-full.txt so the bundled docs stay in sync.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: scrub --probe-models reference (flag not actually wired)
The v0.31.12 CHANGELOG and skills/conventions/model-routing.md both
referenced `gbrain doctor --probe-models` as an integrated probe entry
point. The flag was never implemented — only `gbrain models doctor`
landed as the probe surface. Caught by /document-release subagent.
Drop the references rather than wire an untested flag at the last minute.
The probe is reachable via `gbrain models doctor`; users who want it
in doctor's output run that command separately.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: garrytan-agents <garrytan-agents@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
169 lines
5.6 KiB
TypeScript
169 lines
5.6 KiB
TypeScript
/**
|
|
* v0.28: drift dream phase.
|
|
*
|
|
* Detects takes where the underlying evidence has shifted since the take
|
|
* was made. v0.28 ships the SCAFFOLD: the phase iterates active takes,
|
|
* runs a lightweight check against recent timeline entries on the same
|
|
* page, and writes a drift-report-<date>.md if any takes look stale.
|
|
*
|
|
* The full LLM-driven drift detection (compare each take's claim to recent
|
|
* page evidence and propose a weight adjustment) is the v0.29 follow-up.
|
|
* v0.28 lays the phase orchestration so the contract is stable.
|
|
*
|
|
* Default-disabled. Operator opts in:
|
|
* gbrain config set dream.drift.enabled true
|
|
* gbrain config set dream.drift.lookback_days 30
|
|
*/
|
|
|
|
import type { BrainEngine } from '../engine.ts';
|
|
import { BudgetMeter } from './budget-meter.ts';
|
|
import { resolveModel } from '../model-config.ts';
|
|
import type { DreamPhaseResult } from './auto-think.ts';
|
|
|
|
export interface DriftPhaseOpts {
|
|
brainDir?: string;
|
|
dryRun: boolean;
|
|
/** Override the audit ledger path (tests). */
|
|
auditPath?: string;
|
|
}
|
|
|
|
export interface DriftConfig {
|
|
enabled: boolean;
|
|
lookbackDays: number;
|
|
budgetUsd: number;
|
|
autoUpdate: boolean;
|
|
}
|
|
|
|
async function loadDriftConfig(engine: BrainEngine): Promise<DriftConfig> {
|
|
const enabledStr = await engine.getConfig('dream.drift.enabled');
|
|
const lookbackStr = await engine.getConfig('dream.drift.lookback_days');
|
|
const budgetStr = await engine.getConfig('dream.drift.budget');
|
|
const autoStr = await engine.getConfig('dream.drift.auto_update');
|
|
return {
|
|
enabled: enabledStr === 'true',
|
|
lookbackDays: lookbackStr ? Math.max(1, parseInt(lookbackStr, 10) || 30) : 30,
|
|
budgetUsd: budgetStr ? Math.max(0, parseFloat(budgetStr) || 1.0) : 1.0,
|
|
autoUpdate: autoStr === 'true',
|
|
};
|
|
}
|
|
|
|
interface DriftCandidate {
|
|
takeId: number;
|
|
pageSlug: string;
|
|
rowNum: number;
|
|
claim: string;
|
|
weight: number;
|
|
/** Number of timeline entries within the lookback window for the same page. */
|
|
recentEvidenceCount: number;
|
|
}
|
|
|
|
/**
|
|
* Cheap pre-LLM heuristic: takes that have substantial recent timeline
|
|
* evidence on the same page MAY have drifted. Surface them; the v0.29
|
|
* LLM judge will decide if the weight should move.
|
|
*/
|
|
async function findDriftCandidates(
|
|
engine: BrainEngine,
|
|
lookbackDays: number,
|
|
): Promise<DriftCandidate[]> {
|
|
const cutoffMs = Date.now() - lookbackDays * 86_400_000;
|
|
const cutoffIso = new Date(cutoffMs).toISOString().slice(0, 10);
|
|
// Only consider takes with weight in the "soft" middle band (0.3..0.85)
|
|
// — facts (1.0) don't drift, very-low hunches (<0.3) aren't actionable yet.
|
|
const rows = await engine.executeRaw<{
|
|
take_id: number; page_slug: string; row_num: number;
|
|
claim: string; weight: number; recent_evidence: number;
|
|
}>(`
|
|
SELECT t.id AS take_id, p.slug AS page_slug, t.row_num,
|
|
t.claim, t.weight,
|
|
(SELECT count(*)::int FROM timeline_entries te
|
|
WHERE te.page_id = p.id
|
|
AND te.date >= $1::date)
|
|
AS recent_evidence
|
|
FROM takes t
|
|
JOIN pages p ON p.id = t.page_id
|
|
WHERE t.active
|
|
AND t.weight >= 0.3 AND t.weight <= 0.85
|
|
AND t.resolved_at IS NULL
|
|
ORDER BY recent_evidence DESC, t.weight DESC
|
|
LIMIT 200
|
|
`, [cutoffIso]);
|
|
return rows
|
|
.filter(r => Number(r.recent_evidence) >= 1)
|
|
.map(r => ({
|
|
takeId: Number(r.take_id),
|
|
pageSlug: String(r.page_slug),
|
|
rowNum: Number(r.row_num),
|
|
claim: String(r.claim),
|
|
weight: Number(r.weight),
|
|
recentEvidenceCount: Number(r.recent_evidence),
|
|
}));
|
|
}
|
|
|
|
function skipped(_reason: string, detail: string): DreamPhaseResult {
|
|
return { name: 'drift', status: 'skipped', detail, duration_ms: 0 };
|
|
}
|
|
|
|
export async function runPhaseDrift(
|
|
engine: BrainEngine,
|
|
opts: DriftPhaseOpts,
|
|
): Promise<DreamPhaseResult> {
|
|
const start = Date.now();
|
|
const config = await loadDriftConfig(engine);
|
|
if (!config.enabled) {
|
|
return skipped('not_configured', 'dream.drift.enabled is false');
|
|
}
|
|
|
|
const candidates = await findDriftCandidates(engine, config.lookbackDays);
|
|
if (candidates.length === 0) {
|
|
return {
|
|
name: 'drift',
|
|
status: 'complete',
|
|
detail: 'no candidates: no soft-band takes with recent timeline evidence',
|
|
totals: { candidates: 0 },
|
|
duration_ms: Date.now() - start,
|
|
};
|
|
}
|
|
|
|
// Resolve model for the (future v0.29) LLM judge. For v0.28 we just
|
|
// surface the candidates — the meter call is a no-op when we don't actually
|
|
// submit, but resolveModel sets the right pricing key when v0.29 ships.
|
|
const modelId = await resolveModel(engine, {
|
|
configKey: 'models.drift',
|
|
deprecatedConfigKey: 'dream.drift.model',
|
|
tier: 'reasoning',
|
|
fallback: 'sonnet',
|
|
});
|
|
const meter = new BudgetMeter({
|
|
budgetUsd: config.budgetUsd,
|
|
phase: 'drift',
|
|
auditPath: opts.auditPath,
|
|
});
|
|
|
|
// v0.28 scaffold: write a candidate report. v0.29 wires LLM-driven weight
|
|
// adjustment through autoUpdate. modelId + meter are wired now so the
|
|
// ledger captures the gate state even when we don't submit.
|
|
void modelId; void meter;
|
|
|
|
if (opts.dryRun) {
|
|
return {
|
|
name: 'drift',
|
|
status: 'skipped',
|
|
detail: `dry-run: ${candidates.length} candidates would be evaluated`,
|
|
totals: { candidates: candidates.length },
|
|
duration_ms: Date.now() - start,
|
|
};
|
|
}
|
|
|
|
return {
|
|
name: 'drift',
|
|
status: 'complete',
|
|
detail: `surfaced ${candidates.length} drift candidates (LLM judge: v0.29 follow-up). autoUpdate=${config.autoUpdate}`,
|
|
totals: { candidates: candidates.length },
|
|
duration_ms: Date.now() - start,
|
|
};
|
|
}
|
|
|
|
/** Test helper: expose findDriftCandidates without running the full phase. */
|
|
export const __testing = { findDriftCandidates };
|