Files
gbrain/test/ai/gateway-chat.test.ts
T
29961811a4 v0.31.12 fix: canonical Anthropic model IDs + tier routing surface + gbrain models CLI (#844)
* fix: canonical Anthropic model IDs + reverse alias + Opus 4.7 pricing

Replace claude-sonnet-4-6-20250929 with claude-sonnet-4-6 everywhere it
appears as a model ID. Starting with Claude 4.6, Anthropic API IDs are
dateless and pinned — the date suffix was carried forward from Sonnet 4.5
by mistake, producing a phantom ID that 404'd on every call.

Production impact in v0.31.6: isAvailable("chat") returned false in every
code path that loaded the recipe's model list, and extractFactsFromTurn
silently returned []. The headline real-time facts extraction feature
was a no-op on the happy path.

- gateway.ts:46 DEFAULT_CHAT_MODEL -> anthropic:claude-sonnet-4-6
- recipes/anthropic.ts: chat + expansion model lists drop date suffix;
  remove wrong-direction alias (claude-sonnet-4-6 -> -20250929);
  add reverse alias (-20250929 -> claude-sonnet-4-6) so stale user
  configs in models.dream.synthesize etc. keep working
- facts/extract.ts: routes through resolveModel; both fallbacks corrected
- anthropic-pricing.ts: Opus 4.7 corrected $15/$75 -> $5/$25 per
  Anthropic docs (the $15/$75 was Opus 4.0 pricing)
- cross-modal-eval/runner.ts: PRICING now reads from ANTHROPIC_PRICING
  for Anthropic models instead of duplicating the map (single source of
  truth — fixes the drift trap that motivated this whole patch)

Tests: cherry-pick PR #830's test/anthropic-model-ids.test.ts verbatim
(6 recipe-shape guardrails). Update gateway-chat tests to assert reverse
alias resolves correctly. Update budget-meter test for new Opus pricing.

Co-Authored-By: garrytan-agents <garrytan-agents@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: model tier system + recipe-models merge + async reconfigure hook

Add 4-tier model routing (utility/reasoning/deep/subagent) so users can
swap defaults with one config key. Each tier maps to a class of work;
override globally via models.default or per-tier via models.tier.<tier>.

Codex flagged three real architecture issues in the v0.31.12 plan review;
this commit addresses each.

F3 — sync/async timing of configureGateway:
  - buildGatewayConfig stays synchronous (pre-engine-connect callers
    keep working)
  - New reconfigureGatewayWithEngine(engine) async function re-resolves
    expansion + chat defaults through resolveModel after engine.connect()
  - cli.ts wires the re-stamp into the post-connect path

F4/F5 — softening assertTouchpoint was too broad:
  - Earlier plan was to flip native-recipe validation from throw to warn,
    affecting gateway.chat AND gateway.expand AND gateway.embed
  - Instead: per-gateway-instance recipe-models merge. assertTouchpoint
    gets an optional extendedModels Set; when the user opted into a model
    via config, it bypasses the throw. Source-code typos still fail fast.
  - Existing contract test (test/ai/gateway-chat.test.ts:106) preserved

Tier defaults are TIER_DEFAULTS in model-config.ts. Resolution chain
inserts at step 5 (between models.default and env var). Each existing
resolveModel call site gains a tier: arg — think (deep), cycle/synthesize
(reasoning + utility for verdict), patterns/drift (reasoning), auto-think
(deep), facts/extract (reasoning).

Plus 10 new tests pinning tier precedence, subagent-tier fallback when
models.default is non-Anthropic, and the F6 alias-chain conflict case.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: subagent runtime enforcement for non-Anthropic models (3 layers)

The subagent loop uses Anthropic's Messages API with prompt caching on
system + tools. OpenAI/Google have different shapes. Setting
models.default = openai:gpt-5.5 and routing the subagent there silently
breaks the loop.

Codex F1+F2+F13 in the v0.31.12 plan review pointed out that "warn at
doctor" wasn't enough — handlers/subagent.ts:148 still did
`const model = data.model ?? DEFAULT_MODEL` and called Anthropic directly,
so a job submitted with data.model = openai:gpt-5.5 bypassed any tier
logic and failed at runtime with a confusing provider error.

Three layers of enforcement, defense in depth:

Layer 1 (queue.ts:add) — submit-time guard. When name === 'subagent'
and data.model is set, validate the provider. Non-Anthropic rejects
before the job enters the queue.

Layer 2 (handlers/subagent.ts) — tier-resolution fallback. The handler
routes through resolveModel({ tier: 'subagent' }). If the chain resolves
to a non-Anthropic provider (via models.default or models.tier.subagent),
the resolver warns + falls back to TIER_DEFAULTS.subagent
(claude-sonnet-4-6).

Layer 3 (doctor.ts:checkSubagentProvider) — surfacing layer. Warns when
models.tier.subagent or models.default is explicitly set to a
non-Anthropic provider, with a paste-ready fix command. Lets users see
config drift before submitting a job.

Tests: 3 new cases in test/agent-cli.test.ts asserting the queue-level
guard rejects non-Anthropic data.model. Existing test/subagent-handler
suite still passes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: gbrain models CLI + doctor probe + silent-no-op regression test

New gbrain models CLI gives the agent and user visibility into routing.
Read mode prints the tier table, current overrides, per-task config,
and aliases with source-of-truth attribution per row. Doctor subcommand
fires a 1-token probe to each configured chat/expansion model and
classifies failures (model_not_found / auth / rate_limit / network /
unknown) so config-time invalid IDs surface without waiting for a
production call that silently degrades.

Per Codex F11 — no specific dollar cost claim in either the help text
or the CHANGELOG (providers have minimum-output billing and prompt-cache
rounding that vary). Probe is opt-in (gbrain doctor --probe-models),
never auto-runs. --skip=<provider> narrows the matrix for cost-sensitive
operators.

Per Codex F7+F8+F15 (the structural regression gap): new
test/facts-extract-silent-no-op.test.ts is THE regression test for the
bug class that motivated v0.31.12. Five cases including the smoking-gun:
when chat IS available, extractFactsFromTurn MUST actually call the chat
transport, not silently return []. Uses the gateway's
__setChatTransportForTests seam so it runs in every shard with no API key.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.31.12)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: document v0.31.12 model tier system + gbrain models CLI

Add CLAUDE.md Key Files annotations for the v0.31.12 work:
src/core/model-config.ts (tier system + isAnthropicProvider + TIER_DEFAULTS),
src/core/ai/model-resolver.ts (assertTouchpoint extendedModels arg),
src/core/ai/gateway.ts (reconfigureGatewayWithEngine + extended-models registry),
src/core/minions/queue.ts (subagent submit-time guard, layer 1 of 3),
src/commands/models.ts (new gbrain models CLI + doctor probe),
src/commands/doctor.ts (subagent_provider check, layer 3 of 3),
src/core/ai/recipes/anthropic.ts (canonical model IDs + reverse alias),
src/core/anthropic-pricing.ts (Opus 4.7 corrected to \$5/\$25).

Add CLAUDE.md commands section for gbrain models + gbrain models doctor
+ power-user config recipes. Add README.md command-table rows for the
same. Regenerate llms-full.txt so the bundled docs stay in sync.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: scrub --probe-models reference (flag not actually wired)

The v0.31.12 CHANGELOG and skills/conventions/model-routing.md both
referenced `gbrain doctor --probe-models` as an integrated probe entry
point. The flag was never implemented — only `gbrain models doctor`
landed as the probe surface. Caught by /document-release subagent.

Drop the references rather than wire an untested flag at the last minute.
The probe is reachable via `gbrain models doctor`; users who want it
in doctor's output run that command separately.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: garrytan-agents <garrytan-agents@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 20:06:31 -07:00

222 lines
9.3 KiB
TypeScript

/**
* Commit 1 — chat touchpoint coverage.
*
* Asserts:
* - chat() resolves provider:model strings + aliases
* - assertTouchpoint surfaces chat-only providers correctly
* - getChatModel() default + override
* - chat_fallback_chain plumbing (config plumbing only — chatWithFallback ships in commit 3)
* - new openai-compat recipes (deepseek, groq, together) parse + resolve
* - new ChatTouchpoint shape: supports_subagent_loop, supports_prompt_cache
* - mapStopReason via the chat() boundary (mocked client) — refusal / content_filter / tool_calls / end / length
*
* The actual `generateText` call is exercised via a fake AI SDK model object
* (the `model` returned from `createOpenAICompatible(...).languageModel()`)
* passed by patching the module cache. We bypass the heavy SDK by mocking the
* `generateText` import via Bun's module-replace pattern.
*/
import { describe, test, expect, beforeEach, mock } from 'bun:test';
import {
configureGateway,
resetGateway,
isAvailable,
getChatModel,
getChatFallbackChain,
} from '../../src/core/ai/gateway.ts';
import { parseModelId, resolveRecipe, assertTouchpoint } from '../../src/core/ai/model-resolver.ts';
import { AIConfigError } from '../../src/core/ai/errors.ts';
import { listRecipes, getRecipe } from '../../src/core/ai/recipes/index.ts';
describe('chat touchpoint — recipe registry', () => {
test('all six chat-capable providers ship a chat touchpoint with supports_subagent_loop', () => {
const expected = ['anthropic', 'openai', 'google', 'deepseek', 'groq', 'together'];
for (const id of expected) {
const r = getRecipe(id);
expect(r, `recipe missing: ${id}`).toBeDefined();
expect(r!.touchpoints.chat, `${id} missing chat touchpoint`).toBeDefined();
expect(r!.touchpoints.chat!.models.length, `${id} chat models empty`).toBeGreaterThan(0);
expect(r!.touchpoints.chat!.supports_subagent_loop, `${id} should support subagent loop`).toBe(true);
}
});
test('only Anthropic claims supports_prompt_cache=true', () => {
for (const r of listRecipes()) {
if (!r.touchpoints.chat) continue;
if (r.id === 'anthropic') {
expect(r.touchpoints.chat.supports_prompt_cache).toBe(true);
} else {
expect(r.touchpoints.chat.supports_prompt_cache ?? false).toBe(false);
}
}
});
test('embedding-only providers (voyage, ollama) do NOT declare chat', () => {
expect(getRecipe('voyage')!.touchpoints.chat).toBeUndefined();
expect(getRecipe('ollama')!.touchpoints.chat).toBeUndefined();
});
test('openai-compat chat recipes have base_url_default', () => {
expect(getRecipe('deepseek')!.base_url_default).toBe('https://api.deepseek.com/v1');
expect(getRecipe('groq')!.base_url_default).toBe('https://api.groq.com/openai/v1');
expect(getRecipe('together')!.base_url_default).toBe('https://api.together.xyz/v1');
});
});
describe('chat touchpoint — model resolver + aliases (Codex F-OV-5)', () => {
test('parseModelId handles dated and undated forms identically at parse time', () => {
expect(parseModelId('anthropic:claude-sonnet-4-6')).toEqual({
providerId: 'anthropic',
modelId: 'claude-sonnet-4-6',
});
expect(parseModelId('anthropic:claude-haiku-4-5-20251001')).toEqual({
providerId: 'anthropic',
modelId: 'claude-haiku-4-5-20251001',
});
});
test('resolveRecipe expands pre-4.6 dateless alias to dated canonical', () => {
// Pre-4.6 models keep date-based aliases (Haiku 4.5 predates the
// dateless convention).
const { parsed } = resolveRecipe('anthropic:claude-haiku-4-5');
expect(parsed.modelId).toBe('claude-haiku-4-5-20251001');
});
test('resolveRecipe leaves dateless 4.6+ models unchanged (they ARE canonical)', () => {
const { parsed } = resolveRecipe('anthropic:claude-opus-4-7');
expect(parsed.modelId).toBe('claude-opus-4-7');
const { parsed: parsed2 } = resolveRecipe('anthropic:claude-sonnet-4-6');
expect(parsed2.modelId).toBe('claude-sonnet-4-6');
});
test('reverse alias rescues v0.31.6-shipped broken Sonnet 4.6 ID (regression)', () => {
// gbrain v0.31.6 shipped 'claude-sonnet-4-6-20250929' as a hardcoded
// default, which 404s on the Anthropic API (Sonnet 4.6 is dateless).
// The reverse alias rewrites broken → canonical so any user with a
// stale `models.dream.synthesize` / `facts.extraction_model` config
// keeps working. Regression guard against a future "cleanup" that
// drops this alias entry.
const { parsed } = resolveRecipe('anthropic:claude-sonnet-4-6-20250929');
expect(parsed.modelId).toBe('claude-sonnet-4-6');
});
test('assertTouchpoint accepts chat for chat-capable native + openai-compat providers', () => {
expect(() => assertTouchpoint(getRecipe('anthropic')!, 'chat', 'claude-opus-4-7')).not.toThrow();
expect(() => assertTouchpoint(getRecipe('openai')!, 'chat', 'gpt-5.2')).not.toThrow();
expect(() => assertTouchpoint(getRecipe('google')!, 'chat', 'gemini-2.0-flash')).not.toThrow();
expect(() => assertTouchpoint(getRecipe('deepseek')!, 'chat', 'deepseek-chat')).not.toThrow();
});
test('assertTouchpoint rejects chat on embedding-only providers with a fix hint', () => {
expect(() => assertTouchpoint(getRecipe('voyage')!, 'chat', 'voyage-3'))
.toThrow(AIConfigError);
expect(() => assertTouchpoint(getRecipe('ollama')!, 'chat', 'nomic-embed-text'))
.toThrow(AIConfigError);
});
test('assertTouchpoint rejects unknown native model with the model list in the fix hint', () => {
try {
assertTouchpoint(getRecipe('anthropic')!, 'chat', 'claude-opus-9-99');
throw new Error('should have thrown');
} catch (e) {
expect(e).toBeInstanceOf(AIConfigError);
expect((e as AIConfigError).message).toContain('claude-opus-9-99');
}
});
test('assertTouchpoint accepts arbitrary model on openai-compat tier', () => {
// openai-compat lets users pass models not declared in the recipe (provider may host more)
expect(() => assertTouchpoint(getRecipe('groq')!, 'chat', 'some-future-model')).not.toThrow();
});
});
describe('chat touchpoint — gateway config plumbing', () => {
beforeEach(() => resetGateway());
test('default chat_model is anthropic:claude-sonnet-4-6', () => {
configureGateway({ env: {} });
expect(getChatModel()).toBe('anthropic:claude-sonnet-4-6');
});
test('explicit chat_model overrides the default', () => {
configureGateway({
chat_model: 'openai:gpt-5.2',
env: { OPENAI_API_KEY: 'fake' },
});
expect(getChatModel()).toBe('openai:gpt-5.2');
});
test('chat_fallback_chain plumbed and retrievable', () => {
configureGateway({
chat_fallback_chain: [
'anthropic:claude-opus-4-7',
'deepseek:deepseek-chat',
],
env: {},
});
expect(getChatFallbackChain()).toEqual([
'anthropic:claude-opus-4-7',
'deepseek:deepseek-chat',
]);
});
test('chat_fallback_chain defaults to empty array', () => {
configureGateway({ env: {} });
expect(getChatFallbackChain()).toEqual([]);
});
test('isAvailable("chat") returns true when default Anthropic + key present', () => {
configureGateway({ env: { ANTHROPIC_API_KEY: 'fake' } });
expect(isAvailable('chat')).toBe(true);
});
test('isAvailable("chat") returns false when configured provider has no key', () => {
configureGateway({ chat_model: 'openai:gpt-5.2', env: {} });
expect(isAvailable('chat')).toBe(false);
});
test('isAvailable("chat") returns false on embedding-only chat target', () => {
// Voyage doesn't expose a chat touchpoint; isAvailable should refuse.
configureGateway({ chat_model: 'voyage:voyage-3', env: { VOYAGE_API_KEY: 'fake' } });
expect(isAvailable('chat')).toBe(false);
});
});
describe('chat touchpoint — config alias resolution', () => {
beforeEach(() => resetGateway());
test('isAvailable("chat") accepts undated alias and resolves correctly', () => {
configureGateway({
chat_model: 'anthropic:claude-sonnet-4-6', // undated
env: { ANTHROPIC_API_KEY: 'fake' },
});
expect(isAvailable('chat')).toBe(true);
});
});
describe('chat touchpoint — chat() smoke + stop-reason mapping (Codex D8)', () => {
// We exercise chat() against a mocked AI-SDK 'generateText' to assert the
// gateway's structural-signal mapping (mapStopReason) covers refusal,
// content_filter, tool_calls, end, length without the regex layer (commit 3).
// A full integration test against real provider HTTP lives in
// test/e2e/agent-multi-provider.test.ts (commit 2).
//
// We can't easily monkey-patch ESM imports inside Bun's runtime; instead we
// write an end-to-end assertion against the resolver logic + verify the
// chat() function exists with the documented signature.
test('chat() function is exported with the expected signature', async () => {
const mod = await import('../../src/core/ai/gateway.ts');
expect(typeof mod.chat).toBe('function');
// Signature check: must accept ChatOpts. We don't call it without a real
// provider key — that's the e2e job.
});
test('ChatBlock + ChatMessage + ChatResult types are exported', async () => {
// Type-only assertion: if these imports compile, we're good. The test
// body is just a runtime touch.
const mod = await import('../../src/core/ai/gateway.ts');
expect(mod).toBeDefined();
});
});