fix(ai): fold dashscope + google keys into gateway env, drop the retired Gemini default (#3500, #3510)

Part A (#3500): dashscope_api_key and google_api_key are now GBrainConfig
file-plane fields folded into the gateway env by buildGatewayConfig
(DASHSCOPE_API_KEY / GOOGLE_GENERATIVE_AI_API_KEY), with process-env
GEMINI_API_KEY accepted as an alias for the canonical Google name.
Honors the deferral note in brain-score-recommendations.ts by adding the
matching HOSTED_EMBED_KEY_CONFIG entries in the same change. A new sweep
guard asserts every KNOWN_CONFIG_KEYS *_api_key field reaches the
gateway env (the #121/#2662/#3500 recurring bug class).

Part B (#3510): google:gemini-1.5-pro (retired by Google) removed from
the google recipe chat allowlist first — so recipe-vs-default guard
tests catch every dead default — then replaced at the default sites:
takes-quality DEFAULT_MODEL_PANEL and grade-takes ensemble docstring use
google:gemini-2.0-flash (and openai:gpt-5.2 for the equally-dead
gpt-4o), the takes-quality pricing allowlist swaps in gemini-2.0-flash +
gpt-5.2, and cross-modal slot C moves to deepseek:deepseek-v4-pro (same
replacement as PR #3501). New test/default-model-panels.test.ts pins the
takes-quality panel to recipe chat lists, canonical pricing, the budget
allowlist, and three distinct providers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Garry Tan
2026-07-28 13:11:24 -07:00
co-authored by Claude Opus 5
parent 6920744dd8
commit 398b17aa30
19 changed files with 263 additions and 44 deletions
+1 -1
View File
@@ -85,7 +85,7 @@ per-release `**vX.Y.Z:**` narration — CI enforces this
- `src/core/eval/cycle-default.ts` — single source of truth for the eval cycle-count default. Exports `DEFAULT_CYCLES_TTY = 3`, `DEFAULT_CYCLES_NONTTY = 1`, `resolveCycleDefault(explicit, isTty): {cycles, usedNonTtyDefault}`, and `cycleDefaultSuffix(r)` (returns ` (non-interactive default; --cycles N for more)` only when the non-TTY default was applied, else `''`). Consumed by `eval-cross-modal.ts`, `eval-takes-quality.ts` (run + regress), and `takes-quality-eval/runner.ts` (core uses only the constant — library stays TTY-agnostic; the CLI owns the TTY=3 upgrade + banner annotation). `eval-suspected-contradictions.ts` applies the same transparency to its `$5`/`$1` budget default via a `budgetUsdExplicit` flag (the budget is overwritten in-place so explicitness can't be inferred post-hoc). Not shared with `resolveWorkersWithClamp` (different domain, no engine, no dedup). Pinned by `test/eval/cycle-default.test.ts`, `test/eval-suspected-contradictions-budget-default.test.ts`.
- `src/core/cross-modal-eval/json-repair.ts``parseModelJSON(raw)` named export with a 4-strategy fallback chain (direct parse → fence-strip → trailing-comma + single-quote + embedded-newline repair → regex nuclear option). Adversarial input throws rather than fabricating scores — the aggregator treats a throw as "this model contributed nothing this cycle" so the gate stays correct at >=2/3 successes.
- `src/core/cross-modal-eval/aggregate.ts` — pure verdict logic. Pass criterion: `(successes >= 2) AND (every dim mean >= 7) AND (every dim min across models >= 5)`. Inconclusive when <2/3 models returned parseable scores (regression guard for the v1 `Object.values({}).every(...) === true` empty-array PASS bug).
- `src/core/cross-modal-eval/runner.ts` — orchestrator. Each cycle runs `Promise.allSettled([gwChat(slotA), gwChat(slotB), gwChat(slotC)])` (bare allSettled, no rate-leases for the CLI path). Stops early on PASS or INCONCLUSIVE; runs up to 3 cycles. Default slots: `openai:gpt-4o` / `anthropic:claude-opus-4-7` / `google:gemini-1.5-pro`. `estimateCost()` exports a small per-model pricing table (drifts; refresh alongside model-family bumps).
- `src/core/cross-modal-eval/runner.ts` — orchestrator. Each cycle runs `Promise.allSettled([gwChat(slotA), gwChat(slotB), gwChat(slotC)])` (bare allSettled, no rate-leases for the CLI path). Stops early on PASS or INCONCLUSIVE; runs up to 3 cycles. Default slots: `openai:gpt-5.2` / `anthropic:claude-opus-4-7` / `deepseek:deepseek-v4-pro`. `estimateCost()` prices via the canonical model-pricing table; `test/cross-modal-default-slots.test.ts` pins recipe support, pricing coverage, and three distinct providers.
- `src/core/cross-modal-eval/receipt-name.ts` — receipt filename binds (slug, SKILL.md sha-8). `findReceiptForSkill(skillPath, receiptDir)` returns `'found' | 'stale' | 'missing'`. Skillify-check surfaces the status as informational; the audit does NOT fail on missing/stale receipts.
- `src/core/cross-modal-eval/receipt-write.ts` — wraps `fs.writeFileSync` with `mkdirSync({recursive:true})` ahead of every write (`gbrainPath()` does NOT auto-mkdir).
- `src/commands/eval-export.ts` — streams `eval_candidates` rows as NDJSON to stdout with `schema_version: 1` prefix on every line. EPIPE-safe, progress heartbeats on stderr, stable id-desc tiebreaker so `--since` windows never dupe/miss rows.
+2 -2
View File
@@ -31,7 +31,7 @@ receipt file from disk and re-renders it. The other modes need the brain.
| `--budget-usd N` | unset | Abort before next call's projected cost would exceed cap. Models without a `pricing.ts` entry fail loud (codex #4). |
| `--source db|fs` | `db` | `fs` is reserved for v0.33+. |
| `--slug-prefix P` | unset | Filter takes to pages whose slug starts with P. |
| `--models a,b,c` | `openai:gpt-4o,anthropic:claude-opus-4-7,google:gemini-1.5-pro` | Comma-separated panel. |
| `--models a,b,c` | `openai:gpt-5.2,anthropic:claude-opus-4-7,google:gemini-2.0-flash` | Comma-separated panel. |
| `--json` | off | Emit the full receipt to stdout. |
## Receipt JSON shape (`schema_version: 1`)
@@ -50,7 +50,7 @@ receipt file from disk and re-renders it. The other modes need the brain.
},
"prompt_sha8": "abcd1234",
"models_sha8": "abcd1234",
"models": ["openai:gpt-4o", "anthropic:claude-opus-4-7", "google:gemini-1.5-pro"],
"models": ["openai:gpt-5.2", "anthropic:claude-opus-4-7", "google:gemini-2.0-flash"],
"cycles_run": 3,
"successes_per_cycle": [3, 3, 2],
"verdict": "pass",
+2 -2
View File
@@ -139,9 +139,9 @@ edits writes a new receipt).
| Slot | Default | Provider |
|------|---------|----------|
| A | `openai:gpt-4o` | OpenAI |
| A | `openai:gpt-5.2` | OpenAI |
| B | `anthropic:claude-opus-4-7` | Anthropic |
| C | `google:gemini-1.5-pro` | Google |
| C | `deepseek:deepseek-v4-pro` | DeepSeek |
**These MUST be frontier models from DIFFERENT providers.** Using a single
provider's family or budget models defeats the purpose — different families
+1 -1
View File
@@ -78,7 +78,7 @@ FLAGS:
cycle is 3 model calls; verdict aggregates over them.
--slot-a-model <id> Override default 'openai:gpt-5.2'.
--slot-b-model <id> Override default 'anthropic:claude-opus-4-7'.
--slot-c-model <id> Override default 'google:gemini-1.5-pro'.
--slot-c-model <id> Override default 'deepseek:deepseek-v4-pro'.
--receipt-dir <path> Default: gbrainPath('eval-receipts').
--max-tokens N Output token budget per call. Default: 4000.
--json Emit final aggregate as JSON to stdout (progress to stderr).
+34 -7
View File
@@ -9,7 +9,7 @@
* import it from `../../src/cli.ts`.
*
* The single ownership site for: (a) folding file-plane API keys
* (openai/anthropic/zeroentropy/openrouter/voyage) into the gateway env, and (b) threading
* (openai/anthropic/zeroentropy/openrouter/voyage/dashscope/google) into the gateway env, and (b) threading
* local-server `*_BASE_URL` env vars into base_urls. Both matter for the
* init-time embedding-key probe — without (a) it would false-warn on
* config.json-keyed users, and without (b) a live probe could hit the wrong
@@ -44,6 +44,18 @@ export function buildGatewayConfig(c: GBrainConfig): AIGatewayConfig {
// multimodal/image embeds despite config.json looking complete. process.env
// still wins via the later spread.
if (c.voyage_api_key) envFromConfig.VOYAGE_API_KEY = c.voyage_api_key;
// #3500: same seam for DashScope. The dashscope + dashscope-rerank recipes
// require DASHSCOPE_API_KEY, but the config-plane key was never folded, so
// daemon/launchd/MCP contexts with no process-env export failed auth
// despite config.json looking complete. process.env still wins via the
// later spread.
if (c.dashscope_api_key) envFromConfig.DASHSCOPE_API_KEY = c.dashscope_api_key;
// #3500: same seam for Google Gemini. The google recipe reads
// GOOGLE_GENERATIVE_AI_API_KEY; before this fold, the ONLY way to
// configure Gemini was exporting that exact env var. (This closes the
// deferral noted in src/core/brain-score-recommendations.ts, whose
// HOSTED_EMBED_KEY_CONFIG entry lands in the same change.)
if (c.google_api_key) envFromConfig.GOOGLE_GENERATIVE_AI_API_KEY = c.google_api_key;
// Azure OpenAI (keyless/Entra): fold the non-secret endpoint/deployment + the
// Entra opt-in into the gateway env so the azure-openai recipe works in any
// shell (incl. non-interactive agent shells). The bearer token is minted at
@@ -86,11 +98,26 @@ export function buildGatewayConfig(c: GBrainConfig): AIGatewayConfig {
// every gateway op then throws NO_ANTHROPIC_API_KEY. Drop empty-string /
// undefined entries before the merge. Only '' and undefined are dropped —
// '0' and 'false' are legitimate values and survive.
env: {
...envFromConfig,
...Object.fromEntries(
Object.entries(process.env).filter(([, v]) => v !== undefined && v !== ''),
),
},
env: buildEnv(envFromConfig),
};
}
/**
* Merge config-plane fallbacks with process.env (env wins for keys carrying a
* real value — see #1249 note above), then apply the GEMINI_API_KEY alias:
* Google's own docs/SDKs export GEMINI_API_KEY, but the google recipe (and
* every gateway read site) uses GOOGLE_GENERATIVE_AI_API_KEY. Precedence:
* env GOOGLE_GENERATIVE_AI_API_KEY > env GEMINI_API_KEY > config
* google_api_key — i.e. the alias is still process-env, so it beats the
* config-plane fallback, but never the canonical env name.
*/
function buildEnv(envFromConfig: Record<string, string>): Record<string, string> {
const envReal = Object.fromEntries(
Object.entries(process.env).filter(([, v]) => v !== undefined && v !== ''),
) as Record<string, string>;
const merged = { ...envFromConfig, ...envReal };
if (!envReal.GOOGLE_GENERATIVE_AI_API_KEY && envReal.GEMINI_API_KEY) {
merged.GOOGLE_GENERATIVE_AI_API_KEY = envReal.GEMINI_API_KEY;
}
return merged;
}
+5 -2
View File
@@ -23,11 +23,14 @@ export const google: Recipe = {
price_last_verified: '2026-04-20',
},
chat: {
models: ['gemini-2.0-flash-exp', 'gemini-2.0-flash', 'gemini-1.5-pro'],
// gemini-1.5-pro was retired by Google (#3510) — deliberately NOT
// listed. Default-slot guard tests validate hardcoded defaults against
// this list, so re-adding a dead model here masks dead defaults.
models: ['gemini-2.0-flash-exp', 'gemini-2.0-flash'],
supports_tools: true,
supports_subagent_loop: true,
supports_prompt_cache: false,
max_context_tokens: 1000000, // Gemini 1.5 Pro
max_context_tokens: 1000000, // Gemini 2.0 Flash
cost_per_1m_input_usd: 0.30,
cost_per_1m_output_usd: 1.20,
price_last_verified: '2026-04-20',
+6 -8
View File
@@ -13,17 +13,13 @@ import { parseModelId } from './ai/model-resolver.ts';
*
* Only keys that `buildGatewayConfig` (src/core/ai/build-gateway-config.ts)
* actually folds from config into the gateway env may appear here.
* GOOGLE_GENERATIVE_AI_API_KEY is deliberately absent: its config field is NOT
* threaded to the gateway today, so the producer closures fall through to
* checking `process.env` ONLY for it. That matches what the gateway can
* actually use (the recipe reads that key from env). Counting a config-plane
* google_api_key here would be a false positive: doctor/autopilot would call
* the provider "configured" and dispatch an embed.stale job that then fails
* auth at the gateway. When a future change threads google_api_key into
* buildGatewayConfig, re-add the matching entry here in the same change.
*
* VOYAGE_API_KEY → voyage_api_key was the same kind of gap (#2662) until
* buildGatewayConfig started folding it — now safe to list here too.
* GOOGLE_GENERATIVE_AI_API_KEY → google_api_key and DASHSCOPE_API_KEY →
* dashscope_api_key joined for the same reason (#3500): both are folded by
* buildGatewayConfig now, so a config-plane key is genuinely usable by the
* gateway and counting it here is no longer a false positive.
*
* Caveat inherited from the existing OPENAI_API_KEY/ZEROENTROPY_API_KEY
* entries (unchanged by #2662, noted here for anyone extending this map):
@@ -40,6 +36,8 @@ export const HOSTED_EMBED_KEY_CONFIG: Record<string, string> = {
OPENAI_API_KEY: 'openai_api_key',
ZEROENTROPY_API_KEY: 'zeroentropy_api_key',
VOYAGE_API_KEY: 'voyage_api_key',
GOOGLE_GENERATIVE_AI_API_KEY: 'google_api_key',
DASHSCOPE_API_KEY: 'dashscope_api_key',
};
/**
+19
View File
@@ -63,6 +63,23 @@ export interface GBrainConfig {
* config.json file-plane route is wired through today.
*/
voyage_api_key?: string;
/**
* Alibaba DashScope API key (#3500). File-plane slot so config.json's
* `dashscope_api_key` reaches the dashscope / dashscope-rerank recipes:
* file plane → buildGatewayConfig env dict → recipe reads
* DASHSCOPE_API_KEY. Same fold pattern (and same DB-plane caveat) as
* voyage_api_key above.
*/
dashscope_api_key?: string;
/**
* Google Gemini API key (#3500). File-plane slot folded into the gateway
* env as GOOGLE_GENERATIVE_AI_API_KEY (the name the google recipe reads).
* buildGatewayConfig also accepts process-env GEMINI_API_KEY — the name
* Google's own docs/SDKs use — as an alias for
* GOOGLE_GENERATIVE_AI_API_KEY. Same fold pattern (and same DB-plane
* caveat) as voyage_api_key above.
*/
google_api_key?: string;
/** Azure OpenAI (keyless/Entra). Non-secret endpoint + deployment + Entra opt-in,
* folded into the gateway env so the azure-openai recipe works in any shell.
* The bearer token is minted at request time via `az` — no secret stored here. */
@@ -919,6 +936,8 @@ export const KNOWN_CONFIG_KEYS: readonly string[] = [
'zeroentropy_api_key',
'openrouter_api_key',
'voyage_api_key',
'dashscope_api_key',
'google_api_key',
'azure_openai_endpoint',
'azure_openai_deployment',
'azure_openai_use_entra',
+5 -1
View File
@@ -51,7 +51,11 @@ export const DEFAULT_SLOTS: SlotConfig[] = [
// 2-model quorum without a Google key (verdict: permanently inconclusive).
{ id: 'A', model: 'openai:gpt-5.2' },
{ id: 'B', model: 'anthropic:claude-opus-4-7' },
{ id: 'C', model: 'google:gemini-1.5-pro' },
// gemini-1.5-pro was retired by Google (#3510), so slot C failed even with
// a Google key configured. deepseek:deepseek-v4-pro preserves the
// three-distinct-provider contract with a model registered in both the
// recipe and canonical pricing tables (same replacement as PR #3501).
{ id: 'C', model: 'deepseek:deepseek-v4-pro' },
];
export interface SlotConfig {
+1 -1
View File
@@ -219,7 +219,7 @@ export interface GradeTakesOpts extends BasePhaseOpts {
/**
* E2 ensemble judges. When useEnsemble=true and the single-model verdict
* is borderline, all three judges are called in parallel via Promise.allSettled.
* Defaults to [openai:gpt-4o, anthropic:claude-sonnet-4-6, google:gemini-1.5-pro]
* Defaults to [openai:gpt-5.2, anthropic:claude-sonnet-4-6, google:gemini-2.0-flash]
* via defaultJudge with model-string overrides. Tests inject deterministic
* judges.
*/
+3
View File
@@ -84,6 +84,9 @@ export const CANONICAL_PRICING: Record<string, ModelPricing> = {
'openai:gpt-5.5': { input: 4.00, output: 16.00 },
// ── Google ─────────────────────────────────────────────────────────────
// `gemini-1.5-pro` was retired by Google (#3510); kept so historical
// usage/audit rows still price. Not a valid default — it's deliberately
// absent from the google recipe's chat list.
'google:gemini-1.5-pro': { input: 1.25, output: 5.00 },
// Gemini 2.0 Flash: $0.10 in / $0.40 out (verified 2026-06-03). Reconciled
// from a stale $0.30/$1.20 entry that had drifted in takes-quality-eval.
+4 -1
View File
@@ -34,6 +34,7 @@ export interface ModelPricing {
const SUPPORTED_MODELS = [
'openai:gpt-4o',
'openai:gpt-5',
'openai:gpt-5.2',
'openai:gpt-5.5',
'anthropic:claude-opus-5',
'anthropic:claude-opus-4-8',
@@ -41,7 +42,9 @@ const SUPPORTED_MODELS = [
'anthropic:claude-sonnet-5',
'anthropic:claude-sonnet-4-6',
'anthropic:claude-haiku-4-5',
'google:gemini-1.5-pro',
// gemini-1.5-pro was retired by Google (#3510); gemini-2.0-flash replaces
// it in DEFAULT_MODEL_PANEL. `gemini-2-flash` stays as the legacy alias.
'google:gemini-2.0-flash',
'google:gemini-2-flash',
] as const;
+9 -2
View File
@@ -33,10 +33,17 @@ import type { TakesQualityReceipt } from './receipt.ts';
import { estimateCost, getPricing, PricingNotFoundError } from './pricing.ts';
import { DEFAULT_CYCLES_NONTTY } from '../eval/cycle-default.ts';
/**
* Three distinct providers (uncorrelated judge blind spots). Every entry MUST
* be listed in its recipe's chat touchpoint AND in the SUPPORTED_MODELS
* pricing allowlist — pinned by test/default-model-panels.test.ts.
* google:gemini-1.5-pro (retired by Google) and openai:gpt-4o (dropped from
* the OpenAI recipe's chat list) sat here dead until #3510.
*/
export const DEFAULT_MODEL_PANEL = [
'openai:gpt-4o',
'openai:gpt-5.2',
'anthropic:claude-opus-4-7',
'google:gemini-1.5-pro',
'google:gemini-2.0-flash',
] as const;
export interface RunOpts {
+100 -1
View File
@@ -20,7 +20,7 @@
import { describe, expect, test } from 'bun:test';
import { buildGatewayConfig } from '../../src/cli.ts';
import type { GBrainConfig } from '../../src/core/config.ts';
import { KNOWN_CONFIG_KEYS, type GBrainConfig } from '../../src/core/config.ts';
import { withEnv } from '../helpers/with-env.ts';
const PASSTHROUGHS: Array<{ envVar: string; recipeId: string }> = [
@@ -139,6 +139,105 @@ describe('buildGatewayConfig config-plane API-key folding', () => {
expect(cfg.env.VOYAGE_API_KEY).toBe('pa-env-plane');
});
});
// #3500: dashscope_api_key was accepted at the file plane but never folded,
// so the dashscope/dashscope-rerank recipes (required: DASHSCOPE_API_KEY)
// could only be keyed via a process-env export.
test('dashscope_api_key folds into gateway env as DASHSCOPE_API_KEY', async () => {
await withEnv({ DASHSCOPE_API_KEY: undefined }, async () => {
const cfg = buildGatewayConfig({
dashscope_api_key: 'sk-ds-config-plane',
} as unknown as GBrainConfig);
expect(cfg.env.DASHSCOPE_API_KEY).toBe('sk-ds-config-plane');
});
});
test('a real DASHSCOPE_API_KEY process.env value wins over the config-plane fallback', async () => {
await withEnv({ DASHSCOPE_API_KEY: 'sk-ds-env-plane' }, async () => {
const cfg = buildGatewayConfig({
dashscope_api_key: 'sk-ds-config-plane',
} as unknown as GBrainConfig);
expect(cfg.env.DASHSCOPE_API_KEY).toBe('sk-ds-env-plane');
});
});
// #3500: the google recipe reads GOOGLE_GENERATIVE_AI_API_KEY; before this
// fold the ONLY configuration route was exporting that exact env name.
test('google_api_key folds into gateway env as GOOGLE_GENERATIVE_AI_API_KEY', async () => {
await withEnv(
{ GOOGLE_GENERATIVE_AI_API_KEY: undefined, GEMINI_API_KEY: undefined },
async () => {
const cfg = buildGatewayConfig({
google_api_key: 'AIza-config-plane',
} as unknown as GBrainConfig);
expect(cfg.env.GOOGLE_GENERATIVE_AI_API_KEY).toBe('AIza-config-plane');
},
);
});
// Recurring-class guard: EVERY *_api_key field declared in
// KNOWN_CONFIG_KEYS must reach the gateway env dict. Adding a new
// provider key field to GBrainConfig without folding it in
// buildGatewayConfig fails here — the #121/#2662/#3500 bug class.
test('every KNOWN_CONFIG_KEYS *_api_key field reaches the gateway env', async () => {
const keyFields = KNOWN_CONFIG_KEYS.filter((k) => k.endsWith('_api_key'));
expect(keyFields.length).toBeGreaterThanOrEqual(7);
for (const field of keyFields) {
const sentinel = `sentinel-${field}`;
// Clear the two env names the field could map to so config must win.
await withEnv(
{
[field.replace(/_api_key$/, '').toUpperCase() + '_API_KEY']: undefined,
GOOGLE_GENERATIVE_AI_API_KEY: undefined,
GEMINI_API_KEY: undefined,
},
async () => {
const cfg = buildGatewayConfig({ [field]: sentinel } as unknown as GBrainConfig);
expect(
Object.values(cfg.env).includes(sentinel),
`config field "${field}" never reaches the gateway env — add a fold in buildGatewayConfig`,
).toBe(true);
},
);
}
});
});
describe('buildGatewayConfig GEMINI_API_KEY alias (#3500)', () => {
// GEMINI_API_KEY is the env name Google's own docs and SDKs use; the
// recipe/gateway read GOOGLE_GENERATIVE_AI_API_KEY. Precedence:
// env GOOGLE_GENERATIVE_AI_API_KEY > env GEMINI_API_KEY > config google_api_key.
test('GEMINI_API_KEY aliases to GOOGLE_GENERATIVE_AI_API_KEY', async () => {
await withEnv(
{ GOOGLE_GENERATIVE_AI_API_KEY: undefined, GEMINI_API_KEY: 'AIza-gemini-env' },
async () => {
const cfg = buildGatewayConfig(baseConfig);
expect(cfg.env.GOOGLE_GENERATIVE_AI_API_KEY).toBe('AIza-gemini-env');
},
);
});
test('canonical GOOGLE_GENERATIVE_AI_API_KEY env wins over the GEMINI_API_KEY alias', async () => {
await withEnv(
{ GOOGLE_GENERATIVE_AI_API_KEY: 'AIza-canonical', GEMINI_API_KEY: 'AIza-alias' },
async () => {
const cfg = buildGatewayConfig(baseConfig);
expect(cfg.env.GOOGLE_GENERATIVE_AI_API_KEY).toBe('AIza-canonical');
},
);
});
test('GEMINI_API_KEY (process env) wins over the config-plane google_api_key', async () => {
await withEnv(
{ GOOGLE_GENERATIVE_AI_API_KEY: undefined, GEMINI_API_KEY: 'AIza-gemini-env' },
async () => {
const cfg = buildGatewayConfig({
google_api_key: 'AIza-config-plane',
} as unknown as GBrainConfig);
expect(cfg.env.GOOGLE_GENERATIVE_AI_API_KEY).toBe('AIza-gemini-env');
},
);
});
});
describe('buildGatewayConfig env empty-string clobber guard (#1249)', () => {
+5 -3
View File
@@ -63,9 +63,11 @@ describe('embeddingProviderConfigured (recipe-aware helper)', () => {
// #2662: buildGatewayConfig now folds voyage_api_key → VOYAGE_API_KEY,
// so this producer-facing map must recognize it as gateway-propagated.
expect(HOSTED_EMBED_KEY_CONFIG.VOYAGE_API_KEY).toBe('voyage_api_key');
// Not propagated to the gateway today → must NOT be backed by a config field
// (producer closures fall through to process.env only for this one).
expect(HOSTED_EMBED_KEY_CONFIG.GOOGLE_GENERATIVE_AI_API_KEY).toBeUndefined();
// #3500: buildGatewayConfig now folds google_api_key and
// dashscope_api_key, so both are gateway-propagated and must be mapped
// (a config-plane key is genuinely usable by the gateway).
expect(HOSTED_EMBED_KEY_CONFIG.GOOGLE_GENERATIVE_AI_API_KEY).toBe('google_api_key');
expect(HOSTED_EMBED_KEY_CONFIG.DASHSCOPE_API_KEY).toBe('dashscope_api_key');
});
// #2662: end-to-end regression through the REAL file-plane loader
+48
View File
@@ -0,0 +1,48 @@
/**
* Consistency guard for the takes-quality DEFAULT_MODEL_PANEL — the sibling
* of test/cross-modal-default-slots.test.ts (#3510).
*
* `google:gemini-1.5-pro` sat in the panel after Google retired it, and
* `openai:gpt-4o` after the OpenAI recipe dropped it from its chat list —
* either way the gateway rejects the slot on every default run. The guard
* only works if recipes list LIVE models: removing a dead model from its
* recipe makes every hardcoded default that still names it fail here at
* once. Do not re-add retired models to a recipe to quiet this test.
*/
import { describe, expect, test } from 'bun:test';
import { DEFAULT_MODEL_PANEL } from '../src/core/takes-quality-eval/runner.ts';
import { getPricing } from '../src/core/takes-quality-eval/pricing.ts';
import { getRecipe } from '../src/core/ai/recipes/index.ts';
import { splitProviderModelId } from '../src/core/model-id.ts';
import { canonicalLookup } from '../src/core/model-pricing.ts';
describe('takes-quality DEFAULT_MODEL_PANEL ↔ recipe consistency', () => {
test('every default panel model is listed in its recipe chat touchpoint', () => {
for (const id of DEFAULT_MODEL_PANEL) {
const { provider, model } = splitProviderModelId(id);
expect(provider).not.toBeNull();
const recipe = getRecipe(provider!);
expect(recipe, `unknown recipe "${provider}"`).toBeDefined();
expect(
recipe!.touchpoints.chat?.models ?? [],
`"${model}" not listed for ${provider} chat — the default panel can never run`,
).toContain(model);
}
});
test('every default panel model prices via canonical AND the takes-quality allowlist', () => {
for (const id of DEFAULT_MODEL_PANEL) {
expect(canonicalLookup(id), `"${id}" missing from CANONICAL_PRICING`).toBeDefined();
// getPricing throws PricingNotFoundError if the model is missing from
// SUPPORTED_MODELS — a default that can't be budget-gated aborts every
// `--budget-usd` run before the first call.
expect(getPricing(id)).toBeDefined();
}
});
test('panel spans three distinct providers (uncorrelated blind spots)', () => {
const providers = new Set(DEFAULT_MODEL_PANEL.map((id) => splitProviderModelId(id).provider));
expect(providers.size).toBe(3);
});
});
+8 -8
View File
@@ -45,7 +45,7 @@ afterEach(() => {
function makeChatStub(scoresBySlot: Record<string, number[]>) {
let callIdx = 0;
const order = ['openai:gpt-4o', 'anthropic:claude-opus-4-7', 'google:gemini-1.5-pro'];
const order = ['openai:gpt-5.2', 'anthropic:claude-opus-4-7', 'deepseek:deepseek-v4-pro'];
return mock(async (opts: { model?: string }) => {
const model = opts.model ?? '';
callIdx++;
@@ -74,9 +74,9 @@ function makeChatStub(scoresBySlot: Record<string, number[]>) {
describe('gbrain eval cross-modal — runner verdict contract', () => {
test('PASS: 3 happy responses, all dims >=7', async () => {
const chatStub = makeChatStub({
'openai:gpt-4o': [9, 8],
'openai:gpt-5.2': [9, 8],
'anthropic:claude-opus-4-7': [8, 7],
'google:gemini-1.5-pro': [8, 8],
'deepseek:deepseek-v4-pro': [8, 8],
});
mock.module('../../src/core/ai/gateway.ts', () => ({
chat: chatStub,
@@ -105,9 +105,9 @@ describe('gbrain eval cross-modal — runner verdict contract', () => {
test('FAIL: one dim mean below 7', async () => {
const chatStub = makeChatStub({
'openai:gpt-4o': [9, 6],
'openai:gpt-5.2': [9, 6],
'anthropic:claude-opus-4-7': [8, 6],
'google:gemini-1.5-pro': [8, 6],
'deepseek:deepseek-v4-pro': [8, 6],
});
mock.module('../../src/core/ai/gateway.ts', () => ({
chat: chatStub,
@@ -130,9 +130,9 @@ describe('gbrain eval cross-modal — runner verdict contract', () => {
test('FAIL: min-score floor caught when one model scores <5 (Q2)', async () => {
const chatStub = makeChatStub({
'openai:gpt-4o': [9, 8],
'openai:gpt-5.2': [9, 8],
'anthropic:claude-opus-4-7': [8, 8],
'google:gemini-1.5-pro': [4, 8], // goal=4 trips the floor
'deepseek:deepseek-v4-pro': [4, 8], // goal=4 trips the floor
});
mock.module('../../src/core/ai/gateway.ts', () => ({
chat: chatStub,
@@ -155,7 +155,7 @@ describe('gbrain eval cross-modal — runner verdict contract', () => {
test('INCONCLUSIVE: 2 of 3 mock 5xx -> exit 2 contract (Q3)', async () => {
const chatStub = mock(async (opts: { model?: string }) => {
if (opts.model === 'openai:gpt-4o') {
if (opts.model === 'openai:gpt-5.2') {
return {
text: JSON.stringify({
scores: { goal: { score: 8 } },
+6 -2
View File
@@ -11,9 +11,13 @@ import {
describe('getPricing — fail-closed contract', () => {
test('returns pricing for the default 3-model panel', () => {
expect(getPricing('openai:gpt-4o')).toBeDefined();
expect(getPricing('openai:gpt-5.2')).toBeDefined();
expect(getPricing('anthropic:claude-opus-4-7')).toBeDefined();
expect(getPricing('google:gemini-1.5-pro')).toBeDefined();
expect(getPricing('google:gemini-2.0-flash')).toBeDefined();
});
test('retired google:gemini-1.5-pro is no longer in the allowlist (#3510)', () => {
expect(() => getPricing('google:gemini-1.5-pro')).toThrow(PricingNotFoundError);
});
test('throws PricingNotFoundError on unknown model', () => {
@@ -176,8 +176,10 @@ describe('runner — budget cap (codex review #4)', () => {
const r = await runEval(engine, {
limit: 5,
cycles: 3,
models: ['openai:gpt-4o', 'anthropic:claude-opus-4-7', 'google:gemini-1.5-pro'],
budgetUsd: 0.05, // tighter than projected per-cycle cost
// gemini-2.0-flash (not the retired 1.5-pro) — the budget path
// pre-flights getPricing, which is allowlist-gated (#3510).
models: ['openai:gpt-4o', 'anthropic:claude-opus-4-7', 'google:gemini-2.0-flash'],
budgetUsd: 0.05, // tighter than projected per-cycle cost (~$0.109)
});
// No cycle ever ran successfully because pre-flight aborted cycle 1.
expect(r.budgetAborted).toBe(true);