mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-31 04:07:52 +00:00
* feat(skillpack): enhance skillify with cross-modal eval quality gate Updates skillify from v1.0.0 to v2.0.0 with the key innovation: cross-modal evaluation runs BEFORE tests (step 3) to establish quality, then tests lock in the proven-good behavior. Key changes: - 11-item checklist (was 10) - adds cross-modal eval as step 3 - Cross-modal eval uses 3 models to score output on 5 dimensions - Quality gate: all dimensions ≥ 7 average before proceeding to tests - Prevents locking in mediocrity through tests-first approach - References cross-modal-review skill for eval pipeline - Updated all gbrain-specific paths (bun test, scripts/*.ts) - Maintains compatibility with gbrain check-resolvable workflow The meta-skill for turning raw features into properly-skilled, tested, resolvable capabilities. Cross-modal eval ensures output quality before tests cement the behavior. * feat: skillify hardened via 2 cross-modal eval cycles (8.1/10) Applied top improvements from GPT-5.5 + Opus 4-7 + DeepSeek V4 Pro: - Named 3 frontier models explicitly with provider table - Inlined eval prompt template with CONTEXT param + scoring calibration - Defined aggregation math: mean >= 7 AND no single dim < 5 - Added eval receipt JSON schema - Structured 3-cycle fix loop with before/after delta tracking - Added worked example (summarize-pr, end-to-end) - Added cost guardrails (skip < 200 tokens, max 9 API calls) - Added representative input selection rule - Added SKILL.md frontmatter template (copy-paste ready) - Added Phase 0 decision gate (is this worth skillifying?) Also includes cross-modal-eval runner recipe with robust JSON parsing for LLMs that return malformed JSON (3-tier repair). * chore(recipes): remove cross-modal-eval.mjs Superseded by `gbrain eval cross-modal` (next commit). The .mjs script was the original PR's hand-rolled provider stack; the replacement reuses src/core/ai/gateway.ts so config/auth/model-aliasing comes from the canonical recipe registry instead of a parallel stack. No code references the .mjs (it was invoked by skill prose only), so this delete is independently safe to bisect through. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(eval): cross-modal-eval core module + unit tests Pure-logic foundation for the new `gbrain eval cross-modal` command (wired in the next commit). All five modules are self-contained — no CLI surface, no I/O outside the receipt writer's mkdirSync. Imported from src/core/ai/gateway.ts at runtime via gwChat (no config impact at load time). Modules: - json-repair.ts: parseModelJSON 4-strategy fallback chain. Adversarial nuclear-option throws rather than fabricating scores (Q6 + Q3 in plan). - aggregate.ts: verdict logic. PASS = (>=2 successes) AND (every dim mean >= 7) AND (every dim min across models >= 5). INCONCLUSIVE when <2/3 models returned parseable scores — closes the v1 .mjs `Object.values({}).every(...) === true` empty-array silent-PASS bug (Q2 + Q3). - receipt-name.ts: receipt filename binds (slug, sha8 of SKILL.md) so `gbrain skillify check` can detect stale audits (T10 in plan). - receipt-write.ts: thin wrapper over writeFileSync that auto-mkdirs the parent directory. Standalone module because gbrainPath() does NOT auto-mkdir (T5 plan correction — Codex caught this). - runner.ts: orchestrator. Promise.allSettled across 3 slots per cycle; up to 3 cycles; stops early on PASS or INCONCLUSIVE. Default slots: openai:gpt-4o / anthropic:claude-opus-4-7 / google:gemini-1.5-pro. estimateCost() exports a small per-model pricing table (drifts; refresh alongside model-family bumps). Tests (32 cases total, all green): - json-repair.test.ts: 10 cases (clean JSON, fences, trailing commas, single quotes, embedded newlines, mismatched braces, nuclear-option success + adversarial throws, empty input, numeric-shorthand scores). - aggregate.test.ts: 8 cases pinning Q2/Q3/dedup. The 0-of-3 INCONCLUSIVE case is the regression guard for the v1 silent-PASS bug. - cli.test.ts: 12 cases on receipt-name / receipt-write / GBRAIN_HOME isolation. Uses withEnv() helper for env mutation (R1 isolation rule). Verifies bisect-clean: typecheck passes, all 32 unit cases green. The runner.ts import of gateway.chat() is dead until commit 3 wires the CLI surface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(eval): wire `gbrain eval cross-modal` CLI subcommand User-facing surface for the multi-model quality gate. Three different- provider frontier models score the OUTPUT against the TASK on a 5-dim rubric. Verdict drives exit code: 0 PASS, 1 FAIL, 2 INCONCLUSIVE (<2/3 models returned parseable scores per Q3 in plan). Wiring touches three files: - src/commands/eval-cross-modal.ts (new, ~290 lines) CLI handler. Self-configures the AI gateway from loadConfig() + process.env so it works without `gbrain init` (the cli.ts no-DB branch bypasses connectEngine()). Defaults: cycles=3 in TTY, cycles=1 in non-TTY (T11 partial cost guardrail — limits scripted bulk spend; full --budget-usd hard cap is a v0.27.x TODO). Prints estimated max-cost-per-cycle to stderr before each run. Uses gbrainPath('eval-receipts') for receipt directory. - src/cli.ts (no-DB dispatch branch, 5-line addition) Special-cases `eval cross-modal` BEFORE the existing handleCliOnly path that requires connectEngine(). Mirrors the `dream` no-DB pattern but doesn't even attempt the connect — the command never touches the DB. New users can run the gate before `gbrain init` (T3 in plan). - src/commands/eval.ts (sub-subcommand dispatch) Adds `cross-modal` alongside `export`/`prune`/`replay`. The cli.ts branch takes precedence in the user-facing path; this branch only fires when callers re-enter runEvalCommand with an existing engine. Engine is intentionally unused — the handler self-routes. - test/e2e/cross-modal-eval.test.ts (new, 4 cases) Mocked-fetch E2E. Lives at test/e2e/* (NOT *.serial.test.ts) per plan T8: test/e2e/* is exempt from the test-isolation lint and already runs serially via scripts/run-e2e.sh, so the mock.module() call doesn't need a quarantine rename. Cases: PASS / FAIL (mean<7) / FAIL (min<5 — Q2 floor) / INCONCLUSIVE (2 mock 5xx — Q3 contract). The runner from commit 2 now has live callers. typecheck passes; the 4 E2E cases all green. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(skillify): add informational 11th item (cross-modal eval) Promotes the skillify contract from 10 to 11 items. The 11th item (cross-modal eval) is `required:false` per T7 in the plan — a missing or stale receipt surfaces in the audit output but does not fail the gate. Existing skills keep their current required-score; the bump is additive, not breaking. Changes: - src/commands/skillify.ts Header jsdoc updated 10-item -> 11-item. No code-flow changes. - src/commands/skillify-check.ts (the per-skill audit; not src/commands/skillpack-check.ts which is a different command — plan T6 corrected the conflation in the original plan) New informational item at position 11. Reuses findReceiptForSkill() helper from src/core/cross-modal-eval/receipt-name.ts to detect: * found — receipt matches current SKILL.md sha-8 * stale — receipt exists for an older SKILL.md * missing — no receipt yet Audit output cases pass through to existing pretty/JSON formats. - src/core/skillify/templates.ts Scaffolded SKILL.md now includes a "Phase 3: Cross-modal eval (informational)" section with copy-paste `gbrain eval cross-modal` invocation, pass criteria, and receipt-naming convention. Helps new skill authors discover the gate. - test/skillify-scaffold.test.ts New T9 case verifies the scaffold emits the Phase 3 section, points at the correct command, documents the receipt path, and appends exactly one resolver row. Replaces the original plan's `gbrain skillify scaffold demo-eleven` shell verification (which Codex caught as invalid + repo-mutating). Verifies: typecheck passes; scaffold test 19/19 (was 18, +1 T9 case). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: skillify v1.1.0 + cross-modal-eval references Documentation catches up with the new behavior shipped in commits 1-4. - skills/skillify/SKILL.md (1.0.0 -> 1.1.0) Full rewrite. Frontmatter version is additive (T7 in plan); the 11th item is informational, not breaking. Phase 3 now points at `gbrain eval cross-modal` with copy-paste invocation, default slot table, pass criteria, receipt-naming convention, cycles + cost guardrails (T11 partial cap), provider configuration via the AI gateway, and the cycle-1/2/3 fix loop. Adds Output Format section (skills-conformance.test.ts requires it). Drops the original `(or lib/cross-modal-eval.ts)` parenthetical (Q5 plan correction — that path never existed). - skills/cross-modal-review/SKILL.md Adds 4-line Relationship section pointing at `gbrain eval cross-modal` (D3 plan reciprocal). Distinguishes the manual second-opinion gate (this skill) from the automated multi-model score-and-iterate gate (the new command). - CLAUDE.md Key Files entries for src/commands/eval-cross-modal.ts and the five new src/core/cross-modal-eval/* modules. Commands list gains the `gbrain eval cross-modal` entry under v0.27.x. Notes the non-TTY default 1-cycle behavior + the gbrainPath('eval- receipts') resolution. - TODOS.md Four v0.27.x follow-ups filed under a new "cross-modal-eval" section: full --budget-usd cap (T11 follow-up), subagent integration (recovers cross-process rate-leases T4 deferred), skill adoption telemetry (revisit T7=C with data after 30 days), docs/cross-modal-eval.md user guide. - llms-full.txt Regenerated via `bun run build:llms` to match the CLAUDE.md edits — sync guard at test/build-llms.test.ts requires this. Verifies: typecheck passes; skills-conformance 199/199 green; build-llms 7/7 green; full unit fast loop 3861/3861 green. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: bump version and changelog (v0.28.4) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: garrytan-agents <garrytan-agents@users.noreply.github.com> Co-authored-by: Garry Tan <garrytan@gmail.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
199 lines
6.6 KiB
TypeScript
199 lines
6.6 KiB
TypeScript
/**
|
|
* E2E for `gbrain eval cross-modal` runner via mocked gateway chat().
|
|
*
|
|
* Lives under test/e2e/ so the test-isolation lint (R2 — mock.module quarantine)
|
|
* does not require a *.serial.test.ts rename: test/e2e/* is exempt from the
|
|
* lint, and `scripts/run-e2e.sh` already runs one file per Bun process so
|
|
* `mock.module` leaks are contained.
|
|
*
|
|
* Verifies the verdict / exit-code contract end-to-end:
|
|
* PASS (verdict='pass') when every dim mean >=7 and no model <5
|
|
* FAIL (verdict='fail') when any dim breaches mean OR floor
|
|
* INCONCLUSIVE (verdict='inconclusive') when <2/3 model calls succeed
|
|
*/
|
|
|
|
import { afterEach, beforeEach, describe, expect, mock, test } from 'bun:test';
|
|
import { mkdtempSync, readdirSync, readFileSync, rmSync } from 'fs';
|
|
import { tmpdir } from 'os';
|
|
import { join } from 'path';
|
|
|
|
import { configureGateway } from '../../src/core/ai/gateway.ts';
|
|
|
|
let tempDir: string;
|
|
|
|
beforeEach(() => {
|
|
tempDir = mkdtempSync(join(tmpdir(), 'gbrain-cme-e2e-'));
|
|
// Configure the gateway so our mock can pretend providers are available.
|
|
configureGateway({
|
|
embedding_model: 'openai:text-embedding-3-large',
|
|
embedding_dimensions: 1536,
|
|
expansion_model: 'anthropic:claude-haiku-4-5-20251001',
|
|
chat_model: 'anthropic:claude-sonnet-4-6-20250929',
|
|
base_urls: undefined,
|
|
env: {
|
|
OPENAI_API_KEY: 'sk-test',
|
|
ANTHROPIC_API_KEY: 'sk-ant-test',
|
|
GOOGLE_GENERATIVE_AI_API_KEY: 'sk-google-test',
|
|
},
|
|
});
|
|
});
|
|
|
|
afterEach(() => {
|
|
rmSync(tempDir, { recursive: true, force: true });
|
|
mock.restore();
|
|
});
|
|
|
|
function makeChatStub(scoresBySlot: Record<string, number[]>) {
|
|
let callIdx = 0;
|
|
const order = ['openai:gpt-4o', 'anthropic:claude-opus-4-7', 'google:gemini-1.5-pro'];
|
|
return mock(async (opts: { model?: string }) => {
|
|
const model = opts.model ?? '';
|
|
callIdx++;
|
|
const slotIdx = order.indexOf(model);
|
|
const scores = scoresBySlot[model];
|
|
if (!scores) {
|
|
throw new Error(`mock: no scores configured for model ${model}`);
|
|
}
|
|
const goal = scores[0]!;
|
|
const depth = scores[1]!;
|
|
return {
|
|
text: JSON.stringify({
|
|
scores: { goal: { score: goal }, depth: { score: depth } },
|
|
overall: (goal + depth) / 2,
|
|
improvements: [`${slotIdx + 1}. tighten the intro`],
|
|
}),
|
|
blocks: [],
|
|
stopReason: 'end',
|
|
usage: { input_tokens: 100, output_tokens: 50, cache_read_tokens: 0, cache_creation_tokens: 0 },
|
|
model,
|
|
providerId: model.split(':')[0]!,
|
|
};
|
|
});
|
|
}
|
|
|
|
describe('gbrain eval cross-modal — runner verdict contract', () => {
|
|
test('PASS: 3 happy responses, all dims >=7', async () => {
|
|
const chatStub = makeChatStub({
|
|
'openai:gpt-4o': [9, 8],
|
|
'anthropic:claude-opus-4-7': [8, 7],
|
|
'google:gemini-1.5-pro': [8, 8],
|
|
});
|
|
mock.module('../../src/core/ai/gateway.ts', () => ({
|
|
chat: chatStub,
|
|
configureGateway,
|
|
isAvailable: () => true,
|
|
}));
|
|
|
|
const { runEval } = await import('../../src/core/cross-modal-eval/runner.ts');
|
|
const result = await runEval({
|
|
task: 'sample task',
|
|
output: 'sample output content',
|
|
slug: 'demo',
|
|
receiptDir: tempDir,
|
|
cycles: 1,
|
|
});
|
|
|
|
expect(result.finalAggregate.verdict).toBe('pass');
|
|
expect(result.cycles).toHaveLength(1);
|
|
const files = readdirSync(tempDir);
|
|
expect(files.length).toBeGreaterThan(0);
|
|
expect(files[0]!.startsWith('demo-')).toBe(true);
|
|
const receipt = JSON.parse(readFileSync(join(tempDir, files[0]!), 'utf-8'));
|
|
expect(receipt.schema_version).toBe(1);
|
|
expect(receipt.aggregate.verdict).toBe('pass');
|
|
});
|
|
|
|
test('FAIL: one dim mean below 7', async () => {
|
|
const chatStub = makeChatStub({
|
|
'openai:gpt-4o': [9, 6],
|
|
'anthropic:claude-opus-4-7': [8, 6],
|
|
'google:gemini-1.5-pro': [8, 6],
|
|
});
|
|
mock.module('../../src/core/ai/gateway.ts', () => ({
|
|
chat: chatStub,
|
|
configureGateway,
|
|
isAvailable: () => true,
|
|
}));
|
|
|
|
const { runEval } = await import('../../src/core/cross-modal-eval/runner.ts');
|
|
const result = await runEval({
|
|
task: 'sample task',
|
|
output: 'sample output content',
|
|
slug: 'demo',
|
|
receiptDir: tempDir,
|
|
cycles: 1,
|
|
});
|
|
|
|
expect(result.finalAggregate.verdict).toBe('fail');
|
|
expect(result.finalAggregate.dimensions.depth!.failReason).toBe('mean_below_7');
|
|
});
|
|
|
|
test('FAIL: min-score floor caught when one model scores <5 (Q2)', async () => {
|
|
const chatStub = makeChatStub({
|
|
'openai:gpt-4o': [9, 8],
|
|
'anthropic:claude-opus-4-7': [8, 8],
|
|
'google:gemini-1.5-pro': [4, 8], // goal=4 trips the floor
|
|
});
|
|
mock.module('../../src/core/ai/gateway.ts', () => ({
|
|
chat: chatStub,
|
|
configureGateway,
|
|
isAvailable: () => true,
|
|
}));
|
|
|
|
const { runEval } = await import('../../src/core/cross-modal-eval/runner.ts');
|
|
const result = await runEval({
|
|
task: 'sample task',
|
|
output: 'sample output content',
|
|
slug: 'demo',
|
|
receiptDir: tempDir,
|
|
cycles: 1,
|
|
});
|
|
|
|
expect(result.finalAggregate.verdict).toBe('fail');
|
|
expect(result.finalAggregate.dimensions.goal!.failReason).toBe('min_below_5');
|
|
});
|
|
|
|
test('INCONCLUSIVE: 2 of 3 mock 5xx -> exit 2 contract (Q3)', async () => {
|
|
const chatStub = mock(async (opts: { model?: string }) => {
|
|
if (opts.model === 'openai:gpt-4o') {
|
|
return {
|
|
text: JSON.stringify({
|
|
scores: { goal: { score: 8 } },
|
|
improvements: ['1. ok'],
|
|
}),
|
|
blocks: [],
|
|
stopReason: 'end',
|
|
usage: { input_tokens: 0, output_tokens: 0, cache_read_tokens: 0, cache_creation_tokens: 0 },
|
|
model: opts.model,
|
|
providerId: 'openai',
|
|
};
|
|
}
|
|
throw new Error(`mock: forced 5xx for ${opts.model}`);
|
|
});
|
|
mock.module('../../src/core/ai/gateway.ts', () => ({
|
|
chat: chatStub,
|
|
configureGateway,
|
|
isAvailable: () => true,
|
|
}));
|
|
|
|
const { runEval } = await import('../../src/core/cross-modal-eval/runner.ts');
|
|
const result = await runEval({
|
|
task: 'sample task',
|
|
output: 'sample output content',
|
|
slug: 'demo',
|
|
receiptDir: tempDir,
|
|
cycles: 1,
|
|
});
|
|
|
|
expect(result.finalAggregate.verdict).toBe('inconclusive');
|
|
expect(result.finalAggregate.successes).toBe(1);
|
|
expect(result.finalAggregate.failures).toBe(2);
|
|
// Receipt is still written even on INCONCLUSIVE — forensics path.
|
|
const files = readdirSync(tempDir);
|
|
expect(files.length).toBe(1);
|
|
const receipt = JSON.parse(readFileSync(join(tempDir, files[0]!), 'utf-8'));
|
|
expect(receipt.aggregate.verdict).toBe('inconclusive');
|
|
expect(receipt.aggregate.errors).toHaveLength(2);
|
|
});
|
|
});
|