Files
gbrain/test/routing-eval-cli.test.ts
T
4fc1246606 v0.24.0: production-hardening pass on the skillify loop (#387)
* feat: v0.19.0 — skillify loop + AGENTS.md compat + brain-first convention

This is the v0.19.0 release. The branch ships four new CLI commands, a
refactor to check-resolvable, and an expansion of the brain-first
convention for sub-agent tool discovery. The original commit message
described only the convention expansion, undercounting the scope by ~5x;
this amend captures the full release.

NEW COMMANDS

- gbrain skillify scaffold <name>     — 4 stub files + idempotent resolver row
- gbrain skillify check [path]        — 10-item post-task audit (promoted)
- gbrain skillpack list / install     — curated 25-skill bundle, atomic install
- gbrain skillpack diff <name>        — per-file diff preview
- gbrain routing-eval                 — dedicated CI verb for Check 5 fixtures

CHECK-RESOLVABLE REFACTOR

- Accepts AGENTS.md as a resolver file alongside RESOLVER.md, at either
  the skills directory or one level up (workspace root layout).
- Auto-derives the skill manifest by walking skills/*/SKILL.md when
  manifest.json is missing.
- Splits ResolvableReport into errors[] + warnings[] so advisory checks
  (filing audit, routing gaps, DRY violations) don't break CI by default.
- New --strict opt-in flag promotes warnings to exit 1.

BRAIN-FIRST CONVENTION

- skills/conventions/brain-first.md expanded from 5-step lookup guide to
  full sub-agent reference: tool inventory, lookup chain, score thresholds,
  authority hierarchy, sync rules, entity page conventions, sub-agent
  propagation rule.

PRODUCTION-READINESS HARDENING (this branch's review pass)

- routing-eval --llm: emits stderr placeholder notice + runs structural
  layer only. README, CHANGELOG, CLI help all rewritten consistently.
  Was a silent no-op against documented contract.
- skillpack installer: receipt comment in fence (cumulative-slugs="...")
  preserves single-skill-install accumulation while letting install --all
  prune removed bundle skills cleanly. Unknown rows preserved + stderr
  warning for the operating agent. Pre-v0.19 fences upgrade silently.
- skillify scaffold: resolver-row regex broadened to detect backticked,
  quoted, and bare path forms. No duplicate row on --force after the
  user normalizes formatting.
- scripts/check-privacy.sh: now wired into package.json test chain so
  the wintermute-ban rule is actually enforced. New regression test.
- E2E Tier 2 (LLM skills) promoted from schedule-only to required per-PR
  CI. Local Tier 1 + Tier 2 verified clean.
- Stale v0.17/v0.18 version labels rewritten across new files.

TESTS

- test/routing-eval-cli.test.ts: 4 cases covering --llm warn semantics
- test/privacy-script-wired.test.ts: regression guard for CI wiring
- test/skillpack-install.test.ts: 4 new cases for receipt + cumulative
  + unknown-row preserve+warn + pre-v0.19 upgrade path
- test/skillify-scaffold.test.ts: 4 new cases for broadened regex

VERIFICATION

- bun test: 2237 pass / 18 known PGLite-contention flakes (CI green;
  documented as P3 dev-experience in TODOS.md)
- bun run typecheck: clean
- bun run test:e2e: 18/19 files green (1 pre-existing flake on master,
  not caused by this branch — verified via git stash)
- llms.txt + llms-full.txt regenerated to match README + CHANGELOG

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: scrub banned fork name from public artifacts

The privacy guard wired into the test chain in this branch caught 5
pre-existing references to the banned OpenClaw fork name in CHANGELOG.md
(2x), skills/migrations/v0.19.0.md (1x), src/cli.ts (1x), and
src/commands/sync.ts (1x). All originated in master's v0.19.0 release
notes and migration doc when the privacy script existed but wasn't
wired into CI yet.

Replacements per CLAUDE.md privacy mapping:
- Origin-story copy (CHANGELOG layer narratives, code comments naming
  the production deployment that drove the feature) → "Garry's OpenClaw"
- Reader-facing migration step → "your OpenClaw"

No code semantics changed. Comments + headings only.

Verification: scripts/check-privacy.sh exits 0, full CI guard chain
green (privacy + jsonb + progress + wasm + typecheck).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump VERSION to 0.24.0 + new CHANGELOG entry

Bump branch version above master's v0.21.0 per CLAUDE.md
"CHANGELOG + VERSION are branch-scoped" rule. The new v0.24.0 entry at
the top of CHANGELOG covers what THIS branch adds vs master:

- routing-eval --llm honesty pass (4-surface contract drift fix)
- skillpack installer cumulative-receipt + unknown-row preserve+warn
  (the Codex-caught regression that would have shipped in master if
  the original v0.19.0 had landed without this branch's review pass)
- skillify scaffold resolver-row regex broadening (backtick + quoted
  + bare forms; idempotency contract preserved under hand-editing)
- 5 banned-name leaks scrubbed from public artifacts
- check-privacy.sh wired into CI test chain + regression guard test
- 7 stale v0.17/v0.18 version labels rewritten across 5 files
- Tier 2 (LLM-skills E2E) promoted from schedule-only to required per-PR

VERSION 0.21.0 → 0.24.0
package.json version field synced.
llms.txt + llms-full.txt regenerated (no content drift; sizes match).

Test suite: 62/62 green across the 5 test files this branch added or
extended (routing-eval-cli, privacy-script-wired, skillpack-install,
skillify-scaffold, build-llms).

CI guards: privacy + jsonb + progress + wasm + typecheck all clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: update project documentation for v0.24.0

Auto-discovered drift via /document-release after the v0.24.0 hardening
pass landed. All factual corrections clearly warranted by the diff.

CLAUDE.md:
- Skillpack installer: documented the cumulative-slugs receipt comment,
  install --all prune semantics, unknown-row preserve+warn behavior,
  and pre-v0.24 silent upgrade. Was previously vague about
  "tracks a skill manifest so install --update diffs cleanly" without
  explaining what the receipt is or why it matters.
- routing-eval: replaced the false claim that --llm "opts into a Haiku
  tie-break layer for CI." Now correctly describes the placeholder
  semantic landed in v0.24.0 (stderr notice + structural-only run).

README.md:
- Skillpack section: added one paragraph on the receipt comment + the
  user-visible stderr message for hand-added rows. Connects the safe
  rerun promise to the v0.24.0 implementation that actually enforces it.

CONTRIBUTING.md:
- Running tests section: now recommends `bun run test` (full CI guard
  chain + typecheck + tests) before pushing. Names each guard so new
  contributors understand what catches what. The privacy guard (newly
  wired in v0.24.0) is one of these — without `bun run test` you'd skip
  it locally and find out from CI.

llms-full.txt: regenerated to reflect CLAUDE.md changes.

Verification: full guard chain green locally (privacy + jsonb + progress
+ wasm + typecheck).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Garry Tan <garry@ycombinator.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 23:17:54 -07:00

139 lines
5.1 KiB
TypeScript

/**
* Tests for `gbrain routing-eval` CLI surface — specifically --llm
* placeholder behavior.
*
* v0.19 ships the structural layer only. The --llm flag is accepted
* as a placeholder for a future LLM tie-break layer. This test file
* locks in the contract:
*
* 1. Passing --llm emits a stderr notice ("placeholder" / "structural
* layer only"). Regardless of --json.
* 2. Passing --llm does NOT alter exit code (0 on clean, 1 on issues,
* same as without --llm).
* 3. Passing --llm --json emits valid structural JSON on stdout with
* the warning on stderr only (no stderr-to-stdout bleed).
*/
import { describe, it, expect, afterEach } from 'bun:test';
import { spawnSync } from 'child_process';
import { mkdtempSync, mkdirSync, writeFileSync, rmSync } from 'fs';
import { tmpdir } from 'os';
import { join, resolve } from 'path';
const CLI = resolve(import.meta.dir, '..', 'src', 'cli.ts');
const REPO_ROOT = resolve(import.meta.dir, '..');
function makeFixture(created: string[]): string {
const root = mkdtempSync(join(tmpdir(), 'routing-eval-cli-'));
created.push(root);
const skillsDir = join(root, 'skills');
mkdirSync(skillsDir, { recursive: true });
// Minimal resolver: one skill with a trigger phrase.
const resolver = [
'# Resolver',
'',
'| Trigger | Skill |',
'|---------|-------|',
'| "build the foo" | `skills/foo-builder/SKILL.md` |',
'',
].join('\n');
writeFileSync(join(skillsDir, 'RESOLVER.md'), resolver);
// One skill + one routing fixture that maps to it.
const skillDir = join(skillsDir, 'foo-builder');
mkdirSync(skillDir, { recursive: true });
writeFileSync(
join(skillDir, 'SKILL.md'),
'---\nname: foo-builder\n---\n\nBuilds foos.\n',
);
writeFileSync(
join(skillDir, 'routing-eval.jsonl'),
JSON.stringify({ intent: 'build the foo now please', expected_skill: 'foo-builder' }) + '\n',
);
// Manifest referencing the skill.
writeFileSync(
join(skillsDir, 'manifest.json'),
JSON.stringify({ skills: [{ name: 'foo-builder', path: 'foo-builder/SKILL.md' }] }, null, 2),
);
return skillsDir;
}
const WARNING_NEEDLE = 'placeholder';
describe('gbrain routing-eval CLI — --llm placeholder behavior', () => {
const created: string[] = [];
afterEach(() => {
for (const d of created) try { rmSync(d, { recursive: true, force: true }); } catch { /* best-effort */ }
created.length = 0;
});
it('--llm emits a stderr notice and exits 0 on clean fixtures', () => {
const skillsDir = makeFixture(created);
const proc = spawnSync('bun', [CLI, 'routing-eval', '--skills-dir', skillsDir, '--llm'], {
cwd: REPO_ROOT,
encoding: 'utf-8',
});
expect(proc.status).toBe(0);
expect(proc.stderr).toContain(WARNING_NEEDLE);
// Human-mode stdout still shows the structural results header.
expect(proc.stdout).toContain('routing-eval');
});
it('--llm --json emits warning on stderr AND valid structural JSON on stdout (no bleed)', () => {
const skillsDir = makeFixture(created);
const proc = spawnSync('bun', [CLI, 'routing-eval', '--skills-dir', skillsDir, '--llm', '--json'], {
cwd: REPO_ROOT,
encoding: 'utf-8',
});
expect(proc.status).toBe(0);
expect(proc.stderr).toContain(WARNING_NEEDLE);
// stdout must be clean JSON — no warning text bleed.
expect(proc.stdout).not.toContain(WARNING_NEEDLE);
const envelope = JSON.parse(proc.stdout); // throws if bleed corrupted it
expect(envelope.ok).toBe(true);
expect(envelope.skillsDir).toBe(skillsDir);
expect(envelope.report).not.toBeNull();
});
it('WITHOUT --llm, no placeholder warning on stderr (regression guard)', () => {
const skillsDir = makeFixture(created);
const proc = spawnSync('bun', [CLI, 'routing-eval', '--skills-dir', skillsDir], {
cwd: REPO_ROOT,
encoding: 'utf-8',
});
expect(proc.status).toBe(0);
expect(proc.stderr).not.toContain(WARNING_NEEDLE);
});
it('--llm does NOT alter exit code when fixtures have issues (still 1, not 2)', () => {
const created2: string[] = [];
const root = mkdtempSync(join(tmpdir(), 'routing-eval-cli-fail-'));
created2.push(root);
created.push(root);
const skillsDir = join(root, 'skills');
mkdirSync(skillsDir, { recursive: true });
// Resolver with no row pointing at the expected skill → miss.
writeFileSync(join(skillsDir, 'RESOLVER.md'), '# Resolver\n\n| Trigger | Skill |\n|---|---|\n');
const skillDir = join(skillsDir, 'bar-skill');
mkdirSync(skillDir, { recursive: true });
writeFileSync(join(skillDir, 'SKILL.md'), '---\nname: bar-skill\n---\n');
writeFileSync(
join(skillDir, 'routing-eval.jsonl'),
JSON.stringify({ intent: 'do bar now', expected_skill: 'bar-skill' }) + '\n',
);
writeFileSync(
join(skillsDir, 'manifest.json'),
JSON.stringify({ skills: [{ name: 'bar-skill', path: 'bar-skill/SKILL.md' }] }, null, 2),
);
const proc = spawnSync('bun', [CLI, 'routing-eval', '--skills-dir', skillsDir, '--llm'], {
cwd: REPO_ROOT,
encoding: 'utf-8',
});
expect(proc.status).toBe(1);
expect(proc.stderr).toContain(WARNING_NEEDLE);
});
});