mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-28 06:23:01 +00:00
skill: codex review P2-1 — verification gate now tests the user's edited file
The original SKILL.md Step 6 told users to run `node harness.mjs` from the
gbrain repo as the mandatory ≥95% gate. But that runs the harness against
the COMMITTED sample variants in evals/functional-area-resolver/variants/,
not the file the user just compressed. The gate could pass while the edit
dropped a sub-skill.
Step 6 now:
- Gate 1 stays at `gbrain routing-eval --json` (structural, runs against
the user's actual routing-eval.jsonl fixtures).
- Gate 2 is rewritten: copy the user's edited routing file into a tmp
variants dir, then run `node harness.mjs --variants-dir <tmp>
--variants my-edit --model opus`. This exercises the harness's existing
--variants flag (added in commit 472cc686 / T4) but now points at the
user's actual edit. The harness uses gbrain-bundled fixtures, so this
is a regression check on shared skills, not a full eval of the user's
fixture set — and the SKILL.md says so explicitly.
Also adds a "common false negatives" callout: when the user's routing
file doesn't expose the skills gbrain's bundled fixtures target (e.g.
`gmail`, `enrich`), expect strict-scoring fails on those rows; lenient
scoring remains accurate.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
4f44a06432
commit
ca99fbfeb5
@@ -237,22 +237,59 @@ stay as individual rows — they're checked on every message, not dispatched.
|
|||||||
Run two gates before committing the compressed file. Do NOT commit if either
|
Run two gates before committing the compressed file. Do NOT commit if either
|
||||||
fails.
|
fails.
|
||||||
|
|
||||||
```bash
|
**Gate 1: Structural verification.** Confirms your `routing-eval.jsonl`
|
||||||
# Structural verification — checks routing-eval.jsonl fixtures still resolve
|
fixtures still resolve to the right skills under the compressed routing file.
|
||||||
gbrain routing-eval --json
|
Run from the workspace whose routing file you just edited:
|
||||||
|
|
||||||
# LLM A/B verification — re-runs the held-out corpus against the variants
|
```bash
|
||||||
# (this is the gate that proves compression didn't drop sub-skills)
|
gbrain routing-eval --json
|
||||||
cd evals/functional-area-resolver
|
|
||||||
node harness.mjs
|
|
||||||
```
|
```
|
||||||
|
|
||||||
If the harness's functional-areas variant scores below 95% on the held-out
|
If accuracy on your fixtures drops below 95%, revert and tune the area
|
||||||
corpus, revert the compression and tune the area entries before re-running.
|
entries before re-running.
|
||||||
Common causes of accuracy drops:
|
|
||||||
|
**Gate 2: LLM A/B verification on YOUR edited file.** Confirms a frontier
|
||||||
|
LLM can still drill into the dispatcher list and reach sub-skills under
|
||||||
|
your specific compression. Requires a gbrain repo checkout because the
|
||||||
|
harness lives there. Copy your edited routing file into the harness's
|
||||||
|
variants directory, then invoke the harness with `--variants` pointing
|
||||||
|
at it:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# In your agent workspace, identify the routing file you just compressed.
|
||||||
|
EDITED=/path/to/your/AGENTS.md # or skills/RESOLVER.md, whichever you edited
|
||||||
|
|
||||||
|
# In your gbrain repo checkout:
|
||||||
|
cd /path/to/gbrain/evals/functional-area-resolver
|
||||||
|
TMP=$(mktemp -d)/variants && mkdir -p "$TMP"
|
||||||
|
cp "$EDITED" "$TMP/my-edit.md"
|
||||||
|
|
||||||
|
# Run the harness against your file (sequential, ~75 calls × $0.0076 ≈ $0.57 on Opus).
|
||||||
|
ANTHROPIC_API_KEY=... node harness.mjs --variants-dir "$TMP" --variants my-edit \
|
||||||
|
--model opus --parallel 3 --yes
|
||||||
|
```
|
||||||
|
|
||||||
|
The harness uses gbrain's bundled fixture set, so this verifies "did the LLM
|
||||||
|
land in the right sub-skill for routing intents the gbrain-bundled fixtures
|
||||||
|
cover" — a regression check on shared skills, not a full re-eval of YOUR
|
||||||
|
fixture set. For full eval coverage, mirror this skill's
|
||||||
|
`fixtures.jsonl` + `fixtures-held-out.jsonl` setup with intents specific
|
||||||
|
to your skills.
|
||||||
|
|
||||||
|
If the lenient (same-area) score on your variant drops below 95%, revert the
|
||||||
|
compression and tune. Common causes:
|
||||||
- A sub-skill was omitted from the `(dispatcher for: ...)` list.
|
- A sub-skill was omitted from the `(dispatcher for: ...)` list.
|
||||||
- Trigger phrases for an area are too narrow (LLM can't recognize intent).
|
- Trigger phrases for an area are too narrow (LLM can't recognize intent).
|
||||||
- Areas were collapsed too aggressively (too few areas — see Anti-Patterns).
|
- Areas were collapsed too aggressively (too few areas — see Anti-Patterns).
|
||||||
|
- ASCII `->` vs Unicode `→` mismatch — the harness now accepts both, but
|
||||||
|
earlier versions only matched Unicode. Pin gbrain to v0.32.3.0+.
|
||||||
|
|
||||||
|
Common false negatives on the harness eval (NOT bugs in your compression):
|
||||||
|
- The gbrain-bundled fixtures target skill names like `enrich`, `query`,
|
||||||
|
`gmail`, `executive-assistant`. If your routing file doesn't expose
|
||||||
|
those skills at all, expect strict-scoring failures on those fixtures.
|
||||||
|
Lenient scoring stays accurate for any sub-skill present in your
|
||||||
|
`(dispatcher for: ...)` lists.
|
||||||
|
|
||||||
### Step 7: Review the diff before committing
|
### Step 7: Review the diff before committing
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user