From c39dbdb1ec9eb019e1c917de5f9191bc29892216 Mon Sep 17 00:00:00 2001 From: Garry Tan Date: Thu, 28 May 2026 22:55:09 -0700 Subject: [PATCH] docs(skillopt): hands-on tutorial for auto-improving a skill + discoverability MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit There was no tutorial for skillopt — only a reference guide (docs/guides/skillopt.md) that opens at --bootstrap-from-routing and assumes you already understand benchmarks, and an agent-facing SKILL.md. README had ZERO skillopt mention. The one thing a user must hand-author (the benchmark JSONL) was taught nowhere with a worked example. New: docs/tutorials/improving-skills-with-skillopt.md — Diataxis tutorial (learning-oriented), copy-pasteable end to end: 1. mental model in two sentences (SKILL.md is the trainable param, the agent is frozen) 2. write your first benchmark from scratch — a complete 15-task rule-judge starter you paste and run, with the full check-op table (contains/regex/section_present/max_chars/min_citations/tool_called/ tool_not_called) 3. --dry-run cost preview (and that it exits 2 by convention, not failure) 4. real run + reading accepted(0)/no_improvement(1)/aborted(2) with the actual stderr output shape 5. where output lands (best.md, versions/, history.json, rejected.json, audit jsonl) 6. accept/reject — bundled vs user skills, --no-mutate vs --allow-mutate-bundled 7. iterate by sharpening the benchmark The load-bearing fix the tutorial makes that the reference guide got wrong: the DEFAULT --split 4:1:5 needs ~50 tasks before it runs (sel = N/10, floor 5). A first-time author writing 10-15 tasks hits `D_sel has N task(s) (need >=5)` and bounces. The tutorial ships 15 tasks + `--split 1:1:1` (clean 5/5/5) so the copy-paste path actually works. Verified against the real loadBenchmark + splitBench: the exact shipped block parses 15 unique tasks and splits 5/5/5 with sel>=5; the system's own error message confirms "need ~50 total for 4:1:5". Discoverability (Diataxis cross-linking): - README.md tutorials section: new entry (was zero skillopt mention) - docs/tutorials/README.md: added under ## Shipped - docs/guides/skillopt.md: "New to this? Start with the tutorial" callout Every claim devex-verified against source: exit-code map from skillopt.ts (accepted:0/no_improvement:1/aborted:2/errored:2), stderr format from skillopt.ts:286-292, check ops from score.ts, output paths from SKILL.md, split math from benchmark.ts. Co-Authored-By: Claude Opus 4.7 (1M context) --- README.md | 1 + docs/guides/skillopt.md | 6 + docs/tutorials/README.md | 1 + .../improving-skills-with-skillopt.md | 250 ++++++++++++++++++ 4 files changed, 258 insertions(+) create mode 100644 docs/tutorials/improving-skills-with-skillopt.md diff --git a/README.md b/README.md index a443ed0ee..96d95d312 100644 --- a/README.md +++ b/README.md @@ -208,6 +208,7 @@ Step-by-step walkthroughs for getting the most out of GBrain. Each one takes you - [**Set up your personal AI agent + brain from zero**](docs/tutorials/personal-brain.md) — the canonical full-stack install. Two GitHub repos, a Telegram bot, AlphaClaw on Render, OpenClaw + GBrain + Supabase. End-to-end in about 2 hours. - [**Set up GBrain as your company brain**](docs/tutorials/company-brain.md) — federated, multi-user, OAuth-scoped institutional memory for a 10-50 person team. About 90 minutes end-to-end. +- [**Auto-improve a skill with `gbrain skillopt`**](docs/tutorials/improving-skills-with-skillopt.md) — treat a `SKILL.md` as a trainable parameter. Write a benchmark of realistic tasks, watch the optimizer propose edits and keep only the ones that measurably score higher. ~20 minutes, ~$1 in API calls. Flag + cost + safety reference: [`docs/guides/skillopt.md`](docs/guides/skillopt.md). More walkthroughs in progress: connecting an existing agent (Claude Code, Cursor, OpenClaw, Hermes) to a GBrain memory layer; setting up GBrain for VC dealflow with founder scorecards and meeting prep; migrating an existing Notion or Obsidian vault; indexing a codebase as a queryable code brain. Full tutorial index: [`docs/tutorials/`](docs/tutorials/). diff --git a/docs/guides/skillopt.md b/docs/guides/skillopt.md index d0f27fbbc..bcc0c2510 100644 --- a/docs/guides/skillopt.md +++ b/docs/guides/skillopt.md @@ -8,6 +8,12 @@ that measurably improve the score. Based on [SkillOpt](https://arxiv.org/abs/2605.23904) (Microsoft Research, May 2026). +> **New to this?** Start with the hands-on tutorial: +> [Auto-improve a skill with `gbrain skillopt`](../tutorials/improving-skills-with-skillopt.md). +> It walks you from "I have a skill" to "I accepted a measurably better version" +> in ~20 minutes, including how to write your first benchmark. This page is the +> reference — flags, exit codes, cost model, safety guards. + ## The 30-second pitch ```bash diff --git a/docs/tutorials/README.md b/docs/tutorials/README.md index 649f61c47..75c50912b 100644 --- a/docs/tutorials/README.md +++ b/docs/tutorials/README.md @@ -6,6 +6,7 @@ Step-by-step walkthroughs that take you from zero to a working outcome. Concrete - [**Set up your personal AI agent + brain from zero**](personal-brain.md) — the canonical solo install. Two GitHub repos, a Telegram bot, AlphaClaw on Render, OpenClaw + GBrain + Supabase. End-to-end in about 2 hours; about $100 to $150 a month sustained. The full-stack install I'd run today. - [**Set up GBrain as your company brain**](company-brain.md) — federated, multi-user, OAuth-scoped institutional memory for a 10-50 person team. Three sources (shared / customers / internal-only), per-user scope, first synthesized query as a teammate. About 90 minutes end-to-end, about $5 in API calls for the demo, under $100 a month sustained for a 25-person company. +- [**Auto-improve a skill with `gbrain skillopt`**](improving-skills-with-skillopt.md) — treat a `SKILL.md` as the trainable parameter of a frozen agent. Write your first benchmark from scratch (the part everyone gets stuck on), preview the cost, run the optimizer, read accepted vs no_improvement vs aborted, and accept a measurably better skill. About 20 minutes, about $1 in API calls. Reference: [`../guides/skillopt.md`](../guides/skillopt.md). ## In progress diff --git a/docs/tutorials/improving-skills-with-skillopt.md b/docs/tutorials/improving-skills-with-skillopt.md new file mode 100644 index 000000000..86df0b202 --- /dev/null +++ b/docs/tutorials/improving-skills-with-skillopt.md @@ -0,0 +1,250 @@ +# Auto-improve a skill with `gbrain skillopt` + +You have a `SKILL.md`. Sometimes the agent following it does a great job, sometimes +it forgets a step or pads the output. This tutorial takes you from that skill to a +measurably better version of it, in one session, without you hand-editing the +prose. By the end you'll have written your first benchmark, watched the optimizer +propose and test edits, and accepted an improvement that actually scored higher. + +Time: ~20 minutes. Cost: ~$1 in API calls for the worked example. + +Based on [SkillOpt](https://arxiv.org/abs/2605.23904) (Microsoft Research, May 2026). + +## The mental model (two sentences) + +Your `SKILL.md` is the trainable parameter; the agent that reads it never changes. +SkillOpt runs the agent against a benchmark of realistic tasks, proposes specific +edits to the skill body, re-tests, and keeps a change **only when it measurably +beats the current version** on a held-out slice. + +That's the whole idea. The benchmark is how "better" gets defined — which is why +writing it is the one part you can't skip. Everything else is mechanical. + +## What you'll need + +- `gbrain` installed and a brain initialized (`gbrain --version` works). +- One embedding/chat provider configured. SkillOpt makes real LLM calls. + `gbrain models doctor` should show at least one reachable chat model. +- A skill you want to improve, living at `skills//SKILL.md`. This tutorial + uses a skill called `meeting-prep` — substitute your own name everywhere. +- A clean git working tree for that skill file (SkillOpt refuses to run over + uncommitted changes so it can never clobber your edits; `--force` overrides). + +If you don't have a skill yet, scaffold one first: + +```bash +gbrain skillify scaffold meeting-prep +``` + +## Step 1: Write your first benchmark + +This is the crux. A benchmark is a `.jsonl` file — **one JSON object per line** — +where each line is a task plus a way to score the agent's answer. + +Paste this complete 15-task starter so the rest of the tutorial runs verbatim. +It's deliberately generic — once you've seen the loop work, **replace these tasks +with your skill's real cases** (that's Step 6): + +```bash +cat > skills/meeting-prep/skillopt-benchmark.jsonl <<'EOF' +{"task_id":"mp-001","task":"Prep me for a 1:1 with a direct report I haven't met with in 3 weeks.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"agenda"},{"op":"contains","arg":"follow-up"}]}} +{"task_id":"mp-002","task":"Prep me for a first sales call with a company I know nothing about.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"company"},{"op":"min_citations","arg":1}]}} +{"task_id":"mp-003","task":"Prep me for a board meeting where I present the quarterly numbers.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"metric"}]}} +{"task_id":"mp-004","task":"Prep me for a performance review I'm giving to an underperformer.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"example"}]}} +{"task_id":"mp-005","task":"Prep me for a candidate interview for a senior backend role.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"question"}]}} +{"task_id":"mp-006","task":"Prep me for a vendor renewal negotiation where I want a discount.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"leverage"}]}} +{"task_id":"mp-007","task":"Prep me for a kickoff with a new cross-functional project team.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"goal"},{"op":"contains","arg":"owner"}]}} +{"task_id":"mp-008","task":"Prep me for a difficult conversation about a missed deadline.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"impact"}]}} +{"task_id":"mp-009","task":"Prep me for an investor update call after a flat quarter.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"metric"},{"op":"min_citations","arg":1}]}} +{"task_id":"mp-010","task":"Prep me for a skip-level with someone two reports below me.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"question"}]}} +{"task_id":"mp-011","task":"Prep me for a customer escalation call after an outage.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"timeline"}]}} +{"task_id":"mp-012","task":"Prep me for a partnership exploration call with a competitor-adjacent company.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"company"},{"op":"min_citations","arg":1}]}} +{"task_id":"mp-013","task":"Prep me for a sprint retro where morale has been low.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"action"}]}} +{"task_id":"mp-014","task":"Prep me for a salary negotiation a report initiated.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"market"}]}} +{"task_id":"mp-015","task":"Prep me for an all-hands where I announce a reorg.","judge":{"kind":"rule","checks":[{"op":"max_chars","arg":1800},{"op":"contains","arg":"why"}]}} +EOF +``` + +Each line has three fields: + +- `task_id` — a unique label. Anything; you'll see it in the audit trail. +- `task` — the prompt the agent gets, exactly as a user would phrase it. +- `judge` — how the answer is scored. `kind: "rule"` is deterministic and **free** + (no LLM call): it runs a list of `checks`, and the task's score is the fraction + that pass. + +The rule checks you can use: + +| `op` | `arg` | Passes when the agent's answer… | +|---|---|---| +| `contains` | string | includes that substring | +| `regex` | string | matches that regex (multiline) | +| `section_present` | heading text | has a markdown heading with that text | +| `max_chars` | number | is at most that many characters (punishes padding) | +| `min_citations` | number | has at least N citations (markdown links, `wiki/…` refs, `[1]` footnotes) | +| `tool_called` | tool name | the agent called that tool during the rollout | +| `tool_not_called` | tool name | the agent did NOT call that tool | + +Rule judges are the right place to start. They're free, deterministic, and they +force you to say concretely what a good answer looks like. (`judge.kind` can also +be `"llm"` with a rubric, or `"qrels"` for retrieval tasks — see the +[reference guide](../guides/skillopt.md) once you outgrow rules.) + +### The one gotcha: how many tasks you need + +SkillOpt splits your benchmark three ways — **train** (propose edits against), +**sel** (the held-out gate that decides accept/reject), and **test** (final +score). The sel slice must have **at least 5 tasks** or the run refuses, so noise +can't masquerade as improvement. + +The default split is `4:1:5`, which means sel is 1/10th of your tasks — so the +default needs **~50 tasks** before it'll run. That's too many for a first +benchmark, which is why every command below passes `--split 1:1:1`: with the +15-task starter that's a clean **5 train / 5 sel / 5 test**, and sel hits the +floor exactly. + +```bash +# 15 tasks + --split 1:1:1 → 5 train / 5 sel / 5 test +gbrain skillopt meeting-prep --split 1:1:1 +``` + +If you ever see `D_sel has N task(s) after split (need >=5)`, you either added +fewer than 15 tasks or used a split whose middle number is too small a share. +`--split 1:1:1` on 15+ tasks is the simplest thing that works. + +> When you swap in your own tasks (Step 6), keep at least 15 and cover the boring +> middle, not just the edge cases. The benchmark IS your definition of quality; +> a thin benchmark optimizes for a thin definition. + +## Step 2: Preview the cost (dry run) + +Before spending anything, see what the run will cost: + +```bash +gbrain skillopt meeting-prep --split 1:1:1 --dry-run +``` + +This makes **zero LLM calls** — it just prints the plan and the cost estimate. +A ~15-task benchmark with defaults runs around $0.70–$1.00. The preflight refuses +to start a real run whose estimate exceeds `--max-cost-usd` (default $5.00), so +you can't get surprise-billed mid-run. + +> `--dry-run` exits with code **2** ("aborted"). That's the convention for "did +> not run the optimization," not a failure. The cost line is what you came for. + +## Step 3: Run it for real + +```bash +gbrain skillopt meeting-prep --split 1:1:1 +``` + +You'll watch it work: a baseline eval to set the bar, then per-step forward passes +(run the skill), backward passes (propose edits), and a validation gate that +runs each sel task's judge 3 times and takes the median — accepting only if the +median beats the current best by more than 0.05. + +When it finishes, the last lines tell you everything: + +``` +[skillopt] Outcome: accepted +[skillopt] Best sel-score: 0.840 +[skillopt] Final cost: $0.71 +[skillopt] SKILL.md rewritten with 6 optimization steps. +``` + +### Reading the outcome + +| Outcome | Exit code | What it means | What to do | +|---|---|---|---| +| `accepted` | 0 | A candidate beat the baseline. SKILL.md was rewritten (or a proposed file written — see Step 5). | Review the diff, keep it. | +| `no_improvement` | 1 | Nothing cleared the gate. Your skill is already good, or the benchmark can't tell good from bad. | Strengthen the benchmark (Step 6) or stop. | +| `aborted` | 2 | A gate stopped it: dirty working tree, over budget, `D_sel < 5`, or `--dry-run`. | Read the message — it names the gate. | + +`no_improvement` is not a failure. It's the gate doing its job: it would rather +keep your known-good skill than accept a change it can't prove is better. + +## Step 4: See what changed + +The optimizer leaves a full audit trail under the skill: + +```bash +ls skills/meeting-prep/skillopt/ +``` + +``` +best.md ← the current winning version (== SKILL.md when accepted) +versions/ + v0001_e1_s1.md ← every step's candidate, so you can diff any of them + v0002_e1_s2.md + ... +history.json ← append-only record of every accept/reject + scores +rejected.json ← edits that were tried and didn't help (so it won't retry them) +``` + +The actual change to your skill is a normal git diff: + +```bash +git diff skills/meeting-prep/SKILL.md +``` + +Run-level events (cost, model, scores per run) also land in the rotating audit +log at `~/.gbrain/audit/skillopt-YYYY-Www.jsonl`. + +## Step 5: Accept or reject — and the bundled-skill rule + +**For a skill you own** (your own `skills/` dir): an `accepted` run rewrites +`SKILL.md` in place. It's already a git diff — review it, then `git commit` to +keep it or `git checkout` to throw it away. Nothing is committed for you. + +**For a skill that ships with gbrain** (anything under the gbrain repo's own +`skills/`): SkillOpt refuses to overwrite it by default and writes the winner to +`skills//skillopt/best.md` instead, so an optimization pass can never +silently mutate a skill other people depend on. Two ways to handle that: + +```bash +# See the proposed improvement without touching SKILL.md (works for ANY skill): +gbrain skillopt meeting-prep --split 1:1:1 --no-mutate +# → writes skills/meeting-prep/skillopt/best.md, prints its path. Copy what you want. + +# Actually rewrite a bundled skill (explicit opt-in): +gbrain skillopt brain-ops --split 1:1:1 --allow-mutate-bundled +``` + +Rule of thumb: `--no-mutate` when you want to read the diff before trusting it; +`--allow-mutate-bundled` only when you intend to commit a change to a shared skill. + +## Step 6: Iterate + +The loop that actually makes skills better: + +1. Run it. If `no_improvement`, the benchmark probably can't distinguish good + from bad yet. +2. Add tasks that capture what you wish the skill did differently. Saw the agent + skip citations? Add `{"op":"min_citations","arg":2}`. Saw it ramble? Tighten + `max_chars`. +3. Re-run. A sharper benchmark gives the optimizer a real gradient to climb. +4. When a run lands `accepted`, read the diff, commit it, and bank the win. + +The skill you ship gets better every time the benchmark gets sharper. That's the +whole game: you're not editing prose, you're improving the definition of done and +letting the optimizer chase it. + +## What you built + +You wrote a benchmark that encodes what "good" means for one skill, previewed the +cost, ran the optimizer, and either accepted a measurably better skill or learned +your benchmark needs sharpening. Same loop scales to every skill you own — and +`gbrain skillopt --all` runs it across every skill that has a benchmark, under a +brain-wide cost cap. + +## Where to go next + +- **Full flag + exit-code reference, cost model, safety guards:** + [`docs/guides/skillopt.md`](../guides/skillopt.md) +- **Every flag inline:** `gbrain skillopt --help` +- **Batch + fleet + background runs** (`--all`, `--target-models`, `--background`), + **LLM and qrels judges**, **held-out test sets**, and **resume after a crash** + (`--resume `): all in the reference guide above. +- **Bootstrap a benchmark from existing routing fixtures** instead of hand-writing + it: `gbrain skillopt --bootstrap-from-routing` (then review + delete the + sentinel + re-run with `--bootstrap-reviewed`).