Two fixes surfaced during the Day 3b real-corpus run against Opus 4.5:
**eval/generators/amara-life-gen.ts** — Current Opus rejects
`temperature` and `top_p` together:
```
400 invalid_request_error: `temperature` and `top_p` cannot both be
specified for this model. Please use only one.
```
top_p=1.0 was a no-op (no nucleus truncation), so removing it has zero
semantic effect. The field is still part of MODEL_PARAMS for the cache
key so any past cache entries (none in v1) would invalidate cleanly
on the next schema version bump.
**.gitignore** — `eval/data/amara-life-v1/_cache/` is runtime Opus
cache (398 files, ~1.6MB). Regenerable from seed; no point in source
control. The corpus itself (inbox/slack/calendar/meetings/notes/docs +
corpus-manifest.json with per-item content_sha256) stays committable
for reproducibility, just the cache directory gets excluded.
Real corpus generation ran cleanly after these two fixes: 398 LLM
calls, 84,424 input / 38,062 output tokens, \$4.12 spent (vs \$20 cap,
vs \$12 estimate). All 418 items produced. Poison fixtures use
subtle paraphrased injection ("for anyone on your team who might be
triaging this thread later…") — exactly the pattern that defeats
regex redaction and requires the structured-evidence judge contract
from Day 5.
Corpus itself stays local (will move to the brainbench sibling repo
during the v0.16 split per the design doc). No eval/data/amara-life-v1/
content landing in this PR.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
BrainBench
Public benchmark for personal knowledge brain agent stacks. Ships 4 adapter
configurations scored side-by-side on a 240-page rich-prose fictional corpus
(twin-amara). Measures retrieval, extraction quality, and per-link-type
accuracy.
What this answers: "Does the knowledge graph layer do useful work, or is gbrain just a thin wrapper over vector+keyword hybrid?" Headline: gbrain beats the closest external baseline (hybrid-without-graph, same embedder, same chunking) by +31 points P@5. The graph layer is load-bearing.
5-minute quickstart
# 1. Run the full benchmark (4 adapters × 5 runs, ~15 min wall clock)
bun run eval:run
# 2. Fast iteration (N=1 single run)
bun run eval:run:dev
# 3. Just the type-accuracy report
bun run eval:type-accuracy
# 4. Explore the canonical world (contributor-facing UI)
bun run eval:world:view
What's in the box
eval/
├── data/
│ ├── world-v1/ Canonical world (committed). 240 sharded JSON files.
│ │ One file per entity + _ledger.json metadata.
│ ├── amara-life-v1/ (v0.15+) Fictional-life corpus generated on demand.
│ │ inbox/slack/calendar/meetings/notes/docs +
│ │ corpus-manifest.json. Gitignored; run
│ │ `bun run eval:generate-amara-life` once.
│ └── gold/ (v0.15+) Sealed qrels + perturbation gold.
│ entities, backlinks, qrels, contradictions, poison,
│ personalization-rubric, implicit-preferences, citations.
│ Empty templates in v0.15; filled in v1 Complete.
├── schemas/ (v0.15+) Portable JSON Schema contracts.
│ corpus-manifest, public-probe (PublicQuery with gold
│ stripped), tool-schema (12 read + 3 dry_run, 32K cap),
│ transcript, scorecard (N ∈ {1,5,10}), evidence-contract.
│ Pins the v1→v2 Inspect AI driver-swap boundary.
├── generators/
│ ├── gen.ts Opus-backed world-v1 generator (cached, $80 cap)
│ ├── world.ts World-schema scaffolder
│ ├── world-html.ts World explorer HTML renderer (XSS-safe)
│ ├── amara-life.ts (v0.15+) Deterministic amara-life skeleton.
│ │ Mulberry32 PRNG, 15 contacts, 50+300+20+8+40 items,
│ │ plants 10/5/5/3 perturbations at fixed positions.
│ └── amara-life-gen.ts (v0.15+) Opus prose expansion. Structured cache key
│ (schema_version + template_hash + item_spec_hash),
│ $20 hard-stop, --dry-run for smoke tests.
├── runner/
│ ├── multi-adapter.ts 4-adapter side-by-side scorer (N=5, seeded order)
│ ├── type-accuracy.ts Per-link-type accuracy vs gold from _facts (Cat 2)
│ ├── adversarial.ts Cat 10 robustness — 22 hand-crafted edge cases
│ ├── all.ts Master runner (current: sequential execSync;
│ v1 Complete Day 10: rewrites to async + p-limit(2))
│ ├── before-after.ts Original v1 BEFORE/AFTER retrieval run
│ ├── types.ts Adapter, Page (extended with email|slack|cal|note),
│ Query, RankedDoc. PublicPage/PublicQuery land here
│ when sealed qrels enforcement ships (v1 Complete Day 9).
│ ├── adapters/
│ │ ├── ripgrep-bm25.ts EXT-1: classic IR baseline (BM25 over grep hits)
│ │ ├── vector-only.ts EXT-2: pure cosine similarity, same embedder
│ │ └── hybrid-nograph.ts EXT-3: gbrain hybrid with graph disabled
│ └── queries/
│ ├── tier5-fuzzy.ts 30 vague-recall queries (hand-authored)
│ ├── tier5_5-synthetic.ts 50 synthetic outsider queries (AI-authored, labeled)
│ ├── validator.ts Schema + temporal as_of_date + one-slash slug rule
│ └── index.ts Aggregator + validateAll()
├── cli/
│ ├── world-view.ts Render + open world.html
│ ├── query-validate.ts Validate a Query[] file
│ └── query-new.ts Scaffold a Query template
└── reports/ Benchmark scorecards (gitignored)
Three contributor paths
Path 1: Reproduce a published scorecard
# 1. Check out the specific gbrain commit referenced in the scorecard
git checkout <commit-sha>
# 2. Run the full benchmark
bun run eval:run
# 3. Compare your numbers to the scorecard. Deterministic adapters should
# match exactly. Embedding-based adapters should land within tolerance bands.
Path 2: Submit a new external adapter
See CONTRIBUTING.md for the adapter submission flow. Short version:
- Implement
eval/runner/adapters/<your-adapter>.tsconforming to theAdapterinterface ineval/runner/types.ts. - Add a unit test file alongside.
- Wire your adapter into
eval/runner/multi-adapter.ts(one line). bun run eval:run:devto verify.- Open a PR.
Path 3: Write Tier 5.5 externally-authored queries
The T5.5 queries currently in the repo are AI-authored (author: "synthetic-outsider-v1") as a placeholder. Real outside researchers should:
bun run eval:world:viewto understand the canonical worldbun run eval:query:new --tier externally-authored --author "@your-handle"- Edit the scaffolded template with a real query + gold slugs
bun run eval:query:validate path/to/your.json- Submit via
eval/external-authors/<your-handle>/queries.jsonin a PR
See CONTRIBUTING.md for the query-submission template.
Methodology one-pager
- Corpus: 240 Opus-generated fictional biographical pages. Fixed, committed, zero private data. Reproducibility baseline for any run.
- Gold: Each page's
_factsmetadata defines canonical relationships. The scorer never shows_factsto the adapters — raw pages only cross the ingestion boundary (structural enforcement inAdapter.init). - Metrics: P@5 and R@5 on relational queries (145 canonical from
_facts, 80 tier-5 + tier-5.5). Type accuracy on extracted edges (eval/runner/type-accuracy.ts). - N=5 runs per adapter with page-order shuffle (seeded LCG; runs are reproducible). Stddev surfaces order-dependent adapter bugs. Deterministic adapters correctly show stddev=0.
- Temporal queries require explicit
as_of_date(validated at query authoring time; rejected at load if a temporal verb is present without it).
Adapter scorecard (most recent, N=5)
See docs/benchmarks/2026-04-18-brainbench-v1.md for the full report.
Quick summary from bun run eval:run:
| Adapter | P@5 | R@5 |
|---|---|---|
| gbrain-after | 49.1% | 97.9% |
| hybrid-nograph | 17.8% | 65.1% |
| ripgrep-bm25 | 17.1% | 62.4% |
| vector-only | 10.8% | 40.7% |
The graph layer beats vector+keyword hybrid on relational queries by ~31 points; hybrid-without-graph barely edges BM25. That's the story.