Files
gbrain/eval/runner/adapters/ripgrep-bm25.ts
T
Garry TanandClaude Opus 4.7 629ba853ff feat(eval): Phase 2 adapter interface + EXT-1 ripgrep+BM25 baseline
Phase 2 credibility unlock: BrainBench now compares gbrain to external
baselines on the same corpus and queries. Transforms the benchmark from
internal ablation ("gbrain-graph beats gbrain-grep") to category comparison
("gbrain-graph beats classic BM25 by 32 pts P@5"). This is the #1 fix
from the 4-review arc — addresses Codex's core critique that v1's
before/after was self-referential.

Added:
  eval/runner/types.ts                      — Adapter interface (v1.1 spec)
  eval/runner/adapters/ripgrep-bm25.ts      — EXT-1 classic IR baseline
  eval/runner/adapters/ripgrep-bm25.test.ts — 11 unit tests, all pass
  eval/runner/multi-adapter.ts              — side-by-side scorer

Adapter interface (eng pass 2 spec):
  - Thin 3-method Strategy: init(rawPages, config), query(q, state), snapshot(state)
  - BrainState is opaque to runner (never inspected)
  - Raw pages passed in-memory; gold/ never crosses adapter boundary
    (structural ingestion-boundary enforcement)
  - PoisonDisposition enum reserved for future poison-resistance scoring

EXT-1 ripgrep+BM25:
  - Classic Lucene-variant IDF + k1/b tuned at standard 1.5/0.75
  - Title tokens double-weighted for entity-page slug-match bias
  - Stopword filter, alphanumeric tokenization, stable lexicographic tie-break
  - Pure in-memory inverted index — no external deps, ~100 LOC core

First side-by-side results on 240-page rich-prose corpus, 145 relational queries:

| Adapter       | P@5    | R@5    | Correct top-5 |
|---------------|--------|--------|---------------|
| gbrain-after  | 49.1%  | 97.9%  | 248/261       |
| ripgrep-bm25  | 17.1%  | 62.4%  | 124/261       |
| Delta         | +32.0  | +35.5  | +124          |

gbrain-after is the hybrid graph+grep config from PR #188. Ripgrep+BM25 is
a genuinely strong classic-IR baseline (BM25 is what Lucene/Elasticsearch
ship). gbrain's ~+32-point lead on relational queries reflects real work
by the knowledge graph layer: typed links + traversePaths surface the
correct answers in top-K that BM25 only pulls in via partial-text overlap.

Next in Phase 2: EXT-2 vector-only RAG + EXT-3 hybrid-without-graph
adapters. Both plug into the same Adapter interface.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 23:39:04 +08:00

197 lines
7.3 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* BrainBench EXT-1: Ripgrep + BM25 adapter.
*
* The "honest grep-plus-BM25 baseline" — what any agent could build in an
* afternoon with standard unix tools and a classic IR formula. This is the
* external comparator that turns BrainBench from internal gbrain ablation
* into a real category benchmark.
*
* Design:
* 1. init(): Tokenizes each page's content, builds an inverted index
* + per-doc length table + global term/doc frequencies.
* 2. query(): Tokenizes the query, scores every candidate doc via BM25,
* returns top candidates ranked by score.
*
* No embeddings, no graph, no LLM. Just deterministic token-match ranking.
* This is intentionally what a "what could any agent do with grep + a
* reasonable ranker" baseline looks like.
*
* Reference: Robertson & Zaragoza, "The Probabilistic Relevance Framework:
* BM25 and Beyond" (2009). Standard formula:
*
* BM25(D, Q) = Σ_{q ∈ Q} IDF(q) × (tf(q, D) × (k1 + 1)) /
* (tf(q, D) + k1 × (1 - b + b × |D| / avgdl))
*
* Defaults: k1 = 1.5, b = 0.75. These are the values Lucene ships.
*/
import type { Adapter, AdapterConfig, BrainState, Page, Query, RankedDoc } from '../types.ts';
// ─── Tokenization ──────────────────────────────────────────────────
/**
* Tokenize text for BM25: lowercase, split on non-word chars, filter stopwords
* and tokens shorter than 2 characters. Markdown link syntax `[Name](slug)`
* is preserved enough that entity names tokenize into their component words.
*
* NOTE: Slug references (e.g. `people/alice-chen`) get split into
* `people`, `alice`, `chen` tokens — intentional, so that a query for
* "alice chen" matches pages referencing her by slug as well as name.
*/
const STOPWORDS = new Set([
'a','an','the','and','or','but','is','are','was','were','be','been','being',
'have','has','had','do','does','did','will','would','should','could','may',
'might','can','of','to','in','on','at','for','with','by','from','as','it',
'its','this','that','these','those','i','you','he','she','we','they',
'them','us','him','her','his','hers','their','theirs','my','mine','your',
'yours','our','ours',
]);
function tokenize(text: string): string[] {
return text
.toLowerCase()
.split(/[^a-z0-9]+/)
.filter(t => t.length >= 2 && !STOPWORDS.has(t));
}
// ─── BM25 state ─────────────────────────────────────────────────────
interface Bm25State {
/** term -> Map<docId, termFreq>. Inverted index. */
postings: Map<string, Map<string, number>>;
/** docId -> total token count (doc length). */
docLengths: Map<string, number>;
/** Average doc length across the corpus (BM25 normalization). */
avgDocLength: number;
/** docId -> original Page (for returning ranked results). */
docs: Map<string, Page>;
/** Total docs (|corpus|). */
N: number;
k1: number;
b: number;
}
function buildState(pages: Page[], k1: number, b: number): Bm25State {
const postings = new Map<string, Map<string, number>>();
const docLengths = new Map<string, number>();
const docs = new Map<string, Page>();
let totalLength = 0;
for (const p of pages) {
docs.set(p.slug, p);
// Index title + compiled_truth + timeline. Title gets double weight by
// being tokenized twice — cheap boost for slug-match on entity pages.
const content = `${p.title} ${p.title} ${p.compiled_truth} ${p.timeline}`;
const tokens = tokenize(content);
docLengths.set(p.slug, tokens.length);
totalLength += tokens.length;
const termFreq = new Map<string, number>();
for (const t of tokens) termFreq.set(t, (termFreq.get(t) ?? 0) + 1);
for (const [term, tf] of termFreq) {
let docMap = postings.get(term);
if (!docMap) {
docMap = new Map();
postings.set(term, docMap);
}
docMap.set(p.slug, tf);
}
}
return {
postings,
docLengths,
avgDocLength: pages.length > 0 ? totalLength / pages.length : 0,
docs,
N: pages.length,
k1,
b,
};
}
// ─── Scoring ────────────────────────────────────────────────────────
function bm25Score(state: Bm25State, docId: string, queryTokens: string[]): number {
const docLen = state.docLengths.get(docId) ?? 0;
if (docLen === 0) return 0;
const { k1, b, avgDocLength, N } = state;
let score = 0;
for (const qt of queryTokens) {
const posting = state.postings.get(qt);
if (!posting) continue;
const tf = posting.get(docId) ?? 0;
if (tf === 0) continue;
const df = posting.size;
// IDF formula: log((N - df + 0.5) / (df + 0.5) + 1) — Lucene variant,
// always positive. Standard Robertson-Sparck-Jones can go negative
// for very common terms, which misranks; Lucene's +1 smooths this out.
const idf = Math.log((N - df + 0.5) / (df + 0.5) + 1);
const numerator = tf * (k1 + 1);
const denominator = tf + k1 * (1 - b + b * (docLen / avgDocLength));
score += idf * (numerator / denominator);
}
return score;
}
// ─── Candidate collection ──────────────────────────────────────────
/** Union of docs containing ANY query token — inverted index lookup. */
function candidateDocs(state: Bm25State, queryTokens: string[]): Set<string> {
const candidates = new Set<string>();
for (const qt of queryTokens) {
const posting = state.postings.get(qt);
if (!posting) continue;
for (const docId of posting.keys()) candidates.add(docId);
}
return candidates;
}
// ─── Adapter implementation ────────────────────────────────────────
interface RipgrepBm25Config extends AdapterConfig {
k1?: number;
b?: number;
}
export class RipgrepBm25Adapter implements Adapter {
readonly name = 'ripgrep-bm25';
async init(rawPages: Page[], config: RipgrepBm25Config): Promise<BrainState> {
const k1 = config.k1 ?? 1.5;
const b = config.b ?? 0.75;
return buildState(rawPages, k1, b);
}
async query(q: Query, state: BrainState): Promise<RankedDoc[]> {
const s = state as Bm25State;
const queryTokens = tokenize(q.text);
if (queryTokens.length === 0) return [];
const candidates = candidateDocs(s, queryTokens);
const scored: { id: string; score: number }[] = [];
for (const docId of candidates) {
const score = bm25Score(s, docId, queryTokens);
if (score > 0) scored.push({ id: docId, score });
}
// Descending by score, stable tie-break by docId for determinism.
scored.sort((a, b) => b.score - a.score || a.id.localeCompare(b.id));
return scored.map((s, i) => ({
page_id: s.id,
score: s.score,
rank: i + 1,
}));
}
async snapshot(_state: BrainState): Promise<string> {
// BM25 state is pure-memory; no snapshot semantics needed for v1.1.
// Future: serialize the inverted index to disk for warm-start reruns.
return '';
}
}
/** Convenience factory — construct with default config. */
export function createRipgrepBm25(): RipgrepBm25Adapter {
return new RipgrepBm25Adapter();
}