mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-30 03:12:32 +00:00
* feat: search quality boost — compiled truth ranking, detail parameter, cosine re-scoring Compiled truth chunks now rank 2x higher in hybrid search via RRF normalization + source boost. New --detail flag (low/medium/high) controls timeline inclusion. Cosine re-scoring blends query-chunk similarity before dedup for query-specific ranking. Also: remove DISTINCT ON from keyword search (dedup handles per-page capping), add chunk_id + chunk_index to SearchResult, add getEmbeddingsByChunkIds to BrainEngine interface. Inspired by Ramp Labs' "Latent Briefing" paper (April 2026). * feat: RRF normalization, source-aware dedup, detail param in operations RRF scores normalized to 0-1 before 2.0x compiled truth boost. Source-aware dedup guarantees compiled truth chunk per page. Detail parameter added to query operation, dedupResults added to bare search operation. Debug logging via GBRAIN_SEARCH_DEBUG=1. * chore: bump version and changelog (v0.8.1) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: CJK word count in query expansion CJK text is not space-delimited. A query like "向量搜索优化" was counted as 1 word and silently skipped expansion. Now counts characters for CJK queries instead of space-separated tokens. Co-Authored-By: YIING99 <yiing99@users.noreply.github.com> * feat: retrieval evaluation harness — P@k, R@k, MRR, nDCG@k + gbrain eval Full IR evaluation framework: precisionAtK, recallAtK, mrr, ndcgAtK metrics with runEval() orchestrator. gbrain eval CLI with single-run table and A/B comparison mode (--config-a / --config-b) for parameter tuning. HybridSearchOpts now accepts rrfK and dedupOpts overrides. Co-Authored-By: 4shut0sh <4shut0sh@users.noreply.github.com> * test: search quality tests — RRF boost, dedup guarantee, cosine similarity, E2E benchmark 42 new tests across 3 files: - test/search.test.ts: RRF normalization, compiled truth 2x boost, dedup key collision prevention, cosine similarity edge cases, CJK word count detection - test/dedup.test.ts: source-aware compiled truth guarantee, layer interactions, custom maxPerPage, empty/single result edge cases - test/e2e/search-quality.test.ts: full pipeline against PGLite with basis vector embeddings — chunk_id/chunk_index fields, detail parameter filtering, getEmbeddingsByChunkIds, keyword multi-chunk, vector ordering Also: export rrfFusion + cosineSimilarity for unit testing, fix PGLite getEmbeddingsByChunkIds to parse string vectors from pgvector. * test: search quality benchmark with A/B comparison (baseline vs PR#64) Benchmark measures P@1, MRR, nDCG@5, and source accuracy across 8 queries against 5 seeded pages. Key finding: boost helps entity lookups but over-corrects temporal queries. Validates the --detail parameter as the right control mechanism. Output at docs/benchmarks/2026-04-13.md. * feat: query intent classifier — auto-selects detail level, 100% source accuracy Zero-latency heuristic classifier detects query intent from text patterns: - "Who is Pedro?" → entity → detail=low (compiled truth only) - "When did we last meet?" → temporal → detail=high (no boost, natural ranking) - "Variant fund announcement" → event → detail=high - General queries → detail=medium (default with boost) The key insight: skip the 2.0x compiled truth boost for detail=high queries. Temporal/event queries want natural ranking where timeline entries can win. Benchmark results (source accuracy = does the top chunk match expected type): - Baseline: 100% (already good, no boost needed) - Boost only: 71.4% (boost over-corrects temporal queries) - Boost + intent classifier: 100% (best of both worlds) 35 unit tests for the classifier. 590 total tests pass. * feat: query intent classifier — auto-selects detail level, 100% source accuracy Heuristic classifier detects query intent from text patterns (zero latency, no LLM call). Maps temporal queries ("when did we last meet") to detail=high, entity queries ("who is X") to detail=low, events to detail=high. Benchmark results (29 pages, 20 queries, graded relevance): - Baseline: P@1=0.947, MRR=0.974, source accuracy=89.5% - Boost only: P@1=0.895, MRR=0.939, source accuracy=63.2% (over-correction) - Boost + intent: P@1=0.947, MRR=0.974, source accuracy=89.5% (fully recovered) The intent classifier eliminates the boost's over-correction on temporal queries while preserving its benefits for entity lookups. 35 unit tests for the classifier. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test: search quality benchmark with A/B comparison (baseline vs PR#64) Rich benchmark: 29 pages, 58 chunks, 20 queries with graded relevance. Now measures CHUNK-LEVEL quality, not just page-level retrieval. Key findings (C. Boost+Intent vs A. Baseline): - Unique pages in top-10: 7.2 → 8.7 (+21% broader coverage) - Compiled truth ratio: 51.6% → 66.8% (+15pp more signal) - CT-first rate: 100% (compiled truth leads for entity queries) - Timeline accessible: 100% (temporal queries still find dates) - Source accuracy: 89.5% maintained (intent classifier prevents regression) The boost alone (B) causes -26pp source accuracy regression. Intent classifier (C) recovers it fully. * docs: clean benchmark report — ELI10 search quality analysis for PR#64 Replaces two drafts with one clean report. Explains what changed, why it matters, and what the numbers mean. All fictional data, no private info. Key findings: 21% more page coverage per query, 29% more compiled truth in results. Intent classifier prevents boost from burying timeline for temporal queries. Full per-query breakdown with before/after comparison. * chore: remove auto-generated benchmark file (clean version is 2026-04-14-search-quality.md) * docs: update project documentation for search quality boost CLAUDE.md: added search/intent.ts, search/eval.ts, commands/eval.ts to key files. Added 5 new test files (search, dedup, intent, eval, e2e/search-quality). Updated test count from 23+4 to 28+5. Added docs/benchmarks/ to key files. README.md: updated search pipeline diagram with intent classifier, RRF normalization, compiled truth boost, cosine re-scoring, and 5-layer dedup. Added --detail flag explanation and benchmark instructions. CHANGELOG.md: added search quality entries to v0.9.3 (intent classifier, --detail flag, gbrain eval, CJK fix). Credited @4shut0sh and @YIING99. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * docs: headline benchmark gains in changelog * docs: add community attribution rule to CHANGELOG voice section --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: YIING99 <yiing99@users.noreply.github.com> Co-authored-by: 4shut0sh <4shut0sh@users.noreply.github.com>
218 lines
7.8 KiB
TypeScript
218 lines
7.8 KiB
TypeScript
/**
|
|
* Search Quality E2E Tests
|
|
*
|
|
* Tests the full search pipeline against PGLite with seeded pages and
|
|
* structured mock embeddings (basis vectors). No OpenAI API calls needed.
|
|
*
|
|
* Validates: compiled truth boost, detail parameter, source-aware dedup,
|
|
* chunk_id/chunk_index in results, and getEmbeddingsByChunkIds.
|
|
*/
|
|
|
|
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
|
import { PGLiteEngine } from '../../src/core/pglite-engine.ts';
|
|
import type { ChunkInput, SearchResult } from '../../src/core/types.ts';
|
|
|
|
let engine: PGLiteEngine;
|
|
|
|
// Create a basis vector embedding: dimension `idx` is 1.0, rest are 0.0
|
|
function basisEmbedding(idx: number, dim = 1536): Float32Array {
|
|
const emb = new Float32Array(dim);
|
|
emb[idx % dim] = 1.0;
|
|
return emb;
|
|
}
|
|
|
|
beforeAll(async () => {
|
|
engine = new PGLiteEngine();
|
|
await engine.connect({}); // in-memory
|
|
await engine.initSchema();
|
|
|
|
// Seed test pages with compiled_truth + timeline chunks
|
|
await engine.putPage('people/pedro', {
|
|
type: 'person',
|
|
title: 'Pedro Franceschi',
|
|
compiled_truth: 'Pedro is the co-founder of Brex. Expert in fintech and payments infrastructure.',
|
|
timeline: '2024-03-15: Met Pedro at YC dinner. Discussed AI security.',
|
|
});
|
|
|
|
// Seed chunks with structured embeddings
|
|
const pedroChunks: ChunkInput[] = [
|
|
{
|
|
chunk_index: 0,
|
|
chunk_text: 'Pedro is the co-founder of Brex. Expert in fintech and payments infrastructure.',
|
|
chunk_source: 'compiled_truth',
|
|
embedding: basisEmbedding(0), // direction 0 = fintech/compiled truth
|
|
token_count: 15,
|
|
},
|
|
{
|
|
chunk_index: 1,
|
|
chunk_text: '2024-03-15: Met Pedro at YC dinner. Discussed AI security and Crab Trap.',
|
|
chunk_source: 'timeline',
|
|
embedding: basisEmbedding(1), // direction 1 = meeting/timeline
|
|
token_count: 18,
|
|
},
|
|
];
|
|
await engine.upsertChunks('people/pedro', pedroChunks);
|
|
|
|
await engine.putPage('companies/variant', {
|
|
type: 'company',
|
|
title: 'Variant Fund',
|
|
compiled_truth: 'Variant is a crypto-native investment firm focused on web3 ownership economy.',
|
|
timeline: '2024-06-01: Variant announced new fund.',
|
|
});
|
|
|
|
const variantChunks: ChunkInput[] = [
|
|
{
|
|
chunk_index: 0,
|
|
chunk_text: 'Variant is a crypto-native investment firm focused on web3 ownership economy.',
|
|
chunk_source: 'compiled_truth',
|
|
embedding: basisEmbedding(2),
|
|
token_count: 14,
|
|
},
|
|
{
|
|
chunk_index: 1,
|
|
chunk_text: '2024-06-01: Variant announced new fund. $450M raised.',
|
|
chunk_source: 'timeline',
|
|
embedding: basisEmbedding(3),
|
|
token_count: 12,
|
|
},
|
|
];
|
|
await engine.upsertChunks('companies/variant', variantChunks);
|
|
|
|
await engine.putPage('concepts/ai-philosophy', {
|
|
type: 'concept',
|
|
title: 'AI Changes Who Gets to Build',
|
|
compiled_truth: 'AI democratizes building. The marginal cost of creation approaches zero.',
|
|
timeline: '2024-01-10: First wrote about AI and building access.',
|
|
});
|
|
|
|
const aiChunks: ChunkInput[] = [
|
|
{
|
|
chunk_index: 0,
|
|
chunk_text: 'AI democratizes building. The marginal cost of creation approaches zero. This changes who gets to build.',
|
|
chunk_source: 'compiled_truth',
|
|
embedding: basisEmbedding(4),
|
|
token_count: 20,
|
|
},
|
|
{
|
|
chunk_index: 1,
|
|
chunk_text: '2024-01-10: First wrote about AI and building access. Shared on X.',
|
|
chunk_source: 'timeline',
|
|
embedding: basisEmbedding(5),
|
|
token_count: 15,
|
|
},
|
|
];
|
|
await engine.upsertChunks('concepts/ai-philosophy', aiChunks);
|
|
});
|
|
|
|
afterAll(async () => {
|
|
await engine.disconnect();
|
|
});
|
|
|
|
describe('SearchResult fields', () => {
|
|
test('keyword search returns chunk_id and chunk_index', async () => {
|
|
const results = await engine.searchKeyword('Pedro');
|
|
expect(results.length).toBeGreaterThan(0);
|
|
const r = results[0];
|
|
expect(r.chunk_id).toBeDefined();
|
|
expect(typeof r.chunk_id).toBe('number');
|
|
expect(r.chunk_index).toBeDefined();
|
|
expect(typeof r.chunk_index).toBe('number');
|
|
});
|
|
|
|
test('vector search returns chunk_id and chunk_index', async () => {
|
|
const results = await engine.searchVector(basisEmbedding(0));
|
|
expect(results.length).toBeGreaterThan(0);
|
|
const r = results[0];
|
|
expect(r.chunk_id).toBeDefined();
|
|
expect(typeof r.chunk_id).toBe('number');
|
|
expect(r.chunk_index).toBeDefined();
|
|
expect(typeof r.chunk_index).toBe('number');
|
|
});
|
|
});
|
|
|
|
describe('detail parameter', () => {
|
|
test('detail=low returns only compiled_truth chunks', async () => {
|
|
const results = await engine.searchKeyword('Pedro', { detail: 'low' });
|
|
for (const r of results) {
|
|
expect(r.chunk_source).toBe('compiled_truth');
|
|
}
|
|
});
|
|
|
|
test('detail=high returns all chunk sources', async () => {
|
|
const results = await engine.searchKeyword('Pedro', { detail: 'high' });
|
|
// Should include at least compiled_truth (might include timeline depending on tsvector match)
|
|
expect(results.length).toBeGreaterThan(0);
|
|
});
|
|
|
|
test('detail=low on vector search filters to compiled_truth', async () => {
|
|
// Use a timeline-direction embedding — with detail=low, should get no results
|
|
// or only compiled_truth results
|
|
const results = await engine.searchVector(basisEmbedding(1), { detail: 'low' });
|
|
for (const r of results) {
|
|
expect(r.chunk_source).toBe('compiled_truth');
|
|
}
|
|
});
|
|
|
|
test('default detail (medium) returns all sources', async () => {
|
|
const results = await engine.searchKeyword('Pedro');
|
|
// No filter applied, should return whatever matches
|
|
expect(results.length).toBeGreaterThan(0);
|
|
});
|
|
});
|
|
|
|
describe('getEmbeddingsByChunkIds', () => {
|
|
test('returns embeddings for valid chunk IDs', async () => {
|
|
const searchResults = await engine.searchVector(basisEmbedding(0));
|
|
expect(searchResults.length).toBeGreaterThan(0);
|
|
|
|
const ids = searchResults.map(r => r.chunk_id).filter((id): id is number => id != null);
|
|
const embMap = await engine.getEmbeddingsByChunkIds(ids);
|
|
|
|
expect(embMap.size).toBeGreaterThan(0);
|
|
for (const [id, emb] of embMap) {
|
|
expect(emb).toBeInstanceOf(Float32Array);
|
|
expect(emb.length).toBe(1536);
|
|
}
|
|
});
|
|
|
|
test('returns empty map for empty ID list', async () => {
|
|
const embMap = await engine.getEmbeddingsByChunkIds([]);
|
|
expect(embMap.size).toBe(0);
|
|
});
|
|
|
|
test('returns empty map for non-existent IDs', async () => {
|
|
const embMap = await engine.getEmbeddingsByChunkIds([999999, 999998]);
|
|
expect(embMap.size).toBe(0);
|
|
});
|
|
});
|
|
|
|
describe('keyword search without DISTINCT ON', () => {
|
|
test('returns multiple chunks per page', async () => {
|
|
// Search for something that matches a page with multiple chunks
|
|
const results = await engine.searchKeyword('Pedro', { limit: 10 });
|
|
const pedroChunks = results.filter(r => r.slug === 'people/pedro');
|
|
// Should be able to return more than 1 chunk per page
|
|
// (depends on tsvector matching — Pedro is in page title/search_vector)
|
|
expect(results.length).toBeGreaterThan(0);
|
|
});
|
|
});
|
|
|
|
describe('compiled truth boost (vector search validates ordering)', () => {
|
|
test('compiled_truth chunks rank first with basis vector queries', async () => {
|
|
// Query with the compiled_truth direction for Pedro (basis 0)
|
|
const results = await engine.searchVector(basisEmbedding(0), { limit: 5 });
|
|
expect(results.length).toBeGreaterThan(0);
|
|
// The closest result should be the compiled_truth chunk (basis 0)
|
|
expect(results[0].chunk_source).toBe('compiled_truth');
|
|
expect(results[0].slug).toBe('people/pedro');
|
|
});
|
|
|
|
test('timeline chunks rank first when queried with timeline direction', async () => {
|
|
// Query with the timeline direction for Pedro (basis 1)
|
|
const results = await engine.searchVector(basisEmbedding(1), { limit: 5 });
|
|
expect(results.length).toBeGreaterThan(0);
|
|
expect(results[0].chunk_source).toBe('timeline');
|
|
expect(results[0].slug).toBe('people/pedro');
|
|
});
|
|
});
|