* feat(extract): quarantine lane for auto-extracted entities from untrusted input (#160) extractAndEnrich regex-extracts entity names from arbitrary ingested text and creates people/ + companies/ stub pages. Those writes are now trust- gated end to end: - src/core/extraction-review.ts: new marker module (sibling of quarantine.ts / embed-skip.ts, frontmatter-key pattern, no migration). Untrusted-input stubs carry `provenance: auto-extracted` + `status: unverified`; the shared unverifiedExtractionFragment() is the single SQL source of truth for every consumer. - enrichment-service: enrichEntity/enrichEntities/extractAndEnrich take EnrichmentTrustOptions; only an explicit trusted:true writes authoritative pages (fail-closed, mirrors the OperationContext.remote invariant). Also threads sourceId through the write path. - retrieval: unverified stubs rank as ordinary content — skipped by the compiled-truth fusion boost (stampUnverifiedExtractions pre-fusion on all three hybrid paths + keyword-only opt-out) and by the people// companies/ namespace source-boost (guard inside buildSourceFactorCase, shared by both engines' search SQL). Results carry `unverified: true`. New engine method getUnverifiedExtractionPageIds in BOTH engines. - ops (contract-first): extract_entities (direct write only for ctx.remote === false + --trusted-extraction; everything else quarantines), extraction_pending (read, source-scoped list), extraction_review (owner-only batch promote/reject; promote flips status to verified keeping provenance for audit, reject soft-deletes). - doctor: unverified_extractions check warns on stubs older than N days (default 7) with the exact review commands. Tests: test/extraction-review.test.ts (PGLite: fail-closed matrix incl. remote-unset, fusion boost skip, review queue, doctor, hostile-transcript e2e proving fake entities land quarantined and rank below a verified page of equal lexical relevance) + test/e2e/extraction-review-postgres.test.ts (live Postgres parity, verified against pgvector:pg16). sql-ranking expectations updated to current state. Docs: KEY_FILES + RETRIEVAL + llms rebuild. Closes #160 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(extract): close vector-arm source-boost gap + harden extract_entities (#160 review round) Adversarial review of the quarantine lane found the people//companies/ 1.2x source factor still applied to unverified stubs inside searchVector's pre-LIMIT re-rank (a different multiplier from the fusion-level 2.0x the lane already cancels — and applied early enough to evict legitimate pages from the candidate pool, which nothing downstream can restore). - buildSourceFactorCase gains an optional unverifiedGuardColumn for the bare-slug re-rank form; both engines' hnsw_candidates CTEs now project the guard predicate as `unverified_stub` and the factor CASE checks it first. Wrong "fusion covers the vector arm" comment corrected. - extract_entities resource guards: 200k-char input cap (loud reject), 200-entity cap surfaced as `truncated` + `entities_found`; the library extractAndEnrich gets the same default cap. (OperationContext has no abort signal field — caps are the bound.) - extraction_review promote is now a targeted JSONB-merge UPDATE instead of putPage, so non-carried columns (page_kind, content_hash) can't be reset by the upsert. - extraction_pending applies buildVisibilityClause (archived-source stubs no longer list). - Wording: op description + module header now state the marker-strip assumption plainly (markers are ordinary frontmatter; the boundary against wholesale rewrite is put_page write authz) and document the CREATE-only scope of the lane. Tests: vector-arm factor-1.0 pinned on BOTH engines (PGLite unit + live Postgres e2e, identical basis embeddings → score ratio is the factor); resource-guard test (oversize reject + 300-entity flood capped at 200); guard-column form pinned in the buildSourceFactorCase unit test. search/ suite (340), sql-ranking, searchvector-maxpool, title-retrieval- arm, rrf-source-key, doctor, ops, cli suites all green; JSONB guards clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Garry Tan <garrytan@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
10 KiB
Why the hybrid + graph stack works
Vector search alone underdelivers on real personal-knowledge queries. This doc explains why gbrain layers four strategies together and how they compound.
The four strategies in concert
- Vector (HNSW on pgvector) — semantic similarity. Catches "who works on retrieval quality at YC?" → pages mentioning "Garry Tan + retrieval" even when the user never typed "YC".
- BM25 keyword — lexical match. Catches names, exact phrases, code identifiers, anything where the user remembers the literal token. Survives the cases where vector search drifts into thematic neighbors.
- Reciprocal-rank fusion (RRF) — merges vector + keyword rankings without weighting one over the other globally. Each strategy gets to vote.
- Knowledge graph traversal — follows typed edges. Catches "what did Bob invest in this quarter?" by walking
bob ── invested_in ──> company ── dated ──> Q1. Vector search can't see causal chains; the graph can.
Why each one alone fails
Vector only. Returns chunks semantically close to the query. Misses any factual relationship not directly encoded in the embedding. "Companies in Garry's portfolio" returns essays about portfolios, not company pages.
Keyword only (ripgrep-style). Brittle to phrasing. "Who works on retrieval?" misses pages that say "search ranking" instead of "retrieval." Garbage on synonyms, near-misses, or paraphrases.
Graph only. Excellent at "neighbors of Alice" but blind to anything not yet linked. Sparse on fresh pages until backlinks accumulate.
Hybrid (vector + keyword + RRF), no graph. Decent at "what is X?" type queries. Fails on "what is Y's relationship to X?" — those are graph queries and no amount of embedding tuning recovers them.
The benchmark
BrainBench (corpus + harness in the sibling gbrain-evals repo) measures retrieval P@5, R@5, MRR, nDCG@5 on a 240-page Opus-generated rich-prose corpus.
| Strategy | P@5 | R@5 | Notes |
|---|---|---|---|
| ripgrep BM25 only | ~18 | ~75 | Lexical-only baseline |
| vector-only RAG | ~18 | ~80 | Standard RAG implementation |
| gbrain graph-disabled (hybrid + RRF, no graph traversal) | ~18 | ~85 | Hybrid alone |
| gbrain default (full stack) | 49.1 | 97.9 | Graph + extract-quality lift |
+31 P@5 points from the graph + extract quality work. The graph isn't a marginal feature; it's the load-bearing wall.
Auto-link: why zero-LLM-call edge extraction works
Every put_page runs extractEntityRefs on the markdown body. It matches:
- Standard markdown links:
[Garry Tan](wiki/people/garry-tan) - Obsidian wikilinks:
[[wiki/people/garry-tan|Garry Tan]] - Typed-link blockquotes:
> **Convention:** see [path](path).
Three regexes, zero LLM tokens, single SQL addLinksBatch call with INSERT ... SELECT FROM jsonb_to_recordset(($1::jsonb)->'rows') JOIN pages ON CONFLICT DO NOTHING RETURNING 1 (free-text-safe; the prior unnest(${arr}::text[]) form crashed on calendar/Zoom context per gbrain#1861). The graph grows on every write at near-zero cost. On a 17K-page brain, full graph extract completes in seconds.
Heuristic link-type inference (attended, works_at, invested_in, founded, advises) fires from surrounding sentence context — also LLM-free. Power users who want richer types add them via the typed-link blockquote convention.
ZeroEntropy as reranker: 60% top-1 reshuffle
v0.36.0.0 ships ZeroEntropy's zerank-2 as the default reranker (on for the balanced mode bundle). On a real-corpus benchmark across 20 queries, zerank-2 reshuffles 60% of top-1 results after the hybrid + RRF + graph stack. That's the headline number.
The mechanical reason: hybrid ranking is locally optimal per strategy but globally suboptimal. A cross-encoder reranker reads the query + each candidate document jointly, with full attention. It catches the cases where the vector + keyword + graph signals all agreed on a document that's semantically related but topically wrong.
The cost: +150ms p50 latency, ~$0.025/M tokens. Disabled with gbrain config set search.reranker.enabled false. For agent loops that do downstream LLM work after retrieval, the latency is invisible.
Source-aware ranking
Hybrid search applies a source-factor CASE expression at the SQL layer (lives in src/core/search/sql-ranking.ts). Curated content like originals/, concepts/, writing/ outranks bulk content like your-openclaw/chat/, daily/, media/x/. Hard-exclude prefixes (test/, attachments/, .raw/) filter at retrieval, not post-rank.
archive/ is deliberately NOT hard-excluded (issue #1777): it holds high-signal historical content users expect to find, so it is demoted (0.5x in DEFAULT_SOURCE_BOOSTS), not hidden. The demote is a prior applied in the outer SQL re-rank; the cross-encoder reranker (balanced/tokenmax modes) can still PROMOTE an archive page that survives the demote into the rerank candidate window — it is not an unconditional suppression. gbrain doctor's hidden_by_search_policy check reports how many chunked pages remain hidden by the surviving exclude prefixes.
The boost map is configurable via GBRAIN_SOURCE_BOOST env var or per-call SearchOpts.exclude_slug_prefixes. Temporal queries (detail: 'high') bypass the boost so chat pages re-surface for time-sensitive lookups.
Named-thing retrieval (per-page pool + title + alias + evidence)
A brain organized around chosen names (Mingtang, Hall of Light) needs more than
embedding proximity. Four layers, added after the incident in
RETRIEVAL_MAXPOOL_INCIDENT.md:
- Per-page max-pool —
searchVector(both engines) collapses chunk-grain candidates to the best chunk per page (DISTINCT ON (slug)) over the full candidate set before the userLIMIT, via the sharedbuildBestPerPagePoolCteinsql-ranking.ts. The vector side returns N distinct pages by best chunk, not N chunks that collapse to fewer pages downstream. - Title-phrase boost — when the normalized query is a contiguous token-run
inside
page.title(or an exact full-title match), a floor-ratio-gated, bounded multiplier fires (applyTitleBoost,search.title_boostknob). A query that is a phrase from the title can't lose to a body chunk by luck. - Alias hop — free-text
aliases:frontmatter is projected into apage_aliasestable (separate from theslug_aliaseswikilink redirect) and consulted at query time: a full normalized-query match injects/boosts the canonical page (applyAliasHop). The only layer that bridges true synonyms with zero surface overlap ("Hall of Light" → the Mingtang page). Backfill existing pages withgbrain reindex --aliases. - Evidence contract — every result carries
evidence(alias_hit | exact_title_match | high_vector_match | keyword_exact | weak_semantic) andcreate_safety(exists | probable | unknown). An agent deciding "is this page already here, safe to NOT write a duplicate?" keys offcreate_safety, not a raw blended score.
Extraction quarantine lane (issue #160): pages carrying the unverified
auto-extracted markers (frontmatter provenance: auto-extracted +
status: unverified, see src/core/extraction-review.ts) rank as ordinary
content — they are skipped by the compiled-truth fusion boost and by the
people//companies/ namespace source-boost, and every search result from
such a page carries unverified: true so agents can label the provenance.
Promote or reject them via gbrain extraction-pending / gbrain extraction-review.
The search MCP/CLI op is cheap-hybrid (vector + keyword + RRF + pool +
title + alias, expansion off); query is the full-control variant. NamedThingBench
(gbrain eval retrieval-quality) gates these families on every PR. Diagnose a
specific miss with gbrain search diagnose "<q>" --target <slug>.
Intent-aware query rewriting
src/core/search/intent.ts classifies queries into entity, temporal, event, or general. Each routes through different ranking knobs:
- Entity queries ("who works at X?") apply a higher graph-traversal weight.
- Temporal queries ("what happened last week?") bypass source-boost so chat/daily pages surface.
- Event queries ("Acme AI Series A") engage the timeline index.
- General queries hit the standard hybrid stack.
The classifier is deterministic (no LLM call). Wrong classification degrades gracefully — the hybrid stack still works without it.
Multi-query expansion
For detail: 'high' searches, src/core/search/expansion.ts runs a Haiku-class LLM call to produce 2-3 query variants. Each variant runs through the full hybrid stack; results merge via RRF. Catches synonym misses without recall loss.
Expansion is opt-in per mode bundle (tokenmax on by default; balanced + conservative off). Default off in the cheap tiers because the LLM call adds ~$0.001/query and ~200ms — real money at scale.
Putting it together
The full pipeline for a query op:
intent classify
│
▼
expansion (if enabled)
│
▼
hybrid search:
├── vector (HNSW on chunk embeddings)
├── keyword (BM25 via tsvector)
├── relational (v0.42.34.0: typed-edge recall arm — relational queries only)
├── source-aware re-rank (CASE in SQL)
└── RRF fusion → top 30
│
▼
graph augment (typed-edge traversal from any seed)
│
▼
reranker (zerank-2 cross-encoder, top 30 → reordered)
│
▼
token-budget enforcement (per mode bundle)
│
▼
deduplication (same slug, different chunks → keep best)
│
▼
results
Each stage is testable in isolation. Each stage is replaceable. The whole pipeline is < 1ms of orchestration cost; the latency budget goes to the upstream HTTP calls (embedding, rerank) and the index scans.
How to verify on your own brain
# Run the public LongMemEval benchmark
gbrain eval longmemeval datasets/longmemeval_s.jsonl
# Capture your own queries and replay against retrieval changes
export GBRAIN_CONTRIBUTOR_MODE=1
# ... use gbrain normally ...
gbrain eval export > before.ndjson
# ... change something ...
gbrain eval replay --against before.ndjson
# A/B retrieval strategies on a labeled fixture
gbrain eval --qrels labels.tsv --config balanced.json
Methodology + metric glossary in docs/eval/SEARCH_MODE_METHODOLOGY.md.