mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-28 14:59:47 +00:00
* feat(facts): typed-claim substrate + cycle correctness fixes (v0.35.6 wave 1/3) Schema (migration v67): - Add four optional typed-claim columns to facts: claim_metric TEXT, claim_value DOUBLE PRECISION, claim_unit TEXT, claim_period TEXT - Partial index facts_typed_claim_idx ON (entity_slug, claim_metric, valid_from) WHERE claim_metric IS NOT NULL - All nullable, metadata-only on both engines Fence layer: - ParsedFact (facts-fence.ts) gains optional claimMetric/Value/Unit/Period - Parser tolerates both 10-cell (legacy) and 14-cell (widened) rows - Renderer emits 14 cells iff any row has typed data; otherwise stays 10-cell so existing fences don't widen on unrelated edits - Numeric value cell tolerates comma thousand separators (50,000 -> 50000) Extract pipeline (D-CDX-2, D-ENG-1): - src/core/facts/extract.ts (the actual Haiku call site, NOT extract-facts.ts cycle phase) extends its system prompt to emit typed fields for metric-shaped claims - extractFactsFromFenceText gains optional pageEffectiveDate. Precedence: fence-row validFrom > pageEffectiveDate > undefined (engine defaults to now) - normalizeMetricLabel: 15-entry seed map for common founder metrics (mrr, arr, runway, headcount, team_size, cac, ltv, gross_margin, burn_rate, cash, users, mau, dau, churn_rate, revenue); unknown labels lowercase + space->_ Engine extensions: - NewFact + insertFact + insertFacts in both engines accept the four typed columns (all nullable) - Cycle phase extract-facts.ts threads page.effective_date through AND batch-embeds via gateway.embed() before insertFacts (D-CDX-3 fix for cycle-inserted facts arriving with embedding=NULL) Consolidate fix (D-CDX-4 — Codex F4): - Replace MAX(row_num)+1 INSERT with semantic upsert on (page_id, claim, since_date). Re-running the full cycle on stable input produces zero new takes — fixes the pre-existing duplicate-takes bug after extract_facts wipes consolidated_at - Chronological valid_until writeback per cluster: sort by (valid_from ASC, id ASC), walk pairs, set older.valid_until = newer.valid_from Tests: - test/migrate.test.ts +6 cases for v67 shape + materialization + nullable backward compat - test/facts-fence-typed.test.ts (new, 17 cases): parser+renderer round-trip, normalization seed map coverage, valid_from precedence three-branch - test/consolidate-valid-until.test.ts (new, 4 cases): chronological writeback (R4a), same-day id tiebreaker, cycle re-run zero duplicates (R4b/R7), valid_until idempotency - test/schema-bootstrap-coverage.test.ts: add four typed-claim columns to COLUMN_EXEMPTIONS (migration co-defines the partial index, no forward reference to bootstrap) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(trajectory): find_trajectory MCP op + eval/founder CLIs (v0.35.6 wave 2/3) Engine method (D-CDX-1, D-CDX-6): - BrainEngine.findTrajectory(opts) on both Postgres and PGLite - TrajectoryOpts: scalar sourceId fast path + sourceIds federated array (mirrors v0.34.1.0 search* dual pattern) - opts.remote: when true, SQL adds AND visibility='world' so OAuth read clients see only world-visibility facts (mirrors recall's posture — closes the F7 privacy regression Codex caught in plan review) - Single SQL query, ORDER BY valid_from ASC, id ASC for deterministic output (R3 pin). Returns TrajectoryPoint[] including raw embedding so the caller can compute drift without a second round-trip Pure function library (src/core/trajectory.ts, new): - detectRegressions(points, threshold): walks consecutive (metric, value) pairs per metric; emits when newer drops >= threshold below older. 10% default, override via GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD - computeDriftScore(points): 1 - mean(cosine(emb[i], emb[i-1])) over embedded points; clamped [0,1]; null when <3 embedded points (D-ENG-3 graceful degradation) - computeTrajectoryStats(points): composed shape returning both - TRAJECTORY_SCHEMA_VERSION = 1 — additive-only across releases (R5) MCP op (src/core/operations.ts): - find_trajectory: scope read, NOT localOnly. Routes through sourceScopeOpts(ctx) for federated isolation AND threads ctx.remote for visibility filtering. Strips raw Float32Array embeddings from the wire shape; converts valid_from to YYYY-MM-DD string - Registered in operations array after find_experts - FIND_TRAJECTORY_DESCRIPTION in operations-descriptions.ts CLIs: - gbrain eval trajectory <entity> [--metric M] [--since D] [--until D] [--limit N] [--json] — chronological human view with [REGRESSION] inline annotation; thin-client routing via callRemoteTool(find_trajectory). Dispatched in src/commands/eval.ts sub-subcommand block - gbrain founder scorecard <entity> [--since D] [--until D] [--json] — pure aggregation over Phase 2's substrate. Four signals: claim_accuracy (over resolved takes), consistency, growth_trajectory, red_flags. computeFounderScorecard exported for tests. Registered as top-level command in cli.ts; added to CLI_ONLY set Tests (45 cases across 5 files): - test/engine-find-trajectory.test.ts: 18 cases — chronological order, source scoping (scalar + federated), visibility filter on remote=true, metric + since/until filters, regression detection at threshold boundaries, drift score with various embedding states - test/operations-find-trajectory.test.ts: 9 cases — op registration, param validation, JSON envelope shape, R5 schema_version: 1, embedding stripped from wire, R6 visibility filter, source scoping - test/eval-trajectory.test.ts: 7 cases — arg parsing, --help, --json envelope, regression annotation, --metric filter, empty entity - test/founder-scorecard.test.ts: 9 cases — empty inputs no-NaN (G2), claim_accuracy math, consistency math, growth_trajectory math, red_flags fire for regression / narrative_drift / missed_prediction - test/eval-contradictions/no-valid-until-write.test.ts: 4 cases — R1 (probe never writes valid_until under eval-contradictions/) + R8 (only allow-listed files write valid_until anywhere in src/) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: v0.35.6.0 — CHANGELOG + VERSION + docs + migration note Bumps to v0.35.6.0 (next-minor after master's v0.35.5.1 — typed-claim substrate + trajectory + founder scorecard is a new user-facing feature surface, not a fix). - VERSION + package.json synced - CHANGELOG.md release-summary block in the wave-style voice, lead with what the user can now DO. Sections: typed metric claims in the fence, chronological metric trajectories, founder scorecard, MCP find_trajectory op, cycle re-run idempotency fix, embedding-on-insert fix, valid_from precedence fix. To-take-advantage-of block with verification + opt-in fence syntax example - CLAUDE.md Key Files entry consolidating the wave across eval-trajectory.ts + founder-scorecard.ts + trajectory.ts. Names every D-ENG / D-CDX decision and the Codex outside-voice F-numbers - skills/migrations/v0.35.6.md agent-readable migration note. Includes fence-syntax example for typed-claim rows so downstream agents start emitting them. Iron-rule contracts called out (R1 + R8 + R7 + visibility) - llms-full.txt regenerated to reflect the new CLAUDE.md entry Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: post-ship sync for v0.35.7.0 — trajectory + founder scorecard - README.md: add `gbrain eval trajectory` to EVAL section, add new TEMPORAL block covering `gbrain founder scorecard` + the GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD env override; add v0.35.7 "What's new" paragraph below the v0.28.8 LongMemEval blurb - AGENTS.md: new bullet under Common tasks teaching agents to reach for `gbrain eval trajectory` / `gbrain founder scorecard` / the `find_trajectory` MCP op when asked to evaluate a founder/company over time - docs/contradictions.md: append "Temporal axis follow-on (v0.35.3.1 + v0.35.7)" subsection under See also, cross-linking the trajectory substrate and naming the auto-supersession.ts:4 invariant preserved by both the verdict enum (probe side) and consolidate's valid_until writeback (cycle side) - CLAUDE.md: fix stale (v0.35.4) tag on the trajectory entry to (v0.35.7) — version got rebumped twice during the merge wave - skills/migrations/v0.35.7.md renamed to v0.35.7.0.md for consistency with the v0.35.0.0.md / v0.14.0.md / etc naming convention - llms-full.txt regenerated to reflect the CLAUDE.md edit Coverage map (Diataxis): /eval trajectory CLI ✅ ref (README, AGENTS) ✅ how-to (CHANGELOG) ❌ tutorial /founder scorecard CLI ✅ ref (README, AGENTS) ✅ how-to (CHANGELOG) ❌ tutorial find_trajectory MCP op ✅ ref (CLAUDE.md, AGENTS, contradictions.md) typed-claim fence cols ✅ ref (skills/migrations/v0.35.7.0.md, CHANGELOG) Migration v67 ✅ ref (CLAUDE.md, CHANGELOG) No tutorial / explanation gaps worth filling in this PR — the migration note's fence-syntax example already covers the "first typed claim" walkthrough. ARCHITECTURE diagrams not drifted (the trajectory work extends existing facts/takes infrastructure; no new component boxes). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
167 lines
8.4 KiB
Markdown
167 lines
8.4 KiB
Markdown
# gbrain eval suspected-contradictions (v0.32.6)
|
||
|
||
The contradiction probe samples retrieval results, asks an LLM judge whether
|
||
any pair contradicts on a factual claim relevant to the user's query, and
|
||
aggregates into a calibrated report. The output is data — the operator
|
||
decides what to act on. This doc covers the architecture, severity rubric,
|
||
how to interpret the headline number, and when to act.
|
||
|
||
## Why this exists
|
||
|
||
gbrain handles contradictions for *curated* pages via compiled-truth-plus-
|
||
timeline and source-boost: when `companies/acme.md` says MRR is $2M and a
|
||
chat transcript from 2024 says MRR was $50K, the curated page outranks the
|
||
chat. `takes.active` filtering hides explicitly-superseded takes. Recency
|
||
decay biases ranking toward fresher content per source-tier.
|
||
|
||
What none of those mechanisms measure: how often do unmarked semantic
|
||
contradictions actually surface in retrieval? Without a probe, every
|
||
"should we build the bigger swing (chunk-level `revises` field + ranking
|
||
change)" decision is vibes. The probe produces evidence.
|
||
|
||
## Architecture
|
||
|
||
```
|
||
┌──────────────────────────────────────┐
|
||
│ gbrain eval suspected-contradictions │
|
||
└──────────────────┬───────────────────┘
|
||
│
|
||
┌──────────────────▼───────────────────┐
|
||
│ For each query: hybridSearch top-K │
|
||
│ → cross_slug_chunks + intra_page │
|
||
│ chunk-vs-take pairs │
|
||
└──────────────────┬───────────────────┘
|
||
│
|
||
┌──────────────────▼───────────────────┐
|
||
│ Date pre-filter: skip pairs whose │
|
||
│ dates are >30d apart (Codex fix: │
|
||
│ same-paragraph-dual-date overrides) │
|
||
└──────────────────┬───────────────────┘
|
||
│
|
||
┌──────────────────▼───────────────────┐
|
||
│ Persistent cache lookup │
|
||
│ (chunk_a_hash, chunk_b_hash, model, │
|
||
│ prompt_version, truncation_policy) │
|
||
└────────┬─────────┬────────────────────┘
|
||
hit│ │miss
|
||
│ ▼
|
||
│ ┌─────────────────────────┐
|
||
│ │ LLM judge call │
|
||
│ │ → JudgeVerdict │
|
||
│ │ confidence floor ≥ 0.7 │
|
||
│ └─────────┬───────────────┘
|
||
│ │
|
||
▼ ▼
|
||
┌──────────────────────────────────────┐
|
||
│ Aggregate per-query + global stats │
|
||
│ Wilson 95% CI on headline % │
|
||
│ source-tier breakdown │
|
||
│ hot pages + resolution proposals │
|
||
└──────────────────┬───────────────────┘
|
||
│
|
||
▼
|
||
ProbeReport JSON
|
||
│
|
||
┌──────────────────┼──────────────────────┬───────────────┐
|
||
▼ ▼ ▼ ▼
|
||
doctor (M1) MCP (M3) synthesize (M2) trend (M5)
|
||
surfaces find_contradictions informational persistent
|
||
findings op for agents block in prompt tracking
|
||
```
|
||
|
||
## Severity rubric
|
||
|
||
The judge assigns severity per finding:
|
||
|
||
| Level | Rubric | Example |
|
||
|---|---|---|
|
||
| `low` | naming/format differences | "Alice Smith" vs "A. Smith" |
|
||
| `medium` | factual values that may be stale | revenue figure, headcount, valuation |
|
||
| `high` | identity / structural claims | founder/CEO/CFO role, company status |
|
||
|
||
Doctor sorts findings by severity DESC. The MCP op accepts a severity filter
|
||
so agents can fetch just the high-priority items.
|
||
|
||
## How to interpret the headline number
|
||
|
||
The probe outputs `queries_with_contradiction / queries_evaluated` with a
|
||
Wilson 95% confidence interval:
|
||
|
||
```
|
||
Queries with >=1 contradiction: 12 / 50 (24%) Wilson CI 95%: 14–37%
|
||
```
|
||
|
||
What this says: with 95% confidence, the true rate is between 14% and 37%.
|
||
The 24% point estimate is the most-likely-value but bounded by sampling
|
||
noise. **`small_sample_note` fires when n < 30** — at that scale the CI is
|
||
too wide to act on.
|
||
|
||
Decision criteria for the bigger swing (chunk-level `revises` field):
|
||
|
||
| Wilson CI lower bound | What it says | Action |
|
||
|---|---|---|
|
||
| < 5% | Source-boost + recency-decay + curated pages handle the load | Stop here; this is the right scope |
|
||
| 5–15% | Real but bounded | Operator decides whether the cost justifies the swing |
|
||
| > 15% | Real and substantial | Plan the bigger swing in v0.34+ |
|
||
|
||
## When to act on findings
|
||
|
||
Each finding ships with a `resolution_command` field — paste-ready:
|
||
|
||
- `gbrain takes supersede <slug> --row N` — newer take should replace
|
||
the older chunk text on the same page (intra_page kind).
|
||
- `gbrain dream --phase synthesize --slug <slug>` — compiled_truth for
|
||
the curated entity needs an update (cross_slug curated-vs-bulk).
|
||
- `gbrain takes mark-debate <slug> --row N` — intentional disagreement
|
||
(e.g., two opinions you want to keep both of).
|
||
- `# manual review: <a> vs <b>` — judge wasn't sure; operator decides.
|
||
|
||
Run `gbrain eval suspected-contradictions review --severity high` to
|
||
inspect findings without re-running the probe.
|
||
|
||
## Cost model
|
||
|
||
Default judge is `claude-haiku-4-5` at ~$1/Mtok in, $5/Mtok out. With
|
||
the v0.32.6 truncation at 1500 chars per pair, ~500 input + 80 output
|
||
tokens per judge call. Budget cap defaults to $5 in TTY / $1 non-TTY.
|
||
|
||
- ~$0.0006 per judge call
|
||
- ~$0.005 per query (after date pre-filter + cache hits)
|
||
- ~$0.50 per 100 queries
|
||
|
||
The persistent cache means nightly runs against the same query set
|
||
pay near-zero on re-runs (until you bump PROMPT_VERSION).
|
||
|
||
## Trust posture
|
||
|
||
- Probe never mutates the brain. Runs only read pages/takes/chunks.
|
||
Writes go only to `eval_contradictions_runs` and `eval_contradictions_cache`.
|
||
- MCP `find_contradictions` is read-scope. NOT in the subagent allowlist —
|
||
user-initiated only, not autonomous-action surface.
|
||
- Build-fixture script is local-only. The redactor + `isCleanForCommit`
|
||
gate makes accidental private-data commits hard, but the operator MUST
|
||
inspect every redaction before commit.
|
||
|
||
## See also
|
||
|
||
- Plan: `~/.claude/plans/system-instruction-you-are-working-hashed-dewdrop.md`
|
||
- CHANGELOG: `## [0.32.6]` entry covers the whole release.
|
||
- Cost discipline: `docs/eval-bench.md` for the recommended nightly cadence
|
||
+ trend-tracking workflow.
|
||
- **Temporal axis follow-on (v0.35.3.1 + v0.35.7):** v0.35.3.1 added a
|
||
six-member verdict enum (`no_contradiction | contradiction |
|
||
temporal_supersession | temporal_regression | temporal_evolution |
|
||
negation_artifact`) and threaded `pages.effective_date` into the judge
|
||
prompt so the probe stops crying wolf on legitimate change-over-time.
|
||
v0.35.7 lands the trajectory substrate the probe pointed at:
|
||
`gbrain eval trajectory <entity>` shows the chronological typed-claim
|
||
history with regressions flagged inline; `gbrain founder scorecard
|
||
<entity>` rolls up four signals (accuracy, consistency, growth
|
||
direction, red flags) into a stable JSON contract. MCP op
|
||
`find_trajectory` (read scope, visibility-filtered for remote callers)
|
||
exposes the same data to agents. The probe's `temporal_supersession`
|
||
verdict and the consolidate phase's `valid_until` writeback both
|
||
preserve the `auto-supersession.ts:4` "NEVER auto-applies" invariant
|
||
— the probe still emits paste-ready commands, only `consolidate`
|
||
writes `valid_until` (R1+R8 grep guard pins this).
|