mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* feat(facts): typed-claim substrate + cycle correctness fixes (v0.35.6 wave 1/3) Schema (migration v67): - Add four optional typed-claim columns to facts: claim_metric TEXT, claim_value DOUBLE PRECISION, claim_unit TEXT, claim_period TEXT - Partial index facts_typed_claim_idx ON (entity_slug, claim_metric, valid_from) WHERE claim_metric IS NOT NULL - All nullable, metadata-only on both engines Fence layer: - ParsedFact (facts-fence.ts) gains optional claimMetric/Value/Unit/Period - Parser tolerates both 10-cell (legacy) and 14-cell (widened) rows - Renderer emits 14 cells iff any row has typed data; otherwise stays 10-cell so existing fences don't widen on unrelated edits - Numeric value cell tolerates comma thousand separators (50,000 -> 50000) Extract pipeline (D-CDX-2, D-ENG-1): - src/core/facts/extract.ts (the actual Haiku call site, NOT extract-facts.ts cycle phase) extends its system prompt to emit typed fields for metric-shaped claims - extractFactsFromFenceText gains optional pageEffectiveDate. Precedence: fence-row validFrom > pageEffectiveDate > undefined (engine defaults to now) - normalizeMetricLabel: 15-entry seed map for common founder metrics (mrr, arr, runway, headcount, team_size, cac, ltv, gross_margin, burn_rate, cash, users, mau, dau, churn_rate, revenue); unknown labels lowercase + space->_ Engine extensions: - NewFact + insertFact + insertFacts in both engines accept the four typed columns (all nullable) - Cycle phase extract-facts.ts threads page.effective_date through AND batch-embeds via gateway.embed() before insertFacts (D-CDX-3 fix for cycle-inserted facts arriving with embedding=NULL) Consolidate fix (D-CDX-4 — Codex F4): - Replace MAX(row_num)+1 INSERT with semantic upsert on (page_id, claim, since_date). Re-running the full cycle on stable input produces zero new takes — fixes the pre-existing duplicate-takes bug after extract_facts wipes consolidated_at - Chronological valid_until writeback per cluster: sort by (valid_from ASC, id ASC), walk pairs, set older.valid_until = newer.valid_from Tests: - test/migrate.test.ts +6 cases for v67 shape + materialization + nullable backward compat - test/facts-fence-typed.test.ts (new, 17 cases): parser+renderer round-trip, normalization seed map coverage, valid_from precedence three-branch - test/consolidate-valid-until.test.ts (new, 4 cases): chronological writeback (R4a), same-day id tiebreaker, cycle re-run zero duplicates (R4b/R7), valid_until idempotency - test/schema-bootstrap-coverage.test.ts: add four typed-claim columns to COLUMN_EXEMPTIONS (migration co-defines the partial index, no forward reference to bootstrap) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(trajectory): find_trajectory MCP op + eval/founder CLIs (v0.35.6 wave 2/3) Engine method (D-CDX-1, D-CDX-6): - BrainEngine.findTrajectory(opts) on both Postgres and PGLite - TrajectoryOpts: scalar sourceId fast path + sourceIds federated array (mirrors v0.34.1.0 search* dual pattern) - opts.remote: when true, SQL adds AND visibility='world' so OAuth read clients see only world-visibility facts (mirrors recall's posture — closes the F7 privacy regression Codex caught in plan review) - Single SQL query, ORDER BY valid_from ASC, id ASC for deterministic output (R3 pin). Returns TrajectoryPoint[] including raw embedding so the caller can compute drift without a second round-trip Pure function library (src/core/trajectory.ts, new): - detectRegressions(points, threshold): walks consecutive (metric, value) pairs per metric; emits when newer drops >= threshold below older. 10% default, override via GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD - computeDriftScore(points): 1 - mean(cosine(emb[i], emb[i-1])) over embedded points; clamped [0,1]; null when <3 embedded points (D-ENG-3 graceful degradation) - computeTrajectoryStats(points): composed shape returning both - TRAJECTORY_SCHEMA_VERSION = 1 — additive-only across releases (R5) MCP op (src/core/operations.ts): - find_trajectory: scope read, NOT localOnly. Routes through sourceScopeOpts(ctx) for federated isolation AND threads ctx.remote for visibility filtering. Strips raw Float32Array embeddings from the wire shape; converts valid_from to YYYY-MM-DD string - Registered in operations array after find_experts - FIND_TRAJECTORY_DESCRIPTION in operations-descriptions.ts CLIs: - gbrain eval trajectory <entity> [--metric M] [--since D] [--until D] [--limit N] [--json] — chronological human view with [REGRESSION] inline annotation; thin-client routing via callRemoteTool(find_trajectory). Dispatched in src/commands/eval.ts sub-subcommand block - gbrain founder scorecard <entity> [--since D] [--until D] [--json] — pure aggregation over Phase 2's substrate. Four signals: claim_accuracy (over resolved takes), consistency, growth_trajectory, red_flags. computeFounderScorecard exported for tests. Registered as top-level command in cli.ts; added to CLI_ONLY set Tests (45 cases across 5 files): - test/engine-find-trajectory.test.ts: 18 cases — chronological order, source scoping (scalar + federated), visibility filter on remote=true, metric + since/until filters, regression detection at threshold boundaries, drift score with various embedding states - test/operations-find-trajectory.test.ts: 9 cases — op registration, param validation, JSON envelope shape, R5 schema_version: 1, embedding stripped from wire, R6 visibility filter, source scoping - test/eval-trajectory.test.ts: 7 cases — arg parsing, --help, --json envelope, regression annotation, --metric filter, empty entity - test/founder-scorecard.test.ts: 9 cases — empty inputs no-NaN (G2), claim_accuracy math, consistency math, growth_trajectory math, red_flags fire for regression / narrative_drift / missed_prediction - test/eval-contradictions/no-valid-until-write.test.ts: 4 cases — R1 (probe never writes valid_until under eval-contradictions/) + R8 (only allow-listed files write valid_until anywhere in src/) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: v0.35.6.0 — CHANGELOG + VERSION + docs + migration note Bumps to v0.35.6.0 (next-minor after master's v0.35.5.1 — typed-claim substrate + trajectory + founder scorecard is a new user-facing feature surface, not a fix). - VERSION + package.json synced - CHANGELOG.md release-summary block in the wave-style voice, lead with what the user can now DO. Sections: typed metric claims in the fence, chronological metric trajectories, founder scorecard, MCP find_trajectory op, cycle re-run idempotency fix, embedding-on-insert fix, valid_from precedence fix. To-take-advantage-of block with verification + opt-in fence syntax example - CLAUDE.md Key Files entry consolidating the wave across eval-trajectory.ts + founder-scorecard.ts + trajectory.ts. Names every D-ENG / D-CDX decision and the Codex outside-voice F-numbers - skills/migrations/v0.35.6.md agent-readable migration note. Includes fence-syntax example for typed-claim rows so downstream agents start emitting them. Iron-rule contracts called out (R1 + R8 + R7 + visibility) - llms-full.txt regenerated to reflect the new CLAUDE.md entry Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: post-ship sync for v0.35.7.0 — trajectory + founder scorecard - README.md: add `gbrain eval trajectory` to EVAL section, add new TEMPORAL block covering `gbrain founder scorecard` + the GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD env override; add v0.35.7 "What's new" paragraph below the v0.28.8 LongMemEval blurb - AGENTS.md: new bullet under Common tasks teaching agents to reach for `gbrain eval trajectory` / `gbrain founder scorecard` / the `find_trajectory` MCP op when asked to evaluate a founder/company over time - docs/contradictions.md: append "Temporal axis follow-on (v0.35.3.1 + v0.35.7)" subsection under See also, cross-linking the trajectory substrate and naming the auto-supersession.ts:4 invariant preserved by both the verdict enum (probe side) and consolidate's valid_until writeback (cycle side) - CLAUDE.md: fix stale (v0.35.4) tag on the trajectory entry to (v0.35.7) — version got rebumped twice during the merge wave - skills/migrations/v0.35.7.md renamed to v0.35.7.0.md for consistency with the v0.35.0.0.md / v0.14.0.md / etc naming convention - llms-full.txt regenerated to reflect the CLAUDE.md edit Coverage map (Diataxis): /eval trajectory CLI ✅ ref (README, AGENTS) ✅ how-to (CHANGELOG) ❌ tutorial /founder scorecard CLI ✅ ref (README, AGENTS) ✅ how-to (CHANGELOG) ❌ tutorial find_trajectory MCP op ✅ ref (CLAUDE.md, AGENTS, contradictions.md) typed-claim fence cols ✅ ref (skills/migrations/v0.35.7.0.md, CHANGELOG) Migration v67 ✅ ref (CLAUDE.md, CHANGELOG) No tutorial / explanation gaps worth filling in this PR — the migration note's fence-syntax example already covers the "first typed claim" walkthrough. ARCHITECTURE diagrams not drifted (the trajectory work extends existing facts/takes infrastructure; no new component boxes). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
153 lines
10 KiB
TypeScript
153 lines
10 KiB
TypeScript
/**
|
|
* v0.29 — Tool descriptions, extracted to a constants module so that:
|
|
* 1. The exact LLM-facing strings are pinnable in tests
|
|
* (`test/operations-descriptions.test.ts`).
|
|
* 2. Routing changes ship as data, not buried-in-handler edits.
|
|
* 3. The `salience-llm-routing.test.ts` Tier-2 eval has a stable surface
|
|
* to load tool definitions from.
|
|
*
|
|
* Description style:
|
|
* - Lead with what the tool does in one short sentence.
|
|
* - Include explicit triggers ("Use this when the user asks ...") that
|
|
* the LLM tool-selection prompt can match.
|
|
* - For redirect hints (query/search → salience), be blunt:
|
|
* "Do NOT run a semantic search for these."
|
|
*/
|
|
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
// New v0.29 ops
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
|
|
export const GET_RECENT_SALIENCE_DESCRIPTION =
|
|
"Returns pages recently touched and ranked by emotional + activity salience " +
|
|
"(deterministic 0..1 emotional_weight + take density + recency decay). " +
|
|
"Use this when the user asks what's been going on, what's notable, what's hot, " +
|
|
"anything crazy happening, or for any open-ended 'current state' question " +
|
|
"about themselves or their work. Do NOT run a semantic search for these — " +
|
|
"salience surfaces what's unusual without needing a search term.";
|
|
|
|
export const FIND_ANOMALIES_DESCRIPTION =
|
|
"Returns statistical anomalies in recent page activity, grouped by cohort " +
|
|
"(tag or type). Use this for questions about what stood out, what's unusual, " +
|
|
"or what changed recently. Returns explanatory cohorts (e.g. '15 pages tagged " +
|
|
"wedding touched on 2026-04-28, baseline 0.3/day') so you can speak about " +
|
|
"patterns the user wouldn't have searched for. Cohort kinds: tag, type. " +
|
|
"Year cohort is deferred to a later release.";
|
|
|
|
export const FIND_EXPERTS_DESCRIPTION =
|
|
"Answers 'who in my brain knows about <topic>'. Returns ranked person/company " +
|
|
"pages by expertise depth (sub-linear match score), relationship recency " +
|
|
"(exp decay with 6-month half-life), and salience. Use this for questions " +
|
|
"like 'who should I talk to about X', 'who knows about Y', 'find me someone " +
|
|
"who's worked on Z', or any expertise-routing intent. Filters at SQL to " +
|
|
"person + company pages — does NOT return notes or articles. Pair with " +
|
|
"--explain (CLI) to surface the per-result factor breakdown.";
|
|
|
|
export const GET_RECENT_TRANSCRIPTS_DESCRIPTION =
|
|
"Returns one-line summaries of recent raw conversation transcripts (NOT polished " +
|
|
"reflections). Use this FIRST for questions about 'what's going on with me', " +
|
|
"'what have I been thinking about', or anything personal/emotional. Raw " +
|
|
"transcripts are the canonical source for the user's own state — polished pages " +
|
|
"summarize and flatten. Local-only: rejects remote (MCP/HTTP) callers with a " +
|
|
"clear permission_denied; call via the gbrain CLI.";
|
|
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
// Redirect hints appended to existing op descriptions
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
|
|
export const LIST_PAGES_DESCRIPTION =
|
|
"List pages with optional filters. " +
|
|
"For 'what's recent / what did I touch this week' questions, use list_pages " +
|
|
"with sort=updated_desc instead of semantic search.";
|
|
|
|
export const QUERY_DESCRIPTION =
|
|
"Hybrid search with vector + keyword + multi-query expansion. " +
|
|
"For personal/emotional questions ('what's going on with me', 'anything notable', " +
|
|
"'how am I feeling'), prefer get_recent_salience, find_anomalies, or " +
|
|
"get_recent_transcripts. Semantic search returns polished pages and misses " +
|
|
"recent activity bursts. Do NOT assume words like 'crazy', 'notable', or 'big' " +
|
|
"mean impressive — they often mean difficult or emotionally charged.";
|
|
|
|
export const SEARCH_DESCRIPTION =
|
|
"Keyword search using full-text search. For personal/emotional questions, " +
|
|
"prefer get_recent_salience or find_anomalies — they surface activity bursts " +
|
|
"without needing a search term. " +
|
|
"For code-symbol questions (callers, callees, definitions, blast radius), use " +
|
|
"code_callers / code_callees / code_def / code_refs instead — those return " +
|
|
"structural graph data, not text chunks.";
|
|
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
// v0.32.6 — contradiction probe MCP surface (M3)
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
|
|
export const FIND_CONTRADICTIONS_DESCRIPTION =
|
|
"v0.32.6 — return suspected-contradiction findings from the most recent " +
|
|
"`gbrain eval suspected-contradictions` probe run, optionally filtered by slug " +
|
|
"and/or severity. Use this when the user asks 'what's inconsistent in my " +
|
|
"brain', 'show me contradictions about Acme', 'high-severity issues only', or " +
|
|
"wants to act on the probe's findings without re-running it. Returns " +
|
|
"{contradictions: [{a, b, severity, axis, confidence, resolution_command}]}. " +
|
|
"Reads the cached run row — does NOT trigger a new probe; users run " +
|
|
"`gbrain eval suspected-contradictions` for that.";
|
|
|
|
export const FIND_TRAJECTORY_DESCRIPTION =
|
|
"v0.35.4 — return the chronological claim trajectory for an entity (typed " +
|
|
"metric values over time, plus auto-detected regressions and narrative drift). " +
|
|
"Use this when the user asks 'how has Acme's MRR trended', 'show me what " +
|
|
"alice-example said about runway over time', 'is this founder consistent', " +
|
|
"'find regressions for fund-a's portfolio', or wants a time-series view of an " +
|
|
"entity's structured claims. Returns " +
|
|
"`{points: [{fact_id, valid_from, metric, value, unit, period, text, source_session, source_markdown_slug}], " +
|
|
"regressions: [{metric, from_value, from_date, to_value, to_date, delta_pct}], " +
|
|
"drift_score: number|null, schema_version: 1}`. Drift score 0 = stable narrative, " +
|
|
"1 = every consecutive claim is unrelated; null when fewer than 3 typed points " +
|
|
"exist. Visibility-filtered for remote callers (world-only); source-scoped by " +
|
|
"the caller's OAuth source binding. Pair with `gbrain founder scorecard <slug>` " +
|
|
"for an aggregated rollup of the same data.";
|
|
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
// v0.33.3 Cathedral III foundation — code-intelligence ops (MCP-exposed).
|
|
// Pre-v0.33.3 the callers/callees/def/refs commands were CLI-only — agents
|
|
// reached for grep because the MCP surface didn't expose them. These
|
|
// descriptions are resolver-grade so the LLM tool-selection prompt routes
|
|
// plan-mode questions straight to the right op.
|
|
//
|
|
// Style notes per the v0.34 eng review D10 finding: every description carries
|
|
// an inline example response so agents don't burn first-call context discovering
|
|
// shape. Pin via test/operations-descriptions.test.ts.
|
|
// ──────────────────────────────────────────────────────────────────────────────
|
|
|
|
export const CODE_CALLERS_DESCRIPTION =
|
|
"BEFORE editing any function, run code_callers with the symbol name to find " +
|
|
"every caller (the people who'd be affected by your change). Returns direct " +
|
|
"callers from the v0.20+ tree-sitter call graph. Use during plan-mode to size " +
|
|
"the change. Defaults to source-scoped; for multi-source brains pass source_id " +
|
|
"or all_sources=true. " +
|
|
"Returns: `{symbol, count, callers: [{from_symbol_qualified, to_symbol_qualified, edge_type, resolved}]}`. " +
|
|
"Example: `{symbol:'parseMarkdown', count:4, callers:[{from_symbol_qualified:'callerInA', " +
|
|
"to_symbol_qualified:'parseMarkdown', edge_type:'calls', resolved:true}]}`.";
|
|
|
|
export const CODE_CALLEES_DESCRIPTION =
|
|
"When tracing how a function flows to its dependencies (DB calls, HTTP calls, " +
|
|
"file I/O), run code_callees from the entry point. Forward view of the call " +
|
|
"graph: what does this symbol call? Use this when debugging unexpected behavior " +
|
|
"or when planning to extract / inline a function. Same shape as code_callers " +
|
|
"but the field is `callees` and the edge direction is reversed.";
|
|
|
|
export const CODE_DEF_DESCRIPTION =
|
|
"Where is this symbol defined? Returns one row per definition site (function, " +
|
|
"class, type, interface, enum, struct, trait, module, contract). Use this BEFORE " +
|
|
"reaching for grep when you want to read a definition. Single-result is the common " +
|
|
"case; multiple results indicate same-name symbols across files (which is information " +
|
|
"in itself). " +
|
|
"Returns: `{symbol, count, defs: [{slug, file, language, symbol_type, start_line, end_line, snippet}]}`. " +
|
|
"Filter by --lang to scope a polyglot brain (e.g., lang='typescript').";
|
|
|
|
export const CODE_REFS_DESCRIPTION =
|
|
"Find every reference to a symbol across the codebase (every file, every line). " +
|
|
"Differs from code_callers in two ways: (1) catches references in comments, " +
|
|
"strings, imports, type annotations — not just call sites; (2) returns line " +
|
|
"numbers, not symbol-qualified edges. Use this when planning a rename or " +
|
|
"deprecation where you need to touch every literal mention. " +
|
|
"Returns: `{symbol, count, refs: [{slug, file, language, line, context}]}`.";
|