mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-29 19:01:39 +00:00
* feat(facts): typed-claim substrate + cycle correctness fixes (v0.35.6 wave 1/3) Schema (migration v67): - Add four optional typed-claim columns to facts: claim_metric TEXT, claim_value DOUBLE PRECISION, claim_unit TEXT, claim_period TEXT - Partial index facts_typed_claim_idx ON (entity_slug, claim_metric, valid_from) WHERE claim_metric IS NOT NULL - All nullable, metadata-only on both engines Fence layer: - ParsedFact (facts-fence.ts) gains optional claimMetric/Value/Unit/Period - Parser tolerates both 10-cell (legacy) and 14-cell (widened) rows - Renderer emits 14 cells iff any row has typed data; otherwise stays 10-cell so existing fences don't widen on unrelated edits - Numeric value cell tolerates comma thousand separators (50,000 -> 50000) Extract pipeline (D-CDX-2, D-ENG-1): - src/core/facts/extract.ts (the actual Haiku call site, NOT extract-facts.ts cycle phase) extends its system prompt to emit typed fields for metric-shaped claims - extractFactsFromFenceText gains optional pageEffectiveDate. Precedence: fence-row validFrom > pageEffectiveDate > undefined (engine defaults to now) - normalizeMetricLabel: 15-entry seed map for common founder metrics (mrr, arr, runway, headcount, team_size, cac, ltv, gross_margin, burn_rate, cash, users, mau, dau, churn_rate, revenue); unknown labels lowercase + space->_ Engine extensions: - NewFact + insertFact + insertFacts in both engines accept the four typed columns (all nullable) - Cycle phase extract-facts.ts threads page.effective_date through AND batch-embeds via gateway.embed() before insertFacts (D-CDX-3 fix for cycle-inserted facts arriving with embedding=NULL) Consolidate fix (D-CDX-4 — Codex F4): - Replace MAX(row_num)+1 INSERT with semantic upsert on (page_id, claim, since_date). Re-running the full cycle on stable input produces zero new takes — fixes the pre-existing duplicate-takes bug after extract_facts wipes consolidated_at - Chronological valid_until writeback per cluster: sort by (valid_from ASC, id ASC), walk pairs, set older.valid_until = newer.valid_from Tests: - test/migrate.test.ts +6 cases for v67 shape + materialization + nullable backward compat - test/facts-fence-typed.test.ts (new, 17 cases): parser+renderer round-trip, normalization seed map coverage, valid_from precedence three-branch - test/consolidate-valid-until.test.ts (new, 4 cases): chronological writeback (R4a), same-day id tiebreaker, cycle re-run zero duplicates (R4b/R7), valid_until idempotency - test/schema-bootstrap-coverage.test.ts: add four typed-claim columns to COLUMN_EXEMPTIONS (migration co-defines the partial index, no forward reference to bootstrap) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(trajectory): find_trajectory MCP op + eval/founder CLIs (v0.35.6 wave 2/3) Engine method (D-CDX-1, D-CDX-6): - BrainEngine.findTrajectory(opts) on both Postgres and PGLite - TrajectoryOpts: scalar sourceId fast path + sourceIds federated array (mirrors v0.34.1.0 search* dual pattern) - opts.remote: when true, SQL adds AND visibility='world' so OAuth read clients see only world-visibility facts (mirrors recall's posture — closes the F7 privacy regression Codex caught in plan review) - Single SQL query, ORDER BY valid_from ASC, id ASC for deterministic output (R3 pin). Returns TrajectoryPoint[] including raw embedding so the caller can compute drift without a second round-trip Pure function library (src/core/trajectory.ts, new): - detectRegressions(points, threshold): walks consecutive (metric, value) pairs per metric; emits when newer drops >= threshold below older. 10% default, override via GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD - computeDriftScore(points): 1 - mean(cosine(emb[i], emb[i-1])) over embedded points; clamped [0,1]; null when <3 embedded points (D-ENG-3 graceful degradation) - computeTrajectoryStats(points): composed shape returning both - TRAJECTORY_SCHEMA_VERSION = 1 — additive-only across releases (R5) MCP op (src/core/operations.ts): - find_trajectory: scope read, NOT localOnly. Routes through sourceScopeOpts(ctx) for federated isolation AND threads ctx.remote for visibility filtering. Strips raw Float32Array embeddings from the wire shape; converts valid_from to YYYY-MM-DD string - Registered in operations array after find_experts - FIND_TRAJECTORY_DESCRIPTION in operations-descriptions.ts CLIs: - gbrain eval trajectory <entity> [--metric M] [--since D] [--until D] [--limit N] [--json] — chronological human view with [REGRESSION] inline annotation; thin-client routing via callRemoteTool(find_trajectory). Dispatched in src/commands/eval.ts sub-subcommand block - gbrain founder scorecard <entity> [--since D] [--until D] [--json] — pure aggregation over Phase 2's substrate. Four signals: claim_accuracy (over resolved takes), consistency, growth_trajectory, red_flags. computeFounderScorecard exported for tests. Registered as top-level command in cli.ts; added to CLI_ONLY set Tests (45 cases across 5 files): - test/engine-find-trajectory.test.ts: 18 cases — chronological order, source scoping (scalar + federated), visibility filter on remote=true, metric + since/until filters, regression detection at threshold boundaries, drift score with various embedding states - test/operations-find-trajectory.test.ts: 9 cases — op registration, param validation, JSON envelope shape, R5 schema_version: 1, embedding stripped from wire, R6 visibility filter, source scoping - test/eval-trajectory.test.ts: 7 cases — arg parsing, --help, --json envelope, regression annotation, --metric filter, empty entity - test/founder-scorecard.test.ts: 9 cases — empty inputs no-NaN (G2), claim_accuracy math, consistency math, growth_trajectory math, red_flags fire for regression / narrative_drift / missed_prediction - test/eval-contradictions/no-valid-until-write.test.ts: 4 cases — R1 (probe never writes valid_until under eval-contradictions/) + R8 (only allow-listed files write valid_until anywhere in src/) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: v0.35.6.0 — CHANGELOG + VERSION + docs + migration note Bumps to v0.35.6.0 (next-minor after master's v0.35.5.1 — typed-claim substrate + trajectory + founder scorecard is a new user-facing feature surface, not a fix). - VERSION + package.json synced - CHANGELOG.md release-summary block in the wave-style voice, lead with what the user can now DO. Sections: typed metric claims in the fence, chronological metric trajectories, founder scorecard, MCP find_trajectory op, cycle re-run idempotency fix, embedding-on-insert fix, valid_from precedence fix. To-take-advantage-of block with verification + opt-in fence syntax example - CLAUDE.md Key Files entry consolidating the wave across eval-trajectory.ts + founder-scorecard.ts + trajectory.ts. Names every D-ENG / D-CDX decision and the Codex outside-voice F-numbers - skills/migrations/v0.35.6.md agent-readable migration note. Includes fence-syntax example for typed-claim rows so downstream agents start emitting them. Iron-rule contracts called out (R1 + R8 + R7 + visibility) - llms-full.txt regenerated to reflect the new CLAUDE.md entry Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: post-ship sync for v0.35.7.0 — trajectory + founder scorecard - README.md: add `gbrain eval trajectory` to EVAL section, add new TEMPORAL block covering `gbrain founder scorecard` + the GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD env override; add v0.35.7 "What's new" paragraph below the v0.28.8 LongMemEval blurb - AGENTS.md: new bullet under Common tasks teaching agents to reach for `gbrain eval trajectory` / `gbrain founder scorecard` / the `find_trajectory` MCP op when asked to evaluate a founder/company over time - docs/contradictions.md: append "Temporal axis follow-on (v0.35.3.1 + v0.35.7)" subsection under See also, cross-linking the trajectory substrate and naming the auto-supersession.ts:4 invariant preserved by both the verdict enum (probe side) and consolidate's valid_until writeback (cycle side) - CLAUDE.md: fix stale (v0.35.4) tag on the trajectory entry to (v0.35.7) — version got rebumped twice during the merge wave - skills/migrations/v0.35.7.md renamed to v0.35.7.0.md for consistency with the v0.35.0.0.md / v0.14.0.md / etc naming convention - llms-full.txt regenerated to reflect the CLAUDE.md edit Coverage map (Diataxis): /eval trajectory CLI ✅ ref (README, AGENTS) ✅ how-to (CHANGELOG) ❌ tutorial /founder scorecard CLI ✅ ref (README, AGENTS) ✅ how-to (CHANGELOG) ❌ tutorial find_trajectory MCP op ✅ ref (CLAUDE.md, AGENTS, contradictions.md) typed-claim fence cols ✅ ref (skills/migrations/v0.35.7.0.md, CHANGELOG) Migration v67 ✅ ref (CLAUDE.md, CHANGELOG) No tutorial / explanation gaps worth filling in this PR — the migration note's fence-syntax example already covers the "first typed claim" walkthrough. ARCHITECTURE diagrams not drifted (the trajectory work extends existing facts/takes infrastructure; no new component boxes). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
171 lines
5.7 KiB
TypeScript
171 lines
5.7 KiB
TypeScript
/**
|
|
* v0.35.4 — trajectory derived metrics.
|
|
*
|
|
* Pure functions over `TrajectoryPoint[]` (engine output). Both the
|
|
* `find_trajectory` MCP op and the `gbrain eval trajectory` CLI consume
|
|
* this module so the regression + drift_score definitions stay in one
|
|
* place.
|
|
*
|
|
* Locked specs from the plan:
|
|
*
|
|
* - Regression detection (D-ENG-2): a regression fires for every
|
|
* consecutive (metric, value) pair where the newer value is at least
|
|
* 10% lower than the prior value. Threshold is configurable via the
|
|
* env var `GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD` (default 0.10).
|
|
* Only points with `claim_value !== null` participate.
|
|
*
|
|
* - Drift score (D-ENG-3): `1 - mean(cosine(emb[i], emb[i-1]))` over
|
|
* points with non-null embeddings. Clamped to [0, 1]. Returns null
|
|
* when fewer than 3 points have embeddings (graceful degradation
|
|
* for pre-v0.35.4 facts that arrived without one).
|
|
*/
|
|
|
|
import type { TrajectoryPoint } from './engine.ts';
|
|
|
|
/** Default regression threshold (10% drop). Locked decision D-ENG-2. */
|
|
export const DEFAULT_REGRESSION_THRESHOLD = 0.10;
|
|
|
|
/** Schema version for the trajectory + scorecard JSON contract. Additive-only across releases. */
|
|
export const TRAJECTORY_SCHEMA_VERSION = 1;
|
|
|
|
export interface TrajectoryRegression {
|
|
metric: string;
|
|
from_value: number;
|
|
from_date: string; // YYYY-MM-DD
|
|
to_value: number;
|
|
to_date: string;
|
|
delta_pct: number; // negative for a drop; range typically [-1, 0)
|
|
}
|
|
|
|
export interface TrajectoryStats {
|
|
regressions: TrajectoryRegression[];
|
|
drift_score: number | null;
|
|
}
|
|
|
|
/**
|
|
* Read the regression threshold from `GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD`
|
|
* with fallback to the locked default. Invalid input falls back silently —
|
|
* the threshold is a soft tuning knob, not a correctness gate.
|
|
*/
|
|
export function resolveRegressionThreshold(): number {
|
|
const raw = process.env.GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD;
|
|
if (!raw) return DEFAULT_REGRESSION_THRESHOLD;
|
|
const n = parseFloat(raw);
|
|
if (!Number.isFinite(n) || n <= 0 || n >= 1) return DEFAULT_REGRESSION_THRESHOLD;
|
|
return n;
|
|
}
|
|
|
|
function toISODate(d: Date): string {
|
|
return d.toISOString().slice(0, 10);
|
|
}
|
|
|
|
/**
|
|
* Compute cosine similarity between two equal-length vectors. Returns 0
|
|
* when either vector has length zero (defensive — never throws).
|
|
*/
|
|
function cosineSim(a: Float32Array, b: Float32Array): number {
|
|
if (a.length === 0 || b.length === 0 || a.length !== b.length) return 0;
|
|
let dot = 0;
|
|
let na = 0;
|
|
let nb = 0;
|
|
for (let i = 0; i < a.length; i++) {
|
|
dot += a[i] * b[i];
|
|
na += a[i] * a[i];
|
|
nb += b[i] * b[i];
|
|
}
|
|
if (na === 0 || nb === 0) return 0;
|
|
return dot / (Math.sqrt(na) * Math.sqrt(nb));
|
|
}
|
|
|
|
/**
|
|
* Detect chronological regressions in a sorted trajectory.
|
|
*
|
|
* Iterates per-metric (so trajectories that interleave mrr + arr + team_size
|
|
* don't trip false regressions across metric boundaries). Within each metric,
|
|
* walks consecutive value pairs; a pair fires when
|
|
* `(newer - older) / older <= -threshold`.
|
|
*
|
|
* Pre-condition: caller passed points sorted by (valid_from ASC, fact_id ASC).
|
|
* The engine's `findTrajectory` enforces this. No re-sort here.
|
|
*/
|
|
export function detectRegressions(
|
|
points: TrajectoryPoint[],
|
|
threshold: number = DEFAULT_REGRESSION_THRESHOLD,
|
|
): TrajectoryRegression[] {
|
|
const out: TrajectoryRegression[] = [];
|
|
// Group by metric so each metric's regression detection is independent.
|
|
const byMetric = new Map<string, TrajectoryPoint[]>();
|
|
for (const p of points) {
|
|
if (p.metric === null || p.value === null) continue;
|
|
if (!Number.isFinite(p.value)) continue;
|
|
if (!byMetric.has(p.metric)) byMetric.set(p.metric, []);
|
|
byMetric.get(p.metric)!.push(p);
|
|
}
|
|
|
|
for (const [metric, series] of byMetric) {
|
|
for (let i = 1; i < series.length; i++) {
|
|
const older = series[i - 1];
|
|
const newer = series[i];
|
|
const oldVal = older.value!;
|
|
const newVal = newer.value!;
|
|
// Guard against division-by-zero: a metric starting at exactly 0
|
|
// can't compute a relative delta. Skip.
|
|
if (oldVal === 0) continue;
|
|
const delta = (newVal - oldVal) / oldVal;
|
|
if (delta <= -threshold) {
|
|
out.push({
|
|
metric,
|
|
from_value: oldVal,
|
|
from_date: toISODate(older.valid_from),
|
|
to_value: newVal,
|
|
to_date: toISODate(newer.valid_from),
|
|
delta_pct: delta,
|
|
});
|
|
}
|
|
}
|
|
}
|
|
return out;
|
|
}
|
|
|
|
/**
|
|
* Compute drift score over the trajectory's existing embeddings.
|
|
*
|
|
* `1 - mean(cosine(emb[i], emb[i-1]))` clamped to [0, 1]. Range
|
|
* interpretation: 0 = narrative stable text-wise; 1 = every consecutive
|
|
* claim is unrelated to the prior.
|
|
*
|
|
* Returns null when fewer than 3 points have non-null embeddings — the
|
|
* statistic is meaningless on tiny samples.
|
|
*/
|
|
export function computeDriftScore(points: TrajectoryPoint[]): number | null {
|
|
const withEmb = points.filter(p => p.embedding !== null && p.embedding.length > 0);
|
|
if (withEmb.length < 3) return null;
|
|
let sumCos = 0;
|
|
let pairs = 0;
|
|
for (let i = 1; i < withEmb.length; i++) {
|
|
sumCos += cosineSim(withEmb[i - 1].embedding!, withEmb[i].embedding!);
|
|
pairs += 1;
|
|
}
|
|
if (pairs === 0) return null;
|
|
const meanCos = sumCos / pairs;
|
|
const drift = 1 - meanCos;
|
|
if (drift < 0) return 0;
|
|
if (drift > 1) return 1;
|
|
return drift;
|
|
}
|
|
|
|
/**
|
|
* Compose the two derived metrics into a single TrajectoryStats. The MCP
|
|
* op + CLI both call this so the JSON shape stays consistent.
|
|
*/
|
|
export function computeTrajectoryStats(
|
|
points: TrajectoryPoint[],
|
|
opts: { threshold?: number } = {},
|
|
): TrajectoryStats {
|
|
const threshold = opts.threshold ?? resolveRegressionThreshold();
|
|
return {
|
|
regressions: detectRegressions(points, threshold),
|
|
drift_score: computeDriftScore(points),
|
|
};
|
|
}
|