Files
gbrain/test/facts-fence-typed.test.ts
T
1dadd9ed71 v0.35.7.0 feat: temporal trajectory + founder scorecard (Phases 2-4) (#1131)
* feat(facts): typed-claim substrate + cycle correctness fixes (v0.35.6 wave 1/3)

Schema (migration v67):
- Add four optional typed-claim columns to facts: claim_metric TEXT,
  claim_value DOUBLE PRECISION, claim_unit TEXT, claim_period TEXT
- Partial index facts_typed_claim_idx ON (entity_slug, claim_metric, valid_from)
  WHERE claim_metric IS NOT NULL
- All nullable, metadata-only on both engines

Fence layer:
- ParsedFact (facts-fence.ts) gains optional claimMetric/Value/Unit/Period
- Parser tolerates both 10-cell (legacy) and 14-cell (widened) rows
- Renderer emits 14 cells iff any row has typed data; otherwise stays
  10-cell so existing fences don't widen on unrelated edits
- Numeric value cell tolerates comma thousand separators (50,000 -> 50000)

Extract pipeline (D-CDX-2, D-ENG-1):
- src/core/facts/extract.ts (the actual Haiku call site, NOT extract-facts.ts
  cycle phase) extends its system prompt to emit typed fields for metric-shaped
  claims
- extractFactsFromFenceText gains optional pageEffectiveDate. Precedence:
  fence-row validFrom > pageEffectiveDate > undefined (engine defaults to now)
- normalizeMetricLabel: 15-entry seed map for common founder metrics (mrr,
  arr, runway, headcount, team_size, cac, ltv, gross_margin, burn_rate, cash,
  users, mau, dau, churn_rate, revenue); unknown labels lowercase + space->_

Engine extensions:
- NewFact + insertFact + insertFacts in both engines accept the four typed
  columns (all nullable)
- Cycle phase extract-facts.ts threads page.effective_date through AND
  batch-embeds via gateway.embed() before insertFacts (D-CDX-3 fix for
  cycle-inserted facts arriving with embedding=NULL)

Consolidate fix (D-CDX-4 — Codex F4):
- Replace MAX(row_num)+1 INSERT with semantic upsert on (page_id, claim,
  since_date). Re-running the full cycle on stable input produces zero new
  takes — fixes the pre-existing duplicate-takes bug after extract_facts
  wipes consolidated_at
- Chronological valid_until writeback per cluster: sort by (valid_from ASC,
  id ASC), walk pairs, set older.valid_until = newer.valid_from

Tests:
- test/migrate.test.ts +6 cases for v67 shape + materialization + nullable
  backward compat
- test/facts-fence-typed.test.ts (new, 17 cases): parser+renderer round-trip,
  normalization seed map coverage, valid_from precedence three-branch
- test/consolidate-valid-until.test.ts (new, 4 cases): chronological
  writeback (R4a), same-day id tiebreaker, cycle re-run zero duplicates
  (R4b/R7), valid_until idempotency
- test/schema-bootstrap-coverage.test.ts: add four typed-claim columns to
  COLUMN_EXEMPTIONS (migration co-defines the partial index, no forward
  reference to bootstrap)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(trajectory): find_trajectory MCP op + eval/founder CLIs (v0.35.6 wave 2/3)

Engine method (D-CDX-1, D-CDX-6):
- BrainEngine.findTrajectory(opts) on both Postgres and PGLite
- TrajectoryOpts: scalar sourceId fast path + sourceIds federated array
  (mirrors v0.34.1.0 search* dual pattern)
- opts.remote: when true, SQL adds AND visibility='world' so OAuth read
  clients see only world-visibility facts (mirrors recall's posture —
  closes the F7 privacy regression Codex caught in plan review)
- Single SQL query, ORDER BY valid_from ASC, id ASC for deterministic
  output (R3 pin). Returns TrajectoryPoint[] including raw embedding so
  the caller can compute drift without a second round-trip

Pure function library (src/core/trajectory.ts, new):
- detectRegressions(points, threshold): walks consecutive (metric, value)
  pairs per metric; emits when newer drops >= threshold below older.
  10% default, override via GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD
- computeDriftScore(points): 1 - mean(cosine(emb[i], emb[i-1])) over
  embedded points; clamped [0,1]; null when <3 embedded points (D-ENG-3
  graceful degradation)
- computeTrajectoryStats(points): composed shape returning both
- TRAJECTORY_SCHEMA_VERSION = 1 — additive-only across releases (R5)

MCP op (src/core/operations.ts):
- find_trajectory: scope read, NOT localOnly. Routes through
  sourceScopeOpts(ctx) for federated isolation AND threads ctx.remote
  for visibility filtering. Strips raw Float32Array embeddings from the
  wire shape; converts valid_from to YYYY-MM-DD string
- Registered in operations array after find_experts
- FIND_TRAJECTORY_DESCRIPTION in operations-descriptions.ts

CLIs:
- gbrain eval trajectory <entity> [--metric M] [--since D] [--until D]
  [--limit N] [--json] — chronological human view with [REGRESSION] inline
  annotation; thin-client routing via callRemoteTool(find_trajectory).
  Dispatched in src/commands/eval.ts sub-subcommand block
- gbrain founder scorecard <entity> [--since D] [--until D] [--json] —
  pure aggregation over Phase 2's substrate. Four signals:
  claim_accuracy (over resolved takes), consistency, growth_trajectory,
  red_flags. computeFounderScorecard exported for tests.
  Registered as top-level command in cli.ts; added to CLI_ONLY set

Tests (45 cases across 5 files):
- test/engine-find-trajectory.test.ts: 18 cases — chronological order,
  source scoping (scalar + federated), visibility filter on remote=true,
  metric + since/until filters, regression detection at threshold
  boundaries, drift score with various embedding states
- test/operations-find-trajectory.test.ts: 9 cases — op registration,
  param validation, JSON envelope shape, R5 schema_version: 1,
  embedding stripped from wire, R6 visibility filter, source scoping
- test/eval-trajectory.test.ts: 7 cases — arg parsing, --help,
  --json envelope, regression annotation, --metric filter, empty entity
- test/founder-scorecard.test.ts: 9 cases — empty inputs no-NaN (G2),
  claim_accuracy math, consistency math, growth_trajectory math,
  red_flags fire for regression / narrative_drift / missed_prediction
- test/eval-contradictions/no-valid-until-write.test.ts: 4 cases —
  R1 (probe never writes valid_until under eval-contradictions/) +
  R8 (only allow-listed files write valid_until anywhere in src/)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: v0.35.6.0 — CHANGELOG + VERSION + docs + migration note

Bumps to v0.35.6.0 (next-minor after master's v0.35.5.1 — typed-claim
substrate + trajectory + founder scorecard is a new user-facing
feature surface, not a fix).

- VERSION + package.json synced
- CHANGELOG.md release-summary block in the wave-style voice, lead with
  what the user can now DO. Sections: typed metric claims in the fence,
  chronological metric trajectories, founder scorecard, MCP
  find_trajectory op, cycle re-run idempotency fix, embedding-on-insert
  fix, valid_from precedence fix. To-take-advantage-of block with
  verification + opt-in fence syntax example
- CLAUDE.md Key Files entry consolidating the wave across
  eval-trajectory.ts + founder-scorecard.ts + trajectory.ts. Names every
  D-ENG / D-CDX decision and the Codex outside-voice F-numbers
- skills/migrations/v0.35.6.md agent-readable migration note. Includes
  fence-syntax example for typed-claim rows so downstream agents start
  emitting them. Iron-rule contracts called out (R1 + R8 + R7 + visibility)
- llms-full.txt regenerated to reflect the new CLAUDE.md entry

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: post-ship sync for v0.35.7.0 — trajectory + founder scorecard

- README.md: add `gbrain eval trajectory` to EVAL section, add new
  TEMPORAL block covering `gbrain founder scorecard` + the
  GBRAIN_TRAJECTORY_REGRESSION_THRESHOLD env override; add v0.35.7
  "What's new" paragraph below the v0.28.8 LongMemEval blurb
- AGENTS.md: new bullet under Common tasks teaching agents to reach for
  `gbrain eval trajectory` / `gbrain founder scorecard` / the
  `find_trajectory` MCP op when asked to evaluate a founder/company
  over time
- docs/contradictions.md: append "Temporal axis follow-on (v0.35.3.1 +
  v0.35.7)" subsection under See also, cross-linking the trajectory
  substrate and naming the auto-supersession.ts:4 invariant preserved
  by both the verdict enum (probe side) and consolidate's valid_until
  writeback (cycle side)
- CLAUDE.md: fix stale (v0.35.4) tag on the trajectory entry to
  (v0.35.7) — version got rebumped twice during the merge wave
- skills/migrations/v0.35.7.md renamed to v0.35.7.0.md for consistency
  with the v0.35.0.0.md / v0.14.0.md / etc naming convention
- llms-full.txt regenerated to reflect the CLAUDE.md edit

Coverage map (Diataxis):
  /eval trajectory CLI        ref (README, AGENTS)  how-to (CHANGELOG)  tutorial
  /founder scorecard CLI      ref (README, AGENTS)  how-to (CHANGELOG)  tutorial
  find_trajectory MCP op      ref (CLAUDE.md, AGENTS, contradictions.md)
  typed-claim fence cols      ref (skills/migrations/v0.35.7.0.md, CHANGELOG)
  Migration v67               ref (CLAUDE.md, CHANGELOG)

No tutorial / explanation gaps worth filling in this PR — the migration
note's fence-syntax example already covers the "first typed claim"
walkthrough. ARCHITECTURE diagrams not drifted (the trajectory work
extends existing facts/takes infrastructure; no new component boxes).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 18:52:38 -07:00

322 lines
13 KiB
TypeScript

// v0.35.4 — typed-claim fence parser+renderer round-trip + normalization.
//
// Pins:
// 1. Backward compat — a fence authored without typed fields still
// parses, renders as 10-cell shape, and round-trips byte-identical.
// 2. Typed fields parse from the 14-cell widened fence.
// 3. Renderer widens to 14 cells iff ANY row has a non-undefined typed
// field; otherwise stays at 10-cell (no diff noise on existing fences).
// 4. Round-trip preservation: parse → render → parse produces the same
// ParsedFact array, including typed fields.
// 5. Numeric value cell tolerates thousand separators (`50,000`).
import { test, expect, describe } from 'bun:test';
import {
parseFactsFence,
renderFactsTable,
upsertFactRow,
type ParsedFact,
} from '../src/core/facts-fence.ts';
import {
extractFactsFromFenceText,
normalizeMetricLabel,
METRIC_NORMALIZATION_MAP,
} from '../src/core/facts/extract-from-fence.ts';
function wrap(inner: string): string {
return `## Facts\n\n<!--- gbrain:facts:begin -->\n${inner}\n<!--- gbrain:facts:end -->\n`;
}
describe('v0.35.4 — facts fence typed-claim parser', () => {
test('legacy 10-cell fence parses as before; all typed fields undefined', () => {
const body = wrap(
`| # | claim | kind | confidence | visibility | notability | valid_from | valid_until | source | context |
|---|-------|------|------------|------------|------------|------------|-------------|--------|---------|
| 1 | Founded Acme in 2017 | fact | 1.0 | world | high | 2017-01-01 | | linkedin | |`,
);
const { facts, warnings } = parseFactsFence(body);
expect(warnings).toEqual([]);
expect(facts.length).toBe(1);
expect(facts[0].claim).toBe('Founded Acme in 2017');
expect(facts[0].claimMetric).toBeUndefined();
expect(facts[0].claimValue).toBeUndefined();
expect(facts[0].claimUnit).toBeUndefined();
expect(facts[0].claimPeriod).toBeUndefined();
});
test('14-cell typed fence parses all four typed-claim fields', () => {
const body = wrap(
`| # | claim | kind | confidence | visibility | notability | valid_from | valid_until | source | context | claim_metric | claim_value | claim_unit | claim_period |
|---|-------|------|------------|------------|------------|------------|-------------|--------|---------|--------------|-------------|------------|--------------|
| 1 | MRR hit $50K | fact | 1.0 | private | high | 2026-01-15 | | OH transcript | | mrr | 50000 | USD | monthly |`,
);
const { facts, warnings } = parseFactsFence(body);
expect(warnings).toEqual([]);
expect(facts.length).toBe(1);
expect(facts[0].claimMetric).toBe('mrr');
expect(facts[0].claimValue).toBe(50000);
expect(facts[0].claimUnit).toBe('USD');
expect(facts[0].claimPeriod).toBe('monthly');
});
test('numeric value cell tolerates comma thousand separators', () => {
const body = wrap(
`| # | claim | kind | confidence | visibility | notability | valid_from | valid_until | source | context | claim_metric | claim_value | claim_unit | claim_period |
|---|-------|------|------------|------------|------------|------------|-------------|--------|---------|--------------|-------------|------------|--------------|
| 1 | ARR | fact | 1.0 | private | high | 2026-04-12 | | bo call | | arr | 2,000,000 | USD | annual |`,
);
const { facts } = parseFactsFence(body);
expect(facts[0].claimValue).toBe(2000000);
});
test('empty typed cells parse as undefined (not "")', () => {
const body = wrap(
`| # | claim | kind | confidence | visibility | notability | valid_from | valid_until | source | context | claim_metric | claim_value | claim_unit | claim_period |
|---|-------|------|------------|------------|------------|------------|-------------|--------|---------|--------------|-------------|------------|--------------|
| 1 | bare claim | fact | 1.0 | private | high | 2026-01-01 | | | | | | | |`,
);
const { facts } = parseFactsFence(body);
expect(facts[0].claimMetric).toBeUndefined();
expect(facts[0].claimValue).toBeUndefined();
expect(facts[0].claimUnit).toBeUndefined();
expect(facts[0].claimPeriod).toBeUndefined();
});
});
describe('v0.35.4 — facts fence typed-claim renderer', () => {
test('renders 10-cell shape when no row has typed fields (backward compat)', () => {
const facts: ParsedFact[] = [
{
rowNum: 1,
claim: 'Founded Acme in 2017',
kind: 'fact',
confidence: 1.0,
visibility: 'world',
notability: 'high',
validFrom: '2017-01-01',
source: 'linkedin',
active: true,
},
];
const out = renderFactsTable(facts);
// 10-cell header
expect(out).toContain('| # | claim | kind | confidence | visibility | notability | valid_from | valid_until | source | context |');
// NOT the 14-cell variant
expect(out).not.toContain('claim_metric');
});
test('widens to 14 cells when ANY row has typed fields', () => {
const facts: ParsedFact[] = [
{
rowNum: 1,
claim: 'plain claim',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'medium',
active: true,
},
{
rowNum: 2,
claim: 'MRR hit $50K',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'high',
active: true,
claimMetric: 'mrr',
claimValue: 50000,
claimUnit: 'USD',
claimPeriod: 'monthly',
},
];
const out = renderFactsTable(facts);
expect(out).toContain('claim_metric');
expect(out).toContain('claim_value');
expect(out).toContain('claim_unit');
expect(out).toContain('claim_period');
// Row 1 has empty typed cells; row 2 has the values.
expect(out).toContain('| mrr | 50000 | USD | monthly |');
});
test('round-trip preservation: parse → render → parse is structurally idempotent for typed facts', () => {
const body = wrap(
`| # | claim | kind | confidence | visibility | notability | valid_from | valid_until | source | context | claim_metric | claim_value | claim_unit | claim_period |
|---|-------|------|------------|------------|------------|------------|-------------|--------|---------|--------------|-------------|------------|--------------|
| 1 | MRR $50K | fact | 1.0 | private | high | 2026-01-15 | | OH | | mrr | 50000 | USD | monthly |
| 2 | Team grew to 12 | fact | 1.0 | private | medium | 2026-02-01 | | meeting | | team_size | 12 | people | |
| 3 | Plain non-typed claim | fact | 0.85 | private | low | 2026-03-01 | | inferred | | | | | |`,
);
const first = parseFactsFence(body);
expect(first.warnings).toEqual([]);
expect(first.facts.length).toBe(3);
const rendered = renderFactsTable(first.facts);
const second = parseFactsFence(rendered);
expect(second.warnings).toEqual([]);
expect(second.facts).toEqual(first.facts);
});
test('upsertFactRow threads typed fields when a new row carries them', () => {
// Start from a fence with NO typed fields → 10-cell shape.
const body = wrap(
`| # | claim | kind | confidence | visibility | notability | valid_from | valid_until | source | context |
|---|-------|------|------------|------------|------------|------------|-------------|--------|---------|
| 1 | Founded Acme in 2017 | fact | 1.0 | world | high | 2017-01-01 | | linkedin | |`,
);
const { body: newBody, rowNum } = upsertFactRow(body, {
claim: 'MRR hit $50K',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'high',
validFrom: '2026-01-15',
source: 'OH transcript',
claimMetric: 'mrr',
claimValue: 50000,
claimUnit: 'USD',
claimPeriod: 'monthly',
});
expect(rowNum).toBe(2);
// Adding a typed row widens the table.
expect(newBody).toContain('claim_metric');
const parsed = parseFactsFence(newBody);
expect(parsed.warnings).toEqual([]);
expect(parsed.facts.length).toBe(2);
expect(parsed.facts[1].claimMetric).toBe('mrr');
expect(parsed.facts[1].claimValue).toBe(50000);
});
});
describe('v0.35.4 — metric normalization (D-ENG-4)', () => {
test('known seed-map labels normalize to canonical snake_case', () => {
expect(normalizeMetricLabel('MRR')).toBe('mrr');
expect(normalizeMetricLabel('Monthly Recurring Revenue')).toBe('mrr');
expect(normalizeMetricLabel('ARR')).toBe('arr');
expect(normalizeMetricLabel('annual recurring revenue')).toBe('arr');
expect(normalizeMetricLabel('Team Size')).toBe('team_size');
expect(normalizeMetricLabel('Burn Rate')).toBe('burn_rate');
expect(normalizeMetricLabel('Churn')).toBe('churn_rate');
});
test('unknown labels lowercase + spaces → underscores; non-alphanumeric stripped', () => {
expect(normalizeMetricLabel('Net Promoter Score')).toBe('net_promoter_score');
expect(normalizeMetricLabel(' CAC ')).toBe('cac');
expect(normalizeMetricLabel('Time-to-Hire')).toBe('timetohire');
});
test('empty / null / undefined → undefined (the "no metric set" signal)', () => {
expect(normalizeMetricLabel(undefined)).toBeUndefined();
expect(normalizeMetricLabel(null)).toBeUndefined();
expect(normalizeMetricLabel('')).toBeUndefined();
expect(normalizeMetricLabel(' ')).toBeUndefined();
});
test('METRIC_NORMALIZATION_MAP covers the 15-metric seed list named in the plan', () => {
// Pin the seed map so the docs in CLAUDE.md / CHANGELOG match.
const required = [
'mrr', 'arr', 'runway', 'headcount', 'team_size',
'cac', 'ltv', 'gross_margin', 'burn_rate', 'cash',
'users', 'mau', 'dau', 'churn_rate', 'revenue',
];
const canonicalValues = new Set(METRIC_NORMALIZATION_MAP.values());
for (const r of required) {
expect(canonicalValues.has(r)).toBe(true);
}
});
});
describe('v0.35.4 (D-ENG-1) — extractFactsFromFenceText valid_from precedence', () => {
const fixedToday = new Date('2026-05-17T00:00:00.000Z');
test('Path 1: explicit validFrom in fence row wins', () => {
const facts: ParsedFact[] = [{
rowNum: 1,
claim: 'fence-dated',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'medium',
validFrom: '2026-01-15',
active: true,
}];
const out = extractFactsFromFenceText(facts, 'people/alice-example', 'default', {
nowOverride: fixedToday,
pageEffectiveDate: new Date('2026-04-28'),
});
expect(out[0].valid_from?.toISOString().slice(0, 10)).toBe('2026-01-15');
});
test('Path 2: missing fence validFrom + pageEffectiveDate set → uses page date', () => {
const facts: ParsedFact[] = [{
rowNum: 1,
claim: 'no fence date',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'medium',
active: true,
}];
const out = extractFactsFromFenceText(facts, 'people/alice-example', 'default', {
nowOverride: fixedToday,
pageEffectiveDate: new Date('2026-04-28'),
});
expect(out[0].valid_from?.toISOString().slice(0, 10)).toBe('2026-04-28');
});
test('Path 3: missing fence validFrom AND undefined pageEffectiveDate → undefined (engine defaults to now)', () => {
const facts: ParsedFact[] = [{
rowNum: 1,
claim: 'no dates at all',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'medium',
active: true,
}];
const out = extractFactsFromFenceText(facts, 'people/alice-example', 'default', {
nowOverride: fixedToday,
pageEffectiveDate: null,
});
// valid_from is left undefined; the engine layer applies now() at insert.
expect(out[0].valid_from).toBeUndefined();
});
test('typed-claim fields thread through with metric normalization applied', () => {
const facts: ParsedFact[] = [{
rowNum: 1,
claim: 'MRR $50K',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'high',
active: true,
claimMetric: 'Monthly Recurring Revenue', // unnormalized
claimValue: 50000,
claimUnit: 'USD',
claimPeriod: 'monthly',
}];
const out = extractFactsFromFenceText(facts, 'companies/acme-example', 'default');
expect(out[0].claim_metric).toBe('mrr'); // normalized
expect(out[0].claim_value).toBe(50000);
expect(out[0].claim_unit).toBe('USD');
expect(out[0].claim_period).toBe('monthly');
});
test('rows with no typed-claim fields land with null claim_* columns (backward compat)', () => {
const facts: ParsedFact[] = [{
rowNum: 1,
claim: 'bare claim',
kind: 'fact',
confidence: 1.0,
visibility: 'private',
notability: 'low',
active: true,
}];
const out = extractFactsFromFenceText(facts, 'people/bob-example', 'default');
expect(out[0].claim_metric).toBeNull();
expect(out[0].claim_value).toBeNull();
expect(out[0].claim_unit).toBeNull();
expect(out[0].claim_period).toBeNull();
});
});