Files
gbrain/test/takes-resolution.test.ts
T
1399e519c0 v0.30.0 feat: calibration scorecards (Slice A1 of v0.30 wave) (#731)
* feat(schema): migration v40 — takes_resolved_quality + drift_decisions

Slice A1 of the v0.30 wave. Bundles all wave schema in one migration so
A2/B1/C1 carry no schema of their own (codex F6 schema-first ordering).

- takes.resolved_quality TEXT with CHECK (correct/incorrect/partial).
- takes_resolution_consistency CHECK enforces (quality, outcome) tuple
  consistency at the DB layer. partial → outcome=NULL.
- One-shot backfill maps legacy resolved_outcome → resolved_quality so
  v0.28 brains keep working with no manual reclassification.
- idx_takes_scorecard partial index on (holder, kind, resolved_quality)
  WHERE resolved_quality IS NOT NULL — scorecard hot path.
- drift_decisions audit table (consumed by Slice C1 in v0.30.3).
- PGLite branch via sqlFor.pglite mirrors the same shape; RLS DO-block
  is Postgres-only.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(takes-fence): extend ParsedTake + parser + conditional renderer (codex F3)

A codex consult on the v0.30 plan caught a real bug: the v0.28 parser had
no concept of resolution columns, so every cmdUpdate after a cmdResolve
silently deleted resolution data on the next render. This commit kills
that data-loss path.

ParsedTake gains optional resolvedAt, resolvedQuality, resolvedOutcome,
resolvedEvidence, resolvedValue, resolvedUnit, resolvedBy. parseTakesFence
detects v0.30-shape headers and reads resolution cells when present;
v0.28 7-column fences round-trip byte-identical. renderTakesFence emits
the resolution columns ONLY when at least one row on the page has
resolvedQuality set — pages with no resolved rows keep the narrow shape
exactly as before.

11 new test cases including the round-trip preservation regression gate.
Without those tests, the silent-delete bug returns the moment the parser
shape drifts. Tests cover: parsing v0.30 + v0.28 shapes, conditional
rendering, partial quality round-trip, upsertTakeRow + supersedeRow
preservation when a page already has resolved rows.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(engine): getScorecard + getCalibrationCurve + 3-state TakeResolution

Adds the calibration aggregate methods on BrainEngine. Both engines
implement them with SQL-level allow-list filtering inside the GROUP BY
(D4 fail-closed): hidden-holder rows contribute zero to aggregates.

TakeResolution gains optional `quality` (correct|incorrect|partial). When
both quality and outcome are supplied AND inconsistent, the engine throws
TAKE_RESOLUTION_INVALID rather than silently overwriting. resolveTake
writes both columns: quality directly, outcome derived (correct→true,
incorrect→false, partial→NULL). Schema CHECK is the defense-in-depth
backstop.

Brier scope (D5 + D11): the SQL aggregation excludes partial rows from
the Brier denominator — partial isn't a binary outcome to compare a
probability against. partial_rate is reported alongside as a separate
counter so hedging behavior stays visible. The 20% threshold lives in
src/core/takes-resolution.ts and the CLI surfaces it in v0.30.0's
cmdScorecard.

New module src/core/takes-resolution.ts holds shared pure helpers
(deriveResolutionTuple, finalizeScorecard) consumed by both engines so
the math stays identical across backends. takeRowToTake (utils.ts) reads
resolved_quality through to the Take row shape.

23 new test cases: 16 for the helpers (Brier hand-calc against a 4-bet
reference at 0.205, n=0 no-divide, contradictory-input rejection,
partial-exclusion contract, threshold constant); 7 against PGLite for
the engine path (3-state quality writes, contradictory throws, scorecard
hand-calc, n=0, SQL-level allow-list privacy filter).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(cli): gbrain takes resolve --quality, takes scorecard, takes calibration

cmdResolve widened: --quality correct|incorrect|partial is the new primary
input. --outcome true|false stays as a back-compat alias auto-mapping to
quality, with a stderr deprecation warning on use. Mutually exclusive
with --quality. --evidence is a semantic alias for --source on the
resolve subcommand.

cmdResolve mirrors resolution metadata into the takes-fence on disk via
the page-lock-aware path. Round-trip preservation through parseTakesFence
+ renderTakesFence keeps resolution data intact across unrelated edits to
other rows on the same page. Removes the v0.28 deferred-rendering warning.

cmdScorecard prints `correct | incorrect | partial`, accuracy, Brier
(correct ∨ incorrect only; lower is better; 0.25 = always-50% baseline),
and partial_rate. When partial_rate > 20% the CLI prints
"[!] partial_rate is high — calibration may be optimistic" so hedging
behavior stays visible even though it doesn't enter the math (D11). Small-N
note when resolved < 100. JSON output via --json.

cmdCalibration bins resolved correct/incorrect bets by stated weight
(--bucket-size, default 0.1) and prints observed vs predicted vs delta
per bucket. Diagonal alignment = perfect calibration.

Both new subcommands wire allow-list as undefined for local CLI callers
(trusted); MCP path will thread it from access_tokens.permissions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(mcp): register takes_scorecard + takes_calibration ops

Both ops are read-scope, MCP-callable, allow-list-honoring. Handlers
thread ctx.takesHoldersAllowList into the engine method's required
allowList parameter, which applies WHERE holder = ANY at SQL aggregation
level (D4 fail-closed). Local CLI callers leave the allow-list
undefined and see all holders.

Updates the OperationContext.takesHoldersAllowList contract comment to
list the new aggregate ops alongside takes_list, takes_search, query.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.30.0 release: calibration core (Slice A1 of v0.30 wave)

VERSION + package.json bump. CHANGELOG entry covers the release-summary
(headline + math table + privacy note + data-loss-bug-killed note),
"## To take advantage of v0.30.0" upgrade path, and itemized changes.

llms-full.txt regenerated to capture the v0.28.x annotations that had
been merged but not yet rolled into the docs bundle.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): scorecard + calibration parity on real Postgres + NUMERIC fix

Adds end-to-end coverage for v0.30.0 (Slice A1) against real Postgres:

- test/e2e/takes-scorecard-parity.test.ts (new): seeds the same 6-bet
  fixture (4 binary garry + 1 partial garry + 1 binary harj) into both
  Postgres and PGLite, asserts getScorecard + getCalibrationCurve return
  byte-identical results across engines, runs the 4-bet hand-calc Brier
  reference (0.205) on real PG, and verifies the SQL-level allow-list
  filter strictly subtracts hidden-holder rows on both engines.

- test/e2e/takes-postgres.test.ts: extended with 8 v0.30 cases — quality
  semantics (correct/partial/back-compat) writes the expected (quality,
  outcome) tuple on real PG; the takes_resolution_consistency CHECK
  constraint actually fires on a contradictory raw UPDATE; getScorecard
  + getCalibrationCurve coherent shape + ordered-bucket invariants;
  PRIVACY allow-list filter on real PG; MCP dispatch path for
  takes_scorecard + takes_calibration with allow-list threading.

While writing the parity test, the e2e harness caught a real bug PGLite
tolerated: postgres.js sends scalar `${bucketSize}` params as text by
default, so `FLOOR(weight / $N)` tried to coerce '0.1' to integer and
threw `invalid input syntax for type integer: "0.1"`. The NUMERIC fix
also kills a separate FP-precision divergence — `FLOOR(0.7 / 0.1)`
returns 6 on real PG (IEEE 754 rounds 0.7/0.1 to 6.9999...) and 7 on
PGLite. Both engines now bucket via `weight::numeric / $N::numeric`
which is exact decimal arithmetic and engine-agnostic.

This is the v0.30.0 wave's first cross-engine parity test. Same shape
will guard A2's getTrajectory + getAnnualReview when those land.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(changelog): correct v40→v43 reference and expand v0.30.0 test note

Two fixes to the v0.30.0 entry after the master merge renumbered the
migration:

- The "#### Added" bullet still said "Schema migration v40"; bumped to v43.
- The "#### Tests" section only enumerated unit tests. The PR also ships
  19 E2E cases (11 in takes-scorecard-parity, 8 extending takes-postgres)
  that exercise the calibration math against real Postgres and the
  PG↔PGLite engine parity. Added the count + a note about the two real
  bugs the parity test caught (postgres.js string-typed scalar params
  and IEEE 754 bucketing divergence) that PGLite tolerated.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 23:07:00 -07:00

159 lines
5.5 KiB
TypeScript

/**
* v0.30.0 (Slice A1): tests for the pure resolution + scorecard helpers.
* Covers the (quality, outcome) tuple derivation and the Brier math
* including the partial-exclusion contract (D5 + D11).
*/
import { describe, test, expect } from 'bun:test';
import {
deriveResolutionTuple,
finalizeScorecard,
PARTIAL_RATE_WARNING_THRESHOLD,
} from '../src/core/takes-resolution.ts';
import { GBrainError } from '../src/core/types.ts';
describe('deriveResolutionTuple', () => {
test('quality=correct → (correct, true)', () => {
expect(deriveResolutionTuple({ quality: 'correct', resolvedBy: 'garry' })).toEqual({
quality: 'correct',
outcome: true,
});
});
test('quality=incorrect → (incorrect, false)', () => {
expect(deriveResolutionTuple({ quality: 'incorrect', resolvedBy: 'garry' })).toEqual({
quality: 'incorrect',
outcome: false,
});
});
test('quality=partial → (partial, null)', () => {
expect(deriveResolutionTuple({ quality: 'partial', resolvedBy: 'garry' })).toEqual({
quality: 'partial',
outcome: null,
});
});
test('outcome=true (back-compat alias) → (correct, true)', () => {
expect(deriveResolutionTuple({ outcome: true, resolvedBy: 'garry' })).toEqual({
quality: 'correct',
outcome: true,
});
});
test('outcome=false (back-compat alias) → (incorrect, false)', () => {
expect(deriveResolutionTuple({ outcome: false, resolvedBy: 'garry' })).toEqual({
quality: 'incorrect',
outcome: false,
});
});
test('quality wins when both inputs supplied AND consistent', () => {
// outcome=true is consistent with quality=correct; quality wins.
expect(deriveResolutionTuple({ quality: 'correct', outcome: true, resolvedBy: 'garry' })).toEqual({
quality: 'correct',
outcome: true,
});
});
test('contradictory quality + outcome throws TAKE_RESOLUTION_INVALID', () => {
expect(() =>
deriveResolutionTuple({ quality: 'correct', outcome: false, resolvedBy: 'garry' })
).toThrow(GBrainError);
expect(() =>
deriveResolutionTuple({ quality: 'partial', outcome: true, resolvedBy: 'garry' })
).toThrow(GBrainError);
});
test('neither field set throws TAKE_RESOLUTION_INVALID', () => {
expect(() =>
deriveResolutionTuple({ resolvedBy: 'garry' })
).toThrow(GBrainError);
});
});
describe('finalizeScorecard (Brier math)', () => {
test('n=0: returns nulls, no divide-by-zero', () => {
const card = finalizeScorecard({
total_bets: 0, resolved: 0, correct: 0, incorrect: 0, partial: 0, brier: null,
});
expect(card).toEqual({
total_bets: 0,
resolved: 0,
correct: 0,
incorrect: 0,
partial: 0,
accuracy: null,
brier: null,
partial_rate: null,
});
});
test('all correct: accuracy=1, Brier reflects mean (weight - 1)^2', () => {
// 3 correct bets at weight 0.7. Per-row Brier = (0.7 - 1)^2 = 0.09.
// Mean = 0.09. Accuracy = 3/3 = 1.0.
const card = finalizeScorecard({
total_bets: 3, resolved: 3, correct: 3, incorrect: 0, partial: 0, brier: 0.09,
});
expect(card.accuracy).toBe(1.0);
expect(card.brier).toBeCloseTo(0.09, 5);
expect(card.partial_rate).toBe(0);
});
test('all incorrect: accuracy=0, Brier reflects mean weight^2', () => {
// 3 incorrect bets at weight 0.7. Per-row Brier = (0.7 - 0)^2 = 0.49.
const card = finalizeScorecard({
total_bets: 3, resolved: 3, correct: 0, incorrect: 3, partial: 0, brier: 0.49,
});
expect(card.accuracy).toBe(0.0);
expect(card.brier).toBeCloseTo(0.49, 5);
});
test('hand-calculated reference: 4 mixed bets', () => {
// bets: weight=0.9 correct, weight=0.6 correct, weight=0.7 incorrect, weight=0.4 incorrect
// Brier per row: (0.9-1)^2=0.01, (0.6-1)^2=0.16, (0.7-0)^2=0.49, (0.4-0)^2=0.16
// Mean: (0.01+0.16+0.49+0.16)/4 = 0.205
// Accuracy: 2/4 = 0.5
const card = finalizeScorecard({
total_bets: 4, resolved: 4, correct: 2, incorrect: 2, partial: 0, brier: 0.205,
});
expect(card.accuracy).toBe(0.5);
expect(card.brier).toBeCloseTo(0.205, 5);
expect(card.partial_rate).toBe(0);
});
test('D5: partial excluded from Brier denominator; appears in partial_rate', () => {
// 2 correct, 1 incorrect, 1 partial. Brier reflects only the 3 binary rows
// (the SQL aggregation passes partial=null per row to AVG, so partial
// doesn't affect Brier in finalizeScorecard either).
// partial_rate = 1/4 = 0.25
const card = finalizeScorecard({
total_bets: 4, resolved: 4, correct: 2, incorrect: 1, partial: 1, brier: 0.18,
});
// accuracy uses correct + incorrect denominator (binary), excluding partial.
expect(card.accuracy).toBeCloseTo(2 / 3, 5);
expect(card.partial_rate).toBe(0.25);
expect(card.brier).toBe(0.18);
});
test('D11: partial_rate threshold constant matches plan (20%)', () => {
expect(PARTIAL_RATE_WARNING_THRESHOLD).toBe(0.20);
});
test('all-partial scorecard: Brier null, accuracy null, partial_rate=1', () => {
const card = finalizeScorecard({
total_bets: 5, resolved: 5, correct: 0, incorrect: 0, partial: 5, brier: null,
});
expect(card.brier).toBeNull();
expect(card.accuracy).toBeNull();
expect(card.partial_rate).toBe(1);
});
test('partial_rate = 0 when no partial bets', () => {
const card = finalizeScorecard({
total_bets: 2, resolved: 2, correct: 1, incorrect: 1, partial: 0, brier: 0.5,
});
expect(card.partial_rate).toBe(0);
});
});