Files
gbrain/test/migrations-v99.test.ts
f702ec053b v0.41.16.0 feat: conversation parser cathedral + progressive-batch primitive (closes #1461) (#1510)
* v0.41.15.0 feat: conversation parser cathedral + progressive-batch primitive (closes #1461)

Replaces PR #1461's single-format Telegram regex with a 12-pattern
built-in registry covering iMessage/Slack, Telegram (×2), Discord
(×2), WhatsApp (×2 locales), Signal, Matrix/Element, IRC (×2), Teams.
Each pattern is hand-vetted from public format docs (signal-cli,
DiscordChatExporter, Telegram Desktop, WhatsApp export docs, Element
matrix-archive, irssi/weechat defaults); module-load validation runs
test_positive[] + test_negative[] for every pattern at startup so a
typo makes gbrain refuse to start.

PR #1461 contributor's BRACKET_TIME_RX + cleanSpeaker survive verbatim
as the `telegram-bracket` built-in pattern + DEFAULT_SPEAKER_CLEAN
export. All 33 of their test cases pass against the new orchestrator.

Three layers per page (orchestrator chooses):
  1. Built-in pattern registry (zero-cost, deterministic)
  2. User-declared simple_pattern via config (deferred to v0.42+)
  3. Opt-IN LLM polish + fallback (privacy-first; chat content goes
     to Anthropic only when user explicitly enables)

D18 priority scoring picks the highest-match-rate pattern across the
first 10 lines (not first-wins) so overlapping formats don't silently
mis-route. D5 multi_line per-pattern + D11 quick_reject prefix screen
+ D19 timezone_policy per-pattern complete the registry shape.

Companion: src/core/progressive-batch/ primitive (rule of three
satisfied across 12+ ad-hoc cost-prompt sites). Wintermute-inspired
ramp shape (trial 10 → 100 → 500 → full with verification at each
stage), productionized with verifier+policy injection (callers
describe HOW TO MEASURE SUCCESS, not WHEN TO WAIT FOR CTRL-C). D3
fail-closed budget gate: null tracker + null Policy.maxCostUsd →
abort_cost_cap reason='no_budget_safety_net'. D20 discriminated
Verifier union (output_count | idempotent_mutation | noop).
extract-conversation-facts is the one proven consumer in v0.41.15.0;
9-site retrofit deferred to v0.41.16.0+ per TODOS.md.

Codex outside-voice review absorbed 8 substantive findings:
  - Privacy posture (LLM polish/fallback flipped to opt-IN)
  - ReDoS theater (dropped arbitrary user regex; v0.42+ uses RE2)
  - LLM-inferred-regex persistence as silent-corruption machine
  - Pattern priority scoring across first 10 lines
  - Timezone policy on every PatternEntry
  - Verifier shape discriminated union
  - Behavior parity for sites that "jumped straight to full"
  - Real-corpus-redacted fixture gap (v0.42+ TODO)

CI gates:
  - bun run check:conversation-parser (13 fixtures, --no-llm, deterministic)
  - bun run check:fixture-privacy (banned-token grep)

Doctor surfaces 3 new checks: conversation_format_coverage,
progressive_batch_audit_health, conversation_parser_probe_health.

Tests: 198/198 across primitive + parser + LLM + nightly probe + eval
CLI + debug CLI + doctor checks + migration v97 round-trip + E2E
parser ↔ engine integration. Real bug caught + fixed during gap audit:
IdempotentMutationVerifier was comparing absolute mutated-count vs
per-stage expected (failed silently on stage 2+); now uses per-stage
delta semantics matching OutputCountVerifier.

Schema migration v97: conversation_parser_llm_cache table with
(content_sha256, model_id, call_shape) composite key. NO
inferred_patterns table (D17: silent-corruption machine).

Plan + 23 decisions + codex outside-voice absorption at
~/.claude/plans/system-instruction-you-are-working-cuddly-hollerith.md.

Co-Authored-By: garrytan-agents (PR #1461) <noreply@github.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(check-privacy): allowlist scripts/check-fixture-privacy.sh

The new sibling privacy guard literally names the banned tokens in its
BANNED_TOKENS array — same meta-exception that check-privacy.sh itself
gets. Without this allowlist entry, bun run verify rejects the file
post-merge because the banned name appears in the rule-definition script.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: renumber v0.41.15.0 → v0.41.16.0 (queue drift)

Mechanical rename across all surfaces: VERSION, package.json,
CHANGELOG (header + body refs), CLAUDE.md, TODOS.md, src/core/
migrate.ts (migration v98 comment), all src/core/conversation-parser/*
and src/core/progressive-batch/* file headers, all test/ headers,
scripts/check-privacy.sh allowlist comment, llms-full.txt regenerated.

Audit clean: VERSION + package.json + CHANGELOG header all show
0.41.16.0. verify 24/24, touched tests 179/179.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: garrytan-agents (PR #1461) <noreply@github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 17:31:48 -07:00

160 lines
5.7 KiB
TypeScript

/**
* v0.41.16.0 — Migration v99 round-trip test.
*
* Verifies the `conversation_parser_llm_cache` table:
* - is created on schema init
* - accepts inserts on (content_sha256, model_id, call_shape, value_json)
* - rejects invalid call_shape via CHECK constraint
* - ON CONFLICT DO NOTHING semantics (the llm-base.ts caller's contract)
* - JSONB column round-trips a real object (no double-encode regression)
* - composite primary key prevents duplicate (sha, model, shape)
*
* Hermetic via the canonical PGLite block from CLAUDE.md test-isolation
* rules.
*/
import { describe, expect, test, beforeAll, afterAll, beforeEach } from 'bun:test';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await resetPgliteState(engine);
});
describe('migration v99 — conversation_parser_llm_cache', () => {
test('table exists after schema init', async () => {
const rows = await engine.executeRaw<{ table_name: string }>(
`SELECT table_name FROM information_schema.tables
WHERE table_schema = 'public' AND table_name = 'conversation_parser_llm_cache'`,
);
expect(rows.length).toBe(1);
});
test('insert + select round-trip with polish call_shape', async () => {
await engine.executeRaw(
`INSERT INTO conversation_parser_llm_cache
(content_sha256, model_id, call_shape, value_json)
VALUES ($1, $2, $3, $4::jsonb)`,
[
'abc123',
'anthropic:claude-haiku-4-5',
'polish',
JSON.stringify({ merge_indices: [], drop_indices: [], edits: [] }),
],
);
const rows = await engine.executeRaw<{ value_json: unknown }>(
`SELECT value_json FROM conversation_parser_llm_cache
WHERE content_sha256 = $1 AND model_id = $2 AND call_shape = $3`,
['abc123', 'anthropic:claude-haiku-4-5', 'polish'],
);
expect(rows).toHaveLength(1);
// value_json should round-trip as a parsed object (not a JSON string).
const val =
typeof rows[0].value_json === 'string'
? JSON.parse(rows[0].value_json)
: rows[0].value_json;
expect(val).toEqual({ merge_indices: [], drop_indices: [], edits: [] });
});
test('insert + select round-trip with fallback call_shape', async () => {
await engine.executeRaw(
`INSERT INTO conversation_parser_llm_cache
(content_sha256, model_id, call_shape, value_json)
VALUES ($1, $2, $3, $4::jsonb)`,
[
'def456',
'anthropic:claude-haiku-4-5',
'fallback',
JSON.stringify([
{ speaker: 'Alice', timestamp: '2024-03-15T18:37:00Z', text: 'hi' },
]),
],
);
const rows = await engine.executeRaw<{ value_json: unknown }>(
`SELECT value_json FROM conversation_parser_llm_cache
WHERE content_sha256 = $1 AND call_shape = $2`,
['def456', 'fallback'],
);
expect(rows).toHaveLength(1);
const val =
typeof rows[0].value_json === 'string'
? JSON.parse(rows[0].value_json)
: rows[0].value_json;
expect(Array.isArray(val)).toBe(true);
expect(val).toHaveLength(1);
});
test('CHECK constraint rejects invalid call_shape', async () => {
let threw = false;
try {
await engine.executeRaw(
`INSERT INTO conversation_parser_llm_cache
(content_sha256, model_id, call_shape, value_json)
VALUES ($1, $2, $3, $4::jsonb)`,
['ghi789', 'anthropic:claude-haiku-4-5', 'INVALID_SHAPE', '{}'],
);
} catch {
threw = true;
}
expect(threw).toBe(true);
});
test('composite primary key prevents duplicate (sha, model, shape)', async () => {
await engine.executeRaw(
`INSERT INTO conversation_parser_llm_cache
(content_sha256, model_id, call_shape, value_json)
VALUES ($1, $2, $3, $4::jsonb)`,
['dup1', 'anthropic:claude-haiku-4-5', 'polish', '{}'],
);
// ON CONFLICT DO NOTHING from llm-base.ts writeDbCache — should not throw.
await engine.executeRaw(
`INSERT INTO conversation_parser_llm_cache
(content_sha256, model_id, call_shape, value_json)
VALUES ($1, $2, $3, $4::jsonb)
ON CONFLICT (content_sha256, model_id, call_shape) DO NOTHING`,
['dup1', 'anthropic:claude-haiku-4-5', 'polish', '{"different":true}'],
);
// First write wins on conflict.
const rows = await engine.executeRaw<{ value_json: unknown }>(
`SELECT value_json FROM conversation_parser_llm_cache WHERE content_sha256 = 'dup1'`,
);
expect(rows).toHaveLength(1);
});
test('different call_shape on same (sha, model) coexists', async () => {
await engine.executeRaw(
`INSERT INTO conversation_parser_llm_cache
(content_sha256, model_id, call_shape, value_json)
VALUES ($1, $2, 'polish', $3::jsonb), ($1, $2, 'fallback', $3::jsonb)`,
['co1', 'anthropic:claude-haiku-4-5', '{}'],
);
const rows = await engine.executeRaw<{ call_shape: string }>(
`SELECT call_shape FROM conversation_parser_llm_cache WHERE content_sha256 = 'co1'`,
);
expect(rows).toHaveLength(2);
const shapes = rows.map((r) => r.call_shape).sort();
expect(shapes).toEqual(['fallback', 'polish']);
});
test('created_at index supports time-based pruning queries', async () => {
const rows = await engine.executeRaw<{ indexname: string }>(
`SELECT indexname FROM pg_indexes
WHERE tablename = 'conversation_parser_llm_cache'
AND indexname = 'idx_conversation_parser_llm_cache_created'`,
);
expect(rows.length).toBe(1);
});
});