mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* rfc: temporal axis for contradiction probe Field report on residual HIGH findings from gbrain eval suspected-contradictions and proposal for a 4-phase fix (Phase 1 = judge prompt + verdict enum is the recommended starting point). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(eval): pass effective_date to judge prompt; bump PROMPT_VERSION Lane A1 of the temporal-contradiction-probe wave. Threads page-level effective_date through the search projection into the contradiction judge so the LLM can reason about supersession instead of treating every dated pair as a contradiction. Changes: - SearchResult interface adds optional effective_date + effective_date_source fields; rowToSearchResult populates them from the row data with date-only YYYY-MM-DD normalization (handles both postgres.js Date and PGLite string). - 8 SELECT projection sites (3 in postgres-engine, 5 in pglite-engine) now carry p.effective_date + p.effective_date_source through their inner CTEs and outer SELECTs so search results expose the field on both engines. - PairMember (eval-contradictions/types.ts) gets the two fields as required (string | null) so the type forces every constructor to think about temporal anchoring. Runner's searchResultToMember + takeToMember handle the normalization; takes inherit the chunk's page-level date. - buildJudgePrompt emits `Statement A (from: YYYY-MM-DD)` when effective_date is non-null, else `(date unknown)`. Prompt instructions explain the tag so the model knows what to do with it. - PROMPT_VERSION bumps '1' → '2'. Cache-key tuple shape unchanged; old rows miss naturally on first run against the new prompt. Test fixtures in 5 files updated to include the new required fields. All 205 eval-contradictions unit tests + 101 search-related tests pass. Typecheck clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(eval): replace contradicts:boolean with verdict:enum (6 members) Lane A2 of the temporal-contradiction-probe wave. Expands the judge's classification vocabulary from a binary contradicts:bool to a six-member verdict enum so the probe can distinguish "this changed" from "this is wrong". Verdict taxonomy: no_contradiction — drop from findings contradiction — genuine conflict at same point in time temporal_supersession — newer claim updates/replaces older; not an error temporal_regression — metric/status went backwards over time (signal) temporal_evolution — legitimate change, neither supersession nor regression negation_artifact — judge misread an explicit negation Changes: - types.ts: Verdict union (6 members); Severity gains 'info'; ResolutionKind extended with temporal_supersede, flag_for_review, log_timeline_change; JudgeVerdict.contradicts → verdict; ContradictionFinding now carries verdict; ProbeReport adds queries_with_any_finding + verdict_breakdown (additive). - judge.ts: parseResolutionKind + parseVerdict guards; normalizeVerdict reads the new field and applies the C1 confidence floor only to verdict='contradiction' (the new verdicts are informational classifications, no floor). Prompt rubric rewritten to ask for verdict + extended severity scale. - severity-classify.ts: 'info' joins the rank with value 0; defaultSeverityForVerdict maps each verdict to its baseline severity (D7 — supersession=info, regression=high, etc.). parseSeverity gains a fallback param so consumers can override 'low' default. - auto-supersession.ts: classifyResolution + renderResolutionCommand handle the three new resolution kinds. Probe still NEVER auto-mutates — the new kinds render paste-ready commands or informational lines. - cache.ts: isJudgeVerdict shape check matches the new verdict field; old v1 rows fail the guard and treat as misses. - runner.ts: emit predicate at cache-hit and judge-success branches changes from `verdict.contradicts` to `verdict.verdict !== 'no_contradiction'`. Without this, the new verdicts vanish from the report. Added per-verdict tally + queriesWithAnyFinding alongside the strict queriesWithContradiction. - trends.ts: latest run verdict breakdown surfaces in the trend chart. Test fixtures updated across 8 test files. All 210 eval-contradictions unit tests pass. Typecheck clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(eval): relax date-filter rule 3 when both sides dated Lane B of the temporal-contradiction-probe wave. The v1 date pre-filter skipped pairs whose chunk-text-extracted dates differed by >30 days as a cost-saving heuristic. That heuristic silently killed exactly the cases the new verdict taxonomy exists to surface — role transitions across years (e.g. a 2017 historical record vs. a 2025 current state), MRR claims years apart, status changes recorded over time. Lane A1+A2 made temporal supersession explicit and cheap to classify. The filter no longer needs to skip these pairs; the judge can label them. Changes: - date-filter.ts: shouldSkipForDateMismatch accepts optional effectiveDateA and effectiveDateB. When BOTH are non-null, returns skip=false with the new 'both_have_effective_date' reason — the judge will see the dates via the (from: YYYY-MM-DD) prompt tag from Lane A1. Other rules (same-paragraph dual-date override, missing-date fallback) preserved verbatim and still run first. - runner.ts: threads pair.{a,b}.effective_date into the date-filter call. Pairs that previously vanished into the skip bucket now reach the judge. Tests (R1 IRON RULE regression suite, 6 new cases): - both sides effective_date → not skipped - both sides effective_date overrides >30d chunk-text rule - rule 1 (same-paragraph dual-date) still wins over effective_date relaxation - rule 2 (missing chunk dates) still applies when effective_date partially present - undefined effective_dates fall through to v1 behavior (back-compat) - empty-string effective_date treated as missing (only real dates enable the relaxation) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(cli): cost-estimate prompt + --budget-usd + Haiku routing Lane C of the temporal-contradiction-probe wave. Three layers of cost guardrail, all stacked: (a) cost-estimate prompt at probe-run-time. Before the runner spends any tokens after a PROMPT_VERSION change, eval-suspected-contradictions reads the most recent persisted prompt_version from eval_contradictions_runs and compares. When they differ: - TTY: prints an upper-bound estimate + Ctrl-C window (default 10s, override via GBRAIN_PROBE_PROMPT_GRACE_SECONDS). - non-TTY: prints the estimate + auto-proceeds (autopilot path). - --yes override or GBRAIN_NO_PROBE_PROMPT=1: skip entirely. Mirrors the v0.32.7 runPostUpgradeReembedPrompt pattern. (b) --budget-usd N hard cap (pre-existing; PreFlightBudgetError surfaces when the estimate alone exceeds the cap, and CostTracker halts the run mid-flight when cumulative cost exceeds it). Documented in the help text alongside (a). (c) Judge model now routes through resolveModel() with configKey 'models.eval.contradictions_judge', tier 'utility' (Haiku-class default), and env var GBRAIN_CONTRADICTIONS_JUDGE_MODEL. The legacy --judge CLI flag still wins as the highest-precedence override. Doctor's model touchpoint registry (src/commands/models.ts:50) carries the new key so `gbrain models` and `gbrain models doctor` surface it. Also in this lane: - CLI: --severity accepts 'info' (the new Severity member from Lane A2). - CLI: --severity output shows [verdict] tag alongside slug pairs so operators distinguish genuine contradictions from temporal classifications. - Human summary: prints the new queries_with_any_finding metric and the per-verdict breakdown table. - Help text: explains the cost-prompt + budget-cap + model-routing interactions in one paragraph. New tests (9 cases on the cost-prompt helper): - --yes override skips - GBRAIN_NO_PROBE_PROMPT=1 skips - prompt_version unchanged → skips - non-TTY auto-proceeds with stderr note - TTY proceeds after grace - TTY aborts on Ctrl-C - fresh brain (no prior runs) fires the prompt - GBRAIN_PROBE_PROMPT_GRACE_SECONDS override honored - estimate banner contains query count + judge model + dollar amount All 225 eval-contradictions tests + 25 model-config tests pass. Typecheck clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(eval): R4/R5/R6 IRON-RULE regressions for the verdict-enum wave Lane D of the temporal-contradiction-probe wave. The Lanes A1/A2/B/C lanes landed the behavior; this lane pins the regressions that protect the wave against future drift. R4 (runner emit predicate): five new tests, one per non-no_contradiction verdict, prove the runner.ts emit rule surfaces each one as a finding with the correct verdict tag, and that: - queries_with_contradiction (Wilson-CI denominator) ONLY counts verdict ='contradiction' — the strict metric is preserved - queries_with_any_finding counts every non-no_contradiction verdict - verdict_breakdown tallies correctly Plus one negative case: verdict='no_contradiction' produces zero findings. Without R4, a future runner refactor could collapse the new verdicts back to /dev/null and the report would silently shrink. R5 (cache key shape): direct shape assertion on buildCacheKey output. The key tuple is exactly 5 fields (chunk_a_hash, chunk_b_hash, model_id, prompt_version, truncation_policy). Adding a 6th field would silently break every operator's brain (no migration path). R6 (contradiction severity unchanged): four tests on normalizeVerdict pin the legacy semantics — judge-supplied severity wins (whether 'high' or 'low'), and on garbage severity input the fallback is 'medium' (per defaultSeverityForVerdict('contradiction')) NOT 'low'. The contradiction verdict's severity must never default to 'low', which would silently mask genuine conflicts as cosmetic naming issues. The temporal_regression case is included for parity (garbage → 'high' since regressions are real investor red flags). 236 eval-contradictions tests pass (211 + 6 R4 + 1 R5 + 4 R6 + 9 cost-prompt from Lane C). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(ci): privacy lint for docs/proposals/*.md Captures the residual TODO from the temporal-contradiction-probe wave's plan: prevent the bug class where an RFC lands in docs/proposals/ with PII that should never appear in a public technical artifact. The original RFC had to be scrubbed at force-push time (Step 0); this lint catches the same patterns at CI time so the next one can't slip through. Sibling to scripts/check-privacy.sh: - check-privacy.sh: bans the literal "Wintermute" repo-wide. - check-proposal-pii.sh: focuses on docs/proposals/*.md and the OTHER PII classes — personal-relationship vocabulary, private repo refs. Design contract: the denylist names PATTERNS, not real people. Naming specific real names (deceased relatives, therapist first names, dealflow contacts) inside this script would leak PII into the repo just by appearing here. The structural patterns below catch the SURROUNDING vocabulary that always accompanies such content in personal RFC prose. Trade-off: a future RFC that names a real person without any contextual markers won't be caught — accepted as residual risk handled by human review. Patterns flagged in docs/proposals/*.md: - garrytan/brain (private repo reference) - trial separation, permanent separation - couples session, couples therapist - divorce attorney(s) - grandmother's funeral, aunt's funeral - wintermute (also caught by check-privacy.sh; listed here for proposal-scoped clarity) Bare common words (separation, funeral) are NOT banned — only the combined personal-context phrases. "Separation of concerns" and other software vocabulary survives. Wired into: - `bun run verify` (gates every push) - `bun run check:all` - `bun run check:proposal-pii` (standalone) Tests: 15 cases in test/scripts/check-proposal-pii.test.ts. - Each pattern flagged when present, plus exit-code + stderr signal. - Two negative cases (separation-of-concerns, funeral metaphor) prove the lint doesn't false-positive on legitimate software prose. - No-proposals-dir → exit 0 (not a failure). - Multi-hit case proves all patterns surface together with a summary count. - The two test fixtures that name "Wintermute" / "WINTERMUTE" as sentinel literals are allowlisted in check-test-real-names.sh per the same meta-rule-enforcement exception as check-privacy.sh itself. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(privacy): allowlist new privacy-guard files in check-privacy.sh check-privacy.sh bans the literal Wintermute repo-wide. The two new files from the v0.34 privacy lint (scripts/check-proposal-pii.sh and its test) necessarily name the token to do their job. Same meta-rule-enforcement exception as scripts/check-privacy.sh itself, scripts/check-test-real-names.sh, test/recency-decay.test.ts, and the existing entries — describing what the rule forbids requires naming it. Without this allowlist, `bun run verify` fails on check:privacy. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: bump version and changelog (v0.35.1.0) Temporal-contradiction-probe wave — Phase 1 of the RFC at docs/proposals/temporal-contradiction-probe.md. Headline: the contradiction probe now classifies pairs into a 6-member verdict enum (no_contradiction, contradiction, temporal_supersession, temporal_regression, temporal_evolution, negation_artifact) and sees the page-level effective_date for each chunk via a (from: YYYY-MM-DD) tag in the prompt. The pre-judge date filter no longer skips dated wide-gap pairs, so the role-transition class (e.g. a 2017 historical record vs. a 2025 current state) reaches the judge and gets classified as temporal_supersession instead of vanishing into the skip bucket. PROMPT_VERSION bumped 1 → 2 (cache fully invalidated). Three-layer cost guardrail: TTY-only cost-estimate prompt with Ctrl-C window, --budget-usd hard cap, Haiku-tier routing via new models.eval.contradictions_judge config key. Also adds a CI privacy lint (scripts/check-proposal-pii.sh) wired into bun run verify that catches PII patterns in docs/proposals/*.md so future RFCs can't ship with personal-context vocabulary the way this wave's source RFC did at draft time. Phases 2-4 deferred to follow-up RFCs per the plan. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: garrytan-agents <garrytan-agents@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
322 lines
13 KiB
TypeScript
322 lines
13 KiB
TypeScript
/**
|
|
* E2E — Postgres-specific contradiction-probe behavior (v0.32.6, T1).
|
|
*
|
|
* PGLite covers the contract; this file exercises Postgres-only surfaces
|
|
* that PGLite can't:
|
|
* 1. The actual JSONB round-trip through postgres.js — sql.json() vs
|
|
* double-encode regression class.
|
|
* 2. Migrations v51 + v52 apply cleanly on a real PG instance and the
|
|
* tables come out with the expected column shapes.
|
|
* 3. The P1 batched listActiveTakesForPages uses ANY($1::int[]) which
|
|
* has subtly different semantics on real PG.
|
|
* 4. The full M5 trend write+read with PostgreSQL's TIMESTAMPTZ
|
|
* semantics and ORDER BY ran_at DESC stability.
|
|
* 5. The P2 cache TTL semantics with real `now()` and ON CONFLICT
|
|
* DO UPDATE.
|
|
* 6. The find_contradictions MCP op end-to-end via the dispatch path.
|
|
*
|
|
* Runs only when DATABASE_URL is set. Skips gracefully otherwise.
|
|
*/
|
|
|
|
import { afterAll, beforeAll, beforeEach, describe, expect, test } from 'bun:test';
|
|
import { PostgresEngine } from '../../src/core/postgres-engine.ts';
|
|
import { writeRunRow, loadTrend } from '../../src/core/eval-contradictions/trends.ts';
|
|
import { JudgeCache, buildCacheKey } from '../../src/core/eval-contradictions/cache.ts';
|
|
import type { ProbeReport } from '../../src/core/eval-contradictions/types.ts';
|
|
import { operationsByName, type OperationContext } from '../../src/core/operations.ts';
|
|
|
|
const DATABASE_URL = process.env.DATABASE_URL;
|
|
|
|
let engine: PostgresEngine | null = null;
|
|
|
|
beforeAll(async () => {
|
|
if (!DATABASE_URL) {
|
|
console.log('[e2e/eval-contradictions] DATABASE_URL not set — skipping.');
|
|
return;
|
|
}
|
|
engine = new PostgresEngine();
|
|
await engine.connect({ database_url: DATABASE_URL });
|
|
await engine.initSchema();
|
|
});
|
|
|
|
afterAll(async () => {
|
|
if (engine) await engine.disconnect();
|
|
});
|
|
|
|
beforeEach(async () => {
|
|
if (!engine) return;
|
|
await engine.executeRaw('DELETE FROM eval_contradictions_runs');
|
|
await engine.executeRaw('DELETE FROM eval_contradictions_cache');
|
|
});
|
|
|
|
function mkReport(opts: Partial<ProbeReport> = {}): ProbeReport {
|
|
return {
|
|
schema_version: 1,
|
|
run_id: opts.run_id ?? 'pg-test',
|
|
judge_model: 'anthropic:claude-haiku-4-5',
|
|
prompt_version: '1',
|
|
truncation_policy: '1500-chars-utf8-safe',
|
|
top_k: 5,
|
|
sampling: 'deterministic',
|
|
queries_evaluated: 50,
|
|
queries_with_contradiction: 12,
|
|
queries_with_any_finding: 18,
|
|
total_contradictions_flagged: 18,
|
|
verdict_breakdown: {
|
|
no_contradiction: 282,
|
|
contradiction: 12,
|
|
temporal_supersession: 4,
|
|
temporal_regression: 1,
|
|
temporal_evolution: 1,
|
|
negation_artifact: 0,
|
|
},
|
|
calibration: {
|
|
queries_total: 50,
|
|
queries_judged_clean: 38,
|
|
queries_with_contradiction: 12,
|
|
wilson_ci_95: { point: 0.24, lower: 0.14, upper: 0.37 },
|
|
},
|
|
judge_errors: { parse_fail: 1, refusal: 0, timeout: 0, http_5xx: 2, unknown: 0, total: 3, note: 'n' },
|
|
cost_usd: { judge: 1.18, embedding: 0.005, total: 1.185, estimate_note: 'approx' },
|
|
cache: { hits: 87, misses: 213, hit_rate: 0.29 },
|
|
duration_ms: 45000,
|
|
source_tier_breakdown: { curated_vs_curated: 2, curated_vs_bulk: 11, bulk_vs_bulk: 5, other: 0 },
|
|
per_query: [],
|
|
hot_pages: [],
|
|
...opts,
|
|
};
|
|
}
|
|
|
|
describe('E2E: eval_contradictions migrations applied cleanly', () => {
|
|
test('eval_contradictions_cache and eval_contradictions_runs tables exist', async () => {
|
|
if (!engine) return;
|
|
const rows = await engine.executeRaw<{ table_name: string }>(
|
|
`SELECT table_name FROM information_schema.tables
|
|
WHERE table_schema = 'public'
|
|
AND table_name IN ('eval_contradictions_cache', 'eval_contradictions_runs')
|
|
ORDER BY table_name`,
|
|
);
|
|
expect(rows.length).toBe(2);
|
|
expect(rows[0].table_name).toBe('eval_contradictions_cache');
|
|
expect(rows[1].table_name).toBe('eval_contradictions_runs');
|
|
});
|
|
|
|
test('eval_contradictions_runs has Wilson CI columns', async () => {
|
|
if (!engine) return;
|
|
const cols = await engine.executeRaw<{ column_name: string; data_type: string }>(
|
|
`SELECT column_name, data_type FROM information_schema.columns
|
|
WHERE table_name = 'eval_contradictions_runs'
|
|
AND column_name IN ('wilson_ci_lower', 'wilson_ci_upper')
|
|
ORDER BY column_name`,
|
|
);
|
|
expect(cols.length).toBe(2);
|
|
expect(cols[0].data_type).toBe('real');
|
|
expect(cols[1].data_type).toBe('real');
|
|
});
|
|
|
|
test('eval_contradictions_cache composite PK includes prompt_version + truncation_policy', async () => {
|
|
if (!engine) return;
|
|
const cols = await engine.executeRaw<{ column_name: string }>(
|
|
`SELECT a.attname AS column_name
|
|
FROM pg_index i
|
|
JOIN pg_class c ON c.oid = i.indrelid
|
|
JOIN pg_attribute a ON a.attrelid = c.oid AND a.attnum = ANY(i.indkey)
|
|
WHERE i.indisprimary AND c.relname = 'eval_contradictions_cache'
|
|
ORDER BY a.attname`,
|
|
);
|
|
const names = cols.map((c) => c.column_name);
|
|
expect(names).toContain('prompt_version');
|
|
expect(names).toContain('truncation_policy');
|
|
expect(names).toContain('chunk_a_hash');
|
|
expect(names).toContain('chunk_b_hash');
|
|
expect(names).toContain('model_id');
|
|
});
|
|
});
|
|
|
|
describe('E2E: JSONB round-trip on Postgres (regression class)', () => {
|
|
test('writeContradictionsRun → loadTrend preserves nested objects, not strings', async () => {
|
|
if (!engine) return;
|
|
await writeRunRow(engine, mkReport({
|
|
run_id: 'jsonb-1',
|
|
source_tier_breakdown: { curated_vs_curated: 7, curated_vs_bulk: 8, bulk_vs_bulk: 9, other: 0 },
|
|
}), 100);
|
|
const rows = await loadTrend(engine, 30);
|
|
expect(rows.length).toBe(1);
|
|
// The classic v0.12 double-encode bug stores '{"curated_vs_curated":7,...}' as a string.
|
|
// We must see a parsed object, not a string.
|
|
expect(typeof rows[0].source_tier_breakdown).toBe('object');
|
|
expect(rows[0].source_tier_breakdown.curated_vs_curated).toBe(7);
|
|
expect(rows[0].source_tier_breakdown.curated_vs_bulk).toBe(8);
|
|
expect(typeof rows[0].report_json).toBe('object');
|
|
expect(rows[0].report_json.schema_version).toBe(1);
|
|
});
|
|
|
|
test('postgres jsonb_typeof confirms object shape (defense in depth)', async () => {
|
|
if (!engine) return;
|
|
await writeRunRow(engine, mkReport({ run_id: 'jsonb-2' }), 100);
|
|
const rows = await engine.executeRaw<{ kind: string }>(
|
|
`SELECT jsonb_typeof(source_tier_breakdown) AS kind
|
|
FROM eval_contradictions_runs
|
|
WHERE run_id = 'jsonb-2'`,
|
|
);
|
|
expect(rows[0].kind).toBe('object');
|
|
});
|
|
});
|
|
|
|
describe('E2E: P2 persistent cache with real now()', () => {
|
|
test('lookup returns null for missing key, upsert + lookup round-trips', async () => {
|
|
if (!engine) return;
|
|
const cache = new JudgeCache({ engine, modelId: 'haiku-pg-test' });
|
|
expect(await cache.lookup('text-a', 'text-b')).toBeNull();
|
|
await cache.store('text-a', 'text-b', {
|
|
verdict: 'contradiction', severity: 'high', axis: 'pg-test', confidence: 0.9, resolution_kind: 'dream_synthesize',
|
|
});
|
|
const hit = await cache.lookup('text-a', 'text-b');
|
|
expect(hit).not.toBeNull();
|
|
expect(hit?.verdict).toBe('contradiction');
|
|
expect(hit?.severity).toBe('high');
|
|
});
|
|
|
|
test('expired rows hidden from lookup; sweepContradictionCache deletes', async () => {
|
|
if (!engine) return;
|
|
const cache = new JudgeCache({ engine, modelId: 'haiku-pg-test', ttlSeconds: 60 });
|
|
await cache.store('expire-me-a', 'expire-me-b', {
|
|
verdict: 'no_contradiction', severity: 'info', axis: '', confidence: 0.3, resolution_kind: null,
|
|
});
|
|
// Backdate expires_at by 1 second.
|
|
const key = buildCacheKey({ textA: 'expire-me-a', textB: 'expire-me-b', modelId: 'haiku-pg-test' });
|
|
await engine.executeRaw(
|
|
`UPDATE eval_contradictions_cache
|
|
SET expires_at = now() - interval '1 second'
|
|
WHERE chunk_a_hash = $1 AND chunk_b_hash = $2`,
|
|
[key.chunk_a_hash, key.chunk_b_hash],
|
|
);
|
|
expect(await cache.lookup('expire-me-a', 'expire-me-b')).toBeNull();
|
|
const swept = await engine.sweepContradictionCache();
|
|
expect(swept).toBeGreaterThanOrEqual(1);
|
|
});
|
|
|
|
test('different prompt_version is a different cache key (Codex fix)', async () => {
|
|
if (!engine) return;
|
|
const cache1 = new JudgeCache({ engine, modelId: 'haiku-pg-test' });
|
|
await cache1.store('shared-a', 'shared-b', {
|
|
verdict: 'contradiction', severity: 'medium', axis: '', confidence: 0.85, resolution_kind: 'manual_review',
|
|
});
|
|
// Direct engine call with a different prompt_version should miss.
|
|
const wrong = await engine.getContradictionCacheEntry({
|
|
chunk_a_hash: buildCacheKey({ textA: 'shared-a', textB: 'shared-b', modelId: 'haiku-pg-test' }).chunk_a_hash,
|
|
chunk_b_hash: buildCacheKey({ textA: 'shared-a', textB: 'shared-b', modelId: 'haiku-pg-test' }).chunk_b_hash,
|
|
model_id: 'haiku-pg-test',
|
|
prompt_version: 'OTHER-VERSION',
|
|
truncation_policy: '1500-chars-utf8-safe',
|
|
});
|
|
expect(wrong).toBeNull();
|
|
});
|
|
});
|
|
|
|
describe('E2E: M5 trend semantics on Postgres', () => {
|
|
test('trend ordered newest first with TIMESTAMPTZ', async () => {
|
|
if (!engine) return;
|
|
await writeRunRow(engine, mkReport({ run_id: 'older' }), 100);
|
|
// Add a small delay so the second row gets a strictly-later now().
|
|
await new Promise((r) => setTimeout(r, 50));
|
|
await writeRunRow(engine, mkReport({ run_id: 'newer' }), 100);
|
|
const rows = await loadTrend(engine, 30);
|
|
expect(rows[0].run_id).toBe('newer');
|
|
expect(rows[1].run_id).toBe('older');
|
|
});
|
|
|
|
test('days window filters via ran_at >= cutoff', async () => {
|
|
if (!engine) return;
|
|
await writeRunRow(engine, mkReport({ run_id: 'recent' }), 100);
|
|
// Backdate one row to 10 days ago.
|
|
await engine.executeRaw(
|
|
`UPDATE eval_contradictions_runs SET ran_at = now() - interval '10 days' WHERE run_id = $1`,
|
|
['recent'],
|
|
);
|
|
const oneDayRows = await loadTrend(engine, 1);
|
|
expect(oneDayRows.length).toBe(0);
|
|
const fifteenDayRows = await loadTrend(engine, 15);
|
|
expect(fifteenDayRows.length).toBe(1);
|
|
});
|
|
});
|
|
|
|
describe('E2E: find_contradictions MCP op on Postgres', () => {
|
|
test('returns "no probe runs" note on empty table', async () => {
|
|
if (!engine) return;
|
|
const op = operationsByName['find_contradictions'];
|
|
const ctx: OperationContext = {
|
|
engine,
|
|
config: {} as OperationContext['config'],
|
|
logger: { info: () => {}, warn: () => {}, error: () => {}, debug: () => {} } as unknown as OperationContext['logger'],
|
|
dryRun: false,
|
|
remote: true,
|
|
sourceId: 'default',
|
|
};
|
|
const result = await op.handler(ctx, {}) as { contradictions: unknown[]; note?: string };
|
|
expect(result.contradictions).toEqual([]);
|
|
expect(result.note).toContain('No probe runs');
|
|
});
|
|
|
|
test('returns latest run findings with slug+severity filters', async () => {
|
|
if (!engine) return;
|
|
await writeRunRow(engine, mkReport({
|
|
run_id: 'pg-mcp',
|
|
per_query: [{
|
|
query: 'q',
|
|
result_count: 5,
|
|
pairs_skipped_by_date: 0,
|
|
pairs_cache_hit: 0,
|
|
pairs_judged: 3,
|
|
contradictions: [
|
|
{
|
|
kind: 'cross_slug_chunks',
|
|
a: { slug: 'companies/acme-example', chunk_id: 1, take_id: null, source_tier: 'curated', holder: null, text: 'a', effective_date: null, effective_date_source: null },
|
|
b: { slug: 'openclaw/chat/x', chunk_id: 2, take_id: null, source_tier: 'bulk', holder: null, text: 'b', effective_date: null, effective_date_source: null },
|
|
combined_score: 1.5,
|
|
verdict: 'contradiction',
|
|
severity: 'high',
|
|
axis: 'MRR figure',
|
|
confidence: 0.9,
|
|
resolution_kind: 'dream_synthesize',
|
|
resolution_command: 'gbrain dream --phase synthesize --slug companies/acme-example',
|
|
},
|
|
{
|
|
kind: 'cross_slug_chunks',
|
|
a: { slug: 'people/alice-example', chunk_id: 3, take_id: null, source_tier: 'curated', holder: null, text: 'c', effective_date: null, effective_date_source: null },
|
|
b: { slug: 'people/alice-smith-example', chunk_id: 4, take_id: null, source_tier: 'curated', holder: null, text: 'd', effective_date: null, effective_date_source: null },
|
|
combined_score: 1.2,
|
|
verdict: 'contradiction',
|
|
severity: 'low',
|
|
axis: 'name format',
|
|
confidence: 0.75,
|
|
resolution_kind: 'manual_review',
|
|
resolution_command: 'gbrain takes mark-debate people/alice-example --row 1',
|
|
},
|
|
],
|
|
}],
|
|
}), 100);
|
|
|
|
const op = operationsByName['find_contradictions'];
|
|
const ctx: OperationContext = {
|
|
engine,
|
|
config: {} as OperationContext['config'],
|
|
logger: { info: () => {}, warn: () => {}, error: () => {}, debug: () => {} } as unknown as OperationContext['logger'],
|
|
dryRun: false,
|
|
remote: true,
|
|
sourceId: 'default',
|
|
};
|
|
|
|
const all = await op.handler(ctx, {}) as { contradictions: unknown[]; total_in_run: number };
|
|
expect(all.contradictions.length).toBe(2);
|
|
expect(all.total_in_run).toBe(2);
|
|
|
|
const highOnly = await op.handler(ctx, { severity: 'high' }) as { contradictions: Array<{ severity: string }> };
|
|
expect(highOnly.contradictions.length).toBe(1);
|
|
expect(highOnly.contradictions[0].severity).toBe('high');
|
|
|
|
const slugFiltered = await op.handler(ctx, { slug: 'acme' }) as { contradictions: unknown[] };
|
|
expect(slugFiltered.contradictions.length).toBe(1);
|
|
});
|
|
});
|