mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* feat(dims): OpenAI text-embedding-3 Matryoshka range validation (D13) dimsProviderOptions now fail-loud at the embed boundary when the configured embedding_dimensions is outside the model's native range (1..1536 for -small, 1..3072 for -large). Paste-ready fix hint in the AIConfigError.fix field. Closes the silent-HTTP-400 path that would have bit OpenAI-fallback users on v0.36.0.0 ZE-default installs. 16 new test cases in test/ai/dims-openai.test.ts pinning the contract across native-openai and openai-compatible adapter paths. * feat(ai): flip defaults to ZeroEntropy zembed-1 1280d + zerank-2 reranker Default embedding model is now zeroentropyai:zembed-1 at 1280d via Matryoshka. Real-corpus benchmark: 2.2x faster than OpenAI, 2.6x cheaper at regular pricing, wins 11/20 head-to-head queries. 1280 is the closest valid ZE Matryoshka step to the prior OpenAI 1536d default (valid set: 2560/1280/640/320/160/80/40). 1024 (Voyage's step) is NOT on ZE's list — pinned by AIConfigError fail-loud in dims.ts. balanced mode bundle now defaults reranker_enabled=true. zerank-2 reshuffles 60% of top-1 results in benchmarks. Missing-key fail-open contract in src/core/search/rerank.ts handles unauthenticated cases. Opt out with: gbrain config set search.reranker.enabled false Existing tests updated (gateway.test.ts, search-mode.test.ts) and a new test/balanced-reranker-default.test.ts (10 cases) pins the fail- open invariants. * feat(retrieval-upgrade): RetrievalUpgradePlanner + interactive prompt UX New src/core/retrieval-upgrade-planner.ts is the consolidated planner that computes the brain's pending retrieval-upgrade work (chunker bumps + ZE switch) in one pass and applies the schema transition + config updates atomically. Tagged-union ApplyResult enum (D15): 'applied' | 'skipped_already_ applied' | 'skipped_no_work' | 'declined' | 'planned' | 'failed'. No string-parsing reasons. Three config keys (D12): ze_switch_prompt_shown (UI state), ze_switch_requested (user intent), ze_switch_applied (work done). Plus ze_switch_previous_snapshot (JSON, full prior config for --undo per D16) and ze_switch_declined_at (90-day re-ask window). Schema transition (D18) is atomic: DROP indexes + ALTER COLUMN + CREATE INDEX inside a single engine.transaction(). HNSW recreation is part of the same transaction — no silent slow-search window. C3 eligibility logic: ze_switch_offered iff NOT on ZE + NOT declined recently + NOT applied + (legacy default OR >100 pages). C4 cost math: MAX(chunker_pending, dim_pending) not SUM — one re-embed pass invalidates both surfaces simultaneously. New src/core/retrieval-upgrade-prompt.ts wires the planner to a TTY-only interactive prompt with two-line cost split (D10) and privacy callout for the reranker flip. Tests: test/retrieval-upgrade-planner.test.ts (24 cases) pins the state machine. test/asymmetric-encoding-contract.test.ts (6 cases) pins D17: search read path uses gateway.embedQuery() not embed(), asserted via __setEmbedTransportForTests mock. * feat(cli): gbrain ze-switch — manual lever for the ZE switch New gbrain ze-switch CLI with --dry-run, --json, --resume, --force, --undo, --non-interactive, --confirm-reembed, --ignore-missing-key flags. Mirrors the upgrade prompt's UX symmetry: --undo presents a cost-warning before re-embedding back to the prior width. src/cli.ts: dispatch case + CLI_ONLY entry. ze-switch owns its own engine lifecycle (mirrors the doctor pattern). test/ze-switch-cli.test.ts (11 cases): --help, --dry-run, --json, --non-interactive, --ignore-missing-key, --resume, --undo, --confirm-reembed. Uses captureExit harness to test process.exit() paths without breaking the test process. * feat(doctor): ze_embedding_health + embedding_width_consistency checks Two new doctor checks (D-A5): ze_embedding_health: when embedding_model starts with zeroentropyai:, verify ZEROENTROPY_API_KEY is set (env or config). Paste-ready setup hint with the signup URL on failure. embedding_width_consistency: cross-check that the configured embedding_dimensions matches the actual vector(N) column width on content_chunks.embedding. Catches the half-applied switch state (schema migrated but config write crashed) with a paste-ready gbrain ze-switch --resume hint. Wired into runDoctor between reranker_health and the existing sync_freshness checks. Both checks gracefully no-op on non-ZE embedding configs. test/doctor-ze-checks.test.ts (8 cases) pins both checks across happy + missing-key + missing-config + drift paths. Uses withEnv() helper to clear ZEROENTROPY_API_KEY for the no-key path so tests are hermetic against contributor env state. test/e2e/v0_28_5-fix-wave.test.ts + test/openai-compat-multimodal.test.ts: updated to explicit-configure the gateway when the test depends on specific dims that diverge from the v0.36.0.0 default (1280d). * docs: README zero-based rewrite (884 -> 139 lines) + new docs files Strip 4 months of accreted "New in v0.X.Y" hero blocks and reorganize around what gbrain does today. 33 H2s -> 8. The Commands section (136 lines duplicating gbrain --help) moved out; the 6-table skills enumeration collapsed to a one-paragraph capability description with a link to skills/RESOLVER.md. Hero retains load-bearing facts: OpenClaw + Hermes credit, production numbers (17,888 pages / 4,383 people / 723 companies), BrainBench numbers (P@5 49.1% / R@5 97.9% / +31.4 lift), ZE comparison numbers, 30-min install claim. Adds one paragraph announcing the v0.36.0.0 ZE default with the explicit gbrain config set escape for OpenAI/Voyage users. New files: - docs/INSTALL.md: every install path consolidated (agent platform, CLI standalone, MCP server). Thin-client mode covered. - docs/architecture/RETRIEVAL.md: why the hybrid + graph stack works. BrainBench numbers, why each strategy alone fails, the source-aware ranking + intent classification + multi-query expansion story. - docs/ethos/ORIGIN.md: origin story lifted from the old README so the front door stays factual + concrete. test/readme-hero-anchors.test.ts (5 cases) is the D9 regression guard. Five load-bearing strings: OpenClaw, Hermes, ZE, production-numbers regex, P@5/R@5. Light anchors that let voice/ structure evolve but block accidental loss of headline facts. scripts/check-test-real-names.sh: allowlist entries for OpenClaw + Hermes literals in the anchor test (it explicitly asserts those strings appear in README). * chore: bump version and changelog (v0.36.0.0) ZeroEntropy as the new default for embedding (zembed-1 at 1280d via Matryoshka) and reranker (zerank-2 cross-encoder, on by default in balanced mode bundle). README zero-based rewrite (884 -> 139 lines). 3 new docs files. Two new doctor checks. New gbrain ze-switch CLI with --undo for symmetric reversibility. skills/migrations/v0.36.0.0.md tells the agent how to surface the retrieval-upgrade prompt post-upgrade. llms-full.txt regenerated via bun run build:llms. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(docs): scrub Wintermute from RETRIEVAL.md per privacy rule * chore: rebump version 0.36.0.0 → 0.36.2.0 (queue collision) Three open PRs were claiming v0.36.0.0 (#1130 skillpack, #1139 hindsight, #1136 this PR). Ship-aware queue allocator says this branch lands at v0.36.2.0. Trio audit: VERSION 0.36.2.0 package.json 0.36.2.0 CHANGELOG ## [0.36.2.0] - 2026-05-17 Updates: VERSION, package.json, CHANGELOG header + body refs, README "New default in v0.36.2.0" announcement + credit line, skills/migrations/v0.36.0.0.md renamed to v0.36.2.0.md with frontmatter + body refs updated. llms-full.txt regenerated. * fix(test): pin gateway dim=1536 in cross-file-stateful PGLite tests CI shard 1 reported 10 failures across `query-cache.test.ts` (6) and `consolidate-valid-until.test.ts` (4). Both files hardcode 1536-dim vectors but rely on `PGLiteEngine.initSchema()` to size `vector(__EMBEDDING_DIMS__)` at the right width. Root cause: v0.36.2.0 flipped DEFAULT_EMBEDDING_DIMENSIONS from 1536 to 1280 (ZE Matryoshka step). The gateway module is process-singleton; when ANOTHER test file in the same shard's bun-test process configures the gateway before us, `pglite-engine.ts:216` reads `getEmbeddingDimensions() === 1280` and sizes the schema columns at vector(1280). The hardcoded 1536-dim INSERTs then fail with "expected 1280 dimensions, not 1536". Locally these tests pass in isolation because the gateway falls back through the try/catch at pglite-engine.ts:218 (1536 default). CI runs multiple test files in one process, so cross-file state poisons the schema width. Fix: explicit `resetGateway()` + `configureGateway({embedding_dimensions: 1536, ...})` at the top of `beforeAll`, plus `resetGateway()` in `afterAll`. Pins the schema width regardless of cross-file state. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
241 lines
10 KiB
TypeScript
241 lines
10 KiB
TypeScript
/**
|
|
* v0.35.4 (D-CDX-4) — consolidate semantic upsert + chronological
|
|
* valid_until writeback.
|
|
*
|
|
* Pins:
|
|
* - R4a: a cluster of 3 chronologically-ordered facts produces
|
|
* 2 facts with valid_until set (older) and 1 with NULL (newest).
|
|
* - R4b/R7: running consolidate twice on the same input produces zero
|
|
* NEW takes (semantic upsert by (page_id, claim, since_date)).
|
|
* This is the Codex F4 fix — without it, the second cycle's
|
|
* extract_facts would clear consolidated_at and the second
|
|
* consolidate would append duplicate takes via MAX(row_num)+1.
|
|
*/
|
|
|
|
import { describe, test, expect, beforeAll, afterAll, beforeEach } from 'bun:test';
|
|
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
|
import { runPhaseConsolidate } from '../src/core/cycle/phases/consolidate.ts';
|
|
import { configureGateway, resetGateway } from '../src/core/ai/gateway.ts';
|
|
|
|
let engine: PGLiteEngine;
|
|
|
|
beforeAll(async () => {
|
|
// v0.36.2.0: DEFAULT_EMBEDDING_DIMENSIONS flipped to 1280 (ZE Matryoshka).
|
|
// This test inserts 1536-dim unit vectors (line ~38). If another test file
|
|
// in the shard configured the gateway before us, initSchema() would size
|
|
// facts.embedding at vector(1280) and the inserts below would throw
|
|
// "expected 1280 dimensions, not 1536". Pin the gateway to 1536d so this
|
|
// file is hermetic against cross-file state.
|
|
resetGateway();
|
|
configureGateway({
|
|
embedding_model: 'openai:text-embedding-3-large',
|
|
embedding_dimensions: 1536,
|
|
env: { OPENAI_API_KEY: 'sk-fake' },
|
|
});
|
|
engine = new PGLiteEngine();
|
|
await engine.connect({});
|
|
await engine.initSchema();
|
|
});
|
|
|
|
afterAll(async () => {
|
|
await engine.disconnect();
|
|
resetGateway();
|
|
});
|
|
|
|
beforeEach(async () => {
|
|
await engine.executeRaw(`DELETE FROM facts`);
|
|
await engine.executeRaw(`DELETE FROM takes`);
|
|
await engine.executeRaw(`DELETE FROM pages WHERE slug LIKE 'cdx4-%'`);
|
|
});
|
|
|
|
function unitVec(): string {
|
|
const a = new Float32Array(1536);
|
|
a[0] = 1.0;
|
|
return '[' + Array.from(a).join(',') + ']';
|
|
}
|
|
|
|
async function seedPage(slug: string): Promise<number> {
|
|
await engine.executeRaw(
|
|
`INSERT INTO pages (slug, type, title) VALUES ($1, 'company', 'Test') ON CONFLICT DO NOTHING`,
|
|
[slug],
|
|
);
|
|
const r = await engine.executeRaw<{ id: number }>(
|
|
`SELECT id FROM pages WHERE slug = $1 AND source_id = 'default'`,
|
|
[slug],
|
|
);
|
|
return r[0].id;
|
|
}
|
|
|
|
async function insertFact(args: {
|
|
entity_slug: string;
|
|
text: string;
|
|
valid_from: Date;
|
|
confidence?: number;
|
|
}): Promise<number> {
|
|
const r = await engine.executeRaw<{ id: number }>(
|
|
`INSERT INTO facts (source_id, entity_slug, fact, kind, source, valid_from, confidence, embedding, embedded_at)
|
|
VALUES ('default', $1, $2, 'fact', 'test', $3::timestamptz, $4, $5::vector, $3::timestamptz)
|
|
RETURNING id`,
|
|
[args.entity_slug, args.text, args.valid_from.toISOString(), args.confidence ?? 0.9, unitVec()],
|
|
);
|
|
return r[0].id;
|
|
}
|
|
|
|
describe('R4a — chronological valid_until writeback', () => {
|
|
test('cluster of 3 chronologically-ordered facts: 2 older get valid_until set, newest stays NULL', async () => {
|
|
await seedPage('cdx4-acme-mrr');
|
|
const olderDay = new Date('2026-01-15T00:00:00Z');
|
|
const midDay = new Date('2026-04-12T00:00:00Z');
|
|
const newest = new Date('2026-07-08T00:00:00Z');
|
|
|
|
// All three close enough in vector space to cluster together (identical
|
|
// embeddings via unitVec()). Past the 24h "oldest age" gate.
|
|
const idOlder = await insertFact({
|
|
entity_slug: 'cdx4-acme-mrr',
|
|
text: 'MRR claim',
|
|
valid_from: olderDay,
|
|
});
|
|
const idMid = await insertFact({
|
|
entity_slug: 'cdx4-acme-mrr',
|
|
text: 'MRR claim',
|
|
valid_from: midDay,
|
|
});
|
|
const idNewest = await insertFact({
|
|
entity_slug: 'cdx4-acme-mrr',
|
|
text: 'MRR claim',
|
|
valid_from: newest,
|
|
});
|
|
|
|
const r = await runPhaseConsolidate(engine, {});
|
|
expect(r.details.facts_consolidated).toBe(3);
|
|
expect(r.details.takes_written).toBe(1);
|
|
|
|
const rows = await engine.executeRaw<{ id: number; valid_until: Date | null }>(
|
|
`SELECT id, valid_until FROM facts WHERE entity_slug = 'cdx4-acme-mrr' ORDER BY valid_from ASC`,
|
|
);
|
|
expect(rows.length).toBe(3);
|
|
// Older fact's valid_until = mid.valid_from.
|
|
expect(rows[0].id).toBe(idOlder);
|
|
expect(rows[0].valid_until).not.toBeNull();
|
|
expect(new Date(rows[0].valid_until!).toISOString().slice(0, 10)).toBe('2026-04-12');
|
|
// Mid fact's valid_until = newest.valid_from.
|
|
expect(rows[1].id).toBe(idMid);
|
|
expect(rows[1].valid_until).not.toBeNull();
|
|
expect(new Date(rows[1].valid_until!).toISOString().slice(0, 10)).toBe('2026-07-08');
|
|
// Newest fact's valid_until stays NULL.
|
|
expect(rows[2].id).toBe(idNewest);
|
|
expect(rows[2].valid_until).toBeNull();
|
|
});
|
|
|
|
test('same-day cluster (3 facts, identical valid_from): id tiebreaker establishes chronological order', async () => {
|
|
await seedPage('cdx4-acme-sameday');
|
|
const sameDay = new Date(Date.now() - 30 * 60 * 60 * 1000);
|
|
const idA = await insertFact({ entity_slug: 'cdx4-acme-sameday', text: 'same day', valid_from: sameDay });
|
|
const idB = await insertFact({ entity_slug: 'cdx4-acme-sameday', text: 'same day', valid_from: sameDay });
|
|
const idC = await insertFact({ entity_slug: 'cdx4-acme-sameday', text: 'same day', valid_from: sameDay });
|
|
|
|
await runPhaseConsolidate(engine, {});
|
|
|
|
// All three valid_from values are equal; the (id ASC) tiebreaker
|
|
// makes the lowest-id row the "oldest" chronologically. Pin that
|
|
// contract since the trajectory CLI depends on this ordering.
|
|
const rows = await engine.executeRaw<{ id: number; valid_until: Date | null }>(
|
|
`SELECT id, valid_until FROM facts WHERE entity_slug = 'cdx4-acme-sameday' ORDER BY id ASC`,
|
|
);
|
|
expect(rows.length).toBe(3);
|
|
expect(rows[0].id).toBe(idA);
|
|
expect(rows[1].id).toBe(idB);
|
|
expect(rows[2].id).toBe(idC);
|
|
// First two are "older" by tiebreaker → both get valid_until set
|
|
// (= sameDay, since the next-newer fact has the same valid_from).
|
|
expect(rows[0].valid_until).not.toBeNull();
|
|
expect(rows[1].valid_until).not.toBeNull();
|
|
// Newest by tiebreaker stays NULL.
|
|
expect(rows[2].valid_until).toBeNull();
|
|
});
|
|
});
|
|
|
|
describe('R4b / R7 — cycle idempotency: re-run consolidate produces zero new takes (Codex F4 fix)', () => {
|
|
test('semantic upsert: second consolidate on identical state produces zero NEW takes', async () => {
|
|
await seedPage('cdx4-idempo-1');
|
|
const oldDate = new Date(Date.now() - 30 * 60 * 60 * 1000);
|
|
for (let i = 0; i < 4; i++) {
|
|
await insertFact({
|
|
entity_slug: 'cdx4-idempo-1',
|
|
text: 'stable claim',
|
|
valid_from: new Date(oldDate.getTime() + i * 60 * 60 * 1000),
|
|
});
|
|
}
|
|
|
|
// First run: 1 take, 4 facts consolidated.
|
|
const r1 = await runPhaseConsolidate(engine, {});
|
|
expect(r1.details.takes_written).toBe(1);
|
|
const countAfter1 = await engine.executeRaw<{ n: string }>(
|
|
`SELECT COUNT(*)::text AS n FROM takes WHERE page_id = (SELECT id FROM pages WHERE slug = 'cdx4-idempo-1')`,
|
|
);
|
|
expect(parseInt(countAfter1[0].n, 10)).toBe(1);
|
|
|
|
// Simulate the Codex F4 scenario: clear consolidated_at on every fact
|
|
// (extract_facts cycle phase wipes facts via delete-then-insert, which
|
|
// is functionally identical to NULL-ing consolidated_at). DO NOT touch
|
|
// valid_until — the prior consolidate wrote it; the semantic upsert
|
|
// should still find the take.
|
|
await engine.executeRaw(
|
|
`UPDATE facts SET consolidated_at = NULL, consolidated_into = NULL
|
|
WHERE entity_slug = 'cdx4-idempo-1'`,
|
|
);
|
|
|
|
// Second run: must NOT append another take.
|
|
const r2 = await runPhaseConsolidate(engine, {});
|
|
expect(r2.details.facts_consolidated).toBe(4);
|
|
// takes_written reports the NEW takes inserted this run; on the upsert
|
|
// hit path it's 0 (no new INSERT) but facts still get marked consolidated.
|
|
expect(r2.details.takes_written).toBe(0);
|
|
|
|
const countAfter2 = await engine.executeRaw<{ n: string }>(
|
|
`SELECT COUNT(*)::text AS n FROM takes WHERE page_id = (SELECT id FROM pages WHERE slug = 'cdx4-idempo-1')`,
|
|
);
|
|
expect(parseInt(countAfter2[0].n, 10)).toBe(1); // STILL 1 — no duplicate
|
|
|
|
// Facts were re-consolidated into the existing take.
|
|
const facts = await engine.executeRaw<{ consolidated_into: number }>(
|
|
`SELECT consolidated_into FROM facts WHERE entity_slug = 'cdx4-idempo-1' AND consolidated_into IS NOT NULL`,
|
|
);
|
|
expect(facts.length).toBe(4);
|
|
});
|
|
|
|
test('valid_until idempotency: second run leaves valid_until unchanged (no diff)', async () => {
|
|
await seedPage('cdx4-idempo-2');
|
|
const t1 = new Date('2026-01-15T00:00:00Z');
|
|
const t2 = new Date('2026-04-12T00:00:00Z');
|
|
const t3 = new Date('2026-07-08T00:00:00Z');
|
|
await insertFact({ entity_slug: 'cdx4-idempo-2', text: 'iterable', valid_from: t1 });
|
|
await insertFact({ entity_slug: 'cdx4-idempo-2', text: 'iterable', valid_from: t2 });
|
|
await insertFact({ entity_slug: 'cdx4-idempo-2', text: 'iterable', valid_from: t3 });
|
|
|
|
await runPhaseConsolidate(engine, {});
|
|
const before = await engine.executeRaw<{ id: number; valid_until: Date | null }>(
|
|
`SELECT id, valid_until FROM facts WHERE entity_slug = 'cdx4-idempo-2' ORDER BY valid_from ASC`,
|
|
);
|
|
|
|
// Reset consolidated_at to simulate extract_facts re-run.
|
|
await engine.executeRaw(
|
|
`UPDATE facts SET consolidated_at = NULL, consolidated_into = NULL
|
|
WHERE entity_slug = 'cdx4-idempo-2'`,
|
|
);
|
|
|
|
await runPhaseConsolidate(engine, {});
|
|
const after = await engine.executeRaw<{ id: number; valid_until: Date | null }>(
|
|
`SELECT id, valid_until FROM facts WHERE entity_slug = 'cdx4-idempo-2' ORDER BY valid_from ASC`,
|
|
);
|
|
// Same valid_until values; the IS DISTINCT FROM guard avoided rewrites.
|
|
expect(after.length).toBe(3);
|
|
for (let i = 0; i < before.length; i++) {
|
|
expect(after[i].id).toBe(before[i].id);
|
|
const a = after[i].valid_until ? new Date(after[i].valid_until!).toISOString() : null;
|
|
const b = before[i].valid_until ? new Date(before[i].valid_until!).toISOString() : null;
|
|
expect(a).toBe(b);
|
|
}
|
|
});
|
|
});
|