Files
gbrain/test/query-cache.test.ts
T
cdba533a04 v0.36.2.0 feat: ZeroEntropy as default + zero-based README rewrite (#1136)
* feat(dims): OpenAI text-embedding-3 Matryoshka range validation (D13)

dimsProviderOptions now fail-loud at the embed boundary when the
configured embedding_dimensions is outside the model's native range
(1..1536 for -small, 1..3072 for -large). Paste-ready fix hint in the
AIConfigError.fix field. Closes the silent-HTTP-400 path that would
have bit OpenAI-fallback users on v0.36.0.0 ZE-default installs.

16 new test cases in test/ai/dims-openai.test.ts pinning the contract
across native-openai and openai-compatible adapter paths.

* feat(ai): flip defaults to ZeroEntropy zembed-1 1280d + zerank-2 reranker

Default embedding model is now zeroentropyai:zembed-1 at 1280d via
Matryoshka. Real-corpus benchmark: 2.2x faster than OpenAI, 2.6x
cheaper at regular pricing, wins 11/20 head-to-head queries.

1280 is the closest valid ZE Matryoshka step to the prior OpenAI 1536d
default (valid set: 2560/1280/640/320/160/80/40). 1024 (Voyage's step)
is NOT on ZE's list — pinned by AIConfigError fail-loud in dims.ts.

balanced mode bundle now defaults reranker_enabled=true. zerank-2
reshuffles 60% of top-1 results in benchmarks. Missing-key fail-open
contract in src/core/search/rerank.ts handles unauthenticated cases.
Opt out with: gbrain config set search.reranker.enabled false

Existing tests updated (gateway.test.ts, search-mode.test.ts) and a
new test/balanced-reranker-default.test.ts (10 cases) pins the fail-
open invariants.

* feat(retrieval-upgrade): RetrievalUpgradePlanner + interactive prompt UX

New src/core/retrieval-upgrade-planner.ts is the consolidated planner
that computes the brain's pending retrieval-upgrade work (chunker
bumps + ZE switch) in one pass and applies the schema transition +
config updates atomically.

Tagged-union ApplyResult enum (D15): 'applied' | 'skipped_already_
applied' | 'skipped_no_work' | 'declined' | 'planned' | 'failed'.
No string-parsing reasons.

Three config keys (D12): ze_switch_prompt_shown (UI state),
ze_switch_requested (user intent), ze_switch_applied (work done).
Plus ze_switch_previous_snapshot (JSON, full prior config for --undo
per D16) and ze_switch_declined_at (90-day re-ask window).

Schema transition (D18) is atomic: DROP indexes + ALTER COLUMN +
CREATE INDEX inside a single engine.transaction(). HNSW recreation
is part of the same transaction — no silent slow-search window.

C3 eligibility logic: ze_switch_offered iff NOT on ZE + NOT declined
recently + NOT applied + (legacy default OR >100 pages).

C4 cost math: MAX(chunker_pending, dim_pending) not SUM — one
re-embed pass invalidates both surfaces simultaneously.

New src/core/retrieval-upgrade-prompt.ts wires the planner to a
TTY-only interactive prompt with two-line cost split (D10) and
privacy callout for the reranker flip.

Tests: test/retrieval-upgrade-planner.test.ts (24 cases) pins the
state machine. test/asymmetric-encoding-contract.test.ts (6 cases)
pins D17: search read path uses gateway.embedQuery() not embed(),
asserted via __setEmbedTransportForTests mock.

* feat(cli): gbrain ze-switch — manual lever for the ZE switch

New gbrain ze-switch CLI with --dry-run, --json, --resume, --force,
--undo, --non-interactive, --confirm-reembed, --ignore-missing-key
flags. Mirrors the upgrade prompt's UX symmetry: --undo presents a
cost-warning before re-embedding back to the prior width.

src/cli.ts: dispatch case + CLI_ONLY entry. ze-switch owns its own
engine lifecycle (mirrors the doctor pattern).

test/ze-switch-cli.test.ts (11 cases): --help, --dry-run, --json,
--non-interactive, --ignore-missing-key, --resume, --undo,
--confirm-reembed. Uses captureExit harness to test process.exit()
paths without breaking the test process.

* feat(doctor): ze_embedding_health + embedding_width_consistency checks

Two new doctor checks (D-A5):

ze_embedding_health: when embedding_model starts with zeroentropyai:,
verify ZEROENTROPY_API_KEY is set (env or config). Paste-ready setup
hint with the signup URL on failure.

embedding_width_consistency: cross-check that the configured
embedding_dimensions matches the actual vector(N) column width on
content_chunks.embedding. Catches the half-applied switch state
(schema migrated but config write crashed) with a paste-ready
gbrain ze-switch --resume hint.

Wired into runDoctor between reranker_health and the existing
sync_freshness checks. Both checks gracefully no-op on non-ZE
embedding configs.

test/doctor-ze-checks.test.ts (8 cases) pins both checks across
happy + missing-key + missing-config + drift paths. Uses withEnv()
helper to clear ZEROENTROPY_API_KEY for the no-key path so tests
are hermetic against contributor env state.

test/e2e/v0_28_5-fix-wave.test.ts + test/openai-compat-multimodal.test.ts:
updated to explicit-configure the gateway when the test depends on
specific dims that diverge from the v0.36.0.0 default (1280d).

* docs: README zero-based rewrite (884 -> 139 lines) + new docs files

Strip 4 months of accreted "New in v0.X.Y" hero blocks and reorganize
around what gbrain does today. 33 H2s -> 8. The Commands section
(136 lines duplicating gbrain --help) moved out; the 6-table skills
enumeration collapsed to a one-paragraph capability description with
a link to skills/RESOLVER.md.

Hero retains load-bearing facts: OpenClaw + Hermes credit, production
numbers (17,888 pages / 4,383 people / 723 companies), BrainBench
numbers (P@5 49.1% / R@5 97.9% / +31.4 lift), ZE comparison numbers,
30-min install claim. Adds one paragraph announcing the v0.36.0.0 ZE
default with the explicit gbrain config set escape for OpenAI/Voyage
users.

New files:
- docs/INSTALL.md: every install path consolidated (agent platform,
  CLI standalone, MCP server). Thin-client mode covered.
- docs/architecture/RETRIEVAL.md: why the hybrid + graph stack works.
  BrainBench numbers, why each strategy alone fails, the source-aware
  ranking + intent classification + multi-query expansion story.
- docs/ethos/ORIGIN.md: origin story lifted from the old README so
  the front door stays factual + concrete.

test/readme-hero-anchors.test.ts (5 cases) is the D9 regression
guard. Five load-bearing strings: OpenClaw, Hermes, ZE,
production-numbers regex, P@5/R@5. Light anchors that let voice/
structure evolve but block accidental loss of headline facts.

scripts/check-test-real-names.sh: allowlist entries for OpenClaw +
Hermes literals in the anchor test (it explicitly asserts those
strings appear in README).

* chore: bump version and changelog (v0.36.0.0)

ZeroEntropy as the new default for embedding (zembed-1 at 1280d via
Matryoshka) and reranker (zerank-2 cross-encoder, on by default in
balanced mode bundle). README zero-based rewrite (884 -> 139 lines).
3 new docs files. Two new doctor checks. New gbrain ze-switch CLI
with --undo for symmetric reversibility.

skills/migrations/v0.36.0.0.md tells the agent how to surface the
retrieval-upgrade prompt post-upgrade.

llms-full.txt regenerated via bun run build:llms.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(docs): scrub Wintermute from RETRIEVAL.md per privacy rule

* chore: rebump version 0.36.0.0 → 0.36.2.0 (queue collision)

Three open PRs were claiming v0.36.0.0 (#1130 skillpack, #1139
hindsight, #1136 this PR). Ship-aware queue allocator says this
branch lands at v0.36.2.0.

Trio audit:
  VERSION       0.36.2.0
  package.json  0.36.2.0
  CHANGELOG     ## [0.36.2.0] - 2026-05-17

Updates: VERSION, package.json, CHANGELOG header + body refs,
README "New default in v0.36.2.0" announcement + credit line,
skills/migrations/v0.36.0.0.md renamed to v0.36.2.0.md with
frontmatter + body refs updated. llms-full.txt regenerated.

* fix(test): pin gateway dim=1536 in cross-file-stateful PGLite tests

CI shard 1 reported 10 failures across `query-cache.test.ts` (6) and
`consolidate-valid-until.test.ts` (4). Both files hardcode 1536-dim
vectors but rely on `PGLiteEngine.initSchema()` to size
`vector(__EMBEDDING_DIMS__)` at the right width.

Root cause: v0.36.2.0 flipped DEFAULT_EMBEDDING_DIMENSIONS from 1536
to 1280 (ZE Matryoshka step). The gateway module is process-singleton;
when ANOTHER test file in the same shard's bun-test process configures
the gateway before us, `pglite-engine.ts:216` reads
`getEmbeddingDimensions() === 1280` and sizes the schema columns at
vector(1280). The hardcoded 1536-dim INSERTs then fail with
"expected 1280 dimensions, not 1536".

Locally these tests pass in isolation because the gateway falls back
through the try/catch at pglite-engine.ts:218 (1536 default). CI runs
multiple test files in one process, so cross-file state poisons the
schema width.

Fix: explicit `resetGateway()` + `configureGateway({embedding_dimensions:
1536, ...})` at the top of `beforeAll`, plus `resetGateway()` in
`afterAll`. Pins the schema width regardless of cross-file state.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-18 21:11:02 -07:00

261 lines
9.3 KiB
TypeScript

/**
* v0.32.x search-lite \u2014 semantic query cache.
*
* PGLite-backed test. Confirms:
* - migration v51 creates the query_cache table
* - store + lookup roundtrip with EXACT same embedding \u2192 hit
* - lookup with a similar embedding (cosine > 0.92) \u2192 hit
* - lookup with a far embedding \u2192 miss
* - TTL expiration: a stale row is skipped at read time
* - clear / prune / stats work as advertised
* - source_id isolation: brain A's cache doesn't leak to brain B
* - disabled cache is a pure no-op
*
* Uses synthetic Float32Array embeddings so the test doesn't depend on
* any external embedding provider.
*/
import { describe, test, expect, beforeAll, afterAll, beforeEach } from 'bun:test';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { SemanticQueryCache, cacheRowId } from '../src/core/search/query-cache.ts';
import { configureGateway, resetGateway } from '../src/core/ai/gateway.ts';
import type { SearchResult, HybridSearchMeta } from '../src/core/types.ts';
let engine: PGLiteEngine;
// Build a stable, normalized embedding. PGLite ships pgvector with 1536-dim
// support (the default); a smaller test dim won't match the column. We
// truncate / pad to 1536 to match the migration's resolved dim.
const DIM = 1536;
function makeEmbedding(seed: number, dim = DIM): Float32Array {
const e = new Float32Array(dim);
// Simple deterministic generator with a unique fingerprint per seed
// so similar seeds produce similar (cosine > 0.95) vectors and distinct
// seeds produce orthogonal-ish ones.
for (let i = 0; i < dim; i++) {
e[i] = Math.sin(seed * 0.001 + i * 0.01);
}
// L2-normalize so cosine = dot product.
let mag = 0;
for (let i = 0; i < dim; i++) mag += e[i] * e[i];
mag = Math.sqrt(mag);
if (mag > 0) for (let i = 0; i < dim; i++) e[i] /= mag;
return e;
}
function makeOrthogonalEmbedding(seed: number, dim = DIM): Float32Array {
// Use a totally different basis so cosine is near-zero.
const e = new Float32Array(dim);
for (let i = 0; i < dim; i++) {
e[i] = Math.cos(seed * 13.7 + i * 0.97);
}
let mag = 0;
for (let i = 0; i < dim; i++) mag += e[i] * e[i];
mag = Math.sqrt(mag);
if (mag > 0) for (let i = 0; i < dim; i++) e[i] /= mag;
return e;
}
function makeResult(slug: string): SearchResult {
return {
slug,
page_id: 1,
title: `Title for ${slug}`,
type: 'concept',
chunk_text: `chunk text for ${slug}`,
chunk_source: 'compiled_truth',
chunk_id: 1,
chunk_index: 0,
score: 1.0,
stale: false,
};
}
const META: HybridSearchMeta = {
vector_enabled: true,
detail_resolved: 'medium',
expansion_applied: false,
intent: 'general',
};
beforeAll(async () => {
// v0.36.2.0: DEFAULT_EMBEDDING_DIMENSIONS flipped to 1280 (ZE Matryoshka).
// This test hardcodes DIM=1536 in its embeddings. If another test file in
// the same shard configured the gateway before us, initSchema() would size
// query_cache.embedding at vector(1280) and every insert below would fail
// with "expected 1280 dimensions, not 1536". Pin the gateway to 1536d
// explicitly so this file is hermetic regardless of cross-file state.
resetGateway();
configureGateway({
embedding_model: 'openai:text-embedding-3-large',
embedding_dimensions: 1536,
env: { OPENAI_API_KEY: 'sk-fake' },
});
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
try { await engine.disconnect(); } catch { /* ignore */ }
resetGateway();
});
beforeEach(async () => {
// Wipe the cache between tests so ordering doesn't matter.
await engine.executeRaw(`DELETE FROM query_cache`);
});
describe('migration v51 \u2014 query_cache table exists', () => {
test('table is present and has expected columns', async () => {
const rows = await engine.executeRaw<{ column_name: string }>(
`SELECT column_name FROM information_schema.columns
WHERE table_name = 'query_cache'`,
);
const names = rows.map(r => r.column_name);
expect(names).toContain('id');
expect(names).toContain('query_text');
expect(names).toContain('source_id');
expect(names).toContain('embedding');
expect(names).toContain('results');
expect(names).toContain('meta');
expect(names).toContain('ttl_seconds');
expect(names).toContain('created_at');
expect(names).toContain('hit_count');
});
});
describe('cacheRowId', () => {
test('is deterministic across same input', () => {
expect(cacheRowId('hello', 'default')).toBe(cacheRowId('hello', 'default'));
});
test('differs across source_id', () => {
expect(cacheRowId('hello', 'a')).not.toBe(cacheRowId('hello', 'b'));
});
});
describe('SemanticQueryCache \u2014 store + lookup', () => {
test('roundtrip: exact embedding match returns a hit', async () => {
const cache = new SemanticQueryCache(engine);
const emb = makeEmbedding(1);
const results = [makeResult('a'), makeResult('b')];
await cache.store('what is foo', emb, results, META);
const hit = await cache.lookup(emb);
expect(hit.hit).toBe(true);
expect(hit.results).toHaveLength(2);
expect(hit.results?.[0].slug).toBe('a');
expect(hit.similarity).toBeGreaterThan(0.99);
});
test('similar embedding (cosine > 0.92) is a hit', async () => {
const cache = new SemanticQueryCache(engine);
const base = makeEmbedding(100);
// Construct a near-neighbor: tweak a few dims so cosine stays > 0.92.
const near = new Float32Array(base);
for (let i = 0; i < 10; i++) near[i] += 0.005;
// Re-normalize.
let mag = 0;
for (let i = 0; i < DIM; i++) mag += near[i] * near[i];
mag = Math.sqrt(mag);
for (let i = 0; i < DIM; i++) near[i] /= mag;
await cache.store('what is foo', base, [makeResult('a')], META);
const hit = await cache.lookup(near);
expect(hit.hit).toBe(true);
expect(hit.similarity).toBeGreaterThan(0.92);
});
test('orthogonal embedding is a miss', async () => {
const cache = new SemanticQueryCache(engine);
const a = makeEmbedding(1);
const b = makeOrthogonalEmbedding(2);
await cache.store('q1', a, [makeResult('a')], META);
const hit = await cache.lookup(b);
expect(hit.hit).toBe(false);
});
});
describe('SemanticQueryCache \u2014 TTL', () => {
test('stale row (past TTL) is not returned', async () => {
const cache = new SemanticQueryCache(engine, { ttlSeconds: 1 });
const emb = makeEmbedding(42);
await cache.store('q', emb, [makeResult('a')], META, { ttlSeconds: 1 });
// Manually rewind created_at to simulate expiration.
await engine.executeRaw(
`UPDATE query_cache SET created_at = now() - interval '10 seconds'`,
);
const hit = await cache.lookup(emb);
expect(hit.hit).toBe(false);
});
});
describe('SemanticQueryCache \u2014 source isolation', () => {
test('different source_id cannot read each other\u2019s rows', async () => {
const cache = new SemanticQueryCache(engine);
const emb = makeEmbedding(7);
await cache.store('q', emb, [makeResult('a')], META, { sourceId: 'src-A' });
const hitB = await cache.lookup(emb, { sourceId: 'src-B' });
expect(hitB.hit).toBe(false);
const hitA = await cache.lookup(emb, { sourceId: 'src-A' });
expect(hitA.hit).toBe(true);
});
});
describe('SemanticQueryCache \u2014 management', () => {
test('clear() wipes all rows', async () => {
const cache = new SemanticQueryCache(engine);
const emb = makeEmbedding(9);
await cache.store('q1', emb, [makeResult('a')], META);
await cache.store('q2', makeEmbedding(10), [makeResult('b')], META);
const removed = await cache.clear();
expect(removed).toBeGreaterThanOrEqual(2);
const stats = await cache.stats();
expect(stats.total_rows).toBe(0);
});
test('prune() deletes only stale rows', async () => {
const cache = new SemanticQueryCache(engine);
await cache.store('fresh', makeEmbedding(11), [makeResult('a')], META);
await cache.store('stale', makeEmbedding(12), [makeResult('b')], META, { ttlSeconds: 1 });
await engine.executeRaw(
`UPDATE query_cache SET created_at = now() - interval '10 seconds' WHERE query_text = 'stale'`,
);
const removed = await cache.prune();
expect(removed).toBe(1);
const stats = await cache.stats();
expect(stats.total_rows).toBe(1);
expect(stats.fresh_rows).toBe(1);
});
test('stats() reports fresh / stale / total / hit counters', async () => {
const cache = new SemanticQueryCache(engine);
const emb = makeEmbedding(13);
await cache.store('q', emb, [makeResult('a')], META);
await cache.lookup(emb); // bump hit
// Hit bump is async/fire-and-forget; give it a moment to land.
await new Promise(r => setTimeout(r, 50));
const stats = await cache.stats();
expect(stats.total_rows).toBe(1);
expect(stats.fresh_rows).toBe(1);
expect(stats.stale_rows).toBe(0);
expect(stats.total_hits).toBeGreaterThanOrEqual(1);
});
});
describe('SemanticQueryCache \u2014 disabled', () => {
test('disabled cache is a pure no-op on lookup', async () => {
const cache = new SemanticQueryCache(engine, { enabled: false });
const emb = makeEmbedding(99);
await cache.store('q', emb, [makeResult('a')], META);
// Even after a store call, lookup must miss because enabled=false.
const hit = await cache.lookup(emb);
expect(hit.hit).toBe(false);
});
});