mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* Merge branch 'master' into garrytan/type-taxonomy-unification Resolve VERSION, package.json, CHANGELOG conflicts with v0.41.22.0 on top, preserving master's v0.41.19.0 entry below. * feat: v0.41.22.0 type-unification cathedral — collapse 94 types to 15 (closes #1479) Ships gbrain-base-v2 as the new install default (15 canonical types: 14 + note catch-all) and the unify-types PROTECTED Minion handler that runs the gbrain-base→v2 migration end-to-end on existing brains. What this delivers: - gbrain-base-v2.yaml standalone schema pack (no extends:) with 14 canonical page_types + 9 cluster mapping_rules + catch-all sentinel - 3 new schema-pack primitives: runRetypeCore (chunked UPDATE with legacy_type stamping), runPageToLinkCore (edge-shaped pages → link rows), runPageToAliasCore (concept-redirect → slug_aliases) - rewriteLinksBatch for N-pair atomic FK rewrite - Migration v104 slug_aliases table (forward-bootstrap probed on both engines for safe upgrade chain) - New engine method resolveSlugWithAlias(slug, sourceOrSources) on both Postgres + PGLite with multi-source ambiguity warning - inferTypeAndSubtypeFromPack overload + subtypes: + mapping_rules: + migration_from: schema-pack manifest extensions - findPackSuccessors version-range walker (1.x / 1.0.x / exact match) - expandTypeFilter for --type back-compat (D14): legacy aliases route through mapping_rules → canonical+subtype before the SQL filter fires - 3 new onboard checks: pack_upgrade_available, type_proliferation, dangling_aliases (source-scoped per F12) - unify-types Minion handler (PROTECTED, manual_only via render.ts allowlist per D17): retype-explicit → retype-catch-all → page-to-link → page-to-alias → final sync → active-pack flip - alias_resolved 1.05x post-fusion search boost stage; KNOBS_HASH_VERSION bumped 5→6 (one-time cache miss on upgrade, self-healing in TTL) - ELIGIBLE_TYPES for facts extraction extended with v2 canonicals (codex F-ELIGIBLE: blocker not v0.43 follow-up) Tests: 79 new unit/integration cases + 3 E2E cases covering all 9 production clusters end-to-end. 124-case verification on the cache-key + build-llms fixes. KNOBS_HASH_VERSION assertions updated in 3 tests. Plan: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md (16 locked decisions D1-D17, 12 baseline fixes F7-F21 absorbed from codex outside voice). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: CI verify failures — system-of-record allow-comment + schema-unify manifest registration Two CI failures on PR #1542: 1. check:system-of-record flagged page-to-link.ts:207 addLinksBatch as a direct write to a derived table. The call IS the reconcile surface for page_to_link mapping_rules — it converts edge-shaped pages into canonical link rows under the PROTECTED unify-types Minion handler, source-scoped, atomic per-rule. Added the canonical `// gbrain-allow-direct-insert: <reason>` comment on the same line. 2. check:resolver emitted 11 orphan_trigger warnings for `schema-unify` because the skill was added to skills/RESOLVER.md without a corresponding entry in skills/manifest.json. Added the registration under the existing skills[] array. bun run verify: 28/28 checks pass locally. * fix: CI test failures — schema-unify conformance + eligibility regression Six test failures across shards 2 + 10 on PR #1542: 1. resolver.test.ts: round-trip parser requires frontmatter triggers to be quoted (`- "..."` or `- '...'`). schema-unify shipped with bare YAML strings; quoted the 10 triggers to round-trip correctly. 2. skills-conformance.test.ts (×3): schema-unify SKILL.md was missing the required Contract, Anti-Patterns, and Output Format sections that every conformant skill must declare. Added all three: - Contract: inputs / outputs / side effects / failure modes - Anti-Patterns: 5 DON'Ts including the autopilot trust boundary - Output Format: per-phase stderr lines + celebration summary + JSON envelope shape 3. facts-eligibility.test.ts (×2): the v0.41.22 ELIGIBLE_TYPES expansion added `concept` to the eligible list, but the existing test suite pins concept as rejected (it's `extractable: true` in the schema pack but the v0.41.11 contract documented this as "cosmetic on the backstop path because backstop uses hardcoded ELIGIBLE_TYPES"). Removed `concept` from the expansion; other v2 canonicals (media, tweet, atom, analysis) stay. Comment updated to document the deliberate omission. All 6 failing tests now pass locally (370/370 across the 3 affected files). bun run verify: 28/28 checks green. * fix: harden findPackSuccessors test against shard pollution CI shard 8 reported 1 fail (1.00ms — too fast for any real loadActivePack file I/O) on `finds gbrain-base-v2 as successor of gbrain-base@1.0.0`. Local triple-run passes 9/9 in isolation. Root cause: the existing afterEach reset clears the module-level pack cache AFTER each test, but the FIRST test in the file inherits whatever state sibling files in the same bun shard process left behind. With 24+ schema-pack tests in shard 8 (mutate, mutate-audit, best-effort, registry-reload, manifest-v041_2, etc.) running before this file, the first test can read a poisoned cache. Fix: add `beforeEach(_resetPackCacheForTests)`. Two-sided reset guarantees clean state regardless of file ordering within the shard. bun run verify: 28/28 checks pass. * fix: quarantine two flaky tests to serial runner CI shard 1 + shard 8 each surfaced one intermittent failure: shard 1: buildBrainTools > execute() on put_page with valid namespace shard 8: findPackSuccessors > finds gbrain-base-v2 as successor Both pass cleanly in isolation. Both are concurrency races against shared in-shard state: - brain-allowlist.test.ts shares a singleton PGLiteEngine across 18 tests with a beforeEach DELETE FROM pages. With max-concurrency=4, two put_page tests can interleave their TRUNCATE + write phases, so the auto-link/extract sub-steps inside put_page race against the sibling test's DELETE. - schema-pack-find-pack-successors.test.ts reads bundled YAML packs via loadActivePack. The module-level pack cache is shared across parallel tests in the same shard; the previous beforeEach reset helped but didn't fully isolate against concurrent file reads under CI load. Fix per CLAUDE.md test-isolation lint rule R2 (concurrency-fragile files belong in the .serial.test.ts quarantine): rename both files to *.serial.test.ts. Serial runner picks them up at max-concurrency=1. 49/49 serial files pass locally. 28/28 verify checks pass. * fix: quarantine embed-stale test to serial runner CI shard 9 reported 6 failures, all from the embedStaleForSource describe block, all ~120-150ms each — classic shared-engine concurrency race shape. Passes 7/7 locally in isolation. Root cause: embed-stale.test.ts shares a singleton PGLiteEngine across 7 tests with beforeEach resetPgliteState. Under bun's max-concurrency=4 in the parallel shard, two tests can interleave their TRUNCATE + seedPage + upsertChunks + embedStaleForSource flow, so one test's stale-chunk count sees another test's mid-flight writes. Same fix as brain-allowlist.serial.test.ts and schema-pack-find-pack-successors.serial.test.ts: rename to *.serial.test.ts so the serial runner picks it up at max-concurrency=1. bun run verify: 28/28 checks pass. 7/7 embed-stale tests pass via serial. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
231 lines
7.6 KiB
TypeScript
231 lines
7.6 KiB
TypeScript
/**
|
|
* Tests for src/core/embed-stale.ts (v0.40 D15.2).
|
|
*
|
|
* Hermetic — uses an injected `embedFn` so no network call lands. Validates:
|
|
* - empty stale set → done:true, embedded:0
|
|
* - multi-batch run → embed every stale chunk, advance cursor correctly
|
|
* - kill mid-flight (signal.aborted) → aborted:true, partial progress preserved
|
|
* - resume from cursor → picks up where prior call left off (DB predicate)
|
|
* - per-page embedFn throw → logged + skipped, NOT propagated; chunks stay NULL
|
|
*
|
|
* Why PGLite: validates the engine.listStaleChunks/getChunks/upsertChunks
|
|
* roundtrip the helper depends on, not just the loop control flow.
|
|
*/
|
|
import { describe, test, expect, beforeAll, afterAll, beforeEach } from 'bun:test';
|
|
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
|
import { resetPgliteState } from './helpers/reset-pglite.ts';
|
|
import { embedStaleForSource } from '../src/core/embed-stale.ts';
|
|
import type { ChunkInput } from '../src/core/types.ts';
|
|
|
|
let engine: PGLiteEngine;
|
|
|
|
beforeAll(async () => {
|
|
engine = new PGLiteEngine();
|
|
await engine.connect({});
|
|
await engine.initSchema();
|
|
}, 30000);
|
|
|
|
afterAll(async () => {
|
|
await engine.disconnect();
|
|
});
|
|
|
|
beforeEach(async () => {
|
|
await resetPgliteState(engine);
|
|
});
|
|
|
|
/** Seed a page with N stale chunks (no embedding) into the default source. */
|
|
async function seedPageWithStaleChunks(slug: string, chunkCount: number): Promise<void> {
|
|
await engine.putPage(slug, {
|
|
type: 'note',
|
|
title: slug,
|
|
compiled_truth: `# ${slug}\n\nseeded`,
|
|
});
|
|
const chunks: ChunkInput[] = Array.from({ length: chunkCount }, (_, i) => ({
|
|
chunk_index: i,
|
|
chunk_text: `chunk ${i} of ${slug}`,
|
|
chunk_source: 'compiled_truth',
|
|
token_count: 4,
|
|
embedding: undefined, // NULL = stale
|
|
}));
|
|
await engine.upsertChunks(slug, chunks);
|
|
}
|
|
|
|
/** Deterministic fake embedder — returns unit-length 1536-dim vectors with
|
|
* first dim = text length, so we can assert specific chunks got embedded. */
|
|
function fakeEmbedFn(texts: string[]): Promise<Float32Array[]> {
|
|
return Promise.resolve(
|
|
texts.map((t) => {
|
|
const v = new Float32Array(1536);
|
|
v[0] = t.length;
|
|
v[1] = 1;
|
|
return v;
|
|
}),
|
|
);
|
|
}
|
|
|
|
describe('embedStaleForSource', () => {
|
|
test('empty stale set returns done:true with zero embedded', async () => {
|
|
const result = await embedStaleForSource(engine, 'default', {
|
|
embedFn: fakeEmbedFn,
|
|
});
|
|
expect(result).toEqual({
|
|
embedded: 0,
|
|
chunksProcessed: 0,
|
|
pagesProcessed: 0,
|
|
lastCursor: null,
|
|
done: true,
|
|
aborted: false,
|
|
});
|
|
});
|
|
|
|
test('embeds every stale chunk across multiple pages in one call', async () => {
|
|
await seedPageWithStaleChunks('a', 5);
|
|
await seedPageWithStaleChunks('b', 3);
|
|
|
|
const result = await embedStaleForSource(engine, 'default', {
|
|
embedFn: fakeEmbedFn,
|
|
});
|
|
expect(result.done).toBe(true);
|
|
expect(result.aborted).toBe(false);
|
|
expect(result.embedded).toBe(8);
|
|
expect(result.pagesProcessed).toBe(2);
|
|
|
|
// Verify DB: zero stale remaining for default.
|
|
const stale = await engine.countStaleChunks({ sourceId: 'default' });
|
|
expect(stale).toBe(0);
|
|
});
|
|
|
|
test('respects batchSize for cursor pagination', async () => {
|
|
await seedPageWithStaleChunks('a', 3);
|
|
await seedPageWithStaleChunks('b', 3);
|
|
let batchCount = 0;
|
|
const result = await embedStaleForSource(engine, 'default', {
|
|
embedFn: fakeEmbedFn,
|
|
batchSize: 2,
|
|
onProgress: () => {
|
|
batchCount++;
|
|
},
|
|
});
|
|
expect(result.embedded).toBe(6);
|
|
// 2-chunk batches across 6 stale rows = at least 3 progress callbacks.
|
|
expect(batchCount).toBeGreaterThanOrEqual(3);
|
|
});
|
|
|
|
test('IRON-RULE: aborted mid-flight → aborted:true, partial progress preserved', async () => {
|
|
await seedPageWithStaleChunks('a', 4);
|
|
await seedPageWithStaleChunks('b', 4);
|
|
await seedPageWithStaleChunks('c', 4);
|
|
const controller = new AbortController();
|
|
// Batch size 4 = one page per batch. concurrency 1 = serialize keys.
|
|
// Abort fires inside embedFn for page 'b', so 'a' lands, 'b' aborts mid-call,
|
|
// and the third batch ('c') never starts.
|
|
const result = await embedStaleForSource(engine, 'default', {
|
|
batchSize: 4,
|
|
concurrency: 1,
|
|
signal: controller.signal,
|
|
embedFn: async (texts) => {
|
|
if (texts.some((t) => t.includes(' of b'))) {
|
|
controller.abort();
|
|
throw new Error('aborted'); // simulates HTTP abort throw
|
|
}
|
|
return fakeEmbedFn(texts);
|
|
},
|
|
});
|
|
expect(result.aborted).toBe(true);
|
|
expect(result.done).toBe(false);
|
|
expect(result.embedded).toBe(4); // only 'a' landed
|
|
// 'b' and 'c' (8 chunks) remain stale
|
|
const stale = await engine.countStaleChunks({ sourceId: 'default' });
|
|
expect(stale).toBe(8);
|
|
});
|
|
|
|
test('IRON-RULE: kill + resume — second call picks up via embedding-IS-NULL predicate', async () => {
|
|
await seedPageWithStaleChunks('a', 4);
|
|
await seedPageWithStaleChunks('b', 4);
|
|
|
|
// First call aborts when 'b' is reached
|
|
const controller = new AbortController();
|
|
const first = await embedStaleForSource(engine, 'default', {
|
|
batchSize: 4,
|
|
concurrency: 1,
|
|
signal: controller.signal,
|
|
embedFn: async (texts) => {
|
|
if (texts.some((t) => t.includes(' of b'))) {
|
|
controller.abort();
|
|
throw new Error('aborted');
|
|
}
|
|
return fakeEmbedFn(texts);
|
|
},
|
|
});
|
|
expect(first.aborted).toBe(true);
|
|
expect(first.embedded).toBe(4); // 'a' landed
|
|
|
|
// Second call with NO cursor — predicate excludes already-embedded chunks
|
|
const second = await embedStaleForSource(engine, 'default', {
|
|
embedFn: fakeEmbedFn,
|
|
});
|
|
expect(second.done).toBe(true);
|
|
expect(first.embedded + second.embedded).toBe(8);
|
|
|
|
const stale = await engine.countStaleChunks({ sourceId: 'default' });
|
|
expect(stale).toBe(0);
|
|
});
|
|
|
|
test('per-page embedFn throw is logged but does NOT propagate', async () => {
|
|
await seedPageWithStaleChunks('good', 2);
|
|
await seedPageWithStaleChunks('bad', 2);
|
|
|
|
let badCount = 0;
|
|
const result = await embedStaleForSource(engine, 'default', {
|
|
embedFn: async (texts) => {
|
|
if (texts.some((t) => t.includes('bad'))) {
|
|
badCount++;
|
|
throw new Error('intentional embed failure');
|
|
}
|
|
return fakeEmbedFn(texts);
|
|
},
|
|
});
|
|
|
|
// The helper itself didn't throw
|
|
expect(result.done).toBe(true);
|
|
expect(badCount).toBe(1);
|
|
|
|
// 'good' chunks got embedded; 'bad' chunks stayed NULL
|
|
expect(result.embedded).toBe(2);
|
|
const stale = await engine.countStaleChunks({ sourceId: 'default' });
|
|
expect(stale).toBe(2);
|
|
});
|
|
|
|
test('source-scoped: does not touch other sources', async () => {
|
|
await engine.executeRaw(
|
|
`INSERT INTO sources (id, name, config) VALUES ('other', 'other', '{"federated":true}'::jsonb) ON CONFLICT (id) DO NOTHING`,
|
|
);
|
|
await seedPageWithStaleChunks('a', 3);
|
|
await engine.putPage('b', {
|
|
type: 'note',
|
|
title: 'b',
|
|
compiled_truth: '# b\n\nseeded',
|
|
}, { sourceId: 'other' });
|
|
await engine.upsertChunks(
|
|
'b',
|
|
Array.from({ length: 3 }, (_, i) => ({
|
|
chunk_index: i,
|
|
chunk_text: `other ${i}`,
|
|
chunk_source: 'compiled_truth',
|
|
token_count: 4,
|
|
embedding: undefined,
|
|
})),
|
|
{ sourceId: 'other' },
|
|
);
|
|
|
|
const result = await embedStaleForSource(engine, 'default', {
|
|
embedFn: fakeEmbedFn,
|
|
});
|
|
expect(result.embedded).toBe(3);
|
|
|
|
// 'other' source still has 3 stale chunks
|
|
const otherStale = await engine.countStaleChunks({ sourceId: 'other' });
|
|
expect(otherStale).toBe(3);
|
|
});
|
|
});
|