mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* Merge branch 'master' into garrytan/type-taxonomy-unification Resolve VERSION, package.json, CHANGELOG conflicts with v0.41.22.0 on top, preserving master's v0.41.19.0 entry below. * feat: v0.41.22.0 type-unification cathedral — collapse 94 types to 15 (closes #1479) Ships gbrain-base-v2 as the new install default (15 canonical types: 14 + note catch-all) and the unify-types PROTECTED Minion handler that runs the gbrain-base→v2 migration end-to-end on existing brains. What this delivers: - gbrain-base-v2.yaml standalone schema pack (no extends:) with 14 canonical page_types + 9 cluster mapping_rules + catch-all sentinel - 3 new schema-pack primitives: runRetypeCore (chunked UPDATE with legacy_type stamping), runPageToLinkCore (edge-shaped pages → link rows), runPageToAliasCore (concept-redirect → slug_aliases) - rewriteLinksBatch for N-pair atomic FK rewrite - Migration v104 slug_aliases table (forward-bootstrap probed on both engines for safe upgrade chain) - New engine method resolveSlugWithAlias(slug, sourceOrSources) on both Postgres + PGLite with multi-source ambiguity warning - inferTypeAndSubtypeFromPack overload + subtypes: + mapping_rules: + migration_from: schema-pack manifest extensions - findPackSuccessors version-range walker (1.x / 1.0.x / exact match) - expandTypeFilter for --type back-compat (D14): legacy aliases route through mapping_rules → canonical+subtype before the SQL filter fires - 3 new onboard checks: pack_upgrade_available, type_proliferation, dangling_aliases (source-scoped per F12) - unify-types Minion handler (PROTECTED, manual_only via render.ts allowlist per D17): retype-explicit → retype-catch-all → page-to-link → page-to-alias → final sync → active-pack flip - alias_resolved 1.05x post-fusion search boost stage; KNOBS_HASH_VERSION bumped 5→6 (one-time cache miss on upgrade, self-healing in TTL) - ELIGIBLE_TYPES for facts extraction extended with v2 canonicals (codex F-ELIGIBLE: blocker not v0.43 follow-up) Tests: 79 new unit/integration cases + 3 E2E cases covering all 9 production clusters end-to-end. 124-case verification on the cache-key + build-llms fixes. KNOBS_HASH_VERSION assertions updated in 3 tests. Plan: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md (16 locked decisions D1-D17, 12 baseline fixes F7-F21 absorbed from codex outside voice). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: CI verify failures — system-of-record allow-comment + schema-unify manifest registration Two CI failures on PR #1542: 1. check:system-of-record flagged page-to-link.ts:207 addLinksBatch as a direct write to a derived table. The call IS the reconcile surface for page_to_link mapping_rules — it converts edge-shaped pages into canonical link rows under the PROTECTED unify-types Minion handler, source-scoped, atomic per-rule. Added the canonical `// gbrain-allow-direct-insert: <reason>` comment on the same line. 2. check:resolver emitted 11 orphan_trigger warnings for `schema-unify` because the skill was added to skills/RESOLVER.md without a corresponding entry in skills/manifest.json. Added the registration under the existing skills[] array. bun run verify: 28/28 checks pass locally. * fix: CI test failures — schema-unify conformance + eligibility regression Six test failures across shards 2 + 10 on PR #1542: 1. resolver.test.ts: round-trip parser requires frontmatter triggers to be quoted (`- "..."` or `- '...'`). schema-unify shipped with bare YAML strings; quoted the 10 triggers to round-trip correctly. 2. skills-conformance.test.ts (×3): schema-unify SKILL.md was missing the required Contract, Anti-Patterns, and Output Format sections that every conformant skill must declare. Added all three: - Contract: inputs / outputs / side effects / failure modes - Anti-Patterns: 5 DON'Ts including the autopilot trust boundary - Output Format: per-phase stderr lines + celebration summary + JSON envelope shape 3. facts-eligibility.test.ts (×2): the v0.41.22 ELIGIBLE_TYPES expansion added `concept` to the eligible list, but the existing test suite pins concept as rejected (it's `extractable: true` in the schema pack but the v0.41.11 contract documented this as "cosmetic on the backstop path because backstop uses hardcoded ELIGIBLE_TYPES"). Removed `concept` from the expansion; other v2 canonicals (media, tweet, atom, analysis) stay. Comment updated to document the deliberate omission. All 6 failing tests now pass locally (370/370 across the 3 affected files). bun run verify: 28/28 checks green. * fix: harden findPackSuccessors test against shard pollution CI shard 8 reported 1 fail (1.00ms — too fast for any real loadActivePack file I/O) on `finds gbrain-base-v2 as successor of gbrain-base@1.0.0`. Local triple-run passes 9/9 in isolation. Root cause: the existing afterEach reset clears the module-level pack cache AFTER each test, but the FIRST test in the file inherits whatever state sibling files in the same bun shard process left behind. With 24+ schema-pack tests in shard 8 (mutate, mutate-audit, best-effort, registry-reload, manifest-v041_2, etc.) running before this file, the first test can read a poisoned cache. Fix: add `beforeEach(_resetPackCacheForTests)`. Two-sided reset guarantees clean state regardless of file ordering within the shard. bun run verify: 28/28 checks pass. * fix: quarantine two flaky tests to serial runner CI shard 1 + shard 8 each surfaced one intermittent failure: shard 1: buildBrainTools > execute() on put_page with valid namespace shard 8: findPackSuccessors > finds gbrain-base-v2 as successor Both pass cleanly in isolation. Both are concurrency races against shared in-shard state: - brain-allowlist.test.ts shares a singleton PGLiteEngine across 18 tests with a beforeEach DELETE FROM pages. With max-concurrency=4, two put_page tests can interleave their TRUNCATE + write phases, so the auto-link/extract sub-steps inside put_page race against the sibling test's DELETE. - schema-pack-find-pack-successors.test.ts reads bundled YAML packs via loadActivePack. The module-level pack cache is shared across parallel tests in the same shard; the previous beforeEach reset helped but didn't fully isolate against concurrent file reads under CI load. Fix per CLAUDE.md test-isolation lint rule R2 (concurrency-fragile files belong in the .serial.test.ts quarantine): rename both files to *.serial.test.ts. Serial runner picks them up at max-concurrency=1. 49/49 serial files pass locally. 28/28 verify checks pass. * fix: quarantine embed-stale test to serial runner CI shard 9 reported 6 failures, all from the embedStaleForSource describe block, all ~120-150ms each — classic shared-engine concurrency race shape. Passes 7/7 locally in isolation. Root cause: embed-stale.test.ts shares a singleton PGLiteEngine across 7 tests with beforeEach resetPgliteState. Under bun's max-concurrency=4 in the parallel shard, two tests can interleave their TRUNCATE + seedPage + upsertChunks + embedStaleForSource flow, so one test's stale-chunk count sees another test's mid-flight writes. Same fix as brain-allowlist.serial.test.ts and schema-pack-find-pack-successors.serial.test.ts: rename to *.serial.test.ts so the serial runner picks it up at max-concurrency=1. bun run verify: 28/28 checks pass. 7/7 embed-stale tests pass via serial. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
239 lines
9.4 KiB
TypeScript
239 lines
9.4 KiB
TypeScript
// Commit 1 (Phase 1): cross-modal intent + hybrid routing + knobsHash + RRF.
|
|
//
|
|
// Covers:
|
|
// - suggestedModality regex matches (positive + negative + plural-safe)
|
|
// - isAmbiguousModalityQuery heuristic
|
|
// - SEARCH_MODE_CONFIG_KEYS registry includes new keys (D3)
|
|
// - knobsHash differs across cross-modal knob values (D2)
|
|
// - knobsHash version bumped to 3
|
|
// - MODE_BUNDLES carry cross-modal defaults
|
|
|
|
import { describe, expect, test } from 'bun:test';
|
|
import {
|
|
classifyQuery,
|
|
isAmbiguousModalityQuery,
|
|
type ModalityMode,
|
|
} from '../src/core/search/query-intent.ts';
|
|
import {
|
|
KNOBS_HASH_VERSION,
|
|
MODE_BUNDLES,
|
|
SEARCH_MODE_CONFIG_KEYS,
|
|
knobsHash,
|
|
resolveSearchMode,
|
|
type ResolvedSearchKnobs,
|
|
} from '../src/core/search/mode.ts';
|
|
|
|
describe('query-intent — suggestedModality regex (D6 + D14)', () => {
|
|
test('"show me photos from the hackathon" → image', () => {
|
|
expect(classifyQuery('show me photos from the hackathon').suggestedModality).toBe('image');
|
|
});
|
|
|
|
test('"what is founder mode?" → text (default)', () => {
|
|
expect(classifyQuery('what is founder mode?').suggestedModality).toBe('text');
|
|
});
|
|
|
|
const imagePhrasings: Array<[string, ModalityMode]> = [
|
|
['find images from last week', 'image'],
|
|
['find me images of acme', 'image'],
|
|
['what does the OG photo look like', 'image'],
|
|
['screenshot of the dashboard', 'image'],
|
|
['diagram of the architecture', 'image'],
|
|
['visuals showing the trends', 'image'],
|
|
['whiteboard from the offsite', 'image'],
|
|
['pictures of the team', 'image'],
|
|
['pull me the screenshots', 'image'],
|
|
];
|
|
|
|
for (const [query, expected] of imagePhrasings) {
|
|
test(`image phrasing: "${query}" → ${expected}`, () => {
|
|
expect(classifyQuery(query).suggestedModality).toBe(expected);
|
|
});
|
|
}
|
|
|
|
const textPhrasings = [
|
|
'who is acme corp',
|
|
'tell me about founder mode',
|
|
'what happened at the hackathon',
|
|
'meeting notes from yesterday',
|
|
'most recent take on AI',
|
|
];
|
|
|
|
for (const query of textPhrasings) {
|
|
test(`text phrasing: "${query}" → text`, () => {
|
|
expect(classifyQuery(query).suggestedModality).toBe('text');
|
|
});
|
|
}
|
|
});
|
|
|
|
describe('isAmbiguousModalityQuery (Commit 4 prep)', () => {
|
|
// Genuinely ambiguous = visual noun present + reference marker present BUT
|
|
// CROSS_MODAL_PATTERNS doesn't catch it (otherwise regex already classified
|
|
// confidently and the LLM call would be wasted).
|
|
|
|
test('"any picture during last week" → ambiguous', () => {
|
|
// "picture during" doesn't match (of|from|at|with|...) so CROSS_MODAL
|
|
// doesn't fire; "any pictures" does match the AMBIGUOUS_REFERENCE marker.
|
|
// Actually "any picture" matches /\b(any|some|...)\s+(pics?|photos?|images?...)/ — but
|
|
// the CROSS_MODAL pattern needs "pictures from/of/at/...". This phrasing
|
|
// has neither — so it's genuinely ambiguous.
|
|
expect(isAmbiguousModalityQuery('any picture during last week')).toBe(true);
|
|
});
|
|
|
|
test('"what is founder mode" → not ambiguous (plain text query)', () => {
|
|
expect(isAmbiguousModalityQuery('what is founder mode')).toBe(false);
|
|
});
|
|
|
|
test('"show me photos of acme" → not ambiguous (regex catches it)', () => {
|
|
// Already-confident classification, no LLM needed.
|
|
expect(isAmbiguousModalityQuery('show me photos of acme')).toBe(false);
|
|
});
|
|
|
|
test('"any pictures from the meeting" → not ambiguous (regex catches "pictures from")', () => {
|
|
// CROSS_MODAL fires on "pictures from" — confident classification.
|
|
expect(isAmbiguousModalityQuery('any pictures from the meeting')).toBe(false);
|
|
});
|
|
|
|
test('"chart" without article/determiner → not ambiguous (bare visual noun has no reference marker)', () => {
|
|
// No "any|some|that|the" determiner in front of the visual noun, and no
|
|
// "from last/this/the X" phrase — pure text query.
|
|
expect(isAmbiguousModalityQuery('chart')).toBe(false);
|
|
});
|
|
|
|
test('"the chart" alone → ambiguous (determiner+visual-noun is a real reference marker)', () => {
|
|
// "the chart" is the canonical ambiguous case — user references a
|
|
// specific visual asset without confirming they want image search.
|
|
// LLM tie-break decides.
|
|
expect(isAmbiguousModalityQuery('the chart')).toBe(true);
|
|
});
|
|
|
|
test('"the diagram in last week\'s deck" → ambiguous', () => {
|
|
// "diagram in" doesn't match CROSS_MODAL (of|from|about|showing only).
|
|
// "the diagram" matches AMBIGUOUS_REFERENCE first pattern.
|
|
expect(isAmbiguousModalityQuery("the diagram in last week's deck")).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('D3 — SEARCH_MODE_CONFIG_KEYS registry includes cross-modal keys', () => {
|
|
const expected = [
|
|
'search.cross_modal.both_mode_text_weight',
|
|
'search.cross_modal.both_mode_image_weight',
|
|
'search.image_query.text_refinement_weight',
|
|
'search.image_query.image_refinement_weight',
|
|
'search.unified_multimodal',
|
|
'search.unified_multimodal_only',
|
|
'search.cross_modal.llm_intent',
|
|
];
|
|
|
|
for (const key of expected) {
|
|
test(`registry contains ${key}`, () => {
|
|
expect(SEARCH_MODE_CONFIG_KEYS).toContain(key);
|
|
});
|
|
}
|
|
});
|
|
|
|
describe('D2 — knobsHash differs across cross-modal knob values', () => {
|
|
function baseKnobs(): ResolvedSearchKnobs {
|
|
return resolveSearchMode({ mode: 'balanced' });
|
|
}
|
|
|
|
test('KNOBS_HASH_VERSION is 6 (v=4 graph_signals + schema-pack; v=5 contextual_retrieval; v=6 alias_resolved; cross-modal still appended)', () => {
|
|
// v0.35 ladder: 1→2 reranker, 2→3 floor_ratio. v0.36 piggybacks on v=3
|
|
// with 7 cross-modal knobs + column/provider context. v0.40.4 (salem) +
|
|
// v0.39 T21 (master) bump to v=4 for graph_signals + schema-pack fields.
|
|
// v0.40.3.0 D8 bumps to v=5 (sequenced behind salem's v=4 graph-signals).
|
|
// v0.41.22.0 (type-unification): 5→6 for alias_resolved post-fusion boost.
|
|
expect(KNOBS_HASH_VERSION).toBe(6);
|
|
});
|
|
|
|
test('flipping unified_multimodal changes the hash', () => {
|
|
const k1 = baseKnobs();
|
|
const k2 = { ...k1, unified_multimodal: true };
|
|
expect(knobsHash(k1)).not.toBe(knobsHash(k2));
|
|
});
|
|
|
|
test('flipping unified_multimodal_only changes the hash', () => {
|
|
const k1 = baseKnobs();
|
|
const k2 = { ...k1, unified_multimodal_only: true };
|
|
expect(knobsHash(k1)).not.toBe(knobsHash(k2));
|
|
});
|
|
|
|
test('flipping cross_modal_llm_intent changes the hash', () => {
|
|
const k1 = baseKnobs();
|
|
const k2 = { ...k1, cross_modal_llm_intent: true };
|
|
expect(knobsHash(k1)).not.toBe(knobsHash(k2));
|
|
});
|
|
|
|
test('changing cross_modal_both_text_weight changes the hash', () => {
|
|
const k1 = baseKnobs();
|
|
const k2 = { ...k1, cross_modal_both_text_weight: 0.5 };
|
|
expect(knobsHash(k1)).not.toBe(knobsHash(k2));
|
|
});
|
|
|
|
test('changing image_query_text_refinement_weight changes the hash', () => {
|
|
const k1 = baseKnobs();
|
|
const k2 = { ...k1, image_query_text_refinement_weight: 0.7 };
|
|
expect(knobsHash(k1)).not.toBe(knobsHash(k2));
|
|
});
|
|
|
|
test('identical knobs produce identical hashes (regression sanity)', () => {
|
|
expect(knobsHash(baseKnobs())).toBe(knobsHash(baseKnobs()));
|
|
});
|
|
});
|
|
|
|
describe('D6 — MODE_BUNDLES carry cross-modal defaults', () => {
|
|
test('all three modes default cross_modal_both_text_weight to 0.6', () => {
|
|
expect(MODE_BUNDLES.conservative.cross_modal_both_text_weight).toBe(0.6);
|
|
expect(MODE_BUNDLES.balanced.cross_modal_both_text_weight).toBe(0.6);
|
|
expect(MODE_BUNDLES.tokenmax.cross_modal_both_text_weight).toBe(0.6);
|
|
});
|
|
|
|
test('all three modes default cross_modal_both_image_weight to 0.4', () => {
|
|
expect(MODE_BUNDLES.conservative.cross_modal_both_image_weight).toBe(0.4);
|
|
expect(MODE_BUNDLES.balanced.cross_modal_both_image_weight).toBe(0.4);
|
|
expect(MODE_BUNDLES.tokenmax.cross_modal_both_image_weight).toBe(0.4);
|
|
});
|
|
|
|
test('all three modes default image_query weights (D13: 0.4 text / 0.6 image)', () => {
|
|
expect(MODE_BUNDLES.conservative.image_query_text_refinement_weight).toBe(0.4);
|
|
expect(MODE_BUNDLES.conservative.image_query_image_refinement_weight).toBe(0.6);
|
|
expect(MODE_BUNDLES.tokenmax.image_query_image_refinement_weight).toBe(0.6);
|
|
});
|
|
|
|
test('all three modes default unified_multimodal to false (opt-in)', () => {
|
|
expect(MODE_BUNDLES.conservative.unified_multimodal).toBe(false);
|
|
expect(MODE_BUNDLES.balanced.unified_multimodal).toBe(false);
|
|
expect(MODE_BUNDLES.tokenmax.unified_multimodal).toBe(false);
|
|
});
|
|
|
|
test('all three modes default cross_modal_llm_intent to false (opt-in)', () => {
|
|
expect(MODE_BUNDLES.conservative.cross_modal_llm_intent).toBe(false);
|
|
expect(MODE_BUNDLES.balanced.cross_modal_llm_intent).toBe(false);
|
|
expect(MODE_BUNDLES.tokenmax.cross_modal_llm_intent).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe('resolveSearchMode threads cross-modal overrides', () => {
|
|
test('per-call override beats config override beats mode default', () => {
|
|
const k = resolveSearchMode({
|
|
mode: 'balanced',
|
|
overrides: { cross_modal_both_text_weight: 0.5 },
|
|
perCall: { cross_modal_both_text_weight: 0.8 },
|
|
});
|
|
expect(k.cross_modal_both_text_weight).toBe(0.8);
|
|
});
|
|
|
|
test('config override wins when no per-call override', () => {
|
|
const k = resolveSearchMode({
|
|
mode: 'balanced',
|
|
overrides: { unified_multimodal: true },
|
|
});
|
|
expect(k.unified_multimodal).toBe(true);
|
|
});
|
|
|
|
test('mode default fires when neither override is set', () => {
|
|
const k = resolveSearchMode({ mode: 'balanced' });
|
|
expect(k.cross_modal_both_text_weight).toBe(0.6);
|
|
expect(k.cross_modal_both_image_weight).toBe(0.4);
|
|
});
|
|
});
|