Files
gbrain/test/schema-pack-page-to-link.test.ts
T
5d42f3295e v0.41.22.0 feat: type-unification cathedral — 94 types → 15 canonical (closes #1479) (#1542)
* Merge branch 'master' into garrytan/type-taxonomy-unification

Resolve VERSION, package.json, CHANGELOG conflicts with v0.41.22.0
on top, preserving master's v0.41.19.0 entry below.

* feat: v0.41.22.0 type-unification cathedral — collapse 94 types to 15 (closes #1479)

Ships gbrain-base-v2 as the new install default (15 canonical types: 14
+ note catch-all) and the unify-types PROTECTED Minion handler that
runs the gbrain-base→v2 migration end-to-end on existing brains.

What this delivers:
- gbrain-base-v2.yaml standalone schema pack (no extends:) with 14
  canonical page_types + 9 cluster mapping_rules + catch-all sentinel
- 3 new schema-pack primitives: runRetypeCore (chunked UPDATE with
  legacy_type stamping), runPageToLinkCore (edge-shaped pages →
  link rows), runPageToAliasCore (concept-redirect → slug_aliases)
- rewriteLinksBatch for N-pair atomic FK rewrite
- Migration v104 slug_aliases table (forward-bootstrap probed on both
  engines for safe upgrade chain)
- New engine method resolveSlugWithAlias(slug, sourceOrSources) on
  both Postgres + PGLite with multi-source ambiguity warning
- inferTypeAndSubtypeFromPack overload + subtypes: + mapping_rules:
  + migration_from: schema-pack manifest extensions
- findPackSuccessors version-range walker (1.x / 1.0.x / exact match)
- expandTypeFilter for --type back-compat (D14): legacy aliases route
  through mapping_rules → canonical+subtype before the SQL filter fires
- 3 new onboard checks: pack_upgrade_available, type_proliferation,
  dangling_aliases (source-scoped per F12)
- unify-types Minion handler (PROTECTED, manual_only via render.ts
  allowlist per D17): retype-explicit → retype-catch-all →
  page-to-link → page-to-alias → final sync → active-pack flip
- alias_resolved 1.05x post-fusion search boost stage; KNOBS_HASH_VERSION
  bumped 5→6 (one-time cache miss on upgrade, self-healing in TTL)
- ELIGIBLE_TYPES for facts extraction extended with v2 canonicals
  (codex F-ELIGIBLE: blocker not v0.43 follow-up)

Tests: 79 new unit/integration cases + 3 E2E cases covering all 9
production clusters end-to-end. 124-case verification on the cache-key
+ build-llms fixes. KNOBS_HASH_VERSION assertions updated in 3 tests.

Plan: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md
(16 locked decisions D1-D17, 12 baseline fixes F7-F21 absorbed from
codex outside voice).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: CI verify failures — system-of-record allow-comment + schema-unify manifest registration

Two CI failures on PR #1542:

1. check:system-of-record flagged page-to-link.ts:207 addLinksBatch as
   a direct write to a derived table. The call IS the reconcile surface
   for page_to_link mapping_rules — it converts edge-shaped pages into
   canonical link rows under the PROTECTED unify-types Minion handler,
   source-scoped, atomic per-rule. Added the canonical
   `// gbrain-allow-direct-insert: <reason>` comment on the same line.

2. check:resolver emitted 11 orphan_trigger warnings for `schema-unify`
   because the skill was added to skills/RESOLVER.md without a
   corresponding entry in skills/manifest.json. Added the registration
   under the existing skills[] array.

bun run verify: 28/28 checks pass locally.

* fix: CI test failures — schema-unify conformance + eligibility regression

Six test failures across shards 2 + 10 on PR #1542:

1. resolver.test.ts: round-trip parser requires frontmatter triggers to
   be quoted (`- "..."` or `- '...'`). schema-unify shipped with bare
   YAML strings; quoted the 10 triggers to round-trip correctly.

2. skills-conformance.test.ts (×3): schema-unify SKILL.md was missing
   the required Contract, Anti-Patterns, and Output Format sections
   that every conformant skill must declare. Added all three:
   - Contract: inputs / outputs / side effects / failure modes
   - Anti-Patterns: 5 DON'Ts including the autopilot trust boundary
   - Output Format: per-phase stderr lines + celebration summary +
     JSON envelope shape

3. facts-eligibility.test.ts (×2): the v0.41.22 ELIGIBLE_TYPES
   expansion added `concept` to the eligible list, but the existing
   test suite pins concept as rejected (it's `extractable: true` in
   the schema pack but the v0.41.11 contract documented this as
   "cosmetic on the backstop path because backstop uses hardcoded
   ELIGIBLE_TYPES"). Removed `concept` from the expansion; other v2
   canonicals (media, tweet, atom, analysis) stay. Comment updated
   to document the deliberate omission.

All 6 failing tests now pass locally (370/370 across the 3 affected
files). bun run verify: 28/28 checks green.

* fix: harden findPackSuccessors test against shard pollution

CI shard 8 reported 1 fail (1.00ms — too fast for any real loadActivePack
file I/O) on `finds gbrain-base-v2 as successor of gbrain-base@1.0.0`.
Local triple-run passes 9/9 in isolation.

Root cause: the existing afterEach reset clears the module-level pack
cache AFTER each test, but the FIRST test in the file inherits whatever
state sibling files in the same bun shard process left behind. With
24+ schema-pack tests in shard 8 (mutate, mutate-audit, best-effort,
registry-reload, manifest-v041_2, etc.) running before this file, the
first test can read a poisoned cache.

Fix: add `beforeEach(_resetPackCacheForTests)`. Two-sided reset
guarantees clean state regardless of file ordering within the shard.

bun run verify: 28/28 checks pass.

* fix: quarantine two flaky tests to serial runner

CI shard 1 + shard 8 each surfaced one intermittent failure:

shard 1: buildBrainTools > execute() on put_page with valid namespace
shard 8: findPackSuccessors > finds gbrain-base-v2 as successor

Both pass cleanly in isolation. Both are concurrency races against
shared in-shard state:

- brain-allowlist.test.ts shares a singleton PGLiteEngine across 18
  tests with a beforeEach DELETE FROM pages. With max-concurrency=4,
  two put_page tests can interleave their TRUNCATE + write phases,
  so the auto-link/extract sub-steps inside put_page race against
  the sibling test's DELETE.
- schema-pack-find-pack-successors.test.ts reads bundled YAML packs
  via loadActivePack. The module-level pack cache is shared across
  parallel tests in the same shard; the previous beforeEach reset
  helped but didn't fully isolate against concurrent file reads
  under CI load.

Fix per CLAUDE.md test-isolation lint rule R2 (concurrency-fragile
files belong in the .serial.test.ts quarantine): rename both files
to *.serial.test.ts. Serial runner picks them up at max-concurrency=1.
49/49 serial files pass locally. 28/28 verify checks pass.

* fix: quarantine embed-stale test to serial runner

CI shard 9 reported 6 failures, all from the embedStaleForSource describe
block, all ~120-150ms each — classic shared-engine concurrency race shape.
Passes 7/7 locally in isolation.

Root cause: embed-stale.test.ts shares a singleton PGLiteEngine across 7
tests with beforeEach resetPgliteState. Under bun's max-concurrency=4 in
the parallel shard, two tests can interleave their TRUNCATE + seedPage +
upsertChunks + embedStaleForSource flow, so one test's stale-chunk count
sees another test's mid-flight writes.

Same fix as brain-allowlist.serial.test.ts and
schema-pack-find-pack-successors.serial.test.ts: rename to *.serial.test.ts
so the serial runner picks it up at max-concurrency=1.

bun run verify: 28/28 checks pass. 7/7 embed-stale tests pass via serial.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 07:01:28 -07:00

213 lines
8.0 KiB
TypeScript

// v0.42 Type Unification (T24) — runPageToLinkCore unit tests.
//
// Coverage: resolver variants (frontmatter / body_first_link / explicit field),
// unresolved tracking (no_source / no_target / cycle / parse_failed),
// soft-delete after link insert, source-scoping, regression guard that
// page-to-link does NOT keep the source page (it's converted away).
import { afterAll, beforeAll, beforeEach, describe, expect, it } from 'bun:test';
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';
import { runPageToLinkCore } from '../src/core/schema-pack/page-to-link.ts';
import type { OperationContext } from '../src/core/operations.ts';
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await resetPgliteState(engine);
});
function ctxOf(): OperationContext {
return {
engine,
config: {},
logger: { info: () => {}, warn: () => {}, error: () => {} },
dryRun: false,
remote: false,
} as unknown as OperationContext;
}
async function seed(slug: string, type: string, fm: Record<string, unknown> = {}, body = 'edge body that is long enough') {
await engine.putPage(slug, {
title: slug,
type: type as never,
compiled_truth: body,
timeline: '',
frontmatter: fm,
source_path: `${slug}.md`,
});
}
describe('runPageToLinkCore', () => {
describe('dry-run', () => {
it('counts pages without mutating', async () => {
await seed('atoms/partner-1', 'atom-partner-link',
{ source: 'people/alice', target: 'companies/acme' });
await seed('people/alice', 'person');
await seed('companies/acme', 'company');
const result = await runPageToLinkCore(ctxOf(), {
rules: [{
from_type: 'atom-partner-link',
link_type: 'partner_of',
source_slug_from: { frontmatter_field: 'source' },
target_slug_from: { frontmatter_field: 'target' },
}],
apply: false,
});
expect(result.per_rule[0].would_convert).toBe(1);
expect(result.per_rule[0].converted).toBe(0);
// Source page should still exist
const rows = await engine.executeRaw<{ id: number }>(
`SELECT id FROM pages WHERE slug = 'atoms/partner-1' AND deleted_at IS NULL`,
);
expect(rows.length).toBe(1);
});
});
describe('apply', () => {
it('inserts link row + soft-deletes source page', async () => {
await seed('people/alice', 'person');
await seed('companies/acme', 'company');
await seed('atoms/partner-1', 'atom-partner-link',
{ source: 'people/alice', target: 'companies/acme' });
const result = await runPageToLinkCore(ctxOf(), {
rules: [{
from_type: 'atom-partner-link',
link_type: 'partner_of',
source_slug_from: { frontmatter_field: 'source' },
target_slug_from: { frontmatter_field: 'target' },
}],
apply: true,
});
expect(result.per_rule[0].converted).toBe(1);
expect(result.per_rule[0].soft_deleted).toBe(1);
// Source page soft-deleted
const srcRows = await engine.executeRaw<{ deleted_at: string | null }>(
`SELECT deleted_at FROM pages WHERE slug = 'atoms/partner-1'`,
);
expect(srcRows[0].deleted_at).not.toBeNull();
// Link row inserted
const linkRows = await engine.executeRaw<{ link_type: string }>(
`SELECT link_type FROM links l
JOIN pages p1 ON l.from_page_id = p1.id
JOIN pages p2 ON l.to_page_id = p2.id
WHERE p1.slug = 'people/alice' AND p2.slug = 'companies/acme'`,
);
expect(linkRows.length).toBe(1);
expect(linkRows[0].link_type).toBe('partner_of');
});
it('records unresolved when source frontmatter field is missing', async () => {
await seed('atoms/bad-1', 'atom-partner-link', { target: 'companies/acme' }); // no source
const result = await runPageToLinkCore(ctxOf(), {
rules: [{
from_type: 'atom-partner-link',
link_type: 'partner_of',
source_slug_from: { frontmatter_field: 'source' },
target_slug_from: { frontmatter_field: 'target' },
}],
apply: true,
});
expect(result.per_rule[0].converted).toBe(0);
expect(result.per_rule[0].unresolved.length).toBe(1);
expect(result.per_rule[0].unresolved[0].reason).toBe('no_source');
});
it('records unresolved when target is missing', async () => {
await seed('atoms/bad-1', 'atom-partner-link', { source: 'people/alice' }); // no target
const result = await runPageToLinkCore(ctxOf(), {
rules: [{
from_type: 'atom-partner-link',
link_type: 'partner_of',
source_slug_from: { frontmatter_field: 'source' },
target_slug_from: { frontmatter_field: 'target' },
}],
apply: true,
});
expect(result.per_rule[0].unresolved[0].reason).toBe('no_target');
});
it('rejects self-references (cycle reason)', async () => {
await seed('atoms/loop-1', 'atom-partner-link',
{ source: 'people/alice', target: 'people/alice' });
await seed('people/alice', 'person');
const result = await runPageToLinkCore(ctxOf(), {
rules: [{
from_type: 'atom-partner-link',
link_type: 'partner_of',
source_slug_from: { frontmatter_field: 'source' },
target_slug_from: { frontmatter_field: 'target' },
}],
apply: true,
});
expect(result.per_rule[0].converted).toBe(0);
expect(result.per_rule[0].unresolved[0].reason).toBe('cycle');
});
it('resolves slugs from body_first_link', async () => {
await seed('symlinks/x', 'symlink',
{ target: 'concepts/foo' },
'[[concepts/bar]] this is body first link\nLine 2');
await seed('concepts/foo', 'concept');
await seed('concepts/bar', 'concept');
const result = await runPageToLinkCore(ctxOf(), {
rules: [{
from_type: 'symlink',
link_type: 'relates_to',
source_slug_from: 'body_first_link',
target_slug_from: { frontmatter_field: 'target' },
}],
apply: true,
});
expect(result.per_rule[0].converted).toBe(1);
const links = await engine.executeRaw<{ from_slug: string; to_slug: string }>(
`SELECT p1.slug AS from_slug, p2.slug AS to_slug FROM links l
JOIN pages p1 ON l.from_page_id = p1.id
JOIN pages p2 ON l.to_page_id = p2.id
WHERE l.link_type = 'relates_to'`,
);
expect(links[0].from_slug).toBe('concepts/bar');
expect(links[0].to_slug).toBe('concepts/foo');
});
});
describe('source-scoping (F9)', () => {
it('limits processing to the specified sourceId', async () => {
// Two sources: default + alt
await engine.executeRaw(`INSERT INTO sources (id, name) VALUES ('alt', 'alt') ON CONFLICT DO NOTHING`);
await seed('people/alice', 'person');
await seed('atoms/p1', 'atom-partner-link',
{ source: 'people/alice', target: 'people/alice' }); // default source
// Alt-source page (skipped)
await engine.putPage('atoms/p2', {
title: 'p2', type: 'atom-partner-link' as never,
compiled_truth: 'body that is long enough to pass min char gates around extraction',
timeline: '', frontmatter: { source: 'people/alice', target: 'people/alice' },
source_path: 'atoms/p2.md',
}, { sourceId: 'alt' });
const result = await runPageToLinkCore(ctxOf(), {
rules: [{
from_type: 'atom-partner-link',
link_type: 'partner_of',
source_slug_from: { frontmatter_field: 'source' },
target_slug_from: { frontmatter_field: 'target' },
}],
apply: true,
sourceId: 'default',
});
// Only default-source page processed (and that one is a cycle → unresolved)
expect(result.per_rule[0].would_convert).toBe(1);
});
});
});