* Merge branch 'master' into garrytan/type-taxonomy-unification Resolve VERSION, package.json, CHANGELOG conflicts with v0.41.22.0 on top, preserving master's v0.41.19.0 entry below. * feat: v0.41.22.0 type-unification cathedral — collapse 94 types to 15 (closes #1479) Ships gbrain-base-v2 as the new install default (15 canonical types: 14 + note catch-all) and the unify-types PROTECTED Minion handler that runs the gbrain-base→v2 migration end-to-end on existing brains. What this delivers: - gbrain-base-v2.yaml standalone schema pack (no extends:) with 14 canonical page_types + 9 cluster mapping_rules + catch-all sentinel - 3 new schema-pack primitives: runRetypeCore (chunked UPDATE with legacy_type stamping), runPageToLinkCore (edge-shaped pages → link rows), runPageToAliasCore (concept-redirect → slug_aliases) - rewriteLinksBatch for N-pair atomic FK rewrite - Migration v104 slug_aliases table (forward-bootstrap probed on both engines for safe upgrade chain) - New engine method resolveSlugWithAlias(slug, sourceOrSources) on both Postgres + PGLite with multi-source ambiguity warning - inferTypeAndSubtypeFromPack overload + subtypes: + mapping_rules: + migration_from: schema-pack manifest extensions - findPackSuccessors version-range walker (1.x / 1.0.x / exact match) - expandTypeFilter for --type back-compat (D14): legacy aliases route through mapping_rules → canonical+subtype before the SQL filter fires - 3 new onboard checks: pack_upgrade_available, type_proliferation, dangling_aliases (source-scoped per F12) - unify-types Minion handler (PROTECTED, manual_only via render.ts allowlist per D17): retype-explicit → retype-catch-all → page-to-link → page-to-alias → final sync → active-pack flip - alias_resolved 1.05x post-fusion search boost stage; KNOBS_HASH_VERSION bumped 5→6 (one-time cache miss on upgrade, self-healing in TTL) - ELIGIBLE_TYPES for facts extraction extended with v2 canonicals (codex F-ELIGIBLE: blocker not v0.43 follow-up) Tests: 79 new unit/integration cases + 3 E2E cases covering all 9 production clusters end-to-end. 124-case verification on the cache-key + build-llms fixes. KNOBS_HASH_VERSION assertions updated in 3 tests. Plan: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md (16 locked decisions D1-D17, 12 baseline fixes F7-F21 absorbed from codex outside voice). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: CI verify failures — system-of-record allow-comment + schema-unify manifest registration Two CI failures on PR #1542: 1. check:system-of-record flagged page-to-link.ts:207 addLinksBatch as a direct write to a derived table. The call IS the reconcile surface for page_to_link mapping_rules — it converts edge-shaped pages into canonical link rows under the PROTECTED unify-types Minion handler, source-scoped, atomic per-rule. Added the canonical `// gbrain-allow-direct-insert: <reason>` comment on the same line. 2. check:resolver emitted 11 orphan_trigger warnings for `schema-unify` because the skill was added to skills/RESOLVER.md without a corresponding entry in skills/manifest.json. Added the registration under the existing skills[] array. bun run verify: 28/28 checks pass locally. * fix: CI test failures — schema-unify conformance + eligibility regression Six test failures across shards 2 + 10 on PR #1542: 1. resolver.test.ts: round-trip parser requires frontmatter triggers to be quoted (`- "..."` or `- '...'`). schema-unify shipped with bare YAML strings; quoted the 10 triggers to round-trip correctly. 2. skills-conformance.test.ts (×3): schema-unify SKILL.md was missing the required Contract, Anti-Patterns, and Output Format sections that every conformant skill must declare. Added all three: - Contract: inputs / outputs / side effects / failure modes - Anti-Patterns: 5 DON'Ts including the autopilot trust boundary - Output Format: per-phase stderr lines + celebration summary + JSON envelope shape 3. facts-eligibility.test.ts (×2): the v0.41.22 ELIGIBLE_TYPES expansion added `concept` to the eligible list, but the existing test suite pins concept as rejected (it's `extractable: true` in the schema pack but the v0.41.11 contract documented this as "cosmetic on the backstop path because backstop uses hardcoded ELIGIBLE_TYPES"). Removed `concept` from the expansion; other v2 canonicals (media, tweet, atom, analysis) stay. Comment updated to document the deliberate omission. All 6 failing tests now pass locally (370/370 across the 3 affected files). bun run verify: 28/28 checks green. * fix: harden findPackSuccessors test against shard pollution CI shard 8 reported 1 fail (1.00ms — too fast for any real loadActivePack file I/O) on `finds gbrain-base-v2 as successor of gbrain-base@1.0.0`. Local triple-run passes 9/9 in isolation. Root cause: the existing afterEach reset clears the module-level pack cache AFTER each test, but the FIRST test in the file inherits whatever state sibling files in the same bun shard process left behind. With 24+ schema-pack tests in shard 8 (mutate, mutate-audit, best-effort, registry-reload, manifest-v041_2, etc.) running before this file, the first test can read a poisoned cache. Fix: add `beforeEach(_resetPackCacheForTests)`. Two-sided reset guarantees clean state regardless of file ordering within the shard. bun run verify: 28/28 checks pass. * fix: quarantine two flaky tests to serial runner CI shard 1 + shard 8 each surfaced one intermittent failure: shard 1: buildBrainTools > execute() on put_page with valid namespace shard 8: findPackSuccessors > finds gbrain-base-v2 as successor Both pass cleanly in isolation. Both are concurrency races against shared in-shard state: - brain-allowlist.test.ts shares a singleton PGLiteEngine across 18 tests with a beforeEach DELETE FROM pages. With max-concurrency=4, two put_page tests can interleave their TRUNCATE + write phases, so the auto-link/extract sub-steps inside put_page race against the sibling test's DELETE. - schema-pack-find-pack-successors.test.ts reads bundled YAML packs via loadActivePack. The module-level pack cache is shared across parallel tests in the same shard; the previous beforeEach reset helped but didn't fully isolate against concurrent file reads under CI load. Fix per CLAUDE.md test-isolation lint rule R2 (concurrency-fragile files belong in the .serial.test.ts quarantine): rename both files to *.serial.test.ts. Serial runner picks them up at max-concurrency=1. 49/49 serial files pass locally. 28/28 verify checks pass. * fix: quarantine embed-stale test to serial runner CI shard 9 reported 6 failures, all from the embedStaleForSource describe block, all ~120-150ms each — classic shared-engine concurrency race shape. Passes 7/7 locally in isolation. Root cause: embed-stale.test.ts shares a singleton PGLiteEngine across 7 tests with beforeEach resetPgliteState. Under bun's max-concurrency=4 in the parallel shard, two tests can interleave their TRUNCATE + seedPage + upsertChunks + embedStaleForSource flow, so one test's stale-chunk count sees another test's mid-flight writes. Same fix as brain-allowlist.serial.test.ts and schema-pack-find-pack-successors.serial.test.ts: rename to *.serial.test.ts so the serial runner picks it up at max-concurrency=1. bun run verify: 28/28 checks pass. 7/7 embed-stale tests pass via serial. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
6.0 KiB
Convention: schema evolution — when to add a type vs alias vs prefix
Cross-cutting convention for any skill that proposes a change to the
active schema pack. Read first before invoking schema-author. The
goal: keep the pack small enough that an agent can hold the whole type
graph in its head, but expressive enough that custom domains
(research, legal, founder ops) get first-class types.
Decision tree
You see a cluster of pages that share a domain meaning.
│
▼
How many pages in the cluster?
│
┌─────┴───────┬──────────────┐
▼ ▼ ▼
<20 20-100 100+
│ │ │
▼ ▼ ▼
One-off. Big enough. First-class.
Don't pack- Add an alias Add a new
codify. to an existing page_type with
type OR a its own prefix,
Use the narrow prefix primitive, and
nearest branch. flags.
existing
type +
frontmatter
tag.
Concrete examples
One-off (don't add to pack):
"I have 3 pages under
2026-projects/skunkworks-spec/. Should I add askunkworkstype?"
No. Three pages doesn't justify a permanent pack entry. Type these as
the nearest existing match (concept or note) and use a frontmatter
project: tag. If the cluster grows to 20+, revisit.
20-100 pages — alias OR narrow prefix:
"I have 50 pages under
people/researchers/that overlap with mypersontype. Should I add aresearchertype?"
Two valid options:
- Alias on
person—add-alias person researcher. Closure queries forresearcherwill surfacepersonrows too. - New type sharing the
entityprimitive —add-type researcher --primitive entity --prefix people/researchers/. Distinct type, can be marked--extractableor--expertindependently.
Pick alias when researchers are people first, researchers second (they share enrichment rules, expert-routing semantics, link verbs). Pick new type when researcher-specific behavior diverges (different extractable rules, different link verbs, different rubric).
100+ pages — first-class type:
"I have 4000 pages under
meetings/. I want them typed asmeeting, not the legacy defaultnote."
Add the type:
gbrain schema add-type meeting \
--primitive temporal \
--prefix meetings/ \
--extractable
gbrain schema sync --apply
The sync --apply backfills all 4000 pages. From here forward,
imports under meetings/ infer meeting type via the pack.
Don'ts
- Don't add a type for a directory you imported once for triage. Pack types are permanent decisions; one-time imports are not.
- Don't add a type just to silence
dead_prefixesinschema stats. A dead prefix is a signal that the prefix is mis-declared or the corpus moved. Remove the prefix or migrate the content, don't add an empty type. - Don't promote a candidate from
schema suggestwithout verifying the path prefix matches real content. The suggester is heuristic; it can propose types that overlap existing ones. Runlint --with-dbbeforeadd-typeto catch prefix collisions pre-write. - Don't add
--expertto a type that has nopath_prefixes. Theexpert_routing_without_prefixlint rule warns about this exact shape: an expert-routed type with no prefix never matches a put_page inference, sowhoknowssilently never surfaces it. - Don't mutate
gbrain-baseorgbrain-recommended. Fork first.
When to remove a type
Removing a type is RARE. Only do it when:
- The type was added in error (typo, premature abstraction).
- The corpus the type was meant for has been migrated to a different type.
- The type is dangling (no
path_prefixesactually match pages, no queries reference it, no other type's aliases/link_types reference it).
remove-type is guarded by the STILL_REFERENCED check (codex C14): if
ANY other type's aliases / enrichable_types / link_types / frontmatter_links
references the target, the remove fails loud with the reference list.
Break those references first.
When to commit the pack
If your pack lives in source control (~/.gbrain/schema-packs/<name>/
is a git repo), commit after every batch of mutations. The
mutation_count_anomaly lint rule warns at >50 mutations in 7 days —
that's the hint to start committing rather than relying on disk-only
state.
When to upgrade your pack (v0.42+)
A pack can declare migration_from: {pack: <name>, version: <semver-range>}
to register itself as the successor to another pack. When a brain's
active pack matches the declared from, the pack_upgrade_available
onboard check surfaces the successor + a manual_only RemediationStep
pointing at the unify-types PROTECTED Minion handler.
v0.41.22 ships gbrain-base-v2 as the declared successor to
gbrain-base@1.x — collapses 94 noisy types to 15 canonical via
declarative mapping_rules. Run via gbrain onboard --check --explain
(preview) → gbrain jobs submit unify-types --allow-protected --params '{"target_pack":"gbrain-base-v2"}' (apply). See
skills/schema-unify/SKILL.md for the full playbook.
Authoring a successor pack: declare
migration_from: {pack: <parent>, version: "1.x"} in the manifest
plus mapping_rules: (discriminated union over retype / page_to_link /
page_to_alias kinds). Catch-all sentinel from_type: '*unknown*' MUST
appear last. Subtype_field is restricted to ALLOWED_SUBTYPE_FIELDS
(subtype, legacy_type, origin, format, kind, period, domain) per
codex D9 — third-party packs cannot inject title / slug / type.
When NOT to upgrade:
- Custom types not covered by the successor's mapping_rules → fork the
successor first (
gbrain schema fork gbrain-base-v2 my-pack), edit rules, then target your fork. - Mid-ingest or autopilot maintenance → wait. Unify holds the
gbrain-unifydb-lock for ~10 min on big brains. - Federated brain with sources you don't want to touch → scope per
source via
--params sourceId.