* Merge branch 'master' into garrytan/type-taxonomy-unification Resolve VERSION, package.json, CHANGELOG conflicts with v0.41.22.0 on top, preserving master's v0.41.19.0 entry below. * feat: v0.41.22.0 type-unification cathedral — collapse 94 types to 15 (closes #1479) Ships gbrain-base-v2 as the new install default (15 canonical types: 14 + note catch-all) and the unify-types PROTECTED Minion handler that runs the gbrain-base→v2 migration end-to-end on existing brains. What this delivers: - gbrain-base-v2.yaml standalone schema pack (no extends:) with 14 canonical page_types + 9 cluster mapping_rules + catch-all sentinel - 3 new schema-pack primitives: runRetypeCore (chunked UPDATE with legacy_type stamping), runPageToLinkCore (edge-shaped pages → link rows), runPageToAliasCore (concept-redirect → slug_aliases) - rewriteLinksBatch for N-pair atomic FK rewrite - Migration v104 slug_aliases table (forward-bootstrap probed on both engines for safe upgrade chain) - New engine method resolveSlugWithAlias(slug, sourceOrSources) on both Postgres + PGLite with multi-source ambiguity warning - inferTypeAndSubtypeFromPack overload + subtypes: + mapping_rules: + migration_from: schema-pack manifest extensions - findPackSuccessors version-range walker (1.x / 1.0.x / exact match) - expandTypeFilter for --type back-compat (D14): legacy aliases route through mapping_rules → canonical+subtype before the SQL filter fires - 3 new onboard checks: pack_upgrade_available, type_proliferation, dangling_aliases (source-scoped per F12) - unify-types Minion handler (PROTECTED, manual_only via render.ts allowlist per D17): retype-explicit → retype-catch-all → page-to-link → page-to-alias → final sync → active-pack flip - alias_resolved 1.05x post-fusion search boost stage; KNOBS_HASH_VERSION bumped 5→6 (one-time cache miss on upgrade, self-healing in TTL) - ELIGIBLE_TYPES for facts extraction extended with v2 canonicals (codex F-ELIGIBLE: blocker not v0.43 follow-up) Tests: 79 new unit/integration cases + 3 E2E cases covering all 9 production clusters end-to-end. 124-case verification on the cache-key + build-llms fixes. KNOBS_HASH_VERSION assertions updated in 3 tests. Plan: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md (16 locked decisions D1-D17, 12 baseline fixes F7-F21 absorbed from codex outside voice). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: CI verify failures — system-of-record allow-comment + schema-unify manifest registration Two CI failures on PR #1542: 1. check:system-of-record flagged page-to-link.ts:207 addLinksBatch as a direct write to a derived table. The call IS the reconcile surface for page_to_link mapping_rules — it converts edge-shaped pages into canonical link rows under the PROTECTED unify-types Minion handler, source-scoped, atomic per-rule. Added the canonical `// gbrain-allow-direct-insert: <reason>` comment on the same line. 2. check:resolver emitted 11 orphan_trigger warnings for `schema-unify` because the skill was added to skills/RESOLVER.md without a corresponding entry in skills/manifest.json. Added the registration under the existing skills[] array. bun run verify: 28/28 checks pass locally. * fix: CI test failures — schema-unify conformance + eligibility regression Six test failures across shards 2 + 10 on PR #1542: 1. resolver.test.ts: round-trip parser requires frontmatter triggers to be quoted (`- "..."` or `- '...'`). schema-unify shipped with bare YAML strings; quoted the 10 triggers to round-trip correctly. 2. skills-conformance.test.ts (×3): schema-unify SKILL.md was missing the required Contract, Anti-Patterns, and Output Format sections that every conformant skill must declare. Added all three: - Contract: inputs / outputs / side effects / failure modes - Anti-Patterns: 5 DON'Ts including the autopilot trust boundary - Output Format: per-phase stderr lines + celebration summary + JSON envelope shape 3. facts-eligibility.test.ts (×2): the v0.41.22 ELIGIBLE_TYPES expansion added `concept` to the eligible list, but the existing test suite pins concept as rejected (it's `extractable: true` in the schema pack but the v0.41.11 contract documented this as "cosmetic on the backstop path because backstop uses hardcoded ELIGIBLE_TYPES"). Removed `concept` from the expansion; other v2 canonicals (media, tweet, atom, analysis) stay. Comment updated to document the deliberate omission. All 6 failing tests now pass locally (370/370 across the 3 affected files). bun run verify: 28/28 checks green. * fix: harden findPackSuccessors test against shard pollution CI shard 8 reported 1 fail (1.00ms — too fast for any real loadActivePack file I/O) on `finds gbrain-base-v2 as successor of gbrain-base@1.0.0`. Local triple-run passes 9/9 in isolation. Root cause: the existing afterEach reset clears the module-level pack cache AFTER each test, but the FIRST test in the file inherits whatever state sibling files in the same bun shard process left behind. With 24+ schema-pack tests in shard 8 (mutate, mutate-audit, best-effort, registry-reload, manifest-v041_2, etc.) running before this file, the first test can read a poisoned cache. Fix: add `beforeEach(_resetPackCacheForTests)`. Two-sided reset guarantees clean state regardless of file ordering within the shard. bun run verify: 28/28 checks pass. * fix: quarantine two flaky tests to serial runner CI shard 1 + shard 8 each surfaced one intermittent failure: shard 1: buildBrainTools > execute() on put_page with valid namespace shard 8: findPackSuccessors > finds gbrain-base-v2 as successor Both pass cleanly in isolation. Both are concurrency races against shared in-shard state: - brain-allowlist.test.ts shares a singleton PGLiteEngine across 18 tests with a beforeEach DELETE FROM pages. With max-concurrency=4, two put_page tests can interleave their TRUNCATE + write phases, so the auto-link/extract sub-steps inside put_page race against the sibling test's DELETE. - schema-pack-find-pack-successors.test.ts reads bundled YAML packs via loadActivePack. The module-level pack cache is shared across parallel tests in the same shard; the previous beforeEach reset helped but didn't fully isolate against concurrent file reads under CI load. Fix per CLAUDE.md test-isolation lint rule R2 (concurrency-fragile files belong in the .serial.test.ts quarantine): rename both files to *.serial.test.ts. Serial runner picks them up at max-concurrency=1. 49/49 serial files pass locally. 28/28 verify checks pass. * fix: quarantine embed-stale test to serial runner CI shard 9 reported 6 failures, all from the embedStaleForSource describe block, all ~120-150ms each — classic shared-engine concurrency race shape. Passes 7/7 locally in isolation. Root cause: embed-stale.test.ts shares a singleton PGLiteEngine across 7 tests with beforeEach resetPgliteState. Under bun's max-concurrency=4 in the parallel shard, two tests can interleave their TRUNCATE + seedPage + upsertChunks + embedStaleForSource flow, so one test's stale-chunk count sees another test's mid-flight writes. Same fix as brain-allowlist.serial.test.ts and schema-pack-find-pack-successors.serial.test.ts: rename to *.serial.test.ts so the serial runner picks it up at max-concurrency=1. bun run verify: 28/28 checks pass. 7/7 embed-stale tests pass via serial. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
9.2 KiB
Type Taxonomy (v0.41.22: gbrain-base-v2)
The 14-canonical-type DRY/MECE taxonomy shipped in v0.41.22. Predecessor
gbrain-base(24 types) stays bundled for back-compat; v0.42+ installs default togbrain-base-v2.
Why
A production gbrain brain (186K pages) had accreted 94 distinct
pages.type values in 9 clusters of redundancy. The type system is
the foundation for schema packs, search filtering, extract behavior,
enrichment routing, and expert routing. When types are noisy, every
downstream feature degrades:
- Search filtering is ambiguous —
--type articlemisses 2.2K articles typed asmedia/article,sources/article, etc. - Enrichment routing is incomplete —
enrichable_typescould only list a few canonical types; 80+ legacy types meant most pages never got enriched. - Agent confusion — when ingesting a new article, should it be
article,media/article,sources/article, orsource/article? Four reasonable choices, none of them right. - Orphan inflation — 5,521 concept-redirect pages inflated orphan counts without adding knowledge value.
Issue #1479 catalogues the 9 clusters with exact counts. This doc is the response: a coherent 14-type taxonomy with subtypes/format/origin pushed to frontmatter, alias-table rows for redirects, real link-table rows for edge-shaped pages.
The 14 canonical types (+ note catch-all)
| Type | Primitive | What it holds | Examples |
|---|---|---|---|
person |
entity | People | Founders, partners, individuals |
company |
entity | Companies, products, orgs (subtype-distinguished) | Companies, YC-companies, products |
media |
media | Articles, videos, essays, books, podcasts (subtype-distinguished) | Substack posts, YouTube videos, books |
tweet |
media | Twitter posts (single/bundle/stub subtype) | Single tweets, threads, bundles |
social-digest |
temporal | Period-grouped social summaries (daily/monthly) | X account daily digests |
analysis |
media | Research + competitive intel | Market analysis, pricing analysis |
atom |
annotation | Knowledge units (extraction/manual/lore subtype) | Extracted facts, manual notes, lore |
concept |
concept | Ideas + reference pages | Wiki concepts |
source |
media | Transcripts, references | Interview transcripts |
deal |
temporal | Investment deals | Term sheets, investments |
email |
temporal | Email threads | Email correspondence |
slack |
temporal | Slack messages + threads | Slack conversations |
writing |
media | Original writing | Drafts, essays in progress |
project |
concept | Initiatives, workstreams | Internal projects |
note |
concept | Catch-all for one-offs (legacy_type preserved) | Memos, anecdotes, insights, etc. |
15 types total (14 canonical + note). The catch-all retype rule
binds any uncovered legacy type to note with
frontmatter.legacy_type = <original> preserved for rollback.
Subtypes (declared in frontmatter post-unify)
| Canonical | Subtype field | Values |
|---|---|---|
company |
subtype |
company / product / org |
media |
subtype |
video / article / essay / book / podcast / blog |
tweet |
subtype |
single / bundle / stub |
social-digest |
subtype |
daily / monthly |
atom |
subtype |
extraction / manual / lore |
subtype_field for retype rules is restricted to an allowlist:
{subtype, legacy_type, origin, format, kind, period, domain}. This
prevents third-party packs from injecting title, slug, or type
via mapping_rules (codex D9 security hardening).
Migration flow
gbrain onboard --check # surfaces pack_upgrade_available
↓
gbrain onboard --check --explain # per-cluster narrative dry-run
↓
gbrain jobs submit unify-types \ # PROTECTED + manual_only
--allow-protected \
--params '{"target_pack":"gbrain-base-v2"}'
↓
Handler runs 4 phases:
┌─────────────────────────────────────┐
│ Phase 1: Preflight + lock │ → gbrain-unify db-lock (60min TTL)
├─────────────────────────────────────┤
│ Phase 2: Retype explicit rules │ → chunked UPDATE 1000/batch
├─────────────────────────────────────┤
│ Phase 3: Retype catch-all sentinel │ → 'note' with legacy_type
├─────────────────────────────────────┤
│ Phase 4: Page-to-link conversions │ → insert links + soft-delete
├─────────────────────────────────────┤
│ Phase 5: Page-to-alias conversions │ → insert slug_aliases + soft-delete
├─────────────────────────────────────┤
│ Phase 6: Final sync (residual) │ → path-prefix typing
├─────────────────────────────────────┤
│ Phase 7: Flip active pack (D13) │ → engine.setConfig + saveConfig
├─────────────────────────────────────┤
│ Phase 8: Verify + celebrate │ → assert ≤16 types; stderr summary
└─────────────────────────────────────┘
↓
gbrain onboard --check # pack_upgrade_available cleared
# type_proliferation cleared
Rollback paths
Every primitive ships with a documented rollback:
| Operation | Rollback |
|---|---|
| Retype | frontmatter.legacy_type = <original> preserved on every page (D8). One SQL UPDATE restores types: UPDATE pages SET type = frontmatter->>'legacy_type' WHERE frontmatter ? 'legacy_type'. |
| Page-to-link | Source page soft-deleted with 72h TTL. gbrain pages restore <slug> within 72h. Link row stays harmless if source restored. |
| Page-to-alias | Source page soft-deleted with 72h TTL. gbrain pages restore <slug> within 72h. Alias row stays harmless (or DELETE FROM slug_aliases WHERE alias_slug = <slug> to clean up). |
| Active-pack flip | gbrain schema use gbrain-base reverses the flip. |
What if my brain doesn't fit?
The catch-all retype rule (from_type: '*unknown*') handles long-tail
types automatically — any page whose type isn't covered by an explicit
rule AND isn't a page_to_link / page_to_alias source gets retyped to
note with legacy_type preserved. Guarantees ≤16 distinct types
post-unify on ANY brain.
For brains with substantial custom types that deserve their own canonical
(e.g. researcher for an academic brain), the right move is:
- Fork gbrain-base-v2:
gbrain schema fork gbrain-base-v2 my-pack - Edit your fork to add page_types + mapping_rules covering your custom domain.
- Target your fork:
gbrain jobs submit unify-types --allow-protected --params '{"target_pack":"my-pack"}'
Your fork can also declare migration_from: {pack: gbrain-base-v2, version: "1.x"} to register itself as a successor — future agents
discovering your pack via pack_upgrade_available will offer the
migration.
Wikilink resolution post-unify
The slug_aliases table IS the resolver (D15: codex outside voice —
don't rewrite body-text wikilinks; the alias table is the right
primitive). Wikilinks like [[old-redirect-slug]] keep working post-
unify because:
- The wikilink resolver short-circuits through
engine.resolveSlugWithAlias(slug, sourceId)BEFORE the existing fuzzy/prefix cascade. - The lookup queries
slug_aliasesfor any matching alias_slug in the provided source(s). - If found, returns the canonical_slug. The renderer then resolves the wikilink to the canonical page.
Multi-source ambiguity (same alias_slug in two registered sources)
emits a once-per-process multi_match stderr warning and returns the
first match by source array order. Federated reads pass the full
allowed-source array.
Search ranking signal: alias_resolved_boost
Post-unify, search results whose slug is a canonical_slug in
slug_aliases get a 1.05x score multiplier via the
applyAliasResolvedBoost post-fusion stage. Semantic intent: "user
explicitly disambiguated this as canonical, so it should outrank fuzzy
matches that hit aliases by accident."
SearchResult.alias_resolved_boost is stamped on touched results for
--explain formatter visibility. KNOBS_HASH_VERSION bumped 5→6 to
invalidate pre-v0.42 cache rows that don't reflect the new stage.
Reference
- Issue: https://github.com/garrytan/gbrain/issues/1479
- Pack file:
src/core/schema-pack/base/gbrain-base-v2.yaml - Pack-upgrade mechanism:
docs/architecture/pack-upgrade-mechanism.md - Migration handler:
src/core/schema-pack/unify-types-handler.ts - Onboard checks:
src/core/onboard/checks.ts - Skill:
skills/schema-unify/SKILL.md - Plan + decisions:
~/.claude/plans/system-instruction-you-are-working-transient-elephant.md