Files
gbrain/docs/architecture/type-taxonomy.md
T
5d42f3295e v0.41.22.0 feat: type-unification cathedral — 94 types → 15 canonical (closes #1479) (#1542)
* Merge branch 'master' into garrytan/type-taxonomy-unification

Resolve VERSION, package.json, CHANGELOG conflicts with v0.41.22.0
on top, preserving master's v0.41.19.0 entry below.

* feat: v0.41.22.0 type-unification cathedral — collapse 94 types to 15 (closes #1479)

Ships gbrain-base-v2 as the new install default (15 canonical types: 14
+ note catch-all) and the unify-types PROTECTED Minion handler that
runs the gbrain-base→v2 migration end-to-end on existing brains.

What this delivers:
- gbrain-base-v2.yaml standalone schema pack (no extends:) with 14
  canonical page_types + 9 cluster mapping_rules + catch-all sentinel
- 3 new schema-pack primitives: runRetypeCore (chunked UPDATE with
  legacy_type stamping), runPageToLinkCore (edge-shaped pages →
  link rows), runPageToAliasCore (concept-redirect → slug_aliases)
- rewriteLinksBatch for N-pair atomic FK rewrite
- Migration v104 slug_aliases table (forward-bootstrap probed on both
  engines for safe upgrade chain)
- New engine method resolveSlugWithAlias(slug, sourceOrSources) on
  both Postgres + PGLite with multi-source ambiguity warning
- inferTypeAndSubtypeFromPack overload + subtypes: + mapping_rules:
  + migration_from: schema-pack manifest extensions
- findPackSuccessors version-range walker (1.x / 1.0.x / exact match)
- expandTypeFilter for --type back-compat (D14): legacy aliases route
  through mapping_rules → canonical+subtype before the SQL filter fires
- 3 new onboard checks: pack_upgrade_available, type_proliferation,
  dangling_aliases (source-scoped per F12)
- unify-types Minion handler (PROTECTED, manual_only via render.ts
  allowlist per D17): retype-explicit → retype-catch-all →
  page-to-link → page-to-alias → final sync → active-pack flip
- alias_resolved 1.05x post-fusion search boost stage; KNOBS_HASH_VERSION
  bumped 5→6 (one-time cache miss on upgrade, self-healing in TTL)
- ELIGIBLE_TYPES for facts extraction extended with v2 canonicals
  (codex F-ELIGIBLE: blocker not v0.43 follow-up)

Tests: 79 new unit/integration cases + 3 E2E cases covering all 9
production clusters end-to-end. 124-case verification on the cache-key
+ build-llms fixes. KNOBS_HASH_VERSION assertions updated in 3 tests.

Plan: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md
(16 locked decisions D1-D17, 12 baseline fixes F7-F21 absorbed from
codex outside voice).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: CI verify failures — system-of-record allow-comment + schema-unify manifest registration

Two CI failures on PR #1542:

1. check:system-of-record flagged page-to-link.ts:207 addLinksBatch as
   a direct write to a derived table. The call IS the reconcile surface
   for page_to_link mapping_rules — it converts edge-shaped pages into
   canonical link rows under the PROTECTED unify-types Minion handler,
   source-scoped, atomic per-rule. Added the canonical
   `// gbrain-allow-direct-insert: <reason>` comment on the same line.

2. check:resolver emitted 11 orphan_trigger warnings for `schema-unify`
   because the skill was added to skills/RESOLVER.md without a
   corresponding entry in skills/manifest.json. Added the registration
   under the existing skills[] array.

bun run verify: 28/28 checks pass locally.

* fix: CI test failures — schema-unify conformance + eligibility regression

Six test failures across shards 2 + 10 on PR #1542:

1. resolver.test.ts: round-trip parser requires frontmatter triggers to
   be quoted (`- "..."` or `- '...'`). schema-unify shipped with bare
   YAML strings; quoted the 10 triggers to round-trip correctly.

2. skills-conformance.test.ts (×3): schema-unify SKILL.md was missing
   the required Contract, Anti-Patterns, and Output Format sections
   that every conformant skill must declare. Added all three:
   - Contract: inputs / outputs / side effects / failure modes
   - Anti-Patterns: 5 DON'Ts including the autopilot trust boundary
   - Output Format: per-phase stderr lines + celebration summary +
     JSON envelope shape

3. facts-eligibility.test.ts (×2): the v0.41.22 ELIGIBLE_TYPES
   expansion added `concept` to the eligible list, but the existing
   test suite pins concept as rejected (it's `extractable: true` in
   the schema pack but the v0.41.11 contract documented this as
   "cosmetic on the backstop path because backstop uses hardcoded
   ELIGIBLE_TYPES"). Removed `concept` from the expansion; other v2
   canonicals (media, tweet, atom, analysis) stay. Comment updated
   to document the deliberate omission.

All 6 failing tests now pass locally (370/370 across the 3 affected
files). bun run verify: 28/28 checks green.

* fix: harden findPackSuccessors test against shard pollution

CI shard 8 reported 1 fail (1.00ms — too fast for any real loadActivePack
file I/O) on `finds gbrain-base-v2 as successor of gbrain-base@1.0.0`.
Local triple-run passes 9/9 in isolation.

Root cause: the existing afterEach reset clears the module-level pack
cache AFTER each test, but the FIRST test in the file inherits whatever
state sibling files in the same bun shard process left behind. With
24+ schema-pack tests in shard 8 (mutate, mutate-audit, best-effort,
registry-reload, manifest-v041_2, etc.) running before this file, the
first test can read a poisoned cache.

Fix: add `beforeEach(_resetPackCacheForTests)`. Two-sided reset
guarantees clean state regardless of file ordering within the shard.

bun run verify: 28/28 checks pass.

* fix: quarantine two flaky tests to serial runner

CI shard 1 + shard 8 each surfaced one intermittent failure:

shard 1: buildBrainTools > execute() on put_page with valid namespace
shard 8: findPackSuccessors > finds gbrain-base-v2 as successor

Both pass cleanly in isolation. Both are concurrency races against
shared in-shard state:

- brain-allowlist.test.ts shares a singleton PGLiteEngine across 18
  tests with a beforeEach DELETE FROM pages. With max-concurrency=4,
  two put_page tests can interleave their TRUNCATE + write phases,
  so the auto-link/extract sub-steps inside put_page race against
  the sibling test's DELETE.
- schema-pack-find-pack-successors.test.ts reads bundled YAML packs
  via loadActivePack. The module-level pack cache is shared across
  parallel tests in the same shard; the previous beforeEach reset
  helped but didn't fully isolate against concurrent file reads
  under CI load.

Fix per CLAUDE.md test-isolation lint rule R2 (concurrency-fragile
files belong in the .serial.test.ts quarantine): rename both files
to *.serial.test.ts. Serial runner picks them up at max-concurrency=1.
49/49 serial files pass locally. 28/28 verify checks pass.

* fix: quarantine embed-stale test to serial runner

CI shard 9 reported 6 failures, all from the embedStaleForSource describe
block, all ~120-150ms each — classic shared-engine concurrency race shape.
Passes 7/7 locally in isolation.

Root cause: embed-stale.test.ts shares a singleton PGLiteEngine across 7
tests with beforeEach resetPgliteState. Under bun's max-concurrency=4 in
the parallel shard, two tests can interleave their TRUNCATE + seedPage +
upsertChunks + embedStaleForSource flow, so one test's stale-chunk count
sees another test's mid-flight writes.

Same fix as brain-allowlist.serial.test.ts and
schema-pack-find-pack-successors.serial.test.ts: rename to *.serial.test.ts
so the serial runner picks it up at max-concurrency=1.

bun run verify: 28/28 checks pass. 7/7 embed-stale tests pass via serial.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 07:01:28 -07:00

9.2 KiB

Type Taxonomy (v0.41.22: gbrain-base-v2)

The 14-canonical-type DRY/MECE taxonomy shipped in v0.41.22. Predecessor gbrain-base (24 types) stays bundled for back-compat; v0.42+ installs default to gbrain-base-v2.

Why

A production gbrain brain (186K pages) had accreted 94 distinct pages.type values in 9 clusters of redundancy. The type system is the foundation for schema packs, search filtering, extract behavior, enrichment routing, and expert routing. When types are noisy, every downstream feature degrades:

  • Search filtering is ambiguous--type article misses 2.2K articles typed as media/article, sources/article, etc.
  • Enrichment routing is incompleteenrichable_types could only list a few canonical types; 80+ legacy types meant most pages never got enriched.
  • Agent confusion — when ingesting a new article, should it be article, media/article, sources/article, or source/article? Four reasonable choices, none of them right.
  • Orphan inflation — 5,521 concept-redirect pages inflated orphan counts without adding knowledge value.

Issue #1479 catalogues the 9 clusters with exact counts. This doc is the response: a coherent 14-type taxonomy with subtypes/format/origin pushed to frontmatter, alias-table rows for redirects, real link-table rows for edge-shaped pages.

The 14 canonical types (+ note catch-all)

Type Primitive What it holds Examples
person entity People Founders, partners, individuals
company entity Companies, products, orgs (subtype-distinguished) Companies, YC-companies, products
media media Articles, videos, essays, books, podcasts (subtype-distinguished) Substack posts, YouTube videos, books
tweet media Twitter posts (single/bundle/stub subtype) Single tweets, threads, bundles
social-digest temporal Period-grouped social summaries (daily/monthly) X account daily digests
analysis media Research + competitive intel Market analysis, pricing analysis
atom annotation Knowledge units (extraction/manual/lore subtype) Extracted facts, manual notes, lore
concept concept Ideas + reference pages Wiki concepts
source media Transcripts, references Interview transcripts
deal temporal Investment deals Term sheets, investments
email temporal Email threads Email correspondence
slack temporal Slack messages + threads Slack conversations
writing media Original writing Drafts, essays in progress
project concept Initiatives, workstreams Internal projects
note concept Catch-all for one-offs (legacy_type preserved) Memos, anecdotes, insights, etc.

15 types total (14 canonical + note). The catch-all retype rule binds any uncovered legacy type to note with frontmatter.legacy_type = <original> preserved for rollback.

Subtypes (declared in frontmatter post-unify)

Canonical Subtype field Values
company subtype company / product / org
media subtype video / article / essay / book / podcast / blog
tweet subtype single / bundle / stub
social-digest subtype daily / monthly
atom subtype extraction / manual / lore

subtype_field for retype rules is restricted to an allowlist: {subtype, legacy_type, origin, format, kind, period, domain}. This prevents third-party packs from injecting title, slug, or type via mapping_rules (codex D9 security hardening).

Migration flow

gbrain onboard --check                         # surfaces pack_upgrade_available
        ↓
gbrain onboard --check --explain               # per-cluster narrative dry-run
        ↓
gbrain jobs submit unify-types \               # PROTECTED + manual_only
  --allow-protected \
  --params '{"target_pack":"gbrain-base-v2"}'
        ↓
Handler runs 4 phases:
  ┌─────────────────────────────────────┐
  │ Phase 1: Preflight + lock           │ → gbrain-unify db-lock (60min TTL)
  ├─────────────────────────────────────┤
  │ Phase 2: Retype explicit rules      │ → chunked UPDATE 1000/batch
  ├─────────────────────────────────────┤
  │ Phase 3: Retype catch-all sentinel  │ → 'note' with legacy_type
  ├─────────────────────────────────────┤
  │ Phase 4: Page-to-link conversions   │ → insert links + soft-delete
  ├─────────────────────────────────────┤
  │ Phase 5: Page-to-alias conversions  │ → insert slug_aliases + soft-delete
  ├─────────────────────────────────────┤
  │ Phase 6: Final sync (residual)      │ → path-prefix typing
  ├─────────────────────────────────────┤
  │ Phase 7: Flip active pack (D13)     │ → engine.setConfig + saveConfig
  ├─────────────────────────────────────┤
  │ Phase 8: Verify + celebrate         │ → assert ≤16 types; stderr summary
  └─────────────────────────────────────┘
        ↓
gbrain onboard --check                         # pack_upgrade_available cleared
                                               # type_proliferation cleared

Rollback paths

Every primitive ships with a documented rollback:

Operation Rollback
Retype frontmatter.legacy_type = <original> preserved on every page (D8). One SQL UPDATE restores types: UPDATE pages SET type = frontmatter->>'legacy_type' WHERE frontmatter ? 'legacy_type'.
Page-to-link Source page soft-deleted with 72h TTL. gbrain pages restore <slug> within 72h. Link row stays harmless if source restored.
Page-to-alias Source page soft-deleted with 72h TTL. gbrain pages restore <slug> within 72h. Alias row stays harmless (or DELETE FROM slug_aliases WHERE alias_slug = <slug> to clean up).
Active-pack flip gbrain schema use gbrain-base reverses the flip.

What if my brain doesn't fit?

The catch-all retype rule (from_type: '*unknown*') handles long-tail types automatically — any page whose type isn't covered by an explicit rule AND isn't a page_to_link / page_to_alias source gets retyped to note with legacy_type preserved. Guarantees ≤16 distinct types post-unify on ANY brain.

For brains with substantial custom types that deserve their own canonical (e.g. researcher for an academic brain), the right move is:

  1. Fork gbrain-base-v2: gbrain schema fork gbrain-base-v2 my-pack
  2. Edit your fork to add page_types + mapping_rules covering your custom domain.
  3. Target your fork: gbrain jobs submit unify-types --allow-protected --params '{"target_pack":"my-pack"}'

Your fork can also declare migration_from: {pack: gbrain-base-v2, version: "1.x"} to register itself as a successor — future agents discovering your pack via pack_upgrade_available will offer the migration.

Wikilink resolution post-unify

The slug_aliases table IS the resolver (D15: codex outside voice — don't rewrite body-text wikilinks; the alias table is the right primitive). Wikilinks like [[old-redirect-slug]] keep working post- unify because:

  1. The wikilink resolver short-circuits through engine.resolveSlugWithAlias(slug, sourceId) BEFORE the existing fuzzy/prefix cascade.
  2. The lookup queries slug_aliases for any matching alias_slug in the provided source(s).
  3. If found, returns the canonical_slug. The renderer then resolves the wikilink to the canonical page.

Multi-source ambiguity (same alias_slug in two registered sources) emits a once-per-process multi_match stderr warning and returns the first match by source array order. Federated reads pass the full allowed-source array.

Search ranking signal: alias_resolved_boost

Post-unify, search results whose slug is a canonical_slug in slug_aliases get a 1.05x score multiplier via the applyAliasResolvedBoost post-fusion stage. Semantic intent: "user explicitly disambiguated this as canonical, so it should outrank fuzzy matches that hit aliases by accident."

SearchResult.alias_resolved_boost is stamped on touched results for --explain formatter visibility. KNOBS_HASH_VERSION bumped 5→6 to invalidate pre-v0.42 cache rows that don't reflect the new stage.

Reference

  • Issue: https://github.com/garrytan/gbrain/issues/1479
  • Pack file: src/core/schema-pack/base/gbrain-base-v2.yaml
  • Pack-upgrade mechanism: docs/architecture/pack-upgrade-mechanism.md
  • Migration handler: src/core/schema-pack/unify-types-handler.ts
  • Onboard checks: src/core/onboard/checks.ts
  • Skill: skills/schema-unify/SKILL.md
  • Plan + decisions: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md