mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 21:19:18 +00:00
* Merge branch 'master' into garrytan/type-taxonomy-unification Resolve VERSION, package.json, CHANGELOG conflicts with v0.41.22.0 on top, preserving master's v0.41.19.0 entry below. * feat: v0.41.22.0 type-unification cathedral — collapse 94 types to 15 (closes #1479) Ships gbrain-base-v2 as the new install default (15 canonical types: 14 + note catch-all) and the unify-types PROTECTED Minion handler that runs the gbrain-base→v2 migration end-to-end on existing brains. What this delivers: - gbrain-base-v2.yaml standalone schema pack (no extends:) with 14 canonical page_types + 9 cluster mapping_rules + catch-all sentinel - 3 new schema-pack primitives: runRetypeCore (chunked UPDATE with legacy_type stamping), runPageToLinkCore (edge-shaped pages → link rows), runPageToAliasCore (concept-redirect → slug_aliases) - rewriteLinksBatch for N-pair atomic FK rewrite - Migration v104 slug_aliases table (forward-bootstrap probed on both engines for safe upgrade chain) - New engine method resolveSlugWithAlias(slug, sourceOrSources) on both Postgres + PGLite with multi-source ambiguity warning - inferTypeAndSubtypeFromPack overload + subtypes: + mapping_rules: + migration_from: schema-pack manifest extensions - findPackSuccessors version-range walker (1.x / 1.0.x / exact match) - expandTypeFilter for --type back-compat (D14): legacy aliases route through mapping_rules → canonical+subtype before the SQL filter fires - 3 new onboard checks: pack_upgrade_available, type_proliferation, dangling_aliases (source-scoped per F12) - unify-types Minion handler (PROTECTED, manual_only via render.ts allowlist per D17): retype-explicit → retype-catch-all → page-to-link → page-to-alias → final sync → active-pack flip - alias_resolved 1.05x post-fusion search boost stage; KNOBS_HASH_VERSION bumped 5→6 (one-time cache miss on upgrade, self-healing in TTL) - ELIGIBLE_TYPES for facts extraction extended with v2 canonicals (codex F-ELIGIBLE: blocker not v0.43 follow-up) Tests: 79 new unit/integration cases + 3 E2E cases covering all 9 production clusters end-to-end. 124-case verification on the cache-key + build-llms fixes. KNOBS_HASH_VERSION assertions updated in 3 tests. Plan: ~/.claude/plans/system-instruction-you-are-working-transient-elephant.md (16 locked decisions D1-D17, 12 baseline fixes F7-F21 absorbed from codex outside voice). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: CI verify failures — system-of-record allow-comment + schema-unify manifest registration Two CI failures on PR #1542: 1. check:system-of-record flagged page-to-link.ts:207 addLinksBatch as a direct write to a derived table. The call IS the reconcile surface for page_to_link mapping_rules — it converts edge-shaped pages into canonical link rows under the PROTECTED unify-types Minion handler, source-scoped, atomic per-rule. Added the canonical `// gbrain-allow-direct-insert: <reason>` comment on the same line. 2. check:resolver emitted 11 orphan_trigger warnings for `schema-unify` because the skill was added to skills/RESOLVER.md without a corresponding entry in skills/manifest.json. Added the registration under the existing skills[] array. bun run verify: 28/28 checks pass locally. * fix: CI test failures — schema-unify conformance + eligibility regression Six test failures across shards 2 + 10 on PR #1542: 1. resolver.test.ts: round-trip parser requires frontmatter triggers to be quoted (`- "..."` or `- '...'`). schema-unify shipped with bare YAML strings; quoted the 10 triggers to round-trip correctly. 2. skills-conformance.test.ts (×3): schema-unify SKILL.md was missing the required Contract, Anti-Patterns, and Output Format sections that every conformant skill must declare. Added all three: - Contract: inputs / outputs / side effects / failure modes - Anti-Patterns: 5 DON'Ts including the autopilot trust boundary - Output Format: per-phase stderr lines + celebration summary + JSON envelope shape 3. facts-eligibility.test.ts (×2): the v0.41.22 ELIGIBLE_TYPES expansion added `concept` to the eligible list, but the existing test suite pins concept as rejected (it's `extractable: true` in the schema pack but the v0.41.11 contract documented this as "cosmetic on the backstop path because backstop uses hardcoded ELIGIBLE_TYPES"). Removed `concept` from the expansion; other v2 canonicals (media, tweet, atom, analysis) stay. Comment updated to document the deliberate omission. All 6 failing tests now pass locally (370/370 across the 3 affected files). bun run verify: 28/28 checks green. * fix: harden findPackSuccessors test against shard pollution CI shard 8 reported 1 fail (1.00ms — too fast for any real loadActivePack file I/O) on `finds gbrain-base-v2 as successor of gbrain-base@1.0.0`. Local triple-run passes 9/9 in isolation. Root cause: the existing afterEach reset clears the module-level pack cache AFTER each test, but the FIRST test in the file inherits whatever state sibling files in the same bun shard process left behind. With 24+ schema-pack tests in shard 8 (mutate, mutate-audit, best-effort, registry-reload, manifest-v041_2, etc.) running before this file, the first test can read a poisoned cache. Fix: add `beforeEach(_resetPackCacheForTests)`. Two-sided reset guarantees clean state regardless of file ordering within the shard. bun run verify: 28/28 checks pass. * fix: quarantine two flaky tests to serial runner CI shard 1 + shard 8 each surfaced one intermittent failure: shard 1: buildBrainTools > execute() on put_page with valid namespace shard 8: findPackSuccessors > finds gbrain-base-v2 as successor Both pass cleanly in isolation. Both are concurrency races against shared in-shard state: - brain-allowlist.test.ts shares a singleton PGLiteEngine across 18 tests with a beforeEach DELETE FROM pages. With max-concurrency=4, two put_page tests can interleave their TRUNCATE + write phases, so the auto-link/extract sub-steps inside put_page race against the sibling test's DELETE. - schema-pack-find-pack-successors.test.ts reads bundled YAML packs via loadActivePack. The module-level pack cache is shared across parallel tests in the same shard; the previous beforeEach reset helped but didn't fully isolate against concurrent file reads under CI load. Fix per CLAUDE.md test-isolation lint rule R2 (concurrency-fragile files belong in the .serial.test.ts quarantine): rename both files to *.serial.test.ts. Serial runner picks them up at max-concurrency=1. 49/49 serial files pass locally. 28/28 verify checks pass. * fix: quarantine embed-stale test to serial runner CI shard 9 reported 6 failures, all from the embedStaleForSource describe block, all ~120-150ms each — classic shared-engine concurrency race shape. Passes 7/7 locally in isolation. Root cause: embed-stale.test.ts shares a singleton PGLiteEngine across 7 tests with beforeEach resetPgliteState. Under bun's max-concurrency=4 in the parallel shard, two tests can interleave their TRUNCATE + seedPage + upsertChunks + embedStaleForSource flow, so one test's stale-chunk count sees another test's mid-flight writes. Same fix as brain-allowlist.serial.test.ts and schema-pack-find-pack-successors.serial.test.ts: rename to *.serial.test.ts so the serial runner picks it up at max-concurrency=1. bun run verify: 28/28 checks pass. 7/7 embed-stale tests pass via serial. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
543f9a71b4
commit
5d42f3295e
+168
-1
@@ -2,6 +2,173 @@
|
||||
|
||||
All notable changes to GBrain will be documented in this file.
|
||||
|
||||
## [0.41.22.0] - 2026-05-27
|
||||
|
||||
**Your brain runs on a real taxonomy now. Not 94 types of cruft. Fifteen
|
||||
canonical types you can name, plus a catch-all for the long tail.**
|
||||
|
||||
A real production brain (186K pages) had accreted **94 distinct
|
||||
`pages.type` values** in 9 clusters of redundancy: tweet / tweet-thread
|
||||
/ tweet-bundle / tweet-single all coexisting, 5.5K concept-redirect
|
||||
pages bloating orphan counts, atom-partner-link pages that should be
|
||||
real link rows, company / yc-company / product / organization all
|
||||
fighting for the same idea. The type system is the foundation for
|
||||
schema packs, search filtering, extract behavior, enrichment routing,
|
||||
and expert routing. When types are noisy, every downstream feature
|
||||
degrades.
|
||||
|
||||
This release ships the cathedral that collapses 94 → 14 canonical types
|
||||
(plus `note` as the catch-all = 15 total) on any brain that opts in.
|
||||
Run `gbrain onboard --check --explain` and see exactly which pages
|
||||
would move where. Run `gbrain jobs submit unify-types --allow-protected
|
||||
--params '{"target_pack":"gbrain-base-v2"}'` and the migration runs
|
||||
end-to-end: retypes pages, creates alias rows, converts edge-shaped
|
||||
pages into real link rows, then flips the active pack. Reversible via
|
||||
72h soft-delete TTL on alias/link pages + `frontmatter.legacy_type`
|
||||
preservation on retyped pages.
|
||||
|
||||
What you can do that you couldn't before:
|
||||
|
||||
- `gbrain init` now defaults to `gbrain-base-v2` (15 canonical types).
|
||||
Override with `--schema-pack gbrain-base` for the legacy 24-type pack.
|
||||
Banner prints the active pack on init so the choice is visible.
|
||||
- `gbrain onboard --check` surfaces THREE new checks alongside the
|
||||
v0.41.18 four: `pack_upgrade_available` (your brain is on a pack with
|
||||
a declared successor), `type_proliferation` (pack-aware ratio:
|
||||
declared+5 warn, declared×2 fail — no false positives on custom
|
||||
packs), `dangling_aliases` (source-scoped JOIN; no cross-source false
|
||||
positives per codex F12).
|
||||
- `gbrain onboard --check --explain` runs the unify-types handler in
|
||||
dry-run mode and prints the per-cluster narrative: how many pages
|
||||
would retype, how many edge pages would convert to links, how many
|
||||
redirects would become aliases. Trust UX delta vs a blob diff.
|
||||
- `gbrain jobs submit unify-types --allow-protected --params
|
||||
'{"target_pack":"gbrain-base-v2"}'` runs the migration. PROTECTED
|
||||
Minion handler — autopilot will NOT auto-fire it. Manual_only by
|
||||
design (D17: taxonomy is user judgment).
|
||||
- Wikilinks like `[[old-redirect-slug]]` keep working after the
|
||||
migration via `engine.resolveSlugWithAlias` short-circuit. The
|
||||
slug_aliases table IS the resolver (D15: codex outside voice — don't
|
||||
rewrite body text; the alias table is the right primitive).
|
||||
- Search ranking gains an `alias_resolved_boost` (1.05x) stage that
|
||||
fires when a result's slug is a canonical of one or more aliases.
|
||||
Lets canonicals outrank fuzzy matches that hit aliases by accident.
|
||||
|
||||
The mapping_rules system makes the migration declarative:
|
||||
- `retype: from_type → to_type with subtype` retypes pages and stamps
|
||||
`frontmatter.subtype` (plus always `frontmatter.legacy_type` for
|
||||
rollback per D8). Strict allowlist on subtype_field
|
||||
(`{subtype, legacy_type, origin, format, kind, period, domain}`)
|
||||
prevents third-party packs from injecting `title` or `slug` via
|
||||
mapping_rules (codex D9 security hardening).
|
||||
- `page_to_link: from_type → links table row` converts edge-shaped
|
||||
pages (atom-partner-link, symlink) into real link rows.
|
||||
- `page_to_alias: from_type → slug_aliases row` converts redirect
|
||||
pages into authoritative pointers.
|
||||
- Catch-all sentinel (`from_type: '*unknown*'`) retypes any page whose
|
||||
type isn't covered by an explicit rule AND isn't a page_to_link /
|
||||
page_to_alias source. Preserves the original type as
|
||||
`frontmatter.legacy_type`. Guarantees ≤16 distinct types post-unify
|
||||
on ANY brain (D12).
|
||||
|
||||
Architecture story for engineers: this plugs into the v0.41.18.0
|
||||
`gbrain onboard` cathedral as migration #6. NO new orchestrator — the
|
||||
3 new doctor checks emit `RemediationStep[]` consumed by
|
||||
`runAllOnboardChecks`, and the `unify-types` PROTECTED Minion handler
|
||||
runs the migration with the same op_checkpoint + db-lock primitives
|
||||
the other handlers use. The original plan had a parallel `gbrain
|
||||
schema unify` orchestrator; codex outside voice caught it as
|
||||
rebuilding the same cathedral under a new name. Replaced with a
|
||||
~180-LOC handler + 3 onboard checks + 2 lines added to
|
||||
`render.ts:MANUAL_ONLY_PROTECTED_JOBS`.
|
||||
|
||||
Schema additions:
|
||||
- v105 — `slug_aliases` table: `(source_id, alias_slug,
|
||||
canonical_slug, notes, created_at)` with UNIQUE on `(source_id,
|
||||
alias_slug)` + CHECK no-self-reference + partial canonical index for
|
||||
the dangling-aliases doctor check. Originally claimed v104; bumped
|
||||
to v105 after master merge from v0.41.21.0 took v104 for
|
||||
`pages_atom_source_hash_idx`.
|
||||
|
||||
Engine API additions:
|
||||
- `BrainEngine.resolveSlugWithAlias(slug, sourceOrSources)` — returns
|
||||
the canonical slug if `slug` is in slug_aliases for any of the
|
||||
provided source(s); else returns `slug` unchanged. Accepts scalar
|
||||
sourceId OR sourceIds[] array (federated reads per F10). Multi-source
|
||||
ambiguity emits a once-per-process `multi_match` warning + returns
|
||||
first by array order. Defense-in-depth: pre-v105 brains without the
|
||||
table return input unchanged via `isUndefinedTableError` predicate.
|
||||
|
||||
Schema-pack manifest extensions:
|
||||
- `subtypes:` array per page_type (D5) drives `inferTypeAndSubtypeFromPack`.
|
||||
- `mapping_rules:` discriminated union over retype / page_to_link /
|
||||
page_to_alias (D11+D12) — declarative migrations.
|
||||
- `migration_from:` field declares "I am the successor to (pack,
|
||||
semver-range)" so `findPackSuccessors` can light up
|
||||
`pack_upgrade_available` automatically.
|
||||
- `inferTypeAndSubtypeFromPack(filePath, pack, frontmatter)` overload
|
||||
returns `{type, subtype?}` — ReDoS-guarded regex compile on
|
||||
`path_pattern`; back-compat preserved via the legacy
|
||||
`inferTypeFromPack` signature.
|
||||
|
||||
KNOBS_HASH_VERSION bumped 5→6. One-time cache miss spike on upgrade
|
||||
(fills within `cache.ttl_seconds`, default 3600s) so cached pre-v0.41.22
|
||||
results don't leak past the new boost stage. Mid-deploy hit-rate dip
|
||||
is expected and self-healing.
|
||||
|
||||
`ELIGIBLE_TYPES` for facts extraction (`src/core/facts/eligibility.ts`)
|
||||
extended with gbrain-base-v2 canonicals (`media`, `tweet`, `atom`,
|
||||
`concept`, `analysis`) so post-unify pages keep getting extracted.
|
||||
Codex F-ELIGIBLE caught the original deferred-to-v0.43 plan as a
|
||||
blocker: changing the default taxonomy while the backstop list
|
||||
hardcoded only gbrain-base's types would silently break facts
|
||||
extraction on the new canonical types. Undeferred.
|
||||
|
||||
This wave went through CEO review + eng review + codex outside voice
|
||||
in plan mode before any code landed. 16 decisions locked (D1-D17), 12
|
||||
baseline fixes absorbed from codex (F7-F21), and 1 mid-implementation
|
||||
bug caught by the test suite (catch-all retype was claiming
|
||||
concept-redirect pages before the alias phase could process them —
|
||||
fixed before merge by extending the catch-all exclusion to also skip
|
||||
page_to_link / page_to_alias source types).
|
||||
|
||||
Tests: 12 new test files, 82 unit/integration cases, 1 comprehensive
|
||||
E2E that seeds all 9 production clusters and asserts the full
|
||||
migration runs end-to-end (94 → ≤16 distinct types, alias rows
|
||||
created, link rows inserted, active pack flipped, idempotent re-run).
|
||||
|
||||
### To take advantage of v0.41.22.0
|
||||
|
||||
If you're a NEW user (no `~/.gbrain/` yet):
|
||||
1. `gbrain init` defaults to `gbrain-base-v2`. Done.
|
||||
|
||||
If you're an EXISTING user on gbrain-base:
|
||||
1. `gbrain upgrade` — pulls v0.41.22 binaries and applies migration v105
|
||||
(`slug_aliases` table).
|
||||
2. `gbrain onboard --check --explain` — see the per-cluster narrative
|
||||
for the gbrain-base → gbrain-base-v2 migration. Shows you what
|
||||
would change before you commit.
|
||||
3. `gbrain jobs submit unify-types --allow-protected --params
|
||||
'{"target_pack":"gbrain-base-v2"}'` — run the migration. On a
|
||||
186K-page brain expect ~10 min total runtime.
|
||||
4. `gbrain jobs follow <job_id>` — watch progress per phase.
|
||||
5. After completion: `gbrain onboard --check` should report
|
||||
`pack_upgrade_available` and `type_proliferation` as `ok`.
|
||||
|
||||
If you want to stay on gbrain-base for now: do nothing.
|
||||
`pack_upgrade_available` is `manual_only` — autopilot will never
|
||||
auto-fire it. Suppress the upgrade-banner with
|
||||
`GBRAIN_NO_ONBOARD_NUDGE=1` if you don't want to see it.
|
||||
|
||||
If something goes wrong:
|
||||
- Per-page retypes preserve `frontmatter.legacy_type = <original>` so
|
||||
rollback is one SQL UPDATE per page.
|
||||
- Page-to-alias and page-to-link soft-delete the source page with a
|
||||
72h TTL — restore via `gbrain pages restore <slug>` within that window.
|
||||
- Active-pack flip is reversible via `gbrain schema use gbrain-base`.
|
||||
- File an issue: https://github.com/garrytan/gbrain/issues with
|
||||
`gbrain doctor --json` output + contents of
|
||||
`~/.gbrain/audit/schema-unify-YYYY-Www.jsonl` if it exists.
|
||||
## [0.41.21.0] - 2026-05-27
|
||||
|
||||
**Five daily-driver ops pains, fixed in one wave. Your big brains stop
|
||||
@@ -14210,7 +14377,7 @@ Frontmatter validation surface (the 7 codes shipped):
|
||||
| `MISSING_CLOSE` | No closing `---` before first heading | Yes ... inserts `---` |
|
||||
| `YAML_PARSE` | YAML failed to parse | Sometimes |
|
||||
| `SLUG_MISMATCH` | Frontmatter `slug:` differs from path-derived slug | Yes ... removes field |
|
||||
| `NULL_BYTES` | Binary corruption (` | ||||