mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* Wave A: schema + receipts foundation for v0.42 extract operator surfaces
Foundation layer for the pack-driven extractables + receipt-as-brain-memory
+ operator-discoverability cathedral. Five atomic pieces ship together
because their schema + helpers + module dependencies are tight-coupled:
A1. Widen pack manifest's `extractable` from `boolean` to
`boolean | ExtractableSpec`. ExtractableSpec carries prompt_template,
fixture_corpus, eval_dimensions, benchmark_min_recall, and reserves
verifier_path for v0.43+ pack-shipped verifier code (REFUSE at
runtime in v0.42 per plan D-EXTRACT-37). Back-compat: every pre-v0.42
pack with `extractable: true` continues parsing unchanged. Three new
helpers: extractableSpecsFromPack(), getExtractableSpec(),
refuseVerifierPathInV042().
A2. New page type `extract_receipt` in ALL_PAGE_TYPES. Source-boost map
adds `extracts/` prefix at factor 0.3 — receipts surface in search
when extraction-relevant but never dominate user content (D-EXTRACT-42).
A3. New module src/core/extract/receipt-writer.ts (~190 LOC) exporting
writeReceipt(engine, input). Canonical slug shape
extracts/{date}/{kind}/{source_id}/{run_id_short}/round-{N} per
D-EXTRACT-17. Frontmatter belt+suspenders per D-EXTRACT-19: BOTH
type:extract_receipt AND dream_generated:true stamped on every
receipt, regardless of caller, so the eligibility predicate's
anti-loop guards reject the receipt page from any future extraction
sweep (single-flag bypass requires breaking two unrelated checks).
Idempotent on resume — same run_id+round overwrites cleanly.
A4. Migration v104 creates extract_rollup_7d table (per-day rollup of
extract events keyed on kind+source_id+day). Audit JSONL stays the
SOURCE OF TRUTH per F-OUT-19; this table is a best-effort cache for
doctor's <100ms read budget. Per-day rows mean the 7-day window
auto-evicts on every read. v100 was deliberately skipped on master
(renumbered out during a prior wave); v101/v102/v103 also taken;
v104 is the next clean slot.
A5. Doctor `extract_health` check reads extract_rollup_7d for last 7
days and emits per-kind aggregates: cost_7d_usd, eval_pass_count,
eval_fail_count, halt_count, round_completed_count, halt_rate.
3-state: OK when rollup empty (pre-v0.42 brain or fresh init), WARN
when any per-kind halt rate > 10% (top-3 named in message), WARN
when rollup_write_failures > 0 (audit JSONL is SoT but operator
deserves to know the DB cache is degraded). Pre-v104 brains stay
quiet — the missing-table error path is caught and treated as
OK so doctor doesn't warn during the upgrade window.
Tests added:
- test/extractable-spec-widening.test.ts (22 cases) — back-compat with
boolean shape, new struct parsing, verifier_path REFUSE contract.
- test/extract/receipt-writer.test.ts (12 cases) — slug shape, frontmatter
belt+suspenders, idempotent resume, body human-readability.
- test/doctor-extract-health.test.ts (8 cases) — empty rollup OK, halt
rate WARN, rollup_write_failures WARN, 7-day window inclusion at
boundary, multi-kind top-3 message ordering.
Plus the canonical bootstrap-coverage test passes with the new v104
migration cleanly applied through both engines.
Plan: ~/.claude/plans/system-instruction-you-are-working-stateless-dragonfly.md
Wave A scope. Wave B (hook receipts into existing extractors) follows.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Wave B: hook receipts + rollup row into the 5 shipped extractors
Each LLM-backed extractor surface now records its run in two places when
something actually happened:
1. An extract receipt PAGE at extracts/{date}/{kind}/{source_id}/{run_id_short}/round-{N}
(queryable via gbrain search, citable, surfaces in cross-modal
contradiction probes per the Wave A foundation). Only written when
`total_rows > 0` so no-op runs don't bloat the brain.
2. An UPSERT row in extract_rollup_7d (DB-backed best-effort cache
per F-OUT-19) so the doctor extract_health check from Wave A reads
per-kind aggregates without scanning JSONL.
New module src/core/extract/rollup-writer.ts (~120 LOC) exports
upsertExtractRollup() with PostgreSQL ON CONFLICT DO UPDATE on the
(kind, source_id, day) PK. Concurrency-safe per F-OUT-14 design.
Failure path is best-effort — bumps rollup_write_failures in the
table itself, stderr-warns once per (kind, day, error-class), and
NEVER fails the parent extraction operation. JSONL remains source
of truth.
Wired into 5 extractors:
- extract-conversation-facts (kind: facts.conversation) — both
success path AND BudgetExhausted halt path write receipt+rollup
so partial runs are still observable.
- extract_atoms cycle phase (kind: atoms)
- synthesize_concepts cycle phase (kind: concepts, source_id: default
because concepts are brain-global)
- propose_takes cycle phase (kind: takes.proposed) — scope-aware
source_id from the read scope.
- extract_facts cycle phase (kind: facts.fence) — deterministic
(no LLM cost) but still records reconcile activity so doctor sees
the cycle is alive.
Receipt frontmatter belt+suspenders (D-EXTRACT-19) reused from
Wave A: every receipt stamps BOTH `type: extract_receipt` AND
`dream_generated: true` so the eligibility predicate's anti-loop
guards reject the receipt page from any future extraction sweep.
Test surgery in test/propose-takes.test.ts — one existing assertion
tightened from "no INSERTs" to "no INSERT INTO take_proposals" so
the new rollup UPSERT doesn't falsely fail the cache-hit case test.
Run regression: 85/85 tests pass across extract-conversation-facts,
extract-atoms-synthesize-concepts, extract-facts-phase, propose-takes.
Plan: ~/.claude/plans/system-instruction-you-are-working-stateless-dragonfly.md
Wave B scope. Wave C (pack-author scaffolding + benchmark) follows.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Wave C+D: pack-author scaffolding + operator surfaces for v0.42 extract
Wave C: pack-author authoring loop
- scaffold-extractable mutation primitive declares a kind as extractable
on a pack manifest in one verb (wires through updateTypeOnPack from
the v0.41 mutate library); generates 5 placeholder fixtures + a
pack-supplied prompt template stub
- schema CLI wires gbrain schema scaffold-extractable <type> --pack <pack>
- extract benchmark CLI loads a pack's fixture corpus through strict
D-EXTRACT-21 path validation (rejects absolute paths, .. traversal,
null bytes, symlinks resolving outside pack root); v0.42 ships as a
stub reporter (LLM dispatch deferred to Wave E)
Wave D: operator surfaces
- extract status CLI reads extract_rollup_7d for the last 7 days,
sorts by (halt_rate desc, cost desc); kubectl-style right-aligned
table, top-5 + "more rows" hint by default, --verbose shows all;
stable schema_version: 1 JSON envelope for monitoring pipelines
- extract --explain <kind> CLI prints the active pack's resolution
chain: declaration source (pack-declared vs built-in cycle phase),
prompt_template + fixture_corpus paths with existence checks,
eval_dimensions, benchmark_min_recall, and the last 7d rollup
- extract.ts gains a lifecycle-grouped help text (Extraction /
Inspection / Status) per the original D3 plan goal
Tests:
- test/schema-pack/scaffold-extractable.test.ts (15 cases) including
explicit privacy-rule assertions guarding against real-name leakage
- test/extract/benchmark.test.ts (17 cases) covering path validation
rejections + JSONL fixture parsing
- test/extract/status.test.ts (15 cases) over pure aggregation +
formatting
Housekeeping:
- test/extract/receipt-writer.test.ts refactored to the canonical
PGLite block (beforeAll/afterAll/resetPgliteState in beforeEach)
per CLAUDE.md test-isolation R3+R4; runtime drops from ~30s of
99-migration replay per test to <6s for all 12 cases together
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* v0.42.0.0: extract operator surfaces + pack-driven extractables
Bump VERSION + package.json to 0.42.0.0. CHANGELOG entry covers the
three-wave shipped scope (receipts + rollup + doctor check; receipts
hooked into all 5 shipped extractors; pack-author scaffolding +
benchmark stub-reporter; status + --explain dashboards + lifecycle
help). CLAUDE.md Key Files gains a v0.42 cluster annotation. llms.txt
regenerated.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* v0.41.23.0: re-tag from 0.42.0.0 (patch-channel slot, no scope change)
VERSION + package.json + CHANGELOG header + CLAUDE.md cluster annotation
all moved from 0.42.0.0 to 0.41.23.0. Body text updated in-place: every
"v0.42" / "v0.43+" reference inside this entry's release notes now reads
"v0.41.23" or "follow-up release" as appropriate.
Same scope shipping — the three-wave extract operator surface stays
intact. Just lands in the patch-channel queue (.20/.21/.23 free; .22 is
PR #1542's type-unification cathedral) instead of the minor-channel bump.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: add extract_receipt to gbrain-base.yaml page_types (CI parity gate)
CI shard 5 caught the drift: test/regressions/gbrain-base-equivalence.test.ts
asserts every ALL_PAGE_TYPES seed has a matching page_type entry in the
gbrain-base.yaml pack. Wave A added `extract_receipt` to ALL_PAGE_TYPES
but didn't seed it in the base pack manifest.
Adds the entry under the `annotation` primitive with `extracts/` path
prefix (matches the source-boost demote site) and `extractable: false`
(receipts are written by the framework, never extracted from). Comment
documents the belt+suspenders D-EXTRACT-19 invariant so future readers
understand why receipts carry both `dream_generated: true` AND
`type: extract_receipt`.
Closes the CI gate without changing runtime behavior — the pack-aware
read paths already had the prefix demote wired in src/core/search/source-boost.ts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: bump gbrain-base page-type count 24→25 in schema-cli test
CI shard 4 caught the second drift from the same root cause as the
prior parity-gate fix: v0.41.23's `extract_receipt` addition bumped
gbrain-base.yaml from 24 to 25 page types. The schema-cli smoke test
was pinned at 24 (the count after v0.41.11.0 added `conversation` +
`atom`); update to 25 and note v0.41.23's contribution alongside the
prior version stamp.
Verified hermetic: running test/schema-cli.test.ts with a clean
GBRAIN_HOME tempdir produces 12/12 pass (the local-machine 'schema
active' fail is from a real ~/.gbrain pinning gbrain-base-v2; not a
shipped-code issue, doesn't repro on CI).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(test): pack-locator stub leak between shard 6 test files
CI shard 6 caught three flaky failures in test/onboard-pack-upgrade-checks.test.ts:
- checkPackUpgradeAvailable > fires on gbrain-base brain with gbrain-base-v2
- checkPackUpgradeAvailable > manual_only routing via render.ts allowlist (D17)
- checkTypeProliferation > warns when distinct types exceed declared+5
Root cause: test/schema-pack-sync.test.ts calls
`__setPackLocatorForTests(...)` to stub the disk-loader, but doesn't
restore in afterAll. Bun's CI shard 6 loads multiple test files into
one process; when sync.test.ts runs before onboard-pack-upgrade-checks.test.ts,
the stubbed locator persists at module scope. `loadActivePack` for
gbrain-base / gbrain-base-v2 then returns null and:
- findPackSuccessors returns [] → status='ok' instead of 'warn' (F1+F2)
- declared falls back to 15 → fail threshold becomes 30, 32 > 30 → 'fail'
instead of 'warn' (F3)
Local single-file runs pass because the locator starts at its default.
Two-layer fix:
1. test/schema-pack-sync.test.ts afterAll calls
`_resetPackLocatorForTests()` to undo the mutation (the canonical
fix at the source).
2. test/onboard-pack-upgrade-checks.test.ts beforeEach calls the same
reset (defense-in-depth against any future test file in the shard
that forgets to restore).
Reproduced locally: running the three shard-6 schema-pack files together
fails 3 tests pre-fix and passes 30/30 post-fix. Full shard 6 sweep
(77 files, 1232 tests) now green; bun run verify still 28/28.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(test): pglite-engine — dim-agnostic chunk-embedding test data
CI shard 6 caught two flaky failures in test/pglite-engine.test.ts:
- PGLiteEngine: Chunks > getChunksWithEmbeddings returns embedding data
- PGLiteEngine: stale chunk pagination > countStaleChunks counts chunks
with NULL embedding only
Both failed with `expected 1280 dimensions, not 1536` at the upsert site.
Root cause: pglite-engine.ts:287 initSchema() reads embedding dim from
gw.getEmbeddingDimensions() if the gateway is configured (potentially
left in that state by another shard-6 test file in the same bun process),
falling back to DEFAULT_EMBEDDING_DIMENSIONS otherwise — which is 1280
since v0.36+ when the ZE default landed (zeroentropyai:zembed-1).
Pre-v0.36 defaults were OpenAI's 1536; my test data was pinned to that
stale literal.
The two outcomes that pass:
- gateway happens to be configured for 1536-dim (e.g. master shard 6
run 26515999465 — these tests passed at 20ms + 24ms with no
"dimensions" error)
- gateway happens to be configured for 1280-dim AND test data is 1280
The outcome that fails:
- gateway configured for 1280-dim AND test data hardcoded to 1536
Fix: capture the actual column width after initSchema (probe
pg_attribute.atttypmod for content_chunks.embedding) and use that
captured `CHUNK_EMBED_DIM` constant at the three Float32Array sites.
Test data now matches whatever width the column was created at,
regardless of which shard-6 file ran first.
Local repro: full shard 6 (77 files, 1232 tests, ~6min) green; this
file standalone (100 tests) green; bun run verify 28/28.
Broader pattern: 9 other test files use the same Float32Array(1536)
literal. None land in shard 6 today (so they don't flake), but the
fix shape here can be lifted into a shared helper if the bug class
surfaces elsewhere — filed as a v0.42+ follow-up rather than a
preemptive sweep, since each file's setup shape is slightly different.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
278 lines
11 KiB
TypeScript
278 lines
11 KiB
TypeScript
// v0.40.6.0 — sync.ts contract tests.
|
||
//
|
||
// Dry-run vs apply, chunked-batch correctness, idempotency,
|
||
// soft-delete exclusion, sample-slug payload, dead-prefix hint,
|
||
// per-source write scoping, PGLite parity. Phase 3 of the cathedral.
|
||
|
||
import { afterAll, beforeAll, beforeEach, describe, expect, it } from 'bun:test';
|
||
import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs';
|
||
import { tmpdir } from 'node:os';
|
||
import { join } from 'node:path';
|
||
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
||
import { resetPgliteState } from './helpers/reset-pglite.ts';
|
||
import { runSyncCore } from '../src/core/schema-pack/sync.ts';
|
||
import {
|
||
__setPackLocatorForTests,
|
||
_resetPackLocatorForTests,
|
||
} from '../src/core/schema-pack/load-active.ts';
|
||
import { _resetPackCacheForTests } from '../src/core/schema-pack/registry.ts';
|
||
import type { OperationContext } from '../src/core/operations.ts';
|
||
import { withEnv } from './helpers/with-env.ts';
|
||
|
||
let engine: PGLiteEngine;
|
||
let tmpDir: string;
|
||
|
||
beforeAll(async () => {
|
||
engine = new PGLiteEngine();
|
||
await engine.connect({});
|
||
await engine.initSchema();
|
||
});
|
||
|
||
afterAll(async () => {
|
||
// Restore the disk-loader so we don't leak our test stub into sibling
|
||
// test files in the same shard process (closes the bug where
|
||
// test/onboard-pack-upgrade-checks.test.ts saw a stubbed locator and
|
||
// failed only when this file ran first in CI shard 6).
|
||
_resetPackLocatorForTests();
|
||
await engine.disconnect();
|
||
});
|
||
|
||
beforeEach(async () => {
|
||
await resetPgliteState(engine);
|
||
_resetPackCacheForTests();
|
||
_resetPackLocatorForTests();
|
||
tmpDir = mkdtempSync(join(tmpdir(), 'gbrain-sync-test-'));
|
||
});
|
||
|
||
function ctxOf(remote = false): OperationContext {
|
||
return {
|
||
engine,
|
||
config: {},
|
||
logger: { info: () => {}, warn: () => {}, error: () => {} },
|
||
dryRun: false,
|
||
remote,
|
||
} as unknown as OperationContext;
|
||
}
|
||
|
||
async function ensureSource(id: string): Promise<void> {
|
||
if (id === 'default') return;
|
||
await engine.executeRaw(
|
||
`INSERT INTO sources (id, name) VALUES ($1, $1) ON CONFLICT (id) DO NOTHING`,
|
||
[id],
|
||
);
|
||
}
|
||
|
||
async function seedPage(slug: string, sourcePath: string, opts: { type?: string; sourceId?: string; deleted?: boolean } = {}): Promise<void> {
|
||
const sourceId = opts.sourceId ?? 'default';
|
||
await ensureSource(sourceId);
|
||
await engine.executeRaw(
|
||
`INSERT INTO pages (slug, source_id, source_path, type, title, compiled_truth, timeline, content_hash, deleted_at)
|
||
VALUES ($1, $2, $3, $4, $5, '', '', '', $6)`,
|
||
[slug, sourceId, sourcePath, opts.type ?? '', slug, opts.deleted ? new Date() : null],
|
||
);
|
||
}
|
||
|
||
async function getType(slug: string): Promise<string | null> {
|
||
const rows = await engine.executeRaw<{ type: string }>(
|
||
`SELECT type FROM pages WHERE slug = $1`,
|
||
[slug],
|
||
);
|
||
return rows[0]?.type ?? null;
|
||
}
|
||
|
||
function seedTinyPack(types: Array<{ name: string; prefix: string }>): void {
|
||
const dir = join(tmpDir, 'tiny');
|
||
mkdirSync(dir, { recursive: true });
|
||
const path = join(dir, 'pack.yaml');
|
||
let body = `api_version: gbrain-schema-pack-v1\nname: tiny\nversion: 1.0.0\ndescription: ""\ngbrain_min_version: 0.38.0\nextends: null\nborrow_from: []\npage_types:\n`;
|
||
for (const t of types) {
|
||
body += ` - name: ${t.name}\n primitive: entity\n path_prefixes:\n - ${t.prefix}\n aliases: []\n extractable: false\n expert_routing: false\n`;
|
||
}
|
||
body += `link_types: []\nfrontmatter_links: []\ntakes_kinds:\n - fact\n - take\n - bet\n - hunch\nenrichable_types: []\nfiling_rules: []\n`;
|
||
writeFileSync(path, body, 'utf-8');
|
||
__setPackLocatorForTests((name) => (name === 'tiny' ? path : null));
|
||
}
|
||
|
||
describe('runSyncCore — dry-run', () => {
|
||
it('returns would_apply count + sample_slugs without writing', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md');
|
||
await seedPage('bob', 'people/bob.md');
|
||
const result = await runSyncCore(ctxOf(), { apply: false });
|
||
expect(result.apply).toBe(false);
|
||
expect(result.total_would_apply).toBe(2);
|
||
expect(result.total_applied).toBe(0);
|
||
const personEntry = result.per_prefix.find((p) => p.type === 'person')!;
|
||
expect(personEntry.would_apply).toBe(2);
|
||
expect(personEntry.applied).toBe(0);
|
||
expect(personEntry.sample_slugs).toEqual(['alice', 'bob']);
|
||
// Confirm types still empty after dry-run.
|
||
expect(await getType('alice')).toBe('');
|
||
expect(await getType('bob')).toBe('');
|
||
});
|
||
});
|
||
|
||
it('sample_slugs capped at 10', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
for (let i = 0; i < 15; i++) {
|
||
await seedPage(`p${String(i).padStart(2, '0')}`, `people/p${String(i).padStart(2, '0')}.md`);
|
||
}
|
||
const result = await runSyncCore(ctxOf(), { apply: false });
|
||
const personEntry = result.per_prefix.find((p) => p.type === 'person')!;
|
||
expect(personEntry.would_apply).toBe(15);
|
||
expect(personEntry.sample_slugs.length).toBe(10);
|
||
});
|
||
});
|
||
|
||
it('dead_prefix flag fires when no pages match', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([
|
||
{ name: 'person', prefix: 'people/' },
|
||
{ name: 'company', prefix: 'companies/' },
|
||
]);
|
||
await seedPage('alice', 'people/alice.md');
|
||
const result = await runSyncCore(ctxOf(), { apply: false });
|
||
const companyEntry = result.per_prefix.find((p) => p.type === 'company')!;
|
||
expect(companyEntry.dead_prefix).toBe(true);
|
||
expect(companyEntry.would_apply).toBe(0);
|
||
});
|
||
});
|
||
});
|
||
|
||
describe('runSyncCore — apply', () => {
|
||
it('updates page.type for matching untyped pages', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md');
|
||
await seedPage('bob', 'people/bob.md');
|
||
const result = await runSyncCore(ctxOf(), { apply: true });
|
||
expect(result.total_applied).toBe(2);
|
||
expect(await getType('alice')).toBe('person');
|
||
expect(await getType('bob')).toBe('person');
|
||
});
|
||
});
|
||
|
||
it('idempotent: second apply is a no-op', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md');
|
||
const first = await runSyncCore(ctxOf(), { apply: true });
|
||
const second = await runSyncCore(ctxOf(), { apply: true });
|
||
expect(first.total_applied).toBe(1);
|
||
expect(second.total_applied).toBe(0);
|
||
});
|
||
});
|
||
|
||
it('chunked UPDATE: large set in 1000-row batches (perf-shape verification)', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
// Seed 1500 untyped pages (1.5× the default batch size).
|
||
for (let i = 0; i < 1500; i++) {
|
||
await seedPage(`p${String(i).padStart(4, '0')}`, `people/p${String(i).padStart(4, '0')}.md`);
|
||
}
|
||
const batches: number[] = [];
|
||
const result = await runSyncCore(ctxOf(), {
|
||
apply: true,
|
||
batchSize: 1000,
|
||
onProgress: (info) => batches.push(info.appliedSoFar),
|
||
});
|
||
expect(result.total_applied).toBe(1500);
|
||
// Progress fired at least once for the first batch (1000) and
|
||
// again for the tail (1500 total).
|
||
expect(batches).toContain(1000);
|
||
expect(batches).toContain(1500);
|
||
});
|
||
});
|
||
|
||
it('does NOT touch pages that already have a type', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md', { type: 'old-type' });
|
||
await seedPage('bob', 'people/bob.md'); // untyped
|
||
await runSyncCore(ctxOf(), { apply: true });
|
||
expect(await getType('alice')).toBe('old-type'); // preserved
|
||
expect(await getType('bob')).toBe('person');
|
||
});
|
||
});
|
||
|
||
it('excludes soft-deleted pages', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md');
|
||
await seedPage('zombie', 'people/zombie.md', { deleted: true });
|
||
const result = await runSyncCore(ctxOf(), { apply: true });
|
||
expect(result.total_applied).toBe(1);
|
||
expect(await getType('alice')).toBe('person');
|
||
expect(await getType('zombie')).toBe(''); // untouched (soft-deleted)
|
||
});
|
||
});
|
||
});
|
||
|
||
describe('runSyncCore — source scoping (codex C5 write-side)', () => {
|
||
it('updates only the scoped source', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md', { sourceId: 'src-a' });
|
||
await seedPage('bob', 'people/bob.md', { sourceId: 'src-b' });
|
||
const result = await runSyncCore(ctxOf(), { apply: true, sourceId: 'src-a' });
|
||
expect(result.total_applied).toBe(1);
|
||
expect(await getType('alice')).toBe('person');
|
||
expect(await getType('bob')).toBe('');
|
||
});
|
||
});
|
||
});
|
||
|
||
describe('runSyncCore — pack-load failure', () => {
|
||
it('returns empty result when no pack is loaded', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'nonexistent' }, async () => {
|
||
__setPackLocatorForTests(() => null);
|
||
await seedPage('alice', 'people/alice.md');
|
||
const result = await runSyncCore(ctxOf(), { apply: true });
|
||
expect(result.pack_identity).toBeNull();
|
||
expect(result.per_prefix).toEqual([]);
|
||
expect(result.total_applied).toBe(0);
|
||
// Page untouched (no pack = no inference rules).
|
||
expect(await getType('alice')).toBe('');
|
||
});
|
||
});
|
||
});
|
||
|
||
describe('runSyncCore — JSON envelope shape', () => {
|
||
it('schema_version stays 1', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
const result = await runSyncCore(ctxOf());
|
||
expect(result.schema_version).toBe(1);
|
||
});
|
||
});
|
||
|
||
it('per_prefix entry shape is stable', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md');
|
||
const result = await runSyncCore(ctxOf());
|
||
const entry = result.per_prefix[0]!;
|
||
expect(entry).toHaveProperty('type');
|
||
expect(entry).toHaveProperty('prefix');
|
||
expect(entry).toHaveProperty('would_apply');
|
||
expect(entry).toHaveProperty('sample_slugs');
|
||
expect(entry).toHaveProperty('dead_prefix');
|
||
expect(entry).toHaveProperty('applied');
|
||
});
|
||
});
|
||
});
|
||
|
||
describe('runSyncCore — batch size clamping', () => {
|
||
it('clamps batch size to >=1 and <=10000', async () => {
|
||
await withEnv({ GBRAIN_HOME: tmpDir, GBRAIN_SCHEMA_PACK: 'tiny' }, async () => {
|
||
seedTinyPack([{ name: 'person', prefix: 'people/' }]);
|
||
await seedPage('alice', 'people/alice.md');
|
||
// batchSize:0 and batchSize:99999 both work; result is identical.
|
||
const r1 = await runSyncCore(ctxOf(), { apply: true, batchSize: 0 });
|
||
expect(r1.total_applied).toBe(1);
|
||
});
|
||
});
|
||
});
|