Files
gbrain/test/import-file.test.ts
T
c9652443cf v0.32.7 feat: CJK fix wave — 6 layers from one root cause (closes vinsew + 313094319-sudo PRs) (#898)
* feat: shared CJK detection module (cjk.ts)

Foundation for the CJK fix wave. Single source of truth for CJK ranges
(Han, Hiragana, Katakana, Hangul Syllables), the slug-char string used
by adjacent validators, sentence + clause delimiter sets, the 30%
density threshold for word counting, and a LIKE-pattern escape helper.

Replaces the inline hasCJK regex at expansion.ts:58 so four-place
drift becomes impossible. countCJKAwareWords uses density threshold
(per codex outside-voice C13) so a long English doc with one Japanese
term stays whitespace-tokenized, not char-split.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* feat: migration v51 + pages.chunker_version/source_path columns

Schema-level support for the v0.32.7 CJK wave. Two new columns on pages:

  - chunker_version SMALLINT NOT NULL DEFAULT 1 — bumped to
    MARKDOWN_CHUNKER_VERSION (2) on every new import. The post-upgrade
    gbrain reindex --markdown sweep walks chunker_version < 2 to find
    pre-bump rows and rebuilds them.

  - source_path TEXT — captures the repo-relative path at import time
    so sync's delete/rename code can resolve frontmatter-fallback
    slugs (CJK / emoji / exotic-script files where the path itself
    doesn't derive a slug).

Both columns plumbed through PageInput, partial indexes scoped to
markdown-only / non-null. PGLite + Postgres parity via the standard
ALTER TABLE ... IF NOT EXISTS shape.

Replaces the original PR #599 plan of folding MARKDOWN_CHUNKER_VERSION
into content_hash. Codex outside-voice C2 caught that as a no-op:
performSync gates on actual file change, not hash-would-differ, so
the fold never reached existing pages. Column + sweep is the real fix.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* feat: CJK-aware slugify + SLUG_SEGMENT_PATTERN + adjacent validators

slugifySegment now preserves Han / Hiragana / Katakana / Hangul Syllables
with NFC re-normalization after the NFD-strip-accents pass so Hangul
Jamo recomposes back into precomposed syllables that fall inside the
whitelist. café still slugifies to cafe (regression preserved — iron
rule).

SLUG_SEGMENT_PATTERN (consumed by takes-holder validation) extended
with CJK_SLUG_CHARS in the same commit so CJK slugs aren't rejected by
adjacent validators downstream. Codex outside-voice C4 caught this
exact half-fix in the original plan — leaving the pattern ASCII-only
would have shipped a feature where the slugify produced 品牌圣经 but
adjacent validators flagged it.

src/core/operations.ts: validatePageSlug + validateFilename also
extended with CJK ranges. matchesSlugAllowList is unchanged (works on
string prefixes, no character class).

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* feat: recursive chunker — MARKDOWN_CHUNKER_VERSION + CJK splitting + maxChars cap

Four coordinated chunker changes for the v0.32.7 wave:

  - MARKDOWN_CHUNKER_VERSION = 2 exported. Folded into pages.chunker_version
    so the post-upgrade reindex sweep can find pre-bump pages.

  - countWords delegated to countCJKAwareWords from cjk.ts (30% density
    threshold). Below threshold: whitespace-token count (English-dominant
    docs stay tokenized). At/above: char count (Chinese paragraphs actually
    split instead of being treated as one 8192-token-overflowing word).

  - DELIMITERS extends L2 (sentences) with 。!? and L3 (clauses) with
    ;:,、. CJK punctuation now produces real chunk boundaries.

  - maxChars hard cap (default 6000) with sliding-window splitByChars and
    500-char overlap. Catches pathological whitespace-less inputs that the
    word-level pipeline can't bound (pure-Han paragraphs, base64 blobs,
    long URLs). Applied to both single-short-chunk and merged-chunks
    paths.

  - splitOnWhitespace falls through to char-slice when ANY single "word"
    exceeds target chars (the greedy /\S+/g regex returns a whole CJK
    paragraph as one "word"; without this, the L4 fallback produces one
    huge piece). Pre-fix this was the silent-failure path.

Tests in test/chunkers/recursive.test.ts: 9 new cases — pure Chinese,
Japanese + 。, Korean Hangul, mixed CJK+English, 20KB CJK with overlap,
single-short-chunk maxChars edge, pure-English regression.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* feat: PGLite CJK keyword fallback + engine chunker_version/source_path passthrough

PGLite uses websearch_to_tsquery('english') over to_tsvector('english'),
which can't tokenize CJK. Pre-fix, CJK queries returned empty results
on PGLite brains even with proper embeddings.

searchKeyword + searchKeywordChunks now branch on hasCJK(query):

  - ASCII path: unchanged. websearch_to_tsquery('english') continues
    to drive FTS. No regression risk.

  - CJK path: switches to ILIKE '%' || $qLike || '%' ESCAPE '\\' over
    chunk_text with two distinct param bindings ($qLike escaped for
    the ILIKE clause, $qRaw raw for the ranking arithmetic). Empty
    $qRaw guard bails before binding. Bigram-frequency-count ranking
    via (LENGTH(chunk_text) - LENGTH(REPLACE(chunk_text, $qRaw, ''))) /
    LENGTH($qRaw) approximates ts_rank semantics; position-in-chunk
    tiebreaker so earlier matches outrank later ones at the same
    occurrence count.

Codex outside-voice C8 caught the original plan's one-param shortcut
(escaped chars can't be reused as ranking substrings) + missing
ESCAPE clause + asymmetric whitespace strip. C9 corrected the FTS
dialect (websearch_to_tsquery, not to_tsvector('simple')).

Source-boost CASE, hard-exclude clause, visibility clause, and the
DISTINCT ON (slug) page-dedup all survive on both branches. Postgres
engine path stays untouched (multi-tenant Postgres deployments can
install pgroonga / zhparser for CJK; out of scope for this wave).

Postgres + PGLite putPage both extended to write chunker_version
and source_path columns (with COALESCE(EXCLUDED.x, pages.x) so
auto-link / code-reindex callers that don't supply them don't blank
existing values).

Tests: 8 new cases covering Chinese / Japanese / Korean substring
search, bigram ranking (3-hit > 1-hit), LIKE-meta-char escape
(literal % does not wildcard), English query stays on FTS path.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>
Co-Authored-By: 313094319-sudo <313094319-sudo@users.noreply.github.com>

* feat: import-file frontmatter-slug fallback + audit JSONL

importFromFile gains a fallback branch: when slugifyPath returns
empty (emoji / Thai / Arabic / exotic-script filename — including
post-CJK-wave files that still don't slugify) AND the frontmatter
declares a slug, the frontmatter slug becomes authoritative.

Anti-spoof rule preserved unchanged: when slugifyPath produces a
non-empty path slug AND the frontmatter slug claims a different one,
the file is still rejected. notes/random.md cannot impersonate
people/elon via frontmatter.

D6=B error string when both path slug AND frontmatter slug are empty:
"Filename produces no usable slug. Add a 'slug:' to the frontmatter,
or rename the file to use ASCII / Chinese / Japanese / Korean
characters." Honest about the actually-supported scripts.

Every import now populates pages.chunker_version (set to
MARKDOWN_CHUNKER_VERSION) and pages.source_path (repo-relative). These
drive the post-upgrade reindex sweep + sync's delete/rename slug
resolution.

NEW src/core/audit-slug-fallback.ts — weekly ISO-week-rotated JSONL
at ~/.gbrain/audit/slug-fallback-YYYY-Www.jsonl. Per codex C7, info
events don't belong in sync-failures.jsonl (which gates bookmark
advancement); separate audit surface keeps the failure-handling code
unchanged. logSlugFallback emits a stderr line AND appends to the
audit file (D7=D dual logging).

Tests: 5 new import-file cases (小米 with no frontmatter slug, 🚀.md
with frontmatter fallback, 🌟🚀.md friendly D6=B error, anti-spoof
regression, chunker_version + source_path populated). 6 new audit
cases covering write, weekly rotation, 7-day window, corrupt-row
tolerance.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* feat: git() helper hardening + core.quotepath=false for CJK paths

git CLI emits CJK paths as quoted octal escapes (\345\223\201 ...) by
default in diff --name-status output. Pre-fix, buildSyncManifest
silently dropped these paths because downstream filesystem lookups
saw the literal escape string. gbrain sync reported added=0 while
git had the file committed.

git() helper refactored:
  - New signature: git(repoPath, args: string[], configs?: string[])
  - Config flags emit BEFORE -C and BEFORE the subcommand (git CLI
    requires this order)
  - core.quotepath=false always prepended
  - Future callers needing extra -c config pass configs:[]; no more
    inlining -c into args (the silent-future-drift footgun codex C12
    flagged as a related concern)

New invariant test in test/sync.test.ts pins the emit order.

NEW test/e2e/sync-cjk-git.test.ts — real-git E2E in a tmpdir. Spawns
real git via execFileSync, commits a Chinese-named markdown file,
drives the helper through buildSyncManifest, asserts the manifest
contains the UTF-8 path (not the octal-escape form). Closes the
real-CLI-behavior gap that unit tests can't cover (the helper builds
the right args; only an E2E proves git actually emits UTF-8 under
the flag).

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* feat: gbrain reindex --markdown sweep command

NEW src/commands/reindex.ts — operator-facing markdown re-chunk
sweep. Walks SELECT slug, source_path FROM pages WHERE
page_kind = 'markdown' AND chunker_version < MARKDOWN_CHUNKER_VERSION
in 100-row batches, ordered by id ASC so partial-completion re-runs
pick up where they left off.

For rows with non-null source_path: re-imports via importFromFile
when the file exists on disk. For rows without (legacy pre-migration
backfill): fallback to importFromContent using the stored markdown
body.

Flags: --markdown (target selector), --limit N, --dry-run, --json,
--no-embed (offline / CI / test path that lets the chunker run
without a configured AI gateway), --repo PATH.

Wired into src/cli.ts dispatch table. Will also be invoked
automatically by gbrain upgrade's post-upgrade hook (next commit) so
chunker-version bumps reach existing markdown pages without an
explicit operator action.

Tests in test/reindex.test.ts: 5 cases covering dry-run, actual
sweep, idempotent re-run, --limit cap, skipped-already-at-current.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>
Co-Authored-By: 313094319-sudo <313094319-sudo@users.noreply.github.com>

* feat: post-upgrade chunker-bump cost prompt + auto-reindex sweep

Wires the chunker-version bump into gbrain upgrade so existing brains
heal automatically. Three new pieces:

NEW src/core/embedding-pricing.ts — EMBEDDING_PRICING map keyed
provider:model (OpenAI text-embedding-3-large + 3-small + ada-002,
Voyage 3-large + 3). lookupEmbeddingPrice returns 'known' or
'unknown' shape so the cost-estimate prompt can degrade gracefully
for unknown providers rather than fabricate numbers (codex C3).
estimateCostFromChars uses 3.5 chars/token approximation.

NEW src/core/post-upgrade-reembed.ts — pure-ish functions for the
cost-estimate prompt:
  - computeReembedEstimate: real SQL against
    COUNT(*) + COALESCE(SUM(LENGTH(compiled_truth)) + SUM(LENGTH(timeline))
    on the chunker_version-filtered query. No phantom markdown_body
    column (codex C3 caught the original plan referencing nonexistent
    schema fields).
  - formatReembedPrompt: pure string formatter for the stderr line.
  - runPostUpgradeReembedPrompt: orchestrates the prompt + 10-second
    Ctrl-C window. TTY-only wait so non-TTY upgrades (CI, cron-driven,
    headless) don't hang. GBRAIN_NO_REEMBED=1 bails out entirely
    with a doctor-warning marker; GBRAIN_REEMBED_GRACE_SECONDS=0
    skips the wait.

src/commands/upgrade.ts: after apply-migrations runs, the new prompt
fires through the gateway's configured embedding model, then invokes
gbrain reindex --markdown automatically if the user proceeds.
Wrapped in try-catch so a reindex failure is non-fatal — the user
can re-run manually.

Tests in test/upgrade-reembed-prompt.test.ts: 11 cases covering real
SQL counts, unknown-provider fallback, TTY / non-TTY paths,
GBRAIN_NO_REEMBED bail-out, GBRAIN_REEMBED_GRACE_SECONDS=0 skip-wait.

Codex outside-voice C2 caught the original plan as a no-op
(performSync doesn't re-import unchanged files just because
content_hash would differ). The migration v51 column + this sweep
+ this prompt is the real fix that actually reaches existing pages.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* feat: doctor slug_fallback_audit check + CJK roundtrip E2E

gbrain doctor learns a new slug_fallback_audit check (v0.32.7).
Reads the latest week of ~/.gbrain/audit/slug-fallback-*.jsonl,
counts info-severity entries from the last 7 days, surfaces the
total as an ok-status line. No health-score docking; no warning.

sync-failures.jsonl (which gates bookmark advancement) stays
untouched — info events live in their own surface per codex C7.

NEW test/e2e/cjk-roundtrip.test.ts — proves the wave delivers end-
to-end. PGLite-in-memory fixture with Chinese / Japanese / Korean
content. Each page: importFromContent → chunkText (CJK-aware) →
searchKeyword (LIKE-branch with bigram count). Asserts every CJK
query lands on its source page. ASCII regression: an English query
still uses the FTS path on the same brain. Vector path skips
gracefully without OPENAI_API_KEY.

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>

* chore: bump version and changelog (v0.32.7)

CJK fix wave — six layers from one root cause. Three originating PRs
from @vinsew and one extracted from @313094319-sudo's #765 land
together as a coherent collector. Codex outside-voice review on the
plan caught four critical bugs the eng review missed (no-op
re-embed, SLUG_SEGMENT_PATTERN half-fix, LIKE SQL needing two
distinct param bindings, countCJKAwareWords over-splitting on
English+1-CJK-term docs). All four addressed in the implementation.

TODOS.md: resolved the v0.32.x PGLite CJK keyword fallback entry;
filed five v0.33+ follow-ups (Postgres CJK FTS via pgroonga / wider
Unicode property escapes / -z NUL git framing / CJK overlap context /
other non-Latin scripts / embedding pricing refresh mechanism).

Co-Authored-By: vinsew <vinsew@users.noreply.github.com>
Co-Authored-By: 313094319-sudo <313094319-sudo@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix: review findings — forceRechunk + source_path lookup (codex post-merge)

Two critical issues caught by codex adversarial on the post-merge tree:

F1 — Reindex sweep was a no-op on unchanged-source pages. importFromContent
short-circuits on existing.content_hash === hash BEFORE the chunker runs,
so the v0.32.7 MARKDOWN_CHUNKER_VERSION bump (and master's v0.32.2
stripFactsFence privacy strip) never reached pages whose markdown body
hadn't been edited.

Fix: new `forceRechunk?: boolean` option on importFromContent + importFromFile.
When set, the hash short-circuit is bypassed and the page re-runs the full
chunk + write pipeline. `gbrain reindex --markdown` now passes forceRechunk:
true on every row. This means:
  - The CJK chunker bump actually reaches existing markdown pages.
  - Master's v0.32.2 stripFactsFence applies retroactively too — any
    pre-strip private fact bytes lingering in content_chunks get cleared
    when the v0.32.7 post-upgrade sweep runs.

New test in test/reindex.test.ts seeds a page, runs the sweep, mocks a
stale chunker_version=1 without changing compiled_truth, runs the sweep
again, asserts chunker_version is bumped despite hash match.

F4 — Sync delete/rename still used resolveSlugForPath(path) only, ignoring
the new pages.source_path column added in v52. Frontmatter-fallback pages
(emoji-only / Thai / Arabic filenames where slugifyPath returns empty and
the slug came from the markdown frontmatter) would orphan on delete or
rename because the path-derived slug doesn't match the stored slug.

Fix: new exported helper resolveSlugByPathOrSourcePath(engine, path,
sourceId?) queries pages.source_path first, falls back to
resolveSlugForPath when no row matches. Threaded into 3 call sites in
sync.ts (un-syncable modified cleanup at :531, deletes at :603, rename
oldSlug at :622). Best-effort: query errors fall through to the legacy
path so pre-migration brains still work.

3 new test cases in test/sync.test.ts cover: stored-slug lookup hits,
fallback when no source_path row exists, and source_id scoping when two
sources have the same source_path value.

Codex finding #3 (reindex not in CLI_ONLY) was verified as a false
positive — CLI_ONLY is the set that doesn't need an engine; reindex
correctly belongs to the engine-backed dispatch.

302 wave tests pass / 0 fail. bun run verify green.

* docs: update CLAUDE.md + llms-full.txt for v0.32.7 CJK fix wave

CLAUDE.md Key Files: added entries for the five new modules introduced by
the wave — src/core/cjk.ts (shared detection + delimiters + density
threshold), src/core/audit-slug-fallback.ts (weekly JSONL),
src/core/embedding-pricing.ts (post-upgrade cost lookup table),
src/core/post-upgrade-reembed.ts (prompt + grace window), and
src/commands/reindex.ts (chunker_version sweep with forceRechunk).

Also noted src/commands/sync.ts:resolveSlugByPathOrSourcePath — the
F4 codex post-merge fix that wires the new pages.source_path column
into sync delete/rename so frontmatter-fallback pages don't orphan.

CLAUDE.md Commands: added a v0.32.7 section covering `gbrain reindex
--markdown`, the new doctor slug_fallback_audit check, PGLite CJK
keyword fallback in `gbrain search`, and the post-upgrade
chunker-bump cost prompt with its env-var overrides.

llms-full.txt: regenerated via bun run build:llms (CI gate runs the
generator on every release; commit must include the bundle).

README.md: no changes needed — v0.32.7 is internal correctness
across the existing pipeline, not a new skill or setup story.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: vinsew <vinsew@users.noreply.github.com>
Co-authored-by: 313094319-sudo <313094319-sudo@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-11 22:54:35 -07:00

508 lines
17 KiB
TypeScript

import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
import { writeFileSync, mkdirSync, rmSync, symlinkSync, readdirSync, readFileSync, existsSync } from 'fs';
import { join } from 'path';
import { importFile, importFromContent } from '../src/core/import-file.ts';
import type { BrainEngine } from '../src/core/engine.ts';
import { MARKDOWN_CHUNKER_VERSION } from '../src/core/chunkers/recursive.ts';
const TMP = join(import.meta.dir, '.tmp-import-test');
// Minimal mock engine that tracks calls and supports transaction()
function mockEngine(overrides: Partial<Record<string, any>> = {}): BrainEngine {
const calls: { method: string; args: any[] }[] = [];
const track = (method: string) => (...args: any[]) => {
calls.push({ method, args });
if (overrides[method]) return overrides[method](...args);
return Promise.resolve(null);
};
const engine = new Proxy({} as any, {
get(_, prop: string) {
if (prop === '_calls') return calls;
if (prop === 'getTags') return overrides.getTags || (() => Promise.resolve([]));
if (prop === 'getPage') return overrides.getPage || (() => Promise.resolve(null));
// transaction: just call the fn with the same engine (no real DB transaction in tests)
if (prop === 'transaction') return async (fn: (tx: BrainEngine) => Promise<any>) => fn(engine);
return track(prop);
},
});
return engine;
}
beforeAll(() => {
mkdirSync(TMP, { recursive: true });
});
afterAll(() => {
rmSync(TMP, { recursive: true, force: true });
});
describe('importFile', () => {
test('imports a valid markdown file', async () => {
const filePath = join(TMP, 'test-page.md');
writeFileSync(filePath, `---
type: concept
title: Test Page
tags: [alpha, beta]
---
This is the compiled truth.
---
- 2024-01-01: Something happened.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'concepts/test-page.md', { noEmbed: true });
expect(result.status).toBe('imported');
expect(result.slug).toBe('concepts/test-page');
expect(result.chunks).toBeGreaterThan(0);
// Verify engine was called correctly
const calls = (engine as any)._calls;
const putCall = calls.find((c: any) => c.method === 'putPage');
expect(putCall).toBeTruthy();
expect(putCall.args[0]).toBe('concepts/test-page');
// Tags were added
const tagCalls = calls.filter((c: any) => c.method === 'addTag');
expect(tagCalls.length).toBe(2);
// Chunks were upserted
const chunkCall = calls.find((c: any) => c.method === 'upsertChunks');
expect(chunkCall).toBeTruthy();
});
test('skips files larger than MAX_FILE_SIZE (5MB)', async () => {
const filePath = join(TMP, 'big-file.md');
const bigContent = '---\ntitle: Big\n---\n' + 'x'.repeat(5_100_000);
writeFileSync(filePath, bigContent);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'big-file.md', { noEmbed: true });
expect(result.status).toBe('skipped');
expect(result.error).toContain('too large');
expect((engine as any)._calls.length).toBe(0);
});
test('rejects frontmatter slug that does not match the file path', async () => {
// In a shared brain where contributors can land PRs, this prevents a
// poisoned notes/random.md from declaring `slug: people/elon` in its
// frontmatter and overwriting the legitimate people/elon page on sync.
const filePath = join(TMP, 'hijack.md');
writeFileSync(filePath, `---
type: person
title: Elon Musk
slug: people/elon
---
Poisoned content that would overwrite people/elon.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'notes/random.md', { noEmbed: true });
expect(result.status).toBe('skipped');
expect(result.error).toContain('people/elon');
expect(result.error).toContain('notes/random');
// No writes to the DB — the hijack never reaches putPage/createVersion.
expect((engine as any)._calls.length).toBe(0);
});
test('accepts frontmatter slug that matches the file path', async () => {
// Sanity: a legitimate file whose frontmatter slug happens to equal the
// path-derived slug must still import.
const filePath = join(TMP, 'alice.md');
writeFileSync(filePath, `---
type: person
title: Alice
slug: people/alice-smith
---
Legit content.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'people/alice-smith.md', { noEmbed: true });
expect(result.status).toBe('imported');
expect(result.slug).toBe('people/alice-smith');
});
test('uses path-derived slug when no frontmatter slug is set', async () => {
// The common case: no frontmatter.slug, so the path determines the slug.
const filePath = join(TMP, 'concept-path.md');
writeFileSync(filePath, `---
type: concept
title: From Path
---
Content.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'concepts/from-path.md', { noEmbed: true });
expect(result.status).toBe('imported');
expect(result.slug).toBe('concepts/from-path');
});
test('skips symlinks in importFromFile (defense-in-depth)', async () => {
// Even if the walker somehow passes a symlink through, importFromFile
// should catch it and return skipped.
const realFile = join(TMP, 'real-target.md');
writeFileSync(realFile, `---
type: concept
title: Real
---
Content.
`);
const linkPath = join(TMP, 'symlink-file.md');
try { rmSync(linkPath); } catch { /* may not exist */ }
symlinkSync(realFile, linkPath);
const engine = mockEngine();
const result = await importFile(engine, linkPath, 'symlink-file.md', { noEmbed: true });
expect(result.status).toBe('skipped');
expect(result.error).toContain('symlink');
expect((engine as any)._calls.length).toBe(0);
});
test('skips file when content hash matches (idempotent)', async () => {
const filePath = join(TMP, 'unchanged.md');
writeFileSync(filePath, `---
type: concept
title: Unchanged
---
Same content.
`);
// Hash now includes ALL fields (title, type, frontmatter, tags)
const { createHash } = await import('crypto');
const { parseMarkdown } = await import('../src/core/markdown.ts');
const content = `---
type: concept
title: Unchanged
---
Same content.
`;
const parsed = parseMarkdown(content, 'concepts/unchanged.md');
const hash = createHash('sha256')
.update(JSON.stringify({
title: parsed.title,
type: parsed.type,
compiled_truth: parsed.compiled_truth,
timeline: parsed.timeline,
frontmatter: parsed.frontmatter,
tags: parsed.tags.sort(),
}))
.digest('hex');
const engine = mockEngine({
getPage: () => Promise.resolve({ content_hash: hash }),
});
const result = await importFile(engine, filePath, 'concepts/unchanged.md', { noEmbed: true });
expect(result.status).toBe('skipped');
const calls = (engine as any)._calls;
const putCall = calls.find((c: any) => c.method === 'putPage');
expect(putCall).toBeUndefined();
});
test('reconciles tags: removes old, adds new', async () => {
const filePath = join(TMP, 'retag.md');
writeFileSync(filePath, `---
type: concept
title: Retagged
tags: [new-tag, kept-tag]
---
Content here.
`);
const engine = mockEngine({
getTags: () => Promise.resolve(['old-tag', 'kept-tag']),
getPage: () => Promise.resolve(null),
});
await importFile(engine, filePath, 'concepts/retag.md', { noEmbed: true });
const calls = (engine as any)._calls;
const removeCalls = calls.filter((c: any) => c.method === 'removeTag');
const addCalls = calls.filter((c: any) => c.method === 'addTag');
expect(removeCalls.length).toBe(1);
expect(removeCalls[0].args[1]).toBe('old-tag');
expect(addCalls.length).toBe(2);
});
test('chunks compiled_truth and timeline separately', async () => {
const filePath = join(TMP, 'chunked.md');
writeFileSync(filePath, `---
type: concept
title: Chunked
---
This is compiled truth content that should be chunked as compiled_truth source.
<!-- timeline -->
- 2024-01-01: This is timeline content that should be chunked as timeline source.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'concepts/chunked.md', { noEmbed: true });
expect(result.status).toBe('imported');
expect(result.chunks).toBeGreaterThanOrEqual(2);
const calls = (engine as any)._calls;
const chunkCall = calls.find((c: any) => c.method === 'upsertChunks');
const chunks = chunkCall.args[1];
const ctChunks = chunks.filter((c: any) => c.chunk_source === 'compiled_truth');
const tlChunks = chunks.filter((c: any) => c.chunk_source === 'timeline');
expect(ctChunks.length).toBeGreaterThan(0);
expect(tlChunks.length).toBeGreaterThan(0);
});
test('handles file with minimal content', async () => {
const filePath = join(TMP, 'minimal.md');
writeFileSync(filePath, `---
type: concept
title: Minimal
---
One line.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'concepts/minimal.md', { noEmbed: true });
expect(result.status).toBe('imported');
expect(result.chunks).toBeGreaterThanOrEqual(1);
});
test('skips chunking for empty timeline', async () => {
const filePath = join(TMP, 'empty-tl.md');
writeFileSync(filePath, `---
type: concept
title: No Timeline
---
Just compiled truth, no timeline separator.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'concepts/empty-tl.md', { noEmbed: true });
expect(result.status).toBe('imported');
const calls = (engine as any)._calls;
const chunkCall = calls.find((c: any) => c.method === 'upsertChunks');
if (chunkCall) {
const chunks = chunkCall.args[1];
const tlChunks = chunks.filter((c: any) => c.chunk_source === 'timeline');
expect(tlChunks.length).toBe(0);
}
});
test('noEmbed: true skips embedding', async () => {
const filePath = join(TMP, 'no-embed.md');
writeFileSync(filePath, `---
type: concept
title: No Embed
---
Content to chunk but not embed.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'concepts/no-embed.md', { noEmbed: true });
expect(result.status).toBe('imported');
const calls = (engine as any)._calls;
const chunkCall = calls.find((c: any) => c.method === 'upsertChunks');
if (chunkCall) {
for (const chunk of chunkCall.args[1]) {
expect(chunk.embedding).toBeUndefined();
}
}
});
test('rejects in-memory content larger than MAX_FILE_SIZE', async () => {
// The remote MCP put_page operation hands user-supplied content straight
// to importFromContent, which is the path this guard defends. The guard
// must trigger BEFORE parseMarkdown / chunkText / embedBatch — if it doesn't,
// an authenticated attacker can force the owner to pay for embedding a
// multi-megabyte string.
const bigContent = '---\ntitle: Big\n---\n' + 'x'.repeat(5_100_000);
const engine = mockEngine();
const result = await importFromContent(engine, 'big-slug', bigContent, { noEmbed: true });
expect(result.status).toBe('skipped');
expect(result.error).toContain('too large');
// No engine work at all — confirms the guard short-circuits before any
// parsing or chunking allocation.
expect((engine as any)._calls.length).toBe(0);
});
test('uses UTF-8 byte length, not JS string length, for the size check', async () => {
// 2.6M 4-byte codepoints = ~10.4 MB UTF-8 but only 2.6M JS UTF-16 code units.
// A length-based check would let this through; a byteLength check catches it.
const fourByteChar = '\u{1F600}'; // emoji, 4 bytes in UTF-8
const bigContent = fourByteChar.repeat(2_600_000);
const engine = mockEngine();
const result = await importFromContent(engine, 'emoji-slug', bigContent, { noEmbed: true });
expect(result.status).toBe('skipped');
expect(result.error).toContain('too large');
expect((engine as any)._calls.length).toBe(0);
});
test('accepts in-memory content just under MAX_FILE_SIZE', async () => {
// Sanity: content exactly at the limit must still import. If this test
// fails, the guard is off-by-one and will break legitimate large imports.
const content = '---\ntitle: Borderline\n---\n' + 'x'.repeat(4_900_000);
const engine = mockEngine();
const result = await importFromContent(engine, 'borderline-slug', content, { noEmbed: true });
expect(result.status).toBe('imported');
});
test('assigns sequential chunk_index values', async () => {
const filePath = join(TMP, 'indexed.md');
const longText = Array(50).fill('This is a sentence that adds length to the content.').join(' ');
writeFileSync(filePath, `---
type: concept
title: Indexed
---
${longText}
---
${longText}
`);
const engine = mockEngine();
await importFile(engine, filePath, 'concepts/indexed.md', { noEmbed: true });
const calls = (engine as any)._calls;
const chunkCall = calls.find((c: any) => c.method === 'upsertChunks');
if (chunkCall) {
const chunks = chunkCall.args[1];
for (let i = 0; i < chunks.length; i++) {
expect(chunks[i].chunk_index).toBe(i);
}
}
});
});
describe('importFile — CJK wave (v0.32.7)', () => {
test('REGRESSION: pure-CJK filename with NO frontmatter slug imports cleanly as CJK slug', async () => {
// After #115, slugifyPath('小米.md') = '小米' (CJK preserved). The
// anti-spoof rule is content with no frontmatter slug present.
const filePath = join(TMP, '小米.md');
writeFileSync(filePath, `---
type: company
title: Xiaomi
---
Body text.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, '小米.md', { noEmbed: true });
expect(result.status).toBe('imported');
expect(result.slug).toBe('小米');
const putCall = (engine as any)._calls.find((c: any) => c.method === 'putPage');
expect(putCall.args[1].chunker_version).toBe(MARKDOWN_CHUNKER_VERSION);
expect(putCall.args[1].source_path).toBe('小米.md');
});
test('empty-path-slug + frontmatter slug → fallback path fires (emoji filename)', async () => {
// 🚀.md slugifies empty even after #115 (emoji not in CJK ranges).
// Frontmatter slug must take over. logSlugFallback fires.
const filePath = join(TMP, '🚀.md');
writeFileSync(filePath, `---
type: project
title: Launch
slug: projects/launch
---
Lifting off.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, '🚀.md', { noEmbed: true });
expect(result.status).toBe('imported');
expect(result.slug).toBe('projects/launch');
const putCall = (engine as any)._calls.find((c: any) => c.method === 'putPage');
expect(putCall.args[0]).toBe('projects/launch');
expect(putCall.args[1].source_path).toBe('🚀.md');
});
test('empty-path-slug + NO frontmatter slug → friendly D6=B error message', async () => {
// 🌟🚀 slugifies to '' (both emoji stripped, no remaining chars).
// No frontmatter slug to fall back on → friendly error.
const filePath = join(TMP, '🌟🚀.md');
writeFileSync(filePath, `# Bare body without frontmatter slug
just content.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, '🌟🚀.md', { noEmbed: true });
expect(result.status).toBe('skipped');
expect(result.error).toContain('no usable slug');
expect(result.error).toContain('ASCII / Chinese / Japanese / Korean');
expect((engine as any)._calls.length).toBe(0);
});
test('REGRESSION: anti-spoof still rejects when path DOES derive a slug', async () => {
// notes/random.md derives slug `notes/random`. Frontmatter `slug: people/elon`
// is a mismatch and MUST still be rejected (the original PR #598 + C1 test
// fixture contradiction concern).
const filePath = join(TMP, 'antispoof-cjk-wave.md');
writeFileSync(filePath, `---
type: person
title: Elon
slug: people/elon
---
Hijack.
`);
const engine = mockEngine();
const result = await importFile(engine, filePath, 'notes/antispoof-cjk-wave.md', { noEmbed: true });
expect(result.status).toBe('skipped');
expect(result.error).toContain('does not match');
expect((engine as any)._calls.length).toBe(0);
});
test('chunker_version + source_path populated on every import', async () => {
const filePath = join(TMP, 'cjk-source-path.md');
writeFileSync(filePath, `---
type: concept
title: Has source path
---
Content.
`);
const engine = mockEngine();
await importFile(engine, filePath, 'concepts/cjk-source-path.md', { noEmbed: true });
const putCall = (engine as any)._calls.find((c: any) => c.method === 'putPage');
expect(putCall).toBeTruthy();
expect(putCall.args[1].chunker_version).toBe(MARKDOWN_CHUNKER_VERSION);
expect(putCall.args[1].source_path).toBe('concepts/cjk-source-path.md');
});
});