mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 21:19:18 +00:00
* feat: shared CJK detection module (cjk.ts) Foundation for the CJK fix wave. Single source of truth for CJK ranges (Han, Hiragana, Katakana, Hangul Syllables), the slug-char string used by adjacent validators, sentence + clause delimiter sets, the 30% density threshold for word counting, and a LIKE-pattern escape helper. Replaces the inline hasCJK regex at expansion.ts:58 so four-place drift becomes impossible. countCJKAwareWords uses density threshold (per codex outside-voice C13) so a long English doc with one Japanese term stays whitespace-tokenized, not char-split. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * feat: migration v51 + pages.chunker_version/source_path columns Schema-level support for the v0.32.7 CJK wave. Two new columns on pages: - chunker_version SMALLINT NOT NULL DEFAULT 1 — bumped to MARKDOWN_CHUNKER_VERSION (2) on every new import. The post-upgrade gbrain reindex --markdown sweep walks chunker_version < 2 to find pre-bump rows and rebuilds them. - source_path TEXT — captures the repo-relative path at import time so sync's delete/rename code can resolve frontmatter-fallback slugs (CJK / emoji / exotic-script files where the path itself doesn't derive a slug). Both columns plumbed through PageInput, partial indexes scoped to markdown-only / non-null. PGLite + Postgres parity via the standard ALTER TABLE ... IF NOT EXISTS shape. Replaces the original PR #599 plan of folding MARKDOWN_CHUNKER_VERSION into content_hash. Codex outside-voice C2 caught that as a no-op: performSync gates on actual file change, not hash-would-differ, so the fold never reached existing pages. Column + sweep is the real fix. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * feat: CJK-aware slugify + SLUG_SEGMENT_PATTERN + adjacent validators slugifySegment now preserves Han / Hiragana / Katakana / Hangul Syllables with NFC re-normalization after the NFD-strip-accents pass so Hangul Jamo recomposes back into precomposed syllables that fall inside the whitelist. café still slugifies to cafe (regression preserved — iron rule). SLUG_SEGMENT_PATTERN (consumed by takes-holder validation) extended with CJK_SLUG_CHARS in the same commit so CJK slugs aren't rejected by adjacent validators downstream. Codex outside-voice C4 caught this exact half-fix in the original plan — leaving the pattern ASCII-only would have shipped a feature where the slugify produced 品牌圣经 but adjacent validators flagged it. src/core/operations.ts: validatePageSlug + validateFilename also extended with CJK ranges. matchesSlugAllowList is unchanged (works on string prefixes, no character class). Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * feat: recursive chunker — MARKDOWN_CHUNKER_VERSION + CJK splitting + maxChars cap Four coordinated chunker changes for the v0.32.7 wave: - MARKDOWN_CHUNKER_VERSION = 2 exported. Folded into pages.chunker_version so the post-upgrade reindex sweep can find pre-bump pages. - countWords delegated to countCJKAwareWords from cjk.ts (30% density threshold). Below threshold: whitespace-token count (English-dominant docs stay tokenized). At/above: char count (Chinese paragraphs actually split instead of being treated as one 8192-token-overflowing word). - DELIMITERS extends L2 (sentences) with 。!? and L3 (clauses) with ;:,、. CJK punctuation now produces real chunk boundaries. - maxChars hard cap (default 6000) with sliding-window splitByChars and 500-char overlap. Catches pathological whitespace-less inputs that the word-level pipeline can't bound (pure-Han paragraphs, base64 blobs, long URLs). Applied to both single-short-chunk and merged-chunks paths. - splitOnWhitespace falls through to char-slice when ANY single "word" exceeds target chars (the greedy /\S+/g regex returns a whole CJK paragraph as one "word"; without this, the L4 fallback produces one huge piece). Pre-fix this was the silent-failure path. Tests in test/chunkers/recursive.test.ts: 9 new cases — pure Chinese, Japanese + 。, Korean Hangul, mixed CJK+English, 20KB CJK with overlap, single-short-chunk maxChars edge, pure-English regression. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * feat: PGLite CJK keyword fallback + engine chunker_version/source_path passthrough PGLite uses websearch_to_tsquery('english') over to_tsvector('english'), which can't tokenize CJK. Pre-fix, CJK queries returned empty results on PGLite brains even with proper embeddings. searchKeyword + searchKeywordChunks now branch on hasCJK(query): - ASCII path: unchanged. websearch_to_tsquery('english') continues to drive FTS. No regression risk. - CJK path: switches to ILIKE '%' || $qLike || '%' ESCAPE '\\' over chunk_text with two distinct param bindings ($qLike escaped for the ILIKE clause, $qRaw raw for the ranking arithmetic). Empty $qRaw guard bails before binding. Bigram-frequency-count ranking via (LENGTH(chunk_text) - LENGTH(REPLACE(chunk_text, $qRaw, ''))) / LENGTH($qRaw) approximates ts_rank semantics; position-in-chunk tiebreaker so earlier matches outrank later ones at the same occurrence count. Codex outside-voice C8 caught the original plan's one-param shortcut (escaped chars can't be reused as ranking substrings) + missing ESCAPE clause + asymmetric whitespace strip. C9 corrected the FTS dialect (websearch_to_tsquery, not to_tsvector('simple')). Source-boost CASE, hard-exclude clause, visibility clause, and the DISTINCT ON (slug) page-dedup all survive on both branches. Postgres engine path stays untouched (multi-tenant Postgres deployments can install pgroonga / zhparser for CJK; out of scope for this wave). Postgres + PGLite putPage both extended to write chunker_version and source_path columns (with COALESCE(EXCLUDED.x, pages.x) so auto-link / code-reindex callers that don't supply them don't blank existing values). Tests: 8 new cases covering Chinese / Japanese / Korean substring search, bigram ranking (3-hit > 1-hit), LIKE-meta-char escape (literal % does not wildcard), English query stays on FTS path. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> Co-Authored-By: 313094319-sudo <313094319-sudo@users.noreply.github.com> * feat: import-file frontmatter-slug fallback + audit JSONL importFromFile gains a fallback branch: when slugifyPath returns empty (emoji / Thai / Arabic / exotic-script filename — including post-CJK-wave files that still don't slugify) AND the frontmatter declares a slug, the frontmatter slug becomes authoritative. Anti-spoof rule preserved unchanged: when slugifyPath produces a non-empty path slug AND the frontmatter slug claims a different one, the file is still rejected. notes/random.md cannot impersonate people/elon via frontmatter. D6=B error string when both path slug AND frontmatter slug are empty: "Filename produces no usable slug. Add a 'slug:' to the frontmatter, or rename the file to use ASCII / Chinese / Japanese / Korean characters." Honest about the actually-supported scripts. Every import now populates pages.chunker_version (set to MARKDOWN_CHUNKER_VERSION) and pages.source_path (repo-relative). These drive the post-upgrade reindex sweep + sync's delete/rename slug resolution. NEW src/core/audit-slug-fallback.ts — weekly ISO-week-rotated JSONL at ~/.gbrain/audit/slug-fallback-YYYY-Www.jsonl. Per codex C7, info events don't belong in sync-failures.jsonl (which gates bookmark advancement); separate audit surface keeps the failure-handling code unchanged. logSlugFallback emits a stderr line AND appends to the audit file (D7=D dual logging). Tests: 5 new import-file cases (小米 with no frontmatter slug, 🚀.md with frontmatter fallback, 🌟🚀.md friendly D6=B error, anti-spoof regression, chunker_version + source_path populated). 6 new audit cases covering write, weekly rotation, 7-day window, corrupt-row tolerance. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * feat: git() helper hardening + core.quotepath=false for CJK paths git CLI emits CJK paths as quoted octal escapes (\345\223\201 ...) by default in diff --name-status output. Pre-fix, buildSyncManifest silently dropped these paths because downstream filesystem lookups saw the literal escape string. gbrain sync reported added=0 while git had the file committed. git() helper refactored: - New signature: git(repoPath, args: string[], configs?: string[]) - Config flags emit BEFORE -C and BEFORE the subcommand (git CLI requires this order) - core.quotepath=false always prepended - Future callers needing extra -c config pass configs:[]; no more inlining -c into args (the silent-future-drift footgun codex C12 flagged as a related concern) New invariant test in test/sync.test.ts pins the emit order. NEW test/e2e/sync-cjk-git.test.ts — real-git E2E in a tmpdir. Spawns real git via execFileSync, commits a Chinese-named markdown file, drives the helper through buildSyncManifest, asserts the manifest contains the UTF-8 path (not the octal-escape form). Closes the real-CLI-behavior gap that unit tests can't cover (the helper builds the right args; only an E2E proves git actually emits UTF-8 under the flag). Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * feat: gbrain reindex --markdown sweep command NEW src/commands/reindex.ts — operator-facing markdown re-chunk sweep. Walks SELECT slug, source_path FROM pages WHERE page_kind = 'markdown' AND chunker_version < MARKDOWN_CHUNKER_VERSION in 100-row batches, ordered by id ASC so partial-completion re-runs pick up where they left off. For rows with non-null source_path: re-imports via importFromFile when the file exists on disk. For rows without (legacy pre-migration backfill): fallback to importFromContent using the stored markdown body. Flags: --markdown (target selector), --limit N, --dry-run, --json, --no-embed (offline / CI / test path that lets the chunker run without a configured AI gateway), --repo PATH. Wired into src/cli.ts dispatch table. Will also be invoked automatically by gbrain upgrade's post-upgrade hook (next commit) so chunker-version bumps reach existing markdown pages without an explicit operator action. Tests in test/reindex.test.ts: 5 cases covering dry-run, actual sweep, idempotent re-run, --limit cap, skipped-already-at-current. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> Co-Authored-By: 313094319-sudo <313094319-sudo@users.noreply.github.com> * feat: post-upgrade chunker-bump cost prompt + auto-reindex sweep Wires the chunker-version bump into gbrain upgrade so existing brains heal automatically. Three new pieces: NEW src/core/embedding-pricing.ts — EMBEDDING_PRICING map keyed provider:model (OpenAI text-embedding-3-large + 3-small + ada-002, Voyage 3-large + 3). lookupEmbeddingPrice returns 'known' or 'unknown' shape so the cost-estimate prompt can degrade gracefully for unknown providers rather than fabricate numbers (codex C3). estimateCostFromChars uses 3.5 chars/token approximation. NEW src/core/post-upgrade-reembed.ts — pure-ish functions for the cost-estimate prompt: - computeReembedEstimate: real SQL against COUNT(*) + COALESCE(SUM(LENGTH(compiled_truth)) + SUM(LENGTH(timeline)) on the chunker_version-filtered query. No phantom markdown_body column (codex C3 caught the original plan referencing nonexistent schema fields). - formatReembedPrompt: pure string formatter for the stderr line. - runPostUpgradeReembedPrompt: orchestrates the prompt + 10-second Ctrl-C window. TTY-only wait so non-TTY upgrades (CI, cron-driven, headless) don't hang. GBRAIN_NO_REEMBED=1 bails out entirely with a doctor-warning marker; GBRAIN_REEMBED_GRACE_SECONDS=0 skips the wait. src/commands/upgrade.ts: after apply-migrations runs, the new prompt fires through the gateway's configured embedding model, then invokes gbrain reindex --markdown automatically if the user proceeds. Wrapped in try-catch so a reindex failure is non-fatal — the user can re-run manually. Tests in test/upgrade-reembed-prompt.test.ts: 11 cases covering real SQL counts, unknown-provider fallback, TTY / non-TTY paths, GBRAIN_NO_REEMBED bail-out, GBRAIN_REEMBED_GRACE_SECONDS=0 skip-wait. Codex outside-voice C2 caught the original plan as a no-op (performSync doesn't re-import unchanged files just because content_hash would differ). The migration v51 column + this sweep + this prompt is the real fix that actually reaches existing pages. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * feat: doctor slug_fallback_audit check + CJK roundtrip E2E gbrain doctor learns a new slug_fallback_audit check (v0.32.7). Reads the latest week of ~/.gbrain/audit/slug-fallback-*.jsonl, counts info-severity entries from the last 7 days, surfaces the total as an ok-status line. No health-score docking; no warning. sync-failures.jsonl (which gates bookmark advancement) stays untouched — info events live in their own surface per codex C7. NEW test/e2e/cjk-roundtrip.test.ts — proves the wave delivers end- to-end. PGLite-in-memory fixture with Chinese / Japanese / Korean content. Each page: importFromContent → chunkText (CJK-aware) → searchKeyword (LIKE-branch with bigram count). Asserts every CJK query lands on its source page. ASCII regression: an English query still uses the FTS path on the same brain. Vector path skips gracefully without OPENAI_API_KEY. Co-Authored-By: vinsew <vinsew@users.noreply.github.com> * chore: bump version and changelog (v0.32.7) CJK fix wave — six layers from one root cause. Three originating PRs from @vinsew and one extracted from @313094319-sudo's #765 land together as a coherent collector. Codex outside-voice review on the plan caught four critical bugs the eng review missed (no-op re-embed, SLUG_SEGMENT_PATTERN half-fix, LIKE SQL needing two distinct param bindings, countCJKAwareWords over-splitting on English+1-CJK-term docs). All four addressed in the implementation. TODOS.md: resolved the v0.32.x PGLite CJK keyword fallback entry; filed five v0.33+ follow-ups (Postgres CJK FTS via pgroonga / wider Unicode property escapes / -z NUL git framing / CJK overlap context / other non-Latin scripts / embedding pricing refresh mechanism). Co-Authored-By: vinsew <vinsew@users.noreply.github.com> Co-Authored-By: 313094319-sudo <313094319-sudo@users.noreply.github.com> Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix: review findings — forceRechunk + source_path lookup (codex post-merge) Two critical issues caught by codex adversarial on the post-merge tree: F1 — Reindex sweep was a no-op on unchanged-source pages. importFromContent short-circuits on existing.content_hash === hash BEFORE the chunker runs, so the v0.32.7 MARKDOWN_CHUNKER_VERSION bump (and master's v0.32.2 stripFactsFence privacy strip) never reached pages whose markdown body hadn't been edited. Fix: new `forceRechunk?: boolean` option on importFromContent + importFromFile. When set, the hash short-circuit is bypassed and the page re-runs the full chunk + write pipeline. `gbrain reindex --markdown` now passes forceRechunk: true on every row. This means: - The CJK chunker bump actually reaches existing markdown pages. - Master's v0.32.2 stripFactsFence applies retroactively too — any pre-strip private fact bytes lingering in content_chunks get cleared when the v0.32.7 post-upgrade sweep runs. New test in test/reindex.test.ts seeds a page, runs the sweep, mocks a stale chunker_version=1 without changing compiled_truth, runs the sweep again, asserts chunker_version is bumped despite hash match. F4 — Sync delete/rename still used resolveSlugForPath(path) only, ignoring the new pages.source_path column added in v52. Frontmatter-fallback pages (emoji-only / Thai / Arabic filenames where slugifyPath returns empty and the slug came from the markdown frontmatter) would orphan on delete or rename because the path-derived slug doesn't match the stored slug. Fix: new exported helper resolveSlugByPathOrSourcePath(engine, path, sourceId?) queries pages.source_path first, falls back to resolveSlugForPath when no row matches. Threaded into 3 call sites in sync.ts (un-syncable modified cleanup at :531, deletes at :603, rename oldSlug at :622). Best-effort: query errors fall through to the legacy path so pre-migration brains still work. 3 new test cases in test/sync.test.ts cover: stored-slug lookup hits, fallback when no source_path row exists, and source_id scoping when two sources have the same source_path value. Codex finding #3 (reindex not in CLI_ONLY) was verified as a false positive — CLI_ONLY is the set that doesn't need an engine; reindex correctly belongs to the engine-backed dispatch. 302 wave tests pass / 0 fail. bun run verify green. * docs: update CLAUDE.md + llms-full.txt for v0.32.7 CJK fix wave CLAUDE.md Key Files: added entries for the five new modules introduced by the wave — src/core/cjk.ts (shared detection + delimiters + density threshold), src/core/audit-slug-fallback.ts (weekly JSONL), src/core/embedding-pricing.ts (post-upgrade cost lookup table), src/core/post-upgrade-reembed.ts (prompt + grace window), and src/commands/reindex.ts (chunker_version sweep with forceRechunk). Also noted src/commands/sync.ts:resolveSlugByPathOrSourcePath — the F4 codex post-merge fix that wires the new pages.source_path column into sync delete/rename so frontmatter-fallback pages don't orphan. CLAUDE.md Commands: added a v0.32.7 section covering `gbrain reindex --markdown`, the new doctor slug_fallback_audit check, PGLite CJK keyword fallback in `gbrain search`, and the post-upgrade chunker-bump cost prompt with its env-var overrides. llms-full.txt: regenerated via bun run build:llms (CI gate runs the generator on every release; commit must include the bundle). README.md: no changes needed — v0.32.7 is internal correctness across the existing pipeline, not a new skill or setup story. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: vinsew <vinsew@users.noreply.github.com> Co-authored-by: 313094319-sudo <313094319-sudo@users.noreply.github.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
269 lines
9.3 KiB
TypeScript
269 lines
9.3 KiB
TypeScript
import { describe, test, expect } from 'bun:test';
|
|
import { slugifySegment, slugifyPath, SLUG_SEGMENT_PATTERN } from '../src/core/sync.ts';
|
|
|
|
// Test the validateSlug behavior via the engine
|
|
// We can't import validateSlug directly (it's private), so we test through putPage mock behavior
|
|
// Instead, test the regex logic directly
|
|
|
|
function validateSlug(slug: string): boolean {
|
|
// Mirrors the logic in postgres-engine.ts
|
|
if (!slug || /(^|\/)\.\.($|\/)/.test(slug) || /^\//.test(slug)) return false;
|
|
return true;
|
|
}
|
|
|
|
describe('slugifySegment', () => {
|
|
test('converts spaces to hyphens', () => {
|
|
expect(slugifySegment('hello world')).toBe('hello-world');
|
|
});
|
|
|
|
test('strips special characters', () => {
|
|
expect(slugifySegment('notes (march 2024)')).toBe('notes-march-2024');
|
|
});
|
|
|
|
test('normalizes unicode accents', () => {
|
|
expect(slugifySegment('caf\u00e9')).toBe('cafe');
|
|
});
|
|
|
|
test('collapses multiple hyphens', () => {
|
|
expect(slugifySegment('a - b')).toBe('a-b');
|
|
});
|
|
|
|
test('strips leading and trailing hyphens', () => {
|
|
expect(slugifySegment(' hello ')).toBe('hello');
|
|
});
|
|
|
|
test('preserves dots', () => {
|
|
expect(slugifySegment('v1.0.0')).toBe('v1.0.0');
|
|
});
|
|
|
|
test('preserves underscores', () => {
|
|
expect(slugifySegment('my_file_name')).toBe('my_file_name');
|
|
});
|
|
|
|
test('lowercases', () => {
|
|
expect(slugifySegment('Apple Notes')).toBe('apple-notes');
|
|
});
|
|
|
|
test('returns empty for all-special-chars input', () => {
|
|
expect(slugifySegment('!!!')).toBe('');
|
|
});
|
|
|
|
test('handles curly quotes and ellipsis', () => {
|
|
expect(slugifySegment('she\u2026said \u201chello\u201d')).toBe('shesaid-hello');
|
|
});
|
|
});
|
|
|
|
describe('slugifyPath', () => {
|
|
test('slugifies each path segment independently', () => {
|
|
expect(slugifyPath('Apple Notes/file name.md')).toBe('apple-notes/file-name');
|
|
});
|
|
|
|
test('already-valid slugs unchanged', () => {
|
|
expect(slugifyPath('people/alice-smith.md')).toBe('people/alice-smith');
|
|
});
|
|
|
|
test('strips .md extension case-insensitively', () => {
|
|
expect(slugifyPath('notes/file.MD')).toBe('notes/file');
|
|
});
|
|
|
|
test('strips .mdx extension', () => {
|
|
expect(slugifyPath('components/hero.mdx')).toBe('components/hero');
|
|
expect(slugifyPath('docs/guide.MDX')).toBe('docs/guide');
|
|
});
|
|
|
|
test('normalizes backslashes', () => {
|
|
expect(slugifyPath('notes\\file.md')).toBe('notes/file');
|
|
});
|
|
|
|
test('strips leading ./', () => {
|
|
expect(slugifyPath('./notes/file.md')).toBe('notes/file');
|
|
});
|
|
|
|
test('filters empty segments from all-special-chars dirs', () => {
|
|
expect(slugifyPath('!!!/file.md')).toBe('file');
|
|
});
|
|
|
|
test('preserves dots in filenames', () => {
|
|
expect(slugifyPath('notes/v1.0.0.md')).toBe('notes/v1.0.0');
|
|
});
|
|
|
|
test('handles consecutive slashes', () => {
|
|
expect(slugifyPath('a//b.md')).toBe('a/b');
|
|
});
|
|
|
|
// Bug report example transformations
|
|
test('Apple Notes example 1', () => {
|
|
expect(slugifyPath('Apple Notes/2017-05-03 ohmygreen.md')).toBe('apple-notes/2017-05-03-ohmygreen');
|
|
});
|
|
|
|
test('Apple Notes example 2', () => {
|
|
expect(slugifyPath('Apple Notes/2018-12-14 Team Photo.md')).toBe('apple-notes/2018-12-14-team-photo');
|
|
});
|
|
|
|
test('Apple Notes example 3 (parens and ellipsis)', () => {
|
|
const input = 'Apple Notes/2017-05-05 Today I had a touch base with Kavita for the meeting on Monday. (she\u2026.md';
|
|
const result = slugifyPath(input);
|
|
expect(result).toBe('apple-notes/2017-05-05-today-i-had-a-touch-base-with-kavita-for-the-meeting-on-monday.-she');
|
|
});
|
|
|
|
test('meetings transcript example', () => {
|
|
expect(slugifyPath('meetings/transcripts/2026-01-21 maria - california c4 collaboration discussion.md'))
|
|
.toBe('meetings/transcripts/2026-01-21-maria-california-c4-collaboration-discussion');
|
|
});
|
|
});
|
|
|
|
describe('validateSlug (widened for any filename chars)', () => {
|
|
test('accepts clean slug', () => {
|
|
expect(validateSlug('people/sarah-chen')).toBe(true);
|
|
});
|
|
|
|
test('accepts slug with spaces (Apple Notes)', () => {
|
|
expect(validateSlug('apple-notes/2017-05-03 ohmygreen')).toBe(true);
|
|
});
|
|
|
|
test('accepts slug with parens', () => {
|
|
expect(validateSlug('apple-notes/notes (march 2024)')).toBe(true);
|
|
});
|
|
|
|
test('accepts slug with special chars', () => {
|
|
expect(validateSlug("notes/it's a test")).toBe(true);
|
|
expect(validateSlug('notes/file@2024')).toBe(true);
|
|
expect(validateSlug('notes/50% complete')).toBe(true);
|
|
});
|
|
|
|
test('accepts slug with unicode', () => {
|
|
expect(validateSlug('notes/日本語テスト')).toBe(true);
|
|
expect(validateSlug('notes/café-meeting')).toBe(true);
|
|
});
|
|
|
|
test('rejects empty slug', () => {
|
|
expect(validateSlug('')).toBe(false);
|
|
});
|
|
|
|
test('rejects path traversal', () => {
|
|
expect(validateSlug('../etc/passwd')).toBe(false);
|
|
expect(validateSlug('notes/../../etc')).toBe(false);
|
|
});
|
|
|
|
test('rejects leading slash', () => {
|
|
expect(validateSlug('/absolute/path')).toBe(false);
|
|
});
|
|
|
|
test('accepts slug with dots (not traversal)', () => {
|
|
expect(validateSlug('notes/v1.0.0')).toBe(true);
|
|
expect(validateSlug('notes/file.name.md')).toBe(true);
|
|
});
|
|
|
|
// Ellipsis false positive regression tests (PR #31)
|
|
test('accepts slug with ellipsis (...)', () => {
|
|
expect(validateSlug('ted-talks/i got 99 problems... palsy is just one')).toBe(true);
|
|
expect(validateSlug('huberman-lab/how...works')).toBe(true);
|
|
expect(validateSlug('multiple...dots...here')).toBe(true);
|
|
});
|
|
|
|
test('accepts slug with double dots in non-traversal positions', () => {
|
|
expect(validateSlug('notes/v1..2')).toBe(true);
|
|
expect(validateSlug('file..name')).toBe(true);
|
|
});
|
|
|
|
test('rejects bare .. as slug', () => {
|
|
expect(validateSlug('..')).toBe(false);
|
|
});
|
|
|
|
test('rejects .. at start of path', () => {
|
|
expect(validateSlug('../etc/passwd')).toBe(false);
|
|
});
|
|
|
|
test('rejects .. in middle of path', () => {
|
|
expect(validateSlug('notes/../../etc')).toBe(false);
|
|
expect(validateSlug('a/../b')).toBe(false);
|
|
});
|
|
|
|
test('rejects .. at end of path', () => {
|
|
expect(validateSlug('notes/..')).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe('CJK slug preservation (v0.32.7)', () => {
|
|
test('Han characters preserved (Chinese)', () => {
|
|
expect(slugifySegment('品牌圣经')).toBe('品牌圣经');
|
|
expect(slugifySegment('销售论证文档')).toBe('销售论证文档');
|
|
});
|
|
|
|
test('Hiragana preserved', () => {
|
|
expect(slugifySegment('ひらがなテスト')).toBe('ひらがなテスト');
|
|
});
|
|
|
|
test('Katakana preserved (full-width)', () => {
|
|
expect(slugifySegment('カタカナテスト')).toBe('カタカナテスト');
|
|
});
|
|
|
|
test('Hangul Syllables preserved (Korean)', () => {
|
|
expect(slugifySegment('한글테스트')).toBe('한글테스트');
|
|
});
|
|
|
|
test('NFC re-composition for Hangul', () => {
|
|
// NFD decomposes Hangul Syllables into conjoining Jamo (U+1100 block).
|
|
// Without normalize('NFC') after the accent strip, the result would
|
|
// collapse to empty because Jamo sits outside the Syllables range.
|
|
const decomposed = '한글테스트'.normalize('NFD');
|
|
expect(slugifySegment(decomposed)).toBe('한글테스트');
|
|
});
|
|
|
|
test('mixed CJK + ASCII: lowercase ASCII, preserve CJK', () => {
|
|
expect(slugifySegment('ICP-理想客户画像')).toBe('icp-理想客户画像');
|
|
});
|
|
|
|
test('collision regression: different CJK names produce different slugs', () => {
|
|
expect(slugifySegment('品牌圣经')).not.toBe(slugifySegment('销售论证文档'));
|
|
});
|
|
|
|
test('slugifyPath preserves pure-CJK files', () => {
|
|
expect(slugifyPath('inbox/品牌圣经.md')).toBe('inbox/品牌圣经');
|
|
});
|
|
|
|
test('slugifyPath collision regression at path level', () => {
|
|
expect(slugifyPath('inbox/品牌圣经.md')).not.toBe(slugifyPath('inbox/销售论证文档.md'));
|
|
});
|
|
|
|
test('CJK directory names preserved', () => {
|
|
expect(slugifyPath('档案/2024-记录.md')).toBe('档案/2024-记录');
|
|
});
|
|
|
|
test('REGRESSION: café still slugifies to cafe (NFD-strip-accents chain preserved)', () => {
|
|
// Iron rule: the NFC re-normalize must not break existing Latin-with-accent
|
|
// behavior. café (Latin) decomposes to 'cafe' + combining acute under NFD,
|
|
// strip-combining drops the acute, NFC recomposes 'cafe', then lowercase.
|
|
expect(slugifySegment('café')).toBe('cafe');
|
|
});
|
|
|
|
test('REGRESSION: existing English slugs unchanged', () => {
|
|
expect(slugifySegment('hello world')).toBe('hello-world');
|
|
expect(slugifySegment('notes (march 2024)')).toBe('notes-march-2024');
|
|
});
|
|
});
|
|
|
|
describe('SLUG_SEGMENT_PATTERN (v0.32.7)', () => {
|
|
test('matches pure-CJK slug segments', () => {
|
|
expect(SLUG_SEGMENT_PATTERN.test('品牌圣经')).toBe(true);
|
|
expect(SLUG_SEGMENT_PATTERN.test('한글')).toBe(true);
|
|
});
|
|
|
|
test('matches existing ASCII slug shapes', () => {
|
|
expect(SLUG_SEGMENT_PATTERN.test('hello-world')).toBe(true);
|
|
expect(SLUG_SEGMENT_PATTERN.test('companies/acme.io')).toBe(true);
|
|
expect(SLUG_SEGMENT_PATTERN.test('people/foo_bar')).toBe(true);
|
|
});
|
|
|
|
test('matches mixed CJK + ASCII', () => {
|
|
expect(SLUG_SEGMENT_PATTERN.test('icp-理想客户画像')).toBe(true);
|
|
});
|
|
|
|
test('REGRESSION: rejects non-CJK Unicode (Vietnamese)', () => {
|
|
// Scope is CJK only; Vietnamese with combining diacritics stays rejected
|
|
// until we widen to Unicode property escapes in v0.33+.
|
|
const result = 'người-dùng'.match(new RegExp(`^${SLUG_SEGMENT_PATTERN.source}$`));
|
|
expect(result).toBeNull();
|
|
});
|
|
});
|