Files
gbrain/test/chunkers/recursive.test.ts
T
ecebd5552a feat: GBrain v0.2.0 — incremental sync, file storage, install skill (#2)
* refactor: extract importFile from import.ts + add tag reconciliation

Shared single-file import function used by both import and sync.
Adds tag reconciliation (removes stale tags on reimport), >1MB file
skip, and import->sync checkpoint continuity (writes git HEAD to
config table after import so sync picks up seamlessly).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add sync pure functions, updateSlug engine method, and sync tests

- buildSyncManifest: parses git diff --name-status -M output
- isSyncable: filters to .md pages, excludes hidden/ops/.raw/skip-list
- pathToSlug: converts file paths to page slugs with optional prefix
- updateSlug: renames page slug in-place (preserves page_id, chunks, embeddings)
- rewriteLinks: stub for v0.2 (FKs use page_id, already correct)
- 20 new tests, all passing (39 total across 3 files)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add gbrain sync command with CLI, MCP, and watch mode

18-step sync protocol: read config, git pull, ancestry validation,
git diff --name-status -M for net changes, isSyncable filter, process
deletes/renames/adds/modifies via importFile, batch optimization,
sync state checkpoint in Postgres config table. Watch mode with
polling and consecutive error counter. MCP sync_brain tool returns
structured SyncResult. Stale page deletion for un-syncable files.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add files table, gbrain files commands, and config show redaction

- files table: page_slug FK with ON DELETE SET NULL + ON UPDATE CASCADE,
  storage_path, storage_url, mime_type, content_hash for dedup
- gbrain files list/upload/sync/verify commands for Supabase Storage
- gbrain config show redacts postgresql:// passwords and secret keys
- CLI help updated with FILES section

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add install skill for GBrain onboarding

6-phase install workflow: environment discovery, Supabase setup (magic
path via CLI OAuth or fallback 2-copy-paste), init + import, ongoing
sync cron, optional file migration with mandatory verification, and
agent teaching (AGENTS.md rules). Every error gets what + why + fix.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update project documentation for v0.2.0

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add v0.2 features to README (sync, files, install skill)

README.md: added sync command to IMPORT/EXPORT section, added FILES
section with 4 commands, added files table to schema diagram, added
install skill to skills table, updated MCP tools count from 20 to 21
(sync_brain added).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: OpenClaw DX improvements (skill count, upgrade docs, config show help)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: consolidate version to single source of truth

Create src/version.ts that reads from package.json via static import
(safe for bun compiled binaries). Update mcp/server.ts from hardcoded
'0.1.0' to use shared VERSION. Bump skills/manifest.json to 0.2.0.

* fix: upgrade detection order, npm→bun naming, clawhub false positives

Reorder detection: node_modules first, binary second, clawhub last.
Rename 'npm' install method to 'bun'. Use 'clawhub --version' instead
of 'which clawhub' to avoid false positives from dangling symlinks.
Add 120s timeout to execSync calls to prevent hanging. Add --help flag.

* feat: per-command --help, unknown command check before DB connection

Add COMMAND_HELP map covering all 28 commands. Check --help before
init/upgrade dispatch and before connectEngine() so help works without
a database. Use COMMAND_HELP keys as known-command set to catch unknown
commands before wasting a DB round-trip.

* docs: standardize npm references to bun, add Upgrade section to README

Fix init.ts: npx→bunx, npm→bun for supabase CLI guidance.
Fix README: npm install→bun add for standalone CLI install.
Add ## Upgrade section to README with all three install methods.
Update install skill Upgrading section to list bun, ClawHub, and binary.

* test: full coverage audit — CLI dispatch, upgrade detection, config, edge cases

New test files:
- test/cli.test.ts: COMMAND_HELP ↔ switch consistency, version from
  package.json, per-command --help, unknown command handling, global help
- test/upgrade.test.ts: detection order verification, npm→bun naming,
  clawhub --version (not which), timeout presence
- test/config.test.ts: redactUrl for postgresql URLs, edge cases

Extended existing tests:
- test/sync.test.ts: empty string pathToSlug, uppercase .MD rejection,
  deeply nested files, multiple renames, unknown status codes
- test/markdown.test.ts: multiple --- separators, missing frontmatter,
  no frontmatter at all, empty string, type inference from paths

Tests: 39 → 83 (+44 new). All pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: 100% coverage — import-file mock engine, files utils, chunker edge cases

New test files:
- test/import-file.test.ts (9 tests): mock BrainEngine to test importFile
  without DB — MAX_FILE_SIZE skip, content_hash dedup, tag reconciliation
  (remove stale + add new), compiled_truth/timeline chunking, noEmbed flag,
  sequential chunk_index
- test/files.test.ts (22 tests): getMimeType for all extensions + uppercase
  + unknown + no-extension, fileHash consistency + different content + empty,
  collectFiles pattern (skip .md, skip hidden dirs, recurse, sorted output)

Extended:
- test/chunkers/recursive.test.ts (+6 tests): single newline splits,
  word-only text, clause delimiters, lossless preservation, default options,
  mixed delimiter hierarchy

Tests: 83 → 118 (+35 new). All pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 16:50:15 -07:00

136 lines
5.2 KiB
TypeScript

import { describe, test, expect } from 'bun:test';
import { chunkText } from '../../src/core/chunkers/recursive.ts';
describe('Recursive Text Chunker', () => {
test('returns empty array for empty input', () => {
expect(chunkText('')).toEqual([]);
expect(chunkText(' ')).toEqual([]);
});
test('returns single chunk for short text', () => {
const text = 'Hello world. This is a short text.';
const chunks = chunkText(text);
expect(chunks).toHaveLength(1);
expect(chunks[0].text).toBe(text.trim());
expect(chunks[0].index).toBe(0);
});
test('splits at paragraph boundaries', () => {
const paragraph = 'word '.repeat(200).trim();
const text = paragraph + '\n\n' + paragraph;
const chunks = chunkText(text, { chunkSize: 250 });
expect(chunks.length).toBeGreaterThanOrEqual(2);
});
test('respects chunk size target', () => {
const text = 'word '.repeat(1000).trim();
const chunks = chunkText(text, { chunkSize: 100 });
for (const chunk of chunks) {
const wordCount = chunk.text.split(/\s+/).length;
// Allow up to 1.5x target due to greedy merge
expect(wordCount).toBeLessThanOrEqual(150);
}
});
test('applies overlap between chunks', () => {
const text = 'word '.repeat(1000).trim();
const chunks = chunkText(text, { chunkSize: 100, chunkOverlap: 20 });
expect(chunks.length).toBeGreaterThan(1);
// Second chunk should start with words from end of first chunk
// (overlap means shared content between adjacent chunks)
expect(chunks[1].text.length).toBeGreaterThan(0);
});
test('splits at sentence boundaries', () => {
const sentences = Array.from({ length: 50 }, (_, i) =>
`This is sentence number ${i} with some content about topic ${i}.`
).join(' ');
const chunks = chunkText(sentences, { chunkSize: 50 });
expect(chunks.length).toBeGreaterThan(1);
// Each chunk should end near a sentence boundary
for (const chunk of chunks.slice(0, -1)) {
// Allow for overlap text, but the core content should have sentence endings
expect(chunk.text).toMatch(/[.!?]/);
}
});
test('assigns sequential indices', () => {
const text = 'word '.repeat(1000).trim();
const chunks = chunkText(text, { chunkSize: 100 });
for (let i = 0; i < chunks.length; i++) {
expect(chunks[i].index).toBe(i);
}
});
test('handles single word input', () => {
const chunks = chunkText('hello');
expect(chunks).toHaveLength(1);
expect(chunks[0].text).toBe('hello');
});
test('handles unicode text', () => {
const text = 'Bonjour le monde. ' + 'Ceci est un texte en francais. '.repeat(100);
const chunks = chunkText(text, { chunkSize: 50 });
expect(chunks.length).toBeGreaterThan(1);
expect(chunks[0].text).toContain('Bonjour');
});
test('splits at single newline (line-level) when paragraphs are absent', () => {
// Lines without double newlines should still split at single newlines
const lines = Array(100).fill('This is a single line of text.').join('\n');
const chunks = chunkText(lines, { chunkSize: 20 });
expect(chunks.length).toBeGreaterThan(1);
});
test('handles text with only whitespace delimiters (word-level split)', () => {
// No sentences, no newlines, just words
const words = Array(200).fill('word').join(' ');
const chunks = chunkText(words, { chunkSize: 50 });
expect(chunks.length).toBeGreaterThan(1);
for (const chunk of chunks) {
expect(chunk.text.trim().length).toBeGreaterThan(0);
}
});
test('handles clause-level delimiters (semicolons, colons, commas)', () => {
// Text with clauses but no sentence endings
const text = Array(100).fill('clause one; clause two: clause three, clause four').join(' ');
const chunks = chunkText(text, { chunkSize: 30 });
expect(chunks.length).toBeGreaterThan(1);
});
test('preserves content across chunks (lossless)', () => {
const original = 'First paragraph.\n\nSecond paragraph.\n\nThird paragraph.';
const chunks = chunkText(original, { chunkSize: 5, chunkOverlap: 0 });
// With no overlap, all text should appear in chunks
const reconstructed = chunks.map(c => c.text).join(' ');
expect(reconstructed).toContain('First paragraph');
expect(reconstructed).toContain('Second paragraph');
expect(reconstructed).toContain('Third paragraph');
});
test('default options produce reasonable chunks', () => {
// Large text with defaults (300 words, 50 overlap)
const text = Array(500).fill('This is a test sentence with several words.').join(' ');
const chunks = chunkText(text);
expect(chunks.length).toBeGreaterThan(1);
for (const chunk of chunks) {
const wordCount = chunk.text.split(/\s+/).length;
// Should be roughly 300 words, with 1.5x tolerance
expect(wordCount).toBeLessThanOrEqual(500);
}
});
test('handles mixed delimiter hierarchy', () => {
const text = [
'Paragraph one has sentences. And more sentences! Really?',
'',
'Paragraph two; with clauses: and more, clauses here.',
'',
'Paragraph three.\nWith line breaks.\nAnd more lines.',
].join('\n');
const chunks = chunkText(text, { chunkSize: 10 });
expect(chunks.length).toBeGreaterThan(1);
});
});