mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* fix(sync): remove nested transaction that deadlocks > 10 file syncs sync.ts wraps the add/modify loop in engine.transaction(), and each importFromContent inside opens another one. PGLite's _runExclusiveTransaction is a non-reentrant mutex — the second call queues on the mutex the first is holding, and the process hangs forever in ep_poll. Reproduced with a 15-file commit: unpatched hangs, patched runs in 3.4s. Fix drops the outer wrap; per-file atomicity is correct anyway (one file's failure should not roll back the others). (cherry picked from commit4a1ac00105) * test(sync): regression guard for #132 top-level engine.transaction wrap Reads src/commands/sync.ts verbatim and asserts no uncommented engine.transaction() call appears above the add/modify loop. Protects against silent reintroduction of the nested-mutex deadlock that hung > 10-file syncs forever in ep_poll. * feat(utils): tryParseEmbedding() skip+warn sibling for availability path parseEmbedding() throws on structural corruption — right call for ingest/ migrate paths where silent skips would be data loss. Wrong call for search/rescore paths where one corrupt row in 10K would kill every query that touches it. tryParseEmbedding() wraps parseEmbedding in try/catch: returns null on any shape that would throw, warns once per session so the bad row is visible in logs. Use it anywhere we'd rather degrade ranking than blow up the whole query. Retrofit postgres-engine.getEmbeddingsByChunkIds (the #175 slice call site) — the 5-line rescore loop was the direct motivator. Keep the throwing parseEmbedding() for everything else (pglite-engine rowToChunk, migrate-engine round-trips, ingest). * postgres-engine: scope search statement_timeout to the transaction searchKeyword and searchVector run on a pooled postgres.js client (max: 10 by default). The original code bounded each search with await sql`SET statement_timeout = '8s'` try { await sql`<query>` } finally { await sql`SET statement_timeout = '0'` } but every tagged template is an independent round-trip that picks an arbitrary connection from the pool. The SET, the query, and the reset could all land on DIFFERENT connections. In practice the GUC sticks to whichever connection ran the SET and then gets returned to the pool — the next unrelated caller on that connection inherits the 8s timeout (clipping legitimate long queries) or the reset-to-0 (disabling the guard for whoever expected it). A crash in the middle leaves the state set permanently. Wrap each search in sql.begin(async sql => …). postgres.js reserves a single connection for the transaction body, so the SET LOCAL, the query, and the implicit COMMIT all run on the same connection. SET LOCAL scopes the GUC to the transaction — COMMIT or ROLLBACK restores the previous value automatically, regardless of the code path out. Error paths can no longer leak the GUC. No API change. Timeout value and semantics are identical (8s cap on search queries, no effect on embed --all / bulk import which runs outside these methods). Only one transaction per search — BEGIN + COMMIT round-trips are negligible next to a ranked FTS or pgvector query. Also closes the earlier audit finding R4-F002 which reported the same pattern on searchKeyword. This PR covers both searchKeyword and searchVector so the pool-leak class is fully closed. Tests (test/postgres-engine.test.ts, new file): - No bare SET statement_timeout remains after stripping comments. - searchKeyword and searchVector each wrap their query in sql.begin. - Both use SET LOCAL. - Neither explicitly clears the timeout with SET statement_timeout=0. Source-level guardrails keep the fast unit suite DB-free. Live Postgres coverage of the search path is in test/e2e/search-quality.test.ts, which continues to exercise these methods end-to-end against pgvector when DATABASE_URL is set. (cherry picked from commit6146c3b470) * feat(orphans): add gbrain orphans command for finding under-connected pages Surfaces pages with zero inbound wikilinks. Essential for content enrichment cycles in KBs with 1000+ pages. By default filters out auto-generated pages, raw sources, and pseudo-pages where no inbound links is expected; --include-pseudo to disable. Supports text (grouped by domain), --json, --count outputs. Also exposed as find_orphans MCP operation. Tests cover basic detection, filtering, all output modes. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> (cherry picked from commitf50954f8e0) * feat(extract): support Obsidian wikilinks + wiki-style domain slugs in canonical extractor extractEntityRefs now recognizes both syntaxes equally: [Name](people/slug) -- upstream original [[people/slug|Name]] -- Obsidian wikilink (new) Extends DIR_PATTERN to include domain-organized wiki slugs used by Karpathy-style knowledge bases: - entities (legacy prefix some brains keep during migration) - projects (gbrain canonical, was missing from regex) - tech, finance, personal, openclaw (domain-organized wiki roots) Before this change, a 2,100-page brain with wikilinks throughout extracted zero auto-links on put_page because the regex only matched markdown-style [name](path). After: 1,377 new typed edges on a single extract --source db pass over the same corpus. Matches the behavior of the extract.ts filesystem walker (which already handled wikilinks as of the wiki-markdown-compat fix wave), so the db and fs sources now produce the same link graph from the same content. Both patterns share the DIR_PATTERN constant so adding a new entity dir only requires updating one string. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> (cherry picked from commit1cfb15679a) * feat(doctor): jsonb_integrity + markdown_body_completeness detection Add two v0.12.1-era reliability checks to `gbrain doctor`: - `jsonb_integrity` scans the 4 known write sites from the v0.12.0 double-encode bug (pages.frontmatter, raw_data.data, ingest_log.pages_updated, files.metadata) and reports rows where jsonb_typeof(col) = 'string'. The fix hint points at `gbrain repair-jsonb` (the standalone repair command shipped in v0.12.1). - `markdown_body_completeness` flags pages whose compiled_truth is <30% of the raw source content length when raw has multiple H2/H3 boundaries. Heuristic only; suggests `gbrain sync --force` or `gbrain import --force <slug>`. Also adds test/e2e/jsonb-roundtrip.test.ts — the regression coverage that should have caught the original double-encode bug. Hits all four write sites against real Postgres and asserts jsonb_typeof='object' plus `->>'key'` returns the expected scalar. Detection only: doctor diagnoses, `gbrain repair-jsonb` treats. No overlap with the standalone repair path. * chore: bump to v0.12.3 + changelog (reliability wave) Master shipped v0.12.1 (extract N+1 + migration timeout) and v0.12.2 (JSONB double-encode + splitBody + wiki types + parseEmbedding) while this wave was mid-flight. Ships the remaining pieces as v0.12.3: - sync deadlock (#132, @sunnnybala) - statement_timeout scoping (#158, @garagon) - Obsidian wikilinks + domain patterns (#187 slice, @knee5) - gbrain orphans command (#187 slice, @knee5) - tryParseEmbedding() availability helper - doctor detection for jsonb_integrity + markdown_body_completeness No schema, no migration, no data touch. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs: update project documentation for v0.12.3 CLAUDE.md: - Add src/commands/orphans.ts entry - Expand src/commands/doctor.ts with v0.12.3 jsonb_integrity + markdown_body_completeness check descriptions - Update src/core/link-extraction.ts to mention Obsidian wikilinks + extended DIR_PATTERN (entities/projects/tech/finance/personal/openclaw) - Update src/core/utils.ts to mention tryParseEmbedding sibling - Update src/core/postgres-engine.ts to note statement_timeout scoping + tryParseEmbedding usage in getEmbeddingsByChunkIds - Add Key commands added in v0.12.3 section (orphans, doctor checks) - Add test/orphans.test.ts, test/postgres-engine.test.ts, updated descriptions for test/sync.test.ts, test/doctor.test.ts, test/utils.test.ts - Add test/e2e/jsonb-roundtrip.test.ts with note on intentional overlap - Bump operation count from ~36 to ~41 (find_orphans shipped in v0.12.3) README.md: - Add gbrain orphans to ADMIN commands block Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: sunnnybala <dhruvagarwal5018@gmail.com> Co-authored-by: Gustavo Aragon <gustavoraularagon@gmail.com> Co-authored-by: Clevin Canales <clevin@Clevins-MacBook-Pro.local> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Clevin Canales <clev.canales@gmail.com>
204 lines
6.3 KiB
TypeScript
204 lines
6.3 KiB
TypeScript
import { describe, test, expect } from 'bun:test';
|
|
import {
|
|
shouldExclude,
|
|
deriveDomain,
|
|
formatOrphansText,
|
|
type OrphanPage,
|
|
type OrphanResult,
|
|
} from '../src/commands/orphans.ts';
|
|
|
|
// --- shouldExclude ---
|
|
|
|
describe('shouldExclude', () => {
|
|
test('excludes pseudo-page _atlas', () => {
|
|
expect(shouldExclude('_atlas')).toBe(true);
|
|
});
|
|
|
|
test('excludes pseudo-page _index', () => {
|
|
expect(shouldExclude('_index')).toBe(true);
|
|
});
|
|
|
|
test('excludes pseudo-page _stats', () => {
|
|
expect(shouldExclude('_stats')).toBe(true);
|
|
});
|
|
|
|
test('excludes pseudo-page _orphans', () => {
|
|
expect(shouldExclude('_orphans')).toBe(true);
|
|
});
|
|
|
|
test('excludes pseudo-page _scratch', () => {
|
|
expect(shouldExclude('_scratch')).toBe(true);
|
|
});
|
|
|
|
test('excludes pseudo-page claude', () => {
|
|
expect(shouldExclude('claude')).toBe(true);
|
|
});
|
|
|
|
test('excludes auto-generated _index suffix', () => {
|
|
expect(shouldExclude('companies/_index')).toBe(true);
|
|
expect(shouldExclude('people/_index')).toBe(true);
|
|
});
|
|
|
|
test('excludes auto-generated /log suffix', () => {
|
|
expect(shouldExclude('projects/acme/log')).toBe(true);
|
|
});
|
|
|
|
test('excludes raw source slugs', () => {
|
|
expect(shouldExclude('companies/acme/raw/crustdata')).toBe(true);
|
|
});
|
|
|
|
test('excludes deny-prefix: output/', () => {
|
|
expect(shouldExclude('output/2026-q1')).toBe(true);
|
|
});
|
|
|
|
test('excludes deny-prefix: dashboards/', () => {
|
|
expect(shouldExclude('dashboards/metrics')).toBe(true);
|
|
});
|
|
|
|
test('excludes deny-prefix: scripts/', () => {
|
|
expect(shouldExclude('scripts/ingest-runner')).toBe(true);
|
|
});
|
|
|
|
test('excludes deny-prefix: templates/', () => {
|
|
expect(shouldExclude('templates/meeting-note')).toBe(true);
|
|
});
|
|
|
|
test('excludes deny-prefix: openclaw/config/', () => {
|
|
expect(shouldExclude('openclaw/config/agent')).toBe(true);
|
|
});
|
|
|
|
test('excludes first-segment: scratch', () => {
|
|
expect(shouldExclude('scratch/idea-dump')).toBe(true);
|
|
});
|
|
|
|
test('excludes first-segment: thoughts', () => {
|
|
expect(shouldExclude('thoughts/2026-04-17')).toBe(true);
|
|
});
|
|
|
|
test('excludes first-segment: catalog', () => {
|
|
expect(shouldExclude('catalog/tools')).toBe(true);
|
|
});
|
|
|
|
test('excludes first-segment: entities', () => {
|
|
expect(shouldExclude('entities/product-hunt')).toBe(true);
|
|
});
|
|
|
|
test('does NOT exclude a normal content page', () => {
|
|
expect(shouldExclude('companies/acme')).toBe(false);
|
|
expect(shouldExclude('people/jane-doe')).toBe(false);
|
|
expect(shouldExclude('projects/gbrain')).toBe(false);
|
|
});
|
|
|
|
test('does NOT exclude a page ending with log-like text that is not /log', () => {
|
|
expect(shouldExclude('devlog')).toBe(false);
|
|
expect(shouldExclude('changelog')).toBe(false);
|
|
});
|
|
});
|
|
|
|
// --- deriveDomain ---
|
|
|
|
describe('deriveDomain', () => {
|
|
test('uses frontmatter domain when present', () => {
|
|
expect(deriveDomain('companies', 'companies/acme')).toBe('companies');
|
|
});
|
|
|
|
test('falls back to first slug segment', () => {
|
|
expect(deriveDomain(null, 'people/jane-doe')).toBe('people');
|
|
expect(deriveDomain(undefined, 'projects/gbrain')).toBe('projects');
|
|
});
|
|
|
|
test('returns root for single-segment slugs with no frontmatter', () => {
|
|
expect(deriveDomain(null, 'readme')).toBe('readme');
|
|
});
|
|
|
|
test('ignores empty-string frontmatter domain', () => {
|
|
expect(deriveDomain('', 'people/alice')).toBe('people');
|
|
});
|
|
|
|
test('ignores whitespace-only frontmatter domain', () => {
|
|
expect(deriveDomain(' ', 'people/alice')).toBe('people');
|
|
});
|
|
});
|
|
|
|
// --- formatOrphansText ---
|
|
|
|
describe('formatOrphansText', () => {
|
|
function makeResult(orphans: OrphanPage[], overrides?: Partial<OrphanResult>): OrphanResult {
|
|
return {
|
|
orphans,
|
|
total_orphans: orphans.length,
|
|
total_linkable: orphans.length + 50,
|
|
total_pages: orphans.length + 60,
|
|
excluded: 10,
|
|
...overrides,
|
|
};
|
|
}
|
|
|
|
test('shows summary line', () => {
|
|
const result = makeResult([]);
|
|
const out = formatOrphansText(result);
|
|
expect(out).toContain('0 orphans out of');
|
|
expect(out).toContain('total');
|
|
expect(out).toContain('excluded');
|
|
});
|
|
|
|
test('shows "No orphan pages found." when empty', () => {
|
|
const out = formatOrphansText(makeResult([]));
|
|
expect(out).toContain('No orphan pages found.');
|
|
});
|
|
|
|
test('groups orphans by domain', () => {
|
|
const orphans: OrphanPage[] = [
|
|
{ slug: 'companies/acme', title: 'Acme Corp', domain: 'companies' },
|
|
{ slug: 'people/alice', title: 'Alice', domain: 'people' },
|
|
{ slug: 'companies/beta', title: 'Beta Inc', domain: 'companies' },
|
|
];
|
|
const out = formatOrphansText(makeResult(orphans));
|
|
expect(out).toContain('[companies]');
|
|
expect(out).toContain('[people]');
|
|
// companies section should appear before people (alphabetical)
|
|
const companiesIdx = out.indexOf('[companies]');
|
|
const peopleIdx = out.indexOf('[people]');
|
|
expect(companiesIdx).toBeLessThan(peopleIdx);
|
|
});
|
|
|
|
test('sorts orphans alphabetically within each domain group', () => {
|
|
const orphans: OrphanPage[] = [
|
|
{ slug: 'companies/zeta', title: 'Zeta', domain: 'companies' },
|
|
{ slug: 'companies/alpha', title: 'Alpha', domain: 'companies' },
|
|
{ slug: 'companies/beta', title: 'Beta', domain: 'companies' },
|
|
];
|
|
const out = formatOrphansText(makeResult(orphans));
|
|
const alphaIdx = out.indexOf('companies/alpha');
|
|
const betaIdx = out.indexOf('companies/beta');
|
|
const zetaIdx = out.indexOf('companies/zeta');
|
|
expect(alphaIdx).toBeLessThan(betaIdx);
|
|
expect(betaIdx).toBeLessThan(zetaIdx);
|
|
});
|
|
|
|
test('includes slug and title in output', () => {
|
|
const orphans: OrphanPage[] = [
|
|
{ slug: 'companies/acme', title: 'Acme Corp', domain: 'companies' },
|
|
];
|
|
const out = formatOrphansText(makeResult(orphans));
|
|
expect(out).toContain('companies/acme');
|
|
expect(out).toContain('Acme Corp');
|
|
});
|
|
|
|
test('summary line shows correct numbers', () => {
|
|
const orphans: OrphanPage[] = [
|
|
{ slug: 'a/b', title: 'B', domain: 'a' },
|
|
{ slug: 'a/c', title: 'C', domain: 'a' },
|
|
];
|
|
const result: OrphanResult = {
|
|
orphans,
|
|
total_orphans: 2,
|
|
total_linkable: 100,
|
|
total_pages: 120,
|
|
excluded: 20,
|
|
};
|
|
const out = formatOrphansText(result);
|
|
expect(out).toContain('2 orphans out of 100 linkable pages (120 total; 20 excluded)');
|
|
});
|
|
});
|