mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-29 17:20:29 +00:00
fix(extract+migrate): kill N+1 hang + v0.12.0 migration timeout (v0.12.1) (#198)
* feat(engine): add addLinksBatch + addTimelineEntriesBatch via unnest()
Multi-row INSERT...SELECT FROM unnest() JOIN pages ON CONFLICT DO NOTHING
RETURNING 1. 4 array-typed bound parameters (links) or 5 (timeline)
regardless of batch size, sidesteps Postgres's 65535-parameter cap.
Returns count of rows actually inserted (excluding ON CONFLICT no-ops
and JOIN-dropped rows whose slugs don't exist).
Per-row addLink / addTimelineEntry signatures and SQL behavior unchanged.
All 10 existing call sites compile and behave identically.
Tests: 11 PGLite cases (empty batch, missing optionals, within-batch dedup,
JOIN drops missing slug, half-existing batch, batch of 100) + 9 E2E
postgres-engine cases against real Postgres+pgvector.
* fix(migrate): pre-create btree helper in v8 + v9 dedup; bump phaseASchema timeout
Production bug: v0.12.0 schema migration timed out at Supabase Management API's
60s ceiling on brains with 80K+ duplicate timeline rows. The DELETE...USING
self-join was O(n²) without an index on the dedup columns.
Fix: pre-create idx_links_dedup_helper / idx_timeline_dedup_helper on the
dedup columns BEFORE the DELETE, drop after. Turns O(n²) into O(n log n).
On 80K+ rows the migration completes in <1s instead of timing out.
Also bumps the v0.12.0 orchestrator's phaseASchema timeout 60s -> 600s as
belt-and-suspenders for unforeseen slowness.
Exports MIGRATIONS for structural test assertions.
Tests: 2 structural assertions (helper-index DDL must appear in v8/v9 SQL
in the right order — catches regression even at 0-row scale) + 2 behavioral
regression tests (1000-row dedup completes <5s).
* perf(extract): kill N+1 dedup pre-load; switch to batched writes
Production bug: gbrain extract hung 10+ minutes producing zero output on
47K-page brains. The pre-load loop called engine.getLinks(slug) (or
getTimeline) once per page across engine.listPages({limit: 100000}) — 47K
serial round-trips over the Supabase pooler before the first file was read.
Both engines already enforced uniqueness at the SQL layer
(UNIQUE(from, to, link_type) on links, idx_timeline_dedup on timeline_entries).
The in-memory dedup Set was redundant insurance that became the bottleneck.
Fix: delete the pre-load entirely. Buffer 100 candidates per file walk,
flush via engine.addLinksBatch / engine.addTimelineEntriesBatch. ~99% fewer
DB round-trips per re-extract.
Also fixes counter accuracy: 'created' now counts rows actually inserted
(via batch RETURNING 1 row count). Re-run on a fully-extracted brain
prints 'Done: 0 links' instead of lying.
Dry-run mode keeps a per-run dedup Set so duplicate candidates from N
markdown files print exactly once, not N times.
Batch errors are visible in BOTH json and human modes — silent loss of
100 rows is worse than per-row error visibility.
Tests: extract-fs.test.ts (idempotency + truthful counter + dry-run dedup
+ perf regression guard <2s).
* chore: bump version + changelog (v0.12.1)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: update CLAUDE.md for v0.12.1 (batch engine API, test counts)
Reflect what shipped in v0.12.1:
- New engine methods addLinksBatch + addTimelineEntriesBatch (PGLite via
unnest() + manual $N, postgres-engine via INSERT...SELECT FROM
unnest($1::text[], ...) JOIN pages ON CONFLICT DO NOTHING).
- extract.ts no longer pre-loads dedup set; candidates are buffered 100
at a time and flushed via the new batch methods.
- v0.12.0 orchestrator phaseASchema timeout bumped 60s to 600s.
- Test counts 1297 unit / 105 E2E to 1412 unit / 119 E2E.
- New test/extract-fs.test.ts covers the N+1 regression guard.
- BrainEngine method count 37/38 to 40.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
parent
81b3f7afac
commit
699db50a3d
@@ -11,6 +11,23 @@ import type {
|
||||
EngineConfig,
|
||||
} from './types.ts';
|
||||
|
||||
/** Input row for addLinksBatch. Optional fields default to '' (matches NOT NULL DDL). */
|
||||
export interface LinkBatchInput {
|
||||
from_slug: string;
|
||||
to_slug: string;
|
||||
link_type?: string;
|
||||
context?: string;
|
||||
}
|
||||
|
||||
/** Input row for addTimelineEntriesBatch. Optional fields default to '' (matches NOT NULL DDL). */
|
||||
export interface TimelineBatchInput {
|
||||
slug: string;
|
||||
date: string;
|
||||
source?: string;
|
||||
summary: string;
|
||||
detail?: string;
|
||||
}
|
||||
|
||||
/** Maximum results returned by search operations. Internal bulk operations (listPages) are not clamped. */
|
||||
export const MAX_SEARCH_LIMIT = 100;
|
||||
|
||||
@@ -53,6 +70,13 @@ export interface BrainEngine {
|
||||
|
||||
// Links
|
||||
addLink(from: string, to: string, context?: string, linkType?: string): Promise<void>;
|
||||
/**
|
||||
* Bulk insert links via a single multi-row INSERT...SELECT FROM (VALUES) JOIN pages
|
||||
* statement with ON CONFLICT DO NOTHING. Returns the count of rows actually inserted
|
||||
* (RETURNING clause excludes conflicts and JOIN-dropped rows whose slugs don't exist).
|
||||
* Used by extract.ts to avoid 47K sequential round-trips on large brains.
|
||||
*/
|
||||
addLinksBatch(links: LinkBatchInput[]): Promise<number>;
|
||||
/**
|
||||
* Remove links from `from` to `to`. If linkType is provided, only that specific
|
||||
* (from, to, type) row is removed. If omitted, ALL link types between the pair
|
||||
@@ -98,6 +122,13 @@ export interface BrainEngine {
|
||||
entry: TimelineInput,
|
||||
opts?: { skipExistenceCheck?: boolean },
|
||||
): Promise<void>;
|
||||
/**
|
||||
* Bulk insert timeline entries via a single multi-row INSERT...SELECT FROM (VALUES)
|
||||
* JOIN pages statement with ON CONFLICT DO NOTHING. Returns the count of rows
|
||||
* actually inserted (RETURNING excludes conflicts and JOIN-dropped rows whose
|
||||
* slugs don't exist). Used by extract.ts to avoid sequential round-trips.
|
||||
*/
|
||||
addTimelineEntriesBatch(entries: TimelineBatchInput[]): Promise<number>;
|
||||
getTimeline(slug: string, opts?: TimelineOpts): Promise<TimelineEntry[]>;
|
||||
|
||||
// Raw data
|
||||
|
||||
Reference in New Issue
Block a user