mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* fix(reindex): add-only tag reconciliation + DB-only re-chunk preserves frontmatter (#1621) reindex --markdown and re-import no longer wipe DB-side enrichment tags. Tag reconciliation is now ADD-ONLY (import-file.ts): re-import adds current frontmatter tags and never deletes, so auto/dream/signal-detector tags survive. The reindex DB-only fallback reconstructs full markdown via serializeMarkdown so re-chunking a page with no on-disk source preserves frontmatter/title/timeline. * fix(migrations): v0.13.1 grandfather chunked, source-safe, soft-delete-filtered (#1581) phaseCGrandfather rewritten from a per-page getPage+putPage loop (which hung 70+ min on an 82K-page PGLite brain) to a chunked bulk SQL pass keyed on pages.id (NOT slug — slug isn't globally unique), filtering deleted_at IS NULL, with a batched rollback log carrying source identity. * fix(migrations): run schema phases in-process to fix Windows getaddrinfo ENOTFOUND (#1605) The 9 'gbrain init --migrate-only' execSync spawns died on Windows+bun+Supabase (child DNS resolution). runMigrateOnlyCore (extracted from initMigrateOnly) runs the schema bring-up in-process for all engines, unblocking schema_version advancement. Includes async-call-site audit, a wall-clock guard, and a runGbrainSubprocess stderr-capture wrapper for the remaining backfill spawns. * fix(sync): ReDoS hardening + diagnostics for schema-pack regexes (#1569) Input-length cap in runRegexBounded + route the unbounded link-inference path through it (closes the only no-timeout ReDoS hole); star-height lint rule warns on nested-quantifier patterns; --no-schema-pack sync escape hatch; GBRAIN_SYNC_TRACE per-file begin heartbeat; PGLite serve/sync concurrency doc. Defensive hardening + diagnostics — the deterministic ~3100-file wedge root cause remains open (no repro). * docs(todos): file v0.41.37.0 fix-wave follow-ups (#1621/#1605/#1569) * chore: bump version and changelog (v0.41.37.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: sync README + CLAUDE.md for v0.41.37.0 critical fix wave Add reindex add-only tag reconciliation (#1621), v0.13.1 grandfather + Windows in-process migration (#1581/#1605), and schema-pack ReDoS hardening + sync --no-schema-pack / GBRAIN_SYNC_TRACE triage (#1569) to CLAUDE.md key-files annotations and README Troubleshooting. Regenerated llms-full.txt. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): bump llms-full.txt budget 700KB→750KB (CLAUDE.md crossed 700KB after master merge) The build-llms size-budget test failed: llms-full.txt is 703,244 bytes after the v0.41.37.0 key-files annotations merged on top of master's v0.41.34/35/36 CLAUDE.md additions. Matches the v0.41.9.0 precedent (600→700); the single-fetch bundle still fits comfortably in modern long-context models. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
55 lines
2.1 KiB
Markdown
55 lines
2.1 KiB
Markdown
# `gbrain serve` ↔ `gbrain sync` concurrency (PGLite)
|
|
|
|
**Short version: on a PGLite brain, stop `gbrain serve` before a large sync.**
|
|
|
|
## Why
|
|
|
|
PGLite is a single-writer embedded Postgres (WASM). A running `gbrain serve`
|
|
(stdio or HTTP MCP) holds an open PGLite connection on the brain's data
|
|
directory. `gbrain sync` needs to write to that same data directory. The two
|
|
contend for PGLite's single-writer connection / write-lock — **this is NOT the
|
|
`gbrain-sync` advisory lock** (that's a separate, DB-row coordination lock for
|
|
two concurrent *syncs*). Confusing the two sends you debugging the wrong surface.
|
|
|
|
Symptoms of serve↔sync contention on PGLite:
|
|
|
|
- `gbrain sync` blocks acquiring the PGLite write lock, or makes very slow
|
|
progress, while a `gbrain serve` process is alive on the same brain.
|
|
- Killing stale `gbrain serve` MCP processes frees the lock and sync proceeds.
|
|
|
|
## What to do
|
|
|
|
1. Stop any `gbrain serve` process for this brain before a large sync:
|
|
```bash
|
|
pkill -f 'gbrain serve' # or stop your MCP client / Claude Desktop / Cursor
|
|
gbrain sync --no-pull --no-embed --yes
|
|
```
|
|
2. Restart `gbrain serve` after the sync completes.
|
|
|
|
This contention does **not** apply to the Postgres engine — Postgres tolerates
|
|
concurrent connections, so `serve` and `sync` can run simultaneously there.
|
|
|
|
## Diagnosing a sync hang
|
|
|
|
If a sync wedges (no progress, high CPU), re-run with the per-file begin trace
|
|
so the stalling file is named:
|
|
|
|
```bash
|
|
GBRAIN_SYNC_TRACE=1 gbrain sync --no-pull --no-embed --yes
|
|
```
|
|
|
|
The last `[sync] begin import: <path>` line with no following completion is the
|
|
file being processed when the hang occurred. Under `--workers >1` / `--all`,
|
|
the stuck file is in the set of begin-lines without a matching completion.
|
|
|
|
If you suspect a schema-pack regex is the cause (a pack with a
|
|
catastrophic-backtracking `inference.regex`), complete the sync with the pack
|
|
disabled and re-run extraction afterward:
|
|
|
|
```bash
|
|
gbrain sync --no-schema-pack --no-pull --no-embed --yes
|
|
```
|
|
|
|
`gbrain schema lint` flags the classic nested-quantifier ReDoS shapes
|
|
(`(a+)+`, `(a*)*`, …) in pack regexes as warnings.
|