Files
gbrain/docs/embedding-migrations.md
T
Garry TanandClaude Opus 4.7 306fc0e1ef fix(init): error on existing-brain dim mismatch + embedding-migration recipe
Adds A4 hard-error path: when `gbrain init --embedding-dimensions N` is
run against an existing brain whose `content_chunks.embedding` column is
a different `vector(M)`, init exits 1 with an inline four-step ALTER
recipe and a pointer to docs/embedding-migrations.md.

This kills the silent-corruption pattern surfaced by issue #673: the
v0.27 schema seeded `('embedding_dimensions', '1536')` regardless of the
flag, so users got a config saying 768 but a column at 1536 — first
sync write blew up with "expected 1536, got 768."

A4's contract:
  1. Connect to engine BEFORE saveConfig so we can read the live column type
  2. If column exists AND dim != requested, exit 1 (loud failure)
  3. If column doesn't exist (fresh init) OR dim matches, proceed normally

Recipe in docs/embedding-migrations.md (and inlined in init's error
output) covers all four destructive steps codex's plan-review caught:
  1. DROP INDEX IF EXISTS idx_chunks_embedding (HNSW won't survive ALTER)
  2. ALTER TABLE content_chunks ALTER COLUMN embedding TYPE vector(N)
  3. UPDATE content_chunks SET embedding = NULL, embedded_at = NULL
  4. CREATE INDEX HNSW *only if N <= 2000* (pgvector cap)

Step 4 is conditional: dims > 2000 (e.g. Voyage 4 Large 2048d) cannot
be HNSW-indexed in pgvector; the recipe explicitly says "Skip reindex"
in that case so the user doesn't paste a CREATE INDEX that crashes.

Helper `readContentChunksEmbeddingDim` and message builder
`embeddingMismatchMessage` live in src/core/embedding-dim-check.ts so
doctor 8b (next commit) can reuse the same source of truth.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-06 18:01:18 -07:00

3.7 KiB

Switching embedding models or dimensions on an existing brain

GBrain stores embeddings in a fixed-dimension vector(N) column on content_chunks. If you switch to a model with a different dimension (e.g. text-embedding-3-large 1536 → voyage-multilingual-large-2 2048, or back to a smaller model like nomic-embed-text 768), the on-disk column type doesn't change automatically.

gbrain init and gbrain doctor both detect and refuse to silently proceed in this case. This doc is the recipe they point at.

Why we don't do this automatically

Switching dimensions requires:

  1. Dropping the HNSW vector index (pgvector won't survive an ALTER COLUMN TYPE).
  2. Altering the column type.
  3. Wiping every existing embedding (the old vectors are unusable in the new space).
  4. Re-embedding the entire corpus (can take hours on a 50K-page brain and costs $1-100 in API calls depending on model).
  5. Conditionally recreating the index (HNSW supports up to 2000 dimensions per pgvector; above that you must use exact scans).

That's not an upgrade-time auto-run. It's a deliberate, expensive operation. Run it when you've decided you actually want the new model.

Recipe — manual psql against your brain

Replace <NEW_DIMS> with your target dimension count.

BEGIN;

-- 1. Drop the HNSW index. It can't survive the column type change.
DROP INDEX IF EXISTS idx_chunks_embedding;

-- 2. Alter the column type. (You can DROP COLUMN + ADD COLUMN instead
--    if the existing data is already gone — same end state.)
ALTER TABLE content_chunks ALTER COLUMN embedding TYPE vector(<NEW_DIMS>);

-- 3. Clear stale embeddings so they don't survive into the new space.
--    Either truncate (faster, drops all chunks) or null out (preserves
--    chunk text so re-embed regenerates without re-chunking):
UPDATE content_chunks SET embedding = NULL, embedded_at = NULL;

-- 4. Recreate the HNSW index ONLY IF dims <= 2000. Above that, leave it
--    indexless and rely on exact scans (gbrain searchVector handles this
--    automatically — search just gets slower, not broken).
-- For dims <= 2000 (e.g. 1024, 1536, 768):
CREATE INDEX IF NOT EXISTS idx_chunks_embedding
  ON content_chunks USING hnsw (embedding vector_cosine_ops);
-- For dims > 2000 (e.g. 2048 Voyage 4 Large): skip step 4.

COMMIT;

Then update gbrain's config so it knows the new dim:

gbrain config set embedding_model <model>
gbrain config set embedding_dimensions <NEW_DIMS>

And re-embed the corpus:

gbrain embed --stale

PGLite (local brain)

Same recipe, but you connect to the embedded database differently:

gbrain config get database_url   # confirm engine: pglite
# Open a psql-equivalent — for PGLite, the easiest path is to write a small
# script that imports PGLiteEngine and runs the SQL via engine.executeRaw.
# Or migrate to Postgres temporarily (gbrain migrate --to supabase) if you
# want a real psql connection.

For most PGLite users the simpler path is to wipe and re-init if your corpus is small enough that re-syncing is faster than hand-crafting the migration:

mv ~/.gbrain/brain.pglite ~/.gbrain/brain.pglite.bak
gbrain init --pglite --embedding-dimensions <NEW_DIMS>
gbrain sync   # re-imports your brain repo from disk

Verify

After the recipe lands, gbrain doctor --fast should report green and gbrain doctor (full) should say check 8b passes:

✓ embedding_provider     dim parity: config 768 / column vector(768) / live probe 768

If it doesn't, file an issue with the doctor output and the SQL you ran.

v0.29+ plans

gbrain migrate-embedding-dim --to <N> is a tracked TODO. It will run the recipe above with progress reporting + an explicit confirmation gate. Until that lands, this manual recipe is the canonical path.