* fix(sync): self-heal a never-git-initialized default brain dir (#2964) The dream cycle's sync phase throws unconditionally on a legacy sync.repo_path-anchored default brain dir that was never git init-ed (predates git-backed sync, or was rsync'd without its .git), failing every nightly run with no recovery. doctor's sync_freshness/ sync_consolidation checks report "ok" for this exact brain, but only because they query the sources table (0 rows for a legacy default brain) — a coincidental false-negative, not a real diagnosis. Self-heal by git-initializing the dir and capturing the current on-disk state as the sync baseline, scoped to !opts.sourceId only — gbrain owns this directory outright, unlike a registered local source (sources add --path, no --url) which is the user's own external directory and should keep failing loudly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): dry-run no-write contract, unborn-HEAD recovery, no-gpg-sign (#2964) Codex review on the initial self-heal patch (b1671ee) found 3 real gaps: - P1: the self-heal ran even under --dry-run, mutating the filesystem during what's documented as a preview-only command. Gated the whole self-heal (both discoverGitRoot and the headCommit read) on !opts.dryRun, same as the existing !opts.sourceId ownership check. - P2: if `git init` succeeded but the process died before the baseline commit landed, the next run's discoverGitRoot would succeed (`.git` exists) and skip recovery entirely, permanently wedging on "No commits in repo" forever. Added the same self-heal at the `git rev-parse HEAD` catch site, sharing a new createSyncBaselineCommit helper with the discoverGitRoot catch. - P2: the baseline commit inherited the operator's global commit.gpgSign, which can block headless cron/launchd runs on an unavailable signing agent/pinentry. Added --no-gpg-sign. Two new tests cover dry-run no-mutation and unborn-HEAD recovery. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): restrict git auto-init self-heal to the anchor-resolved path only (#2964) Second Codex review round (a10eeab) found the ownership check still too loose: - P1 (security): !opts.sourceId alone isn't proof gbrain owns repoPath. jobs.ts's `sync` job handler leaves sourceId undefined whenever job.data.repoPath doesn't match a registered source's local_path, so an admin-scope submit_job({name:'sync', data:{repoPath}}) MCP call could point the self-heal at an arbitrary directory and have it silently git-init + commit + ingest it. Gated both self-heal sites on !opts.repoPath too — only the path resolved from gbrain's own sync.repo_path anchor (never a caller-supplied one) is eligible. - P2: the unborn-HEAD recovery site calls discoverGitRoot, which walks UP from repoPath and can resolve to an ANCESTOR repo for a --src-subpath/subdir-as-repoPath sync with an unborn HEAD. Committing there would `git add -A` sibling files well outside the sync scope. Added a check that gitContextRoot === realpathSync(repoPath) before self-healing; refuses (falls through to the original error) otherwise. Tests rewritten to exercise the true self-heal-eligible path (anchor config via engine.setConfig('sync.repo_path', dir), no repoPath/sourceId passed) instead of an explicit repoPath, plus a new test asserting a caller-supplied repoPath with no sourceId still throws. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): prove self-heal ownership by anchor VALUE, not field presence (#2964) Third Codex review round (613ad5b) found the previous round's fix broke the very call site it was meant to repair, plus a second scope gap: - P1 (critical): !opts.repoPath rejected self-heal on the REAL production callers too. runPhaseSync (dream cycle's sync phase, cycle.ts) always passes `repoPath: brainDir` explicitly after resolving it upstream, and the CLI's bare `gbrain sync` resolves sourceId='default'. Both made the ownership gate rethrow, leaving `gbrain dream` and `gbrain sync` wedged on the exact non-git legacy brain this fix targets — only synthetic callers that omitted both fields ever healed. Fixed by proving ownership by VALUE instead of by field absence: a new isAnchorOwnedSyncPath() re-reads gbrain's own persisted sync.repo_path config and requires the resolved repoPath to equal it exactly, regardless of whether the caller passed it explicitly or let it default. An attacker-supplied arbitrary path (e.g. via submit_job({name:'sync', data:{repoPath}})) only self-heals if it happens to already equal gbrain's own anchor — which is the legitimate case, not an escalation. opts.sourceId and opts.srcSubpath still disqualify unconditionally (registered/subpath-scoped syncs are a different ownership context). - P2: a --src-subpath sync with an unborn parent-repo HEAD would commit the whole ancestor root, capturing sibling files outside the scope. isAnchorOwnedSyncPath's opts.srcSubpath check closes this; the existing gitContextRoot === repoPath check stays as defense in depth. - P2: manageGitignore's "warn and return" contract (a deliberate side- effect that must never kill the sync job for its OTHER callers) meant a broken gbrain.yml or unwritable .gitignore would silently let the baseline `git add -A` commit db_only content. createSyncBaselineCommit now recomputes db_only exclusion directly from loadStorageConfig and passes it to `git add` as pathspecs, independent of the .gitignore write's success — true fail-closed. (A redundant pathspec exclude for a path .gitignore ALREADY covers makes git's -A bail with "paths ignored, use -f" even though the negation is correct, so each dir is check-ignore'd first and only pathspec-excluded when NOT already covered.) Tests rewritten around the anchor-VALUE model: the critical regression case (explicit repoPath matching the anchor still heals — the exact scenario Codex proved was broken) plus a true negative (a caller-supplied path that does NOT match the anchor still throws), --src-subpath refusal, and db_only fail-closed exclusion. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): allow default-source ownership, realpath compare, generous timeout, --no-verify (#2964) Fourth Codex review round (9c1d461) found the previous round's ownership gate still didn't match the REAL installed-brain shape, plus 3 more gaps: - P1 (critical): rejecting all non-empty opts.sourceId meant self-heal still never fired on a real brain. Migration sources_table_additive seeds a 'default' source row whose local_path mirrors sync.repo_path on every brain that's run it (virtually all of them), so resolveSourceForDir (dream cycle) and bare `gbrain sync` both resolve sourceId:'default' in practice, never undefined. isAnchorOwnedSyncPath now permits sourceId undefined OR exactly 'default' (gbrain's own bootstrap identity, never something a caller names) and proves ownership by rereading the LIVE anchor for that same identity (sources.default.local_path vs config.sync.repo_path). - P2: compared raw anchor/repoPath strings, so a cosmetic difference (trailing slash, ..) between the stored anchor and dream.ts's path.resolve()-normalized brainDir would defeat the match. Now realpath-compares both sides (fail-closed on ENOENT/dangling). - P2: the shared git() helper's 30s timeout could abort the baseline `git add -A` on a large legacy brain mid-way, after `git init` already created `.git` — leaving an unborn repo every subsequent sync would retry and time out identically forever. Added an optional timeoutMs param (default unchanged at 30s); the baseline add call uses 10min. - P2: the baseline commit could trigger an operator's global core.hooksPath/init.templateDir hooks (pre-commit/commit-msg), breaking headless recovery if those hooks need project tooling or prompt. Added --no-verify. Tests: rewrote the mis-scoped "registered source" test (it used sourceId:'default', which is now correctly permitted) into two — a new regression test proving sourceId='default' + matching local_path heals (the actual production shape), and a corrected non-default-sourceId refusal test. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): defer .gitignore write past first import, rebuild index, fail closed on unparseable db_only (#2964) Fifth Codex review round (c17cd23) found the baseline-commit helper interacting badly with the pre-existing db_only storage-tiering feature: - P1 (data loss): createSyncBaselineCommit called manageGitignore BEFORE performFullSync's collectSyncableFiles ran. collectSyncableFiles enumerates via `git ls-files --cached --others --exclude-standard`, so writing db_only entries into .gitignore first would silently exclude those pages from the DATABASE, not just from git — on a brain's very first sync. This is the exact bug class runSync's existing "manage .gitignore ONLY on successful sync" ordering (this file, ~line 4540, itself a prior Codex P1 fix, comment literally says so) was written to prevent — my new code reintroduced it in a different spot. Fix: stopped calling manageGitignore inside the self-heal at all. db_only exclusion for the COMMIT still happens via the existing pathspec computation (independent of .gitignore); .gitignore itself gets written by the already-existing post-sync flow once this sync completes, same as any other sync. - P1 (data leak): the unborn-HEAD recovery site can reach createSyncBaselineCommit with a repo whose INDEX already has entries staged from some prior operation (manual `git add`, interrupted workflow) before gbrain's self-heal ever touched it. `git add -A` only adds/updates — it doesn't drop an already-staged path our exclusion pathspecs now want excluded. Added `git read-tree --empty` to reset the index before staging (no-op on a freshly-`git init`-ed repo, whose index is already empty). - P2: loadStorageConfig warns-and-returns an EMPTY config (not a throw) for syntactically-valid-but-unsupported YAML (e.g. flow-style `db_only: [dir/]` — the narrow custom parser only handles block-style lists), which would silently resolve zero exclusions from a gbrain.yml that clearly intended some. Added a sniff-test: if gbrain.yml exists and mentions db_only but nothing resolved from it, refuse the baseline commit rather than guess "genuinely empty" vs "syntax silently ignored" (git init may already have run by this point — same "unborn, retry on next sync" recovery path handles it, and will hit this same guard again until the user fixes gbrain.yml). Tests: a positive regression proving db_only markdown IS imported into the DB on first sync (the actual data-loss scenario), the sniff-test refusal, and the stale-staged-content-gets-dropped case for the index rebuild. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): post-heal .gitignore write, supabase_only alias, honor abort signal (#2964) Sixth Codex review round (e687913), 3 P2s: - Dream-cycle callers (cycle.ts:runPhaseSync) invoke performSync directly and never run runSync's CLI-only post-success manageGitignoreAtGitRoot. A brain self-healed only via the dream cycle would have db_only content correctly excluded from the baseline commit (createSyncBaselineCommit's pathspec exclusion) but no .gitignore ever written, leaving the user's own future manual git add/commit unprotected. Added performFullSyncAndMaybeGitignore, a thin wrapper around the 3 post-self-heal performFullSync call sites that writes .gitignore (same success-status gate runSync already uses) only when didSelfHeal is true — a no-op for the normal path, which still relies on runSync exactly as before. - The fail-closed sniff-test only checked the canonical `db_only` key; the deprecated-but-still-supported `supabase_only` alias (same keep-out-of-git semantics) could silently bypass it. Now checks both. - Self-heal didn't check opts.signal?.aborted before starting the (now up to 10-minute) git init + baseline commit, so a cancelled sync could still mutate disk and overrun its budget instead of returning partial. Added the check at both self-heal sites, before any git operation runs. New test proves .gitignore gets written after a bare performSync call (no runSync wrapper) — the actual dream-cycle shape. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): neutralize a leftover .gitignore during the self-heal first sync (#2964) Seventh Codex review round (27337b0) ran an actual repro and caught the primary motivating scenario still broken: a brain rsync'd from another machine without its .git can retain that machine's old auto-managed .gitignore. collectSyncableFiles (inside performFullSync) enumerates via `git ls-files --exclude-standard`, so a leftover db_only ignore rule would silently omit those pages from THIS first sync's DATABASE import — the same bug class the round-6 ordering fix prevented for a .gitignore gbrain would have written itself, just triggered by a pre-existing file this time. Fix: performFullSyncAndMaybeGitignore now neutralizes any existing .gitignore for the duration of the one first-sync call — read, delete, restore byte-for-byte immediately after (even on error) — before manageGitignore re-merges the managed db_only block onto the restored original content. This matches exactly what a truly fresh brain with no .gitignore at all already does on its first sync (nothing to suppress collection there either); db_only content stays out of the git COMMIT independently via createSyncBaselineCommit's pathspec exclusion, which never depended on .gitignore. Test proves both halves: db_only markdown IS imported despite a leftover ignore rule, AND the user's own unrelated .gitignore lines (e.g. .DS_Store) survive the restore intact. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): simplify — drop db_only-import machinery, isolate hooks fully (#2964) Eighth Codex review round (54c1e6f) found MORE problems with round 7's .gitignore-neutralization fix (deleting the whole file loses the user's own unrelated ignore rules; a multi-sync retry scenario could silently skip a still-broken db_only file while advancing the bookmark) plus 2 more issues in existing code. Rather than patch those too, stepped back and checked the actual documented semantics of db_only (docs/storage-tiering.md): it's for "bulk machine-generated content... written to disk as a local cache" — DB is the source of truth, disk is a cache populated FROM the DB (`export --restore-only` restores it), never the other way. Nothing in the docs says `gbrain sync`'s git-diff-based file collection is how db_only content is supposed to reach the database — that's ingest-specific tooling's job. Confirmed directly: `loadStorageConfig` returns the byte-identical `{db_tracked:[], db_only:[]}` for a malformed flow-style array AND a literal empty `db_only: []`, so rounds 6-7's "ensure db_only markdown gets imported on this first sync" chase was solving a problem outside sync's actual scope in the first place, on an increasingly complex, adversarially-discovered- edge-case foundation. Reverted: performFullSyncAndMaybeGitignore (the wrapper + didSelfHeal tracking + .gitignore neutralize/restore dance + post-success manageGitignore call). After self-heal, import and any subsequent .gitignore management now behave EXACTLY like any other brain, self-healed or not — runSync's existing post-success manageGitignoreAtGitRoot covers the CLI path identically either way; the dream cycle not calling it is a separate, pre-existing characteristic of the dream cycle in general (applies equally to an already-git-initialized brain going through the same path), not something this fix introduces. Kept (still correct, self-contained, don't depend on the reverted machinery): createSyncBaselineCommit's pathspec-based db_only exclusion for the COMMIT itself (matches the documented "not committed to git" requirement), the fail-closed sniff-test guard (now documents its known, structurally-unavoidable false-positive on a genuinely-empty `db_only: []` — the trade-off is deliberate: low-cost, self-resolving false positive vs. high-cost, hard-to-undo false negative), the index rebuild, and the 600s add timeout. Improved (round 8, P2): hooks isolation. --no-verify only skips pre-commit/commit-msg; added `-c core.hooksPath=/dev/null` for the baseline commit, which disables prepare-commit-msg and post-commit too (the latter runs synchronously inside the same git invocation and could otherwise hang past the timeout without even being the slow step). Tests: removed the 3 that exercised the reverted db_only-import machinery; the remaining 12 (ownership, dry-run, index rebuild, sniff test, commit-exclusion, unborn-HEAD recovery) are unaffected by the simplification. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS * fix(sync): unconditional db_only exclusion, literal pathspecs, precise sniff test (#2964) Ninth Codex review round (a6e07f6): - P1: the check-ignore pre-filter (skip pathspec-excluding a dir already covered by .gitignore) could be defeated by a pre-existing .gitignore that ignores a db_only tree with a wildcard but re-includes a child via negation (e.g. `private-cache/*` + `!private-cache/index.md`) — check-ignore on the directory still reports "ignored", so the filter skipped the pathspec exclusion, and `git add -A` staged the re-included child anyway. Fixed by making exclusion unconditional: every db_only dir is always pathspec-excluded now, never pre-filtered against .gitignore state at all — our own pathspec doesn't consult .gitignore, so no .gitignore content (negated or not) can defeat it. The advisory "paths ignored... use -f" error this can now trigger when a dir IS also already .gitignore'd (verified: git still stages everything else correctly despite the nonzero exit) is caught and swallowed by matching its exact stderr text; anything else rethrows. - P2: `:!dir` pathspec shorthand reinterprets a dir name that itself starts with a pathspec magic character (e.g. `:private/`) instead of excluding it literally. Switched to `:(exclude,literal)dir`. - P2: the fail-closed sniff-test's bare substring search on gbrain.yml's raw content could trip on a comment or unrelated prose mentioning "db_only" even when there's no real storage section at all, refusing self-heal forever on an unrelated false positive. Now requires an actual YAML key line (`db_only:`/`supabase_only:`, trimmed, ignoring `#` comments) — the round-8-documented "genuinely empty db_only: []" false positive is unchanged and remains an accepted trade-off (still structurally indistinguishable from unsupported syntax at the loadStorageConfig API boundary), but comment/prose mentions no longer false-positive. Not fixed (deliberately, documented trade-off — see PR description): Codex's other P1 this round (refuse baselining when other git refs/ history exist alongside an unborn HEAD) is a narrow, non-destructive scenario — self-heal only ever acts on the current branch ref when it's provably commit-less, never touches or deletes any other ref (remote- tracking, other branches), so at worst it creates a possibly-unexpected extra commit on an otherwise-empty branch the user hadn't checked out yet. Chasing it further trades diminishing real-world risk reduction against unbounded scope growth in what's fundamentally still the self-heal fix from round 1. Two new tests: unconditional exclusion despite a matching pre-existing .gitignore (proves the advisory-swallow path), and a comment-only gbrain.yml no longer false-positives the sniff test. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NCVEvUVm15bFuKqskWgfNS --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
GBrain
Search gives you raw pages. GBrain gives you the answer. It's the brain layer your AI agent has been missing — the only one that does synthesis, graph traversal, and gap analysis in one box. Run a full autonomous agent on top of it, or just wire it into Claude Code or Codex as a supercharged retrieval layer in one command; either way your coding agent stops being amnesiac about everything that isn't code.
I'm Garry Tan, President and CEO of Y Combinator. I built GBrain to run my own AI agents. It's the production brain behind my OpenClaw and Hermes deployments: 146,646 pages, 24,585 people, 5,339 companies, 66 cron jobs running autonomously. My agent ingests meetings, emails, tweets, voice calls, and original ideas while I sleep. It enriches every person and company it encounters. It fixes its own citations and consolidates memory overnight. I wake up smarter than when I went to bed — and so will you.
And now it works as a company brain too. Each person on the team gets their own slice of the brain, scoped by login. When you query, you only see what you're allowed to see — never another person's notes, never another team's data. We fuzz-tested this across every way you can read the brain (search, list, lookup, multi-source reads) and got zero leaks. Drop GBrain in as your team's shared institutional memory — the company-brain shape YC just put on its Request for Startups. If you're building in that space, you might as well build on this. Tutorial: set up GBrain as your company brain →
Lots of personal-knowledge systems give you keyword matching and grep in a box. GBrain does that, and adds two things nobody else ships together:
- A synthesis layer that gives you the actual answer. Synthesized, well-cited prose across people, companies, deals, and ideas. Not "here are 10 chunks that mention your query"; an actual answer with citations and an explicit note on what the brain doesn't know yet. The gap analysis is the part that changes how you use the brain.
- A self-wiring knowledge graph. Every page write extracts entity refs and creates typed edges (
attended,works_at,invested_in,founded,advises) with zero LLM calls. Ask "who works at Acme AI?" or "what did Bob invest in this quarter?" and get answers vector search alone can't reach. Benchmarked: P@5 49.1%, R@5 97.9% on a 240-page Opus-generated rich-prose corpus, +31.4 points P@5 over its graph-disabled variant and over ripgrep-BM25 + vector-only RAG by a similar margin. Full BrainBench scorecards live in the sibling gbrain-evals repo.
The point of building a 100K-page brain is to use it as a strategic moat. To never lose context. To query what's in your own head without re-reading it. The brain layer is what makes the moat usable. The 24/7 dream cycle is what keeps it sharp. Both run on your hardware, your DB, your keys.
It's easier to ship a daemon that runs 24/7 to ingest, enrich, and consolidate than it is to keep an agent in chat working hard. GBrain is that daemon, generalized. Install in 30 minutes. Your agent does the work. As my personal agent gets smarter, so does yours.
~30 minutes to a fully working brain. Database ready in 2 seconds (PGLite, no server). You just answer questions about API keys.
LLMs: fetch
llms.txtfor the documentation map, orllms-full.txtfor the same map with core docs inlined in one fetch. Agents: start withAGENTS.md(orCLAUDE.mdif you're Claude Code).
What this looks like
Say you have a meeting with Alice tomorrow. You want to walk in remembering what she works on, when you last talked, and what's still open between you. Here's what you'd type, and what you'd get back.
You ask:
"What do I need to know before my meeting with Alice tomorrow?"
Most personal-knowledge tools give you back a list of pages. Something like:
1. people/alice — Alice runs engineering at Acme...
2. meetings/2026-03-15-alice-q1 — Q1 product review with Alice...
3. meetings/2026-01-08-acme-kickoff — Kickoff meeting with Acme team...
4. customers/acme — Acme is a series-B fintech we work with...
5. notes/2026-04-22 — Quick chat with Alice about pricing...
Five pages you now have to open and read yourself to actually prepare. The tool found the right material, but it didn't do the work.
GBrain gives you back the answer, with sources:
Alice runs engineering at Acme (a series-B fintech). You last spoke
on April 22 in a quick pricing chat. Three things are still open
from that conversation:
1. She owes you the security review for the new tier
(deadline was May 1; no update since).
2. You committed to pricing for a 500-seat tier
(you sent it April 25; no response yet).
3. She mentioned they're hiring a CISO; you said you'd intro
someone from your network.
Heads up: nothing's been added to the brain about Alice or Acme
since April 22, six weeks ago. She may have replied through email
or Slack DM, channels the brain doesn't see. Worth asking her to
catch up before assuming any of this is still current.
Every claim has a source page behind it. The "heads up" at the end tells you what the brain doesn't know yet, so you can ask Alice about it directly instead of being surprised. The brain just did your meeting prep.
This is the difference between a search engine and a brain. Search finds the pages. The brain reads them for you and writes the answer.
Install
GBrain is designed to be installed and operated by an AI agent. The fastest path is to have your agent do it for you. The CLI and MCP paths below are for people who want to wire it up themselves.
Have your agent install it (recommended)
If you don't already have an AI agent platform running, start with one of these. Both are designed to read GBrain's install protocol and execute it:
- OpenClaw — deploy AlphaClaw on Render (one click, 8GB+ RAM)
- Hermes — deploy on Railway (one click)
Then paste this into your agent:
Retrieve and follow the instructions at:
https://raw.githubusercontent.com/garrytan/gbrain/master/INSTALL_FOR_AGENTS.md
The agent installs GBrain, creates the brain, asks for your API keys, loads 43 skills, configures the dream cycle, and verifies the install end-to-end. ~30 minutes. You answer questions, it does the work.
Never set up an AI agent platform before? The personal-brain tutorial walks the whole path end-to-end — picking OpenClaw vs Hermes, deploying it, pointing it at INSTALL_FOR_AGENTS.md, getting the API keys, and verifying the first query. Start there if any of the above is new.
Quick start: Claude Code or Codex
Already running Claude Code or Codex? There are two ways to wire GBrain in, depending on what you want.
Just want a memory for your coding agent (recommended starting point). Spin up a local brain and connect it in two commands — zero server, zero token, zero tunnel:
gbrain init --pglite # 2-second local brain (no Docker)
claude mcp add gbrain -- gbrain serve # or: codex mcp add gbrain -- gbrain serve
Already have a brain on a remote host (OpenClaw, Hermes, or any gbrain serve --http)? Point your laptop agents at it with one command each — --install wires it up and smoke-tests the token before handoff:
gbrain connect https://your-host/mcp --token gbrain_xxx --install # Claude Code
gbrain connect https://your-host/mcp --token gbrain_xxx --agent codex --install # Codex
→ Full walkthrough: give your coding agent a memory — both paths end to end, plus the brain-first protocol you paste into CLAUDE.md / AGENTS.md and the four habits that make it actually change how you work.
Install the full autonomous setup into your existing agent
Want the whole thing — local brain, 43 skills, the overnight dream cycle that enriches while you sleep? Paste this into Codex, Claude Code, Cursor, or another coding agent:
Retrieve and follow the instructions at:
https://raw.githubusercontent.com/garrytan/gbrain/master/INSTALL_FOR_AGENTS.md
This works in any agent that can read files over HTTPS and execute shell commands. Tested with Codex, Claude Code, Claude Cowork, Cursor, and AlphaClaw.
CLI standalone (no agent)
bun install -g github:garrytan/gbrain
gbrain init --pglite # 2 seconds; no server, no Docker
gbrain doctor # verify health
gbrain import ~/notes/ # index your markdown
gbrain query "what themes show up across my notes?"
Postgres-at-scale, Supabase, and thin-client setup paths live in docs/INSTALL.md.
Connect GBrain to your AI client (MCP)
GBrain exposes 30+ tools over MCP (stdio and HTTP). The specific snippet depends on which client you use:
- Claude Code — local: one command,
claude mcp add gbrain -- gbrain serve(zero server, zero tunnel). Remote with just a bearer token:gbrain connect https://your-host/mcp --token gbrain_xxxprints a paste-ready block (or--installwires it up and smoke-tests the token). - Codex —
gbrain connect https://your-host/mcp --token gbrain_xxx --agent codex(or--install). Codex reads the bearer from$GBRAIN_REMOTE_TOKENat runtime, so the token never lands in Codex config. - Cursor / Windsurf / any stdio MCP client — same shape, add
{"command": "gbrain", "args": ["serve"]}to your MCP config. - Claude Desktop (Cowork) — Settings → Integrations → add the URL of your HTTP server. Remote only; the local
claude_desktop_config.jsondoes not work for remote servers. - Claude Cowork (team plan) — org Owner adds the connector under Organization Settings → Connectors.
- Perplexity Computer —
gbrain connect https://your-host/mcp --agent perplexity --oauth --registermints a least-privilege OAuth client and prints the Issuer/Client ID/Secret to paste into Settings → Connectors (OAuth is the right path for a cloud connector; a bearer token also works for local use). Pro subscription required. - ChatGPT — uses OAuth 2.1 with PKCE (the hard requirement). Register a
chatgptclient from the admin dashboard with grant typeauthorization_code.
For the HTTP server itself:
gbrain serve # stdio MCP (local subprocess; for Claude Code, Cursor, Windsurf)
gbrain serve --http # HTTP MCP with OAuth 2.1 + admin dashboard at /admin
# (required for Claude Desktop, Cowork, Perplexity, ChatGPT)
The HTTP server includes DCR-style client registration, scope-gated access (read / write / admin), and rate limiting. Deployment guides (ngrok, Railway, Fly.io) live under docs/mcp/.
Two ways to query your brain
Raw retrieval (what most personal-knowledge tools ship) and a synthesis layer that gives you an actual answer. They serve different jobs.
# raw retrieval: top pages by hybrid score, fast, no LLM cost
gbrain search "who's working on AI agents at portfolio companies?"
# brain layer: synthesized answer with citations and gap analysis
gbrain think "who's working on AI agents at portfolio companies?"
gbrain search returns the top retrieved pages, ranked by hybrid scoring (vector + keyword + RRF + source-tier boost + reranker). Use it when you want raw material to skim: agent context windows, citation lookups, finding a specific quote.
gbrain think runs the same retrieval, then composes a synthesized answer across the results with explicit citations to the source pages AND an honest note on what the brain doesn't know yet. The gap analysis is the differentiator: the answer tells you when a page is stale, when a claim is uncited, when two pages contradict each other, when there's a hole you should fill.
Why it compounds. Pair the brain layer with find_trajectory and you get answers like "how have the company's metrics changed AND what does the team look like right now AND what did they promise / share AND when did we last meet AND what's the value-add I can offer here": well-scored, well-cited, in one shot. That's the strategic moat. That's why building a 100K-page brain is worth the effort.
gbrain agent run "..." exposes the same surface to a sub-agent through the Minions queue, with crash-safe two-phase persistence. Same answers, durable.
How to get data in
One command, local or hosted, synchronous receipt:
gbrain capture "the thought I want to remember"
gbrain capture --file ./notes/today.md
echo "from a pipe" | gbrain capture --stdin
SLUG=$(gbrain capture "..." --quiet)
The page lands in the database and on disk in one move. Default slug inbox/YYYY-MM-DD-<hash8> so captures cluster in a predictable triage location. On thin-client installs the verb routes through MCP to the server: same command, same UX.
For webhook ingestion (Zapier / IFTTT / Apple Shortcuts):
curl -X POST https://your-brain/ingest \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: text/markdown" \
-d "# a thought from a Shortcut"
For mobile capture, the inbox folder source picks up anything dropped into
~/.gbrain/inbox/ from iOS Shortcuts / AirDrop / Drafts / Finder.
Third-party skillpacks can ship custom ingestion sources (Granola, Linear,
voice, OCR) against the versioned IngestionSource contract at
gbrain/ingestion. See docs/skillpack-anatomy.md.
Your brain's shape (schema packs)
Most personal-knowledge tools force one fixed layout: their idea of "notes" + "people" + "tags." Drop a Notion export or your own years-old Obsidian vault on top, and the agent doesn't know what a Projects/ folder means or whether Reading/ is people or sources.
gbrain doesn't have a fixed layout. It ships with bundled schema packs and lets you author your own when none fit:
gbrain-base-v2(default as of v0.41.22) — 15-type DRY/MECE canonical taxonomy (14 canonical +notecatch-all):person,company,media,tweet,social-digest,analysis,atom,concept,source,deal,email,slack,writing,project,note. Subtypes/format/origin pushed to frontmatter. The taxonomy that responds to issue #1479.gbrain-base(legacy, v0.41 and earlier brains) — the original 24-type layout. Stays bundled for back-compat; brains on it can upgrade viagbrain onboard --check --explain→gbrain jobs submit unify-types --allow-protected --params '{"target_pack":"gbrain-base-v2"}'.gbrain-recommended— extendsgbrain-basewith the 13 additional directories fromdocs/GBRAIN_RECOMMENDED_SCHEMA.md(source, place, trip, conversation, personal, civic, project, etc.). Activate withgbrain schema use gbrain-recommended.- Your own pack —
gbrain schema detectclusters your actual filesystem into proposed types,gbrain schema suggestruns an LLM pass over them, andgbrain schema review-candidates --applypromotes the ones you like. Three commands and the brain knows your shape. Authoring a successor pack (declaresmigration_from:so existing brains can opt in): seedocs/architecture/pack-upgrade-mechanism.md.
gbrain schema active # which pack is running, which tier set it
gbrain schema list # bundled + installed packs
gbrain schema detect # propose types matching your filesystem
gbrain schema suggest # LLM-refined proposals on top of detect
gbrain schema review-candidates # human gate: promote / rename / ignore
gbrain schema use my-pack # activate
The active pack threads through every read + write path: parseMarkdown infers page type from the pack's path prefixes; whoknows scopes expert routing to types declared expert_routing: true; extract_facts runs only on extractable: true types; the search cache folds the pack name + version into its key so cross-pack contamination is structurally impossible. Switch packs and the brain re-interprets itself; switch back and nothing's lost.
Seven-tier resolution chain (per-call flag → env var → per-source DB key → brain-wide DB key → gbrain.yml → ~/.gbrain/config.json → gbrain-base default). Full reference + authoring guide: docs/architecture/schema-packs.md.
Tutorials
Step-by-step walkthroughs for getting the most out of GBrain. Each one takes you from zero to a working outcome, with concrete commands and real numbers.
- Set up your personal AI agent + brain from zero — the canonical full-stack install. Two GitHub repos, a Telegram bot, AlphaClaw on Render, OpenClaw + GBrain + Supabase. End-to-end in about 2 hours.
- Set up GBrain as your company brain — federated, multi-user, OAuth-scoped institutional memory for a 10-50 person team. About 90 minutes end-to-end.
- Auto-improve a skill with
gbrain skillopt— treat aSKILL.mdas a trainable parameter. Generate a starter benchmark straight from the skill with--bootstrap-from-skill(or write your own), strengthen the judges, then watch the optimizer propose edits and keep only the ones that measurably score higher. ~20 minutes, ~$1 in API calls. Flag + cost + safety reference:docs/guides/skillopt.md.
More walkthroughs in progress: connecting an existing agent (Claude Code, Cursor, OpenClaw, Hermes) to a GBrain memory layer; setting up GBrain for VC dealflow with founder scorecards and meeting prep; migrating an existing Notion or Obsidian vault; indexing a codebase as a queryable code brain. Full tutorial index: docs/tutorials/.
Want to see a tutorial that isn't here yet? Open an issue describing the workflow you want documented.
What it does (the loop)
signal → search → respond → write → auto-link → sync
(every (brain-first (informed (page + (typed edges (cron
message) retrieval) by context) timeline) + backlinks) keeps fresh)
- Signal detector runs on every message your agent receives. Captures ideas, entity mentions, time-sensitive todos, names, links.
- Brain-first lookup before any external API call. The cheapest, fastest, most personal information source you have.
- Auto-link fires on every page write. No LLM calls; pure pattern matching on
[[wiki/people/bob]]style references. New entity → new page stub → graph grows. - Cron-driven enrichment runs while you sleep: dedup people pages, fix citations, score salience, find contradictions, prep tomorrow's tasks.
The whole loop is described in docs/architecture/topologies.md with diagrams.
Capabilities
Hybrid search. Vector (HNSW on pgvector) + BM25 keyword + reciprocal-rank fusion + source-tier boost + intent-aware query rewriting. Three named search modes (conservative, balanced, tokenmax) bundle the cost/quality knobs into a single config key. Live cost/recall comparisons in docs/eval/SEARCH_MODE_METHODOLOGY.md. Default: balanced with ZeroEntropy reranker on. Per-query graph signals notice when a top result is a hub for THAT query (adjacency boost), is corroborated across team brains (cross-source boost), or is being crowded out by weak chunks from a chatty session (session demote). Run gbrain search "<query>" --explain to see per-stage attribution: base score, every boost that fired, what it multiplied. gbrain doctor ships a graph_signals_coverage check; gbrain search stats shows fire counts and failure breakdowns. Vector retrieval pools the best chunk per page, so a page surfaces on its strongest evidence instead of losing to a neighbor on one weak chunk. Queries that match a page's title phrase or a declared free-text alias (gbrain reindex --aliases backfills existing pages) get boosted to the page they name. Every result carries an evidence tag (why it matched) and a create_safety hint (exists / probable / unknown) so an agent decides whether a page already exists instead of guessing from a raw score. gbrain search diagnose "<query>" --target <slug> traces which retrieval layer surfaces (or misses) a page.
Self-wiring knowledge graph. Every put_page extracts entity refs from markdown/wikilinks/typed-link syntax and writes edges with zero LLM calls. Typed edges (attended, works_at, invested_in, founded, advises, mentions, …). Multi-hop traversal via gbrain graph-query. The graph is what produces the +31.4 P@5 lift over vector-only RAG. Obsidian-style vaults: bare [[note-name]] wikilinks that point across folders — you wrote [[struktura]] but the page lives at projects/struktura.md — resolve by basename once you opt in with gbrain config set link_resolution.global_basename true. Off by default; gbrain doctor tells you how many edges you'd gain before you flip it. See migrating an Obsidian vault.
Job queue (Minions). BullMQ-shaped, Postgres-native job queue. Durable subagents (LLM tool loops that survive crashes via two-phase pending→done persistence), shell jobs with audit, child jobs with cascading timeouts, rate leases for outbound providers, attachments via S3/Supabase storage. Replaces "spawn subagent as fire-and-forget Promise" with something that recovers from anything.
Non-English brains (FTS language config). The Postgres full-text search tokenizer is configurable via GBRAIN_FTS_LANGUAGE. Defaults to english. Set it to any text-search configuration that exists in your Postgres instance:
export GBRAIN_FTS_LANGUAGE=portuguese # uses built-in portuguese stemmer
export GBRAIN_FTS_LANGUAGE=spanish # built-in spanish stemmer
export GBRAIN_FTS_LANGUAGE=pt_br # custom config (e.g. unaccent + portuguese)
List available configs: psql -c "SELECT cfgname FROM pg_ts_config". Both the query side (websearch_to_tsquery) and the write side (the trigger functions that populate pages.search_vector and content_chunks.search_vector) honor GBRAIN_FTS_LANGUAGE. On first install (or upgrade), the configurable_fts_language schema migration reads the env var and creates trigger functions in the configured language; subsequent inserts/updates tokenize using that setting. To change language on a brain that has already run the migration, use the dedicated CLI command:
export GBRAIN_FTS_LANGUAGE=portuguese
gbrain reindex-search-vector --dry-run # preview row counts
gbrain reindex-search-vector --yes # recreate triggers + backfill
The command is idempotent (re-running with the same language is a no-op for vector content) and uses the same recreate-and-backfill primitives as the migration. For accent-insensitive Portuguese (pt_br), see docs/guides/multi-language-fts.md for the unaccent + portuguese stemmer recipe.
43 curated skills. Routing lives in skills/RESOLVER.md. Covers signal capture, ingest (idea / media / meeting), enrichment, querying, brain ops, citation fixing, daily task management, cron scheduling, reports, voice, soul audit, skill creation, eval framework, and migrations. Skills are markdown files (tool-agnostic), packaged as a single skillpack the installer drops into your agent workspace.
Eval framework. gbrain eval longmemeval runs the public LongMemEval benchmark against your hybrid retrieval. gbrain eval export + gbrain eval replay capture real queries and replay them against code changes (set GBRAIN_CONTRIBUTOR_MODE=1). gbrain eval cross-modal cross-checks an output against the task using three different-provider frontier models. gbrain eval retrieval-quality runs NamedThingBench, which hard-gates the named-thing retrieval families (title-substring, alias-synonym, generic-to-named, multi-chunk-dilution) so a regression in "find the page this query names" fails CI loudly. Full methodology in docs/eval/SEARCH_MODE_METHODOLOGY.md.
Brain consistency. gbrain eval suspected-contradictions samples retrieval pairs, layered date pre-filter, query-conditioned LLM judge, persistent cache. Surfaces conflicts between takes + facts the agent has written. Wired into the daily dream cycle.
Agent-authored schema (v0.40.7.0). Your brain has a shape — what page types exist (person, meeting, paper, case, lab-result), what they link to (attended, authored, prescribed-by), what facts get extracted automatically. The default ships with 22 universal types, but your brain's actual shape is not the default shape. Agents can now evolve that shape on your behalf via 14 gbrain schema CLI verbs + a batched MCP op (schema_apply_mutations, admin scope, NOT localOnly so remote agents reach it over HTTPS). Atomic file locks, audit log with the agent's identity, chunked UPDATE backfill in 1000-row batches that never wedge concurrent writers. The brain stops being a pile of notes and becomes something with structure. Why it matters: docs/what-schemas-unlock.md — 7 killer use cases (4000 invisible meetings, founder ops brain, research brain, legal brain, team brain, agent-as-co-curator). 5-minute walkthrough: docs/schema-author-tutorial.md. Agent skill: skills/schema-author/SKILL.md.
Integrations
Data flowing into the brain. Each integration is a recipe — markdown + setup hints — that ships in recipes/ and is discoverable via gbrain integrations list.
- Voice: Phone calls create brain pages via Twilio + OpenAI Realtime (or DIY STT+LLM+TTS). Setup recipe:
recipes/twilio-voice-brain.md. - Email + calendar: webhook handlers that route to brain signals.
docs/integrations/meeting-webhooks.md. - Embedding providers: 16 recipes covering OpenAI (default fallback), OpenRouter, Voyage, ZeroEntropy (default), Google Gemini, Azure OpenAI, MiniMax, Alibaba DashScope, Zhipu, Ollama (local), llama.cpp llama-server (local), LiteLLM proxy. Pricing matrix + decision tree in
docs/integrations/embedding-providers.md. - Rerankers: ZeroEntropy
zerank-2hosted (default intokenmaxmode) plus the v0.40.6.1llama-server-rerankerrecipe for fully-local cross-encoder rerank via llama.cpp — runs Qwen3-Reranker or self-hosted ZeroEntropy weights against the samegateway.rerank()seam. Setup walkthrough indocs/ai-providers/llama-server-reranker.md. - Credential gateway: vault-aware secret distribution.
docs/integrations/credential-gateway.md. - MCP clients: every major MCP client is supported.
docs/mcp/per-client setup.
Architecture
Two engines, one contract. PGLite (Postgres 17 via WASM, zero-config, default) for personal brains up to ~50K pages. Postgres + pgvector (Supabase or self-hosted) for shared / large / multi-machine deployments. The contract-first BrainEngine interface in src/core/engine.ts defines ~47 operations both engines implement; CLI and MCP server are generated from one source.
Brain repo is the system of record. Your knowledge lives in a regular git repo (your "brain repo") as markdown files. GBrain syncs the repo into Postgres for retrieval; deletes in git become soft-deletes in DB. You can publish public subsets, share team mounts, run thin-client setups pointing at a colleague's brain server. Topologies in docs/architecture/topologies.md.
Two organizational axes (brain ⊥ source). A brain is a database (your personal brain, a team mount you joined). A source is a repo inside that brain (wiki, gstack, an essay, a knowledge base). Routing lives in .gbrain-source dotfiles and resolves via a documented 6-tier precedence chain. Full diagrams in docs/architecture/brains-and-sources.md.
Why the graph matters. Vector search returns chunks that are semantically close. The graph returns chunks that are factually connected. Hybrid search pulls from both; auto-linking on every write keeps the graph fresh. Deep dive: docs/architecture/RETRIEVAL.md.
Troubleshooting
gbrain init --pglite crashes on macOS 26.x (Tahoe)? PGLite's embedded WASM engine is incompatible with macOS 26.x on Apple Silicon. The fix is to use native Homebrew PostgreSQL + pgvector instead. Full step-by-step setup in docs/INSTALL.md — Troubleshooting: PGLite crashes on macOS 26.x.
gbrain import fails with expected N dimensions, not M? Run gbrain doctor. It will print the exact gbrain config set ... or gbrain retrieval-upgrade command to repair the mismatch. You should not need to delete ~/.gbrain. Fresh gbrain init --pglite auto-detects your embedding provider from API keys in your environment: set OPENAI_API_KEY (or ZEROENTROPY_API_KEY / VOYAGE_API_KEY) before running init, or pass --embedding-model <provider>:<model> explicitly. With multiple keys set, init fires an interactive picker. In non-TTY contexts (CI, Docker) with no keys, init exits 1 with a paste-ready setup hint; pass --no-embedding to defer setup until runtime. See docs/integrations/embedding-providers.md for the full provider matrix and docs/operations/headless-install.md for Docker/CI sequencing.
Hourly cron sync keeps timing out on a federated brain? v0.41.13.0 ships
two flags + a recommended pattern. Switch your cron to a per-source loop
with shell timeout(1) doing the OS-level kill and gbrain self-terminating
gracefully half-a-minute earlier:
gbrain sync --break-lock --all --max-age 1800
for src in $(gbrain sources list --json | jq -r '.[].id'); do
timeout 600 gbrain sync --source "$src" --timeout 540 || true
done
When --timeout fires mid-import, gbrain sync exits 0 with status
partial and last_commit UNCHANGED — the next run re-walks the same
diff and content_hash short-circuits already-imported files. The
--max-age 1800 first command self-heals any wedged-but-alive locks
left by a hung previous run, using the v98 last_refreshed_at semantic
(NOT acquired_at) so healthy long-running holders are safe by
construction. See the v0.41.13.0 entry in CHANGELOG.md
for the honest scope notes (extract + embed phases run to completion;
30-min rollout window for --max-age post-migration v98; full-sync
triggers deferred to v0.42+).
Dream cycle silently losing wiki links on Supabase? v0.41.19.0 fixes
the bug class structurally. The engine now self-retries every bulk batch
write (addLinksBatch / addTimelineEntriesBatch / upsertChunks) on
Supavisor pooler blips, with a 12s worst-case wait that covers the full
5-10s circuit-breaker recovery window. gbrain doctor surfaces incidents
via the new batch_retry_health check (reads the last 24h of
~/.gbrain/audit/batch-retry-YYYY-Www.jsonl). To tune for an unusually
slow pooler:
# Defaults: 3 retries, base 1s, max 10s, decorrelated jitter.
# Override per operator without a release:
export GBRAIN_BULK_MAX_RETRIES=5 # int >= 0; 0 disables retries
export GBRAIN_BULK_RETRY_BASE_MS=2000 # int > 0
export GBRAIN_BULK_RETRY_MAX_MS=15000 # int >= base
Bad values surface at gbrain doctor startup with a paste-ready fix
(not at first-retry mid-cycle). PGLite-only installs pay zero cost — the
retry wrap is engine-level, but PGLite has no pooler so retries never
fire in practice.
Dream cycle losing ~150 link rows per run with 'No database connection: connect() has not been called' errors in the log? v0.41.27.0
makes the retry layer self-heal on a nulled-out database singleton. A
new reconnect callback on withRetry rebuilds the connection between
attempts; PostgresEngine.batchRetry injects () => this.reconnect()
so engine-level batch writes survive a mid-cycle disconnect by something
else in the same process. Same release: gbrain capture no longer trails
a 'No database connection' stderr line from a background facts:absorb
worker firing after CLI exit — the op-dispatch finally block awaits
getFactsQueue().drainPending({timeout: 1000}) before
engine.disconnect(). To find which code path is still calling
disconnect mid-process, run gbrain doctor --json | jq '.checks[] | select(.id=="batch_retry_health")'; the extended check now surfaces
24h disconnect-call count and the most-recent caller frame from a new
~/.gbrain/audit/db-disconnect-YYYY-Www.jsonl audit. (Closes #1570.)
gbrain brainstorm returning judge_failed: true with 0 scored
ideas? v0.41.21.0 closes the two bugs that caused it. The judge
hard-coded a 4K-token output cap; for any run past ~40 ideas the call
truncated mid-JSON and the parser threw. Same release closes a slash-
form pricing miss: gbrain brainstorm --judge-model anthropic/claude-sonnet-4-6 --max-cost 5 failed with
BudgetExhausted reason=no_pricing because every pricing site only
matched the colon form. Both shapes work now. No config change, no
schema migration — gbrain upgrade is the whole fix.
gbrain reindex --markdown wiped your auto/dream/signal-detector
tags? v0.41.37.0 makes tag reconciliation add-only. Re-import and
reindex --markdown now ADD current frontmatter tags and never delete,
so enrichment tags written to the DB (auto-tag, dream synthesize,
signal-detector) survive a re-chunk. The reindex DB-only fallback also
reconstructs the full markdown (frontmatter + body + timeline) before
re-chunking, so a page with no on-disk source keeps its frontmatter,
title, and timeline instead of getting overwritten with empty
frontmatter. Trade-off: removing a tag from a page's frontmatter no
longer removes it from the DB on the next sync (frontmatter-tag removal
needs a provenance column, deferred). (Closes #1621.)
gbrain sync wedges on a large brain (no progress, high CPU)?
v0.41.37.0 ships three things. First, name the stalling file:
GBRAIN_SYNC_TRACE=1 gbrain sync --no-pull --no-embed --yes
The last [sync] begin import: <path> line with no following completion
is the file being processed when the hang hit. Second, if you suspect a
schema-pack inference.regex with catastrophic backtracking, complete
the sync with the pack disabled and re-run extraction later:
gbrain sync --no-schema-pack --no-pull --no-embed --yes
gbrain schema lint now warns on the classic nested-quantifier ReDoS
shapes ((a+)+, (a*)*, …) in pack regexes, and the runtime caps
inference-regex input length (override via GBRAIN_MAX_REGEX_INPUT_CHARS).
Third, on a PGLite brain, stop gbrain serve before a large sync —
PGLite is single-writer and a live MCP server contends for the write
lock. See docs/architecture/serve-sync-concurrency.md
for the full triage. (Closes #1569.)
gbrain init --migrate-only / a schema migration fails on Windows
with getaddrinfo ENOTFOUND? v0.41.37.0 runs the 9 schema-bring-up
phases in-process instead of spawning a child gbrain init --migrate-only per phase. The spawned child died on
Windows + bun + Supabase pooler with a DNS-resolution failure even
though the parent connected fine; running in-process removes the spawn
entirely. The v0.13.1 grandfather migration that hung 70+ minutes on an
82K-page PGLite brain is also fixed — it now runs as a chunked bulk SQL
pass (keyed on the page PK, soft-delete-filtered, source-safe) that
completes in ~1-2 seconds. (Closes #1605, #1581.)
Docs
docs/INSTALL.md— every install path, end to enddocs/what-schemas-unlock.md— why schemas matter: 7 killer use cases, the structural argument for typed page kinds, the agent-co-curates pattern (v0.40.7.0)docs/schema-author-tutorial.md— 5-minute walkthrough: fork the bundled pack, add a custom type, backfill existing pages, prove the wiring viagbrain whoknowsdocs/architecture/— system design, topologies, retrieval theorydocs/guides/— how-to runbooks (sub-agent routing, minion deployment, skill development, brain-first lookup, idea capture, diligence ingestion)docs/integrations/— connecting external data sources (voice, email, calendar, embedding providers)docs/mcp/— per-client MCP setup (Claude Desktop, Code, Cursor, ChatGPT, Perplexity, Cowork)docs/eval/— eval framework, metric glossary, methodologydocs/ethos/— philosophy (thin harness, fat skills, markdown as recipes, origin story)AGENTS.md— entry point for non-Claude agentsCLAUDE.md— entry point for Claude Code (deep operating context)CONTRIBUTING.md— contributor guide, test discipline, eval-capture modeSECURITY.md— OAuth threat model, hardening defaults
Contributing
Run bun run test for the fast loop, bun run verify for the pre-push gate, bun run ci:local to run the full Docker-backed CI stack locally. Detailed test discipline in CONTRIBUTING.md.
Community PRs are batched into release waves rather than merged one-by-one — see the "PR wave workflow" section in CLAUDE.md. Contributor attribution stays attached via Co-Authored-By: trailers. We credit every accepted contribution in CHANGELOG.md.
If you find a bug or want a feature: open an issue first. Quick fixes (typo, doc bug, obvious regression) can go straight to a PR. Anything touching schema, retrieval ranking, MCP protocol, or the security boundary needs a design discussion in the issue first.
License + credit
MIT. I built GBrain to run my OpenClaw and Hermes deployments — the production brain behind my AI agents.
Origin story: docs/ethos/ORIGIN.md.
Community PR contributors are credited in CHANGELOG.md per release. ZeroEntropy (@zeroentropy) for the embedding + reranker stack that ships as the default. Voyage AI for the asymmetric-encoding recipe template. Ramp Labs for the search quality improvements lineage.