mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-30 11:22:34 +00:00
A community member reported docs 'have quite a bit of drift and some broken
links' and contradictions like 'says don't use bun but also to use bun.' This
PR is a top-to-bottom audit + fix across every doc file at the repo root and
under docs/. Where docs disagreed with each other, the code was the tie-breaker.
## Categories of fix
### 1. Stale CLI commands (skillpack install → scaffold)
`gbrain skillpack install` was retired in v0.36.0.0 (replaced by the
scaffold/reference/migrate-fence model). The CLI now errors out with a hint:
$ gbrain skillpack install
Error: 'gbrain skillpack install' was removed in v0.33.
Use 'gbrain skillpack scaffold <name>' instead.
But the docs still recommended it:
- README.md line 29 — primary install path
- docs/INSTALL.md lines 12 — primary install path
Both updated to `gbrain skillpack scaffold --all` with the v0.36.0.0 retirement
explained inline + the migrate-fence escape hatch for users upgrading from older
releases.
### 2. The 'bun install -g vs bun link' contradiction
The community member's exact complaint. The drift:
- README.md + docs/INSTALL.md: recommended `bun install -g github:garrytan/gbrain`
- INSTALL_FOR_AGENTS.md line 29: 'Do NOT use `bun install -g github:garrytan/gbrain`.'
Reading the code + CHANGELOG: `bun install -g` IS the canonical path. Bun
occasionally blocks the top-level postinstall hook on global installs (issue #218),
but the postinstall now prints a loud recovery hint when that happens, and
`gbrain doctor` flags `schema_version: 0` and routes users to
`gbrain apply-migrations --yes`. The 'do not use' warning was correct in 2024
when the postinstall silently swallowed errors with `|| true`; it's stale now.
Reconciled:
- INSTALL_FOR_AGENTS.md Step 1: now recommends `bun install -g` as the primary
path, documents #218 as a known issue with the recovery command, and keeps
`git clone + bun link` as a documented fallback.
- AGENTS.md Install (5 min): same reconciliation; clone path is the fallback,
not the default.
- docs/INSTALL.md CLI standalone: added the #218 callout so the deterministic
fallback is one click away when the default fails.
### 3. Broken internal links
- README.md → `docs/integrations/voice.md` (file doesn't exist). The real voice
recipe lives at `recipes/twilio-voice-brain.md` (Twilio + OpenAI Realtime).
Fixed to point there with an accurate one-line summary.
- CONTRIBUTING.md → `docs/SQLITE_ENGINE.md` (file doesn't exist; superseded by
PGLite per docs/ENGINES.md). Replaced with a paragraph explaining the
supersession and pointing at the live ENGINES.md.
- docs/GBRAIN_V0.md → `docs/SQLITE_ENGINE.md` (2 references; same supersession).
Added a historical-doc banner at the top + rewrote both references to point at
the current ENGINES.md.
### 4. Stale API key recommendations
INSTALL_FOR_AGENTS.md Step 2 only mentioned OpenAI + Anthropic. As of v0.36.2.0
ZeroEntropy is the default embedding + reranker stack (README opens with this);
the agent install guide didn't reflect it. Added `ZEROENTROPY_API_KEY` as the
default, kept OpenAI/Voyage as documented fallbacks, noted that keys can live in
`~/.gbrain/config.json` (file plane) or env.
### 5. Stale upgrade workflow
INSTALL_FOR_AGENTS.md 'Upgrade' section assumed the clone+bun-install model
(`cd ~/gbrain && git pull && bun install && gbrain init && gbrain post-upgrade`)
and didn't mention `gbrain upgrade` (the single-command path that exists in the
CLI today: binary self-update + schema migrations + post-upgrade prompts in one).
Split into two paths — `gbrain upgrade` for the bun-install-g case (now the
default per Step 1), clone-path for the fallback case.
Also fixed AGENTS.md 'Migrate' bullet (was `gbrain apply-migrations` only;
now leads with `gbrain upgrade` and keeps apply-migrations as the manual
schema-only path).
### 6. Stale cron-workflow
INSTALL_FOR_AGENTS.md Step 7 referenced cron docs but didn't mention
`gbrain autopilot --install` (the built-in self-maintaining daemon that
exists in the CLI today) or `gbrain sync --watch` (continuous loop). Added
both as alternatives to platform-cron glue.
### 7. ZeroEntropy version typo
docs/INSTALL.md said 'the v0.36.0.0 ZE switch' — ZE landed in v0.36.2.0
(v0.36.0.0 was the skillpack-scaffold retirement). Fixed.
## What I did NOT change
- CHANGELOG.md, CLAUDE.md, TODOS.md prose mentions of historical commands like
`gbrain skillpack install` are correct as history — they're documenting what
was true in past releases. Only forward-looking docs got updated.
- The 'broken link' false-positive matches in CHANGELOG / CLAUDE / TODOS are
inside code-fence examples or regex patterns (`[Name](people/slug)`,
`[a-z0-9](?:[a-z0-9-]{0,30}[a-z0-9])`, `[--json](interrupted)`); they're
illustrative syntax, not real links. Leaving alone.
- llms.txt / llms-full.txt regenerated via `bun run build:llms` so the
agent-fetch documentation map matches the new content.
## Verification
- `bun run src/cli.ts --help` cross-checked against every command/flag the
install docs reference: init, doctor, apply-migrations, upgrade, post-upgrade,
skillpack scaffold/reference/migrate-fence, embed --stale, sync --watch,
autopilot --install, dream, integrations list, extract links/timeline,
graph-query, query, search modes — all real, all current.
- `bun run src/cli.ts skillpack install` confirmed to error out with the
retirement hint pointing at scaffold (proves the README guidance was actively
misleading users into a dead-end).
- Re-ran the broken-internal-link scanner across all root .md + docs/**/*.md;
zero real broken links remain (5 residual matches are illustrative syntax
inside prose, not actionable links).
Co-authored-by: garrytan-agents <agents@garrytan-agents.local>
290 lines
12 KiB
Markdown
290 lines
12 KiB
Markdown
# Contributing to GBrain
|
|
|
|
## Setup
|
|
|
|
```bash
|
|
git clone https://github.com/garrytan/gbrain.git
|
|
cd gbrain
|
|
bun install
|
|
bun test
|
|
```
|
|
|
|
Requires Bun 1.0+.
|
|
|
|
## Project structure
|
|
|
|
```
|
|
src/
|
|
cli.ts CLI entry point
|
|
commands/ CLI-only commands (init, upgrade, import, export, etc.)
|
|
core/
|
|
operations.ts Contract-first operation definitions (the foundation)
|
|
engine.ts BrainEngine interface
|
|
postgres-engine.ts Postgres implementation
|
|
db.ts Connection management + schema loader
|
|
import-file.ts Import pipeline (chunk + embed + tags)
|
|
types.ts TypeScript types
|
|
markdown.ts Frontmatter parsing
|
|
config.ts Config file management
|
|
storage.ts Pluggable storage interface
|
|
storage/ Storage backends (S3, Supabase, local)
|
|
supabase-admin.ts Supabase admin API
|
|
file-resolver.ts MIME detection + content hashing
|
|
migrate.ts Migration helpers
|
|
yaml-lite.ts Lightweight YAML parser
|
|
chunkers/ 3-tier chunking (recursive, semantic, llm)
|
|
search/ Hybrid search (vector, keyword, hybrid, expansion, dedup)
|
|
embedding.ts OpenAI embedding service
|
|
mcp/
|
|
server.ts MCP stdio server (generated from operations)
|
|
schema.sql Postgres DDL
|
|
skills/ Fat markdown skills for AI agents
|
|
test/ Unit tests (bun test, no DB required)
|
|
test/e2e/ E2E tests (requires DATABASE_URL, real Postgres+pgvector)
|
|
fixtures/ Miniature realistic brain corpus (16 files)
|
|
helpers.ts DB lifecycle, fixture import, timing
|
|
mechanical.test.ts All operations against real DB
|
|
mcp.test.ts MCP tool generation verification
|
|
skills.test.ts Tier 2 skill tests (requires OpenClaw + API keys)
|
|
docs/ Architecture docs
|
|
```
|
|
|
|
## Running tests
|
|
|
|
```bash
|
|
# Inner edit loop (~85s on a Mac dev box, 3700+ unit tests)
|
|
bun run test # parallel 8-shard fan-out + serial post-pass
|
|
bun test test/markdown.test.ts # specific unit test
|
|
|
|
# Pre-push gate (matches what CI runs on shard 1 + typecheck)
|
|
bun run verify # privacy + jsonb + progress + test-isolation + wasm + admin-build + typecheck
|
|
|
|
# Pre-merge sanity (everything CI runs)
|
|
bun run test:full # verify + parallel unit + slow + smart e2e
|
|
|
|
# Slow / serial / e2e in isolation
|
|
bun run test:slow # *.slow.test.ts only (cold-path correctness)
|
|
bun run test:serial # *.serial.test.ts only (--max-concurrency=1)
|
|
bun run test:e2e # real-Postgres E2E (requires DATABASE_URL)
|
|
|
|
# E2E setup (Postgres with pgvector)
|
|
docker compose -f docker-compose.test.yml up -d
|
|
DATABASE_URL=postgresql://postgres:postgres@localhost:5434/gbrain_test bun run test:e2e
|
|
|
|
# Or use your own Postgres / Supabase
|
|
DATABASE_URL=postgresql://... bun run test:e2e
|
|
```
|
|
|
|
Use `bun run verify` before pushing. The guard chain catches: banned fork-name
|
|
leaks (`scripts/check-privacy.sh`), `JSON.stringify(x)::jsonb` interpolation
|
|
patterns (`scripts/check-jsonb-pattern.sh`), `\r` progress bleed to stdout
|
|
(`scripts/check-progress-to-stdout.sh`), test-isolation rule violations
|
|
(`scripts/check-test-isolation.sh` — see "Writing tests that survive the parallel
|
|
loop" below), silent fallback to recursive chunking in the compiled binary
|
|
(`scripts/check-wasm-embedded.sh`), and stale admin-dashboard build artifacts
|
|
(`scripts/check-admin-build.sh`). `bun run check:all` runs the full historical
|
|
sweep including the trailing-newline and exports-count checks.
|
|
|
|
### Writing tests that survive the parallel loop
|
|
|
|
`bun run test` shards 92+ unit-test files across 8 worker processes. Files in the
|
|
same shard share a process, so process-global state leaks between them. Four
|
|
lint rules (`scripts/check-test-isolation.sh`, R1-R4) enforce isolation:
|
|
|
|
| Rule | What it bans | Fix |
|
|
|---|---|---|
|
|
| **R1** | Direct `process.env.X = ...` mutation | Use `withEnv()` from `test/helpers/with-env.ts`, or rename to `*.serial.test.ts` |
|
|
| **R2** | `mock.module(...)` anywhere in the file | Rename to `*.serial.test.ts` |
|
|
| **R3** | `new PGLiteEngine(` outside ~50 lines after `beforeAll(` | Use the canonical PGLite block (see below) |
|
|
| **R4** | `new PGLiteEngine(` without paired `afterAll(disconnect)` | Add the `afterAll(() => engine.disconnect())` |
|
|
|
|
Canonical PGLite block (R3 + R4 compliant — paste this verbatim):
|
|
|
|
```ts
|
|
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
|
|
import { resetPgliteState } from './helpers/reset-pglite.ts';
|
|
|
|
let engine: PGLiteEngine;
|
|
|
|
beforeAll(async () => {
|
|
engine = new PGLiteEngine();
|
|
await engine.connect({});
|
|
await engine.initSchema();
|
|
});
|
|
afterAll(async () => { await engine.disconnect(); });
|
|
beforeEach(async () => { await resetPgliteState(engine); });
|
|
```
|
|
|
|
Env-touching tests:
|
|
|
|
```ts
|
|
import { withEnv } from './helpers/with-env.ts';
|
|
|
|
test('reads OPENAI_API_KEY', async () => {
|
|
await withEnv({ OPENAI_API_KEY: 'sk-test' }, async () => {
|
|
expect(loadConfig().openai_key).toBe('sk-test');
|
|
});
|
|
});
|
|
```
|
|
|
|
`withEnv` saves and restores keys via try/finally including when the callback
|
|
throws. Cross-test safe; **NOT** intra-file concurrent-safe (`process.env` is
|
|
process-global). Files using `withEnv` stay outside the future
|
|
`test.concurrent()` codemod's eligibility filter.
|
|
|
|
When to quarantine instead of fix: rename to `*.serial.test.ts` if the file
|
|
uses `mock.module(...)`, is genuinely env-coupled (module-load env readers +
|
|
ESM caching defeat dynamic-import-after-env tricks), or intentionally shares
|
|
state across `it()` boundaries. Quarantine count cap: 10 (informational).
|
|
|
|
Files that violated these rules at the v0.26.7 baseline are listed in
|
|
`scripts/check-test-isolation.allowlist`. **The allow-list MUST shrink over
|
|
time** ... never add new entries. v0.26.8 (env sweep) and v0.26.9 (PGLite sweep
|
|
+ codemod) remove entries as files get fixed.
|
|
|
|
### Local CI gate (recommended before pushing, v0.23.1+)
|
|
|
|
```bash
|
|
bun run ci:local # full gate: gitleaks + unit + ALL 29 E2E files (sequential)
|
|
bun run ci:local:diff # gate with diff-aware E2E selector
|
|
bun run ci:select-e2e # print which E2E files the selector would run
|
|
```
|
|
|
|
`ci:local` spins up `pgvector/pgvector:pg16` + `oven/bun:1` via
|
|
`docker-compose.ci.yml`, runs everything PR CI runs plus the full E2E suite, then
|
|
tears down. Named volumes keep the install warm across runs (~16-20 min sequential
|
|
E2E after the first cold pull). Requires Docker (Docker Desktop, OrbStack, or
|
|
Colima) and `gitleaks` on host (`brew install gitleaks`). Override the postgres
|
|
host port with `GBRAIN_CI_PG_PORT=5435 bun run ci:local` if 5434 collides.
|
|
|
|
Fail-closed selector: an unmapped `src/` change runs all 29 E2E files. Hand-tune
|
|
narrower mappings via `scripts/e2e-test-map.ts`.
|
|
|
|
## Building
|
|
|
|
```bash
|
|
bun build --compile --outfile bin/gbrain src/cli.ts
|
|
```
|
|
|
|
## Adding a new operation
|
|
|
|
GBrain uses a contract-first architecture. Add your operation to one file and it
|
|
automatically appears in the CLI, MCP server, and tools-json:
|
|
|
|
1. Add your operation to `src/core/operations.ts` (define params, handler, cliHints)
|
|
2. Add tests
|
|
3. That's it. The CLI, MCP server, and tools-json are generated from operations.
|
|
|
|
For CLI-only commands (init, upgrade, import, export, files, embed, doctor, sync):
|
|
1. Create `src/commands/mycommand.ts`
|
|
2. Add the case to `src/cli.ts`
|
|
|
|
Parity tests (`test/parity.test.ts`) verify CLI/MCP/tools-json stay in sync.
|
|
|
|
## Adding a new engine
|
|
|
|
See `docs/ENGINES.md` for the full guide. In short:
|
|
|
|
1. Create `src/core/myengine-engine.ts` implementing `BrainEngine`
|
|
2. Add to engine factory in `src/core/engine.ts`
|
|
3. Run the test suite against your engine
|
|
4. Document in `docs/`
|
|
|
|
The original SQLite engine plan was superseded by PGLite (embedded Postgres 17 via WASM), which uses the same SQL dialect as Postgres and eliminates the need for a separate FTS5/sqlite-vss translation layer. See [`docs/ENGINES.md`](docs/ENGINES.md) for the engine architecture and the rationale.
|
|
|
|
## CONTRIBUTOR_MODE — turn on the dev loop
|
|
|
|
gbrain captures retrieval traffic so you can replay real queries against
|
|
your code changes before merging. **This is off by default** (production
|
|
users get a quiet brain, no surprise data accumulation). Contributors turn
|
|
it on with one shell rc line:
|
|
|
|
```bash
|
|
# In ~/.zshrc or ~/.bashrc:
|
|
export GBRAIN_CONTRIBUTOR_MODE=1
|
|
```
|
|
|
|
That's it. Every `query` / `search` you (or agents pointed at your dev
|
|
brain) run from that shell now writes a row to `eval_candidates`, and the
|
|
[replay tool](#running-real-world-eval-benchmarks-touching-retrieval-code)
|
|
has data to work against.
|
|
|
|
What CONTRIBUTOR_MODE actually does:
|
|
|
|
- Turns on `query`/`search` capture into the local `eval_candidates` table.
|
|
Without it the gate is closed and capture is a no-op.
|
|
- That's all. PII scrubbing, retention, and replay are independent.
|
|
|
|
Resolution order (most explicit wins):
|
|
|
|
1. `eval.capture: true` in `~/.gbrain/config.json` → on
|
|
2. `eval.capture: false` in `~/.gbrain/config.json` → off
|
|
3. `GBRAIN_CONTRIBUTOR_MODE=1` → on
|
|
4. otherwise → off
|
|
|
|
Quick check that capture is actually running:
|
|
|
|
```bash
|
|
gbrain query "anything" >/dev/null
|
|
psql $DATABASE_URL -c 'SELECT count(*) FROM eval_candidates'
|
|
# (or `gbrain doctor` — surfaces silent capture failures cross-process)
|
|
```
|
|
|
|
To disable capture even with the env var set, write
|
|
`{"eval": {"capture": false}}` to `~/.gbrain/config.json` — explicit config
|
|
beats the env var both directions.
|
|
|
|
## Running real-world eval benchmarks (touching retrieval code)
|
|
|
|
If your PR touches retrieval — search ranking, RRF fusion, embeddings,
|
|
intent classification, query expansion, source boost, or the `query` /
|
|
`search` op handlers — run `gbrain eval replay` against a snapshot of
|
|
real traffic before merging. Requires `CONTRIBUTOR_MODE` (above) so you
|
|
have captured rows to replay against.
|
|
|
|
Quick loop:
|
|
|
|
```bash
|
|
gbrain eval export --since 7d > baseline.ndjson # snapshot before your change
|
|
# ... make your change ...
|
|
gbrain eval replay --against baseline.ndjson # diff retrieval, get Jaccard@k
|
|
```
|
|
|
|
Three numbers come back: mean Jaccard@k between captured and current slug
|
|
sets, top-1 stability, and mean latency Δ. The replay tool flags the worst
|
|
regressions so you can eyeball whether the change is hurting real queries.
|
|
|
|
Trigger paths (rerun if your diff touches any of these):
|
|
|
|
- `src/core/search/hybrid.ts`
|
|
- `src/core/search/source-boost.ts`, `sql-ranking.ts`
|
|
- `src/core/search/intent.ts`, `expansion.ts`, `dedup.ts`
|
|
- `src/core/embedding.ts`
|
|
- `src/core/operations.ts` (query / search handlers)
|
|
- `src/core/postgres-engine.ts` / `pglite-engine.ts` (searchKeyword /
|
|
searchVector SQL)
|
|
|
|
See [`docs/eval-bench.md`](./docs/eval-bench.md) for the full guide
|
|
including CI integration, hand-crafted NDJSON corpora (so a fresh checkout
|
|
without captured data can still replay), and cost considerations. The
|
|
NDJSON wire format is documented in
|
|
[`docs/eval-capture.md`](./docs/eval-capture.md).
|
|
|
|
For public benchmark coverage on top of replay, `gbrain eval longmemeval
|
|
<dataset.jsonl>` (v0.28.1) runs LongMemEval against gbrain's hybrid
|
|
retrieval. One in-memory PGLite per question, runtime-enumerated
|
|
`TRUNCATE` between questions, ground-truth scoring via LongMemEval's
|
|
published `evaluate_qa.py`. Use it alongside replay when changes affect
|
|
retrieval quality on long-context conversational data — replay catches
|
|
regressions on YOUR queries, LongMemEval catches them on a public set the
|
|
benchmark community already cites. See the "Public benchmarks: LongMemEval"
|
|
section in [`docs/eval-bench.md`](./docs/eval-bench.md).
|
|
|
|
## Welcome PRs
|
|
|
|
- SQLite engine implementation
|
|
- Docker Compose for self-hosted Postgres
|
|
- Additional migration sources
|
|
- New enrichment API integrations
|
|
- Performance optimizations
|