mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-31 04:07:52 +00:00
CLAUDE.md had grown to 592KB / ~147k tokens auto-loaded every session (~77% of
the llms-full.txt single-fetch bundle). The per-file index was append-only by
mandate. This is the exact thin-dispatcher-vs-fat-blob anti-pattern gbrain exists
to fix, so CLAUDE.md becomes a thin orientation + resolver that points at
on-demand docs.
This commit is the VERBATIM move (content-preserving — the next commit compresses):
- docs/architecture/KEY_FILES.md <- ## Key files + the calibration key-files
cluster + Schema Cathedral v3 impl detail
- docs/architecture/thin-client.md <- ## Thin-client routing
- docs/TESTING.md <- ## Testing
- ## Commands DROPPED (18 'added in vX.Y' history blocks; current surface is
gbrain 0.41.38.0 -- personal knowledge brain
USAGE
gbrain <command> [options]
SETUP
init [--pglite|--supabase|--url] Create brain (PGLite default, no server)
migrate --to <supabase|pglite> Transfer brain between engines
upgrade Self-update
check-update [--json] Check for new versions
doctor [--json] [--fast] Health check (resolver, skills, pgvector, RLS, embeddings)
integrations [subcommand] Manage integration recipes (senses + reflexes)
PAGES
get <slug> Read a page
put <slug> [< file.md] Write/update a page
delete <slug> Delete a page
list [--type T] [--tag T] [-n N] List pages
SEARCH
search <query> Keyword search (tsvector)
query <question> [--no-expand] Hybrid search (RRF + expansion)
ask <question> [--no-expand] Alias for query
IMPORT/EXPORT
import <dir> [--no-embed] Import markdown directory
sync [--repo <path>] [flags] Git-to-brain incremental sync
sync --watch [--interval N] Continuous sync (loops until stopped)
sync --install-cron Install persistent sync daemon
export [--dir ./out/] Export to markdown
export --restore-only [--repo <p>] Restore missing supabase-only files
[--type T] [--slug-prefix S] With optional filters
FILES
files list [slug] List stored files
files upload <file> --page <slug> Upload file to storage
files upload-raw <file> --page <s> Smart upload (size routing + .redirect.yaml)
files signed-url <path> Generate signed URL (1-hour)
files sync <dir> Bulk upload directory
files verify Verify all uploads
EMBEDDINGS
embed [<slug>|--all|--stale] Generate/refresh embeddings
LINKS
link <from> <to> [--type T] Create typed link
unlink <from> <to> Remove link
backlinks <slug> Incoming links
graph <slug> [--depth N] Traverse link graph (returns nodes)
graph-query <slug> [--type T] Edge-based traversal with type/direction filters
[--depth N] [--direction in|out|both]
TAGS
tags <slug> List tags
tag <slug> <tag> Add tag
untag <slug> <tag> Remove tag
TIMELINE
timeline [<slug>] View timeline
timeline-add <slug> <date> <text> Add timeline entry
TOOLS
extract <links|timeline|all> Extract links/timeline (idempotent)
[--source fs|db] fs (default) walks .md files; db iterates engine pages
[--dir <brain>] brain dir for fs source
[--type T] [--since DATE] filters (db source)
[--dry-run] [--json]
publish <page.md> [--password] Shareable HTML (strips private data, optional AES-256)
check-backlinks <check|fix> [dir] Find/fix missing back-links across brain
lint <dir|file> [--fix] Catch LLM artifacts, placeholder dates, bad frontmatter
orphans [--json] [--count] Find pages with no inbound wikilinks
salience [--days N] [--kind P] v0.29: pages ranked by emotional + activity salience
anomalies [--since D] [--sigma N] v0.29: cohort-based statistical anomalies (tag, type)
transcripts recent [--days N] v0.29: recent raw .txt transcripts (local-only)
dream [--dry-run] [--json] Run the overnight maintenance cycle once (cron-friendly).
See also: autopilot --install (continuous daemon).
check-resolvable [--json] [--fix] Validate skill tree (reachability/MECE/DRY)
report --type <name> --content ... Save timestamped report to brain/reports/
BRAIN (capture / ideate / explore — v0.37/v0.38)
capture [content] [--file PATH] Single entrypoint for getting content into the brain
[--stdin] [--slug s] [--type t] Inline content / file / stdin; writes to inbox/ by default
[--source ID] [--quiet|--json] Multi-source brains: route to a non-default source
brainstorm <question> [--json] Bisociation idea generator (hybrid search + far-set + judge)
[--save|--no-save] [--limit N]
lsd <question> [--json] Lateral Synaptic Drift: inverted-judge brainstorm
[--save|--no-save] [--limit N] rewarding far-from-obvious + axiomatic inversions
SOURCES (multi-repo / multi-brain)
sources list Show registered sources
sources add <id> --path <p> Register a source (id = short name, e.g. 'wiki')
sources remove <id> Remove a source + its pages
sync --all Sync all sources with a local_path
sync --source <id> Sync one specific source
repos ... DEPRECATED alias for 'sources' (v0.19.0)
CODE INDEXING (v0.19.0 / v0.20.0 Cathedral II)
code-def <symbol> [--lang l] Find the definition of a symbol across code pages
code-refs <symbol> [--lang l] Find all references to a symbol (JSON-first)
code-callers <symbol> Who calls this symbol? (v0.20.0 A1)
code-callees <symbol> What does this symbol call? (v0.20.0 A1)
query <q> --lang <l> Filter hybrid search to one language (v0.20.0)
query <q> --symbol-kind <k> Filter to symbol type (function|class|method|...) (v0.20.0)
reconcile-links [--dry-run] Batch-recompute doc↔impl edges (v0.20.0)
reindex-code [--source id] [--yes] Explicit code-page reindex (v0.20.0)
sync --strategy code Sync code files into the brain
JOBS (Minions)
jobs submit <name> [--params JSON] Submit background job [--follow] [--dry-run]
jobs list [--status S] [--limit N] List jobs
jobs get <id> Job details + history
jobs cancel <id> Cancel job
jobs retry <id> Re-queue failed/dead job
jobs prune [--older-than 30d] Clean old jobs
jobs stats Job health dashboard
jobs work [--queue Q] Start worker daemon (Postgres only)
ADMIN
stats Brain statistics
health Brain health dashboard
history <slug> Page version history
revert <slug> <version-id> Revert to version
features [--json] [--auto-fix] Scan usage + recommend unused features
autopilot [--repo] [--interval N] Self-maintaining brain daemon
config [show|get|set] <key> [val] Brain config
storage status [--repo <path>] Storage tier status and health
[--json] (git-tracked vs supabase-only)
serve MCP server (stdio)
serve --http [--port N] HTTP MCP server with OAuth 2.1
--token-ttl N Access token TTL in seconds (default: 3600)
--enable-dcr Enable Dynamic Client Registration
--public-url URL Public issuer URL (required behind proxy/tunnel)
call <tool> '<json>' Raw tool invocation
version Version info
--tools-json Tool discovery (JSON)
Run gbrain <command> --help for command-specific help. + the per-command KEY_FILES entries; content stays in git)
CLAUDE.md gains: a Reference map (resolver), a Maintaining section (the
anti-disease rule), and a Cross-cutting invariants subsection under Architecture
so the must-never-violate rules (trust fail-closed, sourceScopeOpts isolation,
JSONB trap, engine parity, contract-first, migrations, multi-source) still
auto-load after the index moved out.
Result: CLAUDE.md 592KB -> 61KB; llms-full.txt 740KB -> 210KB (new docs link-only
until compressed). build-llms drift + budget test green; verify 29/29 green.
The pre-move content is recoverable at git show <this^>:CLAUDE.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
302 lines
11 KiB
TypeScript
302 lines
11 KiB
TypeScript
/**
|
|
* llms-config — single source of truth for llms.txt + llms-full.txt.
|
|
*
|
|
* Consumed by scripts/build-llms.ts (emits llms.txt, llms-full.txt) and
|
|
* test/build-llms.test.ts (asserts paths resolve, content contract holds).
|
|
*
|
|
* Adding a doc? Add it here and run `bun run build:llms`. The drift-detection
|
|
* test fails CI if you forget.
|
|
*
|
|
* Fork-friendliness: `rawBaseUrl` reads from `LLMS_REPO_BASE` so forks can
|
|
* regenerate without manual URL rewrites:
|
|
* LLMS_REPO_BASE=https://raw.githubusercontent.com/fork-org/gbrain/main bun run build:llms
|
|
*/
|
|
|
|
export type DocEntry = {
|
|
title: string;
|
|
description: string;
|
|
path: string;
|
|
includeInFull?: boolean;
|
|
};
|
|
|
|
export type DocSection = {
|
|
heading: string;
|
|
optional?: boolean;
|
|
entries: DocEntry[];
|
|
};
|
|
|
|
export const PROJECT = {
|
|
name: "GBrain",
|
|
summary:
|
|
"GBrain is a personal knowledge brain and GStack mod for agent platforms. Pluggable engines (PGLite default, Postgres+pgvector for scale), contract-first operations, 26 fat-markdown skills. Teaches agents brain ops, ingestion, enrichment, scheduling, identity, and access control.",
|
|
repoUrl: "https://github.com/garrytan/gbrain",
|
|
rawBaseUrl:
|
|
process.env.LLMS_REPO_BASE ??
|
|
"https://raw.githubusercontent.com/garrytan/gbrain/master",
|
|
};
|
|
|
|
export const SECTIONS: DocSection[] = [
|
|
{
|
|
heading: "Core entry points",
|
|
entries: [
|
|
{
|
|
title: "AGENTS.md",
|
|
description:
|
|
"Start here if you are not Claude Code. Install order, trust boundary, skill resolver, config/debug/migration pointers.",
|
|
path: "AGENTS.md",
|
|
},
|
|
{
|
|
title: "CLAUDE.md",
|
|
description:
|
|
"Orientation + resolver. North Star, two axes, architecture + cross-cutting invariants, the reference map pointing at on-demand docs, and the inline ship IRON RULES.",
|
|
path: "CLAUDE.md",
|
|
},
|
|
{
|
|
title: "docs/architecture/KEY_FILES.md",
|
|
description:
|
|
"Per-file index for the gbrain repo: what each src/ file does + its load-bearing invariants. The on-demand detail CLAUDE.md's reference map routes to.",
|
|
path: "docs/architecture/KEY_FILES.md",
|
|
// Link-only until compressed to current-state (still large pre-compression).
|
|
// Flip to inlined once the doc-history compression lands and the bundle
|
|
// budget is re-measured.
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
title: "docs/architecture/thin-client.md",
|
|
description:
|
|
"The thin-client / remote-MCP / cross-modal routing seam: isThinClient detection, callRemoteTool, SSRF-hardened URL validation, per-command routing.",
|
|
path: "docs/architecture/thin-client.md",
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
title: "INSTALL_FOR_AGENTS.md",
|
|
description: "9-step agent installation.",
|
|
path: "INSTALL_FOR_AGENTS.md",
|
|
},
|
|
{
|
|
title: "skills/RESOLVER.md",
|
|
description: "Skill dispatcher. Read first for any task.",
|
|
path: "skills/RESOLVER.md",
|
|
},
|
|
{
|
|
title: "README.md",
|
|
description: "Project overview, benchmarks, 30-minute setup.",
|
|
path: "README.md",
|
|
},
|
|
],
|
|
},
|
|
{
|
|
heading: "Configuration",
|
|
entries: [
|
|
{
|
|
title: "docs/ENGINES.md",
|
|
description: "PGLite vs Postgres trade-off and when to migrate.",
|
|
path: "docs/ENGINES.md",
|
|
},
|
|
{
|
|
title: "docs/GBRAIN_RECOMMENDED_SCHEMA.md",
|
|
description:
|
|
"MECE directory structure (people/, companies/, concepts/).",
|
|
path: "docs/GBRAIN_RECOMMENDED_SCHEMA.md",
|
|
// v0.40.6.0: 64KB reference doc. Web index entry stays; the single-fetch
|
|
// bundle gets the README + setup guides instead. Keeps llms-full.txt
|
|
// under the 600KB budget as CLAUDE.md grows with each release.
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
// Excluded from llms-full.txt (stays linked in llms.txt) to keep the
|
|
// single-fetch bundle under FULL_SIZE_BUDGET as CLAUDE.md grows. This is
|
|
// a value-explainer, not load-bearing operational reference for an agent
|
|
// reading the bundle — the right thing to link rather than inline.
|
|
includeInFull: false,
|
|
title: "docs/what-schemas-unlock.md",
|
|
description:
|
|
"Why schemas matter: 7 killer use cases (4000 invisible meetings, founder ops brain, research brain, legal brain, team brain, agent-as-co-curator) + the structural argument for typed page kinds. Read this before pitching schema authoring (v0.40.7.0).",
|
|
path: "docs/what-schemas-unlock.md",
|
|
},
|
|
{
|
|
title: "docs/schema-author-tutorial.md",
|
|
description:
|
|
"5-minute walkthrough: fork the bundled pack, add a custom `researcher` type, backfill existing pages via `gbrain schema sync --apply`, prove the T1.5 wiring via `gbrain whoknows` (v0.40.7.0).",
|
|
path: "docs/schema-author-tutorial.md",
|
|
},
|
|
{
|
|
title: "docs/guides/live-sync.md",
|
|
description: "Incremental markdown sync setup.",
|
|
path: "docs/guides/live-sync.md",
|
|
},
|
|
{
|
|
title: "docs/guides/cron-schedule.md",
|
|
description: "Recurring job scheduling.",
|
|
path: "docs/guides/cron-schedule.md",
|
|
},
|
|
{
|
|
title: "docs/guides/minions-deployment.md",
|
|
description:
|
|
"Deploying the gbrain jobs worker: crontab + watchdog, inline --follow, systemd/Procfile/fly.toml, upgrade checklist.",
|
|
path: "docs/guides/minions-deployment.md",
|
|
// v0.41.8.0: 13KB deployment runbook. Web index entry stays;
|
|
// single-fetch bundle drops it to keep under FULL_SIZE_BUDGET
|
|
// (CLAUDE.md grew past 600KB once master's v0.41.2-v0.41.6 +
|
|
// this wave's annotations landed). Operators read this once;
|
|
// agents rarely need it in context.
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
title: "docs/guides/quiet-hours.md",
|
|
description: "Notification hold + timezone-aware delivery.",
|
|
path: "docs/guides/quiet-hours.md",
|
|
},
|
|
{
|
|
title: "docs/guides/scaling-skills.md",
|
|
description:
|
|
"Three-tier architecture for agents with 300+ skills: always-loaded, resolver-routed, and dormant. Per-turn token math, the v0.41.7.0 compact list-format resolver, and the `gbrain doctor` safety net. 306 skills, ~21K tokens freed per turn, zero capability loss.",
|
|
path: "docs/guides/scaling-skills.md",
|
|
},
|
|
{
|
|
title: "docs/mcp/DEPLOY.md",
|
|
description: "MCP server deployment.",
|
|
path: "docs/mcp/DEPLOY.md",
|
|
},
|
|
],
|
|
},
|
|
{
|
|
heading: "AI providers",
|
|
entries: [
|
|
{
|
|
title: "docs/ai-providers/zeroentropy.md",
|
|
description:
|
|
"ZeroEntropy zembed-1 embedding + zerank-2 reranker (hosted): API key, embedding switch, reranker config.",
|
|
path: "docs/ai-providers/zeroentropy.md",
|
|
// Setup walkthrough — discoverable in the index, not inlined in the
|
|
// single-fetch bundle (keeps llms-full.txt under FULL_SIZE_BUDGET).
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
title: "docs/ai-providers/llama-server-reranker.md",
|
|
description:
|
|
"Local reranker via llama.cpp --reranking: Qwen3-Reranker or self-hosted ZE weights, --alias setup, gbrain config keys, cold-start timeout, budget-cap interaction.",
|
|
path: "docs/ai-providers/llama-server-reranker.md",
|
|
includeInFull: false,
|
|
},
|
|
],
|
|
},
|
|
{
|
|
heading: "Debugging",
|
|
entries: [
|
|
{
|
|
title: "docs/GBRAIN_VERIFY.md",
|
|
description:
|
|
"7-check post-setup verification. Start here when something feels off.",
|
|
path: "docs/GBRAIN_VERIFY.md",
|
|
},
|
|
{
|
|
title: "docs/guides/minions-fix.md",
|
|
description: "Troubleshooting the Minions job queue.",
|
|
path: "docs/guides/minions-fix.md",
|
|
},
|
|
{
|
|
title: "docs/integrations/reliability-repair.md",
|
|
description: "Data integrity recovery.",
|
|
path: "docs/integrations/reliability-repair.md",
|
|
},
|
|
],
|
|
},
|
|
{
|
|
heading: "Migrations",
|
|
entries: [
|
|
{
|
|
title: "docs/UPGRADING_DOWNSTREAM_AGENTS.md",
|
|
description:
|
|
"Patches for downstream agent skill forks. One section per release.",
|
|
path: "docs/UPGRADING_DOWNSTREAM_AGENTS.md",
|
|
// Excluded from inlined bundle (v0.41.7.0): 25KB of release-by-release
|
|
// migration patches that are valuable as a reference but don't need
|
|
// to ride along in every llms-full.txt fetch. Pushes the bundle back
|
|
// under FULL_SIZE_BUDGET after the v0.41.7.0 scaling-skills guide
|
|
// landed.
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
title: "skills/migrations/",
|
|
description:
|
|
"Per-version (v0.5.0 - v0.14.1) agent-executable migration instructions.",
|
|
path: "skills/migrations/",
|
|
},
|
|
{
|
|
title: "CHANGELOG.md",
|
|
description:
|
|
"Release-summary voice + itemized changes + self-repair block per version.",
|
|
path: "CHANGELOG.md",
|
|
includeInFull: false,
|
|
},
|
|
],
|
|
},
|
|
{
|
|
heading: "Contributing",
|
|
optional: true,
|
|
entries: [
|
|
{
|
|
title: "docs/TESTING.md",
|
|
description:
|
|
"Test command tiers, the test-isolation lint (R1-R4), the canonical PGLite block, withEnv, the E2E DB lifecycle, and the file taxonomy. Maintainer-facing.",
|
|
path: "docs/TESTING.md",
|
|
includeInFull: false,
|
|
},
|
|
],
|
|
},
|
|
{
|
|
heading: "Philosophy",
|
|
optional: true,
|
|
entries: [
|
|
{
|
|
title: "docs/ethos/THIN_HARNESS_FAT_SKILLS.md",
|
|
description: "Why skills live in markdown.",
|
|
path: "docs/ethos/THIN_HARNESS_FAT_SKILLS.md",
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
title: "docs/ethos/MARKDOWN_SKILLS_AS_RECIPES.md",
|
|
description: "Homebrew for Personal AI.",
|
|
path: "docs/ethos/MARKDOWN_SKILLS_AS_RECIPES.md",
|
|
includeInFull: false,
|
|
},
|
|
],
|
|
},
|
|
{
|
|
heading: "Optional",
|
|
optional: true,
|
|
entries: [
|
|
{
|
|
title: "docs/designs/",
|
|
description: "Forward-looking designs.",
|
|
path: "docs/designs/",
|
|
includeInFull: false,
|
|
},
|
|
{
|
|
title: "docs/architecture/infra-layer.md",
|
|
description: "Shared infra patterns.",
|
|
path: "docs/architecture/infra-layer.md",
|
|
includeInFull: false,
|
|
},
|
|
],
|
|
},
|
|
];
|
|
|
|
export const INLINE_TIPS = [
|
|
"`gbrain doctor [--json] [--fast] [--fix]` - built-in health checks.",
|
|
"`gbrain orphans [--json]` - pages with zero inbound wikilinks.",
|
|
"`gbrain repair-jsonb [--dry-run]` - repair v0.12.0 double-encoded JSONB rows.",
|
|
"`gbrain upgrade` runs post-upgrade + apply-migrations.",
|
|
];
|
|
|
|
// Target ~750KB so llms-full.txt fits in ~190k-token contexts with room to spare.
|
|
// Bumped 600KB→700KB in v0.41.9.0, then 700KB→750KB once CLAUDE.md crossed 700KB:
|
|
// it's ~540KB (77% of the bundle) and grows ~5-15KB per release with each feature's
|
|
// Key Files annotation. Both master (v0.41.34-38 waves) and this branch (skillopt
|
|
// wave) independently hit the 700KB line and bumped to the same 750KB. CLAUDE.md is
|
|
// the whole point of the one-fetch bundle, so it stays inlined; the budget tracks
|
|
// its legitimate growth. Still fits comfortably in 200k+ context models.
|
|
// Generator prints a WARN if exceeded; ship with includeInFull=false exclusions.
|
|
export const FULL_SIZE_BUDGET = 750_000;
|