mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
82cc7fff90429ef05585705b3d91a74afeffb4ce
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
805814451e |
v0.42.26.0 docs(supabase): update connection-string setup to new UI + Transaction pooler (#1848) (#1875)
* docs(supabase): update connection-string setup to new UI + Transaction pooler Supabase moved the connection string under "Connect" in the top nav and now shows three options (Direct, Transaction pooler, Session pooler). Update the tutorial, gbrain init prompts, the setup skill, the verify runbook, and the live-sync guide to recommend the Transaction pooler (port 6543) — which gbrain is tuned for (prepared statements disabled, DDL/locks routed to a derived direct connection). Document the IPv4 footgun: the derived direct connection is IPv6-only, so on IPv4-only hosts reads work but sync silently skips pages. Tutorial 7c now leads with the free fix (GBRAIN_DIRECT_DATABASE_URL -> Session pooler, port 5432) and keeps the IPv4 add-on as the paid alternative. Removes stale "transaction mode breaks sync (.begin() is not a function)" warnings and the port-6543 "Session pooler" mislabels. Extends PR #1848 by @FilipHarald. * chore: bump version and changelog (v0.42.26.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d4211f4176 |
v0.42.11.0 feat(skillopt): held-out eval gate, honest receipts, ENFORCE + ablation opts (#1759)
* feat(skillopt): wire held-out gate, honest receipts, ENFORCE + ablation opts Wire the F11 held-out gate into the orchestrator at checkpoint acceptance (runHeldOutGate was dead code); parse + thread --held-out through CLI, batch, fleet, background job, and the run_skillopt MCP op. Populate the real receipt.baseline_sel_score (was hardcoded 0) and add a final-test eval (test_score + baseline_test_score) via a shared scoreSkillOnTasks primitive. Fix the --no-mutate proposed.md write (was a stub) and enforce maxRuntimeMin. D16 ENFORCE in core mutation policy (assertBundledMutationHeldOut): mutating a bundled skill in place requires a non-empty (>=5), benchmark-disjoint held-out set or hard-refuses. Add three eval-internal ablation opts (reflectMode, disableValidationGate, optimizerMode='one-shot-rewrite') recorded in the receipt + audit; ROLLOUT_SUCCESS_THRESHOLD named constant. Security: run_skillopt MCP op validates skill_name (kebab-only) and confines caller-supplied benchmark/held-out paths to the skills dir for remote callers. * test(skillopt): held-out gate, ENFORCE, one-shot rewrite, runtime + receipt honesty New test/skillopt/rollout.test.ts (rollout had zero coverage). Held-out ENFORCE unit cases + one-shot-rewrite fence handling (whole-response unwrap, embedded-fence preserved, error path). E2E: F11 held-out BLOCKS/ALLOWS, bundled no-mutate write, reflectMode/disableValidationGate/optimizerMode, maxRuntimeMin abort, receipt baseline/test-score honesty, held-out/benchmark disjointness, D2 no-DB-pollution. * chore: bump version and changelog (v0.42.9.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: document skillopt held-out gate + bundled mutation requirement for v0.42.9.0 Wire --held-out into the skill-optimizer SKILL.md, guide flags/safety tables, and the tutorial's bundled-skill step: mutating a bundled skill in place now requires --allow-mutate-bundled AND --held-out (>=5 benchmark-disjoint tasks) or it hard-refuses. Add the --held-out flag row + F11 held-out gate to the guide; update the receipt contract to the honest baseline/test-score fields. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(gateway): AI SDK v6 toolLoop compat — multi-turn tool calls work again The ai@6.x bump tightened ModelMessage + tool-schema validation, which silently broke every multi-turn tool loop. Both `gbrain skillopt` rollouts and production background `subagent` jobs route through `chat()`/`toolLoop` and crashed the moment the model called a tool ("messages do not match the ModelMessage[] schema" / "schema is not a function"). Surfaced end-to-end by the SkillOpt real-LLM eval. Three fixes: - chat(): wrap tool defs with the SDK's `jsonSchema()` helper instead of a bare `{jsonSchema}` object (v6 asSchema() treated the bare object as a thunk and threw). - chat(): new exported pure `toModelMessages()` converts gbrain's provider-neutral ChatMessage[] into v6 ModelMessage[] — tool results ride a dedicated `role:'tool'` message with structured `{type,value}` output; null output preserved as json null. Load-bearing for the production subagent path, not just skillopt. - rollout.ts: replace the inline params→schema mapper (dropped `items` on array params) with the shared `paramDefToSchema` single source of truth. Pinned by test/gateway-model-messages.test.ts (8 cases). Folds into the open v0.42.9.0 PR (#1759) — these complete the eval-readiness wave by making skillopt actually run against a live model. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(skillopt): budget no-pricing for Haiku silently scored every rollout 0 Surfaced by the SkillOpt real-LLM eval (Track B). Two coupled bugs that made a budget-capped Haiku run report a vacuous "0/N" measurement in ~2ms with zero LLM calls — indistinguishable from a real deficient-skill score: 1. Claude Haiku 4.5's canonical dateless id (`claude-haiku-4-5`) was missing from anthropic-pricing.ts (only the dated `-20251001` was present). With `--max-cost` set, BudgetTracker.reserve() threw no_pricing on the FIRST chat() of every rollout. Added the dateless entry (sonnet already had its dateless form). 2. runValidationGate swallowed that BUDGET_EXHAUSTED error — runWithLimit settled it as {ok:false}, which the gate turned into median:0. A pricing/cap crash became a fake score. The gate now scans settled results for isMustAbortError() and re-throws so the caller aborts loudly; ordinary (non-abort) rollout errors still fail-open to 0 (judge-hiccup posture kept). Pinned by test/skillopt/validate-gate-abort.test.ts (3 cases). Folds into the open v0.42.9.0 PR (#1759). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): llms-full.txt over size budget — drop what-schemas-unlock from full bundle The toolLoop + budget bug-fix annotations grew CLAUDE.md, pushing llms-full.txt to 756KB over the 750KB FULL_SIZE_BUDGET (the `build-llms > size budget` test failed, failing the `test` CI job). CLAUDE.md stays inlined by design (it's the point of the one-fetch bundle), so per the budget comment's own guidance ("ship with includeInFull=false exclusions") this excludes docs/what-schemas-unlock.md (15.4KB value-explainer, not load-bearing operational reference) from llms-full.txt; it stays linked in llms.txt. Bundle now 740KB with ~9KB headroom. No budget bump — 750KB is near the ~190k-token-context fit ceiling. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(ci): re-admit policy docs into ci-cache-hash before doc relocation docs/**/*.md is deny-listed from the CI cache hash (test-irrelevant). The CLAUDE.md restructure moves test/release POLICY into docs/TESTING.md + docs/RELEASING.md, which DO carry contracts the test suite reads. Without re-admitting them, a policy-only edit would produce the same cache hash and skip the test shard that runs the build-llms + doc-history guards (false-pass). Adds an ALLOW_PATTERNS re-admit step after the deny, scoped to the named policy docs (not a blanket docs un-deny). Lands FIRST, before any doc moves. Pinned by 3 new cases in test/scripts/ci-cache-hash.test.ts: TESTING.md + RELEASING.md edits MUST change the hash; docs/guide.md still must not. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(docs): relocate Key files / thin-client / Testing out of CLAUDE.md (verbatim) CLAUDE.md had grown to 592KB / ~147k tokens auto-loaded every session (~77% of the llms-full.txt single-fetch bundle). The per-file index was append-only by mandate. This is the exact thin-dispatcher-vs-fat-blob anti-pattern gbrain exists to fix, so CLAUDE.md becomes a thin orientation + resolver that points at on-demand docs. This commit is the VERBATIM move (content-preserving — the next commit compresses): - docs/architecture/KEY_FILES.md <- ## Key files + the calibration key-files cluster + Schema Cathedral v3 impl detail - docs/architecture/thin-client.md <- ## Thin-client routing - docs/TESTING.md <- ## Testing - ## Commands DROPPED (18 'added in vX.Y' history blocks; current surface is gbrain 0.41.38.0 -- personal knowledge brain USAGE gbrain <command> [options] SETUP init [--pglite|--supabase|--url] Create brain (PGLite default, no server) migrate --to <supabase|pglite> Transfer brain between engines upgrade Self-update check-update [--json] Check for new versions doctor [--json] [--fast] Health check (resolver, skills, pgvector, RLS, embeddings) integrations [subcommand] Manage integration recipes (senses + reflexes) PAGES get <slug> Read a page put <slug> [< file.md] Write/update a page delete <slug> Delete a page list [--type T] [--tag T] [-n N] List pages SEARCH search <query> Keyword search (tsvector) query <question> [--no-expand] Hybrid search (RRF + expansion) ask <question> [--no-expand] Alias for query IMPORT/EXPORT import <dir> [--no-embed] Import markdown directory sync [--repo <path>] [flags] Git-to-brain incremental sync sync --watch [--interval N] Continuous sync (loops until stopped) sync --install-cron Install persistent sync daemon export [--dir ./out/] Export to markdown export --restore-only [--repo <p>] Restore missing supabase-only files [--type T] [--slug-prefix S] With optional filters FILES files list [slug] List stored files files upload <file> --page <slug> Upload file to storage files upload-raw <file> --page <s> Smart upload (size routing + .redirect.yaml) files signed-url <path> Generate signed URL (1-hour) files sync <dir> Bulk upload directory files verify Verify all uploads EMBEDDINGS embed [<slug>|--all|--stale] Generate/refresh embeddings LINKS link <from> <to> [--type T] Create typed link unlink <from> <to> Remove link backlinks <slug> Incoming links graph <slug> [--depth N] Traverse link graph (returns nodes) graph-query <slug> [--type T] Edge-based traversal with type/direction filters [--depth N] [--direction in|out|both] TAGS tags <slug> List tags tag <slug> <tag> Add tag untag <slug> <tag> Remove tag TIMELINE timeline [<slug>] View timeline timeline-add <slug> <date> <text> Add timeline entry TOOLS extract <links|timeline|all> Extract links/timeline (idempotent) [--source fs|db] fs (default) walks .md files; db iterates engine pages [--dir <brain>] brain dir for fs source [--type T] [--since DATE] filters (db source) [--dry-run] [--json] publish <page.md> [--password] Shareable HTML (strips private data, optional AES-256) check-backlinks <check|fix> [dir] Find/fix missing back-links across brain lint <dir|file> [--fix] Catch LLM artifacts, placeholder dates, bad frontmatter orphans [--json] [--count] Find pages with no inbound wikilinks salience [--days N] [--kind P] v0.29: pages ranked by emotional + activity salience anomalies [--since D] [--sigma N] v0.29: cohort-based statistical anomalies (tag, type) transcripts recent [--days N] v0.29: recent raw .txt transcripts (local-only) dream [--dry-run] [--json] Run the overnight maintenance cycle once (cron-friendly). See also: autopilot --install (continuous daemon). check-resolvable [--json] [--fix] Validate skill tree (reachability/MECE/DRY) report --type <name> --content ... Save timestamped report to brain/reports/ BRAIN (capture / ideate / explore — v0.37/v0.38) capture [content] [--file PATH] Single entrypoint for getting content into the brain [--stdin] [--slug s] [--type t] Inline content / file / stdin; writes to inbox/ by default [--source ID] [--quiet|--json] Multi-source brains: route to a non-default source brainstorm <question> [--json] Bisociation idea generator (hybrid search + far-set + judge) [--save|--no-save] [--limit N] lsd <question> [--json] Lateral Synaptic Drift: inverted-judge brainstorm [--save|--no-save] [--limit N] rewarding far-from-obvious + axiomatic inversions SOURCES (multi-repo / multi-brain) sources list Show registered sources sources add <id> --path <p> Register a source (id = short name, e.g. 'wiki') sources remove <id> Remove a source + its pages sync --all Sync all sources with a local_path sync --source <id> Sync one specific source repos ... DEPRECATED alias for 'sources' (v0.19.0) CODE INDEXING (v0.19.0 / v0.20.0 Cathedral II) code-def <symbol> [--lang l] Find the definition of a symbol across code pages code-refs <symbol> [--lang l] Find all references to a symbol (JSON-first) code-callers <symbol> Who calls this symbol? (v0.20.0 A1) code-callees <symbol> What does this symbol call? (v0.20.0 A1) query <q> --lang <l> Filter hybrid search to one language (v0.20.0) query <q> --symbol-kind <k> Filter to symbol type (function|class|method|...) (v0.20.0) reconcile-links [--dry-run] Batch-recompute doc↔impl edges (v0.20.0) reindex-code [--source id] [--yes] Explicit code-page reindex (v0.20.0) sync --strategy code Sync code files into the brain JOBS (Minions) jobs submit <name> [--params JSON] Submit background job [--follow] [--dry-run] jobs list [--status S] [--limit N] List jobs jobs get <id> Job details + history jobs cancel <id> Cancel job jobs retry <id> Re-queue failed/dead job jobs prune [--older-than 30d] Clean old jobs jobs stats Job health dashboard jobs work [--queue Q] Start worker daemon (Postgres only) ADMIN stats Brain statistics health Brain health dashboard history <slug> Page version history revert <slug> <version-id> Revert to version features [--json] [--auto-fix] Scan usage + recommend unused features autopilot [--repo] [--interval N] Self-maintaining brain daemon config [show|get|set] <key> [val] Brain config storage status [--repo <path>] Storage tier status and health [--json] (git-tracked vs supabase-only) serve MCP server (stdio) serve --http [--port N] HTTP MCP server with OAuth 2.1 --token-ttl N Access token TTL in seconds (default: 3600) --enable-dcr Enable Dynamic Client Registration --public-url URL Public issuer URL (required behind proxy/tunnel) call <tool> '<json>' Raw tool invocation version Version info --tools-json Tool discovery (JSON) Run gbrain <command> --help for command-specific help. + the per-command KEY_FILES entries; content stays in git) CLAUDE.md gains: a Reference map (resolver), a Maintaining section (the anti-disease rule), and a Cross-cutting invariants subsection under Architecture so the must-never-violate rules (trust fail-closed, sourceScopeOpts isolation, JSONB trap, engine parity, contract-first, migrations, multi-source) still auto-load after the index moved out. Result: CLAUDE.md 592KB -> 61KB; llms-full.txt 740KB -> 210KB (new docs link-only until compressed). build-llms drift + budget test green; verify 29/29 green. The pre-move content is recoverable at git show <this^>:CLAUDE.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(docs): compress relocated docs to current-state + add recurrence guard Compresses the verbatim-relocated reference docs from append-only release-history to current-state-only (the disease cure), then makes recurrence structurally impossible via a CI guard. Compression (fan-out subagents + adversarial verify, audited mechanically): - KEY_FILES.md 453KB -> 356KB; TESTING.md 42KB -> 38KB; thin-client.md already clean. - 393/393 entries preserved; every src/test/scripts path from the verbatim original survives (mechanical comm-check); zero bolded **v0. markers remain. - Conservative ratio (~22%) because the content is invariant-dense — correctness over brevity. Dropped: **vX.Y.Z (#NNN):** clauses, codex/review tags, contributor credits, PR-numbers-as-ids, pre-fix/then/was-now history deltas. Kept: every exported symbol, invariant, and Pinned-by reference. Verbatim original recoverable at git show <relocation-commit>:docs/architecture/KEY_FILES.md. Recurrence guard (scripts/check-key-files-current-state.sh, wired into verify + check:all): - HARD: bans the bolded **v0.<digit> marker in the reference docs (scoped — plain 'as of pgvector 0.7' prose is fine, no false positives). - HARD: CLAUDE.md size cap (90KB; currently 61KB) — the structural backstop. - Pinned by test/scripts/check-key-files-current-state.test.ts (7 cases). Content contracts (test/build-llms.test.ts, +5 cases per codex outside-voice): CLAUDE.md keeps inline ship IRON RULES (version format, document-release, never-hand-roll); AGENTS.md keeps its boot order; llms indexes the new docs; KEY_FILES stays link-only (not inlined). Privacy: scrubbed the relocated 'wintermute/chat/' source-boost examples + the literal harvest-lint regex to generic placeholders (legitimate in allowlisted CLAUDE.md; genericized for the new public docs per the privacy rule). Reverts the |
||
|
|
7b0d99adb0 |
v0.42.2.0 feat: gbrain connect — one-command Claude Code onboarding from a bearer token (#1683)
* fix: gbrain auth create dropped the name on the bare (no-flag) form Extract parseAuthCreateArgs; only exclude the --takes-holders value from the positional search when the flag is present (rest[takesIdx+1] resolved to rest[0] when takesIdx === -1, silently dropping the name). Add regression test. * feat: gbrain connect — one-command Claude Code onboarding from a bearer token New connect command prints a paste-ready claude-mcp-add block (or --install wires it + smoke-tests the token via a raw-bearer get_brain_identity probe). Direct HTTP MCP, literal-token default, URL normalization, token header-injection guard, --json redaction, execFileSync (no shell). Wired into CLI_ONLY + CLI_ONLY_SELF_HELP + handleCliOnly. 58 unit + 3 PGLite-E2E cases; e2e-test-map updated. * docs: lead CLAUDE_CODE.md with gbrain connect (remote fast path) + README one-liner Regenerate llms-full.txt for the README change. * refactor: pre-landing review fixes for gbrain connect - DRY: single DEFAULT_PROBE_TIMEOUT_MS + shared isAuthErrorMessage predicate - reuse promptLine (shared stdin lifecycle) for the --install confirm - harden redactToken with a Bearer <value> scrub (defense in depth) - +8 tests: orchestrator guard paths, deterministic timeout, invalid --timeout-ms, Bearer-redaction * fix: adversarial-review hardening for gbrain connect - probe: Promise.race the call against a real timer so a stalled connect()/SSE handshake (signal alone doesn't cover it) can't hang --install indefinitely - probe: close transport even if client.connect() throws - parseArgs: reject a missing/flag-shaped value (e.g. --token --install) - block link-local / cloud-metadata hosts (169.254/fe80:/fd00:ec2::254) — keeps localhost + RFC1918 LAN brains working - non-interactive --install now requires --yes - clearer message when --force removed then add failed +8 tests covering each * fix: codex-review P2s for gbrain connect - POSIX single-quote the rendered claude-mcp-add command so a token with shell metacharacters ($(), backticks) can't trigger command substitution on paste - detect IPv4-mapped IPv6 metadata addresses (::ffff:169.254.x.x / ::ffff:a9fe:*) so they don't bypass the link-local guard +3 tests * chore: bump version and changelog (v0.42.2.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: document gbrain connect + connect-probe in CLAUDE.md Key files (v0.42.2.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: gbrain connect — add codex and perplexity agents --agent codex emits 'codex mcp add ... --bearer-token-env-var GBRAIN_REMOTE_TOKEN' (token read from the env var at runtime, never in Codex config; --install runs it). --agent perplexity prints the URL + token for the Settings → Connectors GUI (no --install). Generalized the command file: AGENT_SPECS table, buildCodexMcpAddArgv, cmdString(binary,argv), binary-generic ConnectDeps (hasBinary/runBinary/env), agent-aware buildConnectBlock/buildJson. +25 tests. * docs: codex + perplexity connect paths (new CODEX.md, README, CHANGELOG, CLAUDE.md) Regenerate llms-full.txt for the CLAUDE.md/README edits. * test: real-CLI E2E for connect — drive actual claude + codex against a live server Adds claude-code + codex cases to connect-bearer.test.ts that run the real 'claude mcp add' / 'codex mcp add' through 'gbrain connect --install' against a live 'gbrain serve --http' (sandboxed HOME/CODEX_HOME), then assert via 'claude mcp get' / 'codex mcp get' that the server registered (and codex's token stays out of config). Skips when the binary is absent. Perplexity is GUI-only so it's print-asserted. Regen llms for the CLAUDE.md note. * docs: perplexity OAuth + serve --bind/--public-url footgun (per Perplexity feedback) PERPLEXITY.md now documents the host-side HTTP setup (gbrain serve --http --bind 0.0.0.0 --public-url, the v0.34 ECONNREFUSED footgun) and the OAuth 2.1 client_credentials path (gbrain auth register-client) alongside the legacy bearer token. The 'connect --agent perplexity' output points at the same bind/public-url requirement + PERPLEXITY.md. * feat: gbrain connect --oauth — client-credentials path for perplexity/generic OAuth is the correct path for a third-party cloud connector (Perplexity): instead of a long-lived full-access bearer token, the connector gets Issuer URL + Client ID + Client Secret and mints short-lived scoped tokens. --oauth --register mints a least-privilege client on the host (shells gbrain auth register-client); --oauth --client-id/--client-secret uses an existing one. Rejected for claude-code/codex (bearer) and with --install. Issuer derived from the mcp-url. New E2E proves the full chain: register → connect --oauth → OAuth discovery → /token client_credentials mint → get_brain_identity tool call against a live server. Docs: PERPLEXITY.md leads with OAuth; README + CLAUDE.md updated; +18 unit cases. * docs: add gbrain connect to INSTALL.md MCP section + link CODEX.md The remote-client onboarding command was documented in README/CLAUDE_CODE/CODEX/ PERPLEXITY but missing from INSTALL.md §3 (the natural 'how do I connect a client' home). Add the one-command connect how-to (claude-code/codex/perplexity) and the missing docs/mcp/CODEX.md link. * fix: connect LEARN_INSTRUCTION names put_page, not CLI-only capture The self-orientation block told a connected agent that `capture` is an available MCP tool. It isn't — `capture` is a CLI-only convenience command; the MCP write tool is `put_page`. An agent that followed the instruction hit "unknown tool". Drop capture; put_page was already in the list. Adds a regression block to connect.test.ts. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: serve --http surfaces skill-publishing status (banner + nudge) When mcp.publish_skills is OFF, connected agents can search/write but can't call list_skills/get_skill, so the host's skill catalog is invisible to them. The startup banner now shows a Skills: line, and a stderr nudge fires when off with the paste-ready fix. Pure skillPublishStatus() helper, unit-tested. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: prove the local stdio MCP funnel end-to-end Spawns real `gbrain serve` (stdio) against a freshly init --pglite brain and drives the official MCP SDK client through initialize -> tools/list -> tools/call (get_brain_identity + search). Pins the advertised core-tool set against what the server actually exposes (asserts capture is NOT advertised). This funnel had zero e2e coverage before. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: make batch-retry-audit ENOENT case hermetic The 'no-op when audit dir does not exist' case called pruneOldBatchRetryAuditFiles(30) without a GBRAIN_AUDIT_DIR override, so it read the real ~/.gbrain/audit and flaked (kept:1) on any dev machine with a batch-retry-*.jsonl on disk. Point it at a guaranteed-missing temp subdir, matching this file's own hermetic-header contract. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: two-funnel coding-agent onboarding (Claude Code / Codex) New tutorial docs/tutorials/connect-coding-agent.md: Path A (connect to an existing brain) + Path B (start from nothing, local stdio), the brain-first protocol to paste into CLAUDE.md/AGENTS.md, and the four translatable habits. README gains a 'Quick start: Claude Code or Codex' fork separating lightweight retrieval from the full autonomous install. INSTALL.md shows the one-command wire-up at the standalone CLI section. mcp/CLAUDE_CODE + CODEX cross-link the tutorial + note publish_skills + capture-is-CLI-only. Tutorial promoted to Shipped in the tutorials index. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: changelog + regenerated llms (v0.42.2.0) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: CLAUDE.md Key Files annotation for two-funnel onboarding wave Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
eefe8b5741 |
v0.42.1.0 feat: gbrain skillopt — self-evolving skills (closes #1481) (#1563)
* feat(skillopt): foundation modules — types, lr-schedule, benchmark, score, audit, lock
* feat(skillopt): edit primitives — apply-edits (D5+D9), rejected-buffer LRU, version-store (D8 history-intent-first)
* feat(skillopt): rollout (D2 gateway.toolLoop + D13 read-only allowlist), reflect (D7 two calls), validate-gate (D12 median+epsilon, D4 parallel), preflight (D3), bundled-skill-gate (D16)
* feat(skillopt): orchestrator (D6 slow-update, D10 ASCII diagrams, D11 caching), checkpoint, bootstrap (D15 sentinel), CLI dispatch + help
* feat(skillopt): cycle phase (F1 dream-loop wiring), PROTECTED_JOB_NAMES + MCP op (F6 admin scope + allowlist) + Minion handler (F7 --background)
* feat(skillopt): full cathedral — --all batch (F4), --target-models fleet (F5), write-capture (F10), held-out scaffold (F11), adversarial suite 41 cases (F2), E2E PGLite (F3), meta-skill bundle (T7), reflect+judge evals (F8+F9), docs (T10)
* chore: bump version to v0.42.0.0 (MINOR — significant new feature)
* fix(skillopt): wire trajectories from forward gate to reflect + fix parseEditsResponse parser misuse
Two related v0.42.0.0 bugs that conspired to make `runSkillOpt` structurally
unable to accept any candidate edit. Either alone would have killed self-evolution;
together they made the loop a no-op for every input.
**Bug 1 (orchestrator gap):** `runOptimizationLoop` in orchestrator.ts called
`runReflect({successes: [], failures: []})` with hardcoded empty arrays. The
forward gate's `scoredRollouts` were computed then voided. `runReflect`
short-circuits both modes when their batches are empty, so the optimizer was
never asked to propose an edit. Every step hit the no_edits_applied branch.
Fix: add `scoredRollouts: ScoredRollout[]` to `GateResult` and
`runsPerTask?: number` to `ValidateGateOpts`. Forward pass uses
`runsPerTask: 1`; orchestrator partitions returned rollouts by `score >= 0.5`
and threads real successes + failures into `runReflect`.
**Bug 2 (parser misuse):** `parseEditsResponse` in reflect.ts routed every
optimizer response through `parseJudgeJson` first. `parseJudgeJson` looks for
a `score` key (it's a judge-output parser, not an edits parser) and returns
null for any JSON without one — including the well-formed `{"edits": [...]}`
the optimizer is contractually required to emit. The function then early-
returned `[]` and the actual `tryExtractEdits` path on the next line was
unreachable dead code.
Fix: drop the wrong-typed guard. `parseEditsResponse` now calls
`tryExtractEdits` directly. Export it so `reflect.test.ts` can pin the
contract independently of the chat transport.
**Why this slipped through 152 prior skillopt tests:** zero unit coverage
of `parseEditsResponse` or `runReflect`. The existing E2E `all-reject` case
asserted no_improvement (which was true for the wrong reason — empty edits,
not gate rejection). Both bugs were structurally invisible to the existing
test surface.
**New coverage:**
- `test/skillopt/reflect.test.ts` (15 cases):
- 8 `parseEditsResponse` cases including the IRON-RULE regression pin
for the v0.42.0.1 fix (`{"edits": [...]}` JSON must survive the parser).
- 7 `runReflect` D7 contract cases: both modes fire, empty-batch skips,
additive token usage, one-mode-throws-other-still-works, rejected-buffer
flows into anti-bias prompt.
- Documents the trailing-comma limitation as an explicit out-of-scope pin
(so a future tightening of `tryExtractEdits` lights this test up
intentionally).
- `test/e2e/skillopt-loop.serial.test.ts` (7 cases):
- HAPPY PATH: stubbed `gateway.chat` acts as both target agent (emits
sections based on skill content) and optimizer (proposes a real
add-Citations edit). Drives `runSkillOpt` end-to-end against PGLite.
Asserts outcome=accepted, SKILL.md mutated with new section,
frontmatter preserved (D5), history has one committed row,
best.md mirrors disk, delta > epsilon, receipt fields populated.
- 5 broken cases (each isolates a distinct orchestrator-visible failure):
1. Below-baseline regression: optimizer proposes a destructive edit;
gate rejects with reason=below_baseline; SKILL.md unchanged;
rejected-buffer captures the bad edit for anti-bias context.
2. Malformed reflect JSON: orchestrator degrades gracefully to
no_improvement without crashing.
3. Anchor-not-found: applyEditBatch rejects all; sel gate skipped;
rejected-buffer captures with reason=apply_failed.
4. Budget exhausted mid-step: outcome=aborted, no pending rows survive.
5. Converged-skill re-run: starting from already-perfect skill →
no_improvement (no thrash on a well-tuned starting point).
- IDEMPOTENT RE-RUN: drive runSkillOpt twice in sequence. Run 1 accepts.
Run 2 sees improved baseline, no failures, returns no_improvement.
SKILL.md byte-identical to post-run-1; history still has exactly 1
committed row. Proves stability at the fixed point.
All hermetic (no DATABASE_URL, no API keys). PGLite in-memory engine,
tempdir SKILL.md + benchmark, stubbed gateway.chat via
`__setChatTransportForTests`. `.serial.test.ts` because the stub installs
module state and the loop walks shared disk state across epochs.
Test counts after fix: 174 skillopt-surface tests pass (149 pre-existing
unit + 15 new reflect unit + 3 existing E2E + 7 new E2E). Typecheck clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(cycle): align ALL_PHASES skillopt position with actual dispatch order
v0.42.0.0 added skillopt to ALL_PHASES right after `patterns` (line 127), but
the dispatch block in runCycle (line ~1912) actually runs skillopt between
`conversation_facts_backfill` and `embed`. The two were inconsistent, and the
serial test `report.phases.map(p => p.phase)).toEqual(ALL_PHASES)` was failing
on master because of it.
A second pre-existing failure: the two phase-count assertions in
`test/core/cycle.serial.test.ts` still said `toBe(20)` even though
ALL_PHASES grew to 21 when skillopt was added. The author bumped the array
but forgot the test.
Two fixes, one commit:
1. Move `'skillopt'` in ALL_PHASES from after `patterns` to between
`conversation_facts_backfill` and `embed`, matching where runCycle
actually dispatches it. Runtime behavior is unchanged — only the
declaration order moves. Updated the surrounding comment to call out
the position invariant and reference the test that pins it.
2. Update both `toBe(20)` assertions in cycle.serial.test.ts to `toBe(21)`
with a v0.42.0.0 history line in the running comments.
Why declaration follows runtime (not the other way around): the comment
intent ("Runs AFTER patterns — graph-fresh") is still satisfied because
"after the entire main graph-mutating cluster" is strictly fresher than
"right after patterns". No design intent is lost.
Test result: cycle.serial.test.ts is now 28/28 (was 27/28 on master + my
prior commit). Skillopt suite still 174/174.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(ci): bump PHASE_SCOPE assertion to 21 + fix skill-optimizer Anti-Patterns case
Two CI failures pre-existing on this branch since the v0.42.0.0 skillopt
cathedral landed; master is green because skillopt didn't exist there yet.
1. test/phase-scope-coverage.test.ts asserted ALL_PHASES.length === 20.
skillopt is the 21st phase. Bumped to 21 with v0.42.0.0 history line
in the comment chain. Sibling fix to the cycle.serial.test.ts bump
in commit
|
||
|
|
374deff579 |
v0.41.7.0 feat: compact list-format resolver + 300-skill scaling tutorial (#1407)
* feat(check-resolvable): parseResolverEntries accepts compact list format
Add the second parser branch alongside the existing markdown-table branch
so RESOLVER.md and AGENTS.md can use the OpenClaw-native list shape:
- **skill-name**: trigger1 | trigger2 | trigger3
- skill-name: trigger1 | trigger2
Constraints:
- Skill names must be kebab-lowercase ([a-z][a-z0-9-]+). Bold names
starting with an uppercase letter (e.g. **Note**, **Convention**)
are deliberately skipped so prose bullets in real-world AGENTS.md
files don't get mis-parsed as fake skill rows.
- skillPath is always derived as skills/<name>/SKILL.md. An optional
arrow suffix (Unicode -> or ASCII ->) is stripped from the trigger
string but NOT honored as a path. Downstream consumers
(routing-eval.ts skillSlugFromPath, the manifest check at line 367)
assume the convention. For non-conventional paths, use the table
format.
- Multiple triggers fan out to one entry per trigger. checkResolvable
dedupes by skillPath downstream, so the reachability count counts
each skill once regardless of trigger fan-out.
The parser body is restructured to an if/else-if shape so the existing
'continue' on non-table rows no longer short-circuits the list branch.
Unit tests cover 11 new cases: bold + plain name shapes, multi-trigger
fan-out, Unicode and ASCII path-suffix strip, ellipsis filter, empty
pipe segments, mixed-shape files, section tracking, and two D4
regression cases (prose-bullet rejection + convention-violation
silent-skip).
Closes #1370 — credit @garrytan-agents for the original PR that flagged
the parser gap.
* test(check-resolvable): integration fixtures + regression suite for compact format
Two fixtures pin the v0.41.7.0 parser fix at the integration layer:
test/fixtures/openclaw-compact-resolver/
List-format only RESOLVER.md with 10 fictional skills (gift-advisor,
flight-tracker, email-triage, etc.), each with valid frontmatter
triggers. A trailing 'Notes' section embeds 4 prose bullets
(- **Note**:, - **Convention**:, - **TODO**:, - **Important**:)
that pin the D4 kebab-lowercase regex tighten: if the regex ever
regresses to permissive [\w-]+, those prose bullets would surface
as orphan_trigger warnings and the test fails loudly.
test/fixtures/openclaw-mixed-merge/
Tests the v0.31.7 D-CX-14 multi-resolver merge: workspace-root
AGENTS.md (compact list, 3 skills) + skills/RESOLVER.md (table
format, 5 skills). The merge dedups by skillPath and counts each
skill once.
The regression test (test/check-resolvable-openclaw-compact.test.ts)
runs 8 assertions across both fixtures:
1. unreachable === 0 on the compact fixture (the 'pre-v0.41.7.0
reported 238 FAILs on a 306-skill OpenClaw, post-fix 0' headline).
2. zero error-severity issues; report.ok === true.
3. zero mece_gap warnings (every stub ships valid triggers).
4. zero orphan_trigger warnings for the 4 prose-bullet names — D4
regex regression guard at integration level.
5. zero missing_file warnings.
6. mixed-merge: total_skills === 8 (5 table + 3 list), all reachable.
7. mixed-merge: errors.length === 0; report.ok === true.
8. mixed-merge: each expected skill from BOTH shapes is non-unreachable
(catches the bug where one shape silently swallows the other via
dedup-by-skillPath).
* docs(guides): scaling-skills.md walkthrough for 300-skill agents
Three-tier architecture for agents that have outgrown the always-loaded
skill manifest:
Tier A — always loaded (~35 skills, in the system prompt every turn)
Tier B — resolver-routed (~85 skills, looked up via RESOLVER.md/AGENTS.md
only when no Tier A match)
Tier C — dormant (~180 skills, on disk but not injected into the prompt)
Real numbers from Garry's 306-skill OpenClaw: 25K tokens of skill
descriptions per turn collapsed to 4K tokens (~21K tokens freed per
turn) with zero capability loss. The compact list-format resolver
(v0.41.7.0) is the parser-level enabler for this pattern.
The guide covers:
- The scaling wall (when the always-loaded manifest stops working)
- The three tiers + per-turn token math
- What the resolver actually does (routing-table-but-cheaper pattern)
- The compact list format (kebab-lowercase contract, optional path
suffix, mixed-shape support)
- The 'gbrain doctor' / 'gbrain check-resolvable --strict' safety net
- Implementation walkthrough (audit → tier → disable → resolver →
doctor)
- The scaling curve (50 → 100 → 200 → 300 → 1000, no ceiling)
Voice + privacy cleanup applied per CLAUDE.md rules:
- Wintermute → 'Garry's OpenClaw' / 'your OpenClaw'
- Unicode em dashes stripped; ASCII '--' preserved in command flags
- Made-up 'check_resolvable' invocation replaced with real
'gbrain doctor' and 'gbrain check-resolvable --json'/'--strict'
- Blog-style 'Previous in this series' footer dropped
Wiring:
- scripts/llms-config.ts registers the new guide in the curated
array so 'bun run build:llms' picks it up. docs/UPGRADING_
DOWNSTREAM_AGENTS.md excluded from the inlined bundle to stay
under the 600KB FULL_SIZE_BUDGET after adding the new content.
- docs/tutorials/README.md gains a one-line entry pointing at the
guide under Related documentation.
- llms.txt + llms-full.txt regenerated.
* chore: bump version and changelog (v0.41.7.0)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs: update CLAUDE.md for v0.41.7.0 compact-format resolver
Annotate the src/core/check-resolvable.ts entry with the v0.41.7.0
parseResolverEntries compact list-format support: kebab-lowercase name
gate (closes the prose-bullet false-positive class), path-suffix strip
contract (skillPath always derived as skills/<name>/SKILL.md so
routing-eval and the manifest check don't drift), multi-trigger fan-out
plus checkResolvable downstream dedupe, the 238 FAILs to 0 OpenClaw
headline, the two integration fixtures pinning the regression, and the
docs/guides/scaling-skills.md pointer for the tutorial context.
Regenerate llms-full.txt to match (CLAUDE.md edit chaser, per the
CLAUDE.md own rule about test/build-llms.test.ts catching drift).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
||
|
|
af5ee1eb5a |
v0.40.8.1 docs: README rewrite + personal-brain + company-brain tutorials (#1345)
* docs: rewrite README lead around search-vs-think differentiator The current README opened with a generic "smart but forgetful" tagline that buried the actual differentiator. Garry's 2026-05-23 X thread crystallized the positioning: "Search gives you raw pages. Think gives you the answer." That plus graph traversal plus gap analysis is what nobody else ships in one box. Changes: - README lead now leads with the search-vs-think frame, the "nobody else does this" claim, and the "strategic moat / so you don't lose context" framing. - Collapsed five stacked "New in vX.Y.Z" paragraphs in the lead into one Recent Releases section after the install path, freeing the first viewport to be about what gbrain IS, not what shipped last week. - New ## Search vs think section with side-by-side CLI example, gap-analysis explanation, and the find_trajectory + think compounding story. - ORIGIN.md closing paragraph names think as the reason the brain is worth building. - No em dashes used as connectors (per humanizer rules). - All factual claims (page counts, benchmark numbers, version references) preserved verbatim. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: add v0.40.6.0 to README Recent Releases Picked up sync --all + per-source locks + sources status dashboard from the v0.40.6.0 merge. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): add visceral "what this looks like" before/after table Pulled verbatim from BrainBench Cat 29 — same question, same brain, Haiku judge. Shows what a typical personal-knowledge brain (top-K vector retrieval, what MemPalace / Mem0 / Hindsight ship) returns vs what gbrain think returns. Search hallucinates three people who actually work at OTHER companies; think correctly identifies what's known + flags the gap. Score: search 1/10 vs think 9/10. The before/after lands above Install so the reader sees concrete differentiation before they decide to install. Backs the abstract "search vs think" claim from the lead with a real receipt. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): swap before/after example to verbatim Cat 29 receipt (Q2 ARR) Per @garrytan — the prior example used the synthetic Q1 (employees of Horizon TECH 6) with paraphrased answer text. Replaced with the verbatim Cat 29 Q2 receipt: actual question (with the in-question typo that exists in the eval), actual truncated search-answer text from the JSON receipt, actual think-answer text with all three ARR readings + citations, and the actual Haiku judge verdicts pasted verbatim. Also strips Mem0 + Hindsight references from the comparator phrasing — Mem0 is a YC company we don't want to single out, and Hindsight was a hackathon-stage project that never launched. The comparison phrasing now reads "MemPalace and most peer AI-memory stacks" — accurate without naming systems we shouldn't be benchmarking against. The three things gbrain think did that a typical top-K retrieval cannot are listed below the table for readers who want the takeaway in plain English: 1. Caught the name typo (question said "Acme AI 0", brain has "Acme CO 0") 2. Walked the typed-claim Facts fence to build a chronological trajectory 3. Cited every claim Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): add "Your brain's shape (schema packs)" section Per @garrytan — the README had nothing on schema packs (v0.38/v0.39 dynamic-schema cathedral). New section lands between "How to get data in" and "Recent releases" so the narrative flow reads: what gbrain is → concrete example → install → query (search vs think) → get data in → schema packs (how the brain understands your shape) → recent releases → loop → capabilities → ... The section opens by naming the differentiator out loud: "Most personal- knowledge tools force one fixed layout: their idea of notes + people + tags. Drop a Notion export or your own years-old Obsidian vault and the agent doesn't know what your folders mean." Three options surfaced: - gbrain-base (default, zero-config Garry layout) - gbrain-recommended (extends base with 13 more dirs) - your own pack via the schema detect → suggest → review-candidates three-command magical moment Six representative CLI verbs shown verbatim. Closes with one paragraph explaining the threading through every read/write path (parseMarkdown, whoknows, extract_facts, search cache) and a one-line summary of the 7-tier resolution chain pointing at docs/architecture/schema-packs.md for the full reference. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): de-brand "Think" in the lead; reframe as GBrain's brain layer Per @garrytan — "we don't need to brand it Think, we want to say the think command later but we don't lead on it!" Changes: - Lead sentence: "Search gives you raw pages. Think gives you the answer" → "Search gives you raw pages. GBrain gives you the answer, through a brain layer." The product is GBrain; think is the CLI verb that runs the brain layer, introduced later. - Two-bullet differentiator list: the "gbrain think" bullet now leads with the capability ("A synthesis layer that gives you the actual answer.") rather than the CLI command. The bullet body still names what the layer does (synthesized prose, citations, gap analysis). - Strategic-moat paragraph: "`gbrain think` is what makes the moat usable" → "The brain layer is what makes the moat usable." - "What this looks like" table header: "GBrain `think`" → "GBrain's brain layer (one synthesized answer, run via `gbrain think`)". CLI command stays in the cell so the example is reproducible; the framing leads with what it IS, not the verb. - Three-bullet takeaway under the table: "Three things `gbrain think` did" → "Three things the brain layer did". Aggregate sentence: "gbrain think averages 5.60/10" → "GBrain's synthesis layer averages 5.60/10". - Section heading: "## Search vs think" → "## Two ways to query your brain". The section still introduces `gbrain search` and `gbrain think` as the two CLI verbs side by side; the heading no longer brands "think" as the thing. The "think command" comes through naturally where it appears as a CLI example. The PRODUCT is GBrain. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): switch to first-person voice (Garry speaking directly) Per @garrytan — drop "Built by..." third-person framing and write in first person. Edits: - Lead credibility paragraph: "Built by the President and CEO of Y Combinator to run his actual AI agents" → "I'm Garry Tan, President and CEO of Y Combinator. I built GBrain to run my own AI agents." Subsequent sentences switch "his deployments" → "my deployments", "the agent ingests... you wake up smarter" → "my agent ingests... I wake up smarter — and so will you" (the closing "and so will you" connects Garry's experience to the reader's). - Compounding paragraph: "As Garry's personal agent gets smarter, so does yours" → "As my personal agent gets smarter, so does yours." - Schema packs section: "the layout used by Garry's production brain" → "the layout my production brain uses". - License + credit: "Built by Garry Tan to run his OpenClaw and Hermes deployments — the production brain behind his actual AI agents" → "I built GBrain to run my OpenClaw and Hermes deployments — the production brain behind my AI agents." The whole top of the README now reads as Garry talking directly to the reader about what he built and why. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): DRY the lead — drop redundant "through a brain layer" + repeat GBrain "GBrain gives you the answer, through a brain layer. GBrain is the brain layer your AI agent has been missing..." — two GBrains, two brain layers. Tightened to: "GBrain gives you the answer. It's the brain layer your AI agent has been missing — the only one that does synthesis, graph traversal, and gap analysis in one box." Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): frame GBrain as a company brain too, link to YC RFS Per @garrytan — GBrain is now usable as a company brain (federated sync, OAuth scoping, Cat 22 source isolation), and YC just put company-brain on its Request for Startups. Added a paragraph after the personal-brain lead that names the three v0.34+ features that make multi-user safe (federated sync, per-source OAuth scoping, the Cat 22 leak-free source isolation), then links to https://www.ycombinator.com/rfs#company-brain with a one-line pitch: "if you're building in that space, you might as well build on this." The framing carries forward the personal-brain story while opening the aperture: GBrain works for one person (Garry's production deployment) AND for a team (per the features Cat 22 just verified). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): rewrite company-brain paragraph in plain English @garrytan caught me writing internal eval-suite jargon ("Cat 22 proves source isolation is leak-free across hybrid search, listPages, getPage, and federated reads") in a paragraph aimed at someone deciding whether to use GBrain. Rewritten in plain English: "Each person on the team gets their own slice of the brain, scoped by login. When you query, you only see what you're allowed to see — never another person's notes, never another team's data. We fuzz- tested this across every way you can read the brain (search, list, lookup, multi-source reads) and got zero leaks." Same factual content, zero internal vocabulary. The "Cat 22" name belongs in the benchmark page, not the front-door README. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(tutorials): add company-brain tutorial (Diataxis tutorial quadrant) End-to-end walkthrough for setting up GBrain as a multi-user company brain. Audience: founder / CTO / head of ops at a 10-50 person company who has heard about GBrain (from the YC RFS company-brain page or my tweets) and wants to set it up as their team's shared institutional memory. ~3700 words, written for a learner with zero prior gbrain knowledge. Twelve parts walk the reader from "I've never run gbrain" to "three teammates each query the brain through their own AI agent and see the correctly scoped answer": 1. The mental model (personal brain vs company brain, federated sources, OAuth scoping, what you get) 2. Prerequisites table (Postgres, embedding key, Anthropic key, git repo, Bun, host machine + cost projection) 3. Install + Postgres + API keys + doctor verify 4. Create three sources (shared / customers / internal) + sync 5. Spin up HTTP MCP server with --bind 0.0.0.0 + --public-url 6. Register one OAuth client per teammate with --source + --federated-read 7. Verify scoping works (alice can't see internal, bob can't see customers) 8. Connect each teammate's AI agent via thin-client install 9. First real `gbrain think` query showing sourced + synthesized + gap-analysis answer 10. Operating the brain (autopilot, doctor --remediate, sources status, admin dashboard) 11. Cost + speed expectations from the v0.40.6.0 benchmark 12. Common gotchas + troubleshooting Voice: - First-person Garry in intro / motivation paragraphs - Zero internal jargon. No "Cat 22", "P@5", "knobsHash", "RRF", "MRR" - Plain English throughout - No em dashes used as connectors (zero in final draft) - "Brain layer" / "synthesized answer" framing, not "Think" branding - No Hindsight, no Mem0 (per repeated session feedback) - Every example uses placeholder names (alice-example, bob-example, acme-co) Cross-linked from: - README.md company-brain paragraph - docs/INSTALL.md (under the migrate --to supabase block) - docs/architecture/topologies.md (See also section) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): drop version chatter, add Tutorials section, expand tutorial roadmap Per @garrytan: the README should read as the current docs written for people who have never known GBrain before. Versions are what the CHANGELOG is for. Version-chatter sweep across the README: - Killed the entire ## Recent releases section. That's a changelog summary, not docs. CHANGELOG.md owns it. - Stripped "(v0.38+)" from the ## How to get data in heading. - Rewrote "(the v0.38 put_page write-through plumbing)" as "(the database and on disk in one move)" — describes WHAT happens, not WHEN it shipped. - Stripped "(The legacy gbrain skillpack install managed-block model was retired in v0.36.0.0; run gbrain skillpack migrate-fence once if you're upgrading from an older release.)" from the skillpack paragraph. Upgrade history goes in the CHANGELOG. - Stripped "New in v0.40.4.0:" from the hybrid-search graph-signals description. Just describes what the feature does. - Stripped "As of v0.37," from the embedding-provider auto-detect paragraph in the Troubleshooting section. - Stripped "the embedding + reranker stack that became the v0.36.2.0 default" → "ships as the default" in the License + credit section. - Stripped "in Cat 29" from the synthesis-benchmark sentence (same internal-jargon class as Cat 22 which @garrytan called out earlier). Replaced ## Recent releases with ## Tutorials: - Links to the one shipped tutorial (company-brain.md) with a one-line description. - Names the next several planned tutorials in prose (not as broken links): personal brain quickstart, connect your agent, VC dealflow, vault migration, code brain. No fake links. - Points at the tutorial index page for the full roadmap. - Closes with "Open an issue describing the workflow you want documented" to invite prioritization input from real users. New tutorial index at docs/tutorials/README.md: - Shipped section: company-brain.md - In-progress roadmap with 7 candidate tutorials (personal brain quickstart, connect your agent, VC dealflow, vault migration, code brain, fully local, dream cycle setup) - Each roadmap entry names the persona, the core commands the tutorial would demonstrate, and the differentiator - "Want to write one?" section invites community PRs and points at company-brain.md as the model Reverted the obscure topologies.md "See also" link from pointing at company-brain.md specifically to pointing at the tutorials index overall — the tutorial belongs in the README's Tutorials section where users actually look, not buried in an architecture doc. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(tutorials): land personal-brain tutorial (full-stack install) The canonical solo install walkthrough, adapted from Garry's live setup session notes (the "Apple I, soldering breadboards" session). Builds the full stack: 2 GitHub repos, Telegram bot, AlphaClaw on Render, OpenClaw + GBrain + Supabase. About 2 hours end-to-end, $100-150/month sustained. Replaces the "Set up your personal brain in 30 minutes" stub on the tutorials roadmap with something genuinely complete. Edits to the source draft: - Stripped brain-page YAML frontmatter (type/access/links/etc — not needed for a public docs file) - Privacy sweep per CLAUDE.md: removed real-name references to the collaborator and the agents involved in the session, replaced with "a collaborator" / "my main agent" / generic placeholder names - First-person voice consistency: the source draft slipped between first person and third person ("Garry walked through"); rewrote everything in first person - Zero em-dashes used as connectors (verified by grep) - Added ZeroEntropy to the providers list (it's the default; not mentioning it would leave readers paying 2.6× more on embeddings) - Opened with a router paragraph that points brain-layer-only readers at INSTALL.md and team readers at company-brain.md, so each audience finds the right walkthrough fast Tutorials index updated: - personal-brain.md promoted from In Progress to Shipped (it IS the "set up your personal brain in 30 minutes" entry that was on the roadmap, just better-named and more complete) README.md Tutorials section now lists both shipped tutorials side by side. Reads naturally: solo install first (broader audience), team install second. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(tutorials): rewrite company-brain as a true superset of personal-brain Per @garrytan: the company brain tutorial should pick up from where the personal brain tutorial leaves off, not duplicate the install. Pedagogical flow now reads: "you already did personal-brain; here is what to add to make it multi-user." Restructure: - Opens with explicit "this tutorial picks up where the personal brain tutorial leaves off" + a router for readers who haven't done that one yet. No duplicated install steps. - Part 1 (mental model) reframed as "what changes when you go from personal to company" + "what this is NOT" (not a different install, not a thin-client-everywhere replacement). - Part 2 is now "switch the brain backend to multi-user Postgres" — surfacing the migrate --to supabase path for readers who started on PGLite, skip-to-Part-3 for readers already on Postgres. - Part 3 adds the new content @garrytan called for: per-person folder structure inside each source. customers/alice-example/, internal/ alice-example/, internal/bob-example/, internal/legal/ etc. So teammates' writes don't collide and the per-person/per-role abstraction is real on disk, not just in OAuth scope. - NEW Part 6: per-person crons. Each teammate gets their own scheduled tasks (7am customer digest for alice, 9am ops status for bob, weekly contract compliance for carol) scoped to their OAuth client so the cron can only touch their slice. - NEW Part 7: per-person skills. The 60+ shipped skills are generic; teams want a few specific ones (onboarding-new-hire, customer-success-followup, weekly-team-digest). Scaffolded via gbrain skillify scaffold, scoped via allowed_clients in frontmatter. - Existing strong parts retained: OAuth scoping + verify, per- teammate AI client connect, first synthesized query, operating notes, gotchas. Reworded where needed to refer back to personal- brain steps instead of re-explaining them. Voice + style sweeps: - Zero em-dashes used as connectors (verified by grep) - Zero internal jargon (no "Cat 22", "P@5", "RRF", "MRR", etc.) - Zero version chatter in body text (only the one acceptable reference to the dated benchmark filename in a docs link) - "Brain layer" / "synthesized answer" framing, not "Think" branding - First-person Garry voice in intro/motivation paragraphs - All examples use placeholder names (alice-example, bob-example, carol-example, acme-co, diana-example) - No mentions of Mem0 / Hindsight (per session-long convention) Word count: 3608 (vs 3717 before — about the same length, but more of those words now describe the NEW content rather than re-explaining prereq install). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(tutorials): enrich company-brain with real-world patterns from production deployment Mined production patterns from my own running company-brain deployment to ground the tutorial in shapes that actually work. Added four substantive sections. Part 3 (sources) gains "Two scoping models" sidebar: - Model A: separate sources with OAuth scoping (SQL-enforced isolation, right for multi-user with different AI clients per person — what the tutorial walks you through). - Model B: one source with partners/<slug>/ directory convention (simpler ops, scoping is convention-only, right when one agent serves everyone over Telegram — what I actually run in production). - Mix-and-match guidance: separate sources for the obviously-different ones AND partners/<slug>/ inside the shared source for per-person workspace. Part 7 (skills) gains "Shared rule files at the skills root" subsection: - _brain-filing-rules.md: iron-rule decision tree for where new pages belong. Every ingest skill consults it before creating a page. - _output-rules.md: output quality standards (deterministic links built from API data not LLM-composed strings, citation format, no AI-slop). - _excluded-people.md: privacy gate naming people the brain must never reference even when they appear in source material. Re-attribute or discard. The file that prevents accidental publication of things about people who aren't fair game. - _operating-rules.md, _x-ingestion-rules.md, _x-api-rules.md. - These turn into the de facto company policy for the agent. Edit one, every skill picks it up next request. NEW Part 8 "Wire Slack carefully": - Two crons, two jobs (scan every 5-15min for live signals + archive nightly for full history). - Channel-to-task-ID mapping via topic-registry.json (don't reference raw Slack channel IDs in skills; friendly names that resolve at runtime). - Deterministic links rule (LLM-composed Slack URLs hallucinate constantly; build from API data only). - Dismissed-items state so re-scans don't surface noise that was already triaged. - Per-channel scoping mirrors per-person scoping. Sensitive channels scope by OAuth client. - Names the actual production skills (slack, slack-scan, slack-archive) for scaffold reference. NEW Part 9 "Onboard each teammate yourself (the botmaster pattern)": - The load-bearing UX gate for adoption. Don't hand a teammate an OAuth credential and tell them to "try it out." That's how internal tools die. - Step 1: pre-populate their slice (partners/<their-slug>/USER.md with role/focus/priorities/preferences, 5-10 concepts that are theirs, 2-3 example brain entries that demonstrate the shape). About 20 minutes per teammate. - Step 2: walk them through 2-3 wow flows personally. A synthesis query (show the brain layer). A gap-analysis query (build trust). A write-back flow (show capture value). About 15 minutes. - Step 3: graduate to DM only after the wow moment lands. The order flips the conversion rate. - About 45 minutes per person total. Cheaper than an unadopted tool. Parts 8-12 renumbered to 10-14 to make room. Cross-references in body text checked (Part 3 ref in Part 5, Part 4 ref in operating notes, Part 5 ref in Slack section all still correct). Word count: 4938 (vs 3608 before). Still readable in one sitting per the original target; added content is all load-bearing patterns from production. Voice gates: - Zero em-dashes used as connectors (sed-replaced all 8 introduced by the additions with periods) - Zero internal jargon (no Cat N, no P@5, no RRF, no MRR) - Zero banned names (no Brad, Gessler, Wintermute, Zion, Straylight, Seibel, Caldwell, Mem0, Hindsight) - "Brain layer" / "synthesized answer" framing preserved - First-person Garry voice throughout - All examples use placeholder names (alice-example, bob-example, carol-example, diana-example, acme-co) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(tutorials): expand personal-brain Step 7 with the three real Supabase gotchas @garrytan: the personal-brain tutorial needs to surface the specific Supabase setup gotchas I hit the hard way. Rewrote Step 7 with the operational detail. Three new subsections: - 7a: Turn on pgvector. The vector extension has to be toggled in Database → Extensions before GBrain's schema migrations will run. Five seconds in the dashboard, an hour of debugging if you forget. - 7b: Use the CONNECTION POOLER string, not the direct connection. Direct is port 5432, IPv6-only. Pooler is port 6543 via pgbouncer, IPv4-compatible, survives connection storms from parallel workers. Shows the exact pooler hostname format and the gbrain config set command. - 7c: Buy the IPv4 add-on. About $4/month. Even with the pooler, some Supabase regions / Render plans hit IPv6 resolution snags. Symptom: network-unreachable errors or connect hangs in gbrain doctor. Toggle on in Project Settings → Add-ons. Saves debugging time on multiple installs. - 7d: Verify with gbrain doctor — names which of 7a / 7b / 7c to revisit if a check fails. The old "Operating note" about Supabase being the scaling bottleneck preserved at the end of the section since it's a different concern (scale, not setup). Voice gates: 0 em-dashes (verified), first-person Garry voice ("I hit the hard way"), no internal jargon, no banned names. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): rewrite "What this looks like" for someone arriving cold @garrytan: the prior version assumed too much. "Eval receipt", "top-K vector retrieval", "MemPalace", "Haiku judge", "Facts fence", "synthesis layer" — all jargon that requires reading the rest of the project to parse. A new reader bounces. New version requires zero prior context: - Sets up a universal scenario anyone gets: "you have a meeting with alice tomorrow, what do you need to know?" - Shows what a typical tool returns (a list of 5 pages with snippets) so the reader sees the gap themselves - Shows what gbrain returns (a real briefing with the open items surfaced, plus a "heads up" about what's missing from the brain) - Lets the two outputs speak for themselves, no judge scores or benchmark numbers in the body - Closes with one plain-English sentence on the difference: "Search finds the pages. The brain reads them for you and writes the answer." What got cut: - The eval-receipt path reference (means nothing to a new reader) - The Haiku judge scores (0/10 vs 9/10) — not useful out of context - The verbatim judge verdict quotes (long, internal vocabulary) - The "Facts fence", "typed-claim", "synthesis layer" feature names - The MemPalace name-drop - The link to the comprehensive benchmark page (it's still reachable from the Tutorials and benchmark sections; doesn't need to be in the intro) - The numbered "three things gbrain think did" breakdown The example content is illustrative (same scenario shape as a real production query) but stripped of the internal benchmark wrapper. Reads naturally as "imagine you're about to ask gbrain something." Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: capitalize proper-noun names in prose across README + tutorials @garrytan caught lowercase "alice" in prose — proper nouns should be capitalized. The lowercase was leaking from the slug convention (people/alice as the actual storage slug) into descriptive text. Rule applied: capitalize Alice / Bob / Carol / Diana / Acme when used as a person or company name in prose. Keep lowercase in: - Slugs and file paths (people/alice, customers/acme-co) - Code identifiers in fenced blocks where the slug IS the value - URLs and hostnames (brain.acme-co.com) - Channel-name-style references (#alice-customers) Files swept: README.md, docs/tutorials/company-brain.md, docs/tutorials/personal-brain.md, docs/tutorials/README.md. The README sweep was the primary target (Garry's actual call-out); the tutorial sweep keeps voice consistent across all the front-door docs. Hand-fixed four occurrences inside code-fenced blocks that were human comments rather than code (terminal session comments "# Terminal 1, as Alice", "# Terminal 2, as Bob", directory-tree inline comments "← Alice's customer notebook"). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(readme-hero-anchors): rotate ZeroEntropy anchor → search-vs-answer headline CI caught one expected regression after v0.40.8.1's README rewrite. The D9 hero-anchors guard required "ZeroEntropy" in the first 50 lines of README.md (the v0.36.0.0 default story). The post-rewrite hero intentionally rotated that out per Garry's "no version chatter in README" directive — ZeroEntropy still appears further down (line 211, 231, 279) but no longer in the hero. The guard's docstring explicitly handles this case: "did we deliberately rotate the headline? If yes: update the anchors here." Rotation: - Dropped: regex /ZeroEntropy|\bZE\b/ in the first 50 lines - Added: regex matching the new headline "Search gives you raw pages. GBrain gives you the answer." which is the load-bearing differentiator of the post-rewrite hero. If a future cleanup PR accidentally rewords the search-vs-answer framing, the new anchor catches it the same way the ZeroEntropy anchor caught accidental drops before. Other 4 anchors unchanged (OpenClaw + Hermes + production-number + P@5/R@5 — all still load-bearing). Updated docstring records the v0.40.8.1 rotation as the audit trail for the next time this happens. Verified: `bun test test/readme-hero-anchors.test.ts` → 5 pass / 0 fail. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): restore agent-led install path + per-client MCP guides The README install section had collapsed to three generic shapes ("agent platform", "CLI", "MCP server") that buried the load-bearing flow: paste a URL pointing at INSTALL_FOR_AGENTS.md into your agent and let it do the work. That's the path most users actually take. Restored the original three-tier structure with an explicit second tier for "install it into your existing agent" (Codex, Claude Code, Cursor), plus surfaced the per-client MCP guides individually so users see the command shape they actually need instead of one generic docs/mcp/ link. Six per-client MCP links now in the README itself: Claude Code, Cursor/Windsurf (stdio), Claude Desktop, Claude Cowork, Perplexity Computer, ChatGPT. Each carries the one-line shape that matters (claude mcp add, Settings > Integrations, OAuth 2.1, etc.). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(README): link personal-brain tutorial from the agent-install path People landing on the README without an existing OpenClaw or Hermes deployment need a starting point that walks the whole flow, not just "paste this URL into your agent." The personal-brain tutorial already covers picking a platform, deploying it, pointing it at INSTALL_FOR_AGENTS.md, and verifying the first query. Surfaced as a callout right under the agent-install snippet so the path is visible to first-time users without burying the experienced-user flow. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |