mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-30 19:49:14 +00:00
* feat(dims): OpenAI text-embedding-3 Matryoshka range validation (D13) dimsProviderOptions now fail-loud at the embed boundary when the configured embedding_dimensions is outside the model's native range (1..1536 for -small, 1..3072 for -large). Paste-ready fix hint in the AIConfigError.fix field. Closes the silent-HTTP-400 path that would have bit OpenAI-fallback users on v0.36.0.0 ZE-default installs. 16 new test cases in test/ai/dims-openai.test.ts pinning the contract across native-openai and openai-compatible adapter paths. * feat(ai): flip defaults to ZeroEntropy zembed-1 1280d + zerank-2 reranker Default embedding model is now zeroentropyai:zembed-1 at 1280d via Matryoshka. Real-corpus benchmark: 2.2x faster than OpenAI, 2.6x cheaper at regular pricing, wins 11/20 head-to-head queries. 1280 is the closest valid ZE Matryoshka step to the prior OpenAI 1536d default (valid set: 2560/1280/640/320/160/80/40). 1024 (Voyage's step) is NOT on ZE's list — pinned by AIConfigError fail-loud in dims.ts. balanced mode bundle now defaults reranker_enabled=true. zerank-2 reshuffles 60% of top-1 results in benchmarks. Missing-key fail-open contract in src/core/search/rerank.ts handles unauthenticated cases. Opt out with: gbrain config set search.reranker.enabled false Existing tests updated (gateway.test.ts, search-mode.test.ts) and a new test/balanced-reranker-default.test.ts (10 cases) pins the fail- open invariants. * feat(retrieval-upgrade): RetrievalUpgradePlanner + interactive prompt UX New src/core/retrieval-upgrade-planner.ts is the consolidated planner that computes the brain's pending retrieval-upgrade work (chunker bumps + ZE switch) in one pass and applies the schema transition + config updates atomically. Tagged-union ApplyResult enum (D15): 'applied' | 'skipped_already_ applied' | 'skipped_no_work' | 'declined' | 'planned' | 'failed'. No string-parsing reasons. Three config keys (D12): ze_switch_prompt_shown (UI state), ze_switch_requested (user intent), ze_switch_applied (work done). Plus ze_switch_previous_snapshot (JSON, full prior config for --undo per D16) and ze_switch_declined_at (90-day re-ask window). Schema transition (D18) is atomic: DROP indexes + ALTER COLUMN + CREATE INDEX inside a single engine.transaction(). HNSW recreation is part of the same transaction — no silent slow-search window. C3 eligibility logic: ze_switch_offered iff NOT on ZE + NOT declined recently + NOT applied + (legacy default OR >100 pages). C4 cost math: MAX(chunker_pending, dim_pending) not SUM — one re-embed pass invalidates both surfaces simultaneously. New src/core/retrieval-upgrade-prompt.ts wires the planner to a TTY-only interactive prompt with two-line cost split (D10) and privacy callout for the reranker flip. Tests: test/retrieval-upgrade-planner.test.ts (24 cases) pins the state machine. test/asymmetric-encoding-contract.test.ts (6 cases) pins D17: search read path uses gateway.embedQuery() not embed(), asserted via __setEmbedTransportForTests mock. * feat(cli): gbrain ze-switch — manual lever for the ZE switch New gbrain ze-switch CLI with --dry-run, --json, --resume, --force, --undo, --non-interactive, --confirm-reembed, --ignore-missing-key flags. Mirrors the upgrade prompt's UX symmetry: --undo presents a cost-warning before re-embedding back to the prior width. src/cli.ts: dispatch case + CLI_ONLY entry. ze-switch owns its own engine lifecycle (mirrors the doctor pattern). test/ze-switch-cli.test.ts (11 cases): --help, --dry-run, --json, --non-interactive, --ignore-missing-key, --resume, --undo, --confirm-reembed. Uses captureExit harness to test process.exit() paths without breaking the test process. * feat(doctor): ze_embedding_health + embedding_width_consistency checks Two new doctor checks (D-A5): ze_embedding_health: when embedding_model starts with zeroentropyai:, verify ZEROENTROPY_API_KEY is set (env or config). Paste-ready setup hint with the signup URL on failure. embedding_width_consistency: cross-check that the configured embedding_dimensions matches the actual vector(N) column width on content_chunks.embedding. Catches the half-applied switch state (schema migrated but config write crashed) with a paste-ready gbrain ze-switch --resume hint. Wired into runDoctor between reranker_health and the existing sync_freshness checks. Both checks gracefully no-op on non-ZE embedding configs. test/doctor-ze-checks.test.ts (8 cases) pins both checks across happy + missing-key + missing-config + drift paths. Uses withEnv() helper to clear ZEROENTROPY_API_KEY for the no-key path so tests are hermetic against contributor env state. test/e2e/v0_28_5-fix-wave.test.ts + test/openai-compat-multimodal.test.ts: updated to explicit-configure the gateway when the test depends on specific dims that diverge from the v0.36.0.0 default (1280d). * docs: README zero-based rewrite (884 -> 139 lines) + new docs files Strip 4 months of accreted "New in v0.X.Y" hero blocks and reorganize around what gbrain does today. 33 H2s -> 8. The Commands section (136 lines duplicating gbrain --help) moved out; the 6-table skills enumeration collapsed to a one-paragraph capability description with a link to skills/RESOLVER.md. Hero retains load-bearing facts: OpenClaw + Hermes credit, production numbers (17,888 pages / 4,383 people / 723 companies), BrainBench numbers (P@5 49.1% / R@5 97.9% / +31.4 lift), ZE comparison numbers, 30-min install claim. Adds one paragraph announcing the v0.36.0.0 ZE default with the explicit gbrain config set escape for OpenAI/Voyage users. New files: - docs/INSTALL.md: every install path consolidated (agent platform, CLI standalone, MCP server). Thin-client mode covered. - docs/architecture/RETRIEVAL.md: why the hybrid + graph stack works. BrainBench numbers, why each strategy alone fails, the source-aware ranking + intent classification + multi-query expansion story. - docs/ethos/ORIGIN.md: origin story lifted from the old README so the front door stays factual + concrete. test/readme-hero-anchors.test.ts (5 cases) is the D9 regression guard. Five load-bearing strings: OpenClaw, Hermes, ZE, production-numbers regex, P@5/R@5. Light anchors that let voice/ structure evolve but block accidental loss of headline facts. scripts/check-test-real-names.sh: allowlist entries for OpenClaw + Hermes literals in the anchor test (it explicitly asserts those strings appear in README). * chore: bump version and changelog (v0.36.0.0) ZeroEntropy as the new default for embedding (zembed-1 at 1280d via Matryoshka) and reranker (zerank-2 cross-encoder, on by default in balanced mode bundle). README zero-based rewrite (884 -> 139 lines). 3 new docs files. Two new doctor checks. New gbrain ze-switch CLI with --undo for symmetric reversibility. skills/migrations/v0.36.0.0.md tells the agent how to surface the retrieval-upgrade prompt post-upgrade. llms-full.txt regenerated via bun run build:llms. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(docs): scrub Wintermute from RETRIEVAL.md per privacy rule * chore: rebump version 0.36.0.0 → 0.36.2.0 (queue collision) Three open PRs were claiming v0.36.0.0 (#1130 skillpack, #1139 hindsight, #1136 this PR). Ship-aware queue allocator says this branch lands at v0.36.2.0. Trio audit: VERSION 0.36.2.0 package.json 0.36.2.0 CHANGELOG ## [0.36.2.0] - 2026-05-17 Updates: VERSION, package.json, CHANGELOG header + body refs, README "New default in v0.36.2.0" announcement + credit line, skills/migrations/v0.36.0.0.md renamed to v0.36.2.0.md with frontmatter + body refs updated. llms-full.txt regenerated. * fix(test): pin gateway dim=1536 in cross-file-stateful PGLite tests CI shard 1 reported 10 failures across `query-cache.test.ts` (6) and `consolidate-valid-until.test.ts` (4). Both files hardcode 1536-dim vectors but rely on `PGLiteEngine.initSchema()` to size `vector(__EMBEDDING_DIMS__)` at the right width. Root cause: v0.36.2.0 flipped DEFAULT_EMBEDDING_DIMENSIONS from 1536 to 1280 (ZE Matryoshka step). The gateway module is process-singleton; when ANOTHER test file in the same shard's bun-test process configures the gateway before us, `pglite-engine.ts:216` reads `getEmbeddingDimensions() === 1280` and sizes the schema columns at vector(1280). The hardcoded 1536-dim INSERTs then fail with "expected 1280 dimensions, not 1536". Locally these tests pass in isolation because the gateway falls back through the try/catch at pglite-engine.ts:218 (1536 default). CI runs multiple test files in one process, so cross-file state poisons the schema width. Fix: explicit `resetGateway()` + `configureGateway({embedding_dimensions: 1536, ...})` at the top of `beforeAll`, plus `resetGateway()` in `afterAll`. Pins the schema width regardless of cross-file state. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
158 lines
6.5 KiB
Bash
Executable File
158 lines
6.5 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# CI guard: fail if any test fixture references a real person's name.
|
|
#
|
|
# CLAUDE.md's "Privacy rule" section is unambiguous: never reference real
|
|
# people, companies, funds, or private agent names in any public-facing
|
|
# artifact. Tests are checked-in code distributed with every release and
|
|
# indexed by GitHub search. This guard catches the patterns the rule names.
|
|
#
|
|
# Design (post-Codex F4 review):
|
|
# - Banned names: exact-string allowlist of known real identifiers. Adding
|
|
# a name when CLAUDE.md flags one is a one-line edit.
|
|
# - Banned emails: specific addresses that identify real contacts. NOT a
|
|
# broad corporate-email regex — those would catch legitimate fixture
|
|
# domains in billing/auth tests (`customer@stripe.com` etc.).
|
|
# - Allowlist: exact "file:offending-string" pairs that are intentional
|
|
# and pre-existing (e.g., the user's own email is not a "contact").
|
|
#
|
|
# Scope: test/**/*.test.ts only. Historical CHANGELOG entries, doc examples,
|
|
# and skill READMEs each have their own scrub status and are out of scope
|
|
# for this guard.
|
|
#
|
|
# Usage: scripts/check-test-real-names.sh
|
|
# Exit: 0 clean, 1 banned reference found, 2 setup error (rg + grep missing).
|
|
|
|
set -euo pipefail
|
|
|
|
ROOT="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
|
|
cd "$ROOT"
|
|
|
|
# Banned real-name strings (matched as whole words, case-insensitive).
|
|
# Add an entry when CLAUDE.md flags a new real-person name.
|
|
BANNED_NAMES=(
|
|
'Diana' # Diana Hu, named in CLAUDE.md privacy example
|
|
'Wintermute' # private OpenClaw fork name (CLAUDE.md rule)
|
|
'Hermes' # downstream agent fork name
|
|
'Technium' # real GP handle
|
|
'McGrew' # ex-OpenAI exec
|
|
'YC Labs' # internal team name
|
|
)
|
|
|
|
# Banned specific email addresses. NOT a generic corporate-email regex —
|
|
# those would catch legitimate fixture domains in billing/auth tests
|
|
# (`customer@stripe.com`, `account@openai.com` etc).
|
|
BANNED_EMAILS=(
|
|
'diana@ycombinator.com'
|
|
)
|
|
|
|
# Exact "file:offending-string" pairs that are intentional and pre-existing.
|
|
# These pre-date the rule, the file's own author confirmed the use, the
|
|
# string identifies the user themselves (not a contact), OR the reference
|
|
# is structural (e.g., a regression test that ASSERTS the banned name does
|
|
# NOT appear in production code — the name MUST be in the test file as a
|
|
# literal).
|
|
ALLOWLIST=(
|
|
"test/writer.test.ts:garry@ycombinator.com" # user's own email — CLAUDE.md rule does not apply
|
|
"test/integrations.test.ts:Wintermute" # regex pattern in personal-info filter test (structural)
|
|
"test/recency-decay.test.ts:Wintermute" # regression-prevention test asserting wintermute is absent (structural)
|
|
"test/scripts/check-proposal-pii.test.ts:Wintermute" # privacy-guard test asserting docs/proposals/ rejects wintermute (structural; same meta-rule exception as check-privacy.sh)
|
|
"test/scripts/check-proposal-pii.test.ts:WINTERMUTE" # case-insensitive sentinel literal for the same privacy-guard test
|
|
"test/serve-stdio-lifecycle.test.ts:Hermes" # comment naming a downstream-agent scenario — pre-existing, low signal
|
|
"test/extract.test.ts:Hermes" # markdown-link extraction test fixture — pre-existing, ambiguous (Greek god vs fork)
|
|
"test/readme-hero-anchors.test.ts:Hermes" # v0.36.0.0 D9 anchor test — asserts README mentions Hermes as a credit
|
|
"test/readme-hero-anchors.test.ts:OpenClaw" # v0.36.0.0 D9 anchor test — asserts README mentions OpenClaw as a credit
|
|
# v0.36.0.0: skillpack-harvest privacy linter tests structurally
|
|
# require the literal "Wintermute" to verify the linter catches it.
|
|
# Same meta-rule exception as integrations.test.ts and the proposal-pii
|
|
# privacy guard test above.
|
|
"test/skillpack-harvest.test.ts:Wintermute"
|
|
"test/skillpack-harvest-lint.test.ts:Wintermute"
|
|
"test/e2e/skillpack-flow.test.ts:Wintermute"
|
|
)
|
|
|
|
# Build the combined regex. Names matched as whole words (\b), emails matched
|
|
# literally with dot escapes.
|
|
PATTERN_PARTS=()
|
|
for n in "${BANNED_NAMES[@]}"; do
|
|
# Escape any regex metacharacters in the name (defensive — most are bare
|
|
# words but YC Labs has a space).
|
|
escaped="${n//./\\.}"
|
|
escaped="${escaped// /\\s}"
|
|
PATTERN_PARTS+=("\\b${escaped}\\b")
|
|
done
|
|
for e in "${BANNED_EMAILS[@]}"; do
|
|
escaped="${e//./\\.}"
|
|
PATTERN_PARTS+=("${escaped}")
|
|
done
|
|
|
|
# Join with |.
|
|
IFS='|' eval 'PATTERN="${PATTERN_PARTS[*]}"'
|
|
|
|
# Find tool.
|
|
if command -v rg >/dev/null 2>&1; then
|
|
matches="$(rg -niH --no-heading -t ts "$PATTERN" test/ 2>/dev/null || true)"
|
|
elif command -v grep >/dev/null 2>&1; then
|
|
matches="$(grep -rniE --include='*.test.ts' "$PATTERN" test/ 2>/dev/null || true)"
|
|
else
|
|
echo "check-test-real-names: ERROR: neither rg nor grep available." >&2
|
|
exit 2
|
|
fi
|
|
|
|
if [ -z "$matches" ]; then
|
|
exit 0
|
|
fi
|
|
|
|
# Apply allowlist. Each line is "file:lineno:content"; check whether
|
|
# "file:<needle>" appears in ALLOWLIST for any needle in BANNED_EMAILS+NAMES
|
|
# that matches the content.
|
|
filtered=""
|
|
while IFS= read -r line; do
|
|
[ -z "$line" ] && continue
|
|
# Extract filename and content (everything after second :).
|
|
file="${line%%:*}"
|
|
rest="${line#*:}"
|
|
# rest is "lineno:content" — strip lineno.
|
|
content="${rest#*:}"
|
|
|
|
matched_needle=""
|
|
for needle in "${BANNED_EMAILS[@]}" "${BANNED_NAMES[@]}"; do
|
|
if echo "$content" | grep -qi -- "$needle"; then
|
|
matched_needle="$needle"
|
|
break
|
|
fi
|
|
done
|
|
|
|
allow_key="${file}:${matched_needle}"
|
|
allowed=0
|
|
for allow_entry in "${ALLOWLIST[@]}"; do
|
|
if [ "$allow_entry" = "$allow_key" ]; then
|
|
allowed=1
|
|
break
|
|
fi
|
|
done
|
|
|
|
if [ "$allowed" = "0" ]; then
|
|
filtered+="${line}"$'\n'
|
|
fi
|
|
done <<< "$matches"
|
|
|
|
if [ -z "$filtered" ]; then
|
|
exit 0
|
|
fi
|
|
|
|
echo "check-test-real-names: banned real-name references found in test/ fixtures." >&2
|
|
echo "" >&2
|
|
echo "$filtered" >&2
|
|
echo "" >&2
|
|
echo "Fix: replace with canonical placeholders per CLAUDE.md 'Name mapping' table." >&2
|
|
echo " alice-example / @alice-example for people" >&2
|
|
echo " bob-example / charlie-example for additional people" >&2
|
|
echo " alice@example.com for emails (example.com is RFC 6761 reserved)" >&2
|
|
echo " acme-example / widget-co for companies" >&2
|
|
echo " fund-a / fund-b for funds" >&2
|
|
echo " a-team / agent-fork for teams / OpenClaw forks" >&2
|
|
echo "" >&2
|
|
echo "If the match is intentional (e.g., the user's own identifier, not a contact)," >&2
|
|
echo "add an exact 'file:string' entry to ALLOWLIST in scripts/check-test-real-names.sh." >&2
|
|
exit 1
|