Files
gbrain/test/claw-test-cli.test.ts
T
83e55ffcdb v0.22.16 feat: gbrain claw-test — end-to-end fresh-install friction harness (#522)
* feat: hermeticity migration — every $GBRAIN_HOME write site honors the env override

configDir() in src/core/config.ts already implemented $GBRAIN_HOME as a
parent-dir override (returns <override>/.gbrain), but ~12 consumers built paths
from os.homedir() directly and bypassed it. Critically, loadConfig/saveConfig
themselves used a private getConfigDir() that ignored the env. Fixed.

Migrated every write site to gbrainPath() — fail-improve, validator-lint, cycle
lock, shell-audit, backpressure-audit, sync-failures, integrity logs,
integrations heartbeat, init pglite path, migrate-engine manifest, import
checkpoint, v0_13_1 rollback, v0_14_0 host-work. Read-side host-detection in
init.ts (~/.claude / ~/.openclaw probes) intentionally NOT migrated; that's a
v1.1 follow-up under a separate $GBRAIN_HOST_HOME override.

Adds gbrainPath(...segments) sugar plus path validation: $GBRAIN_HOME must be
absolute and contain no '..' segments (throws GbrainHomeInvalidError).

test/gbrain-home-isolation.test.ts proves write-isolation across all migrated
sites. test/migrations-v0_14_0.test.ts updated to use $GBRAIN_HOME instead of
the old HOME-swap pattern.

Closes part of the claw-test E2E harness preconditions (D13 + D21).

* feat: gbrain friction {log,render,list,summary} — agent friction reporter

Append-only JSONL writer at $GBRAIN_HOME/friction/<run-id>.jsonl. Schema is a
flat extension of StructuredAgentError (D20), one envelope shape across both
agent-emitted entries and harness-wrapped command failures. Run-id resolves
from --run-id > $GBRAIN_FRICTION_RUN_ID > 'standalone'.

Subcommands stay ≤30 LOC each; core lives in src/core/friction.ts (writer +
reader + renderer + redactor). render --redact (default for md output) strips
\$HOME / \$CWD to placeholders so reports paste safely in PRs/issues.

Severity: confused | error | blocker | nit. Kind: friction | delight (D7) |
phase-marker | interrupted. Readers tolerate malformed lines (skip + warn).

40 unit tests; this is the channel the claw-test harness writes to and that
agents emit through during live-mode runs.

* feat: gbrain claw-test — end-to-end fresh-install friction harness

Two modes: scripted (CI gate, no agent) and --live (real agent subprocess).
Phases: setup → install_brain (gbrain init --pglite) → import (--no-embed) →
query → extract all --source fs → verify (gbrain doctor --json, asserts
status==='ok' and progress.jsonl phase coverage).

AgentRunner interface + registry — interface stays narrow (detect, invoke,
optional postInstallHook). v1 ships only OpenClawRunner; the registry pattern
lets v1.1 land hermes/codex as ~50-line additions without refactoring callers.
OpenClaw invocation: 'openclaw agent --local --agent <name> --message <brief>'
matching test/e2e/skills.test.ts (NOT --prompt-file, which doesn't exist).

transcript-capture: spawns child with piped stdio, async-drains via
fs.createWriteStream + 'drain' events so 256KB+ bursts don't stall the child
(D17 backpressure). Writes <run>/transcript.jsonl with schema_version + ts +
channel + byte_offset + bytes_b64. Friction entries' transcript_offset field
references byte offsets here so render --transcripts can resolve back.

progress-tail: parses gbrain's --progress-json events out of child stderr.
Phase verification asserts each scenario.expected_phases entry (dotted names
like import.files, extract.links_fs, doctor.db_checks) saw at least one event
from the actual command — proves the COMMAND ran, not that the agent obeyed
prompts.

seed-pglite: ~50 LOC SQL replay primitive for the upgrade-from-v0.18 scenario.
Existing migration helpers (test/e2e/helpers.ts) are Postgres-only; PGLite has
no equivalent. seedPglite opens a fresh PGLite, executes each statement
individually (errors name the failing one), then disconnects so gbrain init
can take over and walk forward.

53 unit tests covering registry selection, runner detection, multi-byte UTF-8
chunk-boundary safety, PIPE buffer drain, scenario load+validate, progress
event parsing, and SQL splitter.

* feat: claw-test scenario fixtures + friction-protocol skills convention

Two scenarios ship in v1 — fresh-install and upgrade-from-v0.18. Each is a
self-contained directory: brain/ (markdown pages), BRIEF.md (live-mode prompt),
expected.json (scripted-mode assertions), scenario.json (kind, expected_phases,
optional from_version + seed paths). Schema is owned by src/core/claw-test/
scenarios.ts.

upgrade-from-v0.18 ships scaffolded — seed/dump.sql is the v1.1 follow-up
(needs a real v0.18-shape PGLite dump; seed/README.md documents the gen
procedure). The harness gracefully no-ops the seed phase when dump.sql is
absent.

skills/_friction-protocol.md is a cross-cutting convention skill (like
_brain-filing-rules.md). Tells agents when to call gbrain friction log and how
to choose severity. Skills the claw-test exercises will gain a > Convention:
callout pointing here in a v1.1 sweep.

13 unit tests for the scenario loader + 'shipped scenarios load cleanly' for
both.

* feat: register gbrain claw-test + gbrain friction; CLAUDE.md + llms sync

Wires both commands into src/cli.ts CLI_ONLY allow-list and adds dispatch
in handleCliOnly so neither command requires a brain engine connection.

CLAUDE.md gains entries for src/commands/{friction,claw-test}.ts +
src/core/claw-test/ + skills/_friction-protocol.md, and a Commands section
listing all 8 new gbrain claw-test ... and gbrain friction ... invocations
with the v0.23 marker. Documents the GBRAIN_HOME write-isolation contract
and the v1 caveat (read-side host-fingerprint detection deferred to v1.1).
llms.txt + llms-full.txt regenerated via 'bun run build:llms' so the
committed generator-output gate passes.

test/e2e/claw-test.test.ts is the scripted-mode E2E. Builds a tiny shim that
delegates to 'bun run src/cli.ts' (NOT bun --compile, which doesn't bundle
PGLite's runtime assets), points the harness at it via GBRAIN_BIN_OVERRIDE,
runs --scenario fresh-install end-to-end. Asserts exit 0, zero error/blocker
friction. Includes a deliberate-break test that proves the friction signal
fires when a phase command rejects.

test/claw-test-cli.test.ts covers shipped-scenario load + agent registry +
OpenClawRunner detection (relative-path / .. / missing-bin guards) + the
GBRAIN_FRICTION_RUN_ID env handoff between harness and friction CLI.

Closes the v0.23 claw-test E2E feature.

* chore: bump version and changelog (v0.24.0)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(tests): typecheck failures + spawnWithCapture timeout headroom in CI

Three CI fixes after PR #522 landed:

1. test/agent-runner.test.ts:89 — UnavailableRunner.invoke() returns
   Promise<void> by default but the AgentRunner contract requires
   Promise<InvokeResult>. Annotate the throw-only invoke explicitly so tsc
   sees the contract is satisfied (the throw makes the body unreachable as
   far as the return type is concerned).

2. test/seed-pglite.test.ts — bun:test signature is test(name, fn, timeoutMs:
   number), not test(name, opts: {timeout}, fn). The {timeout: 30_000} object
   form was a guess that tsc on bun 1.3.13 rejects. Move the 30s cap to the
   trailing positional number arg on each PGLite-using test.

3. test/transcript-capture.test.ts — `spawnWithCapture > timeout fires
   SIGTERM/SIGKILL` blew the 10s outer cap on the GitHub runner. Two fixes:
   (a) use `exec sleep` so the child we spawn IS sleep — SIGTERM goes
   directly to it, no `/bin/sh` fork-vs-exec process-group ambiguity that
   could orphan the sleep and force the SIGKILL grace path. (b) bump outer
   cap to 30s for headroom even when the runner is slow and SIGKILL after
   the 5s grace is what actually ends the child.

* chore: rebump to v0.22.16 (next free 0.22.x patch slot per queue)

PR #506 claims v0.22.15, PR #521 claims v0.22.10, intermediate slots
(.11/.12/.13/.14) are claimed by other open PRs. v0.22.16 is the next
clean PATCH slot. v0.23.0 is claimed by PR #462 so MINOR isn't free.
This release fits the 0.22.x train; v0.23.0 lands when #462 ships.

Updates VERSION, package.json, CHANGELOG.md header, TODOS.md follow-up
labels. Code is unchanged.

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-29 23:46:36 -07:00

166 lines
6.1 KiB
TypeScript

/**
* gbrain claw-test CLI dispatch tests.
*
* These tests exercise the harness's argument parsing, scenario loading,
* agent registry resolution, and friction-report path. They do NOT spawn
* real gbrain commands (no built binary in CI yet); the canonical scripted
* E2E that walks `gbrain init → import → query → extract → verify` lives
* in test/e2e/claw-test.test.ts and gates on a built binary.
*/
import { describe, test, expect, beforeEach, afterEach } from 'bun:test';
import { mkdtempSync, rmSync, existsSync, readFileSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';
import { runFriction } from '../src/commands/friction.ts';
import { listScenarios, loadScenario } from '../src/core/claw-test/scenarios.ts';
import {
registerAgentRunner, resolveAgentRunner, listRegisteredAgents,
_resetRegistryForTests,
type AgentRunner, type DetectResult, type InvokeOpts, type InvokeResult,
} from '../src/core/claw-test/agent-runner.ts';
let tmp: string;
const ORIG_HOME = process.env.GBRAIN_HOME;
beforeEach(() => {
tmp = mkdtempSync(join(tmpdir(), 'claw-test-cli-'));
process.env.GBRAIN_HOME = tmp;
_resetRegistryForTests();
});
afterEach(() => {
process.env.GBRAIN_HOME = ORIG_HOME;
rmSync(tmp, { recursive: true, force: true });
});
describe('shipped scenarios are loadable', () => {
test('default fixtures root contains both v1 scenarios', () => {
delete process.env.GBRAIN_CLAW_SCENARIOS_DIR;
const names = listScenarios();
expect(names).toContain('fresh-install');
expect(names).toContain('upgrade-from-v0.18');
});
test('fresh-install has expected_phases', () => {
delete process.env.GBRAIN_CLAW_SCENARIOS_DIR;
const cfg = loadScenario('fresh-install');
expect(cfg.expectedPhases).toContain('import.files');
expect(cfg.expectedPhases).toContain('extract.links_fs');
expect(cfg.expectedPhases).toContain('doctor.db_checks');
});
test('upgrade-from-v0.18 declares from_version', () => {
delete process.env.GBRAIN_CLAW_SCENARIOS_DIR;
const cfg = loadScenario('upgrade-from-v0.18');
expect(cfg.kind).toBe('upgrade');
expect(cfg.fromVersion).toBe('0.18.0');
expect(cfg.seedRelative).toBe('seed');
});
});
describe('agent registry — fake-runner integration', () => {
test('a fake runner can be registered, resolved, and detect/invoke called', async () => {
let invokeCount = 0;
class FakeRunner implements AgentRunner {
readonly name = 'fake';
async detect(): Promise<DetectResult> { return { available: true, binPath: '/usr/bin/fake' }; }
async invoke(_opts: InvokeOpts): Promise<InvokeResult> {
invokeCount++;
return { exitCode: 0, durationMs: 1 };
}
}
registerAgentRunner('fake', () => new FakeRunner());
expect(listRegisteredAgents()).toContain('fake');
const r = resolveAgentRunner('fake');
const detected = await r.detect();
expect(detected.available).toBe(true);
const result = await r.invoke({
cwd: tmp,
brief: 'test',
env: {},
timeoutMs: 1000,
transcriptSink: { write: () => {}, nextOffset: () => 0, close: async () => {} },
});
expect(result.exitCode).toBe(0);
expect(invokeCount).toBe(1);
});
test('resolveAgentRunner with unknown name throws with registered list', () => {
registerAgentRunner('alpha', () => ({} as AgentRunner));
expect(() => resolveAgentRunner('unknown')).toThrow(/registered: alpha/);
});
});
describe('friction CLI integrates with harness run-id env', () => {
test('GBRAIN_FRICTION_RUN_ID populates harness-style run-ids', () => {
process.env.GBRAIN_FRICTION_RUN_ID = 'claw-test-20260428-fake-abcd1234';
try {
const code = runFriction(['log', '--phase', 'install', '--message', 'simulated harness write']);
expect(code).toBe(0);
const expectedFile = join(tmp, '.gbrain', 'friction', 'claw-test-20260428-fake-abcd1234.jsonl');
expect(existsSync(expectedFile)).toBe(true);
const raw = readFileSync(expectedFile, 'utf-8');
const entry = JSON.parse(raw.split('\n')[0]);
expect(entry.run_id).toBe('claw-test-20260428-fake-abcd1234');
expect(entry.message).toBe('simulated harness write');
} finally {
delete process.env.GBRAIN_FRICTION_RUN_ID;
}
});
});
describe('OpenClawRunner detection (reliable on box without openclaw)', () => {
test('detect returns unavailable when OPENCLAW_BIN missing', async () => {
const orig = process.env.OPENCLAW_BIN;
delete process.env.OPENCLAW_BIN;
try {
const { OpenClawRunner } = await import('../src/core/claw-test/runners/openclaw.ts');
const r = new OpenClawRunner();
const d = await r.detect();
// Either unavailable, or available if openclaw IS on PATH for the dev — both states are valid.
// We only assert the contract shape.
expect(typeof d.available).toBe('boolean');
if (!d.available) {
expect(typeof d.reason).toBe('string');
} else {
expect(d.binPath?.startsWith('/')).toBe(true);
}
} finally {
if (orig !== undefined) process.env.OPENCLAW_BIN = orig;
}
});
test('detect rejects relative OPENCLAW_BIN', async () => {
const orig = process.env.OPENCLAW_BIN;
process.env.OPENCLAW_BIN = 'relative/openclaw';
try {
const { OpenClawRunner } = await import('../src/core/claw-test/runners/openclaw.ts');
const r = new OpenClawRunner();
const d = await r.detect();
expect(d.available).toBe(false);
expect(d.reason).toMatch(/absolute/);
} finally {
if (orig !== undefined) process.env.OPENCLAW_BIN = orig;
else delete process.env.OPENCLAW_BIN;
}
});
test("detect rejects '..' segments in OPENCLAW_BIN", async () => {
const orig = process.env.OPENCLAW_BIN;
process.env.OPENCLAW_BIN = '/tmp/foo/../bar';
try {
const { OpenClawRunner } = await import('../src/core/claw-test/runners/openclaw.ts');
const r = new OpenClawRunner();
const d = await r.detect();
expect(d.available).toBe(false);
expect(d.reason).toMatch(/'\.\.' segments/);
} finally {
if (orig !== undefined) process.env.OPENCLAW_BIN = orig;
else delete process.env.OPENCLAW_BIN;
}
});
});