Files
gbrain/test/scripts/run-unit-parallel.test.ts
T
d97f159793 v0.26.4 test: parallel unit-test loop (12x speedup, failure-first logging) (#605)
* test: parallel unit-test wrapper + failure-first logging (commit 1/8)

Lay foundation for v0.26.4 parallel test loop:

- scripts/run-unit-parallel.sh: spawns N shards (default min(8, cpu_count))
  via run-unit-shard.sh, captures per-shard logs, post-shard single-writer
  failure-log aggregation at .context/test-failures.log, 10s heartbeat to
  stderr, per-shard 600s timeout (gtimeout/timeout/bg-pid fallback chain),
  loud final banner with absolute path + tail-30 of failures, summary file
  for at-a-glance status. Single writer eliminates concurrent-write hazards
  on the failure log.
- scripts/run-serial-tests.sh: discovers *.serial.test.ts files (concurrency-
  unsafe by design), runs them with --max-concurrency=1. Invoked after the
  parallel pass.
- scripts/run-unit-shard.sh: now accepts --max-concurrency=N (forwarded to
  bun test); --dry-run-list moved into argv parsing alongside; excludes
  *.serial.test.ts in addition to *.slow.test.ts.
- bunfig.toml: trim stale comment about typecheck-chained timeout.
- .gitignore: add .context/ (Conductor workspace artifacts directory; the
  failure log + summary + per-shard logs all live here).

No package.json changes yet (commit 2). No test reorganization yet
(commits 4-7).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: split package.json scripts; bun run test = parallel fast loop (commit 2/8)

Per Codex Tension #4 (verify scope), distinguish three tiers cleanly:

- `bun run test` = fast loop, file-level parallel fan-out via the new wrapper
  (scripts/run-unit-parallel.sh). No pre-checks, no typecheck, no wasm
  compile in the hot path. ~15s of pre-test gates removed.
- `bun run verify` = CI's authoritative gate set: check:jsonb +
  check:progress + check:wasm + typecheck. Matches what
  .github/workflows/test.yml runs on shard 1, no scope drift. The 4
  checks not in CI (privacy, no-legacy-getconnection, trailing-newline,
  exports-count) move to `bun run check:all` for opt-in local use.
- `bun run test:full` = verify + parallel + slow + smart e2e (runs e2e
  only if DATABASE_URL is set; else loud skip notice to stderr per Open
  Item #7). The local equivalent of "everything CI runs."

Adds `bun run test:serial` for the *.serial.test.ts subset (concurrency-
unsafe files run with --max-concurrency=1).

Bumps VERSION + package.json to 0.26.4. Both move together per the CI
version-gate contract in CLAUDE.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: fix-wave for parallel wrapper + tighten privacy gate (commit 3/5)

Wave: makes the new wrapper actually green and tightens the CI gate it
exposed.

Wrapper bug fixes (scripts/run-unit-parallel.sh):
- grep_count helper: avoids the `grep -c | echo 0` double-output bug
  where 0 matches yields a 2-line "0\n0" string and breaks arithmetic.
- bun_summary_count helper: parses Bun's actual end-of-shard summary
  format (`N pass` / `N fail` / `N skip`), not the per-test markers
  (which are `✓` / `(fail)`, never `(pass)` / `(skip)`).
- Heartbeat now reads `^\s+✓` (Bun's per-test pass marker) for live
  progress mid-run; final summary still uses the summary-line counts
  for accuracy.

Privacy gate tightening:
- Move scripts/check-privacy.sh into `bun run verify` (was previously
  only in the now-removed `bun run test` chain). Without this, after
  commit 2 the privacy check ran in nothing automatic.
- .github/workflows/test.yml now calls `bun run verify` instead of
  inlining the gate list. Single source of truth for "what's the ship
  gate." This is what verify == CI was supposed to mean per Codex T#4.
- Pre-existing `Wintermute` references in src/core/mounts-cache.ts:6
  and :324 caught by the now-running gate; replaced with `your OpenClaw`
  per CLAUDE.md privacy rule (verify gate now passes on master HEAD).
- test/privacy-script-wired.test.ts updated: regression guard now
  asserts verify includes check:privacy AND that test.yml runs
  `bun run verify`, replacing the obsolete "test script includes
  check-privacy.sh" assertion.

Quarantine 2 cross-file-contention flakes:
- test/brain-registry.test.ts: 28 tests pass alone (41ms); 1 test
  ("empty/null/undefined id routes to host") fails when run alongside
  other files in the same shard. Renamed → *.serial.test.ts so it
  runs in scripts/run-serial-tests.sh's serial pass after the parallel
  pass completes.
- test/reconcile-links.test.ts: 6 tests pass alone (1s); a beforeEach
  hook times out (~896s) under cross-file contention. Same treatment.

Both flakes are bun-process-level shared-state leaks (PGLite singletons
or top-level imports). Fixing them properly is the v0.27.0+ intra-file
parallelism project (TODO P0 — see commit 5).

Measurement after this commit:
  bun run test = 94s (was 18 min sequential)
  3639 pass, 0 fail, 0 skip across 8 parallel shards + 34 serial tests
  Failure-log + heartbeat + summary all working

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: regression tests for parallel wrapper + serial-test contracts (commit 4/5)

Three regression suites pin the v0.26.4 contracts. Without these,
future refactors of the wrapper or shard scripts could silently
regress the work in commits 1-3.

test/scripts/run-unit-shard.test.ts (4 cases — gap b):
- Asserts the unit-shard `--dry-run-list` output excludes every
  *.slow.test.ts and *.serial.test.ts file, plus the test/e2e/ subtree.
- Catches a future `find` expression that drops one of the `-not -name`
  clauses and silently un-quarantines slow/serial files into the
  parallel pass.

test/scripts/serial-files.test.ts (3 cases — gap e):
- Every checked-in *.serial.test.ts (via `git ls-files`) is listed by
  scripts/run-serial-tests.sh's `--dry-run-list`.
- The script's source contains `bun test --max-concurrency=1` (the
  serial-pass guarantee that quarantined files don't run intra-file
  concurrent and reintroduce the contention they were quarantined for).
- Disjoint set: a file is never in both the unit-shard list AND the
  serial list — pins the carve-out contract.

test/scripts/run-unit-parallel.test.ts (6 cases — gaps a + d):
- Exit-code propagation (a): wrapper exits non-zero when ANY shard
  has a failing test; exits zero when all pass. The hardest contract
  to silently break in a fan-out wrapper (`for ... &; wait` returns
  the LAST child's status, not any failure's).
- Failure-log contract (d): on failure, .context/test-failures.log
  exists, is non-empty, contains the `--- shard N:` prefix and the
  failing test's describe text. Stderr banner contains the absolute
  log path. On success, the log is cleared (no stale content).
- Summary file format: `shard N/M: pass=X fail=Y skip=Z rc=W` per
  shard, machine-parseable for future tooling.

The wrapper test runs against a 4-file tempdir (3 pass + 1 fail) so
it executes in ~500ms; spawning the wrapper against the real test
suite would take ~90s and isn't worth the cost in a regression suite.

All 13 cases pass on first run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(v0.26.4): testing tier docs + CHANGELOG + intra-file P0 TODO (commit 5/5)

Closes the v0.26.4 ship.

CLAUDE.md Testing section rewritten:
- New tier table: test (fast loop, 85s) / verify (CI gates, 12s) /
  test:full (everything local) / test:slow / test:serial / test:e2e /
  check:all. Each row names its scope, wallclock, and when to use.
- Intentional CI vs local divergence section: CI matrix (test-shard.sh,
  hash-bucketed, includes slow) vs local fast loop (run-unit-shard.sh,
  round-robin, excludes slow + serial). Codex correctly flagged that a
  parity test would always fail by design — this is the documentation
  that explains why.
- Failure-first logging contract: .context/test-failures.log format,
  stderr banner, summary file, wedge handling.
- File taxonomy: *.test.ts / *.slow.test.ts / *.serial.test.ts /
  test/e2e/. Names the two currently-quarantined files and points at the
  intra-file P0 TODO for the proper fix.

CHANGELOG.md `## [0.26.4]` entry per voice rules:
- Two-line headline: "bun run test finishes in 85 seconds. Was 18
  minutes." + failure-log directive.
- Lead paragraph names what shipped and why.
- Numbers-that-matter table: BEFORE / AFTER / Δ for wallclock, pre-test
  gates, failure visibility, shards, pipe-survival.
- "What this means for you" closing tied to the inner-loop user.
- "To take advantage of v0.26.4" block per the v0.13+ self-repair
  template (gbrain upgrade + contributor steps).
- Itemized changes by area (new scripts, script extensions, package.json
  tier split, CI tightening, failure-first logging, quarantine, regression
  tests, bunfig).
- "What did NOT ship" section names the intra-file project + E2E
  template-DB project as P0/P1 follow-ups with concrete acceptance
  criteria.
- Process section names the codex review + scope-correction loop
  honestly: "snapped back to ship today once empirical measurement showed
  Bun's --max-concurrency does nothing on tests not marked
  test.concurrent()."
- For-contributors note on portability + single-writer + fallback paths.

TODOS.md adds two P-rated entries:
- P0: intra-file parallelism via --concurrent flag. Sweep ~58 PGLite
  sites + ~40 env mutations + 2 mock.module sites. Target: bun run test
  < 30s. ~1-2 weeks. Detailed acceptance criteria. References Codex
  findings and plan-file rationale.
- P1: E2E parallelism via Postgres template databases. CREATE DATABASE
  TEMPLATE gbrain_template per test file. ~1-2 days.

llms.txt + llms-full.txt regenerated via `bun run build:llms` to absorb
the CLAUDE.md changes (per CLAUDE.md's "After any release ship that
touches the Key Files annotations in CLAUDE.md, run bun run build:llms"
rule). The build-llms regression test was firing in shard 7 of the
parallel pass — caught the drift, regeneration cleared it. Final
measurement after fix: 94s wallclock, 3652 pass, 0 fail across 8
parallel shards + 34 serial tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-03 20:16:15 -07:00

157 lines
6.7 KiB
TypeScript

/**
* Regression tests (a) + (d) for scripts/run-unit-parallel.sh:
* (a) Exit-code propagation: a failing test in any shard MUST cause the
* wrapper to exit non-zero. The hardest contract to silently break
* in a fan-out wrapper (`for ... &; wait` returns the LAST child's
* status, not any failure's).
* (d) Failure-log contract: when any test fails, the wrapper writes
* extracted failure block(s) to .context/test-failures.log with
* `--- shard $i:` prefixes, and prints a loud stderr banner with
* the absolute path. Empty log ⇔ exit 0.
*
* The wrapper takes ~1.5 minutes against the real test suite. To keep
* this regression test fast and hermetic, we point it at a tiny tempdir
* containing one passing and one failing test, override the discovery
* roots via env-vars, and run with --shards=2.
*
* NOT covered here: the heartbeat (timing-sensitive, not load-bearing
* for correctness) and timeout / WEDGED markers (require synthesizing a
* hung test which is fragile across machines). Those rely on the live
* smoke tests captured in CHANGELOG measurements.
*/
import { describe, it, expect, beforeAll, afterAll } from 'bun:test';
import { execFileSync, spawnSync } from 'child_process';
import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, existsSync, rmSync, copyFileSync, chmodSync } from 'fs';
import { tmpdir } from 'os';
import { join, resolve } from 'path';
const REPO_ROOT = resolve(import.meta.dir, '..', '..');
const PARALLEL_SH_SRC = resolve(REPO_ROOT, 'scripts/run-unit-parallel.sh');
const SHARD_SH_SRC = resolve(REPO_ROOT, 'scripts/run-unit-shard.sh');
const SERIAL_SH_SRC = resolve(REPO_ROOT, 'scripts/run-serial-tests.sh');
let TMPROOT: string;
beforeAll(() => {
// Build a tiny repo-shaped tempdir with the wrapper scripts copied in
// and 4 fixture test files (3 pass, 1 fail). The wrapper's `find test`
// expression will pick them up via cwd.
TMPROOT = mkdtempSync(join(tmpdir(), 'gbrain-parallel-test-'));
mkdirSync(join(TMPROOT, 'scripts'), { recursive: true });
mkdirSync(join(TMPROOT, 'test'), { recursive: true });
copyFileSync(PARALLEL_SH_SRC, join(TMPROOT, 'scripts', 'run-unit-parallel.sh'));
copyFileSync(SHARD_SH_SRC, join(TMPROOT, 'scripts', 'run-unit-shard.sh'));
copyFileSync(SERIAL_SH_SRC, join(TMPROOT, 'scripts', 'run-serial-tests.sh'));
chmodSync(join(TMPROOT, 'scripts', 'run-unit-parallel.sh'), 0o755);
chmodSync(join(TMPROOT, 'scripts', 'run-unit-shard.sh'), 0o755);
chmodSync(join(TMPROOT, 'scripts', 'run-serial-tests.sh'), 0o755);
// 3 passing + 1 failing test file. Round-robin sharding will land
// them across 2 shards so we exercise the multi-shard merge path.
const passing = `import { describe, it, expect } from 'bun:test';
describe('passing', () => {
it('arithmetic works', () => { expect(1 + 1).toBe(2); });
});`;
const failing = `import { describe, it, expect } from 'bun:test';
describe('failing-on-purpose', () => {
it('expects 1 to equal 2 (this should fail)', () => { expect(1).toBe(2); });
});`;
writeFileSync(join(TMPROOT, 'test', 'a-pass.test.ts'), passing);
writeFileSync(join(TMPROOT, 'test', 'b-pass.test.ts'), passing);
writeFileSync(join(TMPROOT, 'test', 'c-pass.test.ts'), passing);
writeFileSync(join(TMPROOT, 'test', 'd-fail.test.ts'), failing);
});
afterAll(() => {
if (TMPROOT) rmSync(TMPROOT, { recursive: true, force: true });
});
function runWrapper(extraArgs: string[] = []): { code: number; stdout: string; stderr: string } {
const result = spawnSync(
'bash',
[join(TMPROOT, 'scripts', 'run-unit-parallel.sh'), '--shards', '2', ...extraArgs],
{ cwd: TMPROOT, encoding: 'utf-8', env: { ...process.env } },
);
return {
code: result.status ?? -1,
stdout: result.stdout || '',
stderr: result.stderr || '',
};
}
describe('run-unit-parallel.sh exit-code propagation (a)', () => {
it('exits non-zero when any shard contains a failing test', () => {
const r = runWrapper();
expect(r.code).not.toBe(0);
});
it('exits zero when all shards pass (after removing the failing fixture)', () => {
rmSync(join(TMPROOT, 'test', 'd-fail.test.ts'));
try {
const r = runWrapper();
expect(r.code).toBe(0);
} finally {
// Restore the failing fixture for any downstream tests in the same
// describe block (afterAll cleans the whole tempdir; this is belt-
// and-suspenders).
const failing = `import { describe, it, expect } from 'bun:test';
describe('failing-on-purpose', () => {
it('expects 1 to equal 2', () => { expect(1).toBe(2); });
});`;
writeFileSync(join(TMPROOT, 'test', 'd-fail.test.ts'), failing);
}
});
});
describe('run-unit-parallel.sh failure-log contract (d)', () => {
it('writes failures to .context/test-failures.log with --- shard prefix on failure', () => {
const r = runWrapper();
expect(r.code).not.toBe(0);
const failureLog = join(TMPROOT, '.context/test-failures.log');
expect(existsSync(failureLog)).toBe(true);
const contents = readFileSync(failureLog, 'utf-8');
expect(contents.length).toBeGreaterThan(0);
expect(contents).toMatch(/--- shard \d+:/);
expect(contents).toContain('failing-on-purpose');
});
it('prints loud stderr banner with absolute failure-log path on failure', () => {
const r = runWrapper();
expect(r.code).not.toBe(0);
expect(r.stderr).toContain('TEST FAILURES');
// Banner includes the absolute path so users can `cat` it directly.
expect(r.stderr).toContain(join(TMPROOT, '.context', 'test-failures.log'));
});
it('clears .context/test-failures.log to empty when all shards pass', () => {
// Pre-seed a stale failure log to prove it gets cleared.
mkdirSync(join(TMPROOT, '.context'), { recursive: true });
writeFileSync(join(TMPROOT, '.context', 'test-failures.log'), 'STALE\n');
rmSync(join(TMPROOT, 'test', 'd-fail.test.ts'));
try {
const r = runWrapper();
expect(r.code).toBe(0);
const contents = readFileSync(join(TMPROOT, '.context', 'test-failures.log'), 'utf-8');
expect(contents).toBe('');
} finally {
const failing = `import { describe, it, expect } from 'bun:test';
describe('failing-on-purpose', () => {
it('expects 1 to equal 2', () => { expect(1).toBe(2); });
});`;
writeFileSync(join(TMPROOT, 'test', 'd-fail.test.ts'), failing);
}
});
it('writes per-shard summary lines to .context/test-summary.txt', () => {
runWrapper();
const summary = readFileSync(join(TMPROOT, '.context', 'test-summary.txt'), 'utf-8');
// Format: `shard 1/2: pass=N fail=N skip=N rc=N`
expect(summary).toMatch(/shard 1\/2: pass=\d+ fail=\d+ skip=\d+ rc=\d+/);
expect(summary).toMatch(/shard 2\/2: pass=\d+ fail=\d+ skip=\d+ rc=\d+/);
});
});