mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-31 04:07:52 +00:00
Schema-migration matrix + fuzz harness + RSS budget gate + read-latency
under sync + sync lock regression + tests/heavy convention + nightly CI
workflow + BFS frontier cap on traverseGraph.
CI infra (T1-T7):
- tests/heavy/ directory convention + scripts/run-heavy.sh + bun run test:heavy
- tests/heavy/pg_upgrade_matrix.sh: walk pre-v0.13 + pre-v0.18 brain shapes
forward to head via bootstrap → SCHEMA_SQL → migrations → verifySchema
- test/fuzz/{pure,mixed,filesystem}-validators.test.ts: 1000-run fast-check
property tests across 8 trust-boundary validators
- scripts/check-fuzz-purity.sh: bun-bundle + grep guard, wired into verify
- tests/heavy/measure_rss.sh: in-memory PGLite workload + peak RSS measurement
via /proc/self/status (Linux) or process.memoryUsage().rss fallback (macOS,
refuses to write baseline)
- tests/heavy/read_latency_under_sync.sh: phase A baseline + phase B under
parallel writer load, reports p50/p95/p99 + delta_pct
- tests/heavy/sync_lock_regression.sh: N concurrent gbrain sync against one
DB, asserts 1 winner + N-1 lock-busy + zero leaked gbrain_cycle_locks rows
- .github/workflows/heavy-tests.yml: cron '17 8 * * *' + heavy-tests label
trigger + Postgres service + artifact upload on failure
Engine (T8):
- BrainEngine.traverseGraph opts gain frontierCap?: number + onTruncation?:
(info: TruncationInfo) => void callback. Return shape preserved
(Promise<GraphNode[]>) for MCP wire stability.
- Postgres CTE: parenthesized LIMIT N ORDER BY (slug, id) inside recursive term.
- PGLite: same SQL with positional params.
- Per-call callback closure — not engine-instance state — so concurrent
traversals on the same engine don't cross-talk. 5 contracts pinned in
test/regressions/v0_36_frontier_cap.test.ts.
Three plan-review passes ran before any code: CEO scope review (Approach C),
Eng dual-voice review (Claude subagent + Codex), and Codex 2nd-pass against
the revised plan. The 2nd pass caught issues the first two missed (Bun ESM
vs require.cache; engine-instance metadata stomping under concurrency;
fixture-size inconsistency). All addressed.
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
74 lines
2.6 KiB
Markdown
74 lines
2.6 KiB
Markdown
# tests/heavy/
|
|
|
|
Heavy ops-shape tests. Shell scripts that exercise gbrain end-to-end against
|
|
real infrastructure (Postgres, large fixtures, concurrent processes). Cost
|
|
minutes per run; NOT in default `bun test`.
|
|
|
|
## When to add a script here
|
|
|
|
Put a test here if it:
|
|
- Costs more than ~30s wallclock per run
|
|
- Needs real Postgres (not PGLite in-memory)
|
|
- Spins up multiple processes or measures concurrency
|
|
- Measures system metrics (RSS, latency under load, lock contention)
|
|
- Tests an upgrade / migration matrix against committed historical states
|
|
|
|
## When to use `*.slow.test.ts` instead
|
|
|
|
Put a slow test in `test/` with the `.slow.test.ts` suffix if it:
|
|
- Runs under `bun test` (TypeScript, uses bun:test imports)
|
|
- Is correctness-shaped, not ops-shaped (asserts behavior of one function)
|
|
- Can stub external dependencies
|
|
|
|
The two patterns coexist intentionally. `*.slow.test.ts` is per-file
|
|
correctness for cold paths; `tests/heavy/` is ops-shape scripts that don't
|
|
fit bun's test runner.
|
|
|
|
## How to run
|
|
|
|
```bash
|
|
# Run every script in this directory, sequentially:
|
|
bun run test:heavy
|
|
|
|
# Run a single script:
|
|
tests/heavy/<script>.sh
|
|
```
|
|
|
|
The runner is `scripts/run-heavy.sh`. It discovers every `tests/heavy/*.sh`
|
|
file at this directory's top level (NOT recursive), runs them in lexical
|
|
order, fails on the first non-zero exit.
|
|
|
|
## Naming convention
|
|
|
|
- `tests/heavy/<name>.sh` — top-level test script, picked up by the runner.
|
|
- `tests/heavy/_<name>.sh` — library/helper invoked by a sibling test.
|
|
The leading underscore tells the runner to SKIP this file. Use this
|
|
pattern for fixture builders, shared setup, anything that needs a
|
|
required argument or isn't standalone-runnable.
|
|
- `tests/heavy/fixtures/<name>` — committed input data (SQL, JSON, etc).
|
|
|
|
## CI scheduling
|
|
|
|
Heavy tests run nightly at 08:17 UTC via `.github/workflows/heavy-tests.yml`,
|
|
and on PRs labeled `heavy-tests`. They are NOT part of the default PR CI
|
|
matrix — that gate stays fast.
|
|
|
|
## Failure output convention
|
|
|
|
Each script writes a per-run log to `~/.gbrain/audit/heavy-<script>-<ts>.log`
|
|
containing subprocess stdout/stderr, environment state, and any captured
|
|
metrics. The CI workflow uploads these as artifacts on failure for triage
|
|
without re-running locally.
|
|
|
|
## Style
|
|
|
|
- `#!/usr/bin/env bash`
|
|
- `set -euo pipefail`
|
|
- Explicit array argv for execs (no `eval`, no unquoted globs)
|
|
- Print a one-line `[<script>] <action>` log per major step
|
|
- Exit non-zero on any failure path; print enough context to diagnose
|
|
- Honor `$GBRAIN_HOME` / `$TMP_ROOT` env overrides where relevant
|
|
|
|
See `scripts/check-jsonb-pattern.sh` and `scripts/run-slow-tests.sh` for the
|
|
in-tree style reference.
|