mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* feat: diff-aware E2E test selector Adds scripts/select-e2e.ts: reads git diff vs origin/master, classifies the change set (EMPTY/DOC_ONLY/SRC), and emits the relevant E2E test files on stdout. Fail-closed by design: any unmapped src/ change runs all E2E. - scripts/e2e-test-map.ts: hand-tuned path-glob -> test files map - scripts/select-e2e.ts: pure-function selector with three explicit cases - scripts/run-e2e.sh: accepts optional file list from argv + --dry-run-list - test/select-e2e.test.ts: 24 cases including 3 codex regression guards (skills/, untracked files, unmapped src/) * feat: local CI gate via docker compose Adds bun run ci:local — runs every check GH Actions runs (gitleaks + unit + 29 E2E files) inside a Docker container that bind-mounts the repo. Pure bind-mount + named volumes (gbrain-ci-node-modules, gbrain-ci-bun-cache, gbrain-ci-pg-data) for fast warm restarts. - docker-compose.ci.yml: pgvector/pgvector:pg16 + oven/bun:1 - scripts/ci-local.sh: orchestrator with --diff, --no-pull, --clean - gitleaks runs on host (scoped to working dir + branch commits) - DATABASE_URL unset for unit phase (matches GH Actions split) - git installed in container at startup (oven/bun:1 omits it) - Postgres host port via GBRAIN_CI_PG_PORT env (default 5434) Stronger than PR CI: runs all 29 E2E files vs CI's 2-file Tier 1. * chore: bump version and changelog (v0.23.1) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: document local CI gate for v0.23.1 CLAUDE.md gains key-files entries for docker-compose.ci.yml, scripts/ci-local.sh, scripts/select-e2e.ts + e2e-test-map.ts, and the scripts/run-e2e.sh argv tweak. Pre-ship requirements section now lists the Docker-based local gate as Path A alongside the manual lifecycle. CONTRIBUTING.md tests section adds the bun run ci:local / ci:local:diff / ci:select-e2e block with prerequisites (Docker engine + gitleaks) and the GBRAIN_CI_PG_PORT override. AGENTS.md "Before shipping" promotes ci:local as the easiest path and keeps the manual lifecycle as a fallback. README.md Contributing section points to ci:local for the full gate. CHANGELOG.md untouched — v0.23.1 entry already finalized. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat: SHARD=N/M env support in scripts/run-e2e.sh Filters the E2E file list to every M-th file starting at index N (1-indexed). Sequential execution within a shard preserves the TRUNCATE CASCADE no-race property documented at the top of the file. Empty-shard handling under `set -u` uses ${arr[@]:-} fallback. Standalone change; not yet wired up in ci-local.sh. * feat: 4-way parallel E2E shards in ci:local Replaces the single postgres service with 4 (postgres-1..4) on host ports 5434-5437. scripts/ci-local.sh fans 4 workers via xargs -P4 inside the runner container; each pinned to its own DATABASE_URL via SHARD=N/4. Wall-time on a 16-core host: ~6 min sequential -> ~1.5-2 min sharded. Total full-gate wall-time goes from ~25 min to ~3-5 min warm. Also handles git-worktree (Conductor) layouts: when /app/.git is a file instead of a directory, parse the gitdir + commondir and bind-mount the shared host gitdir at its absolute path. Without this, in-container `git ls-files` (used by scripts/check-trailing-newline.sh and friends) exits 128 with "not a git repository". Also runs `git config --global --add safe.directory '*'` inside the container so the root-uid container can read host-uid gitdir without "dubious ownership" rejection. CHANGELOG entry updated to cover the speedup. - docker-compose.ci.yml: 4 pgvector services + per-shard named volumes - scripts/ci-local.sh: parallel xargs orchestration + worktree mount fix - CHANGELOG.md v0.23.1: 4-way sharded wall-time, 36 E2E files, --no-shard flag * chore: regenerate llms-full.txt for v0.23.1 doc updates Required by test/build-llms.test.ts case 4 — committed llms-full.txt must match `bun run build:llms` output. The CHANGELOG + CLAUDE.md updates in this branch shifted bytes; regen catches up. * feat: scripts/run-unit-shard.sh + slow-test convention Tier 1 + Tier 4 plumbing: - scripts/run-unit-shard.sh: SHARD=N/M filter for unit files (excludes test/e2e/*). Excludes *.slow.test.ts (Tier 4 convention) so the fast shard fan-out skips known-slow files; CI's `bun run test` still includes them via default discovery. - scripts/run-slow-tests.sh: companion that runs ONLY *.slow.test.ts. Wired as `bun run test:slow`. - scripts/profile-tests.sh: portable awk parser that extracts the top-N slowest tests from any captured `bun test` output. Wired as `bun run test:profile`. Use it to pick demotion candidates. * feat: PGLite snapshot fixture for ~4.5x faster cold init (Tier 3) scripts/build-pglite-snapshot.ts boots a fresh PGLite, runs the full initSchema() (forward bootstrap + 30 migrations), and dumps the post-init state to test/fixtures/pglite-snapshot.tar plus a SHA-256 schema hash sidecar (.version). Both gitignored — built on demand via `bun run build:pglite-snapshot`. PGLiteEngine.connect() reads GBRAIN_PGLITE_SNAPSHOT env: validates the sidecar hash against the in-process MIGRATIONS hash, loads via PGLite's loadDataDir blob, sets _snapshotLoaded so initSchema() short-circuits. Measured per-file cold init drops from 828ms → 181ms. Bootstrap-correctness tests (bootstrap.test.ts, schema-bootstrap-coverage.test.ts) explicitly delete the env at file top so they keep exercising the cold path they verify. * feat: --classify-only + heartbeat tolerance fix (Tiers 2 + flake fix) - scripts/select-e2e.ts: --classify-only flag emits EMPTY|DOC_ONLY|SRC. Used by ci-local.sh's --diff fast-path to skip the heavy gate when only docs changed. - test/progress.test.ts: startHeartbeat tolerance widened to 1-20 over 200ms (was 2-6 over 85ms). Under 4-way parallel shard load on a contended host, setTimeout's effective quantum balloons and the tight bound flakes. The test still verifies "fires multiple times, stops cleanly" — exact count was never load-bearing. * feat: 4-way unit + E2E sharding in ci-local.sh + CHANGELOG (Tiers 1-4) ci-local.sh ties the four tiers together: - Tier 2: pre-flight diff classification on host. DOC_ONLY exits in ~5s (gitleaks only, no postgres, no container). - Tier 1: guards + typecheck run ONCE before fan-out. xargs -P4 then spawns 4 shards inside the runner container, each running unit phase (env -u DATABASE_URL bash run-unit-shard.sh) followed by E2E phase (DATABASE_URL=postgres-N bash run-e2e.sh) — both sharded N/4. Per-shard logs in /tmp/shard-logs/shard-N.log; printed in shard order at the end. - Tier 3: snapshot fixture built once at runner startup if missing, GBRAIN_PGLITE_SNAPSHOT exported so all shards inherit. - Tier 4: run-unit-shard.sh excludes *.slow.test.ts; run-slow-tests.sh + test:slow npm script handle the demoted set. - --no-shard preserves the legacy single-process flow for debug. package.json: build:pglite-snapshot, test:slow, test:profile scripts. Measured wall-time on 16-core host: 100s warm (down from ~22 min cold single-process). 4 shards × ~640-1024 unit tests each, plus 9 E2E files each. PGLite snapshot saves 4.5× per cold init (828ms → 181ms). CHANGELOG.md updated with measured numbers + four-tier breakdown. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
130 lines
4.2 KiB
Bash
Executable File
130 lines
4.2 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Run E2E tests ONE FILE AT A TIME.
|
|
#
|
|
# Bun's default is to run test files in parallel (each in its own worker).
|
|
# Our E2E suite shares one Postgres database across all 13 files, and
|
|
# `setupDB()` does TRUNCATE CASCADE + fixture import. When files run in
|
|
# parallel, file A's TRUNCATE can race with file B's fixture import,
|
|
# producing observed fails like "expected 16 pages, got 8", missing
|
|
# links, orphaned timeline entries, etc. The flakiness was visible on
|
|
# ~3 of every 5 runs pre-fix.
|
|
#
|
|
# Running files sequentially eliminates the race entirely. It also costs
|
|
# some startup overhead (each file spins up a fresh bun process) but for
|
|
# a suite this size that is measured in ~1-2s per file, amortized under
|
|
# the natural per-file test time of 5-10s.
|
|
#
|
|
# Exits non-zero on the first failing file so CI fails fast.
|
|
#
|
|
# `--timeout=60000` matches the unit test suite. Bun's default is 5s,
|
|
# which is too tight for setupDB's TRUNCATE CASCADE on ~30 tables on
|
|
# CI runners under load (one CI flake observed on PR #475 hitting
|
|
# exactly 5000.09ms in the Tags beforeAll).
|
|
|
|
set -euo pipefail
|
|
|
|
cd "$(dirname "$0")/.."
|
|
|
|
# --dry-run-list: print the resolved file list (one per line) and exit. Used
|
|
# by scripts/ci-local.sh to smoke-test the argv branching at startup.
|
|
DRY_RUN_LIST=0
|
|
if [ "${1:-}" = "--dry-run-list" ]; then
|
|
DRY_RUN_LIST=1
|
|
shift
|
|
fi
|
|
|
|
# Argv-driven file list (used by `ci:local:diff`); fall back to the full glob.
|
|
if [ "$#" -gt 0 ]; then
|
|
files=("$@")
|
|
else
|
|
files=(test/e2e/*.test.ts)
|
|
fi
|
|
|
|
# SHARD env (e.g. SHARD=1/4) keeps every M-th file starting at index N (1-indexed).
|
|
# Used by scripts/ci-local.sh to fan 4 shards in parallel against 4 postgres
|
|
# containers. Sequential execution within a shard is preserved (the TRUNCATE
|
|
# CASCADE no-race rationale at the top of this file still holds).
|
|
if [ -n "${SHARD:-}" ]; then
|
|
shard_n=${SHARD%/*}
|
|
shard_m=${SHARD#*/}
|
|
if ! printf '%s' "$shard_n" | grep -qE '^[0-9]+$' || \
|
|
! printf '%s' "$shard_m" | grep -qE '^[0-9]+$' || \
|
|
[ "$shard_n" -lt 1 ] || [ "$shard_m" -lt 1 ] || [ "$shard_n" -gt "$shard_m" ]; then
|
|
echo "ERROR: invalid SHARD=$SHARD (expected N/M with 1<=N<=M, both integers)" >&2
|
|
exit 1
|
|
fi
|
|
filtered=()
|
|
i=0
|
|
for f in "${files[@]}"; do
|
|
if [ $((i % shard_m + 1)) -eq "$shard_n" ]; then
|
|
filtered+=("$f")
|
|
fi
|
|
i=$((i + 1))
|
|
done
|
|
# ${filtered[@]:-} avoids "unbound variable" under `set -u` when no files matched.
|
|
files=("${filtered[@]:-}")
|
|
# If the empty placeholder slipped in, drop it.
|
|
if [ "${#files[@]}" -eq 1 ] && [ -z "${files[0]}" ]; then
|
|
files=()
|
|
fi
|
|
fi
|
|
|
|
if [ "$DRY_RUN_LIST" = "1" ]; then
|
|
if [ "${#files[@]}" -eq 0 ]; then
|
|
exit 0
|
|
fi
|
|
printf '%s\n' "${files[@]}"
|
|
exit 0
|
|
fi
|
|
|
|
if [ "${#files[@]}" -eq 0 ]; then
|
|
# Empty shard (e.g. SHARD=4/4 with only 3 files): nothing to do.
|
|
echo "No files for shard ${SHARD:-(unsharded)}; exiting clean."
|
|
exit 0
|
|
fi
|
|
|
|
pass_files=0
|
|
fail_files=0
|
|
fail_list=()
|
|
total_pass=0
|
|
total_fail=0
|
|
|
|
for f in "${files[@]}"; do
|
|
name=$(basename "$f")
|
|
echo ""
|
|
echo "=== $name ==="
|
|
if output=$(bun test --timeout=60000 "$f" 2>&1); then
|
|
pass_files=$((pass_files + 1))
|
|
# Extract pass/fail counts from bun's summary (e.g., "123 pass")
|
|
p=$(echo "$output" | grep -oE '[0-9]+ pass' | tail -1 | grep -oE '[0-9]+' || echo 0)
|
|
total_pass=$((total_pass + p))
|
|
echo "$output" | tail -8
|
|
else
|
|
fail_files=$((fail_files + 1))
|
|
fail_list+=("$name")
|
|
p=$(echo "$output" | grep -oE '[0-9]+ pass' | tail -1 | grep -oE '[0-9]+' || echo 0)
|
|
fl=$(echo "$output" | grep -oE '[0-9]+ fail' | tail -1 | grep -oE '[0-9]+' || echo 0)
|
|
total_pass=$((total_pass + p))
|
|
total_fail=$((total_fail + fl))
|
|
echo "$output"
|
|
echo ""
|
|
echo "FAILED: $name"
|
|
# Continue so we see all failures; exit nonzero at the end.
|
|
fi
|
|
done
|
|
|
|
echo ""
|
|
echo "========================================"
|
|
echo "E2E SUMMARY (sequential execution)"
|
|
echo "========================================"
|
|
echo "Files: $((pass_files + fail_files)) total, $pass_files passed, $fail_files failed"
|
|
echo "Tests: $total_pass passed, $total_fail failed"
|
|
if [ ${#fail_list[@]} -gt 0 ]; then
|
|
echo ""
|
|
echo "Failing files:"
|
|
for f in "${fail_list[@]}"; do
|
|
echo " - $f"
|
|
done
|
|
exit 1
|
|
fi
|