mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* feat(engine): add deletePages + resolveSlugsByPaths to BrainEngine (v0.41.21.0 T1) Two new REQUIRED methods on the BrainEngine interface, implemented on both Postgres and PGLite engines. Closes the per-file N+1 query pattern that PR #1538 batched on Postgres only. deletePages(slugs: string[], opts: { sourceId: string }): Promise<string[]> — Single SQL round-trip: DELETE FROM pages WHERE slug = ANY($1::text[]) AND source_id = $2 RETURNING slug — Returns slugs ACTUALLY DELETED (D6, codex CDX-8) so callers can filter pagesAffected to exclude phantom slugs (paths in the deletion list but with no DB row). — Single-batch primitive: caller chunks input to DELETE_BATCH_SIZE. Throws if input exceeds the cap. — sourceId is REQUIRED at the type level (D5, codex CDX-10). Asymmetric with single-row deletePage which keeps the optional 'default' fallback for back-compat. v0.42+ TODO to tighten. resolveSlugsByPaths(paths, opts): Promise<Map<path, slug>> — Batch path → slug lookup. Single SQL round-trip: SELECT slug, source_path FROM pages WHERE source_path = ANY($1::text[]) AND source_id = $2 — Missing paths absent from the Map (caller falls back to path-derived slug, same contract as resolveSlugByPathOrSourcePath). — Empty input short-circuits to empty Map (no SQL). src/core/engine-constants.ts (NEW) — Single source of truth for DELETE_BATCH_SIZE = 500. — Both engines import; no engine-from-engine coupling. — Lives outside engine.ts (the interface module) to avoid circular imports. Also updates the deletePage JSDoc (CDX-11): drops the misleading "hard delete is admin-only" framing. `gbrain sync` hard-deletes on every run that sees a deleted file; not admin-only. Co-Authored-By: garrytan-agents <noreply@anthropic.com> Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(sync): batched delete + rename + DRY refactor (v0.41.21.0 T2/T3/T4) Replaces the per-file delete loop (sync.ts:1241-1257) and per-file rename slug-resolve (sync.ts:1263-1295) with interleaved per-batch flows using engine.resolveSlugsByPaths + engine.deletePages. Also refactors resolveSlugByPathOrSourcePath (sync.ts:267) to delegate to the new batch helper when sourceId is set — one owner of the SQL + fallback semantics (D8). ROUND-TRIP COUNTS (73K-delete commit): pre-fix: 73,000 SELECTs + 73,000 DELETEs = 146,000 (~5 hours) post-fix: 146 SELECTs + 146 DELETEs = 292 (~2 minutes) Headline win: a single commit deleting 73K files no longer jams the sync pipeline for hours, no longer cascades staleness across every other source on the brain. Shape (T2 delete loop, per the plan's ASCII diagram): filtered.deleted (73K paths) │ ▼ slice into batches of DELETE_BATCH_SIZE (500) │ ▼ for each batch: abort-check ──► partial('timeout') │ ▼ engine.resolveSlugsByPaths(batch, {sourceId}) ◀── 1 SQL round-trip │ ▼ slugs = batch.map(path => map.get(path) ?? resolveSlugForPath(path)) ◀── pure-JS fallback │ for frontmatter- ▼ fallback slugs try { deleted = engine.deletePages(slugs, opts) ◀── 1 SQL round-trip pagesAffected.push(...deleted) ◀── D6 confirmed only } catch { // D7 decompose: per-slug deletePage, // unrecoverable failures → failedFiles } Per-batch try-catch (D7) decomposes batch DELETE failures to per-slug deletePage so a transient blip on batch 73 doesn't lose 500 deletes — it self-heals to one-at-a-time for that batch only. Unrecoverable per-slug failures land in failedFiles (matching the existing import-loop pattern at sync.ts:~1350). failedFiles declaration hoisted above the delete loop so both delete decompose and import loops feed the same sync-bookmark gate. T4 rename loop: pre-resolves all `from` slugs in batches via resolveSlugsByPaths BEFORE iterating. Per-file updateSlug + importFile calls stay (those are inherently per-file). The try/catch around updateSlug for slug-doesn't-exist preserves verbatim. T3 DRY refactor: resolveSlugByPathOrSourcePath delegates to resolveSlugsByPaths via a single-element array when sourceId is set. When sourceId is undefined (legacy unscoped callers), falls back to the original executeRaw shape — the batch engine surface requires sourceId per D5 (multi-source-bug-class defense). Atomicity coarsening (D3): each batch is one transaction. A mid-batch abort or connection failure rolls back up to DELETE_BATCH_SIZE - 1 successful deletes from the in-flight batch. Sync is idempotent so the next run picks them up via git diff regenerating the deletion list. Documented at the call site + in the deletePages JSDoc. Co-Authored-By: garrytan-agents <noreply@anthropic.com> Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(schema): global page-generation clock + statement-level trigger (v0.41.21.0 T5) Migration v104: page_generation_clock_and_statement_trigger. The pre-v0.41.21.0 query-cache Layer 1 bookmark read MAX(generation) FROM pages to detect "writes happened since cache-store". Two bugs in that contract — independent of any sync work, surfaced by codex outside-voice on the /plan-eng-review pass: 1. The row-level bump_page_generation_trg (migration v91) sets NEW.generation = OLD.generation + 1 on UPDATE. Updating a NON-MAX page didn't advance MAX(generation). Cache silently served stale for any UPDATE-to-non-max page. (CDX-2) 2. The trigger is BEFORE INSERT OR UPDATE — DELETE doesn't fire it at all. Even an AFTER DELETE wouldn't move MAX (surviving rows are untouched). (CDX-1) Fix: single-row page_generation_clock counter, bumped per-statement (FOR EACH STATEMENT — per-row would turn a 73K-row batch DELETE into 73K UPDATEs on the same counter, recreating the bottleneck this PR fixes elsewhere — codex CDX-4). Layer 1 reads the clock value directly (T6, separate commit). Per-row pages.generation stays for Layer 2 (per-page snapshot via jsonb_each + LEFT JOIN pages) which doesn't care about MAX, only per-page advancement. Seeded with COALESCE(MAX(pages.generation), 0) so existing query_cache rows stored under the old MAX semantics aren't all instantly invalidated on upgrade. Their max_generation_at_store stamp compares cleanly against the seeded clock; future writes bump the clock and the bookmark fires correctly. CREATE TABLE page_generation_clock ( id INTEGER PRIMARY KEY CHECK (id = 1), value BIGINT NOT NULL DEFAULT 0 ); CREATE TRIGGER bump_page_generation_clock_trg AFTER INSERT OR UPDATE OR DELETE ON pages FOR EACH STATEMENT EXECUTE FUNCTION bump_page_generation_clock_fn(); Mirror in src/core/pglite-schema.ts so fresh PGLite installs get the table + trigger via SCHEMA_SQL replay. The forward-reference bootstrap probe doesn't need an entry: page_generation_clock is created directly by SCHEMA_SQL (no separate index or FK references it), so the schema-bootstrap-coverage gate is satisfied as-is. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(cache): move Layer 1 to global clock + invalidate empty snapshots (v0.41.21.0 T6) Closes the silent stale-cache bug class that's been live in master since the bookmark feature shipped. Pre-fix, gbrain search would silently serve stale cached results in three independent scenarios: 1. UPDATE to a non-max-generation page (CDX-2) — the row-level trigger advanced per-page generation but didn't move MAX(generation), so the bookmark passed. 2. DELETE of any page (CDX-1) — the trigger didn't fire at all, and even an AFTER DELETE wouldn't move MAX. 3. Empty-result cache row + subsequent matching INSERT (CDX-6 / D20) — page_generations = '{}'::jsonb was "vacuously valid" via Layer 2, surviving any clock bump. Fix: buildPageGenerationsSnapshot (store path) — Replaces the SELECT MAX(generation) FROM pages reads at cache-write time with SELECT value FROM page_generation_clock WHERE id = 1. — Empty pageIds path: only need the clock value (D20 contract). — Combined non-empty path: per-page generation (Layer 2 substrate) + clock value, both folded in one round trip via UNION ALL. CACHE_GATE_WHERE_CLAUSE (lookup path) — Layer 1 reads page_generation_clock.value (single-row O(1) lookup, faster than the pre-fix MAX(generation) backward index scan). — Layer 2 stricter: requires page_generations <> '{}'::jsonb AND the per-page check (not OR with the vacuously-valid `= '{}'` shortcut). Empty snapshots can no longer survive a Layer 1 miss. validateCacheRowAgainstPages (pure validator) — Layer 2 returns false for empty snapshots when Layer 1 fails. — Documented contract change. Backward compat: pre-v0.40.3.0 cache rows have max_generation_at_store = 0 AND page_generations = '{}'::jsonb. On a populated brain, Layer 1 fails (clock > 0). Layer 2 is now stricter so legacy rows invalidate once on first post-upgrade lookup, then the cache fills back correctly. Acceptable one-time miss spike; post-upgrade cache is structurally sound. The clock seed (COALESCE(MAX(pages.generation), 0)) from migration v104 keeps NON-empty legacy rows passing Layer 1 until the next write — they don't all invalidate at once. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test: cover v0.41.21.0 delete-batch + global clock + cache contract (T7+T8+T9) Tests for every behavior the v0.41.21.0 wave introduces or changes. New test files: test/sync-delete-batch.test.ts (PGLite hermetic) — engine.deletePages: empty input short-circuit, returns confirmed slugs (D6), multi-source isolation, cascade integrity (chunks + links cleared via FK), rejects oversized input. — engine.resolveSlugsByPaths: empty input, present + missing rows, D10 exotic-filename substrate (🌟.md / ทดสอบ.md / عربي.md), source isolation. — D13 pagesAffected filter: 100 deletable + 10 ghost paths → deletePages returns 100 (regression-pin: pre-fix would return all 110 via D6's pre-RETURNING shape). test/sync-delete-batch.slow.test.ts (.slow suffix keeps it out of the fast loop) — 10K-page batched delete completes in <5s on PGLite. Measured 277ms on dev hardware (18x under the gate); pins the headline perf promise. test/sync-rename-batch.test.ts (PGLite hermetic) — 500-rename batch slug-resolve in 1 round-trip (exactly at DELETE_BATCH_SIZE boundary). — Frontmatter-fallback rename: exotic source_paths resolve via the batch SELECT. — Mixed present + missing: partial Map (missing → caller falls back to path-derived). test/page-generation-counter.test.ts (PGLite hermetic) — Statement-level trigger fires once per INSERT statement (raw SQL — NOT putPage, which uses ON CONFLICT DO UPDATE and bumps by 2 in PG semantics). — Statement-level trigger fires once per UPDATE statement. — Headline contract: batch DELETE bumps clock by 1, NOT by row count (25-row batch → +1). — CDX-1 regression: DELETE of non-max page bumps clock. — CDX-2 regression: UPDATE of non-max page bumps clock (raw SQL). — D14 end-to-end: clock advances after batch DELETE → cache rows stamped at the prior clock value are now stale by Layer 1. — CDX-6/D20: empty-result cache + INSERT matching page → clock advances (Layer 1 fires). — Documents the PG quirk: putPage's INSERT...ON CONFLICT DO UPDATE bumps clock by 2 (both INSERT and UPDATE triggers fire). Test-helper update: test/helpers/reset-pglite.ts — Added page_generation_clock to PRESERVE_TABLES so the seeded single-row counter survives resetPgliteState between tests (same treatment as schema_version). Production never truncates. Existing test contract inversions (CDX-6 / D20 fix): test/query-cache-gate.test.ts — Pre-v0.41.21.0 "vacuously valid for legacy empty snapshot" assertion inverted: empty snapshot now invalidates when Layer 1 fires. Add positive CDX-6 regression test (empty-result + INSERT matching page). — SQL shape regression: page_generation_clock in Layer 1 (negative regression guard: MAX(generation) FROM pages MUST be gone). — Empty-snapshot reject guard: `qc.page_generations <> '{}'::jsonb` present; the old `qc.page_generations = '{}'::jsonb OR` shortcut MUST be gone. test/e2e/cache-gate-pglite.test.ts — Pre-v0.41.21.0 "legacy row serves vacuously" test inverted: legacy rows now invalidate on first clock advance post-upgrade. — CDX-1 regression: DELETE bumps clock → cached query for surviving pages invalidates. — CDX-2 regression: UPDATE-to-non-max-page bumps clock → cache invalidates. — CDX-11 comment fix: drop misleading "hard delete is admin-only" framing; gbrain sync hard-deletes on every run. Engine parity extension: test/e2e/engine-parity.test.ts — deletePages parity: same input set, both engines return same string[] of confirmed-deleted slugs (D6). — resolveSlugsByPaths parity: same Map on both engines. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(release): v0.41.21.0 — batched sync deletes + global page-generation clock (T10) VERSION bump (0.41.18.0 → 0.41.21.0; master is at 0.41.20.0 so next free slot per the queue allocator). CHANGELOG entry with the ELI10 lead per CLAUDE.md voice rules. CLAUDE.md annotations on engine.ts, postgres-engine.ts, pglite-engine.ts, sync.ts, and query-cache-gate.ts plus a new entry for engine-constants.ts. llms-full.txt regenerated to match CLAUDE.md (per CLAUDE.md mandatory rule). Co-Authored-By: garrytan-agents <noreply@anthropic.com> Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * dx(test-runner): heartbeat shows real progress instead of 0p 0f Bun's default test reporter doesn't print per-test markers — only a single shard-end summary block when you pass it a file list. The existing heartbeat tried to count `^[[:space:]]+✓` lines as a live pass-count proxy, but bun never emits them in the multi-file mode this runner uses, so every mid-run heartbeat showed `0p 0f` for the entire 12-20 minute wallclock. Users (and agents polling the runner) couldn't distinguish "still bootstrapping" from "wedged" from "almost done." Fix: parse three complementary real-time signals instead. 1. Total files this shard was assigned — parsed from the `[unit-shard N/M] running X files` banner that run-unit-shard.sh echoes before invoking bun test. Available from second 1. 2. PGLite initSchema() count — proxy for "test files started so far." Each PGLite-using test file's beforeAll triggers one initSchema(), which logs `Schema version 1 → 106 (101 migration(s) pending)`. Undercounts because not every test file opens a PGLite engine (covers ~30-60% of files in practice), but it's the only real-time progress signal bun's default reporter leaves in the log. The output uses a `~` prefix to convey "approximate count." 3. Log size in KB — strictly monotonic liveness signal that works even when the PGLite count is still 0 (early-shard startup before the first initSchema fires). 4. Per-shard elapsed time — formatted as MmSSs. New mid-run heartbeat line: [heartbeat] [s1: ~62/190f 476KB 12m31s] [s2: ~63/190f 513KB 12m31s] ... When a shard finishes, the heartbeat upgrades to its final summary including pass/fail counts from bun's end-of-shard summary block: [heartbeat] [s1: done ✓ 2807p 0f] [s2: done ✓ 2784p 0f] ... Portability: BSD awk on macOS doesn't support `match($0, /re/, arr)` with the array sink — that's a gawk extension. The total-files parser uses sed instead so the runner stays portable to the default Mac toolchain. Helpers are pure functions and unit-testable in isolation: pass a log file path, get the parsed number. No mocking. No bun runtime required. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(release): rebump v0.41.23.0 → v0.41.25.0 Per user request — skip v0.41.23.0 / v0.41.24.0 slots to land at v0.41.25.0. Master is at v0.41.22.1, no version-trio collision. Touches VERSION, package.json, CHANGELOG header, CLAUDE.md annotations, src/core/engine-constants.ts header, src/core/migrate.ts migration v106 comment, regenerated llms-full.txt + llms.txt. Migration version (v106) and CDX1-6 trigger semantics unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(schema): reorder migration_impact_log AFTER minion_jobs (CI green) Pre-existing bug in master's SCHEMA_SQL ordering, surfaced by CI on this PR but lived silently on master since v0.41.18.0. migration_impact_log declares `job_id BIGINT REFERENCES minion_jobs(id)`, but its CREATE TABLE was at line 658 while minion_jobs's CREATE TABLE was at line 778. On any fresh-install initSchema() the FK target didn't exist yet: psql:/tmp/schema.sql:672: ERROR: relation "minion_jobs" does not exist postgres-js's `unsafe()` aborts the multi-statement batch on the first error response, so every CREATE TABLE after migration_impact_log (including minion_jobs itself) never ran. Every subsequent CLI subprocess that opened a connection then crashed with `relation "minion_jobs" does not exist` on its first query. Why master CI sometimes passed: the per-shard advisory lock + the test setup's `engine.initSchema()` second pass (which runs the migrations array) would eventually create minion_jobs via the v5 `minion_jobs_table` migration. From there migration_impact_log would land via migration v103 with its FK resolving correctly. But CLI subprocesses spawned by mechanical.test.ts's Parallel Import block open their OWN connections and run a fresh `engine.connect() → initSchema()` — that path runs SCHEMA_SQL FIRST and aborted at the same forward-reference error before the migrations array could repair. Fix: relocate the migration_impact_log CREATE TABLE + its two indexes to AFTER the minion_jobs CREATE TABLE block (lines ~865), keeping the rest of the schema layout intact. PGLite schema (pglite-schema.ts) already had the correct ordering — only Postgres SCHEMA_SQL needed the move. Verified: fresh-DB local repro that previously failed 31/34 tests with `relation minion_jobs does not exist` now passes 78/78 in test/e2e/mechanical.test.ts. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: garrytan-agents <noreply@anthropic.com>
430 lines
20 KiB
Bash
Executable File
430 lines
20 KiB
Bash
Executable File
#!/usr/bin/env bash
|
||
# scripts/run-unit-parallel.sh — fast unit-test loop, parallel fan-out.
|
||
#
|
||
# Spawns N parallel `bun test` processes, each running a hash-disjoint shard
|
||
# of the unit-test set (files only — no e2e, no .slow, no .serial). After
|
||
# all shards complete, runs serial-only files (*.serial.test.ts) with
|
||
# --max-concurrency=1. Failure-first logging: extracts failure blocks from
|
||
# each shard's log, writes to .context/test-failures.log with --- shard $i:
|
||
# prefixes, prints loud stderr banner if any failures, exit non-zero.
|
||
#
|
||
# Usage:
|
||
# bash scripts/run-unit-parallel.sh [--shards N] [--max-concurrency N] [--dry-run]
|
||
#
|
||
# Env overrides:
|
||
# SHARDS=N same as --shards
|
||
# GBRAIN_TEST_SHARD_TIMEOUT per-shard wallclock cap, seconds (default 600)
|
||
# GBRAIN_TEST_MAX_CONCURRENCY passed through to bun test (default 4)
|
||
#
|
||
# Output files (workspace-local; falls back to /tmp if .context/ unwritable):
|
||
# .context/test-failures.log failure blocks (cleared at start)
|
||
# .context/test-summary.txt per-shard pass/fail/skip/duration (cleared at start)
|
||
# .context/test-shards/ per-shard logs + exit codes (cleared at start)
|
||
|
||
set -uo pipefail
|
||
|
||
cd "$(dirname "$0")/.."
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# CPU detection: Apple Silicon perf cores → Mac total physical → nproc → 4.
|
||
# Returns a single positive integer.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
detect_cpus() {
|
||
local n=""
|
||
n=$(sysctl -n hw.perflevel0.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
|
||
n=$(sysctl -n hw.physicalcpu 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
|
||
n=$(nproc 2>/dev/null) && [ -n "$n" ] && [ "$n" -gt 0 ] && echo "$n" && return
|
||
echo 4
|
||
}
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Argument parsing. --shards N override wins over $SHARDS; both are clamped.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
SHARDS_OVERRIDE=""
|
||
MAX_CONCURRENCY_OVERRIDE=""
|
||
DRY_RUN=0
|
||
while [ $# -gt 0 ]; do
|
||
case "$1" in
|
||
--shards) SHARDS_OVERRIDE="$2"; shift 2 ;;
|
||
--shards=*) SHARDS_OVERRIDE="${1#*=}"; shift ;;
|
||
--max-concurrency) MAX_CONCURRENCY_OVERRIDE="$2"; shift 2 ;;
|
||
--max-concurrency=*) MAX_CONCURRENCY_OVERRIDE="${1#*=}"; shift ;;
|
||
--dry-run) DRY_RUN=1; shift ;;
|
||
*) echo "ERROR: unknown arg: $1" >&2; exit 2 ;;
|
||
esac
|
||
done
|
||
|
||
N="${SHARDS_OVERRIDE:-${SHARDS:-$(detect_cpus)}}"
|
||
if ! printf '%s' "$N" | grep -qE '^[0-9]+$' || [ "$N" -lt 1 ]; then
|
||
echo "ERROR: invalid shard count: $N" >&2; exit 2
|
||
fi
|
||
# v0.40.10 flake-hardening: clamp default to 4 (was 8) to match CI's
|
||
# test-shard.sh fan-out. At 8-shard parallel on Apple Silicon we observed
|
||
# shard 5 SIGKILL during source-health.test.ts's PGLite migration replay —
|
||
# 8 parallel PGLite WASM inits contend severely on the lockfile, and the
|
||
# 92-migration replay × 8 simultaneous can wedge past even 900s. CI uses
|
||
# 4 and is stable. Trade ~2x wallclock for reliability + parity with CI's
|
||
# fan-out. Override via --shards N or SHARDS=N (still capped at 8).
|
||
[ "$N" -gt 8 ] && N=8
|
||
if [ -z "${SHARDS_OVERRIDE:-}" ] && [ -z "${SHARDS:-}" ] && [ "$N" -gt 4 ]; then
|
||
N=4
|
||
fi
|
||
|
||
INTRA_CONC="${MAX_CONCURRENCY_OVERRIDE:-${GBRAIN_TEST_MAX_CONCURRENCY:-4}}"
|
||
# v0.40.10 flake-hardening: bump per-shard cap 600 → 1500 (was 900). At
|
||
# 4-shard default each shard runs 159 files / ~2420 tests with internal
|
||
# wallclock 960-1020s. The 900s value (sized for 8-shard's ~80 files /
|
||
# 1100 tests at 620-770s) false-killed shard 1 at 900s even though it
|
||
# had completed in 968s. 1500s cap gives ~55% headroom over observed
|
||
# 4-shard wallclock; real hangs still hit it. Override via
|
||
# GBRAIN_TEST_SHARD_TIMEOUT=N.
|
||
SHARD_TIMEOUT="${GBRAIN_TEST_SHARD_TIMEOUT:-1500}"
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Output directories. Prefer workspace-local .context/, fall back to /tmp.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
LOG_DIR=""
|
||
if mkdir -p .context/test-shards 2>/dev/null; then
|
||
LOG_DIR=".context/test-shards"
|
||
FAILURES_LOG=".context/test-failures.log"
|
||
SUMMARY_FILE=".context/test-summary.txt"
|
||
else
|
||
LOG_DIR="/tmp/gbrain-test-shards-$$"
|
||
FAILURES_LOG="/tmp/gbrain-test-failures.log"
|
||
SUMMARY_FILE="/tmp/gbrain-test-summary.txt"
|
||
mkdir -p "$LOG_DIR" || { echo "ERROR: cannot create log dir" >&2; exit 2; }
|
||
fi
|
||
# Clear from prior run.
|
||
rm -f "$LOG_DIR"/shard-*.log "$LOG_DIR"/shard-*.exit "$LOG_DIR"/shard-*.wedged 2>/dev/null
|
||
: > "$FAILURES_LOG"
|
||
: > "$SUMMARY_FILE"
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Resolve `timeout` command. macOS without coreutils has neither; we degrade
|
||
# to bg-pid + sleep cap. For now, prefer gtimeout (brew coreutils) → timeout.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
TIMEOUT_BIN=""
|
||
if command -v gtimeout >/dev/null 2>&1; then TIMEOUT_BIN="gtimeout"
|
||
elif command -v timeout >/dev/null 2>&1; then TIMEOUT_BIN="timeout"
|
||
fi
|
||
|
||
START_TS=$(date +%s)
|
||
echo "[unit-parallel] N=$N shards | --max-concurrency=$INTRA_CONC | timeout=${SHARD_TIMEOUT}s | logs=$LOG_DIR" >&2
|
||
|
||
if [ "$DRY_RUN" = "1" ]; then
|
||
echo "[unit-parallel] dry-run: would spawn $N shards with the above settings."
|
||
for i in $(seq 1 "$N"); do
|
||
SHARD="$i/$N" bash scripts/run-unit-shard.sh --dry-run-list 2>/dev/null \
|
||
| sed "s|^| [s$i] |"
|
||
done
|
||
exit 0
|
||
fi
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Spawn shards. Each child captures its own exit code into a sentinel file
|
||
# so $? is recoverable per-shard (we never trust `wait`'s aggregate value).
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
SHARD_PIDS=()
|
||
for i in $(seq 1 "$N"); do
|
||
(
|
||
SHARD_LOG="$LOG_DIR/shard-$i.log"
|
||
if [ -n "$TIMEOUT_BIN" ]; then
|
||
"$TIMEOUT_BIN" "${SHARD_TIMEOUT}s" \
|
||
env SHARD="$i/$N" \
|
||
bash scripts/run-unit-shard.sh --max-concurrency="$INTRA_CONC" \
|
||
> "$SHARD_LOG" 2>&1
|
||
else
|
||
env SHARD="$i/$N" \
|
||
bash scripts/run-unit-shard.sh --max-concurrency="$INTRA_CONC" \
|
||
> "$SHARD_LOG" 2>&1 &
|
||
pid=$!
|
||
( sleep "$SHARD_TIMEOUT" && kill -TERM "$pid" 2>/dev/null && \
|
||
sleep 5 && kill -KILL "$pid" 2>/dev/null ) &
|
||
cap_pid=$!
|
||
wait "$pid" 2>/dev/null
|
||
kill "$cap_pid" 2>/dev/null
|
||
wait "$cap_pid" 2>/dev/null
|
||
fi
|
||
rc=$?
|
||
echo "$rc" > "$LOG_DIR/shard-$i.exit"
|
||
[ "$rc" = "124" ] && echo "WEDGED" > "$LOG_DIR/shard-$i.wedged"
|
||
) &
|
||
SHARD_PIDS+=($!)
|
||
done
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Heartbeat: every 10s, print per-shard progress to stderr by tailing logs
|
||
# and counting Bun's `(pass)` / `(fail)` / `(skip)` markers. Read-only.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# grep_count: returns 0 (single integer) if file is missing or zero matches,
|
||
# otherwise the match count. Avoids the `grep -c | echo 0` double-output bug
|
||
# where 0 matches produces a 2-line "0\n0" string that breaks arithmetic.
|
||
grep_count() {
|
||
local pattern="$1"; local file="$2"
|
||
if [ ! -f "$file" ]; then echo 0; return; fi
|
||
local n
|
||
n=$(grep -cE "$pattern" "$file" 2>/dev/null) || n=0
|
||
echo "${n:-0}"
|
||
}
|
||
|
||
# bun_summary_count: parses Bun's summary lines (one per `bun test` invocation
|
||
# inside a shard — there's only one when we pass an explicit file list).
|
||
# Looks for ` N pass` / ` N fail` / ` N skip` patterns and sums them across
|
||
# all summary blocks the shard emitted. `bun test` prints these near the end
|
||
# of its output. Format: leading whitespace + integer + space + label.
|
||
bun_summary_count() {
|
||
local label="$1"; local file="$2"
|
||
if [ ! -f "$file" ]; then echo 0; return; fi
|
||
awk -v label="$label" '
|
||
$1 ~ /^[0-9]+$/ && $2 == label { total += $1 }
|
||
END { print total + 0 }
|
||
' "$file"
|
||
}
|
||
|
||
# shard_total_files: parse the "[unit-shard N/M] running X files" line that
|
||
# run-unit-shard.sh echoes before invoking bun test. Returns the file count
|
||
# the shard was given, or 0 if the line isn't there yet (shard still
|
||
# bootstrapping). Uses sed-then-grep so it's portable to macOS awk (BSD awk
|
||
# doesn't support `match($0, /re/, arr)` with the array sink — that's gawk-only).
|
||
shard_total_files() {
|
||
local file="$1"
|
||
[ -f "$file" ] || { echo 0; return; }
|
||
local n
|
||
n=$(sed -n 's/^\[unit-shard [0-9][0-9]*\/[0-9][0-9]*\] running \([0-9][0-9]*\) files.*/\1/p' "$file" 2>/dev/null | head -1)
|
||
echo "${n:-0}"
|
||
}
|
||
|
||
# shard_pglite_init_count: count "Schema version" lines as a proxy for "test
|
||
# files initialized so far." Each PGLite-using test file's beforeAll triggers
|
||
# one initSchema() which prints this. Undercounts because not every test file
|
||
# opens a PGLite engine, but it's the only real-time progress signal bun's
|
||
# default reporter leaves in the log (bun has no per-file progress markers,
|
||
# only a final shard-end summary).
|
||
shard_pglite_init_count() {
|
||
local file="$1"
|
||
[ -f "$file" ] || { echo 0; return; }
|
||
grep -cE 'Schema version [0-9]+ → [0-9]+' "$file" 2>/dev/null || echo 0
|
||
}
|
||
|
||
# log_size_kb: total stderr+stdout written by the shard so far. Strictly
|
||
# monotonic — useful as a "definitely alive" signal when other heuristics
|
||
# read 0 (e.g. very early in shard startup before initSchema fires).
|
||
log_size_kb() {
|
||
local file="$1"
|
||
[ -f "$file" ] || { echo 0; return; }
|
||
local b
|
||
b=$(wc -c < "$file" 2>/dev/null | tr -d ' ')
|
||
echo $(( ${b:-0} / 1024 ))
|
||
}
|
||
|
||
# fmt_elapsed: pretty-print seconds → "Mm:SS" or "SSs" for short.
|
||
fmt_elapsed() {
|
||
local s=$1
|
||
if [ "$s" -ge 60 ]; then
|
||
printf '%dm%02ds' $((s / 60)) $((s % 60))
|
||
else
|
||
printf '%ds' "$s"
|
||
fi
|
||
}
|
||
|
||
heartbeat() {
|
||
local hb_start=$(date +%s)
|
||
while true; do
|
||
sleep 10
|
||
local line=""
|
||
local now; now=$(date +%s)
|
||
local hb_elapsed=$((now - hb_start))
|
||
for i in $(seq 1 "$N"); do
|
||
if [ -f "$LOG_DIR/shard-$i.exit" ]; then
|
||
local rc; rc=$(cat "$LOG_DIR/shard-$i.exit" 2>/dev/null || echo "?")
|
||
local status="✓"
|
||
[ "$rc" != "0" ] && status="✗"
|
||
local f
|
||
f=$(bun_summary_count "fail" "$LOG_DIR/shard-$i.log")
|
||
local p
|
||
p=$(bun_summary_count "pass" "$LOG_DIR/shard-$i.log")
|
||
line="$line [s$i: done $status ${p}p ${f}f]"
|
||
else
|
||
local lf="$LOG_DIR/shard-$i.log"
|
||
if [ -f "$lf" ]; then
|
||
# Bun's default reporter has no per-file progress markers, only a
|
||
# final shard-end summary, so we surface three complementary signals
|
||
# mid-run: (1) PGLite initSchema() count as a "files started" proxy,
|
||
# (2) total files this shard was assigned (from the runner banner),
|
||
# (3) log size in KB as a strictly-monotonic liveness signal.
|
||
local total; total=$(shard_total_files "$lf")
|
||
local pglite; pglite=$(shard_pglite_init_count "$lf")
|
||
local kb; kb=$(log_size_kb "$lf")
|
||
local et; et=$(fmt_elapsed "$hb_elapsed")
|
||
if [ "$total" -gt 0 ]; then
|
||
line="$line [s$i: ~${pglite}/${total}f ${kb}KB ${et}]"
|
||
else
|
||
line="$line [s$i: starting ${kb}KB ${et}]"
|
||
fi
|
||
else
|
||
line="$line [s$i: spawning]"
|
||
fi
|
||
fi
|
||
done
|
||
printf '[heartbeat] %s\n' "$line" >&2
|
||
done
|
||
}
|
||
heartbeat &
|
||
HB_PID=$!
|
||
# v0.41.11.0 cleanup: pkill children FIRST, then kill heartbeat. If we
|
||
# kill the heartbeat shell first, its current `sleep 10` is reparented
|
||
# to init/launchd and pkill -P can no longer find it (orphan). Order:
|
||
# children first while the parent PID is still findable, then parent.
|
||
# Known bash quirk: SIGTERM to a shell sleeping inside `sleep` doesn't
|
||
# propagate to the sleep child before the wait returns. Without this,
|
||
# each invocation of this script leaks ONE orphan sleep; CI's "orphan
|
||
# process cleanup" at end-of-job reports them as (unnamed) test failures.
|
||
# Seen on the garrytan/port-pr-1406 PR, 2 CI runs in a row, 6 orphans
|
||
# matching the 6 invocations in test/scripts/run-unit-parallel.test.ts.
|
||
trap 'pkill -P "$HB_PID" 2>/dev/null; kill "$HB_PID" 2>/dev/null; wait "$HB_PID" 2>/dev/null' EXIT
|
||
|
||
# Wait for every shard. Don't care about wait's exit code.
|
||
for pid in "${SHARD_PIDS[@]}"; do wait "$pid" 2>/dev/null || true; done
|
||
|
||
pkill -P "$HB_PID" 2>/dev/null
|
||
kill "$HB_PID" 2>/dev/null
|
||
wait "$HB_PID" 2>/dev/null
|
||
trap - EXIT
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Aggregate failures (single writer; serial; never concurrent).
|
||
# Bun failure block format: from `(fail) ...` line through next `(pass)`,
|
||
# `(skip)`, blank line, or `__bun_test_summary__` marker.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
TOTAL_FAILURES=0
|
||
TOTAL_PASS=0
|
||
TOTAL_SKIP=0
|
||
TOTAL_RC=0
|
||
for i in $(seq 1 "$N"); do
|
||
SHARD_LOG="$LOG_DIR/shard-$i.log"
|
||
EXIT_FILE="$LOG_DIR/shard-$i.exit"
|
||
WEDGED_FILE="$LOG_DIR/shard-$i.wedged"
|
||
rc=1
|
||
[ -f "$EXIT_FILE" ] && rc=$(cat "$EXIT_FILE" 2>/dev/null || echo 1)
|
||
|
||
pass_count=$(bun_summary_count "pass" "$SHARD_LOG")
|
||
fail_count=$(bun_summary_count "fail" "$SHARD_LOG")
|
||
skip_count=$(bun_summary_count "skip" "$SHARD_LOG")
|
||
TOTAL_PASS=$((TOTAL_PASS + pass_count))
|
||
TOTAL_FAILURES=$((TOTAL_FAILURES + fail_count))
|
||
TOTAL_SKIP=$((TOTAL_SKIP + skip_count))
|
||
|
||
if [ -f "$WEDGED_FILE" ]; then
|
||
TOTAL_RC=1
|
||
{
|
||
echo "--- shard $i: WEDGED after ${SHARD_TIMEOUT}s ---"
|
||
[ -f "$SHARD_LOG" ] && tail -50 "$SHARD_LOG"
|
||
echo ""
|
||
} >> "$FAILURES_LOG"
|
||
echo "shard $i/$N: WEDGED after ${SHARD_TIMEOUT}s (rc=$rc)" >> "$SUMMARY_FILE"
|
||
continue
|
||
fi
|
||
|
||
echo "shard $i/$N: pass=$pass_count fail=$fail_count skip=$skip_count rc=$rc" >> "$SUMMARY_FILE"
|
||
|
||
if [ "$rc" != "0" ]; then
|
||
TOTAL_RC=1
|
||
if [ "$fail_count" -gt 0 ] && [ -f "$SHARD_LOG" ]; then
|
||
# Extract each (fail) block: from `(fail)` line through next `(pass)`,
|
||
# `(skip)`, blank line, or `__bun_test_summary__`. Single awk pass.
|
||
awk -v shard="$i" '
|
||
/^\(fail\) / { in_block=1; print "--- shard " shard ": " $0; next }
|
||
in_block {
|
||
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
|
||
print $0
|
||
}
|
||
' "$SHARD_LOG" >> "$FAILURES_LOG"
|
||
elif [ -f "$SHARD_LOG" ]; then
|
||
# Non-zero rc but no (fail) line found — extraction couldn't pinpoint.
|
||
# Dump the full shard log so we never silently lose the failure cause.
|
||
{
|
||
echo "--- shard $i: rc=$rc, no (fail) markers — full log follows ---"
|
||
cat "$SHARD_LOG"
|
||
echo ""
|
||
} >> "$FAILURES_LOG"
|
||
fi
|
||
fi
|
||
done
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Print each shard's full output to stdout (developer expects to scroll
|
||
# through it). Print summary file last for one-glance overview.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
for i in $(seq 1 "$N"); do
|
||
SHARD_LOG="$LOG_DIR/shard-$i.log"
|
||
echo ""
|
||
echo "════════════ shard $i/$N ════════════"
|
||
[ -f "$SHARD_LOG" ] && cat "$SHARD_LOG"
|
||
done
|
||
echo ""
|
||
echo "════════════ summary ════════════"
|
||
cat "$SUMMARY_FILE"
|
||
echo ""
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Serial pass: any *.serial.test.ts files run after parallel pass.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
SERIAL_RC=0
|
||
SERIAL_FILES_COUNT=0
|
||
SERIAL_FILES_COUNT=$(find test -name '*.serial.test.ts' -not -path 'test/e2e/*' 2>/dev/null | wc -l | tr -d ' ')
|
||
if [ "$SERIAL_FILES_COUNT" -gt 0 ]; then
|
||
echo "════════════ serial pass ($SERIAL_FILES_COUNT files) ════════════"
|
||
bash scripts/run-serial-tests.sh > "$LOG_DIR/serial.log" 2>&1
|
||
SERIAL_RC=$?
|
||
cat "$LOG_DIR/serial.log"
|
||
if [ "$SERIAL_RC" != "0" ]; then
|
||
TOTAL_RC=1
|
||
s_fail=$(bun_summary_count "fail" "$LOG_DIR/serial.log")
|
||
TOTAL_FAILURES=$((TOTAL_FAILURES + s_fail))
|
||
if [ "$s_fail" -gt 0 ]; then
|
||
awk '
|
||
/^\(fail\) / { in_block=1; print "--- shard serial: " $0; next }
|
||
in_block {
|
||
if (/^\(pass\)/ || /^\(skip\)/ || /^[[:space:]]*$/ || /__bun_test_summary__/) { in_block=0; print ""; next }
|
||
print $0
|
||
}
|
||
' "$LOG_DIR/serial.log" >> "$FAILURES_LOG"
|
||
else
|
||
{
|
||
echo "--- shard serial: rc=$SERIAL_RC, no (fail) markers — full log follows ---"
|
||
cat "$LOG_DIR/serial.log"
|
||
echo ""
|
||
} >> "$FAILURES_LOG"
|
||
fi
|
||
echo "serial: rc=$SERIAL_RC fail=$s_fail" >> "$SUMMARY_FILE"
|
||
else
|
||
s_pass=$(bun_summary_count "pass" "$LOG_DIR/serial.log")
|
||
TOTAL_PASS=$((TOTAL_PASS + s_pass))
|
||
echo "serial: pass=$s_pass rc=0" >> "$SUMMARY_FILE"
|
||
fi
|
||
fi
|
||
|
||
END_TS=$(date +%s)
|
||
ELAPSED=$((END_TS - START_TS))
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# Loud banner if anything failed. To stderr so it survives `| head`/`| tail`.
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
if [ "$TOTAL_RC" != "0" ]; then
|
||
ABS_FAIL=$(cd "$(dirname "$FAILURES_LOG")" && pwd)/$(basename "$FAILURES_LOG")
|
||
{
|
||
echo ""
|
||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||
echo "❌ $TOTAL_FAILURES TEST FAILURES — full details:"
|
||
echo " $ABS_FAIL"
|
||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||
tail -30 "$FAILURES_LOG"
|
||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||
echo "[unit-parallel] elapsed=${ELAPSED}s | pass=$TOTAL_PASS fail=$TOTAL_FAILURES skip=$TOTAL_SKIP"
|
||
} >&2
|
||
exit 1
|
||
fi
|
||
|
||
echo "[unit-parallel] elapsed=${ELAPSED}s | pass=$TOTAL_PASS fail=$TOTAL_FAILURES skip=$TOTAL_SKIP" >&2
|
||
exit 0
|