src/core/eval-capture-graph.ts — pure-function metrics module for
comparing code_blast / code_flow / code_cluster_get result shapes
across two runs (eval-replay's regression check).
Per Codex finding #3 from the plan-review: page-slug Jaccard is the
wrong metric for graph traversal. v0.34 W7 ships proper per-op metrics:
- nodeSetJaccard(a, b): set Jaccard over (file, line, symbol)
tuples. Right metric for code_blast/code_flow node sets.
- depthGroupStability(a, b): 1 - (displaced / |union|). Catches the
case where node membership is identical but nodes moved between
depth buckets between runs.
- truncationMatch(a, b): boolean match on the truncation enum.
Discrete signal that pairs with Jaccard.
- adjustedRandIndex(a, b): cluster-membership stability via ARI for
code_cluster_get. v0.34.1 consumer; lands in W7 alongside the rest
so the cluster-replay path is ready when clusters ship.
- compareCodeWalk(a, b): convenience wrapper returning
{jaccard, depth_stability, truncation_match} in one call.
Hermetic — no engine, no DB, fully unit-testable. 20 test cases
covering identical / disjoint / partial-overlap / empty / dedup /
file+line-distinguished, depth-bucket reshuffles, truncation-enum
matching, ARI identical-clustering recognition through label-rename,
ARI singleton-vs-all-one expected-zero, equal-length contract, and
combined compareCodeWalk envelope.
Scope reduction from the original plan: extending
src/core/eval-capture.ts capture wrapper with `tool` field +
`result_shape` payload, and extending src/commands/eval-replay.ts to
dispatch on tool — both deferred to v0.34.1. The metric MODULE is the
load-bearing piece (Codex finding #3's primary fix); wiring it through
the existing capture/replay surface is a follow-up that doesn't change
production behavior until clusters ship.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>