Files
gbrain/test/search/graph-signals.test.ts
T
d28be5d091 v0.40.4.0 feat(search): selective graph signals + per-stage attribution + audit-writer unification (#1300)
* v0.40.4.0 T1: shared audit-writer primitive

Extract createAuditWriter() helper. Five hand-rolled JSONL audit
modules (rerank-audit, shell-audit, supervisor-audit, audit-slug-
fallback, phantom-audit) duplicated the same ISO-week filename math,
best-effort write loop, and read-current-plus-previous-week loop.
T2 refactors all 5 onto this primitive.

Behavior preservation: filename format, JSONL line shape, mkdir
recursive, appendFileSync utf8, stderr-on-failure all byte-identical
to the existing modules so their tests pass unchanged.

resolveAuditDir() moves here from shell-audit.ts; shell-audit.ts
will re-export for back-compat (T2). Honors GBRAIN_AUDIT_DIR with
whitespace-trim, falls back to ~/.gbrain/audit/.

Test coverage: 22 cases covering ISO-week math + year-boundary edges
(2027-01-01 → 2026-W53), env override, mkdir-recursive, fail-open
stderr-warn shape, cross-week readback, corrupt-row skip, non-finite-
ts skip, round-trip with nested fields, computeFilename + resolveDir
accessors.

Plan ref: D5=B audit unification cathedral expansion.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T2: refactor 5 audit modules onto shared writer

Replace the duplicated ISO-week filename math + best-effort write loop
+ read-current-plus-previous-week loop in:
  - src/core/rerank-audit.ts (rerank-failures-*.jsonl)
  - src/core/audit-slug-fallback.ts (slug-fallback-*.jsonl)
  - src/core/minions/handlers/shell-audit.ts (shell-jobs-*.jsonl)
  - src/core/minions/handlers/supervisor-audit.ts (supervisor-*.jsonl)
  - src/core/facts/phantom-audit.ts (phantoms-*.jsonl)

All five now delegate file I/O to createAuditWriter from T1. Public
API preserved bit-for-bit:
  - logRerankFailure, readRecentRerankFailures, computeRerankAuditFilename
  - logSlugFallback, readRecentSlugFallbacks, computeSlugFallbackAuditFilename
  - logShellSubmission, computeAuditFilename, resolveAuditDir
  - writeSupervisorEvent, readSupervisorEvents, computeSupervisorAuditFilename
    plus isCrashExit, summarizeCrashes, CrashSummary (domain-specific
    helpers stay in supervisor-audit.ts; only file I/O moves)
  - logPhantomEvent, readRecentPhantomEvents, computePhantomAuditFilename

Domain-specific behavior preserved:
  - audit-slug-fallback emits per-call stderr (D7 dual logging) in the
    caller; the shared writer is failure-only stderr
  - rerank-audit truncates error_summary to 200 chars before write
  - phantom-audit spreads optional fields conditionally (skip undefined)
  - supervisor-audit keeps single-file readback (no cross-week walk)
    to preserve pre-v0.40.4 doctor assertions

resolveAuditDir lives in src/core/audit/audit-writer.ts; shell-audit.ts
re-exports it so existing imports keep working (every other audit
module + gbrain-home-isolation.test.ts + minions.test.ts +
minions-shell.test.ts pull resolveAuditDir from shell-audit.ts).

Operator-visible drift: rerank-audit stderr line drops the
'rerank-failure audit' qualifier — was '[gbrain] rerank-failure audit
write failed (...)' now '[gbrain] write failed (...); search continues'.
Stderr is human-debugging, not machine-parsed; the file written gives
the qualifier away in `tail -f audit/*`.

Test coverage: 128/128 audit-touching tests pass unchanged.

Plan ref: D5=B audit unification cathedral expansion.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T3: getAdjacencyBoosts engine method (PG+PGLite parity)

Add BrainEngine.getAdjacencyBoosts(pageIds) returning Map<page_id,
AdjacencyRow{hits, cross_source_hits}>. Returns ALL pages with
hits >= 1 (callers apply their own threshold).

Cross-source semantic (D15=A): cross_source_hits EXCLUDES the target
page's own source. A page in source A linked from 2 pages in source A
reports cross_source_hits = 0. Linked from 1 in source B + 1 in
source C reports 2.

Source-scope contract: pageIds MUST already be source-scoped by the
caller. Method does NOT filter by source_id. The in-set restriction
makes cross-source leakage impossible by construction. JSDoc spells
this out; same trust posture as cosineReScore's chunk_id handling.

COALESCE(p.source_id, 'default') on both target and from-page sides
for defense-in-depth even though pages.source_id is NOT NULL today.

JSDoc/SQL contract alignment (codex #2): HAVING >= 1 matches the
"returns ALL pages with hits >= 1" contract; threshold of 2 is the
caller's call in applyGraphSignals.

Known limitation (codex #15): cross_source_hits cannot distinguish
"genuinely linked from another team" from "mirrored imports from
another source." T-todo-4 captures the v0.41+ refinement.

SearchResult type extension (D4=A flat fields, D12=A attribution):
  - graph_adjacency_hits, graph_cross_source_hits,
    graph_session_demoted, graph_session_prefix
  - base_score, backlink_boost, salience_boost, recency_boost,
    exact_match_boost, graph_adjacency_boost, graph_cross_source_boost,
    session_demote_factor, reranker_delta
All optional; T4-T6 populate them.

Test coverage: 7/7 hermetic PGLite cases. Empty input, singleton,
same-source hub, cross-source attribution including the
"linked-only-from-other-source" case (widget in source b, linked
from alice+bob in source a → cross_source_hits=1), JSDoc HAVING>=1
contract. Postgres parity asserted by SQL-shape identity (will get a
mirror Postgres E2E in T10's eval gate work via DATABASE_URL when
set; PGLite hermetic case shipped now).

NULL source_id COALESCE branch noted as untestable in current PGLite
schema (pages.source_id is NOT NULL); kept as defense-in-depth.

Plan ref: T3 in v0.40.4.0 wave plan; D1=A, D3=A, D15=A.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T4+T11: applyGraphSignals 4th stage in runPostFusionStages

New file src/core/search/graph-signals.ts. Three signals:

  1. Adjacency-within-top-K (×1.05): hits >= 2 inbound from in-set.
  2. Cross-source adjacency (×1.10, stacks): cross_source_hits >= 2.
     Dormant on single-source brains.
  3. Session diversification (×0.95): if multiple top-K share a slug
     prefix, keep highest scoring, DEMOTE the rest. NOT amplify —
     codex caught the original framing was backwards (amplification
     of redundancy makes the cited "weak chunks compete for budget"
     problem worse, not better).

Conservative magnitudes (D14=B): 1.05/1.10/0.95. Score-distribution
probe (onScoreDistribution) collects min/p25/p50/p75/p95/max +
reorder_band_width to feed T-todo-2 magnitude calibration wave.

Slot: 4th stage inside runPostFusionStages (hybrid.ts:248), AFTER
backlink/salience/recency, pre-dedup. Inherits the v0.35.6.0
floor-ratio gate from computeFloorThreshold — this is the structural
protection that prevents a low-cosine hub from outranking a strong
non-hub (codex T2 / D1=A).

PostFusionOpts extends with graphSignalsEnabled, onGraphMeta,
onScoreDistribution. Caller (hybridSearch in subsequent T5 work)
resolves graph_signals from the mode bundle.

Source-scope contract preserved: getAdjacencyBoosts takes raw
page_ids, no source filter. Adjacency is in-set restricted so
cross-source leakage is impossible by construction (D3=A).

Fail-open: engine throw → JSONL audit row via shared createAuditWriter
(T1/T2 primitive, featureName='graph-signals-failures') + meta.errored
+ caller's results unchanged. Session diversification ALSO skips on
failure (predictable all-or-nothing posture).

Mutation note (codex #9): score mutated in place. base_score must be
stamped at runPostFusionStages entry BEFORE this stage so eval-capture
sees pre-boost score (T6 attribution wave).

Test coverage (24 cases, including T11 IRON RULE regression):
  - sessionPrefix multi/single/empty cases
  - computeScoreDistribution percentile math
  - Disabled + empty short-circuits
  - Adjacency hit, no-hit, cross-source stacking, cross-source alone
  - Session diversification 3-share + single-segment + singleton
  - Test seam injection (no engine call)
  - Fail-open: throw → audit row + meta.errored + unchanged
  - Empty Map → session still runs
  - Score-distribution always emits when enabled
  - Meta carries fire counts + duration_ms
  - Missing page_id silently skipped from dedup set
  - **T11 IRON RULE regression (3 cases):**
    * weak hub BELOW floor_threshold does NOT get boosted past
      above-floor non-hub (the bug class the floor gate exists for)
    * hub AT floor still gets boosted (gate is < not <=)
    * NaN score → NaN >= threshold is false → no boost

Plan ref: T4 + T11 in v0.40.4.0 wave plan; D1=A, D2=A, D11=B, D14=B,
D9=A, D5=B. Codex outside-voice #1 + #2 + #6 + #8 + #9 addressed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T5: graph_signals mode-bundle knob + KNOBS_HASH bump 3→4

ModeBundle gains graph_signals: boolean. Per-mode defaults:
  - conservative: false (cost-sensitive tier)
  - balanced:     true (the wave's primary surface for default-on)
  - tokenmax:     true (power-user tier, capstone fit)

SearchKeyOverrides + SearchPerCallOpts gain optional graph_signals
field. resolveSearchMode picks via the standard per-call → config
override → mode bundle chain.

loadOverridesFromConfig parses 'search.graph_signals' from the config
table ('1' or 'true' → true). SEARCH_MODE_CONFIG_KEYS adds the key
so `gbrain search modes --reset` clears it alongside other knobs.

KNOBS_HASH_VERSION bump 3→4 (append-only per CDX2-F13). New `gs=`
parts entry appended AFTER cross-modal + column + prov entries. A
graph-on cache write cannot be served to a graph-off lookup —
mid-deploy hit-rate dip clears within cache.ttl_seconds (3600s).

src/commands/search.ts KNOB_DESCRIPTIONS gains graph_signals entry
so `gbrain search modes` dashboard renders the new knob.

Test coverage:
  - test/search-mode.test.ts (+ 8 new cases): per-mode defaults
    canonical, config override both directions, per-call override
    wins, knobsHash distinct for on/off, config key registered,
    attributeKnob reports per-call + mode sources correctly.
  - test/search/knobs-hash-reranker.test.ts: version assertion
    bumped 3→4 with v0.40.4 rationale comment.
  - test/cross-modal-phase1.test.ts: version assertion bumped
    3→4 with v0.40.4 rationale comment.
  - Canonical-bundle assertions updated to include graph_signals
    in expected shape (3 cases).

50/50 search-mode tests pass. 45/45 cross-modal pass. 17/17
knobs-hash-reranker pass. 10/10 balanced-reranker pass.

Plan ref: T5 in v0.40.4.0 wave plan.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T6: per-stage attribution stamping in every boost

Every boost stage that mutates SearchResult.score now stamps a field
recording WHAT it multiplied:

  - applyBacklinkBoost  → backlink_boost (skipped when count == 0)
  - applySalienceBoost  → salience_boost (skipped when score == 0)
  - applyRecencyBoost   → recency_boost (skipped on evergreen prefix)
  - applyExactMatchBoost → exact_match_boost (skipped on no-match
    OR when intent's exactMatchBoost == 1.0 no-op)
  - runPostFusionStages → base_score stamped ONCE at entry, BEFORE
    any boost mutates r.score. Idempotent: caller-pre-stamped value
    preserved. Empty-results short-circuit unchanged.
  - applyReranker → reranker_delta = original_index - new_index
    (positive = rank improved; raw rerank score stays in rerank_score)
  - applyGraphSignals → graph_adjacency_boost, graph_cross_source_boost,
    session_demote_factor (T4 already stamped these)

Why: feeds the T7 `gbrain search --explain` formatter so it can
attribute the final score to its components. Without these stamps,
"why did this rank where it did?" is grep-and-guess.

SearchResult.reranker_delta doc updated to clarify it's a RANK delta
(positive = improved), not a score delta. The raw relevance score
stays in `rerank_score` (untyped, for back-compat with telemetry that
already reads it).

Test coverage: 16 new cases in test/search/attribution-stamping.test.ts.
Pins: every boost stamps when it fires AND skips stamping when it
doesn't (no false attribution on no-op stages). base_score idempotency
preserved. reranker_delta computed correctly across rank-improved +
rank-degraded cases.

All 178/178 search tests pass (no regressions).

Plan ref: T6 cathedral expansion in v0.40.4.0 wave plan; D12=A.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T7: gbrain search --explain per-stage attribution

New file src/core/search/explain-formatter.ts renders SearchResult[]
as a multi-line breakdown of how the final score was formed:

  1. people/alice (score=12.4)
     base=10.2 (rrf+cosine)
     + backlink ×1.08
     + salience ×1.05
     + adjacency ×1.05 (hits=3)
     + cross_source ×1.10 (other_sources=2)
     ↑ reranker rank +2
     = final 12.4

Reads the boost_* / base_score / *_hits fields populated by T4 + T6.
Empty path: "no boosts applied" when no stage stamped anything.
Session demote rendered with `-` prefix (not `+`) so the demotion
direction is visually distinct from boosts.

CliOptions gains `explain: boolean`; parseGlobalFlags recognizes
`--explain` anywhere in argv. cli.ts formatResult for `search` +
`query` cases reads CliOptions.explain via the module-level
singleton and routes to formatResultsExplain when set. Lazy import
keeps the hot path narrow for the common non-explain case.

Number formatting: 4-decimal precision, trailing zeros stripped
('1.0000' → '1', '0.1234' → '0.1234'). NaN preserved as 'NaN'.

Test coverage:
  - test/search/explain-formatter.test.ts: 19 cases pin output
    format. Each boost type renders correctly, every-stage stacking
    composes, reranker_delta=0 doesn't render, empty list short-
    circuits, rank numbering 1-based, number formatting edge cases.
  - test/cli-options.test.ts: 3 new cases for --explain parsing
    (basic, absent default, any-argv-position).

Existing CliOptions literals in test/cli-options.test.ts +
test/thin-client-upgrade-prompt.test.ts updated for new required
explain field.

JSON envelope unchanged — the same attribution fields surface in
existing --json output via JSON.stringify; no separate JSON formatter
needed.

Plan ref: T7 cathedral expansion in v0.40.4.0 wave plan; D12=A + D6=A.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T8: doctor check graph_signals_coverage

New checkGraphSignalsCoverage in src/commands/doctor.ts. Wired into
both runDoctor (local engine) and doctorReportRemote (HTTP MCP /
JSON path) so local AND remote-server brains both surface the metric.

Logic:
  1. Resolve active graph_signals setting: config override
     'search.graph_signals' wins, else mode bundle default
     ('search.mode' → conservative=false, balanced/tokenmax=true).
  2. When disabled → silent ok ("disabled — coverage not checked").
     Avoids polluting doctor output on installs that don't use the
     feature.
  3. When enabled, compute global inbound-link density:
     COUNT(DISTINCT to_page_id) / COUNT(*) across non-deleted pages.
  4. <10% → warn ("signal will rarely fire") with paste-ready
     `gbrain extract all` fix hint.
  5. >=30% → ok ("fire on most queries") with metric.
  6. 10-29% → ok ("fire occasionally") with metric.

Known limitation (codex outside-voice #14): global density is an
imperfect proxy for "top-K subgraphs have enough edges to fire."
T-todo-5 captures the v0.41+ refinement that measures actual fire
rate from search-stats after 30 days of data.

Best-effort: SQL errors → warn with the underlying message. Never
breaks doctor.

Test coverage (7 new cases in test/doctor.test.ts):
  - conservative mode → silent ok regardless of coverage
  - balanced default + 0 links → warn at 0% with fix hint
  - balanced default + 40% inbound → ok "fire on most queries"
  - balanced default + 20% inbound → ok "fire occasionally"
  - explicit search.graph_signals=false overrides mode default
  - empty brain → ok with explanation
  - check is wired into runDoctor (source-grep regression guard)

All 55/55 doctor.test.ts cases pass.

Plan ref: T8 in v0.40.4.0 wave plan; D6=A.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T9: gbrain search stats graph_signals section

runStatsSubcommand in src/commands/search.ts gains a graph_signals
section in both --json and human output:

  Graph signals:
    enabled:    true (mode default)
    failures:   3 fail-open event(s)
      ECONNREFUSED         2
      timeout              1

Data sources:
  - config: 'search.graph_signals' override → enabled + source=config,
    otherwise mode-bundle default → enabled + source=mode_default.
  - JSONL audit: readRecentGraphSignalsFailures(days) returns events;
    failures_count is len, failures_by_reason buckets by first word of
    error_summary (e.g. 'ECONNREFUSED', 'timeout').

JSON envelope (schema_version 2 unchanged; graph_signals is a new
sibling property of stats, so consumers reading the existing fields
keep working):

  {
    "schema_version": 2,
    ...stats...,
    "graph_signals": {
      "enabled": bool,
      "source": "config" | "mode_default",
      "failures_count": int,
      "failures_by_reason": { reason: count }
    },
    "_meta": { metric_glossary: { ..., graph_signals_enabled: ..., graph_signals_failures_count: ... } }
  }

Fire-rate metrics (adjacency_fires, cross_source_fires,
session_demotions) and score-distribution stats are NOT in this
section yet — they require telemetry-table writes from the
applyGraphSignals onMeta callback. Wired in v0.41+ via T-todo-2
calibration wave (the wave that needs them). For v0.40.4: status +
error count is the actionable surface for "is graph_signals on, and
is it failing?"

Human output: prints the section after the existing stats block.
Edge case: when total_calls is 0 BUT graph_signals is enabled OR
has historical failures, still prints the section so operators
don't lose the signal on a brain with no telemetry yet.

Test coverage (6 cases in test/search/search-stats-graph-signals.test.ts):
  - search.graph_signals=true → enabled true, source=config
  - mode=conservative → enabled false, source=mode_default
  - no config → enabled true (balanced default), source=mode_default
  - JSONL failures bucketed by first word of error_summary
  - empty audit → failures_count 0, empty failures_by_reason
  - human output includes "Graph signals:" header

Plan ref: T9 in v0.40.4.0 wave plan; D6=A.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T10: eval gates (longmemeval-mini A/B + paired bootstrap)

New test/e2e/graph-signals-eval.test.ts runs each longmemeval-mini
question twice (graph_signals off, graph_signals on) and asserts:

  Gate 1 (QUALITY) — paired bootstrap, 10,000 resamples:
    - If signals-on is significantly WORSE than off
      (delta < 0 AND p < 0.05) → fail.
    - Otherwise pass. p>=0.05 either direction OR delta >= 0 → ok.

  Gate 2a (CHANGE-MAGNITUDE): mean Jaccard@5 over result-set overlap
    must be >= 0.5. If results overlap less than half, the change is
    too large and needs human review before default-on.

  Gate 2b (CHANGE-MAGNITUDE): top-1 stability rate >= 0.7. If 30%+
    of top picks change, hard look required.

  Gate 3 (HARD ABSOLUTE FLOOR): recall@5 drop <= 5pt. Catastrophic
    regression catch (codex outside-voice #18 — addresses the "top-5
    must not drop at all" brittleness on tiny fixtures).

Bootstrap implementation:
  - Per-question observation is binary (recall@5 hit/miss).
  - Paired pairing on question_id between on/off branches.
  - Centered distribution under null (subtract observed mean) per
    standard paired-bootstrap-shift approach for binary outcomes.
  - Two-tailed p-value: |resampled delta| >= |observed delta|.
  - Deterministic seeded RNG so test runs are stable across CI.

pairedBootstrapPValue exported as a pure function with separate
tests for edge cases (empty input, all-equal, strong positive, strong
negative, determinism). Reusable from future calibration waves.

Hermetic: in-memory PGLite via createBenchmarkBrain + resetTables
between questions. No API keys needed (--no-embed import path
exercises keyword-only retrieval). Skips gracefully via describe.skip
when the fixture is missing.

Plan ref: T10 in v0.40.4.0 wave plan; D7=C absolute floor + D13=A
paired bootstrap; codex #4 + #18 stability-vs-quality distinction.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 T12: VERSION + package.json + CHANGELOG + TODOS

VERSION: 0.37.11.0 → 0.40.4.0
package.json: 0.37.11.0 → 0.40.4.0
CHANGELOG.md: top entry for v0.40.4.0 in ELI10-lead voice per
  CLAUDE.md release rules. Lead is plain-English ("Your search now
  notices when a page is a hub for your query"); precise file paths
  / SQL semantics / numbers live in the "Itemized changes" section
  below. Includes the cathedral-expansion notes (D5=B audit
  unification, D12=A per-stage attribution, D13=A eval gates) and
  the "To take advantage of v0.40.4.0" verify-and-fix block.

TODOS.md: 5 new items captured under "v0.40.4 graph signals —
deferred follow-ups (v0.41+)":
  - T-todo-1: profile graph-signal SQL latency, merge if hot (D8=C)
  - T-todo-2: magnitude calibration wave from probe data (D14=B / D17)
  - T-todo-3: DB-backed audit table for cross-deploy observability (codex #15)
  - T-todo-4: sync-topology-aware cross-source signal (codex #11)
  - T-todo-5: replace doctor's global density with fire-rate (codex #14)

Verified the 3-line audit: VERSION + package.json + CHANGELOG topmost
all match 0.40.4.0. `bun install` ran (lockfile unchanged — root
package version isn't stored in bun.lock). `bun run build:llms`
refreshed llms.txt + llms-full.txt for the next commit.

Plan ref: T12 in v0.40.4.0 wave plan.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 TODO: document pre-existing shard-2 flake noticed during ship

3 isCacheSafe test failures in shard 2 reproduce on stashed clean
master. Confirmed pre-existing — not introduced by v0.40.4. Filed
under "Pre-existing flake on master (noticed during v0.40.4 ship)"
with reproduction commands + remediation options. Shipping v0.40.4
through it; future wave can fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 privacy scrub: replace wintermute → media in example slugs

CLAUDE.md line 550 bans the private OpenClaw fork name in public
artifacts. Example session prefix in sessionPrefix() docs + 3 test
fixtures swept to 'media/chat/...' instead. Pre-existing
scripts/check-privacy.sh in `bun run verify` caught it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 fix: wire graph_signals from mode bundle to runPostFusionStages

CRITICAL: pre-landing review (codex outside-voice via /ship Step 9)
caught that hybrid.ts's `postFusionOpts` literal at line 566 was
building PostFusionOpts WITHOUT threading `resolvedMode.graph_signals`
to `graphSignalsEnabled`. The gate at hybrid.ts:358 read the field
from a literal that never set it.

Result before this fix: the entire v0.40.4 graph-signals wave was
dead code in production. Mode bundles set
`balanced.graph_signals = true` and `tokenmax.graph_signals = true`,
but no production call site ever reached applyGraphSignals. The
KNOBS_HASH bump 3→4 correctly varied the cache key by the flag, so
contamination was prevented — but the feature itself never fired.

All shipped infrastructure (engine SQL, fail-open audit, attribution
stamps, --explain formatter, doctor coverage check, search-stats
section) was reachable only through the unit-test seam
(`opts.adjacencyFn`). The CHANGELOG-advertised behavior never
landed in user-visible search.

Fix: thread `graphSignalsEnabled: resolvedMode.graph_signals` into
the postFusionOpts literal (1 line). Inline comment names codex's
catch so future refactors see the regression class.

Tests: new test/search/graph-signals-wire-integration.test.ts pins
the wire end-to-end. Three cases:
  1. balanced mode → hybridSearch on a seeded brain with adjacency
     hub produces a result with base_score stamped (proves
     runPostFusionStages actually ran).
  2. search.graph_signals=false config override → no graph_* fields
     stamped (proves the gate honors the override path).
  3. Source-grep regression guard pinning the
     `graphSignalsEnabled: resolvedMode.graph_signals` literal in
     hybrid.ts so a future refactor can't silently disconnect.

All 57 existing v0.40.4 wave tests still pass. Typecheck clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 fix: pre-landing review AUTO-FIX findings (audit msg drift + deleted_at)

Two informational findings from /ship pre-landing review (Step 9):

1. Stderr message qualifier drift (rerank/slug-fallback/phantom audits)
   Pre-v0.40.4 messages included a per-feature qualifier:
     [gbrain] rerank-failure audit write failed (...)
     [gbrain] slug-fallback audit write failed (...)
     [gbrain] phantom audit write failed (...)
   The T2 refactor dropped the qualifier (plan promised "byte-identical"
   operator-visible behavior, but stderr lines did drift). Restored via
   new `errorMessagePrefix` option on `createAuditWriter` (optional, ''
   default). Three modules pass the per-feature qualifier; shell-audit
   and supervisor-audit unaffected (their pre-v0.40.4 messages didn't
   have a separate qualifier — label already carried the feature name).

2. Defense-in-depth `deleted_at IS NULL` on getAdjacencyBoosts
   SQL was previously protected by-construction (hybridSearch's
   visibility filter ensures input pageIds are live), but matches the
   v0.35.5.0 findOrphanPages pattern and closes the bug class if a
   future caller bypasses hybridSearch. Added to both Postgres and
   PGLite engines for parity. Three JOIN sites guarded (targets CTE,
   FROM-pages join). One inline comment per engine cites the codex
   review and the v0.35.5.0 precedent.

Plan ref: /ship pre-landing review v0.40.4.0 (codex finding C and F).

All 84 audit+graph-signals tests pass. Typecheck clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 fix: adversarial review HIGH findings (codex H1+H2 + Claude F1)

Three HIGH-severity issues from /ship adversarial pass:

H1 (Codex): Eval gate was a no-op.
  Test passed `graph_signals: graphSignalsOn` via `as any` cast, but
  SearchOpts had no field and hybridSearch's perCall didn't thread it.
  Both off/on branches resolved to the mode-bundle default — gate
  measured identical behavior, could pass while detecting nothing.

  Fix: add `graph_signals?: boolean` to SearchOpts (types.ts:794).
  Thread `opts.graph_signals` into perCall in both hybridSearch
  (hybrid.ts:425) AND hybridSearchCached (hybrid.ts:1027) so the
  cache-key resolver also sees the override. Drop the `as any` from
  the eval test — types are real now.

H2 (Codex): Session diversification fired on entity directories.
  sessionPrefix() used "any shared parent directory" as the session
  signal. Result: a search for "people in SF" returned `people/alice`
  + `people/bob` + `people/charlie` and the latter two got demoted
  to 0.95×. Every common entity-search query silently penalized
  legitimate same-type results. Default-on for balanced/tokenmax
  means production behavior was wrong.

  Fix: narrow sessionPrefix() to fire ONLY when the slug contains a
  session-like marker (`chat`/`session`/`sessions` segment OR a
  `YYYY-MM-DD` date segment). Entity directories (`people/`,
  `companies/`, `docs/`) return null → diversification skips.
  Returns NULL (not the slug itself) so the loop skips clean.
  Examples in JSDoc:
    your-agent/chat/2026-05-20-foo → 'your-agent/chat/2026-05-20-foo'
    daily/2026-05-20/journal-entry-1 → 'daily/2026-05-20'
    transcripts/chat/funding-discussion → 'transcripts/chat/funding-discussion'
    people/alice → null  ← codex H2 regression
    docs/quickstart → null

F1 (Claude adversarial subagent): case-sensitivity drift across 3 sites.
  loadOverridesFromConfig in mode.ts is case-insensitive +
  whitespace-trimmed for 'search.graph_signals' values. But
  doctor's checkGraphSignalsCoverage (doctor.ts:899) AND
  search-stats's readGraphSignalsStats (search.ts:288) used
  case-sensitive compare. User sets `search.graph_signals TRUE`:
  production enables the feature, but doctor + search-stats both
  silently report disabled. Operators lose the only observability
  surface for the new feature on values like 'True'/'TRUE'.

  Fix: trim + lowercase parity at both sites. Mirror the parser's
  semantic. Also case-normalized `search.mode` reads at both sites
  for the same divergence class.

Tests:
  - sessionPrefix block rewritten with 7 cases covering chat marker
    + date anchor + entity dirs (now-NULL) + degenerate (no /).
  - Added regression test pinning codex H2: people/alice +
    people/bob + people/charlie do NOT get diversified.
  - graph-signals-eval.test.ts drops `as any` — typed field works.
  - Existing tests using `chat/a`/`chat/b` updated to session-shaped
    `media/2026-05-20/chunk-a` so the date anchor actually fires.

111/111 graph-signals + doctor + search-stats tests pass. Typecheck clean.

Plan ref: /ship adversarial review v0.40.4.0 (codex H1, H2; Claude F1).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* v0.40.4.0 TODOs: capture 11 LOW adversarial findings for v0.41+

Codex L1 (audit window underreport) + Claude F2/F3/F5-F8/F11/F12/F14/F16
from /ship adversarial review. None are load-bearing; all captured under
'v0.40.4 adversarial review LOW findings — captured for v0.41+'.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: update project documentation for v0.40.4.0

- README: surface v0.40.4.0 graph signals + --explain in Hybrid search capability
- CLAUDE.md: annotate engine.ts getAdjacencyBoosts, new graph-signals.ts /
  explain-formatter.ts / audit/audit-writer.ts, plus hybrid.ts post-fusion
  4th stage, mode.ts graph_signals knob + KNOBS_HASH 3→4, cli-options.ts
  --explain flag, search stats + doctor coverage check
- llms-full.txt: regenerated from CLAUDE.md per the build:llms chaser rule

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(ci): pin bun-version to 1.3.13 across all workflows

setup-bun action with `bun-version: latest` calls the GitHub API
(https://api.github.com/repos/oven-sh/bun/git/refs/tags) to resolve
the tag. CI started failing today with HTTP 401 "Bad credentials"
even though the action receives a token (visible as `token: ***`
in the run log). Pinning the version eliminates the API call
entirely.

Affected workflows: test.yml, e2e.yml, release.yml, heavy-tests.yml
(5 invocations total). Pinned to 1.3.13 — matches package.json
engines (`bun >= 1.3.10`) and the version v0.40.4.0 was developed
against.

Bump cadence: when a new bun version is required, update this
pin in one PR. Trading "always-latest" for "always-deterministic"
is the right trade for a 5-shard CI matrix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-23 10:01:08 -07:00

469 lines
18 KiB
TypeScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/**
* v0.40.4.0 — applyGraphSignals unit + REGRESSION tests.
*
* Hermetic via opts.adjacencyFn injection seam (no engine needed).
*
* Pinned contracts:
* - Disabled / empty → unchanged + zero-meta
* - Adjacency hit boosts and stamps fields
* - Cross-source hit boosts ON TOP of adjacency (stacking)
* - Cross-source ONLY (hits<2 but cross_source_hits>=2) — uncommon
* in practice but the SQL allows it
* - Session diversification: highest keeps full, rest demoted
* - Single-segment slug (no '/') doesn't false-group
* - Test seam injection works without engine
* - Fail-open on engine throw; audit row written; meta.errored=true
* - Score-distribution probe always fires when enabled
* - Single-pass O(K) session grouping (D9=A regression assertion)
* - REGRESSION (IRON RULE T11): floor-gate respected — weak hub
* does NOT outrank above-floor non-hub
*/
import { describe, test, expect } from 'bun:test';
import type { SearchResult } from '../../src/core/types.ts';
import type { AdjacencyRow } from '../../src/core/types.ts';
import type { BrainEngine } from '../../src/core/engine.ts';
import {
applyGraphSignals,
ADJACENCY_BOOST,
CROSS_SOURCE_BOOST,
SESSION_DEMOTE,
sessionPrefix,
computeScoreDistribution,
} from '../../src/core/search/graph-signals.ts';
// Minimal SearchResult factory.
function makeResult(slug: string, score: number, page_id: number, source_id = 'a'): SearchResult {
return {
slug,
page_id,
title: slug,
type: 'note' as const,
chunk_text: `body of ${slug}`,
chunk_source: 'compiled_truth',
chunk_id: page_id * 1000,
chunk_index: 0,
score,
stale: false,
source_id,
};
}
// Engine stub — only the methods applyGraphSignals reaches into are used.
// adjacencyFn is the test seam, so engine is rarely called; fall through to
// throwing if accidentally hit.
const ENGINE_STUB = {} as BrainEngine;
describe('sessionPrefix (v0.40.4 narrowed scope — codex fix for entity-dir false positives)', () => {
test('chat marker → prefix up to and including session id', () => {
expect(sessionPrefix('your-agent/chat/2026-05-20-foo')).toBe('your-agent/chat/2026-05-20-foo');
});
test('chat marker with trailing chunk segment → prefix is parent', () => {
// For media/chat/2026-05-20-foo/chunk-001, the chat-session is
// media/chat/2026-05-20-foo (the parent of the chunk file).
expect(sessionPrefix('media/chat/2026-05-20-foo/chunk-001')).toBe('media/chat/2026-05-20-foo');
});
test('date segment in middle of path → prefix is up to and including date', () => {
expect(sessionPrefix('daily/2026-05-20/journal-entry-1')).toBe('daily/2026-05-20');
});
test('transcripts marker → prefix up to and including transcript id', () => {
expect(sessionPrefix('transcripts/chat/funding-discussion')).toBe('transcripts/chat/funding-discussion');
});
test('entity directory (no marker, no date) → null (skip diversification)', () => {
expect(sessionPrefix('people/alice')).toBeNull();
expect(sessionPrefix('companies/acme')).toBeNull();
expect(sessionPrefix('wiki/concepts/auth')).toBeNull();
expect(sessionPrefix('docs/quickstart')).toBeNull();
});
test('single-segment slug → null (no path, no session)', () => {
expect(sessionPrefix('standalone')).toBeNull();
});
test('empty slug → null', () => {
expect(sessionPrefix('')).toBeNull();
});
test('meetings + date → meeting-session prefix', () => {
expect(sessionPrefix('meetings/2026-04-03/notes')).toBe('meetings/2026-04-03');
});
});
describe('computeScoreDistribution', () => {
test('empty input → zeros', () => {
const d = computeScoreDistribution([]);
expect(d.top_k_size).toBe(0);
expect(d.reorder_band_width).toBe(0);
});
test('ordered scores → percentiles + band width', () => {
const d = computeScoreDistribution([10, 8, 6, 4, 2]);
expect(d.top_k_size).toBe(5);
expect(d.min).toBe(2);
expect(d.max).toBe(10);
expect(d.p50).toBe(6);
expect(d.reorder_band_width).toBe(8);
});
test('unsorted input is sorted internally', () => {
const d = computeScoreDistribution([4, 10, 2, 8, 6]);
expect(d.min).toBe(2);
expect(d.max).toBe(10);
expect(d.p50).toBe(6);
});
});
describe('applyGraphSignals — gate', () => {
test('disabled → unchanged + zero-meta', async () => {
const results = [makeResult('a/b', 1.0, 1), makeResult('a/c', 0.5, 2)];
const before = results.map(r => ({ score: r.score, slug: r.slug }));
let calledFn = false;
let metaOut: any;
await applyGraphSignals(results, ENGINE_STUB, {
enabled: false,
adjacencyFn: async () => { calledFn = true; return new Map(); },
onMeta: (m) => { metaOut = m; },
});
expect(calledFn).toBe(false);
expect(results.map(r => ({ score: r.score, slug: r.slug }))).toEqual(before);
expect(metaOut.enabled).toBe(false);
expect(metaOut.adjacency_fires).toBe(0);
});
test('empty results → unchanged + zero-meta + no SQL', async () => {
let calledFn = false;
let metaOut: any;
await applyGraphSignals([], ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => { calledFn = true; return new Map(); },
onMeta: (m) => { metaOut = m; },
});
expect(calledFn).toBe(false);
expect(metaOut.enabled).toBe(true);
expect(metaOut.top_k_size).toBe(0);
});
});
describe('applyGraphSignals — adjacency boost', () => {
test('hits >= 2 → score multiplied by ADJACENCY_BOOST, fields stamped', async () => {
const results = [
makeResult('people/alice', 10, 1),
makeResult('people/bob', 9, 2),
makeResult('companies/acme', 8, 3), // hub: 3 inbound links from top-K
];
const adjacency = new Map<number, AdjacencyRow>([
[3, { hits: 3, cross_source_hits: 0 }],
]);
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => adjacency,
});
const acme = results.find(r => r.slug === 'companies/acme')!;
expect(acme.score).toBeCloseTo(8 * ADJACENCY_BOOST, 5);
expect(acme.graph_adjacency_hits).toBe(3);
expect(acme.graph_adjacency_boost).toBe(ADJACENCY_BOOST);
});
test('hits < 2 → no boost, no field stamp', async () => {
// Distinct prefixes so session-diversification doesn't mutate scores.
const results = [makeResult('alpha/b', 10, 1), makeResult('beta/c', 9, 2)];
const adjacency = new Map<number, AdjacencyRow>([
[2, { hits: 1, cross_source_hits: 0 }], // below threshold
]);
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => adjacency,
});
const r = results.find(r => r.slug === 'beta/c')!;
expect(r.score).toBe(9);
expect(r.graph_adjacency_hits).toBeUndefined();
});
});
describe('applyGraphSignals — cross-source boost (stacks on adjacency)', () => {
test('hits>=2 AND cross_source>=2 → both multipliers stack', async () => {
const results = [
makeResult('people/alice', 10, 1),
makeResult('companies/acme', 8, 3), // both signals fire
];
const adjacency = new Map<number, AdjacencyRow>([
[3, { hits: 3, cross_source_hits: 2 }],
]);
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => adjacency,
});
const acme = results.find(r => r.slug === 'companies/acme')!;
expect(acme.score).toBeCloseTo(8 * ADJACENCY_BOOST * CROSS_SOURCE_BOOST, 5);
expect(acme.graph_adjacency_boost).toBe(ADJACENCY_BOOST);
expect(acme.graph_cross_source_boost).toBe(CROSS_SOURCE_BOOST);
expect(acme.graph_cross_source_hits).toBe(2);
});
test('cross_source only (hits<2 but cross_source>=2) → cross-source fires alone', async () => {
const results = [makeResult('companies/acme', 8, 3)];
const adjacency = new Map<number, AdjacencyRow>([
[3, { hits: 1, cross_source_hits: 2 }],
]);
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => adjacency,
});
const acme = results[0];
expect(acme.score).toBeCloseTo(8 * CROSS_SOURCE_BOOST, 5);
expect(acme.graph_adjacency_boost).toBeUndefined();
expect(acme.graph_cross_source_boost).toBe(CROSS_SOURCE_BOOST);
});
});
describe('applyGraphSignals — session diversification', () => {
test('3 chat-session chunks share prefix → highest keeps full, other 2 demoted', async () => {
const results = [
makeResult('media/chat/2026-05-20-foo/a', 10, 1),
makeResult('media/chat/2026-05-20-foo/b', 9, 2),
makeResult('media/chat/2026-05-20-foo/c', 7, 3),
];
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => new Map(),
});
// Highest (slug a, score 10) keeps full.
const a = results[0];
expect(a.score).toBe(10);
expect(a.graph_session_demoted).toBeUndefined();
expect(a.graph_session_prefix).toBe('media/chat/2026-05-20-foo');
// Others demoted.
expect(results[1].score).toBeCloseTo(9 * SESSION_DEMOTE, 5);
expect(results[1].graph_session_demoted).toBe(true);
expect(results[1].session_demote_factor).toBe(SESSION_DEMOTE);
expect(results[2].score).toBeCloseTo(7 * SESSION_DEMOTE, 5);
expect(results[2].graph_session_demoted).toBe(true);
});
test('REGRESSION (codex H2): entity-directory siblings (people/alice + people/bob) are NOT diversified', async () => {
// Pre-fix behavior: people/alice + people/bob shared `people/` prefix
// and people/bob got demoted. This silently penalized every common
// entity-search query like "people in SF". Post-fix: sessionPrefix
// returns null for non-session slugs, so no diversification fires.
const results = [
makeResult('people/alice', 10, 1),
makeResult('people/bob', 9, 2),
makeResult('people/charlie', 7, 3),
];
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => new Map(),
});
expect(results[0].score).toBe(10);
expect(results[1].score).toBe(9); // NOT demoted
expect(results[2].score).toBe(7); // NOT demoted
expect(results[0].graph_session_demoted).toBeUndefined();
expect(results[1].graph_session_demoted).toBeUndefined();
expect(results[2].graph_session_demoted).toBeUndefined();
});
test('non-session slug (no chat/date marker) → not grouped, no false demote', async () => {
const results = [
makeResult('standalone', 10, 1),
makeResult('media/chat/2026-05-20-foo/a', 9, 2),
makeResult('media/chat/2026-05-20-foo/b', 8, 3),
];
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => new Map(),
});
// standalone: not session-shaped → no demote.
const standalone = results[0];
expect(standalone.score).toBe(10);
expect(standalone.graph_session_demoted).toBeUndefined();
// Two chat pages still group + demote.
expect(results[1].score).toBe(9); // highest in group
expect(results[2].score).toBeCloseTo(8 * SESSION_DEMOTE, 5);
});
test('singleton session group → no demote', async () => {
const results = [
makeResult('media/chat/foo-session', 10, 1),
makeResult('media/chat/bar-session', 9, 2),
];
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => new Map(),
});
expect(results[0].score).toBe(10);
expect(results[1].score).toBe(9);
expect(results[0].graph_session_demoted).toBeUndefined();
expect(results[1].graph_session_demoted).toBeUndefined();
});
});
describe('applyGraphSignals — IRON RULE (T11): floor-gate respected', () => {
test('weak hub below floor threshold does NOT get boosted past strong above-floor non-hub', async () => {
// Top result at score 100; weak hub at score 30. Floor = 50.
// Without the gate: weak hub × 1.05 × 1.10 = 34.65 — still below
// strong, BUT the regression class is that future stacked boosts
// OR more aggressive magnitudes could shove a hub past a non-hub.
// The gate is the structural protection — weak hub gets ZERO graph
// boost regardless of magnitude.
const results = [
makeResult('strong/result', 100, 1), // above floor, no hub signal
makeResult('weak/hub', 30, 2), // BELOW floor, hub
];
const adjacency = new Map<number, AdjacencyRow>([
[2, { hits: 5, cross_source_hits: 5 }], // would be massive boost
]);
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
floorThreshold: 50, // weak hub at 30 is BELOW
adjacencyFn: async () => adjacency,
});
const weak = results[1];
// Score MUST be unchanged — no adjacency, no cross-source boost.
expect(weak.score).toBe(30);
expect(weak.graph_adjacency_hits).toBeUndefined();
expect(weak.graph_cross_source_hits).toBeUndefined();
expect(weak.graph_adjacency_boost).toBeUndefined();
// Strong result also untouched (no hub).
const strong = results[0];
expect(strong.score).toBe(100);
});
test('hub AT or ABOVE floor still gets boosted (gate is < not <=)', async () => {
const results = [makeResult('hub', 50, 1)]; // exactly at floor
const adjacency = new Map<number, AdjacencyRow>([
[1, { hits: 3, cross_source_hits: 0 }],
]);
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
floorThreshold: 50,
adjacencyFn: async () => adjacency,
});
expect(results[0].score).toBeCloseTo(50 * ADJACENCY_BOOST, 5);
expect(results[0].graph_adjacency_hits).toBe(3);
});
test('NaN score result skips the gate AND the boost (NaN < x is false in JS)', async () => {
const results = [makeResult('weird', NaN, 1)];
const adjacency = new Map<number, AdjacencyRow>([
[1, { hits: 3, cross_source_hits: 0 }],
]);
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
floorThreshold: 50,
adjacencyFn: async () => adjacency,
});
// NaN >= threshold is FALSE → gate kicks in → no boost.
expect(Number.isNaN(results[0].score)).toBe(true);
expect(results[0].graph_adjacency_hits).toBeUndefined();
});
});
describe('applyGraphSignals — fail-open', () => {
test('adjacencyFn throws → meta.errored=true, results unchanged', async () => {
const results = [makeResult('a/b', 10, 1), makeResult('c/d', 9, 2)];
const before = results.map(r => r.score);
let metaOut: any;
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => { throw new Error('boom'); },
onMeta: (m) => { metaOut = m; },
});
expect(results.map(r => r.score)).toEqual(before);
expect(metaOut.errored).toBe(true);
// Session diversification also skips (predictable all-or-nothing).
expect(metaOut.session_demotions).toBe(0);
});
test('adjacencyFn returns empty Map → no boosts, session diversification still runs', async () => {
// Use session-shaped slugs (date anchor) so sessionPrefix returns a
// real session id, not null. Post-v0.40.4 codex fix: bare `chat/a`
// + `chat/b` would no longer group because sessionPrefix returns
// 'chat/a' and 'chat/b' respectively (each treated as own session).
const results = [
makeResult('media/2026-05-20/chunk-a', 10, 1),
makeResult('media/2026-05-20/chunk-b', 9, 2),
];
let metaOut: any;
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => new Map(),
onMeta: (m) => { metaOut = m; },
});
expect(metaOut.errored).toBe(false);
expect(metaOut.adjacency_fires).toBe(0);
// Session demotion DOES fire — both share session 'media/2026-05-20'.
expect(metaOut.session_demotions).toBe(1);
});
});
describe('applyGraphSignals — score-distribution probe', () => {
test('always emitted when enabled, even with no fires', async () => {
const results = [makeResult('a/b', 10, 1), makeResult('a/c', 9, 2)];
let dist: any;
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => new Map(),
onScoreDistribution: (d) => { dist = d; },
});
expect(dist.top_k_size).toBe(2);
expect(dist.max).toBe(10);
expect(dist.min).toBe(9);
expect(dist.reorder_band_width).toBe(1);
});
test('not emitted when disabled', async () => {
let called = false;
await applyGraphSignals([makeResult('a/b', 10, 1)], ENGINE_STUB, {
enabled: false,
onScoreDistribution: () => { called = true; },
});
expect(called).toBe(false);
});
});
describe('applyGraphSignals — meta + timing', () => {
test('meta carries fire counts and duration_ms', async () => {
// Session-shaped slugs (date anchor) so sessionPrefix returns the
// same session for chunk-a + chunk-b. Pre-v0.40.4 codex fix: bare
// 'chat/a' had session 'chat' but post-fix that returns null.
const results = [
makeResult('media/2026-05-20/chunk-a', 10, 1),
makeResult('media/2026-05-20/chunk-b', 9, 2),
makeResult('hub', 8, 3),
];
const adjacency = new Map<number, AdjacencyRow>([
[3, { hits: 3, cross_source_hits: 2 }],
]);
let metaOut: any;
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async () => adjacency,
onMeta: (m) => { metaOut = m; },
});
expect(metaOut.adjacency_fires).toBe(1);
expect(metaOut.cross_source_fires).toBe(1);
expect(metaOut.session_demotions).toBe(1); // chat/a + chat/b → demote chat/b
expect(metaOut.duration_ms).toBeGreaterThanOrEqual(0);
expect(metaOut.top_k_size).toBe(3);
});
});
describe('applyGraphSignals — page_id invariant', () => {
test('result with missing page_id is silently skipped in dedup set', async () => {
const r = makeResult('weird', 10, 0); // invalid page_id
(r as any).page_id = undefined;
const results = [r, makeResult('normal', 9, 1)];
const pageIdsSeen: number[][] = [];
await applyGraphSignals(results, ENGINE_STUB, {
enabled: true,
adjacencyFn: async (ids) => { pageIdsSeen.push(ids); return new Map(); },
});
// The invalid page_id should NOT appear in the SQL input set.
expect(pageIdsSeen[0]).toEqual([1]);
});
});