Files
gbrain/test/e2e/search-swamp.test.ts
T
4fc1246606 v0.24.0: production-hardening pass on the skillify loop (#387)
* feat: v0.19.0 — skillify loop + AGENTS.md compat + brain-first convention

This is the v0.19.0 release. The branch ships four new CLI commands, a
refactor to check-resolvable, and an expansion of the brain-first
convention for sub-agent tool discovery. The original commit message
described only the convention expansion, undercounting the scope by ~5x;
this amend captures the full release.

NEW COMMANDS

- gbrain skillify scaffold <name>     — 4 stub files + idempotent resolver row
- gbrain skillify check [path]        — 10-item post-task audit (promoted)
- gbrain skillpack list / install     — curated 25-skill bundle, atomic install
- gbrain skillpack diff <name>        — per-file diff preview
- gbrain routing-eval                 — dedicated CI verb for Check 5 fixtures

CHECK-RESOLVABLE REFACTOR

- Accepts AGENTS.md as a resolver file alongside RESOLVER.md, at either
  the skills directory or one level up (workspace root layout).
- Auto-derives the skill manifest by walking skills/*/SKILL.md when
  manifest.json is missing.
- Splits ResolvableReport into errors[] + warnings[] so advisory checks
  (filing audit, routing gaps, DRY violations) don't break CI by default.
- New --strict opt-in flag promotes warnings to exit 1.

BRAIN-FIRST CONVENTION

- skills/conventions/brain-first.md expanded from 5-step lookup guide to
  full sub-agent reference: tool inventory, lookup chain, score thresholds,
  authority hierarchy, sync rules, entity page conventions, sub-agent
  propagation rule.

PRODUCTION-READINESS HARDENING (this branch's review pass)

- routing-eval --llm: emits stderr placeholder notice + runs structural
  layer only. README, CHANGELOG, CLI help all rewritten consistently.
  Was a silent no-op against documented contract.
- skillpack installer: receipt comment in fence (cumulative-slugs="...")
  preserves single-skill-install accumulation while letting install --all
  prune removed bundle skills cleanly. Unknown rows preserved + stderr
  warning for the operating agent. Pre-v0.19 fences upgrade silently.
- skillify scaffold: resolver-row regex broadened to detect backticked,
  quoted, and bare path forms. No duplicate row on --force after the
  user normalizes formatting.
- scripts/check-privacy.sh: now wired into package.json test chain so
  the wintermute-ban rule is actually enforced. New regression test.
- E2E Tier 2 (LLM skills) promoted from schedule-only to required per-PR
  CI. Local Tier 1 + Tier 2 verified clean.
- Stale v0.17/v0.18 version labels rewritten across new files.

TESTS

- test/routing-eval-cli.test.ts: 4 cases covering --llm warn semantics
- test/privacy-script-wired.test.ts: regression guard for CI wiring
- test/skillpack-install.test.ts: 4 new cases for receipt + cumulative
  + unknown-row preserve+warn + pre-v0.19 upgrade path
- test/skillify-scaffold.test.ts: 4 new cases for broadened regex

VERIFICATION

- bun test: 2237 pass / 18 known PGLite-contention flakes (CI green;
  documented as P3 dev-experience in TODOS.md)
- bun run typecheck: clean
- bun run test:e2e: 18/19 files green (1 pre-existing flake on master,
  not caused by this branch — verified via git stash)
- llms.txt + llms-full.txt regenerated to match README + CHANGELOG

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: scrub banned fork name from public artifacts

The privacy guard wired into the test chain in this branch caught 5
pre-existing references to the banned OpenClaw fork name in CHANGELOG.md
(2x), skills/migrations/v0.19.0.md (1x), src/cli.ts (1x), and
src/commands/sync.ts (1x). All originated in master's v0.19.0 release
notes and migration doc when the privacy script existed but wasn't
wired into CI yet.

Replacements per CLAUDE.md privacy mapping:
- Origin-story copy (CHANGELOG layer narratives, code comments naming
  the production deployment that drove the feature) → "Garry's OpenClaw"
- Reader-facing migration step → "your OpenClaw"

No code semantics changed. Comments + headings only.

Verification: scripts/check-privacy.sh exits 0, full CI guard chain
green (privacy + jsonb + progress + wasm + typecheck).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump VERSION to 0.24.0 + new CHANGELOG entry

Bump branch version above master's v0.21.0 per CLAUDE.md
"CHANGELOG + VERSION are branch-scoped" rule. The new v0.24.0 entry at
the top of CHANGELOG covers what THIS branch adds vs master:

- routing-eval --llm honesty pass (4-surface contract drift fix)
- skillpack installer cumulative-receipt + unknown-row preserve+warn
  (the Codex-caught regression that would have shipped in master if
  the original v0.19.0 had landed without this branch's review pass)
- skillify scaffold resolver-row regex broadening (backtick + quoted
  + bare forms; idempotency contract preserved under hand-editing)
- 5 banned-name leaks scrubbed from public artifacts
- check-privacy.sh wired into CI test chain + regression guard test
- 7 stale v0.17/v0.18 version labels rewritten across 5 files
- Tier 2 (LLM-skills E2E) promoted from schedule-only to required per-PR

VERSION 0.21.0 → 0.24.0
package.json version field synced.
llms.txt + llms-full.txt regenerated (no content drift; sizes match).

Test suite: 62/62 green across the 5 test files this branch added or
extended (routing-eval-cli, privacy-script-wired, skillpack-install,
skillify-scaffold, build-llms).

CI guards: privacy + jsonb + progress + wasm + typecheck all clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: update project documentation for v0.24.0

Auto-discovered drift via /document-release after the v0.24.0 hardening
pass landed. All factual corrections clearly warranted by the diff.

CLAUDE.md:
- Skillpack installer: documented the cumulative-slugs receipt comment,
  install --all prune semantics, unknown-row preserve+warn behavior,
  and pre-v0.24 silent upgrade. Was previously vague about
  "tracks a skill manifest so install --update diffs cleanly" without
  explaining what the receipt is or why it matters.
- routing-eval: replaced the false claim that --llm "opts into a Haiku
  tie-break layer for CI." Now correctly describes the placeholder
  semantic landed in v0.24.0 (stderr notice + structural-only run).

README.md:
- Skillpack section: added one paragraph on the receipt comment + the
  user-visible stderr message for hand-added rows. Connects the safe
  rerun promise to the v0.24.0 implementation that actually enforces it.

CONTRIBUTING.md:
- Running tests section: now recommends `bun run test` (full CI guard
  chain + typecheck + tests) before pushing. Names each guard so new
  contributors understand what catches what. The privacy guard (newly
  wired in v0.24.0) is one of these — without `bun run test` you'd skip
  it locally and find out from CI.

llms-full.txt: regenerated to reflect CLAUDE.md changes.

Verification: full guard chain green locally (privacy + jsonb + progress
+ wasm + typecheck).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Garry Tan <garry@ycombinator.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 23:17:54 -07:00

149 lines
5.7 KiB
TypeScript

/**
* Search Swamp Resistance E2E
*
* Reproduces the v3-plan repro case: a curated article (originals/) competes
* with two chat-log pages (openclaw/chat/) on similar ts_rank. With v0.21+
* source-aware ranking, the article must rank #0.
*
* Mirrors the structure of search-quality.test.ts. Uses PGLite in-memory.
*/
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
import { PGLiteEngine } from '../../src/core/pglite-engine.ts';
import type { ChunkInput } from '../../src/core/types.ts';
let engine: PGLiteEngine;
function basisEmbedding(idx: number, dim = 1536): Float32Array {
const emb = new Float32Array(dim);
emb[idx % dim] = 1.0;
return emb;
}
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
// Curated article — short, dense, opinionated. The page that should win.
await engine.putPage('originals/talks/article-outline-fat-code', {
type: 'writing',
title: 'Fat Code Thin Harness — Part 3',
compiled_truth:
'Fat code thin harness is the architectural pattern where business logic ' +
'lives in fat skill files and the runtime stays thin. Part 3 covers the ' +
'production case studies.',
timeline: '2026-04-10: Drafted Part 3 outline.',
});
await engine.upsertChunks('originals/talks/article-outline-fat-code', [
{
chunk_index: 0,
chunk_text:
'Fat code thin harness — the pattern where business logic lives in fat skill files. Part 3.',
chunk_source: 'compiled_truth',
embedding: basisEmbedding(7),
token_count: 20,
},
] satisfies ChunkInput[]);
// Chat swamp #1 — long page, mentions the phrase repeatedly.
await engine.putPage('openclaw/chat/2026-04-15', {
type: 'note',
title: '2026-04-15 chat',
compiled_truth: '',
timeline:
'fat code thin harness fat code thin harness — discussed at length. ' +
'fat code thin harness came up again. ' +
'The fat code thin harness pattern is something we keep returning to. ' +
'fat code thin harness fat code thin harness fat code thin harness.',
});
await engine.upsertChunks('openclaw/chat/2026-04-15', [
{
chunk_index: 0,
chunk_text:
'fat code thin harness fat code thin harness discussed at length, ' +
'the fat code thin harness pattern keeps coming back, ' +
'fat code thin harness fat code thin harness fat code thin harness.',
chunk_source: 'timeline',
embedding: basisEmbedding(8),
token_count: 30,
},
] satisfies ChunkInput[]);
// Chat swamp #2 — same shape.
await engine.putPage('openclaw/chat/2026-04-16', {
type: 'note',
title: '2026-04-16 chat',
compiled_truth: '',
timeline:
'fat code thin harness once more. fat code thin harness fat code thin harness. ' +
'still talking about fat code thin harness. fat code thin harness.',
});
await engine.upsertChunks('openclaw/chat/2026-04-16', [
{
chunk_index: 0,
chunk_text:
'fat code thin harness once more, fat code thin harness fat code thin harness, ' +
'still talking about fat code thin harness fat code thin harness.',
chunk_source: 'timeline',
embedding: basisEmbedding(9),
token_count: 25,
},
] satisfies ChunkInput[]);
}, 60_000);
afterAll(async () => {
await engine.disconnect();
});
describe('searchKeyword swamp resistance', () => {
test('curated originals/ page outranks chat swamp on multi-word query', async () => {
const results = await engine.searchKeyword('fat code thin harness');
expect(results.length).toBeGreaterThan(0);
const top = results[0];
expect(top.slug).toBe('originals/talks/article-outline-fat-code');
});
test('detail=high (temporal bypass) lets chat swamp re-surface', async () => {
// With source-boost disabled, raw ts_rank wins → chat pages, which have
// many more keyword hits, are allowed back to the top. This guards the
// temporal-query workflow ("what did we discuss about X").
const results = await engine.searchKeyword('fat code thin harness', { detail: 'high' });
expect(results.length).toBeGreaterThan(0);
// Top result should be a chat page (more keyword density per chunk).
const topSlugs = results.slice(0, 2).map(r => r.slug);
const anyChat = topSlugs.some(s => s.startsWith('openclaw/chat/'));
expect(anyChat).toBe(true);
});
});
describe('searchVector swamp resistance', () => {
test('curated originals/ page outranks chat swamp when boost is meaningful', async () => {
// Query vector is close to all three pages (mixed direction). Without
// source-boost the chat pages would tie or win on raw cosine; with
// source-boost the originals/ page dominates.
const queryVec = new Float32Array(1536);
queryVec[7] = 0.6; // article direction
queryVec[8] = 0.55; // chat-1 direction (slightly higher, simulating swamp)
queryVec[9] = 0.55; // chat-2 direction
// Normalize so cosine math is well-formed.
const norm = Math.sqrt(0.6 * 0.6 + 0.55 * 0.55 + 0.55 * 0.55);
for (let i = 0; i < queryVec.length; i++) queryVec[i] = queryVec[i] / norm;
const results = await engine.searchVector(queryVec);
expect(results.length).toBeGreaterThan(0);
expect(results[0].slug).toBe('originals/talks/article-outline-fat-code');
});
test('two-stage CTE returns p.source_id (regression for v0.18 multi-source)', async () => {
const queryVec = basisEmbedding(7);
const results = await engine.searchVector(queryVec);
expect(results.length).toBeGreaterThan(0);
// source_id is added by v0.18 multi-source brains; carrying it through
// the inner→outer CTE is one of the v3 plan's pass-4 findings.
for (const r of results) {
expect(r.source_id).toBeDefined();
}
});
});