mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 22:15:33 +00:00
* fix(subagent): bind Anthropic SDK messages.create() correctly The makeSubagentHandler was casting `new Anthropic()` directly to MessagesClient, but MessagesClient.create() maps to sdk.messages.create(), not sdk.create(). Every subagent job immediately died with: client.create is not a function Fix: wrap the SDK instance so .create() delegates to .messages.create() with proper `this` binding via .bind(sdk.messages). Discovered on first production run of gbrain agent against Supabase. Co-Authored-By: Wintermute <wintermute@openclaw.ai> * chore(ci): add typescript typecheck to test pipeline + clean up baseline errors Root cause infra gap that let the v0.16.0 subagent bug ship: CI ran only `bun test`, which transpiles types without checking them. Type errors only surfaced at runtime, in production. Changes: - Add `typescript` devDep and a `typecheck` npm script (`tsc --noEmit`). - Chain `bun run typecheck` into `bun run test` so developers get the same pipeline locally that CI runs. - Flip `.github/workflows/test.yml` to invoke `bun run test` (the npm script, including typecheck) instead of `bun test` (runner only). - Clean up 100+ pre-existing type errors across 30+ files so the first run of `tsc --noEmit` is green. Root causes were: - `databaseUrl` → `database_url` rename drift in test fixtures (9 files) - `PageType` union missing `'meeting'` / `'note'` entries that are already used in both src and tests (link-extraction.ts comments acknowledged the gap) - `GBrainConfig.storage` field never declared despite being read in files.ts and operations.ts - `ErrorCode` union missing `'permission_denied'` - `OrchestratorOpts` shape changed; test callers not updated - Dead-code comparisons in migration orchestrators against narrowed status types - postgres.js `Row`-callback type drift on several `.map()` calls - Buffer-as-BodyInit assignment in supabase.ts (real but non-fatal runtime bug; Uint8Array slice works and is type-correct) - Various `as X` single-step casts that now need `as unknown as X` per TS's stricter structural-conversion rules - Bump `beforeAll` hook timeout to 30s on four PGLite-heavy tests that were flaky under parallel test execution: wait-for-completion, extract-fs, e2e/search-quality, e2e/graph-quality. All pass in isolation; timeouts only happened when dozens of PGLite instances init'd simultaneously. The new CI pipeline now fails on any type error across src/ or test/, giving us the compile-time regression guard the subagent fix depends on. * fix(subagent): bind Anthropic SDK messages.create() correctly Shipped bug: v0.16.0 cast `new Anthropic()` to `MessagesClient`, but `.create()` lives at `sdk.messages.create`, not on the top-level client. Every subagent job in production died on first LLM call with `client.create is not a function`. Discovered on the first `gbrain agent run` against Supabase. Fix: assign `sdk.messages` directly to the `MessagesClient` slot. `sdk.messages` IS the object with a callable `.create()`; the original bug was picking the wrong entry point on the SDK. No helper, no wrapper, no `.bind()` — JS method-call semantics preserve `this` at the call site because `subagent.ts:336` invokes `client.create(...)` with `client === sdk.messages`. The one-line assignment also typechecks cleanly against the existing `MessagesClient` interface (SDK's first `create` overload: `(MessageCreateParamsNonStreaming, Core.RequestOptions?) => APIPromise<Message>` is assignable structurally). This gives us compile-time regression protection: anyone reverting to `new Anthropic()` would fail tsc because `Anthropic` has no top-level `.create`. (The companion chore commit puts `tsc --noEmit` in CI so this guard is enforced.) Also adds a `makeAnthropic?: () => Anthropic` dep-injection seam so the factory default construction branch is testable without real API calls. Regression test drives one handler turn through a fake SDK, asserting `sdk.messages.create` is actually called. If someone later reverts to `new Anthropic()`, both guards fire: tsc fails AND the test fails. Co-Authored-By: Wintermute <wintermute@garrytan.com> * chore(tests): add bunfig.toml + 60s hook timeouts to stabilize PGLite-heavy suites After turning on tsc in CI (previous commit), running the full `bun run test` suite in one shot triggered flaky `beforeEach/afterEach hook timed out` failures on 8+ test files. Every failure traced to PGLite WASM init contention when many test files spin up fresh PGLite instances in parallel; each one alone passes in isolation. - `bunfig.toml` sets the global test hook timeout to 60s (default is 5s), covering every test file without per-file edits. - Individual `beforeAll(fn, 60_000)` / `beforeEach(fn, 15_000)` calls on the 8 tests that flaked most stay in place as explicit safety nets so a future bunfig config change doesn't silently re-introduce the flake. Result: 1997 pass, 0 fail on `bun run test` (117 tests added since the prior baseline by picking up typecheck-gated passes). No infrastructure flake tolerated in CI. * chore: bump version and changelog (v0.16.3) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Wintermute <wintermute@garrytan.com> Co-authored-by: Wintermute <wintermute@openclaw.ai> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
218 lines
7.8 KiB
TypeScript
218 lines
7.8 KiB
TypeScript
/**
|
|
* Search Quality E2E Tests
|
|
*
|
|
* Tests the full search pipeline against PGLite with seeded pages and
|
|
* structured mock embeddings (basis vectors). No OpenAI API calls needed.
|
|
*
|
|
* Validates: compiled truth boost, detail parameter, source-aware dedup,
|
|
* chunk_id/chunk_index in results, and getEmbeddingsByChunkIds.
|
|
*/
|
|
|
|
import { describe, test, expect, beforeAll, afterAll } from 'bun:test';
|
|
import { PGLiteEngine } from '../../src/core/pglite-engine.ts';
|
|
import type { ChunkInput, SearchResult } from '../../src/core/types.ts';
|
|
|
|
let engine: PGLiteEngine;
|
|
|
|
// Create a basis vector embedding: dimension `idx` is 1.0, rest are 0.0
|
|
function basisEmbedding(idx: number, dim = 1536): Float32Array {
|
|
const emb = new Float32Array(dim);
|
|
emb[idx % dim] = 1.0;
|
|
return emb;
|
|
}
|
|
|
|
beforeAll(async () => {
|
|
engine = new PGLiteEngine();
|
|
await engine.connect({}); // in-memory
|
|
await engine.initSchema();
|
|
|
|
// Seed test pages with compiled_truth + timeline chunks
|
|
await engine.putPage('people/pedro', {
|
|
type: 'person',
|
|
title: 'Pedro Franceschi',
|
|
compiled_truth: 'Pedro is the co-founder of Brex. Expert in fintech and payments infrastructure.',
|
|
timeline: '2024-03-15: Met Pedro at YC dinner. Discussed AI security.',
|
|
});
|
|
|
|
// Seed chunks with structured embeddings
|
|
const pedroChunks: ChunkInput[] = [
|
|
{
|
|
chunk_index: 0,
|
|
chunk_text: 'Pedro is the co-founder of Brex. Expert in fintech and payments infrastructure.',
|
|
chunk_source: 'compiled_truth',
|
|
embedding: basisEmbedding(0), // direction 0 = fintech/compiled truth
|
|
token_count: 15,
|
|
},
|
|
{
|
|
chunk_index: 1,
|
|
chunk_text: '2024-03-15: Met Pedro at YC dinner. Discussed AI security and Crab Trap.',
|
|
chunk_source: 'timeline',
|
|
embedding: basisEmbedding(1), // direction 1 = meeting/timeline
|
|
token_count: 18,
|
|
},
|
|
];
|
|
await engine.upsertChunks('people/pedro', pedroChunks);
|
|
|
|
await engine.putPage('companies/variant', {
|
|
type: 'company',
|
|
title: 'Variant Fund',
|
|
compiled_truth: 'Variant is a crypto-native investment firm focused on web3 ownership economy.',
|
|
timeline: '2024-06-01: Variant announced new fund.',
|
|
});
|
|
|
|
const variantChunks: ChunkInput[] = [
|
|
{
|
|
chunk_index: 0,
|
|
chunk_text: 'Variant is a crypto-native investment firm focused on web3 ownership economy.',
|
|
chunk_source: 'compiled_truth',
|
|
embedding: basisEmbedding(2),
|
|
token_count: 14,
|
|
},
|
|
{
|
|
chunk_index: 1,
|
|
chunk_text: '2024-06-01: Variant announced new fund. $450M raised.',
|
|
chunk_source: 'timeline',
|
|
embedding: basisEmbedding(3),
|
|
token_count: 12,
|
|
},
|
|
];
|
|
await engine.upsertChunks('companies/variant', variantChunks);
|
|
|
|
await engine.putPage('concepts/ai-philosophy', {
|
|
type: 'concept',
|
|
title: 'AI Changes Who Gets to Build',
|
|
compiled_truth: 'AI democratizes building. The marginal cost of creation approaches zero.',
|
|
timeline: '2024-01-10: First wrote about AI and building access.',
|
|
});
|
|
|
|
const aiChunks: ChunkInput[] = [
|
|
{
|
|
chunk_index: 0,
|
|
chunk_text: 'AI democratizes building. The marginal cost of creation approaches zero. This changes who gets to build.',
|
|
chunk_source: 'compiled_truth',
|
|
embedding: basisEmbedding(4),
|
|
token_count: 20,
|
|
},
|
|
{
|
|
chunk_index: 1,
|
|
chunk_text: '2024-01-10: First wrote about AI and building access. Shared on X.',
|
|
chunk_source: 'timeline',
|
|
embedding: basisEmbedding(5),
|
|
token_count: 15,
|
|
},
|
|
];
|
|
await engine.upsertChunks('concepts/ai-philosophy', aiChunks);
|
|
}, 60_000);
|
|
|
|
afterAll(async () => {
|
|
await engine.disconnect();
|
|
});
|
|
|
|
describe('SearchResult fields', () => {
|
|
test('keyword search returns chunk_id and chunk_index', async () => {
|
|
const results = await engine.searchKeyword('Pedro');
|
|
expect(results.length).toBeGreaterThan(0);
|
|
const r = results[0];
|
|
expect(r.chunk_id).toBeDefined();
|
|
expect(typeof r.chunk_id).toBe('number');
|
|
expect(r.chunk_index).toBeDefined();
|
|
expect(typeof r.chunk_index).toBe('number');
|
|
});
|
|
|
|
test('vector search returns chunk_id and chunk_index', async () => {
|
|
const results = await engine.searchVector(basisEmbedding(0));
|
|
expect(results.length).toBeGreaterThan(0);
|
|
const r = results[0];
|
|
expect(r.chunk_id).toBeDefined();
|
|
expect(typeof r.chunk_id).toBe('number');
|
|
expect(r.chunk_index).toBeDefined();
|
|
expect(typeof r.chunk_index).toBe('number');
|
|
});
|
|
});
|
|
|
|
describe('detail parameter', () => {
|
|
test('detail=low returns only compiled_truth chunks', async () => {
|
|
const results = await engine.searchKeyword('Pedro', { detail: 'low' });
|
|
for (const r of results) {
|
|
expect(r.chunk_source).toBe('compiled_truth');
|
|
}
|
|
});
|
|
|
|
test('detail=high returns all chunk sources', async () => {
|
|
const results = await engine.searchKeyword('Pedro', { detail: 'high' });
|
|
// Should include at least compiled_truth (might include timeline depending on tsvector match)
|
|
expect(results.length).toBeGreaterThan(0);
|
|
});
|
|
|
|
test('detail=low on vector search filters to compiled_truth', async () => {
|
|
// Use a timeline-direction embedding — with detail=low, should get no results
|
|
// or only compiled_truth results
|
|
const results = await engine.searchVector(basisEmbedding(1), { detail: 'low' });
|
|
for (const r of results) {
|
|
expect(r.chunk_source).toBe('compiled_truth');
|
|
}
|
|
});
|
|
|
|
test('default detail (medium) returns all sources', async () => {
|
|
const results = await engine.searchKeyword('Pedro');
|
|
// No filter applied, should return whatever matches
|
|
expect(results.length).toBeGreaterThan(0);
|
|
});
|
|
});
|
|
|
|
describe('getEmbeddingsByChunkIds', () => {
|
|
test('returns embeddings for valid chunk IDs', async () => {
|
|
const searchResults = await engine.searchVector(basisEmbedding(0));
|
|
expect(searchResults.length).toBeGreaterThan(0);
|
|
|
|
const ids = searchResults.map(r => r.chunk_id).filter((id): id is number => id != null);
|
|
const embMap = await engine.getEmbeddingsByChunkIds(ids);
|
|
|
|
expect(embMap.size).toBeGreaterThan(0);
|
|
for (const [id, emb] of embMap) {
|
|
expect(emb).toBeInstanceOf(Float32Array);
|
|
expect(emb.length).toBe(1536);
|
|
}
|
|
});
|
|
|
|
test('returns empty map for empty ID list', async () => {
|
|
const embMap = await engine.getEmbeddingsByChunkIds([]);
|
|
expect(embMap.size).toBe(0);
|
|
});
|
|
|
|
test('returns empty map for non-existent IDs', async () => {
|
|
const embMap = await engine.getEmbeddingsByChunkIds([999999, 999998]);
|
|
expect(embMap.size).toBe(0);
|
|
});
|
|
});
|
|
|
|
describe('keyword search without DISTINCT ON', () => {
|
|
test('returns multiple chunks per page', async () => {
|
|
// Search for something that matches a page with multiple chunks
|
|
const results = await engine.searchKeyword('Pedro', { limit: 10 });
|
|
const pedroChunks = results.filter(r => r.slug === 'people/pedro');
|
|
// Should be able to return more than 1 chunk per page
|
|
// (depends on tsvector matching — Pedro is in page title/search_vector)
|
|
expect(results.length).toBeGreaterThan(0);
|
|
});
|
|
});
|
|
|
|
describe('compiled truth boost (vector search validates ordering)', () => {
|
|
test('compiled_truth chunks rank first with basis vector queries', async () => {
|
|
// Query with the compiled_truth direction for Pedro (basis 0)
|
|
const results = await engine.searchVector(basisEmbedding(0), { limit: 5 });
|
|
expect(results.length).toBeGreaterThan(0);
|
|
// The closest result should be the compiled_truth chunk (basis 0)
|
|
expect(results[0].chunk_source).toBe('compiled_truth');
|
|
expect(results[0].slug).toBe('people/pedro');
|
|
});
|
|
|
|
test('timeline chunks rank first when queried with timeline direction', async () => {
|
|
// Query with the timeline direction for Pedro (basis 1)
|
|
const results = await engine.searchVector(basisEmbedding(1), { limit: 5 });
|
|
expect(results.length).toBeGreaterThan(0);
|
|
expect(results[0].chunk_source).toBe('timeline');
|
|
expect(results[0].slug).toBe('people/pedro');
|
|
});
|
|
});
|