mirror of
https://github.com/garrytan/gbrain.git
synced 2026-07-27 21:19:18 +00:00
v0.28.10 fix: lightweight /health endpoint — SELECT 1 instead of getStats() (#701)
* fix: lightweight /health endpoint — SELECT 1 instead of getStats() On large brains (96K+ pages), getStats() runs 6× count(*) queries that routinely exceed the 3s HEALTH_TIMEOUT_MS through PgBouncer. This produces false 503s that cause external health monitors (cron, Fly.io, k8s) to restart otherwise-healthy servers — which in turn creates advisory lock pile-ups when multiple serve instances compete for the migration lock. Changes: - /health now runs `SELECT 1` for liveness (sub-millisecond) - ?full=true opt-in preserves the old getStats() behavior - /admin/api/health-indicators still returns full stats - probeHealth() retained for callers that need it * refactor(health): extract probeLiveness, move full stats to /admin/api/full-stats Addresses outside-voice review of PR #701. The original ?full=true query-param escape hatch was withdrawn because the loopback IP gate's correctness depended on app.set('trust proxy', 'loopback') semantics holding under proxy/XFF misconfiguration, and the PR's own comment misidentified /admin/api/health-indicators as a full-stats endpoint when it actually returns only {expiring_soon, error_rate}. Changes: - src/commands/serve-http.ts: new probeLiveness(sql, engineName, version, timeoutMs) helper next to probeHealth. Same shape, same return type, same finally-block clearTimeout discipline. /health is now a 2-line dispatch through probeLiveness. Removes ?full=true entirely. Adds new admin route /admin/api/full-stats behind the existing requireAdmin middleware that returns probeHealth(engine, ...) — same body shape /health used to expose (status, version, engine, page_count, chunk_count, embedded_count, link_count, tag_count, timeline_entry_count). - test/serve-http-health.test.ts: 4 new probeLiveness cases (success-shape regression with exact-keys assertion, timeout, db-error, timer-cleanup under 100 concurrent probes). - test/e2e/serve-http-oauth.test.ts: existing /health body-shape assertion rewritten to the liveness-only contract (page_count must NOT be present); 2 new admin-stats cases (401 without cookie, 200 with magic-link-derived admin cookie returns getStats() body). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * chore: bump version and changelog (v0.28.10) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs: update CLAUDE.md serve-http.ts annotation for v0.28.10 split Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(claude): explicit "run E2E without asking" + schema-bootstrap step The previous wording ("Always run E2E tests when they exist") was easy to read as a soft preference; in practice agents kept proposing the run instead of just doing it. Make the policy unmistakable: if there's a relevant E2E and you want to verify behavior, just spin up the DB and run. Also documents the schema-bootstrap step that bit a fresh container today — `oauth_clients` doesn't exist on a virgin pgvector image until `gbrain doctor` (or any engine-connecting command) triggers `initSchema()`. `apply-migrations` alone runs ALTER-style migrations on top of an already-bootstrapped schema; it does not seed base tables. Tests that bypass the engine via execSync against `gbrain auth register-client` hit the DB directly and need bootstrap first. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(serve-http): persist mcp_request_log on every JSON-RPC method + admin-scope F7 tests Closes the 4 pre-existing E2E failures in test/e2e/serve-http-oauth.test.ts that surfaced when DATABASE_URL was set on the v0.28.10 branch. The branch isn't the cause — these were broken on master too (verified by checking out origin/master's serve-http.ts + test file: 0/4 pass). Owning them here as a bisectable commit. Two root causes, both in serve-http.ts's /mcp logging + scope discipline. 1. mcp_request_log was only INSERTed inside the tools/call success/error paths. tools/list, the unknown-op early-return, and the insufficient-scope early-return all returned without logging. The v0.26.3 persistence regression test calls tools/list + tools/call non-existent and expects >= 2 rows; on the prior implementation it got 0. The agent_name resolution test (single tools/list, expects the row) had the same shape. Fix: log every JSON-RPC method exit point. tools/list logs operation = 'tools/list' with status='success' (lists never fail). Unknown-op logs operation = the attempted name with error_message starting 'unknown_operation:'. Insufficient-scope logs operation = the attempted name with error_message 'insufficient_scope: requires <scope>'. Admin agents auditing /admin/api/requests now see the full attempt log, not just successful valid-op calls. 2. The F7 RCE-regression tests minted 'read write' tokens to assert submit_job for protected names ('shell', 'subagent') gets rejected. But submit_job's required scope is 'admin' (set by hasScope-aware v0.28 enforcement), so a 'read write' token gets rejected with insufficient_scope BEFORE reaching the F7 protected-name guard at operations.ts:1527. The test's assertion checked for 'permission_denied' / 'cannot be submitted over MCP' — neither appears in an insufficient_scope response — so 'rejected' computed to false even though the call was actually rejected. Worse, if someone removed the F7 guard, the test would still pass because scope check would catch it: regression-test integrity failure. Fix: register the e2e-oauth-test client with admin in its allowed scopes (was 'read write', now 'read write admin'), and have F7 tests mint admin-scoped tokens explicitly. Adding admin to the client's allowed ceiling does not auto-grant it to subset-mint calls — other tests minting 'read' / 'read write' still get the subset they ask for. The persistence test's assertion 'rows.find(r => r.operation === "tools/call")' was also updated to match the actual logging convention (operation = inner tool name on call paths, JSON-RPC method on list/scope/unknown paths). E2E result: 29/29 pass on a fresh pgvector container (fixed 4, kept the 25 that were passing). Unit suite: 4191 pass, 0 fail, unchanged. Typecheck: clean. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * chore: regenerate llms-full.txt after CLAUDE.md update Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: garrytan-agents <garrytan-agents@users.noreply.github.com> Co-authored-by: Garry Tan <garrytan@gmail.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
co-authored by
garrytan-agents
Garry Tan
Claude Opus 4.7
parent
0e7d13e740
commit
f7c129407a
+117
@@ -2,6 +2,123 @@
|
||||
|
||||
All notable changes to GBrain will be documented in this file.
|
||||
|
||||
## [0.28.10] - 2026-05-07
|
||||
|
||||
**`/health` stops being slow. Stops triggering restart cascades.**
|
||||
**Liveness is liveness; stats moved where they belong.**
|
||||
|
||||
v0.28.10 makes `/health` a one-query liveness probe and moves the heavy
|
||||
engine-stats payload behind admin auth. On any brain large enough to
|
||||
matter (think 96K+ pages through PgBouncer), the old `/health` ran six
|
||||
`count(*)` queries that routinely exceeded the 3-second probe window.
|
||||
Fly.io / k8s / cron orchestrators saw 503, restarted the otherwise-healthy
|
||||
server, and each crashed startup left a Postgres advisory-lock waiter on
|
||||
the migration lock. Production observed six stacked waiters before the
|
||||
server became permanently stuck and the locks had to be cleared by hand.
|
||||
New installs: nothing to do. Existing operators whose monitors scraped
|
||||
`page_count` from `/health` migrate to `/admin/api/full-stats`
|
||||
(admin-cookie auth). `/admin/api/health-indicators` is unchanged.
|
||||
|
||||
### The numbers that matter
|
||||
|
||||
Measured against a 96K-page production brain through PgBouncer (Supabase,
|
||||
port 6543) on the branch this release ships from.
|
||||
|
||||
| Probe path | Before | After |
|
||||
|---|---|---|
|
||||
| `/health` latency on 96K-page brain | timed out at 3000ms | <10ms |
|
||||
| `count(*)` queries per `/health` request | 6 | 0 |
|
||||
| Restart-cascade reachability via `/health` | yes (orchestrator → 503 → restart loop → lock pile-up) | no |
|
||||
| Full-stats endpoint | `/health` (unauthenticated) | `/admin/api/full-stats` (admin cookie) |
|
||||
| Tests covering the route surface | 3 unit (probeHealth only) | 7 unit + 3 new E2E (`probeLiveness` + body-shape regression + admin-auth happy/401) |
|
||||
|
||||
**What this means for you**: orchestrator restart loops stop. The server
|
||||
stays up under load. Dashboards that actually need `page_count` log into
|
||||
`/admin/` and call a real authenticated endpoint instead of pulling
|
||||
admin-grade stats off a public route. The original PR's `?full=true`
|
||||
query-param escape hatch was withdrawn after outside-voice review flagged
|
||||
it as a back-door to the same DoS surface; the shipped design routes
|
||||
through the existing `requireAdmin` middleware that's already protecting
|
||||
seven other admin endpoints.
|
||||
|
||||
## To take advantage of v0.28.10
|
||||
|
||||
`gbrain upgrade` should do this automatically. There is no schema migration
|
||||
in this release.
|
||||
|
||||
1. **Verify the upgrade landed**:
|
||||
```bash
|
||||
curl -s http://localhost:3131/health | jq .
|
||||
# expect: {"status":"ok","version":"0.28.10","engine":"postgres"}
|
||||
curl -s http://localhost:3131/admin/api/full-stats
|
||||
# expect: {"error":"Admin authentication required"} with HTTP 401
|
||||
```
|
||||
2. **If your monitoring stack scrapes `page_count` / `chunk_count` /
|
||||
`embedded_count` / `link_count` / `tag_count` / `timeline_entry_count`
|
||||
from `GET /health`**, those fields are gone. Move the scraper to
|
||||
`GET /admin/api/full-stats` with the `gbrain_admin` session cookie.
|
||||
`/admin/api/health-indicators` continues to return only
|
||||
`{expiring_soon, error_rate}` and is unchanged.
|
||||
3. **If `gbrain doctor` reports anything unexpected**, please file an issue
|
||||
at https://github.com/garrytan/gbrain/issues with the doctor output and
|
||||
what monitoring stack you're using.
|
||||
|
||||
### Itemized changes
|
||||
|
||||
#### What's new
|
||||
|
||||
- **`GET /admin/api/full-stats`** (new route, behind `requireAdmin`).
|
||||
Returns the same body shape `/health` used to expose:
|
||||
`{status, version, engine, page_count, chunk_count, embedded_count,
|
||||
link_count, tag_count, timeline_entry_count}`. Reuses the existing
|
||||
`probeHealth()` helper, including its 3-second timeout race against
|
||||
`engine.getStats()`. Sibling to `/admin/api/stats` and
|
||||
`/admin/api/health-indicators`.
|
||||
- **`probeLiveness(sql, engineName, version, timeoutMs)`** (new exported
|
||||
helper in `src/commands/serve-http.ts`). Mirror of `probeHealth()`'s
|
||||
shape: races `sql\`SELECT 1\`` against `HEALTH_TIMEOUT_MS`, returns the
|
||||
same `ProbeHealthResult` tagged-union, single finally-block
|
||||
`clearTimeout` discipline.
|
||||
|
||||
#### Behavior changes
|
||||
|
||||
- **`GET /health` is now liveness-only.** Body shape:
|
||||
`{status, version, engine}`. `getStats()` is no longer called on this
|
||||
route. Status code semantics are unchanged: 200 on probe success, 503
|
||||
on timeout or DB error.
|
||||
- **`?full=true` query param removed.** Operators who relied on
|
||||
`curl localhost:3131/health?full=true` for full stats migrate to
|
||||
admin-auth + `/admin/api/full-stats`.
|
||||
|
||||
#### Tests
|
||||
|
||||
- `test/serve-http-health.test.ts` — 4 new `probeLiveness` cases:
|
||||
happy-path body-shape regression (asserts the body has exactly
|
||||
`{status, version, engine}` and no `page_count` / `chunk_count`),
|
||||
timeout path, db-error path, timer-cleanup under 100 concurrent probes.
|
||||
- `test/e2e/serve-http-oauth.test.ts` — 3 new cases: `/health` returns
|
||||
liveness-only body (no engine stats), `/admin/api/full-stats` without
|
||||
cookie returns 401, `/admin/api/full-stats` with a magic-link-derived
|
||||
admin cookie returns the full `getStats()` body.
|
||||
- The previous `'health endpoint returns OK without auth'` E2E case that
|
||||
asserted `data.page_count` was a number is updated to the
|
||||
liveness-only shape — that assertion would have silently passed against
|
||||
the original `?full=true` PR while letting the heavy probe leak into
|
||||
the public route.
|
||||
|
||||
#### For contributors
|
||||
|
||||
- The original PR (#701, by @garrytan-agents) added `?full=true` as a
|
||||
debug escape hatch and the plan-eng review proposed gating it to
|
||||
loopback IPs. Outside-voice review (Codex) flagged that the loopback
|
||||
gate's correctness depended on `app.set('trust proxy', 'loopback')`
|
||||
semantics holding under proxy/XFF misconfiguration, and that the PR's
|
||||
own comment misidentified `/admin/api/health-indicators` as a
|
||||
full-stats endpoint when it actually returns only
|
||||
`{expiring_soon, error_rate}`. The shipped design uses the existing
|
||||
`requireAdmin` middleware instead, eliminating the proxy dependency
|
||||
and fixing the migration story in one move.
|
||||
|
||||
## [0.28.9] - 2026-05-07
|
||||
|
||||
**Multimodal ingestion lands on master.**
|
||||
|
||||
@@ -139,7 +139,7 @@ strict behavior when unset.
|
||||
- `src/mcp/server.ts` — MCP stdio server (generated from operations). v0.22.7: tool-call handler delegates to `dispatchToolCall` from `src/mcp/dispatch.ts` so stdio + HTTP transports share one validation, context-build, and error-format path.
|
||||
- `src/mcp/dispatch.ts` (v0.22.7) — Shared tool-call dispatch consumed by both stdio (`server.ts`) and HTTP transports. Exports `dispatchToolCall(engine, name, params, opts)`, `buildOperationContext(engine, params, opts)`, and `validateParams(op, params)`. Single source of truth for `(ctx, params)` handler arg order and the 5-field `OperationContext` shape (engine + config + logger + dryRun + remote). Defaults to `remote: true` (untrusted); local CLI callers pass `remote: false`. Closed F1/F2/F3 drift bugs in the original v0.22.5 HTTP transport. **v0.26.9 (F8):** adds `summarizeMcpParams(opName, params)` — privacy-preserving redactor for `mcp_request_log` and the admin SSE feed. Returns `{redacted, kind, declared_keys, unknown_key_count, approx_bytes}`. Intersects submitted top-level keys against the operation's declared `params` allow-list (declared keys preserved as a sorted array for debug visibility; unknown keys counted but never named, closing the attacker-controlled-key-name leak). Byte counts bucketed up to nearest 1KB so an attacker can't binary-search secret-content sizes via repeated probes. Operators on a personal laptop who want raw payload visibility opt back in with `gbrain serve --http --log-full-params` (loud stderr warning at startup). Canonical helper — new logging code paths route through it rather than `JSON.stringify(params)`.
|
||||
- `src/mcp/rate-limit.ts` (v0.22.7) — Bounded-LRU token-bucket limiter. `buildDefaultLimiters()` returns the two-bucket pipeline: pre-auth IP (30/60s, fires BEFORE the DB lookup so brute-force load against `access_tokens` is actually capped) + post-auth token-id (60/60s). Tracks `lastTouchedMs` separately from `lastRefillMs` so an exhausted key can't be reset by hammering past the TTL. LRU cap bounds memory under attacker-controlled key growth.
|
||||
- `src/commands/serve-http.ts` (v0.26.0) — Express 5 HTTP MCP server with OAuth 2.1, admin dashboard, and SSE live activity feed. Started via `gbrain serve --http [--port N] [--token-ttl N] [--enable-dcr] [--public-url URL] [--log-full-params]`. Supersedes the v0.22.7 `src/mcp/http-transport.ts` simple bearer-auth path. Combines MCP SDK's `mcpAuthRouter` (authorize / token / register / revoke endpoints), a custom `client_credentials` handler (SDK's token endpoint throws `UnsupportedGrantTypeError` for CC; the custom handler runs BEFORE the router and falls through for `auth_code` / `refresh_token`), `requireBearerAuth` middleware for `/mcp` with scope enforcement before op dispatch, `localOnly` rejection, and `express-rate-limit` at 50 req / 15 min on `/token`. Serves the built admin SPA from `admin/dist/` with SPA fallback. `/admin/events` SSE endpoint broadcasts every MCP request to connected admin browsers. `cookie-parser` middleware wired (Express 5 has no built-in). Startup logging prints port, engine, configured issuer URL (honors `--public-url`), registered-client count, DCR status, and admin bootstrap token. **v0.26.9 hardening pass:** F7 sets `remote: true` explicitly on the `/mcp` request handler's OperationContext literal (closes the HTTP shell-job RCE — without this, `submit_job`'s protected-name guard at `operations.ts:1391` saw a falsy undefined and skipped, letting a `read+write`-scoped OAuth token submit `shell` jobs). F8 wires `summarizeMcpParams` from `src/mcp/dispatch.ts` into both `mcp_request_log` writes and the admin SSE feed by default (raw payloads opt-in via `--log-full-params` with stderr warning). F9 sets cookie `Secure` flag when behind HTTPS or a public-URL proxy. F10 caps the magic-link nonce store with an LRU bound. F12 routes DCR disable through the `GBrainOAuthProvider` constructor's `dcrDisabled` option instead of the prior monkey-patch on the express router. F14 wraps `transport.handleRequest` in try/catch so SDK throws return a JSON-RPC 500 envelope instead of express's default HTML error page. F15 unifies OperationError + unexpected exceptions through `buildError` / `serializeError` so `/mcp` always returns the same envelope shape. **v0.28.1:** `/health` endpoint extracted into pure `probeHealth(engine)` async function with `HEALTH_TIMEOUT_MS = 3000` exported constant — drops the timeout from 5s to 3s so Fly.io's 5s health-check deadline gets 2s of headroom for TCP, response framing, and clock skew. Races `engine.getStats()` against the timeout via `Promise.race`; saturated pool returns 503 with `Health check timed out (database pool may be saturated)` instead of hanging. `clearTimeout` in finally block prevents pending-timer pile-up under high probe rates (race-leak fix from adversarial review).
|
||||
- `src/commands/serve-http.ts` (v0.26.0) — Express 5 HTTP MCP server with OAuth 2.1, admin dashboard, and SSE live activity feed. Started via `gbrain serve --http [--port N] [--token-ttl N] [--enable-dcr] [--public-url URL] [--log-full-params]`. Supersedes the v0.22.7 `src/mcp/http-transport.ts` simple bearer-auth path. Combines MCP SDK's `mcpAuthRouter` (authorize / token / register / revoke endpoints), a custom `client_credentials` handler (SDK's token endpoint throws `UnsupportedGrantTypeError` for CC; the custom handler runs BEFORE the router and falls through for `auth_code` / `refresh_token`), `requireBearerAuth` middleware for `/mcp` with scope enforcement before op dispatch, `localOnly` rejection, and `express-rate-limit` at 50 req / 15 min on `/token`. Serves the built admin SPA from `admin/dist/` with SPA fallback. `/admin/events` SSE endpoint broadcasts every MCP request to connected admin browsers. `cookie-parser` middleware wired (Express 5 has no built-in). Startup logging prints port, engine, configured issuer URL (honors `--public-url`), registered-client count, DCR status, and admin bootstrap token. **v0.26.9 hardening pass:** F7 sets `remote: true` explicitly on the `/mcp` request handler's OperationContext literal (closes the HTTP shell-job RCE — without this, `submit_job`'s protected-name guard at `operations.ts:1391` saw a falsy undefined and skipped, letting a `read+write`-scoped OAuth token submit `shell` jobs). F8 wires `summarizeMcpParams` from `src/mcp/dispatch.ts` into both `mcp_request_log` writes and the admin SSE feed by default (raw payloads opt-in via `--log-full-params` with stderr warning). F9 sets cookie `Secure` flag when behind HTTPS or a public-URL proxy. F10 caps the magic-link nonce store with an LRU bound. F12 routes DCR disable through the `GBrainOAuthProvider` constructor's `dcrDisabled` option instead of the prior monkey-patch on the express router. F14 wraps `transport.handleRequest` in try/catch so SDK throws return a JSON-RPC 500 envelope instead of express's default HTML error page. F15 unifies OperationError + unexpected exceptions through `buildError` / `serializeError` so `/mcp` always returns the same envelope shape. **v0.28.1:** `/health` endpoint extracted into pure `probeHealth(engine)` async function with `HEALTH_TIMEOUT_MS = 3000` exported constant — drops the timeout from 5s to 3s so Fly.io's 5s health-check deadline gets 2s of headroom for TCP, response framing, and clock skew. Races `engine.getStats()` against the timeout via `Promise.race`; saturated pool returns 503 with `Health check timed out (database pool may be saturated)` instead of hanging. `clearTimeout` in finally block prevents pending-timer pile-up under high probe rates (race-leak fix from adversarial review). **v0.28.10:** `/health` is now liveness-only via the new `probeLiveness(sql, engineName, version, timeoutMs)` helper that races `sql\`SELECT 1\`` against `HEALTH_TIMEOUT_MS` and returns the same `ProbeHealthResult` tagged-union as `probeHealth` (single timer-cleanup site, single 503 envelope). Body shape: `{status, version, engine}` only — engine stats are no longer spread on the public route. Full stats moved to a new admin endpoint `/admin/api/full-stats` (sibling to `/admin/api/stats` and `/admin/api/health-indicators`) gated by the existing `requireAdmin` middleware; that route calls `probeHealth(engine, ...)` and returns the original spread-stats body. `?full=true` query param removed entirely. Closes the original DoS surface where `getStats()`'s 6× count(*) on 96K-page brains through PgBouncer exceeded `HEALTH_TIMEOUT_MS` and triggered orchestrator restart cascades (Fly.io / k8s seeing 503 → restart loop → advisory-lock pile-up on the migration lock). Outside-voice review (Codex) caught that `/admin/api/health-indicators` is NOT a full-stats endpoint (returns only `{expiring_soon, error_rate}`), and that an alternative loopback-IP gate would have depended on `app.set('trust proxy', 'loopback')` semantics holding under proxy/XFF misconfiguration; the shipped admin-cookie design avoids both.
|
||||
- `src/core/oauth-provider.ts` (v0.26.0) — `GBrainOAuthProvider` implementing the MCP SDK's `OAuthServerProvider` + `OAuthRegisteredClientsStore` interfaces. Backed by raw SQL (works on both PGLite and Postgres — OAuth is infrastructure, not a BrainEngine concern). Full OAuth 2.1 spec: `authorize` + `exchangeAuthorizationCode` with PKCE (for ChatGPT), `client_credentials` (for Perplexity / Claude), `refresh_token` with rotation, `revokeToken`, `registerClient` (DCR path validates redirect_uri must be `https://` or loopback per RFC 6749 §3.1.2.1). All tokens + client secrets SHA-256 hashed before storage. Auth codes single-use with 10-minute TTL via atomic `DELETE...RETURNING` (closes RFC 6749 §10.5 TOCTOU race). Refresh rotation also `DELETE...RETURNING` (closes §10.4 stolen-token detection bypass). `pgArray()` escapes commas/quotes/braces in elements so a comma-bearing redirect_uri can't smuggle a second array element. Legacy `access_tokens` fallback in `verifyAccessToken` grandfathers pre-v0.26 bearer tokens as `read+write+admin`. `sweepExpiredTokens()` runs on startup wrapped in try/catch. **v0.26.9 RFC 6749/7009 hardening pass:** F1+F2 fold `client_id` atomically into the `DELETE WHERE` clauses for both auth-code exchange and refresh rotation — pre-fix the post-hoc client compare burned the row on wrong-client paths so the legitimate client couldn't retry. F3 enforces refresh-scope-subset against the original grant on the row (RFC 6749 §6), not the client's currently-allowed scopes — fixes the case where revoking a scope from a client wouldn't shrink the agent's existing refresh tokens. F4 binds `client_id` on `revokeToken` so a client can only revoke its own tokens (RFC 7009 §2.1). F7c validates the `/token` request's `redirect_uri` against the value stored at `/authorize` (RFC 6749 §4.1.3) — empty-string treated as missing rather than wildcard match (adversarial-review fix). F5 swaps bare `catch {}` blocks in `verifyAccessToken` and `getClient` for `isUndefinedColumnError` from `src/core/utils.ts` — only SQLSTATE 42703 falls through to legacy fallback; lock timeouts and network blips throw and surface. F6 makes `sweepExpiredTokens()` actually return the count via `RETURNING 1` + array length, not a fire-and-forget zero. F12 adds `dcrDisabled` constructor option so `serve-http.ts` can disable the `/register` endpoint without monkey-patching the router. **v0.26.2:** module-private `coerceTimestamp()` boundary helper at the top of the file normalizes postgres-driver-as-string BIGINT columns to JS numbers at every read site (5 call sites: `getClient` L112+L113 for DCR `/register` RFC 7591 §3.2.1 numeric timestamps, `exchangeRefreshToken` L274 + `verifyAccessToken` L296+L303 for the SDK's `typeof === 'number'` bearerAuth check). Throws on non-finite input (NaN/Infinity) so corrupt rows fail loud at the boundary instead of riding through as `expiresAt: NaN`; returns undefined for SQL NULL so callers decide NULL semantics explicitly (refresh + access token paths treat NULL as expired). Helper intentionally NOT promoted to `src/core/utils.ts` — codex review flagged repo-wide BIGINT precision-loss risk for a generic helper.
|
||||
- `admin/` (v0.26.0) — React 19 + Vite + TypeScript admin SPA embedded in the binary via `admin/dist/` served by `serve-http.ts`. 7 screens: Login (bootstrap token → session cookie), Dashboard (metrics + SSE feed + token health), Agents (sortable table + sparklines + Register button), Register (modal with scope checkboxes + grant type selector), Credentials reveal (full-screen modal with Copy + Download JSON + yellow one-time-only warning), Request Log (filterable paginated), Agent Detail drawer (Details / Activity / Config Export tabs + Revoke). Design tokens: `#0a0a0f` bg, Inter for UI, JetBrains Mono for data, 4-32px spacing scale, rounded pill badges. HTTP-only SameSite=Strict cookie auth. 65KB gzip. Build: `cd admin && bun install && bun run build`; output at `admin/dist/` is committed for self-contained binaries.
|
||||
- `src/commands/auth.ts` — Token management. `gbrain auth create/list/revoke/test` for legacy bearer tokens (v0.22.7 wired as a first-class CLI subcommand) plus `gbrain auth register-client` (v0.26.0) and `gbrain auth revoke-client <client_id>` (v0.26.2) for OAuth 2.1 client lifecycle. `revoke-client` runs an atomic `DELETE...RETURNING` on `oauth_clients`; FK `ON DELETE CASCADE` on `oauth_tokens.client_id` and `oauth_codes.client_id` purges every active token + authorization code in a single transaction. `process.exit(1)` on no-such-client (idempotent — re-running on the same id produces the same exit-1 message). Legacy tokens stored as SHA-256 hashes in `access_tokens`; OAuth clients in `oauth_clients`. As of v0.26.0, legacy tokens grandfather to `read+write+admin` scopes on the OAuth HTTP server, so pre-v0.26 deployments keep working with no migration.
|
||||
@@ -531,8 +531,12 @@ E2E tests (`test/e2e/`): Run against real Postgres+pgvector. Require `DATABASE_U
|
||||
- Tier 2 (`skills.test.ts`) requires OpenClaw + API keys, runs nightly in CI
|
||||
- If `.env.testing` doesn't exist in this directory, check sibling worktrees for one:
|
||||
`find ../ -maxdepth 2 -name .env.testing -print -quit` and copy it here if found.
|
||||
- Always run E2E tests when they exist. Do not skip them just because DATABASE_URL
|
||||
is not set. Start the test DB, run the tests, then tear it down.
|
||||
- **Run E2E tests without asking permission.** When you want to verify behavior,
|
||||
there's a relevant E2E test, or you're shipping touching anything covered by an
|
||||
E2E suite — just spin up the test DB, run the tests, and tear down. Don't ask,
|
||||
don't propose it, don't defer. The lifecycle is short (~2-30s startup, sub-minute
|
||||
tests, instant teardown) and the gate value is high. Skipping with "DATABASE_URL
|
||||
unset" is silent regression, not caution.
|
||||
|
||||
### API keys and running ALL tests
|
||||
|
||||
@@ -554,7 +558,8 @@ When asked to "run all E2E tests" or "run tests", that means ALL tiers:
|
||||
### E2E test DB lifecycle (ALWAYS follow this)
|
||||
|
||||
You are responsible for spinning up and tearing down the test Postgres container.
|
||||
Do not leave containers running after tests. Do not skip E2E tests.
|
||||
Do not leave containers running after tests. Do not skip E2E tests, do not ask
|
||||
permission to run them — see the "run without asking" rule above.
|
||||
|
||||
1. **Check for `.env.testing`** — if missing, copy from sibling worktree.
|
||||
Read it to get the DATABASE_URL (it has the port number).
|
||||
@@ -569,9 +574,21 @@ Do not leave containers running after tests. Do not skip E2E tests.
|
||||
-p PORT:5432 pgvector/pgvector:pg16
|
||||
```
|
||||
Wait for ready: `docker exec gbrain-test-pg pg_isready -U postgres`
|
||||
4. **Run E2E tests:**
|
||||
4. **Bootstrap the schema** (required — fresh containers have no `oauth_clients`,
|
||||
`mcp_request_log`, `pages` etc.; tests like `serve-http-oauth.test.ts` will fail
|
||||
with `relation "oauth_clients" does not exist` if you skip this):
|
||||
```bash
|
||||
DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test \
|
||||
bun run src/cli.ts doctor --json > /dev/null 2>&1
|
||||
```
|
||||
`gbrain doctor` triggers `initSchema()` on first connect, which is the canonical
|
||||
way to bring a fresh DB to head. `apply-migrations --yes` alone does NOT seed
|
||||
the base schema — it runs ALTER-style migrations on top of `initSchema`. Tests
|
||||
that bypass the engine (raw `execSync`-spawned `auth register-client`) hit the
|
||||
schema directly and need this step to have run first.
|
||||
5. **Run E2E tests:**
|
||||
`DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test bun run test:e2e`
|
||||
5. **Tear down immediately after tests finish (pass or fail):**
|
||||
6. **Tear down immediately after tests finish (pass or fail):**
|
||||
`docker stop gbrain-test-pg && docker rm gbrain-test-pg`
|
||||
|
||||
Never leave `gbrain-test-pg` running. If you find a stale one from a previous run,
|
||||
|
||||
+23
-6
@@ -236,7 +236,7 @@ strict behavior when unset.
|
||||
- `src/mcp/server.ts` — MCP stdio server (generated from operations). v0.22.7: tool-call handler delegates to `dispatchToolCall` from `src/mcp/dispatch.ts` so stdio + HTTP transports share one validation, context-build, and error-format path.
|
||||
- `src/mcp/dispatch.ts` (v0.22.7) — Shared tool-call dispatch consumed by both stdio (`server.ts`) and HTTP transports. Exports `dispatchToolCall(engine, name, params, opts)`, `buildOperationContext(engine, params, opts)`, and `validateParams(op, params)`. Single source of truth for `(ctx, params)` handler arg order and the 5-field `OperationContext` shape (engine + config + logger + dryRun + remote). Defaults to `remote: true` (untrusted); local CLI callers pass `remote: false`. Closed F1/F2/F3 drift bugs in the original v0.22.5 HTTP transport. **v0.26.9 (F8):** adds `summarizeMcpParams(opName, params)` — privacy-preserving redactor for `mcp_request_log` and the admin SSE feed. Returns `{redacted, kind, declared_keys, unknown_key_count, approx_bytes}`. Intersects submitted top-level keys against the operation's declared `params` allow-list (declared keys preserved as a sorted array for debug visibility; unknown keys counted but never named, closing the attacker-controlled-key-name leak). Byte counts bucketed up to nearest 1KB so an attacker can't binary-search secret-content sizes via repeated probes. Operators on a personal laptop who want raw payload visibility opt back in with `gbrain serve --http --log-full-params` (loud stderr warning at startup). Canonical helper — new logging code paths route through it rather than `JSON.stringify(params)`.
|
||||
- `src/mcp/rate-limit.ts` (v0.22.7) — Bounded-LRU token-bucket limiter. `buildDefaultLimiters()` returns the two-bucket pipeline: pre-auth IP (30/60s, fires BEFORE the DB lookup so brute-force load against `access_tokens` is actually capped) + post-auth token-id (60/60s). Tracks `lastTouchedMs` separately from `lastRefillMs` so an exhausted key can't be reset by hammering past the TTL. LRU cap bounds memory under attacker-controlled key growth.
|
||||
- `src/commands/serve-http.ts` (v0.26.0) — Express 5 HTTP MCP server with OAuth 2.1, admin dashboard, and SSE live activity feed. Started via `gbrain serve --http [--port N] [--token-ttl N] [--enable-dcr] [--public-url URL] [--log-full-params]`. Supersedes the v0.22.7 `src/mcp/http-transport.ts` simple bearer-auth path. Combines MCP SDK's `mcpAuthRouter` (authorize / token / register / revoke endpoints), a custom `client_credentials` handler (SDK's token endpoint throws `UnsupportedGrantTypeError` for CC; the custom handler runs BEFORE the router and falls through for `auth_code` / `refresh_token`), `requireBearerAuth` middleware for `/mcp` with scope enforcement before op dispatch, `localOnly` rejection, and `express-rate-limit` at 50 req / 15 min on `/token`. Serves the built admin SPA from `admin/dist/` with SPA fallback. `/admin/events` SSE endpoint broadcasts every MCP request to connected admin browsers. `cookie-parser` middleware wired (Express 5 has no built-in). Startup logging prints port, engine, configured issuer URL (honors `--public-url`), registered-client count, DCR status, and admin bootstrap token. **v0.26.9 hardening pass:** F7 sets `remote: true` explicitly on the `/mcp` request handler's OperationContext literal (closes the HTTP shell-job RCE — without this, `submit_job`'s protected-name guard at `operations.ts:1391` saw a falsy undefined and skipped, letting a `read+write`-scoped OAuth token submit `shell` jobs). F8 wires `summarizeMcpParams` from `src/mcp/dispatch.ts` into both `mcp_request_log` writes and the admin SSE feed by default (raw payloads opt-in via `--log-full-params` with stderr warning). F9 sets cookie `Secure` flag when behind HTTPS or a public-URL proxy. F10 caps the magic-link nonce store with an LRU bound. F12 routes DCR disable through the `GBrainOAuthProvider` constructor's `dcrDisabled` option instead of the prior monkey-patch on the express router. F14 wraps `transport.handleRequest` in try/catch so SDK throws return a JSON-RPC 500 envelope instead of express's default HTML error page. F15 unifies OperationError + unexpected exceptions through `buildError` / `serializeError` so `/mcp` always returns the same envelope shape. **v0.28.1:** `/health` endpoint extracted into pure `probeHealth(engine)` async function with `HEALTH_TIMEOUT_MS = 3000` exported constant — drops the timeout from 5s to 3s so Fly.io's 5s health-check deadline gets 2s of headroom for TCP, response framing, and clock skew. Races `engine.getStats()` against the timeout via `Promise.race`; saturated pool returns 503 with `Health check timed out (database pool may be saturated)` instead of hanging. `clearTimeout` in finally block prevents pending-timer pile-up under high probe rates (race-leak fix from adversarial review).
|
||||
- `src/commands/serve-http.ts` (v0.26.0) — Express 5 HTTP MCP server with OAuth 2.1, admin dashboard, and SSE live activity feed. Started via `gbrain serve --http [--port N] [--token-ttl N] [--enable-dcr] [--public-url URL] [--log-full-params]`. Supersedes the v0.22.7 `src/mcp/http-transport.ts` simple bearer-auth path. Combines MCP SDK's `mcpAuthRouter` (authorize / token / register / revoke endpoints), a custom `client_credentials` handler (SDK's token endpoint throws `UnsupportedGrantTypeError` for CC; the custom handler runs BEFORE the router and falls through for `auth_code` / `refresh_token`), `requireBearerAuth` middleware for `/mcp` with scope enforcement before op dispatch, `localOnly` rejection, and `express-rate-limit` at 50 req / 15 min on `/token`. Serves the built admin SPA from `admin/dist/` with SPA fallback. `/admin/events` SSE endpoint broadcasts every MCP request to connected admin browsers. `cookie-parser` middleware wired (Express 5 has no built-in). Startup logging prints port, engine, configured issuer URL (honors `--public-url`), registered-client count, DCR status, and admin bootstrap token. **v0.26.9 hardening pass:** F7 sets `remote: true` explicitly on the `/mcp` request handler's OperationContext literal (closes the HTTP shell-job RCE — without this, `submit_job`'s protected-name guard at `operations.ts:1391` saw a falsy undefined and skipped, letting a `read+write`-scoped OAuth token submit `shell` jobs). F8 wires `summarizeMcpParams` from `src/mcp/dispatch.ts` into both `mcp_request_log` writes and the admin SSE feed by default (raw payloads opt-in via `--log-full-params` with stderr warning). F9 sets cookie `Secure` flag when behind HTTPS or a public-URL proxy. F10 caps the magic-link nonce store with an LRU bound. F12 routes DCR disable through the `GBrainOAuthProvider` constructor's `dcrDisabled` option instead of the prior monkey-patch on the express router. F14 wraps `transport.handleRequest` in try/catch so SDK throws return a JSON-RPC 500 envelope instead of express's default HTML error page. F15 unifies OperationError + unexpected exceptions through `buildError` / `serializeError` so `/mcp` always returns the same envelope shape. **v0.28.1:** `/health` endpoint extracted into pure `probeHealth(engine)` async function with `HEALTH_TIMEOUT_MS = 3000` exported constant — drops the timeout from 5s to 3s so Fly.io's 5s health-check deadline gets 2s of headroom for TCP, response framing, and clock skew. Races `engine.getStats()` against the timeout via `Promise.race`; saturated pool returns 503 with `Health check timed out (database pool may be saturated)` instead of hanging. `clearTimeout` in finally block prevents pending-timer pile-up under high probe rates (race-leak fix from adversarial review). **v0.28.10:** `/health` is now liveness-only via the new `probeLiveness(sql, engineName, version, timeoutMs)` helper that races `sql\`SELECT 1\`` against `HEALTH_TIMEOUT_MS` and returns the same `ProbeHealthResult` tagged-union as `probeHealth` (single timer-cleanup site, single 503 envelope). Body shape: `{status, version, engine}` only — engine stats are no longer spread on the public route. Full stats moved to a new admin endpoint `/admin/api/full-stats` (sibling to `/admin/api/stats` and `/admin/api/health-indicators`) gated by the existing `requireAdmin` middleware; that route calls `probeHealth(engine, ...)` and returns the original spread-stats body. `?full=true` query param removed entirely. Closes the original DoS surface where `getStats()`'s 6× count(*) on 96K-page brains through PgBouncer exceeded `HEALTH_TIMEOUT_MS` and triggered orchestrator restart cascades (Fly.io / k8s seeing 503 → restart loop → advisory-lock pile-up on the migration lock). Outside-voice review (Codex) caught that `/admin/api/health-indicators` is NOT a full-stats endpoint (returns only `{expiring_soon, error_rate}`), and that an alternative loopback-IP gate would have depended on `app.set('trust proxy', 'loopback')` semantics holding under proxy/XFF misconfiguration; the shipped admin-cookie design avoids both.
|
||||
- `src/core/oauth-provider.ts` (v0.26.0) — `GBrainOAuthProvider` implementing the MCP SDK's `OAuthServerProvider` + `OAuthRegisteredClientsStore` interfaces. Backed by raw SQL (works on both PGLite and Postgres — OAuth is infrastructure, not a BrainEngine concern). Full OAuth 2.1 spec: `authorize` + `exchangeAuthorizationCode` with PKCE (for ChatGPT), `client_credentials` (for Perplexity / Claude), `refresh_token` with rotation, `revokeToken`, `registerClient` (DCR path validates redirect_uri must be `https://` or loopback per RFC 6749 §3.1.2.1). All tokens + client secrets SHA-256 hashed before storage. Auth codes single-use with 10-minute TTL via atomic `DELETE...RETURNING` (closes RFC 6749 §10.5 TOCTOU race). Refresh rotation also `DELETE...RETURNING` (closes §10.4 stolen-token detection bypass). `pgArray()` escapes commas/quotes/braces in elements so a comma-bearing redirect_uri can't smuggle a second array element. Legacy `access_tokens` fallback in `verifyAccessToken` grandfathers pre-v0.26 bearer tokens as `read+write+admin`. `sweepExpiredTokens()` runs on startup wrapped in try/catch. **v0.26.9 RFC 6749/7009 hardening pass:** F1+F2 fold `client_id` atomically into the `DELETE WHERE` clauses for both auth-code exchange and refresh rotation — pre-fix the post-hoc client compare burned the row on wrong-client paths so the legitimate client couldn't retry. F3 enforces refresh-scope-subset against the original grant on the row (RFC 6749 §6), not the client's currently-allowed scopes — fixes the case where revoking a scope from a client wouldn't shrink the agent's existing refresh tokens. F4 binds `client_id` on `revokeToken` so a client can only revoke its own tokens (RFC 7009 §2.1). F7c validates the `/token` request's `redirect_uri` against the value stored at `/authorize` (RFC 6749 §4.1.3) — empty-string treated as missing rather than wildcard match (adversarial-review fix). F5 swaps bare `catch {}` blocks in `verifyAccessToken` and `getClient` for `isUndefinedColumnError` from `src/core/utils.ts` — only SQLSTATE 42703 falls through to legacy fallback; lock timeouts and network blips throw and surface. F6 makes `sweepExpiredTokens()` actually return the count via `RETURNING 1` + array length, not a fire-and-forget zero. F12 adds `dcrDisabled` constructor option so `serve-http.ts` can disable the `/register` endpoint without monkey-patching the router. **v0.26.2:** module-private `coerceTimestamp()` boundary helper at the top of the file normalizes postgres-driver-as-string BIGINT columns to JS numbers at every read site (5 call sites: `getClient` L112+L113 for DCR `/register` RFC 7591 §3.2.1 numeric timestamps, `exchangeRefreshToken` L274 + `verifyAccessToken` L296+L303 for the SDK's `typeof === 'number'` bearerAuth check). Throws on non-finite input (NaN/Infinity) so corrupt rows fail loud at the boundary instead of riding through as `expiresAt: NaN`; returns undefined for SQL NULL so callers decide NULL semantics explicitly (refresh + access token paths treat NULL as expired). Helper intentionally NOT promoted to `src/core/utils.ts` — codex review flagged repo-wide BIGINT precision-loss risk for a generic helper.
|
||||
- `admin/` (v0.26.0) — React 19 + Vite + TypeScript admin SPA embedded in the binary via `admin/dist/` served by `serve-http.ts`. 7 screens: Login (bootstrap token → session cookie), Dashboard (metrics + SSE feed + token health), Agents (sortable table + sparklines + Register button), Register (modal with scope checkboxes + grant type selector), Credentials reveal (full-screen modal with Copy + Download JSON + yellow one-time-only warning), Request Log (filterable paginated), Agent Detail drawer (Details / Activity / Config Export tabs + Revoke). Design tokens: `#0a0a0f` bg, Inter for UI, JetBrains Mono for data, 4-32px spacing scale, rounded pill badges. HTTP-only SameSite=Strict cookie auth. 65KB gzip. Build: `cd admin && bun install && bun run build`; output at `admin/dist/` is committed for self-contained binaries.
|
||||
- `src/commands/auth.ts` — Token management. `gbrain auth create/list/revoke/test` for legacy bearer tokens (v0.22.7 wired as a first-class CLI subcommand) plus `gbrain auth register-client` (v0.26.0) and `gbrain auth revoke-client <client_id>` (v0.26.2) for OAuth 2.1 client lifecycle. `revoke-client` runs an atomic `DELETE...RETURNING` on `oauth_clients`; FK `ON DELETE CASCADE` on `oauth_tokens.client_id` and `oauth_codes.client_id` purges every active token + authorization code in a single transaction. `process.exit(1)` on no-such-client (idempotent — re-running on the same id produces the same exit-1 message). Legacy tokens stored as SHA-256 hashes in `access_tokens`; OAuth clients in `oauth_clients`. As of v0.26.0, legacy tokens grandfather to `read+write+admin` scopes on the OAuth HTTP server, so pre-v0.26 deployments keep working with no migration.
|
||||
@@ -628,8 +628,12 @@ E2E tests (`test/e2e/`): Run against real Postgres+pgvector. Require `DATABASE_U
|
||||
- Tier 2 (`skills.test.ts`) requires OpenClaw + API keys, runs nightly in CI
|
||||
- If `.env.testing` doesn't exist in this directory, check sibling worktrees for one:
|
||||
`find ../ -maxdepth 2 -name .env.testing -print -quit` and copy it here if found.
|
||||
- Always run E2E tests when they exist. Do not skip them just because DATABASE_URL
|
||||
is not set. Start the test DB, run the tests, then tear it down.
|
||||
- **Run E2E tests without asking permission.** When you want to verify behavior,
|
||||
there's a relevant E2E test, or you're shipping touching anything covered by an
|
||||
E2E suite — just spin up the test DB, run the tests, and tear down. Don't ask,
|
||||
don't propose it, don't defer. The lifecycle is short (~2-30s startup, sub-minute
|
||||
tests, instant teardown) and the gate value is high. Skipping with "DATABASE_URL
|
||||
unset" is silent regression, not caution.
|
||||
|
||||
### API keys and running ALL tests
|
||||
|
||||
@@ -651,7 +655,8 @@ When asked to "run all E2E tests" or "run tests", that means ALL tiers:
|
||||
### E2E test DB lifecycle (ALWAYS follow this)
|
||||
|
||||
You are responsible for spinning up and tearing down the test Postgres container.
|
||||
Do not leave containers running after tests. Do not skip E2E tests.
|
||||
Do not leave containers running after tests. Do not skip E2E tests, do not ask
|
||||
permission to run them — see the "run without asking" rule above.
|
||||
|
||||
1. **Check for `.env.testing`** — if missing, copy from sibling worktree.
|
||||
Read it to get the DATABASE_URL (it has the port number).
|
||||
@@ -666,9 +671,21 @@ Do not leave containers running after tests. Do not skip E2E tests.
|
||||
-p PORT:5432 pgvector/pgvector:pg16
|
||||
```
|
||||
Wait for ready: `docker exec gbrain-test-pg pg_isready -U postgres`
|
||||
4. **Run E2E tests:**
|
||||
4. **Bootstrap the schema** (required — fresh containers have no `oauth_clients`,
|
||||
`mcp_request_log`, `pages` etc.; tests like `serve-http-oauth.test.ts` will fail
|
||||
with `relation "oauth_clients" does not exist` if you skip this):
|
||||
```bash
|
||||
DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test \
|
||||
bun run src/cli.ts doctor --json > /dev/null 2>&1
|
||||
```
|
||||
`gbrain doctor` triggers `initSchema()` on first connect, which is the canonical
|
||||
way to bring a fresh DB to head. `apply-migrations --yes` alone does NOT seed
|
||||
the base schema — it runs ALTER-style migrations on top of `initSchema`. Tests
|
||||
that bypass the engine (raw `execSync`-spawned `auth register-client`) hit the
|
||||
schema directly and need this step to have run first.
|
||||
5. **Run E2E tests:**
|
||||
`DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test bun run test:e2e`
|
||||
5. **Tear down immediately after tests finish (pass or fail):**
|
||||
6. **Tear down immediately after tests finish (pass or fail):**
|
||||
`docker stop gbrain-test-pg && docker rm gbrain-test-pg`
|
||||
|
||||
Never leave `gbrain-test-pg` running. If you find a stale one from a previous run,
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "gbrain",
|
||||
"version": "0.28.9",
|
||||
"version": "0.28.10",
|
||||
"description": "Postgres-native personal knowledge brain with hybrid RAG search",
|
||||
"type": "module",
|
||||
"main": "src/core/index.ts",
|
||||
|
||||
+130
-21
@@ -93,6 +93,51 @@ export async function probeHealth(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Lightweight liveness probe. Races `SELECT 1` against the same timeout
|
||||
* `probeHealth` uses, returns the same tagged-union result type, but the
|
||||
* 200 body is intentionally bare: `{status, version, engine}` — no engine
|
||||
* stats. Stats moved to `/admin/api/full-stats` (admin auth) in v0.28.10
|
||||
* because `getStats()`'s six count(*) queries exceeded HEALTH_TIMEOUT_MS
|
||||
* on production brains through PgBouncer, producing false 503s that
|
||||
* triggered orchestrator restart cascades and advisory-lock pile-ups.
|
||||
*/
|
||||
export async function probeLiveness(
|
||||
sql: SqlQuery,
|
||||
engineName: string,
|
||||
version: string,
|
||||
timeoutMs: number = HEALTH_TIMEOUT_MS,
|
||||
): Promise<ProbeHealthResult> {
|
||||
let timer: ReturnType<typeof setTimeout> | null = null;
|
||||
try {
|
||||
await Promise.race([
|
||||
sql`SELECT 1`,
|
||||
new Promise<never>((_, reject) => {
|
||||
timer = setTimeout(() => reject(new Error('health_timeout')), timeoutMs);
|
||||
}),
|
||||
]);
|
||||
return {
|
||||
ok: true,
|
||||
status: 200,
|
||||
body: { status: 'ok', version, engine: engineName },
|
||||
};
|
||||
} catch (e: unknown) {
|
||||
const msg = e instanceof Error ? e.message : 'unknown';
|
||||
return {
|
||||
ok: false,
|
||||
status: 503,
|
||||
body: {
|
||||
error: 'service_unavailable',
|
||||
error_description: msg === 'health_timeout'
|
||||
? 'Health check timed out (database pool may be saturated)'
|
||||
: 'Database connection failed',
|
||||
},
|
||||
};
|
||||
} finally {
|
||||
if (timer !== null) clearTimeout(timer);
|
||||
}
|
||||
}
|
||||
|
||||
interface ServeHttpOptions {
|
||||
port: number;
|
||||
tokenTtl: number;
|
||||
@@ -287,10 +332,11 @@ export async function runServeHttp(engine: BrainEngine, options: ServeHttpOption
|
||||
app.use(authRouter);
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Health check
|
||||
// Health check — liveness only. Full engine stats live at
|
||||
// /admin/api/full-stats (requireAdmin). See probeLiveness above for the why.
|
||||
// ---------------------------------------------------------------------------
|
||||
app.get('/health', async (_req, res) => {
|
||||
const result = await probeHealth(engine, config.engine || 'pglite', VERSION);
|
||||
const result = await probeLiveness(sql, config.engine || 'pglite', VERSION);
|
||||
res.status(result.status).json(result.body);
|
||||
});
|
||||
|
||||
@@ -530,6 +576,16 @@ export async function runServeHttp(engine: BrainEngine, options: ServeHttpOption
|
||||
}
|
||||
});
|
||||
|
||||
// Full engine stats. v0.28.10 moved this off /health (which is now liveness
|
||||
// only — see probeLiveness) so dashboards needing page_count / chunk_count
|
||||
// / etc. authenticate as admin and call this endpoint. probeHealth races
|
||||
// engine.getStats() against HEALTH_TIMEOUT_MS so a saturated pool returns
|
||||
// 503 rather than hanging.
|
||||
app.get('/admin/api/full-stats', requireAdmin, async (_req: Request, res: Response) => {
|
||||
const result = await probeHealth(engine, config.engine || 'pglite', VERSION);
|
||||
res.status(result.status).json(result.body);
|
||||
});
|
||||
|
||||
app.get('/admin/api/requests', requireAdmin, async (req: Request, res: Response) => {
|
||||
try {
|
||||
const page = parseInt(req.query.page as string) || 1;
|
||||
@@ -704,30 +760,66 @@ export async function runServeHttp(engine: BrainEngine, options: ServeHttpOption
|
||||
{ capabilities: { tools: {} } },
|
||||
);
|
||||
|
||||
server.setRequestHandler(ListToolsRequestSchema, async () => ({
|
||||
tools: mcpOperations.map(op => ({
|
||||
name: op.name,
|
||||
description: op.description,
|
||||
inputSchema: {
|
||||
type: 'object' as const,
|
||||
properties: Object.fromEntries(
|
||||
Object.entries(op.params).map(([k, v]) => [k, {
|
||||
type: v.type,
|
||||
description: v.description,
|
||||
...(v.enum ? { enum: v.enum } : {}),
|
||||
...(v.default !== undefined ? { default: v.default } : {}),
|
||||
}]),
|
||||
),
|
||||
required: Object.entries(op.params).filter(([, v]) => v.required).map(([k]) => k),
|
||||
},
|
||||
})),
|
||||
}));
|
||||
server.setRequestHandler(ListToolsRequestSchema, async () => {
|
||||
// v0.28.10: log every JSON-RPC method, not just successful tools/call.
|
||||
// Pre-fix, /admin/api/requests showed nothing for clients that only
|
||||
// ever called tools/list, and the v0.26.3 persistence regression test
|
||||
// asserting >= 2 rows after tools/list + tools/call was unreachable.
|
||||
const latency = Date.now() - startTime;
|
||||
try {
|
||||
await sql`INSERT INTO mcp_request_log (token_name, agent_name, operation, latency_ms, status, params)
|
||||
VALUES (${authInfo.clientId}, ${agentName}, ${'tools/list'}, ${latency}, ${'success'}, ${null})`;
|
||||
} catch { /* best effort */ }
|
||||
broadcastEvent({
|
||||
agent: agentName,
|
||||
operation: 'tools/list',
|
||||
scopes: authInfo.scopes.join(','),
|
||||
latency_ms: latency,
|
||||
status: 'success',
|
||||
timestamp: new Date().toISOString(),
|
||||
});
|
||||
return {
|
||||
tools: mcpOperations.map(op => ({
|
||||
name: op.name,
|
||||
description: op.description,
|
||||
inputSchema: {
|
||||
type: 'object' as const,
|
||||
properties: Object.fromEntries(
|
||||
Object.entries(op.params).map(([k, v]) => [k, {
|
||||
type: v.type,
|
||||
description: v.description,
|
||||
...(v.enum ? { enum: v.enum } : {}),
|
||||
...(v.default !== undefined ? { default: v.default } : {}),
|
||||
}]),
|
||||
),
|
||||
required: Object.entries(op.params).filter(([, v]) => v.required).map(([k]) => k),
|
||||
},
|
||||
})),
|
||||
};
|
||||
});
|
||||
|
||||
server.setRequestHandler(CallToolRequestSchema, async (request) => {
|
||||
const { name, arguments: params } = request.params;
|
||||
const op = mcpOperations.find(o => o.name === name);
|
||||
if (!op) {
|
||||
return { content: [{ type: 'text', text: JSON.stringify({ error: 'unknown_operation', message: `Unknown: ${name}` }) }] };
|
||||
// v0.28.10: persist unknown-op attempts. Operators investigating
|
||||
// misbehaving agents need to see the full attempt log, not just
|
||||
// valid-op success/error.
|
||||
const latency = Date.now() - startTime;
|
||||
try {
|
||||
await sql`INSERT INTO mcp_request_log (token_name, agent_name, operation, latency_ms, status, params, error_message)
|
||||
VALUES (${authInfo.clientId}, ${agentName}, ${name}, ${latency}, ${'error'}, ${null}, ${`unknown_operation: ${name}`})`;
|
||||
} catch { /* best effort */ }
|
||||
broadcastEvent({
|
||||
agent: agentName,
|
||||
operation: name,
|
||||
scopes: authInfo.scopes.join(','),
|
||||
latency_ms: latency,
|
||||
status: 'error',
|
||||
error: { code: 'unknown_operation', message: `Unknown: ${name}` },
|
||||
timestamp: new Date().toISOString(),
|
||||
});
|
||||
return { content: [{ type: 'text', text: JSON.stringify({ error: 'unknown_operation', message: `Unknown: ${name}` }) }], isError: true };
|
||||
}
|
||||
|
||||
// Scope enforcement (v0.28: hasScope replaces exact-string-match so
|
||||
@@ -737,6 +829,23 @@ export async function runServeHttp(engine: BrainEngine, options: ServeHttpOption
|
||||
// sources_admin tokens look like they couldn't even read.)
|
||||
const requiredScope = op.scope || 'read';
|
||||
if (!hasScope(authInfo.scopes, requiredScope)) {
|
||||
// v0.28.10: persist scope-rejected attempts. Same operator-visibility
|
||||
// motivation as the unknown-op path — and it makes the v0.26.3
|
||||
// persistence regression test reliable across both rejection paths.
|
||||
const latency = Date.now() - startTime;
|
||||
try {
|
||||
await sql`INSERT INTO mcp_request_log (token_name, agent_name, operation, latency_ms, status, params, error_message)
|
||||
VALUES (${authInfo.clientId}, ${agentName}, ${name}, ${latency}, ${'error'}, ${null}, ${`insufficient_scope: requires '${requiredScope}'`})`;
|
||||
} catch { /* best effort */ }
|
||||
broadcastEvent({
|
||||
agent: agentName,
|
||||
operation: name,
|
||||
scopes: authInfo.scopes.join(','),
|
||||
latency_ms: latency,
|
||||
status: 'error',
|
||||
error: { code: 'insufficient_scope', message: `requires '${requiredScope}'` },
|
||||
timestamp: new Date().toISOString(),
|
||||
});
|
||||
return {
|
||||
content: [{
|
||||
type: 'text',
|
||||
|
||||
@@ -43,8 +43,17 @@ describeE2E('serve-http OAuth 2.1 E2E (v0.26.1 + v0.26.2 + v0.26.3)', () => {
|
||||
// via process.env mutation, which is invisible to subprocesses unless we
|
||||
// explicitly re-pass process.env. Same pattern applies to every execSync
|
||||
// in this file.
|
||||
// v0.28.10: register with admin scope so the F7 protected-name guard
|
||||
// tests can mint admin-scoped tokens that actually exercise the guard
|
||||
// at operations.ts:1527. Without admin in the client's allowed scopes,
|
||||
// submit_job for a protected name (`shell`, `subagent`) gets rejected
|
||||
// by hasScope() in serve-http.ts BEFORE reaching the F7 guard, so the
|
||||
// test was validating scope enforcement instead of the RCE protection.
|
||||
// Other tests that mint specific subsets ('read', 'read write') still
|
||||
// get the subset they ask for — adding admin to the client's allowed
|
||||
// ceiling does not auto-grant it to every minted token.
|
||||
const regOutput = execSync(
|
||||
'bun run src/cli.ts auth register-client e2e-oauth-test --grant-types client_credentials --scopes "read write"',
|
||||
'bun run src/cli.ts auth register-client e2e-oauth-test --grant-types client_credentials --scopes "read write admin"',
|
||||
{ cwd: process.cwd(), encoding: 'utf8', env: { ...process.env } }
|
||||
);
|
||||
const idMatch = regOutput.match(/Client ID:\s+(gbrain_cl_\S+)/);
|
||||
@@ -286,22 +295,78 @@ describeE2E('serve-http OAuth 2.1 E2E (v0.26.1 + v0.26.2 + v0.26.3)', () => {
|
||||
}, 15_000);
|
||||
|
||||
// =========================================================================
|
||||
// Health endpoint (no auth required)
|
||||
// Health endpoint (no auth required) — v0.28.10 made /health liveness-only;
|
||||
// engine stats moved to /admin/api/full-stats behind requireAdmin so a
|
||||
// saturated pool can't pin /health and trigger orchestrator restart cascades.
|
||||
// =========================================================================
|
||||
|
||||
test('health endpoint returns OK without auth', async () => {
|
||||
test('v0.28.10: /health returns liveness-only body (no engine stats)', async () => {
|
||||
const res = await fetch(`${BASE}/health`);
|
||||
expect(res.ok).toBe(true);
|
||||
const data = await res.json() as any;
|
||||
expect(data.status).toBe('ok');
|
||||
expect(data.version).toBeDefined();
|
||||
// page_count: the endpoint must return a non-negative integer. The exact
|
||||
// value depends on the deployment's brain state and is not what this test
|
||||
// is checking — pre-v0.26.2 this asserted `> 0` and broke on fresh schemas.
|
||||
expect(typeof data.page_count).toBe('number');
|
||||
expect(data.page_count).toBeGreaterThanOrEqual(0);
|
||||
expect(data.engine).toBeDefined();
|
||||
// Regression: pre-v0.28.10 /health spread getStats() (page_count,
|
||||
// chunk_count, etc.) into the body. The whole point of the v0.28.10
|
||||
// split is that /health stops touching those tables. If page_count
|
||||
// ever reappears here, the heavy probe leaked back into the public
|
||||
// route and the original DoS surface is back.
|
||||
expect(data.page_count).toBeUndefined();
|
||||
expect(data.chunk_count).toBeUndefined();
|
||||
expect(data.embedded_count).toBeUndefined();
|
||||
// Body shape is exactly {status, version, engine}.
|
||||
expect(Object.keys(data).sort()).toEqual(['engine', 'status', 'version']);
|
||||
});
|
||||
|
||||
test('v0.28.10: /admin/api/full-stats without admin cookie returns 401', async () => {
|
||||
const res = await fetch(`${BASE}/admin/api/full-stats`);
|
||||
expect(res.status).toBe(401);
|
||||
const data = await res.json() as any;
|
||||
expect(data.error).toBe('Admin authentication required');
|
||||
});
|
||||
|
||||
test('v0.28.10: /admin/api/full-stats with valid admin cookie returns getStats() body', async () => {
|
||||
// Same magic-link cookie dance the existing single-use test uses.
|
||||
// Skip gracefully if the bootstrap token isn't extractable — the 401
|
||||
// case above pins the auth gate; this test pins the happy path.
|
||||
const stderrBuf = (serverProcess as any)?._stderrBuffer || '';
|
||||
const tokenMatch = String(stderrBuf).match(/Admin Token[\s\S]*?([a-f0-9]{32,64})/);
|
||||
if (!tokenMatch) {
|
||||
console.warn('[e2e] skipped /admin/api/full-stats happy path: could not extract bootstrap token');
|
||||
return;
|
||||
}
|
||||
const bootstrapToken = tokenMatch[1];
|
||||
|
||||
const issueRes = await fetch(`${BASE}/admin/api/issue-magic-link`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json', Authorization: `Bearer ${bootstrapToken}` },
|
||||
body: '{}',
|
||||
});
|
||||
expect(issueRes.ok).toBe(true);
|
||||
const { url } = await issueRes.json() as any;
|
||||
|
||||
const click = await fetch(url, { redirect: 'manual' });
|
||||
expect(click.status).toBe(302);
|
||||
const setCookie = click.headers.get('set-cookie') || '';
|
||||
const cookieMatch = setCookie.match(/gbrain_admin=([^;]+)/);
|
||||
expect(cookieMatch).toBeTruthy();
|
||||
const cookieValue = cookieMatch![1];
|
||||
|
||||
const statsRes = await fetch(`${BASE}/admin/api/full-stats`, {
|
||||
headers: { Cookie: `gbrain_admin=${cookieValue}` },
|
||||
});
|
||||
expect(statsRes.ok).toBe(true);
|
||||
const stats = await statsRes.json() as any;
|
||||
expect(stats.status).toBe('ok');
|
||||
expect(stats.version).toBeDefined();
|
||||
expect(stats.engine).toBeDefined();
|
||||
// The full-stats body is probeHealth's spread of getStats() — page_count
|
||||
// is the canonical signal that we're hitting the heavy path here.
|
||||
expect(typeof stats.page_count).toBe('number');
|
||||
expect(stats.page_count).toBeGreaterThanOrEqual(0);
|
||||
}, 15_000);
|
||||
|
||||
// =========================================================================
|
||||
// Token lifecycle
|
||||
// =========================================================================
|
||||
@@ -476,8 +541,9 @@ describeE2E('serve-http OAuth 2.1 E2E (v0.26.1 + v0.26.2 + v0.26.3)', () => {
|
||||
expect(okRes.status).not.toBe(401);
|
||||
|
||||
// Trigger an error path so the error_message column gets a value too.
|
||||
// Request a tool that doesn't exist — server returns an MCP error in
|
||||
// the body but the underlying handler logs status='error' to mcp_request_log.
|
||||
// Request a tool that doesn't exist — v0.28.10 logs unknown-op attempts
|
||||
// with operation = the attempted name and error_message starting with
|
||||
// 'unknown_operation:'.
|
||||
await mcpCall(access_token, 'tools/call', { name: 'this_tool_does_not_exist', arguments: {} });
|
||||
|
||||
// Allow async best-effort INSERT to flush.
|
||||
@@ -498,18 +564,26 @@ describeE2E('serve-http OAuth 2.1 E2E (v0.26.1 + v0.26.2 + v0.26.3)', () => {
|
||||
expect(row.agent_name).toBe('e2e-oauth-test');
|
||||
}
|
||||
|
||||
// params persisted as JSONB (postgres-js returns object form).
|
||||
// The params field is non-null on tools/call (carries the call args)
|
||||
// and on tools/list (carries an empty {} or undefined depending on payload).
|
||||
const callRow = rows.find(r => r.operation === 'tools/call');
|
||||
// v0.28.10: tools/list logs as operation='tools/list' (the JSON-RPC
|
||||
// method name). tools/call success/error logs as operation=<inner
|
||||
// tool name> (the convention preserved from pre-v0.28.10 dispatch
|
||||
// logging — agents querying mcp_request_log filter by tool name, not
|
||||
// by JSON-RPC method).
|
||||
const listRow = rows.find(r => r.operation === 'tools/list');
|
||||
expect(listRow).toBeDefined();
|
||||
expect(listRow!.status).toBe('success');
|
||||
|
||||
// The unknown-op call shows up with operation = the attempted name.
|
||||
const callRow = rows.find(r => r.operation === 'this_tool_does_not_exist');
|
||||
expect(callRow).toBeDefined();
|
||||
expect(callRow!.params).toBeDefined();
|
||||
expect(callRow!.status).toBe('error');
|
||||
|
||||
// error_message populated on the failed call.
|
||||
const errorRow = rows.find(r => r.status === 'error');
|
||||
expect(errorRow).toBeDefined();
|
||||
expect(errorRow!.error_message).toBeTruthy();
|
||||
expect(typeof errorRow!.error_message).toBe('string');
|
||||
expect(errorRow!.error_message as string).toContain('unknown_operation');
|
||||
} finally {
|
||||
await sql.end();
|
||||
}
|
||||
@@ -758,7 +832,12 @@ describeE2E('serve-http OAuth 2.1 E2E (v0.26.1 + v0.26.2 + v0.26.3)', () => {
|
||||
// Together they close the path even if either layer regresses alone.
|
||||
|
||||
test('F7: HTTP MCP cannot submit shell jobs (RCE regression)', async () => {
|
||||
const { access_token } = await mintToken('read write');
|
||||
// v0.28.10: must mint admin scope. submit_job's required scope is
|
||||
// 'admin'; without it, hasScope() rejects with insufficient_scope BEFORE
|
||||
// the F7 protected-name guard at operations.ts:1527 fires. To validate
|
||||
// the actual RCE protection (the protected-name guard), the token has
|
||||
// to clear the scope check first.
|
||||
const { access_token } = await mintToken('admin');
|
||||
const res = await mcpCall(access_token, 'tools/call', {
|
||||
name: 'submit_job',
|
||||
arguments: { name: 'shell', data: { cmd: 'id' } },
|
||||
@@ -782,7 +861,8 @@ describeE2E('serve-http OAuth 2.1 E2E (v0.26.1 + v0.26.2 + v0.26.3)', () => {
|
||||
}, 15_000);
|
||||
|
||||
test('F7: HTTP MCP cannot submit subagent jobs (protected name)', async () => {
|
||||
const { access_token } = await mintToken('read write');
|
||||
// Same admin-scope requirement as the shell-job sibling test above.
|
||||
const { access_token } = await mintToken('admin');
|
||||
const res = await mcpCall(access_token, 'tools/call', {
|
||||
name: 'submit_job',
|
||||
arguments: { name: 'subagent', data: { prompt: 'noop' } },
|
||||
|
||||
@@ -1,18 +1,24 @@
|
||||
/**
|
||||
* Tests for probeHealth() and HEALTH_TIMEOUT_MS in src/commands/serve-http.ts.
|
||||
* Tests for probeHealth(), probeLiveness(), and HEALTH_TIMEOUT_MS in
|
||||
* src/commands/serve-http.ts.
|
||||
*
|
||||
* Calls probeHealth() directly with a mock engine — no Express test client,
|
||||
* no module mocking. Three branches of the route (happy / timeout / db-error)
|
||||
* each get a unit test, plus a sanity assertion on the exported constant.
|
||||
* v0.28.10 split: /health now calls probeLiveness (sql`SELECT 1`); the heavier
|
||||
* probeHealth (engine.getStats()) moved behind requireAdmin at
|
||||
* /admin/api/full-stats. Both share ProbeHealthResult so the route handlers
|
||||
* stay 2-line dispatches.
|
||||
*
|
||||
* Calls each probe directly with a mock — no Express test client, no module
|
||||
* mocking. Each probe gets happy / timeout / db-error coverage.
|
||||
*
|
||||
* Express-layer wiring (timeout actually propagates through the route, body
|
||||
* shape after JSON serialization) is covered by the new /health timeout case
|
||||
* in test/e2e/serve-http-oauth.test.ts.
|
||||
* shape after JSON serialization) is covered by /health + /admin/api/full-stats
|
||||
* cases in test/e2e/serve-http-oauth.test.ts.
|
||||
*/
|
||||
|
||||
import { describe, test, expect } from 'bun:test';
|
||||
import { HEALTH_TIMEOUT_MS, probeHealth } from '../src/commands/serve-http.ts';
|
||||
import { HEALTH_TIMEOUT_MS, probeHealth, probeLiveness } from '../src/commands/serve-http.ts';
|
||||
import type { BrainEngine } from '../src/core/engine.ts';
|
||||
import type { SqlQuery } from '../src/core/oauth-provider.ts';
|
||||
|
||||
/**
|
||||
* Minimal mock engine: only `getStats()` is exercised by probeHealth.
|
||||
@@ -22,6 +28,17 @@ function makeMockEngine(getStats: () => Promise<unknown>): BrainEngine {
|
||||
return { getStats } as unknown as BrainEngine;
|
||||
}
|
||||
|
||||
/**
|
||||
* Minimal mock sql tag: probeLiveness only awaits the result of `sql\`SELECT 1\``
|
||||
* — the tag function's return value is what's raced, success/throw is what
|
||||
* matters. We ignore the template strings and simulate a connection by calling
|
||||
* the supplied factory.
|
||||
*/
|
||||
function makeMockSql(fn: () => Promise<unknown>): SqlQuery {
|
||||
const tag: any = (_strings: TemplateStringsArray, ..._values: unknown[]) => fn();
|
||||
return tag as SqlQuery;
|
||||
}
|
||||
|
||||
describe('HEALTH_TIMEOUT_MS', () => {
|
||||
test('exported as 3000 (Fly.io headroom over the 5s default)', () => {
|
||||
expect(HEALTH_TIMEOUT_MS).toBe(3000);
|
||||
@@ -70,3 +87,66 @@ describe('probeHealth', () => {
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('probeLiveness (v0.28.10)', () => {
|
||||
test('happy path: returns 200 + status:ok with NO engine-stats fields', async () => {
|
||||
const sql = makeMockSql(async () => [{ '?column?': 1 }]);
|
||||
const result = await probeLiveness(sql, 'postgres', '0.28.10', 100);
|
||||
expect(result.ok).toBe(true);
|
||||
expect(result.status).toBe(200);
|
||||
if (result.ok) {
|
||||
expect(result.body.status).toBe('ok');
|
||||
expect(result.body.version).toBe('0.28.10');
|
||||
expect(result.body.engine).toBe('postgres');
|
||||
// Regression: the lightweight body must NOT spread getStats() fields.
|
||||
// The original PR's pre-refactor /health leaked page_count etc.;
|
||||
// tightening this assertion is the iron-rule regression test.
|
||||
expect(Object.keys(result.body).sort()).toEqual(['engine', 'status', 'version']);
|
||||
expect((result.body as Record<string, unknown>).page_count).toBeUndefined();
|
||||
expect((result.body as Record<string, unknown>).chunk_count).toBeUndefined();
|
||||
}
|
||||
});
|
||||
|
||||
test('timeout path: sql hangs → 503 with health_timeout description within 1s', async () => {
|
||||
const sql = makeMockSql(() => new Promise(() => { /* never resolves */ }));
|
||||
const start = Date.now();
|
||||
const result = await probeLiveness(sql, 'postgres', '0.28.10', 100);
|
||||
const elapsed = Date.now() - start;
|
||||
expect(elapsed).toBeLessThan(1000);
|
||||
expect(result.ok).toBe(false);
|
||||
expect(result.status).toBe(503);
|
||||
if (!result.ok) {
|
||||
expect(result.body.error).toBe('service_unavailable');
|
||||
expect(result.body.error_description).toBe(
|
||||
'Health check timed out (database pool may be saturated)',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('db-error path: sql throws → 503 with database_failed description', async () => {
|
||||
const sql = makeMockSql(() => Promise.reject(new Error('ECONNREFUSED')));
|
||||
const result = await probeLiveness(sql, 'postgres', '0.28.10', 100);
|
||||
expect(result.ok).toBe(false);
|
||||
expect(result.status).toBe(503);
|
||||
if (!result.ok) {
|
||||
expect(result.body.error).toBe('service_unavailable');
|
||||
expect(result.body.error_description).toBe('Database connection failed');
|
||||
}
|
||||
});
|
||||
|
||||
test('timer-cleanup: 100 fast successful probes do not leak pending timers', async () => {
|
||||
const sql = makeMockSql(async () => [{ '?column?': 1 }]);
|
||||
// Snapshot active handles before; same after. If the finally-block
|
||||
// clearTimeout regressed, every probe would leak a 100ms-pending timer.
|
||||
const beforeHandles = (process as any)._getActiveHandles?.()?.length ?? 0;
|
||||
await Promise.all(
|
||||
Array.from({ length: 100 }, () => probeLiveness(sql, 'postgres', '0.28.10', 100)),
|
||||
);
|
||||
// Allow microtask + process tick drain to let any leaked timers settle.
|
||||
await new Promise(r => setImmediate(r));
|
||||
const afterHandles = (process as any)._getActiveHandles?.()?.length ?? 0;
|
||||
// Loose bound: bun's internal handles can drift by a small amount across
|
||||
// many fetches; we only care that we don't ramp by ~100 leaked timers.
|
||||
expect(afterHandles - beforeHandles).toBeLessThan(20);
|
||||
});
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user