Files
gbrain/recipes/agent-voice/code/lib/audio-convert.mjs
T
e9fa51d46e v0.40.0.0 feat: agent-voice (Mars + Venus) + copy-into-host-repo skillpack paradigm (#1128)
* feat: agent-voice reference skillpack (Mars + Venus) + copy-into-host-repo install paradigm

Ships a new skillpack paradigm: gbrain holds the REFERENCE content;
`gbrain integrations install agent-voice --target <repo>` COPIES it into
the operator's host agent repo where it becomes user-owned and mutable.
Future refresh is diff-and-propose against per-file SHA-256 hashes from
.gbrain-source.json, not blind overwrite.

What ships:
- recipes/agent-voice.md entrypoint + recipes/agent-voice/ bundle
- Two voice personas (Mars dual-mode SOLO/DEMO, Venus executive assistant)
  with PII / private-agent-name / hardcoded-path scrubbed out
- WebRTC-first browser client (call.html) with ?test=1 gated instrumentation
  and Web Audio API tee -> MediaRecorder capture for E2E roundtrip testing
- Read-only tool router (D14-A allow-list: search, query, get_page,
  list_pages, find_experts, get_recent_salience, get_recent_transcripts,
  read_article). Write ops permanently denylisted; opt-in via local override
- Persona-aware prompt builder with identity-first composition + Unicode
  sanitization for OpenAI Realtime API safety
- Upstream-error classifier (HTTP 429/500/503 -> soft-fail, plumbing -> hard)
- Three SKILL.md skills (voice-persona-mars, voice-persona-venus,
  voice-post-call) with routing-eval.jsonl fixtures
- 99 host-side tests (vitest-compatible, runs in bun) covering registry,
  prompt-shape privacy guards, tool allow-list, upstream classifier
- install/manifest.json + refresh-algorithm.md + post-install-hint.md

Privacy infrastructure:
- scripts/check-no-pii-in-agent-voice.sh wired into bun run verify
  Shape regex (phone/email/SSN/JWT/bearer/credit-card) + path patterns +
  $AGENT_VOICE_PII_BLOCKLIST env-driven name blocklist
- scripts/import-from-upstream.sh + scripts/upstream-scrub-table.txt
  Deterministic refresh from upstream voice-agent source. Placeholder-
  driven (envsubst-expanded at run time) so no private names land in
  checked-in files
- recipes/agent-voice/code/lib/personas/private-name-blocklist.json
  Single source of truth for the regex contract (shape categories +
  path patterns + env-var contract for operator-specific names)

src/ surface:
- src/commands/integrations.ts gains `install <recipe-id>` subcommand
  with install_kind: 'local-managed' | 'copy-into-host-repo' discriminator.
  Path-traversal hardening (rejects '..', absolute paths, symlink escapes).
  Refuses target == gbrain itself, missing .git, existing files (without
  --overwrite). Writes .gbrain-source.json with per-file SHA-256. Appends
  resolver rows to host repo's RESOLVER.md or AGENTS.md.
- test/integrations-install.test.ts: 11 cases (happy path, manifest shape,
  no upstream_repo field per D11-A, resolver appending, file modes,
  refusal cases, dry-run)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: bump version and changelog (v0.36.0.0)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(privacy): scrub literal private agent names from prompt-shape tests + guard script

The prompt-shape tests carried regex patterns naming the literal banned terms
(Garry/Steph/Garrison/Solomon/Herbert/Wintermute) inline. CLAUDE.md's
"never use Wintermute in any public artifact" applies to test source files
too. Master's check-privacy.sh correctly caught this.

Replaced with env-driven check that reads AGENT_VOICE_PII_BLOCKLIST (the
single source of truth from private-name-blocklist.json). Same enforcement
guarantee via the env var, zero literal names in shipped source.

Also scrubbed the literal /data/.openclaw/ from the guard script's comment
and the literal 'tell_wintermute' from the venus write-tools test.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat: ship all v0.36.0.0 deferred items in this PR (E2E + evals + pipeline + refresh + multilingual + twilio deprecation)

Closes the "deferred to follow-up" section of the v0.36.0.0 CHANGELOG.

E2E tests + harness (env-gated):
- tests/e2e/voice-roundtrip.test.mjs — spawns server, drives puppeteer + fake-audio, three-tier assertions (CONNECTION hard, NON-SILENT hard, SEMANTIC soft via Whisper + LLM judge). Upstream errors (429/500/503, WS 1011/1013) soft-fail via lib/upstream-classifier.mjs.
- tests/e2e/voice-full-flow.test.mjs — wraps openclaw doing the install, then runs the roundtrip. Friction-discovery flavor, NOT a ship gate.
- tests/e2e/lib/browser-audio.mjs — puppeteer + fake-audio harness; reads window._gbrainTest namespace; PCM RMS-variance helper.
- tests/e2e/lib/whisper-judge.mjs — Whisper transcription + LLM-judge for SEMANTIC tier.
- tests/e2e/audio-fixtures/utterance-{add,joke,brain-query}.wav — 16kHz mono WAV via `say` + ffmpeg, committed for reproducibility.
- test/fixtures/claw-test-scenarios/voice-agent-install/{BRIEF.md, scenario.json, expected.json} — labeled BENCHMARK_FRICTION, blocks_ship=false.

LLM-judge persona evals + synthetic canonical baselines:
- tests/evals/judge.mjs — gateway-routed 3-model (Claude + GPT + Gemini) harness with 4-strategy JSON repair + 2/3-quorum aggregation (per the v0.27.x cross-modal pattern). Pass criterion: every axis mean ≥7 AND no model <5.
- tests/evals/fixtures/{mars-solo,mars-demo,venus,persona-routing,mars-multilingual}.jsonl — 5 fixture sets covering all axes.
- tests/evals/{mars-eval,venus-eval,mars-multilingual-eval,persona-routing-eval}.mjs — per-axis drivers.
- tests/evals/baseline-runs/canonical/*.json — agent-authored synthetic exemplars (PII-impossible by construction; demonstrate expected pass shape; never overwrite with live model output).
- tests/evals/baseline-runs/.gitignore — live receipts excluded.

DIY pipeline (Option B):
- code/pipeline.mjs — streaming STT (Deepgram nova-2) + LLM (Claude Sonnet 4.6 streaming SSE with sentence-boundary TTS dispatch) + TTS (Cartesia primary, OpenAI TTS fallback). 20-turn history cap, exponential-backoff reconnects, 25s keepalives, VAD presets (quiet/normal/noisy/very_noisy), barge-in via STT speechStart → LLM interrupt. Modular adapters for swapping providers.

--refresh mode (D3-A diff-and-propose):
- src/commands/integrations.ts: refreshRecipeIntoHostRepo() + classifyForRefresh() implementing the five states from refresh-algorithm.md (unchanged-identical, unchanged-stale, locally-modified, source-deleted, host-deleted, new-in-manifest). Transaction journal at .gbrain-source.refresh.log. Default policy: preserve operator's local edits (keep-mine); --auto take-theirs to overwrite; --dry-run for preview.
- test/integrations-install.test.ts: 7 new test cases pinning each classification state + default-preserve behavior + take-theirs overwrite + transaction journal + refusal on uninstalled target.

Mars multilingual restore:
- code/lib/personas/mars.mjs: explicit cross-lingual rule (Mandarin, Spanish, French, Japanese, Korean default to English but follow the speaker). Voice (Orus) supports the languages natively.
- tests/unit/mars-prompt-shape.test.mjs: assertion flipped from "MUST NOT claim multilingual" to "declares cross-lingual capability with English bias."
- tests/evals/fixtures/mars-multilingual.jsonl: 5 fixtures across Mandarin/Spanish/Japanese/French + explicit switch-back, pinned by mars-multilingual-eval.mjs.

Twilio recipe deprecation:
- recipes/twilio-voice-brain.md: deprecation banner pointing at agent-voice.md. Frontmatter version bumped to 0.8.2. Will be removed in v0.37.

Verify: bun run verify clean, 6736+ unit tests pass, 18/18 install+refresh tests pass, 96/98 host-side persona/tool/classifier tests pass (2 skipped env-gated).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: update CHANGELOG — all v0.36.0.0 deferred items now shipped in this PR

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: rebump version v0.36.0.0 → v0.37.0.0

Captures the wave-1 + wave-2 scope at the v0.37 slot. The bump reflects
the size of what this PR ships: copy-into-host-repo install paradigm
(new install_kind discriminator + new install/refresh subcommand) +
Mars/Venus voice agent reference + 5,500+ LOC of vendored scrubbed
code + 4 LLM-judge eval suites + 2 env-gated E2E test suites + DIY
Option B pipeline + 18-case install subcommand test coverage. A minor
bump felt too small.

Side fix: privacy guard caught two stale literal "wintermute" and
"/data/.openclaw/" references in wave-2 files
(voice-full-flow.test.mjs comment, expected.json blocklist payload).
Both replaced with env-driven references to $AGENT_VOICE_PII_BLOCKLIST
matching the D15-A pattern from the original review.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: rebump v0.37.0.0 → v0.40.0.0

Jumps past the v0.37/v0.38/v0.39 slots master might claim in subsequent
PRs. The wave's scope (copy-into-host-repo skillpack paradigm + agent-voice
+ install/refresh + LLM-judge evals + DIY pipeline + Mars multilingual)
justifies a larger version arithmetic step.

Files bumped:
- VERSION 0.37.0.0 → 0.40.0.0
- package.json 0.37.0.0 → 0.40.0.0
- CHANGELOG.md header + "To take advantage of v0.40.0.0" block
- recipes/twilio-voice-brain.md deprecation banner (now "removed in v0.41")
- recipes/agent-voice/tests/evals/mars-multilingual-eval.mjs comment

Left alone: master's pre-existing "v0.37+" roadmap labels in src/core/calibration/*,
src/core/cycle/*, DESIGN.md, CLAUDE.md, etc. Those are master's author-intent
references to "the next planned release" relative to master's frame at the time —
rewriting them just to keep numbering consistent would overreach.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-22 21:54:13 -07:00

217 lines
6.8 KiB
JavaScript

/**
* audio-convert.mjs — µ-law ↔ PCM conversion for Twilio ↔ Gemini bridge
*
* Twilio sends: µ-law 8kHz mono base64 (20ms chunks = 160 bytes)
* Gemini wants: PCM 16-bit 16kHz mono base64 (buffered ~300ms)
* Gemini sends: PCM 16-bit 24kHz mono base64 (variable chunks)
* Twilio wants: µ-law 8kHz mono base64
*/
// ── µ-law decode table (ITU-T G.711) ─────────────────────
const ULAW_DECODE = new Int16Array(256);
for (let i = 0; i < 256; i++) {
let u = ~i & 0xFF;
let sign = u & 0x80;
let exponent = (u >> 4) & 0x07;
let mantissa = u & 0x0F;
let sample = (mantissa << 3) + 0x84;
sample <<= exponent;
sample -= 0x84;
ULAW_DECODE[i] = sign ? -sample : sample;
}
// ── PCM → µ-law encode ───────────────────────────────────
const ULAW_MAX = 0x1FFF;
const ULAW_BIAS = 0x84;
function pcmToUlaw(sample) {
let sign = 0;
if (sample < 0) { sign = 0x80; sample = -sample; }
if (sample > ULAW_MAX) sample = ULAW_MAX;
sample += ULAW_BIAS;
let exponent = 7;
for (let mask = 0x4000; (sample & mask) === 0 && exponent > 0; exponent--, mask >>= 1) {}
let mantissa = (sample >> (exponent + 3)) & 0x0F;
return (~(sign | (exponent << 4) | mantissa)) & 0xFF;
}
// ── Stateless converters (for unit tests + simple cases) ──
/**
* Decode µ-law bytes to PCM 16-bit samples (no resampling)
*/
export function ulawToPcm8k(ulawBuf) {
const pcm = new Int16Array(ulawBuf.length);
for (let i = 0; i < ulawBuf.length; i++) {
pcm[i] = ULAW_DECODE[ulawBuf[i]];
}
return pcm;
}
/**
* Simple stateless: µ-law 8kHz base64 → PCM 16kHz base64
* Uses linear interpolation. OK for testing, not ideal for production.
*/
export function ulawToGemini(base64Ulaw) {
const ulawBuf = Buffer.from(base64Ulaw, 'base64');
if (ulawBuf.length === 0) return '';
const pcm8k = ulawToPcm8k(ulawBuf);
const pcm16k = new Int16Array(pcm8k.length * 2);
for (let i = 0; i < pcm8k.length; i++) {
pcm16k[i * 2] = pcm8k[i];
pcm16k[i * 2 + 1] = i < pcm8k.length - 1 ? (pcm8k[i] + pcm8k[i + 1]) >> 1 : pcm8k[i];
}
return Buffer.from(pcm16k.buffer).toString('base64');
}
/**
* PCM 24kHz base64 → µ-law 8kHz base64 (downsample 3:1)
*/
export function geminiToUlaw(base64Pcm) {
const pcmBuf = Buffer.from(base64Pcm, 'base64');
const pcm24k = new Int16Array(pcmBuf.buffer, pcmBuf.byteOffset, pcmBuf.length / 2);
const numOut = Math.floor(pcm24k.length / 3);
const ulawBuf = Buffer.alloc(numOut);
for (let i = 0; i < numOut; i++) {
ulawBuf[i] = pcmToUlaw(pcm24k[i * 3]);
}
return ulawBuf.toString('base64');
}
/**
* PCM 16kHz base64 → µ-law 8kHz base64 (downsample 2:1)
*/
export function gemini16kToUlaw(base64Pcm) {
const pcmBuf = Buffer.from(base64Pcm, 'base64');
const pcm16k = new Int16Array(pcmBuf.buffer, pcmBuf.byteOffset, pcmBuf.length / 2);
const numOut = Math.floor(pcm16k.length / 2);
const ulawBuf = Buffer.alloc(numOut);
for (let i = 0; i < numOut; i++) {
ulawBuf[i] = pcmToUlaw(pcm16k[i * 2]);
}
return ulawBuf.toString('base64');
}
// ── Stateful resampler for production use ─────────────────
// Proper linear interpolation with state across chunk boundaries
/**
* Create a stateful 8kHz→16kHz upsampler.
* Tracks the last sample across chunks for smooth interpolation.
*/
export function createUpsampler() {
let lastSample = 0;
return function upsample(pcm8k) {
const pcm16k = new Int16Array(pcm8k.length * 2);
for (let i = 0; i < pcm8k.length; i++) {
const prev = i === 0 ? lastSample : pcm8k[i - 1];
pcm16k[i * 2] = (prev + pcm8k[i]) >> 1; // Interpolated sample
pcm16k[i * 2 + 1] = pcm8k[i]; // Original sample
}
lastSample = pcm8k[pcm8k.length - 1] || 0;
return pcm16k;
};
}
/**
* Create a stateful 24kHz→8kHz downsampler.
* Averages 3 samples for each output (low-pass filter).
*/
export function createDownsampler24to8() {
let remainder = new Int16Array(0);
return function downsample(pcm24k) {
// Prepend any remainder from last chunk
let input;
if (remainder.length > 0) {
input = new Int16Array(remainder.length + pcm24k.length);
input.set(remainder);
input.set(pcm24k, remainder.length);
} else {
input = pcm24k;
}
const numOut = Math.floor(input.length / 3);
const leftover = input.length - numOut * 3;
const out = new Int16Array(numOut);
for (let i = 0; i < numOut; i++) {
// Average 3 samples (simple low-pass)
const idx = i * 3;
out[i] = Math.round((input[idx] + input[idx + 1] + input[idx + 2]) / 3);
}
// Save leftover samples for next chunk
remainder = leftover > 0 ? input.slice(input.length - leftover) : new Int16Array(0);
return out;
};
}
/**
* Create a buffered audio processor for Twilio→Gemini.
* Buffers µ-law chunks and flushes PCM 16kHz every ~300ms.
*
* @param {Function} onFlush - (base64Pcm16k) => void
* @param {number} flushMs - buffer duration before flushing (default 200ms)
*/
export function createTwilioToGeminiProcessor(onFlush, flushMs = 200) {
const upsample = createUpsampler();
// 16kHz * 2 bytes * flushMs/1000 = buffer threshold
const FLUSH_BYTES = Math.floor(16000 * 2 * flushMs / 1000);
let pcmBuffer = [];
let totalBytes = 0;
return {
/** Process a base64 µ-law chunk from Twilio */
push(base64Ulaw) {
const ulawBuf = Buffer.from(base64Ulaw, 'base64');
const pcm8k = ulawToPcm8k(ulawBuf);
const pcm16k = upsample(pcm8k);
pcmBuffer.push(Buffer.from(pcm16k.buffer));
totalBytes += pcm16k.length * 2;
if (totalBytes >= FLUSH_BYTES) {
this.flush();
}
},
/** Force flush any buffered audio */
flush() {
if (pcmBuffer.length === 0) return;
const combined = Buffer.concat(pcmBuffer);
pcmBuffer = [];
totalBytes = 0;
onFlush(combined.toString('base64'));
},
/** Get current buffer size in bytes */
get bufferedBytes() { return totalBytes; },
};
}
/**
* Create a Gemini→Twilio audio processor.
* Converts PCM 24kHz to µ-law 8kHz with proper downsampling.
*/
export function createGeminiToTwilioProcessor() {
const downsample = createDownsampler24to8();
return {
/** Process base64 PCM 24kHz from Gemini → base64 µ-law 8kHz for Twilio */
process(base64Pcm) {
const pcmBuf = Buffer.from(base64Pcm, 'base64');
const pcm24k = new Int16Array(pcmBuf.buffer, pcmBuf.byteOffset, pcmBuf.length / 2);
const pcm8k = downsample(pcm24k);
const ulawBuf = Buffer.alloc(pcm8k.length);
for (let i = 0; i < pcm8k.length; i++) {
ulawBuf[i] = pcmToUlaw(pcm8k[i]);
}
return ulawBuf.toString('base64');
}
};
}