* feat(voice): add dedicated voice assistance module for STT/TTS
Extracts speech-to-text (whisper.cpp) and text-to-speech (piper) into a
dedicated `src/openhuman/voice/` domain module with its own RPC namespace
(`openhuman.voice_*`). Adds proactive availability checking via
`voice_status` so the UI can show clear errors when binaries/models are
missing instead of failing silently at transcription time.
- New module: voice/types.rs, voice/ops.rs, voice/schemas.rs, voice/mod.rs
- 4 RPC endpoints: voice_status, voice_transcribe, voice_transcribe_bytes, voice_tts
- 21 unit tests + 1 integration test (json_rpc_e2e)
- Frontend updated to use voice_* endpoints with status check on mode switch
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: fix cargo fmt in voice/ops.rs
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test(e2e): add voice mode integration spec
Tests switching to voice input mode, verifying status check fires,
recording button renders, and switching back to text mode restores
text input. Also checks reply mode toggle visibility.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: remove unused waitForText import in voice-mode e2e spec
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(voice): in-process whisper engine and LLM post-processing
- Add whisper-rs (0.16) for in-process whisper.cpp inference, eliminating
cold-start latency from subprocess-per-call (~1-3s) to warm inference
(~50ms). Model is loaded once during bootstrap and reused across calls.
Falls back to whisper-cli subprocess if in-process loading fails.
- Add LLM post-processing layer that passes raw transcription through
Ollama to fix grammar, punctuation, and filler words. Accepts optional
conversation context to disambiguate names and technical terms.
Gracefully degrades to raw whisper output if Ollama is unavailable.
- Update voice RPC endpoints with new optional params (context,
skip_cleanup) and return both cleaned text and raw_text.
- Update frontend to pass conversation history as context for voice
transcription cleanup, and update TypeScript interfaces to match.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: apply cargo fmt formatting fixes
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(build): make whisper-rs optional behind `whisper` feature flag
The whisper-rs crate requires cmake to compile whisper.cpp from source,
which is not available in the CI environment. Move it behind an optional
cargo feature so CI builds succeed without cmake.
The whisper_engine module now compiles as a no-op stub when the feature
is disabled, returning "whisper feature not compiled in" errors. Desktop
builds can opt in with `--features whisper`.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: cargo fmt whisper_engine.rs
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): make whisper-rs mandatory and install cmake in CI
Revert whisper-rs from optional to mandatory dependency. Add cmake
installation to all CI workflows (build, typecheck, test, release) and
the CI Docker image so whisper-rs can compile whisper.cpp from source.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: cargo fmt whisper_engine.rs
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address code review findings across voice module
- whisper_engine: validate WAV sample rate (must be 16kHz) and channel
count (1 or 2) before feeding audio to whisper
- speech: offload load_engine and transcribe_in_process to
tokio::task::spawn_blocking to avoid blocking the Tokio runtime
- ops: use RAII guard for WHISPER_BIN env var in test to prevent races
and ensure restore on panic; log temp file cleanup failures instead
of silently ignoring; sanitize paths in debug logs to basenames only
- postprocess: add test for disabled cleanup config returning raw text
- voice-mode.spec: assert failure when neither voice CTA nor
unavailable message appears; make reply mode test runnable in
isolation with auth/nav setup
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: apply cargo fmt
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Split monolithic inline bash from release.yml and release-packages.yml
into standalone scripts under scripts/release/ for easier debugging
and manual execution.
New scripts:
- bump-version.js: version bumping across package.json/tauri/Cargo
- stage-sidecar.sh: stage + verify sidecar binary for Tauri bundler
- sign-and-notarize-macos.sh: macOS code signing and notarization
- repackage-dmg.sh: re-create and notarize DMG post-signing
- upload-macos-artifacts.sh: re-upload notarized artifacts to release
- package-cli-tarball.sh: package CLI binary into release tarball
- build-linux-arm64.sh: build Linux arm64 CLI tarball
- update-homebrew.sh: render and commit Homebrew formula to tap
- build-apt-packages.sh: build .deb packages and apt repository
- publish-npm.sh: stamp version and publish npm package
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>