Commit Graph
100 Commits
Author SHA1 Message Date
07140e9c0b feat(mcp): complete the MCP Servers page — curated catalog, UX, lifecycle & E2E (#4272) (#4300)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: M3gA-Mind <megamind@mahadao.com>
2026-06-30 22:37:41 +05:30
5319479182 fix(composio): surface actionable re-auth CTA when triggers fail with SESSION_EXPIRED (#4296)
Co-authored-by: M3gA-Mind <megamind@mahadao.com>
2026-06-30 22:09:36 +05:30
2ff765b60a fix(mcp): spawn npx/uvx MCP servers with the user's real shell PATH (#4295)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 19:53:49 +05:30
3ed93c1a38 feat(skills): activation persistence, latency instrumentation, registry browse cache + pagination (#4288)
Co-authored-by: M3gA-Mind <megamind@mahadao.com>
2026-06-30 19:53:41 +05:30
08a9bff841 fix(voice): in-process Whisper STT works without an external binary on macOS (#3425) (#3915)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: M3gA-Mind <megamind@mahadao.com>
2026-06-30 17:17:28 +05:30
01c35414b5 MCP catalog — namespace-correct curation, data-driven auth, Official badge, Smithery opt-in (#4120)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 13:19:04 -07:00
5756271335 feat(meet): expandable recent-call detail with transcript + summary (#4113)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 09:31:42 -07:00
2b90955482 feat(chat): install skills and connect integrations inline (#4062)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 12:53:18 -07:00
YellowSnnowmannandGitHub 8b6b5fdbe8 fix(skills): surface registry-installed skills in Installed tab + raise catalog browse timeout (#3954) (#3987) 2026-06-23 11:50:36 -07:00
6ce4f52972 fix(channels): repair Discord & Telegram messaging end-to-end (#3712, #3763) (#3794)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-22 12:17:21 -07:00
56f9dac91b feat: add in-app feedback board (#3834)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-22 10:11:34 -07:00
1f00ab79b1 feat(rewards): Discord connect + disconnect/re-link on the Rewards page (#3766)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 11:38:08 +05:30
3cf3c5443b feat: wire up Connect Discord OAuth on the Rewards page (#3748)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 12:50:00 +05:30
YellowSnnowmannandGitHub decfc5b89b feat(agent_meetings): post-call meeting summary + context label (#3709) 2026-06-17 17:25:35 +05:30
YellowSnnowmannandGitHub 5265525761 feat(meetings): calendar-triggered auto-join prompt + Meeting Assistant settings UI (#3721) 2026-06-17 17:23:57 +05:30
YellowSnnowmannandGitHub eaad68ecaa fix(meetings): populate Recent calls for backend meet flow (owner, duration, participants) (#3710) 2026-06-17 13:31:30 +05:30
1e0985f576 [merge 3/5] test: agent harness behaviors JSON-RPC E2E suite (#3471) (#3667)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-15 19:41:43 -07:00
5b80f25cd4 [merge 1/5] feat(approval): debug-only OPENHUMAN_APPROVAL_TTL_SECS TTL override (#3471) (#3665)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-15 18:26:54 -07:00
93c7b3d6a3 fix(agent): parse Claude-native <invoke name> attribute tool calls (#3622)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-15 17:20:08 -07:00
7d71b06711 [merge 5/5] docs: agent harness E2E coverage matrix + plan (#3471) (#3668)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-15 17:01:37 -07:00
a54854af03 [merge 4/5] test(e2e): agent harness behaviors browser spec (#3471) (#3669)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-15 17:01:32 -07:00
6363790b1f [merge 2/5] feat(channels/web): per-runtime approval-surface bridge for E2E tests (#3471) (#3666)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-15 17:00:37 -07:00
c5087d83da feat(meet): in-call meeting agency (active replies, streaming, approvals) (#3677)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-15 11:34:01 -07:00
YellowSnnowmannandGitHub 25cbeb6376 test(e2e): restore chat-harness-subagent strong assertions + diagnostics (#3584) 2026-06-11 18:08:45 -07:00
8e2b1cf2cc feat(mcp): auto-reconnect installed servers via a background supervisor (#3312) (#3571)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 18:08:12 -07:00
4c79343a84 feat(health): granular /health liveness — a degraded component no longer 503s the container (#3312) (#3572)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 13:48:15 -07:00
188f0fb7eb fix(meet): block join modal until backend admits + hide wake-phrase field (#3597)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-11 19:04:29 +05:30
YellowSnnowmannandGitHub 3f0d9f6ef8 feat(meetings): respondTo required + wake-phrase filtering UI (#3555) 2026-06-09 13:31:37 -07:00
YellowSnnowmannandGitHub 2f324812e6 fix(test): stabilize flaky mcp-tab-flow E2E tests (#3480 follow-up) (#3515) 2026-06-08 10:08:37 -07:00
YellowSnnowmannandGitHub 0d9cfde6e6 fix(observability): classify ollama 403 subscription-gate as config-rejection (TAURI-RUST-4XK) (#3502) 2026-06-08 07:20:56 -07:00
b3be665937 feat(agent_meetings): wire mascotId through full stack + fix tests (#3363)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-05 10:15:26 -04:00
YellowSnnowmannandGitHub 77cfea2b61 fix(composio): show connection ID in picker when no identity cached (#3405) 2026-06-05 10:10:29 -04:00
YellowSnnowmannandGitHub 05785981d8 fix(composio): enrich connection picker with cached account identity (#3356) (#3400) 2026-06-05 16:30:13 +05:30
YellowSnnowmannandGitHub 505b04783b fix(intelligence): deduplicate and label connections in integration source picker (#3361) 2026-06-04 19:12:23 +05:30
YellowSnnowmannandGitHub 68d0da5311 fix(composio): defer on_connection_created sync until onboarding completes (#3283) 2026-06-03 15:39:14 +05:30
YellowSnnowmannandGitHub a83a2a1439 fix(onboarding,billing): surface completeAndExit errors and suppress near-limit banner for custom providers (#3278) 2026-06-03 15:38:51 +05:30
75ac1f24fa fix(routines): deduplicate morning_briefing jobs on seed boot (#3126) (#3141)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-01 22:56:32 -07:00
d666ce61d7 fix(email): add date_local field with host timezone to email reshaper (#3128) (#3143)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-01 22:03:34 -07:00
e1a8f16778 fix: Homebrew install error and Discord channel listing (#3085) (#3144)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-01 22:03:18 -07:00
6ba2ba2717 fix(cron): make cron_run non-blocking, enqueue async execution (#3127) (#3142)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-06-01 21:06:42 -07:00
fb8068e92e feat(cron): complete smart scheduling flow from chat + polish cron management UI (#2943)
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-30 08:25:50 -07:00
1144274844 fix(observability): demote "Config loading timed out" out of Sentry (#2920)
Co-authored-by: Cyrus Gray <144336577+graycyrus@users.noreply.github.com>
2026-05-29 19:47:53 +05:30
YellowSnnowmannandGitHub 9f3e161670 fix(agent): replay reasoning_content across native tool-call turns to prevent DeepSeek 400s (#2918) 2026-05-29 17:12:48 +05:30
YellowSnnowmannandGitHub 823b52a8e9 fix(observability): demote WhatsApp ingest SQLite lock noise (#2917) 2026-05-29 14:39:45 +05:30
YellowSnnowmannandGitHub a211cac9fd fix(providers): dedup tool specs at wire boundary to prevent 400 "Tool names must be unique" (#2846) 2026-05-29 05:01:58 +05:30
71d9ecb6bc fix(socket): recover from stale disconnected socket to prevent permanent "Connecting..." (#2487)
## Summary

- Fixed a reconnect edge case where socketService.connect() could get stuck when a stale disconnected socket instance existed for the same auth token.
- Prevented false-positive connecting UI state by clearing stale disconnected socket references before async reconnect guards run.
- Preserved existing safety behavior for active sockets (connected) and in-flight sockets (!disconnected) to avoid duplicate connections.
- Added a regression unit test to verify same-token reconnect creates a fresh socket instead of silently no-oping.

## Problem

- Users could be stuck on Connecting... and unable to chat after a disconnect/reconnect cycle.
- Root cause: reconnect logic set state to connecting, but returned early because this.socket was still non-null (stale/disconnected), so no new socket was created and no connect/connect_error transition fired.
- This left connection state stranded and blocked chat flows.

## Solution

- In socketService.connectAsync, when token is unchanged and this.socket.disconnected === true, explicitly clear stale runtime references (this.socket, this.mcpTransport) before continuing.
- Keep existing early returns for:
  - same-token + already connected
  - same-token + currently connecting (!disconnected)
- Added test coverage for the stale-socket scenario: second same-token connect() now creates a new socket instance (verifies io(...) called twice).

## Submission Checklist

- If a section does not apply to this change, mark the item as N/A with a one-line reason. Do not delete items.
- Tests added or updated (happy path + at least one failure / edge case) per Testing Strategy
- Diff coverage ≥ 80% — changed lines (Vitest + cargo-llvm-cov merged via diff-cover) meet the gate enforced by `.github/workflows/coverage.yml`. Run pnpm test:coverage and pnpm test:rust locally; PRs below 80% on changed lines will not merge.
- Coverage matrix updated — added/removed/renamed feature rows in `docs/TEST-COVERAGE-MATRIX.md` reflect this change (or N/A: behaviour-only change)
- All affected feature IDs from the matrix are listed in the PR description under ## Related
- No new external network dependencies introduced (mock backend used per Testing Strategy)
- Manual smoke checklist updated if this touches release-cut surfaces (`docs/RELEASE-MANUAL-SMOKE.md`)
- Linked issue closed via Closes #NNN in the ## Related section

## Impact

- Runtime/platform impact: frontend connection management path (app/src/services/socketService.ts) affecting desktop/web UI behavior where this service is used.
- Compatibility: no API/schema changes.
- Performance: negligible; only additional stale-reference cleanup in reconnect edge case.
- Security: no new credential surface or transport changes.


Co-authored-by: M3gA-Mind <megamind@mahadao.com>
2026-05-29 04:49:44 +05:30
YellowSnnowmannandGitHub a6fd156131 fix(inference): default empty model to reasoning-v1 on OpenHuman backend (#2837) 2026-05-29 04:12:29 +05:30
cc088d7e70 fix(providers): include body snippet on list_models JSON parse failure (#2838)
## Summary

- Replace `response.json()` with read-as-text + `serde_json::from_str` in `list_configured_models_from_config` so the response body is preserved when JSON decoding fails.
- Append a sanitized + truncated body snippet to the `[providers][list_models] failed to parse JSON` error so the failure is diagnosable from the log/Sentry line alone.
- Add three async unit tests: HTML body returns a diagnostic snippet, empty body still surfaces the parse error, and a valid `/models` response still lists models (regression guard for the new text-then-parse path).

## Problem

Sentry issue **TAURI-RUST-12** — `[providers][list_models] failed to parse JSON: error decoding response body` — 376 events / 14d on the `tauri-rust` project.

`response.json()` in [src/openhuman/inference/provider/ops.rs:125](src/openhuman/inference/provider/ops.rs:125) (pre-change) consumes the body in the process of decoding it. When the decode fails — typically because the server returned HTML from a captive portal / corporate proxy login page, an upstream load-balancer 502 served as HTML with `200 OK`, or a wrong-path endpoint returning a non-JSON response — the body is gone by the time we format the error, so Sentry receives `error decoding response body` with no payload context.

We can't fix this server-side. We *can* stop discarding the diagnostic information at the call site so users and devs can identify the real cause from the error string instead of guessing.

## Solution

`src/openhuman/inference/provider/ops.rs`:

- After the `status.is_success()` check, call `response.text().await` instead of `response.json()`. The text path returns the raw body verbatim, which we can then both parse *and* embed in the diagnostic message.
- `serde_json::from_str(&raw_body)` reproduces the previous decode behaviour exactly — same JSON parser, same `serde_json::Error` shape. On failure, the closure sanitizes the body via the existing `sanitize_api_error` helper and truncates it through the existing `crate::openhuman::util::truncate_with_ellipsis(_, 300)` before appending it as `(body: …)`.
- Adds an explicit error for `response.text()` failure (`failed to read response body`) — a transport-layer concern distinct from JSON parsing.

**Design choices**

- Re-use the existing `sanitize_api_error` (strips ANSI / control chars, caps at `MAX_API_ERROR_CHARS`) and `truncate_with_ellipsis` helpers — same sanitization the non-2xx branch already applies a few lines above. No new redaction policy.
- Keep the canonical error prefix `[providers][list_models] failed to parse JSON:` so any existing log greps / Sentry classifiers continue to match.
- 300-character snippet cap matches the existing non-2xx branch's `truncated` cap and the codebase convention for "include enough for triage, not enough to flood logs."
- No change to the JSON parser, no change to what counts as a valid `/models` response, and no change to error semantics — the new branch returns `Err(...)` in exactly the same shape and code path as before. Callers see one extra clause appended to the message string.
- Body is only read on the success path (`status.is_success()`). The non-2xx branch already had its own `response.text()` + sanitize chain, untouched.

## Submission Checklist

- Tests added or updated (happy path + at least one failure / edge case) per [Testing Strategy](../gitbooks/developing/testing-strategy.md#failure-path-requirement)
- **Diff coverage ≥ 80%** — pending local `pnpm test:rust` run
- Coverage matrix updated — `N/A: diagnostic-only change, no new feature row`
- All affected feature IDs from the matrix are listed in the PR description under `## Related`
- No new external network dependencies introduced (uses existing axum-based mock pattern from `spawn_openrouter_probe_server`)
- Manual smoke checklist updated — `N/A: no release-cut surface touched`
- Linked issue closed via `Closes #NNN` — `N/A: Sentry-tracked issue, no GitHub issue yet`

## Impact

- **Runtime**: desktop (Rust core). No mobile / web / CLI surface change.
- **Performance**: negligible — one extra `String` allocation for the body (length already bounded by reqwest's response size limits) and one extra `serde_json::from_str` instead of `response.json()`'s internal equivalent. Happy path serialization cost is identical.
- **Security**: no new surface. Body is sanitized via the same helper the non-2xx branch already trusts; `truncate_with_ellipsis(_, 300)` caps the leak window. No PII redaction policy changes.
- **Migration / compatibility**: none. RPC schema, return type, error-string prefix all preserved. Callers that previously matched on `"failed to parse JSON"` still match — only a `(body: …)` suffix is added.

## Related

- Closes: Sentry [TAURI-RUST-12](https://sentry.tinyhumans.ai/organizations/tinyhumans/issues/100/?project=4&referrer=issue-list&statsPeriod=14d)


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **Bug Fixes**
  * Improved model-listing response parsing and diagnostics: parsing now uses the raw response text and, on failure, error messages include a sanitized, truncated snippet of the body to aid troubleshooting. Non-2xx handling and subsequent response validation remain unchanged.

* **Tests**
  * Added tests covering HTML responses, empty bodies, and valid model-listing payloads.

<!-- review_stack_entry_start -->

[![Review Change Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/tinyhumansai/openhuman/pull/2838?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: M3gA-Mind <megamind@mahadao.com>
2026-05-29 04:04:48 +05:30
ad6d600469 fix(agent): drop bisected tool_call/tool_result pairs before sending to provider (#2840)
## Summary

* Filter unpaired `AssistantToolCalls` and orphan `ToolResults` out of the wire payload in `NativeToolDispatcher::to_provider_messages` so a bisected tool cycle can no longer reach the provider.
* Strip trailing `assistant` messages that carry `tool_calls` without paired tool responses from a resumed transcript in `bound_cached_transcript_messages` (symmetric to the existing leading-orphan strip).
* Add `assistant_message_has_tool_calls` helper that peeks into the JSON-encoded `assistant` ChatMessage content to detect the tool_calls field at the `ChatMessage` boundary.
* Add 5 regression tests for `NativeToolDispatcher::to_provider_messages`: paired cycle, trailing unpaired, mid-history unpaired, orphan ToolResults, and multiple paired cycles back-to-back.

## Problem

Sentry issue **TAURI-RUST-7** — `OpenHuman API error (400 Bad Request): {"success":false,"error":"400 An assistant message with 'tool_calls' must be followed by tool messages responding to each 'tool_call_id'. (insufficient tool messages following tool_calls message)"}` — 589 events / 14d on the `tauri-rust` project.

The OpenAI chat-completions contract requires every `assistant{tool_calls}` message to be immediately followed by `tool` messages — one for every `tool_call_id`. Any other ordering produces `400 An assistant message with 'tool_calls' must be followed by tool messages`.

Two reachable code paths in our harness produce a bisected pair:

1. **Mid-turn abort / resume**. `ConversationMessage::AssistantToolCalls` is pushed into `Agent::history` before the tool runner finishes. If the turn aborts (user cancel, error, process kill, max-iter cap) before the matching `ConversationMessage::ToolResults` is appended, history ends on a tool_calls with no follow-up. The next turn submits the whole history and the backend rejects it.
2. **Cached transcript bound**. `bound_cached_transcript_messages` (`turn.rs:1660`) keeps the tail of a resumed transcript. The existing leading-orphan strip handles a window that opens on a stray `tool` message, but it never stripped a *trailing* assistant tool_calls. A resume from a mid-cycle snapshot ships an unpaired tool_calls and the backend rejects.

`trim_history` already had a leading-orphan guard ([`[turn.rs:1637](https://chatgpt.com/c/src/openhuman/agent/harness/session/turn.rs:1637)`](src/openhuman/agent/harness/session/turn.rs:1637)). The symmetric trailing / unpaired case was uncovered.

## Solution

### `src/openhuman/agent/dispatcher.rs` — Native dispatcher only

`to_provider_messages` is the boundary that serialises `ConversationMessage` → `ChatMessage` (wire format). The Native dispatcher is the only impl affected — XML and PFormat dispatchers render tool results into a `user` role with inline `<tool_result>` tags, so the tool_calls/tool ordering rule doesn't apply to them and their impls are untouched.

Two-pass pairing in a single loop:

```rust
let mut paired_indices: Vec<usize> = Vec::with_capacity(history.len());
for (i, msg) in history.iter().enumerate() {
    match msg {
        AssistantToolCalls { .. } => {
            // Keep only if next is ToolResults.
            if matches!(history.get(i + 1), Some(ToolResults(_))) {
                paired_indices.push(i);
            } else { log::debug!("dropping unpaired AssistantToolCalls at index {i}"); }
        }
        ToolResults(_) => {
            // Keep only if the previous *emitted* index is a kept AssistantToolCalls.
            let preceded_by_kept = i > 0
                && matches!(history.get(i - 1), Some(AssistantToolCalls { .. }))
                && paired_indices.last() == Some(&(i - 1));
            if preceded_by_kept { paired_indices.push(i); }
            else { log::debug!("dropping orphan ToolResults at index {i}"); }
        }
        Chat(_) => paired_indices.push(i),
    }
}
```

The key invariant is `paired_indices.last() == Some(&(i - 1))` — that captures emitted-not-raw lookbehind. A bisected assistant tool_calls is dropped, which means the now-orphan tool results that physically followed it are also dropped, even though `history[i-1]` is still `AssistantToolCalls`. Symmetric drop in a single pass.

Second pass: `flat_map` over `paired_indices` runs the existing serialisation logic verbatim (assistant tool_calls → JSON-encoded ChatMessage; tool results → one tool ChatMessage per result). No serialisation change — only the input set is filtered.

### `src/openhuman/agent/harness/session/turn.rs` — cached transcript bound

This layer operates on `Vec<ChatMessage>`, not `ConversationMessage`. Can't pattern-match on enum variants because the dispatcher's serialisation packs both content and tool_calls into a single JSON-encoded string in the assistant `ChatMessage.content` (see `dispatcher.rs:484-490`). To detect tool_calls at this boundary the helper peeks inside the JSON:

```rust
fn assistant_message_has_tool_calls(msg: &ChatMessage) -> bool {
    if msg.role != "assistant" { return false; }
    let Ok(value) = serde_json::from_str::<serde_json::Value>(&msg.content) else { return false; };
    value.get("tool_calls").and_then(|tc| tc.as_array())
        .map(|arr| !arr.is_empty()).unwrap_or(false)
}
```

Non-assistant role → false. Non-JSON content (a plain text reply) → false. Missing field or empty array → false. Message kept in all those cases.

Then a `pop_while` after the existing leading-orphan strip:

```rust
while bounded.last().map(assistant_message_has_tool_calls).unwrap_or(false) {
    bounded.pop();
    dropped_tail += 1;
}
```

Symmetric to the existing leading-orphan strip — both ends of the bounded window now end on a clean turn boundary.

## Design choices

- Filter at the wire boundary, not at write time. Both write sites (turn loop, transcript restore) push into history without coordinating with each other. Centralising the guard at the serialisation boundary catches every code path that ends up shipping to a provider, present and future.

- Native dispatcher only. XML / PFormat dispatchers render tool results into user role with inline tags and aren't subject to the OpenAI tool ordering rule. Leaving their `to_provider_messages` untouched avoids unrelated behaviour change.

- What we drop the backend would have rejected anyway. Every dropped message would have triggered the same 400. Dropping client-side turns a hard failure into a recoverable turn — the rest of the well-formed history still flies.

- `log::debug!` on drop. Visibility into how often the guard fires without flooding warn-level logs.

-JSON-content inspection helper isolated and tested by the integration tests above. A `serde_json::from_str` parse failure means the content isn't the dispatcher's JSON envelope, so the message is a plain text assistant reply and must be kept — false return is correct.

## Submission Checklist

- Tests added or updated (happy path + at least one failure / edge case) per Testing Strategy

- Diff coverage ≥ 80% — pending local pnpm test:rust run

- Coverage matrix updated — N/A: bug-fix behaviour-only change, no new feature row

- All affected feature IDs from the matrix are listed in the PR description under ## Related

- No new external network dependencies introduced

- Manual smoke checklist updated — N/A: no release-cut surface touched

- Linked issue closed via Closes #NNN — N/A: Sentry-tracked issue, no GitHub issue yet

## Impact

- Runtime: desktop (Rust core). No mobile / web / CLI surface change.

- Performance: one extra `Vec<usize>` allocation sized to `history.len()` in the dispatcher and one O(n) pass over the bounded transcript tail. Happy-path cost is two extra `matches!` checks per message. No measurable overhead on the chat hot path.

- Security: none — no new network surface, no new inputs trusted, no auth path touched.

- Migration / compatibility: none. Trait signatures, RPC schemas, and dispatcher output for fully-paired histories are unchanged. Previously-failing turns that hit the backend's 400 now succeed with a `log::debug!` line per dropped orphan.


## Related

- Closes: [TAURI-RUST-7](https://sentry.tinyhumans.ai/organizations/tinyhumans/issues/47/?project=4&referrer=issue-list&statsPeriod=14d)


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **Bug Fixes**
  * Sanitized provider-bound message history to keep only well-formed assistant opener + matching tool-results (exact ID-set match, order-insensitive); orphaned or mismatched tool messages are dropped.
  * Hardened transcript trimming to strip partial/native tool-call envelopes at window edges so turns end on clean boundaries and log the count of stripped envelopes.

* **Tests**
  * Expanded regression suite covering pairing semantics, orphan/drop cases, ordering tolerance, and transcript-bounding behavior.

<!-- review_stack_entry_start -->

[![Review Change Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/tinyhumansai/openhuman/pull/2840?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: M3gA-Mind <megamind@mahadao.com>
2026-05-29 03:54:09 +05:30
YellowSnnowmannandGitHub 59c66bc385 fix(embeddings): recover from Ollama NaN-encoding 500 instead of failing whole batch (#2834) 2026-05-28 18:02:18 +05:30
442fef3195 refactor(tools): update GitHub tool slugs for clarity and remove deprecated entry (#2766)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 17:55:32 +05:30
YellowSnnowmannandGitHub d578b57e38 fix(observability): demote expected-error Sentry buckets across embeddings, provider, memory-store, FS, and thinking-mode wire shapes (#2830) 2026-05-28 16:26:27 +05:30
YellowSnnowmannandGitHub 4c1194241d feat(cost): add settings cost dashboard with global tracker, dashboard RPCs, and charts (#2762) 2026-05-27 18:36:58 -07:00
YellowSnnowmannandGitHub 129abfe8a4 feat(chat): render agent-bubble LaTeX with KaTeX and safe math detection (#2697) 2026-05-27 18:36:31 -07:00
YellowSnnowmannandGitHub 99b3affff5 docs(readme): fix Homebrew tap namespace in macOS install instructions (#2739) 2026-05-27 16:06:22 +05:30
YellowSnnowmannandGitHub fae16082d3 fix(memory-workspace): show all native Composio memory-sync sources (not Gmail-only) (#2685) 2026-05-27 15:54:10 +05:30
YellowSnnowmannandGitHub ee63a9ddbc fix(e2e/oauth): gate Redux store exposure and fix loopback listener timeout race (#2670) 2026-05-27 15:24:23 +05:30
e6192e242e fix(security): expand default allowed_commands and auto_approve (#2486) (#2673)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-27 12:51:16 +05:30
e2601a20d1 feat(doctor): surface embedding model health to user (#2474) (#2674)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-27 12:50:45 +05:30
YellowSnnowmannandGitHub f946eda516 fix(tauri): harden Windows pre-CEF single-instance mutex handling (#2669) 2026-05-27 12:46:37 +05:30
YellowSnnowmannandGitHub a6b19ce239 fix(agent-harness): dedup visible tool specs in all provider-call paths (#2446) 2026-05-22 17:40:34 -07:00
YellowSnnowmannandGitHub b0facb5e78 fix(config): preserve built-in reserved-slug cloud_providers across settings saves (#2457) 2026-05-22 17:40:09 -07:00
03d1e2512e feat(e2e): complete E2E v2 suite — 66 specs, orchestrator, bug fixes (#2353)
Co-authored-by: Steven Enamakel <31011319+senamakel@users.noreply.github.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-22 17:23:01 -07:00
YellowSnnowmannandGitHub 7aa1bf1f88 fix(agent): handle config rejection in streaming_chat path (#2346) 2026-05-21 23:24:07 +05:30
YellowSnnowmannandGitHub f51f140234 fix(prompt-injection): rebalance detector + classify rejections as expected (#2429) 2026-05-21 22:57:55 +05:30
YellowSnnowmannandGitHub 208a64483b fix(auth-profiles): tolerate legacy kind values on load (#2439) 2026-05-21 22:57:49 +05:30
b9925e0bf4 feat(memory): on-device multilingual PII redaction
## Summary

- Added an on-device multilingual PII redactor covering 15 identifier types — Brazilian CPF/CNPJ (mod-11), Argentine CUIT, Mexican RFC, Japanese マイナンバー, US SSN with reserved-range filter, Luhn-validated credit cards, mod-97 IBAN, Verhoeff-validated Aadhaar, Indian PAN, UK NINO, Spanish DNI/NIE, Korean RRN, E.164 and NANP phones — and wired it into the existing `sanitize_text` pipeline that runs on every memory write.

- Added a Unicode-normalization pre-pass that strips zero-width characters and folds fullwidth/Arabic-Indic digits + punctuation, defeating common regex-bypass tricks before matching while preserving non-PII bytes in the original text.

- Added `has_likely_pii()` and wired it into the namespace/key boundary checks in `documents.rs` and `kv.rs` so PII-bearing inputs are rejected outright, mirroring how `has_likely_secret()` is enforced today.

- Extended `SanitizationReport` with a `pii_redactions` counter and surfaced it in the existing `[memory:safety]` audit-log lines across `documents.rs`, `kv.rs`, and `fts5.rs`.

## Problem

- OpenHuman ingests data from 118+ integrations into Memory Tree. The existing `memory::safety` module redacted API keys, tokens, and PEM blocks but had no coverage for personal PII — national IDs, financial identifiers, phone numbers.

- Issue #2017 proposed sending raw content to a third-party HTTP endpoint (`api.trustboost.dev`) for "sanitization". That approach directly contradicts OpenHuman's privacy-first, on-device posture — it would exfiltrate the exact PII it claims to protect to an unaffiliated vendor.

- A naive regex implementation would still leave two real gaps: (a) the multilingual identifier formats that motivated the original issue (LATAM, JP, IN, EU, KR) and (b) trivial bypass via fullwidth-digit or zero-width-character obfuscation, which any motivated attacker (or accidentally-pasted Japanese-locale data) will trigger.

## Solution

- Built `src/openhuman/memory/safety/pii.rs` with 15 PII categories. Where checksums exist (CPF/CNPJ mod-11, CUIT, credit-card Luhn, IBAN mod-97, Aadhaar Verhoeff, Spanish DNI/NIE check-letter, SSN reserved-range), false-positives are rejected at the algorithm level — no LLM, no network. Where checksums don't exist (RFC, PAN-IN, NANP, E.164, RRN), structural format rules carry the discrimination.

- Added a `NormalizedView` that strips U+200B/200C/200D/FEFF/2060/180E and folds fullwidth (`0-9`, `.-/:`) plus Arabic-Indic / Eastern Arabic-Indic digits to ASCII before matching. Match offsets are mapped back to the original byte positions so only PII bytes are replaced — surrounding text (including any intentional fullwidth glyphs) is byte-identical to input.

- Patterns run in priority order (formatted before bare, IBAN before credit-card, etc.) with overlap-deduplication so a single span can't be redacted twice or partially counted as multiple types. A `RegexSet` pre-filter short-circuits PII-free text in one scan instead of ~18 per-pattern scans.

- `has_likely_pii()` mirrors `has_likely_secret()` and is wired into the same boundary checks in `unified/documents.rs` (both `upsert_document` and `upsert_document_metadata_only`) and `unified/kv.rs` (both `kv_set_global` and `kv_set_namespace`).

- Added 37 new tests in `pii.rs` and 2 integration tests in `safety/mod.rs`: positive + negative per pattern, checksum-failing rejection cases, Unicode/zero-width bypass attempts, `has_likely_pii` gating, and an aggressive mixed-language end-to-end test covering 13 PII types in one document. Full safety suite: 53 tests passing. Full memory module: 1007 tests passing, zero regressions.

## Submission Checklist

- If a section does not apply to this change, mark the item as N/A with a one-line reason. Do not delete items.

- [x] Tests added or updated (happy path + at least one failure / edge case) per Testing Strategy — 39 new tests; checksum-failing and bypass-attempt negatives included
- [x] Diff coverage ≥ 80% — new file is ~100% line-covered by inline tests; integration sites have direct assertions in `safety/mod.rs::tests`
- [x] All affected feature IDs from the matrix listed in ## Related — N/A no existing matrix IDs cover memory safety redaction; the new row above will be the first
- [x] No new external network dependencies introduced — deliberately on-device; no new crates; `regex` and `once_cell` already in tree
- [x] Manual smoke checklist updated — N/A not a release-cut UI surface
- [x] Linked issue closed via Closes #NNN — see ## Related

## Impact

- Runtime/platform impact: Rust core memory ingestion path only. Desktop app behavior change is two-fold for users — (1) memory writes now produce additional `[REDACTED_PII_*]` tokens in stored content when format-matching PII is present, (2) a new error return (`document/kv namespace/key cannot contain personal identifiers`) on the rare case where a caller tries to use a PII-shaped string as a namespace or key.

- Performance: `RegexSet` pre-filter makes PII-free text a single-scan no-op. On text containing PII, adds one normalized-string allocation plus a handful of regex scans gated by the screen — negligible compared to the embedding/SQLite/markdown-sidecar costs already on the write path. No measurable impact on ingestion latency in local testing.

- Security/migration/compatibility: no schema changes, no new dependencies. The boundary-gate rejection is a behavior change for any caller that previously stored namespace/keys *containing* identifier-shaped strings; expected impact is zero in practice because real namespace/keys are paths like `memory/global/preferences`. Privacy posture is strictly improved — every byte stays on device; no telemetry, no outbound calls.

## Related

- Closes: #2017 





<!-- This is an auto-generated comment: release notes by coderabbit.ai -->

## Summary by CodeRabbit

* **New Features**
  * Personal information detection and redaction integrated into the memory system
  * Write operations on namespace and key fields now validate against personal identifiers

* **Improvements**
  * Enhanced sanitization reports with additional metrics on personal identifier redactions

<!-- review_stack_entry_start -->

[![Review Change Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/tinyhumansai/openhuman/pull/2310?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Shanu <shanu@tinyhumans.ai>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-20 16:54:38 -07:00
YellowSnnowmannandGitHub fa8d75fb5b fix(tauri): skip single-instance plugin when D-Bus session bus is unreachable (#2352) 2026-05-21 00:55:52 +05:30
YellowSnnowmannandGitHub f24dbc6653 fix(observability): demote transient OpenAI embeddings 429s to expected and reduce Sentry noise (#2294) 2026-05-20 18:37:14 +05:30
9ec2ae765f fix(e2e): sync E2E specs with current codebase (#2220)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-19 22:46:56 -07:00
a40272ef45 fix(observability,database): silence expected provider/channel errors and add SQLite busy timeout for WhatsApp store (#2107)
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-19 21:08:25 -07:00
ff8d60c3bb fix(observability): classify Kimi access_terminated 403 as expected provider user-state (#2090)
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-19 19:42:51 -07:00
094d482210 fix(credentials): recover from leaked auth-profile lock on Windows (Sentry OPENHUMAN-TAURI-H1) (#2085)
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-19 16:52:42 -07:00
YellowSnnowmannandGitHub c37f459dc3 Fix/embeddings OpenAI expected error reporting (#2190) 2026-05-19 14:18:10 -07:00
YellowSnnowmannandGitHub 868ad8d5d5 fix(tauri): resolve Linux CEF init panic — root/container + SingletonLock + display-server guards (OPENHUMAN-TAURI-K1) (#2103) 2026-05-19 14:17:54 -07:00
YellowSnnowmannandGitHub ca379c7f9f fix(config): guard env overrides during config load (#2201) 2026-05-19 14:17:35 -07:00
YellowSnnowmannandGitHub d6a99fcba4 fix(reliable): fail fast on SESSION_EXPIRED in provider retry loop (#2200) 2026-05-19 14:17:24 -07:00
YellowSnnowmannandGitHub 525d7c7ec5 fix(agent): bound cached resume transcript by max_history_messages (#2224) 2026-05-19 14:17:12 -07:00
YellowSnnowmannandGitHub 1cadb986fc fix(credentials): diagnose + recover from H8 auth-profile-lock create failures (#2180) 2026-05-19 14:16:54 -07:00
YellowSnnowmannandGitHub 13500dbb37 fix(inference): propagate temperature_unsupported_models to local routing providers (#2221) 2026-05-19 21:22:33 +05:30
YellowSnnowmannandGitHub 52203454ce fix(security): self-repair locked .secret_key on Windows (OPENHUMAN-TAURI-GN) (#2061) 2026-05-18 15:58:47 +05:30
YellowSnnowmannandGitHub bc371bdea2 docs: align Claude/Codex context with current main (#1789) 2026-05-15 15:02:02 -07:00
b4c19b105b fix(auth): scope clearAllAppData to active user; fix re-onboarding race; drop dead API call (#1816)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-15 14:57:30 -07:00
c37a122a0f fix(jira): collect Atlassian subdomain and handle ConnectedAccount_MissingRequiredFields (#1726)
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-15 04:14:23 -07:00
YellowSnnowmannandGitHub 1550e9f428 fix: pre-CEF single-instance mutex guard on Windows + provider retry for 502s (#1723) 2026-05-14 21:06:39 -07:00
8ae921dd68 fix(normalization): guard function.arguments against malformed JSON and default to {} (#1645)
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-13 20:07:04 -07:00
YellowSnnowmannandGitHub 10a726d007 feat(voice): add mic input device selector and stabilize media capture for composer (#1616) 2026-05-13 16:35:24 +05:30
386b0025ed fix(observability/auth): handle JWT-required expiry path and normalize backend API base URLs (#1551)
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-12 19:56:24 -07:00
706edf5f44 fix(composio): collect WABA ID before WhatsApp Business OAuth flow (#1550)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Steven Enamakel <enamakel@tinyhumans.ai>
2026-05-12 19:51:54 -07:00
a2ced40ef9 feat(memory): add tool-scoped memory with durable rule capture, RPC APIs, and prompt pinning (#1487)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-11 12:07:59 -07:00
YellowSnnowmannandGitHub fc573b05f6 fix(core-rpc): normalize config method wiring and harden startup against config/SQLite edge cases (#1497) 2026-05-11 09:30:05 -07:00
YellowSnnowmannandGitHub e58b5abfcf fix(meet_call): abort scanner on close to unblock 60-second navigation stall (#1380) 2026-05-08 19:04:31 -07:00
YellowSnnowmannandGitHub dcd4f97f00 fix(ci): staging builds resolve to prod API URL — bake VITE vars into build.yml (#1371) 2026-05-08 19:04:00 -07:00
YellowSnnowmannandGitHub 0a749f5cec feat(heartbeat): deliver durable proactive meeting/reminder notifications with dedup + category controls (#1369) 2026-05-08 19:03:31 -07:00
YellowSnnowmannandGitHub 9c1df8d91b fix(webview): LinkedIn "Sign in with Google" — keep GSI popup in-app (#1368) 2026-05-08 19:03:02 -07:00
YellowSnnowmannandGitHub 2c047a2153 feat(composio): granular trigger triage settings — per-toolkit + global toggle (#1334) 2026-05-07 12:31:40 -07:00
YellowSnnowmannandGitHub c2f2d8497c feat(agent): enforce subagent role contract with concise delegated outputs (#1336) 2026-05-07 12:30:10 -07:00
YellowSnnowmannandGitHub 694ae4e677 fix(chat): sanitize agent/cron failures and add user-safe error fallback with Sentry reporting (#1332) 2026-05-07 12:28:41 -07:00
YellowSnnowmannandGitHub a3b2fb814c feat(memory-security): prevent secret leakage into agent memory with redaction, validation, and diagnostics (#1224) 2026-05-05 11:08:25 -07:00
YellowSnnowmannandGitHub 995e5ccccb feat(security): enforce prompt-injection guard before model and tool execution (#1175) 2026-05-04 02:30:22 -07:00