From b02916d0bebed11960e226c9816636749d5860f5 Mon Sep 17 00:00:00 2001 From: Steven Enamakel <31011319+senamakel@users.noreply.github.com> Date: Sat, 9 May 2026 00:21:51 -0700 Subject: [PATCH] Feat/docs v8 (#1391) --- gitbooks/README.md | 6 +-- gitbooks/features/integrations/triggers.md | 7 ++- gitbooks/features/mascot/meeting-agents.md | 52 +++++++++++----------- 3 files changed, 32 insertions(+), 33 deletions(-) diff --git a/gitbooks/README.md b/gitbooks/README.md index 113a75e38..fd52bd231 100644 --- a/gitbooks/README.md +++ b/gitbooks/README.md @@ -1,8 +1,8 @@ --- description: >- - OpenHuman is a personal AI assistant that runs on your desktop, connects to - 118+ services, builds a local-first memory of your life from them, - self-reflects and can do interact with you in audio/video + Personal AI assistant for your desktop. Connects to 118+ services, builds a + local-first memory of your life, self-reflects, and can interact with you + over audio and video. icon: diamond --- diff --git a/gitbooks/features/integrations/triggers.md b/gitbooks/features/integrations/triggers.md index f6dc22b33..47143ed96 100644 --- a/gitbooks/features/integrations/triggers.md +++ b/gitbooks/features/integrations/triggers.md @@ -1,9 +1,8 @@ --- description: >- - External events from connected integrations (a new Gmail message, a Notion - edit, a Stripe charge, an inbound Slack DM) flow into OpenHuman as - triggers, get classified by a triage agent, and can fire off agent - actions automatically. + Live events from connected integrations (a new Gmail message, a Notion edit, + a Stripe charge) arrive as triggers, get classified by a triage agent, and + can fire agent actions automatically. icon: bolt --- diff --git a/gitbooks/features/mascot/meeting-agents.md b/gitbooks/features/mascot/meeting-agents.md index 6c39b34d1..fd4df4241 100644 --- a/gitbooks/features/mascot/meeting-agents.md +++ b/gitbooks/features/mascot/meeting-agents.md @@ -1,8 +1,8 @@ --- description: >- - The mascot can join meetings as a real participant - it listens, takes notes, - speaks back into the call, animates its face into the camera grid, and uses - tools mid-meeting. More than just a notetaker. + The mascot joins meetings as a real participant: listens, takes notes, speaks + back into the call, animates its face into the camera grid, and uses tools + mid-meeting. More than a notetaker. icon: video --- @@ -48,10 +48,10 @@ This is the difference between a transcription bot and a meeting _agent_. While the call is happening, the mascot has access to the same tool surface it has on your desktop: -* [**Memory Tree**](../obsidian-wiki/memory-tree.md) - recall prior meetings, decisions, open threads, who said what last time, what's been promised. -* [**Auto-fetch from integrations**](../obsidian-wiki/auto-fetch.md) and [**third-party integrations**](../integrations/) - pull a thread from Slack, an email, a Linear ticket, a Notion doc, a calendar entry, a file from Drive. -* [**Native tools**](../native-tools/) - search the web, scrape a page, run a quick code/data lookup, all without leaving the call. -* [**Subconscious Loop**](../subconscious.md) outputs - anything it has been working on in the background is already on hand. +- [**Memory Tree**](../obsidian-wiki/memory-tree.md) - recall prior meetings, decisions, open threads, who said what last time, what's been promised. +- [**Auto-fetch from integrations**](../obsidian-wiki/auto-fetch.md) and [**third-party integrations**](../integrations/) - pull a thread from Slack, an email, a Linear ticket, a Notion doc, a calendar entry, a file from Drive. +- [**Native tools**](../native-tools/) - search the web, scrape a page, run a quick code/data lookup, all without leaving the call. +- [**Subconscious Loop**](../subconscious.md) outputs - anything it has been working on in the background is already on hand. So when someone in the call asks "wait, didn't we decide to drop the Q3 launch last month?", the mascot doesn't just transcribe the question. It answers it - with the actual decision, the meeting it was made in, and who agreed. @@ -61,35 +61,35 @@ That moves it from _notetaker_ to _the most informed participant in the room_. A meeting agent that only transcribes is a tool. A meeting agent that participates is a presence. The Meet integration is deliberately built to make the mascot feel like a real attendee, not a recording device: -* It has a **face on the camera grid** that lip-syncs and reacts, not a black square or a logo. -* It has its **own voice** that plays into the call, not into your speakers. -* It has **persistent memory** of the people in the room, the project, the prior decisions - so it can be addressed by name and answer in context. -* It has **tools** so it can act on what's said, not just record it. -* It runs the **subconscious loop** between meetings - so when it joins your next call, it has already done the homework on what was promised in the last one. +- It has a **face on the camera grid** that lip-syncs and reacts, not a black square or a logo. +- It has its **own voice** that plays into the call, not into your speakers. +- It has **persistent memory** of the people in the room, the project, the prior decisions - so it can be addressed by name and answer in context. +- It has **tools** so it can act on what's said, not just record it. +- It runs the **subconscious loop** between meetings - so when it joins your next call, it has already done the homework on what was promised in the last one. The result, in practice, is that participants stop treating it like a bot and start treating it like a colleague who happens to be very fast at looking things up. ## Setup, controls, privacy -* **Joining a call.** You can hand the mascot a Google Meet link from the desktop app; it will open the embedded Meet webview, join with the configured display name, and switch its camera tile to the mascot canvas. -* **Mic and camera control.** The agent's mic is the TTS injection stream, not your real microphone. The agent's camera is the mascot frame producer, not your real webcam. You can mute the agent's mic from the app at any time, the same way you'd mute yourself in Meet. -* **Transcripts and memory.** Live transcripts land in the [Memory Tree](../obsidian-wiki/memory-tree.md) the same way any other source does - under the people in the call, the project, and the topics that came up. They are local-first and follow the project's [Privacy & Security](../privacy-and-security.md) rules. -* **No covert recording.** The agent appears as a normal participant in the grid; everyone in the call can see it and see when it's speaking. +- **Joining a call.** You can hand the mascot a Google Meet link from the desktop app; it will open the embedded Meet webview, join with the configured display name, and switch its camera tile to the mascot canvas. +- **Mic and camera control.** The agent's mic is the TTS injection stream, not your real microphone. The agent's camera is the mascot frame producer, not your real webcam. You can mute the agent's mic from the app at any time, the same way you'd mute yourself in Meet. +- **Transcripts and memory.** Live transcripts land in the [Memory Tree](../obsidian-wiki/memory-tree.md) the same way any other source does - under the people in the call, the project, and the topics that came up. They are local-first and follow the project's [Privacy & Security](../privacy-and-security.md) rules. +- **No covert recording.** The agent appears as a normal participant in the grid; everyone in the call can see it and see when it's speaking. ## Implementation pointers (for developers) Curious how this is wired up: -* Brain - `src/openhuman/meet_agent/brain.rs` (LLM turns, speak/no-speak decisions, tool calls). -* Voice plumbing - `src/openhuman/voice/` (STT in, TTS out, hallucination filter, postprocess). See [Native Voice](../native-tools/voice.md). -* Mascot canvas as outbound camera - `app/src/features/meet/MascotFrameProducer.tsx` and the Tauri-side `mascot_native_window.rs` window. -* Embedded Meet webview - see [Chromium Embedded Framework](../../developing/cef.md). The Meet child webview ships with **zero injected JavaScript**; everything host-side runs natively via CDP. -* Notable commits to read for context - `0bc74575` (live note-taking), `f1203479` (real LLM turns + tuned TTS), `b6d05cb4` (mascot canvas as outbound camera), `f5dce783` (mascot frame pipeline + off-screen meet window). +- Brain - `src/openhuman/meet_agent/brain.rs` (LLM turns, speak/no-speak decisions, tool calls). +- Voice plumbing - `src/openhuman/voice/` (STT in, TTS out, hallucination filter, postprocess). See [Native Voice](../native-tools/voice.md). +- Mascot canvas as outbound camera - `app/src/features/meet/MascotFrameProducer.tsx` and the Tauri-side `mascot_native_window.rs` window. +- Embedded Meet webview - see [Chromium Embedded Framework](../../developing/cef.md). The Meet child webview ships with **zero injected JavaScript**; everything host-side runs natively via CDP. +- Notable commits to read for context - `0bc74575` (live note-taking), `f1203479` (real LLM turns + tuned TTS), `b6d05cb4` (mascot canvas as outbound camera), `f5dce783` (mascot frame pipeline + off-screen meet window). ## See also -* [The Mascot](./) - the on-screen character itself, outside of meetings. -* [Native Voice](../native-tools/voice.md) - STT / TTS that the meeting agent rides on. -* [Memory Tree](../obsidian-wiki/memory-tree.md) - where transcripts and decisions land. -* [Native Tools](../native-tools/) - what the mascot can reach for mid-call. -* [Automatic Model Routing](../model-routing/) - why conversational turns feel low-latency. +- [The Mascot](./) - the on-screen character itself, outside of meetings. +- [Native Voice](../native-tools/voice.md) - STT / TTS that the meeting agent rides on. +- [Memory Tree](../obsidian-wiki/memory-tree.md) - where transcripts and decisions land. +- [Native Tools](../native-tools/) - what the mascot can reach for mid-call. +- [Automatic Model Routing](../model-routing/) - why conversational turns feel low-latency.