Docs/overhaul readme gitbooks (#4225)

This commit is contained in:
Steven Enamakel
2026-06-26 13:02:55 -07:00
committed by GitHub
parent 1a676c6d44
commit 5a41a4f0d2
76 changed files with 1900 additions and 7755 deletions
+14 -8
View File
@@ -39,7 +39,7 @@
</p>
<p align="center">
🇺🇸 <a href="./README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
🇺🇸 <a href="./README.md">English</a> | 🇨🇳 <a href="./docs/README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./docs/README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./docs/README.ko.md">한국어</a> | 🇩🇪 <a href="./docs/README.de.md">Deutsch</a> | 🇵🇰 <a href="./docs/README.ur-pk.md">اردو</a>
</p>
@@ -49,11 +49,11 @@
<a href="https://github.com/tinyhumansai/openhuman/releases/latest"><img src="https://img.shields.io/github/v/release/tinyhumansai/openhuman?label=latest" alt="Latest Release" /></a>
<a href="https://github.com/tinyhumansai/openhuman/stargazers"><img src="https://img.shields.io/github/stars/tinyhumansai/openhuman?style=flat" alt="GitHub Stars" /></a>
<a href="./LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="License" /></a>
<a href="./README.zh-CN.md"><img src="https://img.shields.io/badge/lang-简体中文-blue" alt="简体中文" /></a>
<a href="./README.ja-JP.md"><img src="https://img.shields.io/badge/lang-日本語-blue" alt="日本語" /></a>
<a href="./README.ko.md"><img src="https://img.shields.io/badge/lang-한국어-blue" alt="한국어" /></a>
<a href="./README.de.md"><img src="https://img.shields.io/badge/lang-Deutsch-blue" alt="Deutsch" /></a>
<a href="./README.ur-pk.md"><img src="https://img.shields.io/badge/lang-اردو-blue" alt="اردو" /></a>
<a href="./docs/README.zh-CN.md"><img src="https://img.shields.io/badge/lang-简体中文-blue" alt="简体中文" /></a>
<a href="./docs/README.ja-JP.md"><img src="https://img.shields.io/badge/lang-日本語-blue" alt="日本語" /></a>
<a href="./docs/README.ko.md"><img src="https://img.shields.io/badge/lang-한국어-blue" alt="한국어" /></a>
<a href="./docs/README.de.md"><img src="https://img.shields.io/badge/lang-Deutsch-blue" alt="Deutsch" /></a>
<a href="./docs/README.ur-pk.md"><img src="https://img.shields.io/badge/lang-اردو-blue" alt="اردو" /></a>
</p>
> **Early Beta**: Under active development. Expect rough edges.
@@ -118,12 +118,18 @@ OpenHuman is an open-source agentic assistant designed to integrate with you in
- **Simple, UI-first & Human** A clean desktop experience and short onboarding paths take you from install to a working agent in a few clicks — no config-first setup, no terminal required. The agent has [a face](https://tinyhumans.gitbook.io/openhuman/features/mascot): a desktop mascot that speaks, reacts to its surroundings, [joins your Google Meets](https://tinyhumans.gitbook.io/openhuman/features/mascot/meeting-agents) as a real participant, remembers you across weeks, and keeps thinking in the background even when you've stopped typing.
- **[118+ third-party integrations](https://tinyhumans.gitbook.io/openhuman/features/integrations) with [auto-fetch](https://tinyhumans.gitbook.io/openhuman/features/obsidian-wiki/auto-fetch)**: plug into Gmail, Notion, GitHub, Slack, Stripe, Calendar, Drive, Linear, Jira and the rest of your stack with **one-click OAuth**. Every connection is exposed to the agent as a typed tool, and every twenty minutes the core walks each active connection and pulls fresh data into the [memory tree](https://tinyhumans.gitbook.io/openhuman/features/integrations/auto-fetch). No prompts, no polling loops you have to write, so the agent already has tomorrow's context this morning.
- **100+ one-click OAuth integrations, 5,000+ MCP servers, 90,000+ Skills**: plug into Gmail, Notion, GitHub, Slack, Stripe, Calendar, Drive, Linear, Jira and the rest of your stack with [**one-click OAuth**](https://tinyhumans.gitbook.io/openhuman/features/integrations) — 100+ curated connectors brokered through the Composio layer. Beyond that, OpenHuman browses the open **Model Context Protocol** ecosystem (Smithery + the official MCP registry — thousands of servers) and a **90,000-entry Skills catalog**, so the agent can install new typed tools and skills on demand. Every connection becomes a typed tool, and every twenty minutes [auto-fetch](https://tinyhumans.gitbook.io/openhuman/features/obsidian-wiki/auto-fetch) walks each active connection and pulls fresh data into the [memory tree](https://tinyhumans.gitbook.io/openhuman/features/integrations/auto-fetch). No prompts, no polling loops you have to write, so the agent already has tomorrow's context this morning.
Managed integrations use OpenHuman's Composio connector layer. OAuth handshakes and integration tool calls are proxied through the managed backend by default. If you want to run Composio directly instead, configure direct mode with your own Composio API key; real-time trigger webhooks then need to be hosted and wired by you.
- **[Memory Tree](https://tinyhumans.gitbook.io/openhuman/features/memory-tree) + [Obsidian Wiki](https://tinyhumans.gitbook.io/openhuman/features/obsidian-wiki)**: a local-first knowledge base built from your data and your activity. Everything you connect is canonicalized into ≤3k-token Markdown chunks, scored, and folded into hierarchical summary trees stored in **SQLite on your machine**. The same chunks land as `.md` files in an Obsidian-compatible vault you can open, browse and edit, inspired by Karpathy's [obsidian-wiki workflow](https://x.com/karpathy/status/2039805659525644595).
- **[SuperContext](https://tinyhumans.gitbook.io/openhuman/features/super-context)**: a fresh chat shouldn't start cold. With SuperContext enabled, the harness deterministically spawns a read-only `context_scout` on the **first turn of every new thread** — it sweeps your memory tree, files, and connected data, assembles a bounded context bundle, and prepends it to your message before the model ever reads it. No tool call to wait on, no "let me look that up" round-trip: the agent answers your first message already knowing the relevant background. Toggle it from the composer or `context.super_context_enabled`.
- **[Goals & Todos](https://tinyhumans.gitbook.io/openhuman/features/goals-and-todos)**: OpenHuman keeps the agent pointed at what matters. A short, human-editable list of **long-term goals** (`MEMORY_GOALS.md`) rides along in memory and can self-reflect against your recent activity; each **thread** can carry a single durable **goal** with an optional token budget that the agent works across turns, interrupts and idle periods (autonomous idle continuation); and every conversation hosts a **kanban task board** of todos that you and the agent build together — plans, acceptance criteria, approval gates and all. There's also a personal task list you own outright.
- **[Themes & Theme Studio](https://tinyhumans.gitbook.io/openhuman/features/theming)**: make it yours. Five built-in theme families (Classic, Ocean, Sepia, Matrix, HAL 9000) ship in light/dark/auto variants, and the **Theme Studio** in Settings is a full visual editor — adjust every colour token with live contrast warnings, swap fonts per role, pick an animated WebGL-mesh / flat / custom-image backdrop, and export or import themes as JSON to share. Editing a preset auto-forks a custom theme so the originals stay pristine; everything applies instantly and persists locally.
- **Batteries included**: web search, a web-fetch [scraper](https://tinyhumans.gitbook.io/openhuman/features/native-tools), a full coder toolset (filesystem, git, lint, test, grep), and [native voice](https://tinyhumans.gitbook.io/openhuman/features/voice) (STT in, ElevenLabs TTS out, mascot lip-sync, live Google Meet agent) are wired in by default. By default, [model routing](https://tinyhumans.gitbook.io/openhuman/features/model-routing) uses the OpenHuman backend to select and proxy the right LLM for each workload (reasoning, fast, or vision). One subscription includes all models. No "install a plugin to read files" friction. Use [optional local AI via Ollama](https://tinyhumans.gitbook.io/openhuman/features/model-routing/local-ai) for supported on-device workloads.
- **[Smart token compression (TokenJuice)](https://tinyhumans.gitbook.io/openhuman/features/token-compression)**: every tool call, scrape result, email body, and search payload is run through a token compression layer before it touches any LLM Model. HTML is converted to Markdown, long URLs are shortened, and verbose tool output is deduped and summarized via a configurable rule overlay etc... CJK, emoji, and other multi-byte text are preserved grapheme-by-grapheme — never stripped. You get the same information but at a fraction of the tokens. Reducing cost &amp; latency by up to 80%.
@@ -166,7 +172,7 @@ High-level comparison (products evolve, so verify against each vendor). OpenHuma
| **Simple to start** | ✅ Desktop + CLI | ⚠️ Terminal-first | ⚠️ Terminal-first | ✅ Clean UI, minutes |
| **Cost** | ⚠️ Sub + add-ons | ⚠️ BYO models | ⚠️ BYO models | ✅ One sub + TokenJuice |
| **Memory** | ✅ Chat-scoped | ⚠️ Plugin-reliant | ✅ Self-learning | 🚀 Memory Tree + Obsidian vault, optional [agentmemory](https://github.com/rohitg00/agentmemory) backend |
| **Integrations** | ⚠️ Few connectors | ⚠️ BYO | ⚠️ BYO | 🚀 118+ via OAuth |
| **Integrations** | ⚠️ Few connectors | ⚠️ BYO | ⚠️ BYO | 🚀 100+ OAuth · 5k+ MCP · 90k+ Skills |
| **Auto-fetch** | 🚫 None | 🚫 None | 🚫 None | ✅ 20-min sync into memory |
| **API sprawl** | 🚫 Extra keys | 🚫 BYOK | 🚫 Multi-vendor | ✅ One account |
| **Model routing** | 🚫 Single model | ⚠️ Manual | ⚠️ Manual | ✅ Built-in |
+8 -8
View File
@@ -1,12 +1,12 @@
<p align="center">
🇺🇸 <a href="./README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
🇺🇸 <a href="../README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
</p>
<h1 align="center">OpenHuman</h1>
<p align="center">
<img src="./gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
<img src="../gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
</p>
<p align="center" style="display: inline-block">
@@ -37,8 +37,8 @@
<img src="https://img.shields.io/badge/status-early%20beta-orange" alt="Frühe Beta" />
<a href="https://github.com/tinyhumansai/openhuman/releases/latest"><img src="https://img.shields.io/github/v/release/tinyhumansai/openhuman?label=latest" alt="Aktuellste Version" /></a>
<a href="https://github.com/tinyhumansai/openhuman/stargazers"><img src="https://img.shields.io/github/stars/tinyhumansai/openhuman?style=flat" alt="GitHub Stars" /></a>
<a href="./LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="Lizenz" /></a>
<a href="./README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="../LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="Lizenz" /></a>
<a href="../README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="./README.zh-CN.md"><img src="https://img.shields.io/badge/lang-简体中文-blue" alt="简体中文" /></a>
<a href="./README.ja-JP.md"><img src="https://img.shields.io/badge/lang-日本語-blue" alt="日本語" /></a>
<a href="./README.ko.md"><img src="https://img.shields.io/badge/lang-한국어-blue" alt="한국어" /></a>
@@ -64,7 +64,7 @@ irm https://raw.githubusercontent.com/tinyhumansai/openhuman/main/scripts/instal
<!-- TODO: translate (de) — English source mirrored from README.md so non-EN readers get the same install caveats. Please translate. -->
> **Linux:** the AppImage can crash on launch under Wayland (and on Arch-based distros with `sharun: Interpreter not found!`) — see [#2463](https://github.com/tinyhumansai/openhuman/issues/2463) for the cause and env-var workarounds.
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](./packages/arch/openhuman-bin/);
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](../packages/arch/openhuman-bin/);
once published, Arch users can install it with `yay -S openhuman-bin`.
<!-- /TODO -->
@@ -88,20 +88,20 @@ OpenHuman ist ein quelloffener, agentenbasierter Assistent, der sich in deinen A
## Beitragen aus dem Quellcode
Neu hier? Beginne mit [`CONTRIBUTING.md`](./CONTRIBUTING.md) für den Fork-/PR-Workflow und die lokalen Prüfbefehle. Der kurze Weg:
Neu hier? Beginne mit [`CONTRIBUTING.md`](../CONTRIBUTING.md) für den Fork-/PR-Workflow und die lokalen Prüfbefehle. Der kurze Weg:
1. Installiere Git, Node.js 24+, pnpm 10.10.0, Rust 1.93.0 (`rustfmt` + `clippy`), CMake, Ninja, ripgrep sowie die plattformspezifischen Desktop-Build-Voraussetzungen.
2. Forke und klone das Repo, führe dann `git submodule update --init --recursive` aus, bevor du `pnpm install` startest, damit die mitgelieferten Tauri/CEF-Quellen vorhanden sind.
3. Nutze `pnpm dev` für reine Web-UI-Arbeit, `pnpm --filter openhuman-app dev:app` für die Desktop-Shell sowie gezielte Checks wie `pnpm typecheck`, `pnpm format:check` und `cargo check -p openhuman --lib`, bevor du einen PR öffnest.
Tiefer einsteigen: [Architektur](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [Einrichtung](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [Cloud-Deployment](./gitbooks/features/cloud-deploy.md).
Tiefer einsteigen: [Architektur](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [Einrichtung](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [Cloud-Deployment](../gitbooks/features/cloud-deploy.md).
## Kontext in Minuten, nicht in Wochen
OpenHuman ist das erste Agent-Harness, das dich in Minuten kennenlernt. Inspiriert von [Karpathys LLM-Knowledgebase](https://x.com/karpathy/status/2039805659525644595). Die meisten Agenten starten aus dem Kalten. Hermes lernt, indem er dir bei der Arbeit zusieht; OpenClaw wartet darauf, dass Plugins Kontext einspielen. So oder so vergehen Tage oder Wochen, bevor der Agent genug über deinen Stack weiß, um wirklich nützlich zu sein.
<p align="center">
<img src="./gitbooks/.gitbook/assets/image (1).png" alt="Diagramm zum OpenHuman-Kontextaufbau" />
<img src="../gitbooks/.gitbook/assets/image (1).png" alt="Diagramm zum OpenHuman-Kontextaufbau" />
</p>
> OpenHuman fasst all deine Dokumente, E-Mails und Chats zusammen, komprimiert sie und legt einen Memory Graph an, mit dem dein Agent sich alles über dich merken kann.
+8 -8
View File
@@ -1,7 +1,7 @@
<h1 align="center">OpenHuman</h1>
<p align="center">
<img src="./gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
<img src="../gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
</p>
<p align="center" style="display: inline-block">
@@ -29,7 +29,7 @@
</p>
<p align="center">
🇺🇸 <a href="./README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
🇺🇸 <a href="../README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
</p>
@@ -38,8 +38,8 @@
<img src="https://img.shields.io/badge/status-early%20beta-orange" alt="Early Beta" />
<a href="https://github.com/tinyhumansai/openhuman/releases/latest"><img src="https://img.shields.io/github/v/release/tinyhumansai/openhuman?label=latest" alt="最新リリース" /></a>
<a href="https://github.com/tinyhumansai/openhuman/stargazers"><img src="https://img.shields.io/github/stars/tinyhumansai/openhuman?style=flat" alt="GitHub Stars" /></a>
<a href="./LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="ライセンス" /></a>
<a href="./README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="../LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="ライセンス" /></a>
<a href="../README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="./README.zh-CN.md"><img src="https://img.shields.io/badge/lang-简体中文-blue" alt="简体中文" /></a>
<a href="./README.ko.md"><img src="https://img.shields.io/badge/lang-한국어-blue" alt="한국어" /></a>
<a href="./README.de.md"><img src="https://img.shields.io/badge/lang-Deutsch-blue" alt="Deutsch" /></a>
@@ -64,7 +64,7 @@ irm https://raw.githubusercontent.com/tinyhumansai/openhuman/main/scripts/instal
<!-- TODO: translate (ja-JP) — English source mirrored from README.md so non-EN readers get the same install caveats. Please translate. -->
> **Linux:** the AppImage can crash on launch under Wayland (and on Arch-based distros with `sharun: Interpreter not found!`) — see [#2463](https://github.com/tinyhumansai/openhuman/issues/2463) for the cause and env-var workarounds.
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](./packages/arch/openhuman-bin/);
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](../packages/arch/openhuman-bin/);
once published, Arch users can install it with `yay -S openhuman-bin`.
<!-- /TODO -->
@@ -88,20 +88,20 @@ OpenHuman は、あなたの日常生活に統合されるよう設計された
## ソースからのコントリビュート
新しいコントリビューターの方は、まず [`CONTRIBUTING.md`](./CONTRIBUTING.md) で fork/PR ワークフローとローカル検証コマンドを確認してください。最短経路は以下のとおりです:
新しいコントリビューターの方は、まず [`CONTRIBUTING.md`](../CONTRIBUTING.md) で fork/PR ワークフローとローカル検証コマンドを確認してください。最短経路は以下のとおりです:
1. Git、Node.js 24+、pnpm 10.10.0、Rust 1.93.0(`rustfmt` + `clippy`)、CMake、Ninja、ripgrep、プラットフォーム向けデスクトップビルドの前提条件をインストールします。
2. リポジトリを fork してクローンし、`pnpm install` の前に `git submodule update --init --recursive` を実行して、ベンダー化された Tauri/CEF のソースを取得します。
3. ウェブのみの UI 作業には `pnpm dev` を、デスクトップシェルには `pnpm --filter openhuman-app dev:app` を使用し、PR を出す前に `pnpm typecheck``pnpm format:check``cargo check -p openhuman --lib` などの集中チェックを実行してください。
詳細なドキュメント: [アーキテクチャ](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [セットアップガイド](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [クラウドデプロイ](./gitbooks/features/cloud-deploy.md)。
詳細なドキュメント: [アーキテクチャ](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [セットアップガイド](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [クラウドデプロイ](../gitbooks/features/cloud-deploy.md)。
## コンテキストを数週間ではなく数分で
OpenHuman は、数分であなたのことを理解する初めてのエージェントハーネスです。[Karpathy 氏の LLM ナレッジベース](https://x.com/karpathy/status/2039805659525644595)にインスパイアされました。ほとんどのエージェントは冷えた状態から始まります。Hermes はあなたの作業を見て学習し、OpenClaw はプラグインがコンテキストを運び込むのを待ちます。いずれにせよ、エージェントがあなたのスタックを十分理解して本当に役立つようになるまで、数日から数週間を費やすことになります。
<p align="center">
<img src="./gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman のコンテキスト構築図">
<img src="../gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman のコンテキスト構築図">
</p>
> OpenHuman はあなたのすべてのドキュメント、メール、チャットを要約・圧縮し、エージェントがあなたについてすべてを覚えていられるメモリーグラフを作成します。
+8 -8
View File
@@ -1,7 +1,7 @@
<h1 align="center">OpenHuman</h1>
<p align="center">
<img src="./gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
<img src="../gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
</p>
<p align="center" style="display: inline-block">
@@ -29,7 +29,7 @@
</p>
<p align="center">
🇺🇸 <a href="./README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
🇺🇸 <a href="../README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
</p>
@@ -38,8 +38,8 @@
<img src="https://img.shields.io/badge/status-early%20beta-orange" alt="얼리 베타" />
<a href="https://github.com/tinyhumansai/openhuman/releases/latest"><img src="https://img.shields.io/github/v/release/tinyhumansai/openhuman?label=latest" alt="최신 릴리스" /></a>
<a href="https://github.com/tinyhumansai/openhuman/stargazers"><img src="https://img.shields.io/github/stars/tinyhumansai/openhuman?style=flat" alt="GitHub Stars" /></a>
<a href="./LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="라이선스" /></a>
<a href="./README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="../LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="라이선스" /></a>
<a href="../README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="./README.zh-CN.md"><img src="https://img.shields.io/badge/lang-简体中文-blue" alt="简体中文" /></a>
<a href="./README.ja-JP.md"><img src="https://img.shields.io/badge/lang-日本語-blue" alt="日本語" /></a>
<a href="./README.ko.md"><img src="https://img.shields.io/badge/lang-한국어-blue" alt="한국어" /></a>
@@ -65,7 +65,7 @@ irm https://raw.githubusercontent.com/tinyhumansai/openhuman/main/scripts/instal
<!-- TODO: translate (ko) — English source mirrored from README.md so non-EN readers get the same install caveats. Please translate. -->
> **Linux:** the AppImage can crash on launch under Wayland (and on Arch-based distros with `sharun: Interpreter not found!`) — see [#2463](https://github.com/tinyhumansai/openhuman/issues/2463) for the cause and env-var workarounds.
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](./packages/arch/openhuman-bin/);
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](../packages/arch/openhuman-bin/);
once published, Arch users can install it with `yay -S openhuman-bin`.
<!-- /TODO -->
@@ -87,20 +87,20 @@ OpenHuman은 일상 생활에 통합되도록 설계된 오픈 소스 에이전
## 소스에서 기여하기
새로운 기여자인가요? 포크/PR 워크플로우 및 로컬 검증 명령에 대해서는 [`CONTRIBUTING.md`](./CONTRIBUTING.md)에서 시작하세요. 빠른 경로는 다음과 같습니다.
새로운 기여자인가요? 포크/PR 워크플로우 및 로컬 검증 명령에 대해서는 [`CONTRIBUTING.md`](../CONTRIBUTING.md)에서 시작하세요. 빠른 경로는 다음과 같습니다.
1. Git, Node.js 24+, pnpm 10.10.0, Rust 1.93.0(`rustfmt` + `clippy`), CMake, Ninja, ripgrep 및 플랫폼 데스크톱 빌드 필수 구성 요소를 설치합니다.
2. 저장소를 포크하고 클론한 다음, `pnpm install` 전에 `git submodule update --init --recursive`를 실행하여 벤더링된 Tauri/CEF 소스가 존재하는지 확인합니다.
3. 웹 전용 UI 작업에는 `pnpm dev`를, 데스크톱 쉘에는 `pnpm --filter openhuman-app dev:app`을 사용하고, PR을 열기 전에 `pnpm typecheck`, `pnpm format:check`, `cargo check -p openhuman --lib`와 같은 집중 점검을 수행합니다.
상세 문서: [아키텍처](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [설정하기](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [클라우드 배포](./gitbooks/features/cloud-deploy.md).
상세 문서: [아키텍처](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [설정하기](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [클라우드 배포](../gitbooks/features/cloud-deploy.md).
## 몇 주가 아닌 몇 분 만에 구축되는 컨텍스트
OpenHuman은 몇 분 만에 당신을 알게 되는 최초의 에이전트 하네스입니다. [Karpathy의 LLM 지식 베이스](https://x.com/karpathy/status/2039805659525644595)에서 영감을 받았습니다. 대부분의 에이전트는 아무런 정보 없이 시작합니다. Hermes는 당신의 작업을 지켜보며 학습하고, OpenClaw는 플러그인이 컨텍스트를 가져오기를 기다립니다. 어느 쪽이든 에이전트가 당신의 스택에 대해 충분히 알고 정말 유용해지기까지는 며칠 또는 몇 주가 걸립니다.
<p align="center">
<img src="./gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman 컨텍스트 구축 다이어그램">
<img src="../gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman 컨텍스트 구축 다이어그램">
</p>
> OpenHuman은 당신의 모든 문서, 이메일 및 채팅을 요약하고 압축합니다. 그리고 에이전트가 당신에 대한 모든 것을 기억할 수 있도록 메모리 그래프를 생성합니다.
+8 -8
View File
@@ -1,7 +1,7 @@
<div dir="rtl" lang="ur">
<p align="center">
🇺🇸 <a href="./README.md">انگریزی</a> | 🇨🇳 <a href="./README.zh-CN.md">چینی (آسان)</a> | 🇯🇵 <a href="./README.ja-JP.md">جاپانی</a> | 🇰🇷 <a href="./README.ko.md">کورین</a> | 🇩🇪 <a href="./README.de.md">جرمن</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
🇺🇸 <a href="../README.md">انگریزی</a> | 🇨🇳 <a href="./README.zh-CN.md">چینی (آسان)</a> | 🇯🇵 <a href="./README.ja-JP.md">جاپانی</a> | 🇰🇷 <a href="./README.ko.md">کورین</a> | 🇩🇪 <a href="./README.de.md">جرمن</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
</p>
</div>
@@ -11,7 +11,7 @@
<h1 align="center">OpenHuman</h1>
<p align="center">
<img src="./gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
<img src="../gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
</p>
<p align="center" style="display: inline-block">
@@ -48,8 +48,8 @@
<img src="https://img.shields.io/badge/status-early%20beta-orange" alt="ابتدائی آزمائشی نسخہ" />
<a href="https://github.com/tinyhumansai/openhuman/releases/latest"><img src="https://img.shields.io/github/v/release/tinyhumansai/openhuman?label=latest" alt="تازہ ترین نسخہ" /></a>
<a href="https://github.com/tinyhumansai/openhuman/stargazers"><img src="https://img.shields.io/github/stars/tinyhumansai/openhuman?style=flat" alt="ستارے" /></a>
<a href="./LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="لائسنس" /></a>
<a href="./README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="../LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="لائسنس" /></a>
<a href="../README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="./README.ja-JP.md"><img src="https://img.shields.io/badge/lang-日本語-blue" alt="日本語" /></a>
<a href="./README.ko.md"><img src="https://img.shields.io/badge/lang-한국어-blue" alt="한국어" /></a>
<a href="./README.de.md"><img src="https://img.shields.io/badge/lang-Deutsch-blue" alt="Deutsch" /></a>
@@ -108,7 +108,7 @@ sudo apt-get install -y openhuman
<div dir="rtl" lang="ur">
**لینکس (آرچ — اے یو آر):** ریپو میں [`openhuman-bin` AUR recipe](./packages/arch/openhuman-bin/) موجود ہے۔ شائع ہونے کے بعد، آرچ صارفین `yay -S openhuman-bin` سے انسٹال کر سکتے ہیں۔
**لینکس (آرچ — اے یو آر):** ریپو میں [`openhuman-bin` AUR recipe](../packages/arch/openhuman-bin/) موجود ہے۔ شائع ہونے کے بعد، آرچ صارفین `yay -S openhuman-bin` سے انسٹال کر سکتے ہیں۔
**ونڈوز:** [تازہ ترین ریلیز](https://github.com/tinyhumansai/openhuman/releases/latest) سے سائنڈ `.msi` ڈاؤن لوڈ کریں اور چلائیں۔
@@ -160,13 +160,13 @@ OpenHuman ایک اوپن سورس ایجنٹک اسسٹنٹ ہے جو آپ کی
## سورس سے تعاون
نیا تعاون کنندہ؟ fork/PR ورک فلو اور مقامی تصدیقی کمانڈز کے لیے [`CONTRIBUTING.md`](./CONTRIBUTING.md) سے شروع کریں۔ مختصر راستہ:
نیا تعاون کنندہ؟ fork/PR ورک فلو اور مقامی تصدیقی کمانڈز کے لیے [`CONTRIBUTING.md`](../CONTRIBUTING.md) سے شروع کریں۔ مختصر راستہ:
1. Git، Node.js 24+، pnpm 10.10.0، Rust 1.93.0 (`rustfmt` + `clippy`)، CMake، Ninja، ripgrep، اور پلیٹ فارم ڈیسک ٹاپ بلڈ کی ضروریات انسٹال کریں۔
2. ریپو کو fork اور کلون کریں، پھر `pnpm install` سے پہلے `git submodule update --init --recursive` چلائیں تاکہ وینڈرڈ Tauri/CEF سورس موجود ہوں۔
3. ویب صرف UI کام کے لیے `pnpm dev`، ڈیسک ٹاپ شیل کے لیے `pnpm --filter openhuman-app dev:app`، اور PR کھولنے سے پہلے فوکسڈ چیکس جیسے `pnpm typecheck`، `pnpm format:check`، اور `cargo check -p openhuman --lib` استعمال کریں۔
مزید دستاویزات: [آرکیٹیکچر](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [سیٹ اپ](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [کلاؤڈ ڈیپلائے](./gitbooks/features/cloud-deploy.md)۔
مزید دستاویزات: [آرکیٹیکچر](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [سیٹ اپ](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [کلاؤڈ ڈیپلائے](../gitbooks/features/cloud-deploy.md)۔
## منٹوں میں سیاق و سباق، ہفتوں میں نہیں
@@ -177,7 +177,7 @@ OpenHuman پہلا ایجنٹ ہارنس ہے جو منٹوں میں آپ کو
<div dir="ltr">
<p align="center">
<img src="./gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman سیاق و سباق بنانے کا خاکہ" />
<img src="../gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman سیاق و سباق بنانے کا خاکہ" />
</p>
</div>
+8 -8
View File
@@ -1,11 +1,11 @@
<p align="center">
🇺🇸 <a href="./README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
🇺🇸 <a href="../README.md">English</a> | 🇨🇳 <a href="./README.zh-CN.md">简体中文</a> | 🇯🇵 <a href="./README.ja-JP.md">日本語</a> | 🇰🇷 <a href="./README.ko.md">한국어</a> | 🇩🇪 <a href="./README.de.md">Deutsch</a> | 🇵🇰 <a href="./README.ur-pk.md">اردو</a>
</p>
<h1 align="center">OpenHuman</h1>
<p align="center">
<img src="./gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
<img src="../gitbooks/.gitbook/assets/demo.png" alt="The Tet" />
</p>
<p align="center" style="display: inline-block">
@@ -36,8 +36,8 @@
<img src="https://img.shields.io/badge/status-early%20beta-orange" alt="早期测试版" />
<a href="https://github.com/tinyhumansai/openhuman/releases/latest"><img src="https://img.shields.io/github/v/release/tinyhumansai/openhuman?label=latest" alt="最新版本" /></a>
<a href="https://github.com/tinyhumansai/openhuman/stargazers"><img src="https://img.shields.io/github/stars/tinyhumansai/openhuman?style=flat" alt="GitHub Stars" /></a>
<a href="./LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="许可证" /></a>
<a href="./README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="../LICENSE"><img src="https://img.shields.io/github/license/tinyhumansai/openhuman" alt="许可证" /></a>
<a href="../README.md"><img src="https://img.shields.io/badge/lang-English-blue" alt="English" /></a>
<a href="./README.ja-JP.md"><img src="https://img.shields.io/badge/lang-日本語-blue" alt="日本語" /></a>
<a href="./README.ko.md"><img src="https://img.shields.io/badge/lang-한국어-blue" alt="한국어" /></a>
<a href="./README.de.md"><img src="https://img.shields.io/badge/lang-Deutsch-blue" alt="Deutsch" /></a>
@@ -62,7 +62,7 @@ irm https://raw.githubusercontent.com/tinyhumansai/openhuman/main/scripts/instal
<!-- TODO: translate (zh-CN) — English source mirrored from README.md so non-EN readers get the same install caveats. Please translate. -->
> **Linux:** the AppImage can crash on launch under Wayland (and on Arch-based distros with `sharun: Interpreter not found!`) — see [#2463](https://github.com/tinyhumansai/openhuman/issues/2463) for the cause and env-var workarounds.
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](./packages/arch/openhuman-bin/);
Arch Linux package maintainers can use the [`openhuman-bin` AUR recipe](../packages/arch/openhuman-bin/);
once published, Arch users can install it with `yay -S openhuman-bin`.
<!-- /TODO -->
@@ -86,20 +86,20 @@ OpenHuman 是一个开源智能助手,旨在融入你的日常生活。以下
## 从源码贡献
新贡献者?从 [`CONTRIBUTING.md`](./CONTRIBUTING.md) 了解 fork/PR 工作流和本地验证命令。快速路径:
新贡献者?从 [`CONTRIBUTING.md`](../CONTRIBUTING.md) 了解 fork/PR 工作流和本地验证命令。快速路径:
1. 安装 Git、Node.js 24+、pnpm 10.10.0、Rust 1.93.0`rustfmt` + `clippy`)、CMake、Ninja、ripgrep,以及各平台桌面构建的前置依赖。
2. Fork 并克隆仓库,然后运行 `git submodule update --init --recursive` 之后再执行 `pnpm install`,确保内置的 Tauri/CEF 源码就位。
3. 使用 `pnpm dev` 进行纯 Web UI 开发,`pnpm --filter openhuman-app dev:app` 用于桌面壳,在提交 PR 之前运行针对性的检查如 `pnpm typecheck``pnpm format:check``cargo check -p openhuman --lib`
更多文档:[架构](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [环境搭建](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [云端部署](./gitbooks/features/cloud-deploy.md)。
更多文档:[架构](https://tinyhumans.gitbook.io/openhuman/developing/architecture) · [环境搭建](https://tinyhumans.gitbook.io/openhuman/developing/getting-set-up) · [云端部署](../gitbooks/features/cloud-deploy.md)。
## 几分钟内建立上下文,而非数周
OpenHuman 是首个能在几分钟内了解你的智能体框架。灵感来源于 [Karpathy 的 LLM 知识库](https://x.com/karpathy/status/2039805659525644595)。大多数智能体从零开始——Hermes 通过观察你的工作来学习;OpenClaw 等待插件输送上下文。无论哪种方式,你都需要花费数天甚至数周时间,智能体才能对你的技术栈有足够的了解从而真正发挥作用。
<p align="center">
<img src="./gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman 上下文构建示意图" />
<img src="../gitbooks/.gitbook/assets/image (1).png" alt="OpenHuman 上下文构建示意图" />
</p>
> OpenHuman 将你的所有文档、邮件和聊天记录进行摘要和压缩,并创建一个记忆图谱,让你的智能体记住关于你的一切。
-30
View File
@@ -1,30 +0,0 @@
---
description: >-
你的桌面级个人 AI 助手。连接 118+ 服务,构建本地优先的记忆树,
自我反思,并可通过音频和视频与你交互。
icon: diamond
---
# 欢迎使用 OpenHuman
<figure><img src=".gitbook/assets/demo.png" alt=""><figcaption></figcaption></figure>
OpenHuman 是一款开源 AI 助手,旨在成为你跨工具协作时的**记忆**与**执行者**。基于 Rust + Tauri 构建,采用 GNU GPL3 许可证,它弥合了 AI 模型能做什么与它们实际上了解**你**多少之间的差距。
世界上所有的模型,200 多个,都面临同一个根本性限制:它们是无状态的。你输入一段提示,得到回复,然后上下文就消失了。即便那些自称有"记忆"的模型,也只存储了几条要点。几条要点是便利贴,不是智能。
OpenHuman 通过一套冷静、刻意不同的技术栈解决了这个问题:
* **本地优先的** [**记忆树**](features/obsidian-wiki/memory-tree.zh-CN.md)**。** 你连接的每一个来源——Gmail、Slack、GitHub、Notion、你自己的笔记——都会流经一个确定性流水线:规范 Markdown、≤3k token 的块、评分、折叠成按来源 / 按主题 / 按日期的摘要树。存储在你机器上的 SQLite 中。没有向量黑盒。
* **其上的** [**Obsidian 风格 Wiki**](features/obsidian-wiki/)**。智能体用于推理的同样的块,会作为 `.md` 文件落在一个你可以用 [Obsidian](https://obsidian.md) 打开的仓库中,手动浏览、编辑和链接。灵感来自 [Karpathy 的 obsidian-wiki 工作流](https://x.com/karpathy/status/2039805659525644595)。你无法信任一个你无法阅读的记忆。
* [**118+ 第三方集成**](features/integrations/README.zh-CN.md)**。** 一键 OAuth 接入 Gmail、GitHub、Slack、Notion、Stripe、Calendar、Drive、Linear、Jira 等——无需手动配置 API key,无需在插件市场中翻找。
* [**自动拉取**](features/obsidian-wiki/auto-fetch.zh-CN.md)**。** 每二十分钟,OpenHuman 会从每一个活跃连接中拉取最新数据,并在你无需开口的情况下将其折叠进记忆树,这样智能体在早上就已经拥有了明天的上下文。
* **为大数据而生的智能体。** [智能 Token 压缩(TokenJuice](features/token-compression.zh-CN.md) 在冗长的工具输出进入模型上下文之前对其进行压缩,因此横扫过去六个月邮件的成本仅为个位数美元。[自动模型路由](features/model-routing/) 将每个任务发送给合适的模型——`hint:reasoning` 交给前沿模型,`hint:fast` 交给廉价模型,视觉任务交给视觉模型——全部在一个订阅下完成。可选的 [本地 AI(通过 Ollama 或 LM Studio](features/model-routing/local-ai.zh-CN.md) 让支持的负载保留在设备上。
* [**开箱即用**](features/native-tools/)**。一套完整的智能体工具链默认已接入:[网页搜索](features/native-tools/web-search.zh-CN.md)、[网页抓取](features/native-tools/web-scraper.zh-CN.md)、全套[编程工具集](features/native-tools/coder.zh-CN.md)(文件系统、git、lint、test、grep)、[浏览器与电脑控制](features/native-tools/browser-and-computer.zh-CN.md)、[定时任务与调度](features/native-tools/cron.zh-CN.md)、[记忆工具](features/native-tools/memory-tools.zh-CN.md)、用于生成子智能体的[智能体协调](features/native-tools/agent-coordination.zh-CN.md),以及[原生语音](features/native-tools/voice.zh-CN.md)——STT 输入、TTS 输出、吉祥物口型同步,还有一个实时 Google Meet 智能体,它可以加入会议、将会议转录进你的记忆树,并在通话中回话。没有"装个插件才能读文件"的摩擦。
* **简洁,UI 优先。** 干净的桌面体验和简短的引导路径,让你在几次点击内从安装到拥有一个可用的智能体——无需先配 config,无需终端。智能体[有一张脸](features/mascot/README.zh-CN.md):一个桌面吉祥物,会说话、对周围环境作出反应、作为真实参与者加入你的 Google Meet、在数周内记住你,甚至在你停止打字后仍在后台思考。
这些特性合在一起,让 OpenHuman 从根本上不同于聊天机器人。它是一个能够低成本消费大量个人数据、对你的世界保持持久且不断演进的理解、并代表你采取主动行动的 AI 智能体。
{% hint style="warning" %}
OpenHuman 不是 AGI。但它在更好的记忆、更好的编排和更好的工具链方面,是一个有意义的架构进步。
{% endhint %}
+21 -3
View File
@@ -10,12 +10,22 @@
* [Realtime Mascot](features/mascot/README.md)
* [Meeting Agents](features/mascot/meeting-agents.md)
* [Obsidian-Style Memory](features/obsidian-wiki/README.md)
* [Memory Trees](features/obsidian-wiki/memory-tree.md)
* [agentmemory backend](features/obsidian-wiki/agentmemory-backend.md)
* [Memory](features/obsidian-wiki/README.md)
* [Memory Tree](features/obsidian-wiki/memory-tree.md)
* [Memory Sources & Scoping](features/obsidian-wiki/sources.md)
* [Auto-fetch from Integrations](features/obsidian-wiki/auto-fetch.md)
* [Scoring & Ranking](features/obsidian-wiki/scoring.md)
* [Retrieval & Recall](features/obsidian-wiki/retrieval.md)
* [Memory Diff (Git-Backed)](features/obsidian-wiki/memory-diff.md)
* [agentmemory backend](features/obsidian-wiki/agentmemory-backend.md)
* [Third-party Integrations (118+)](features/integrations/README.md)
* [Triggers](features/integrations/triggers.md)
* [MCP Servers & Skills](features/integrations/mcp-and-skills.md)
* [Messaging Channels](features/channels.md)
* [SuperContext](features/super-context.md)
* [Goals & Todos](features/goals-and-todos.md)
* [Personalization & Self-Learning](features/personalization.md)
* [Themes & Theme Studio](features/theming.md)
* [Smart Token Compression](features/token-compression.md)
* [Automatic Model Routing](features/model-routing/README.md)
* [Local AI (optional)](features/model-routing/local-ai.md)
@@ -33,6 +43,13 @@
* [Agent Coordination](features/native-tools/agent-coordination.md)
* [System & Utilities](features/native-tools/system-and-utilities.md)
* [Subconscious Loop](features/subconscious.md)
* [Notifications & Activity](features/notifications-and-activity.md)
* [Screen Intelligence](features/screen-intelligence.md)
* [Wallet](features/wallet.md)
* [Billing, Cost & Usage](features/billing-and-usage.md)
* [Rewards & Referrals](features/rewards-and-referrals.md)
* [iOS Companion](features/ios-companion.md)
* [Approval Gate](features/approval-gate.md)
* [Privacy & Security](features/privacy-and-security.md)
* [OS Keyring & Secret Storage](features/os-keyring-and-secret-storage.md)
* [Platform & Availability](features/platform.md)
@@ -48,6 +65,7 @@
* [Release Policy](developing/release-policy.md)
* [Polymarket Integration (v1 Read + Trading)](developing/integrations/polymarket.md)
* [Chromium Embedded Framework](developing/cef.md)
* [Theming (Token System)](developing/theming.md)
* [Agent Observability](developing/agent-observability.md)
* [Architecture](developing/architecture/README.md)
* [Agent Harness](developing/architecture/agent-harness.md)
-58
View File
@@ -1,58 +0,0 @@
# 目录
## 概述
* [欢迎使用 OpenHuman](README.zh-CN.md)
* [快速开始](overview/getting-started.zh-CN.md)
* [登录故障排查](overview/troubleshooting-sign-in.zh-CN.md)
## 功能
* [实时吉祥物](features/mascot/README.zh-CN.md)
* [会议智能体](features/mascot/meeting-agents.zh-CN.md)
* [Obsidian 风格记忆](features/obsidian-wiki/README.zh-CN.md)
* [记忆树](features/obsidian-wiki/memory-tree.zh-CN.md)
* [agentmemory 后端](features/obsidian-wiki/agentmemory-backend.zh-CN.md)
* [从集成自动拉取](features/obsidian-wiki/auto-fetch.zh-CN.md)
* [第三方集成(118+](features/integrations/README.zh-CN.md)
* [触发器](features/integrations/triggers.zh-CN.md)
* [智能 Token 压缩](features/token-compression.zh-CN.md)
* [自动模型路由](features/model-routing/README.zh-CN.md)
* [本地 AI(可选)](features/model-routing/local-ai.zh-CN.md)
* [可用工具](features/native-tools/README.zh-CN.md)
* [网页搜索](features/native-tools/web-search.zh-CN.md)
* [网页抓取](features/native-tools/web-scraper.zh-CN.md)
* [编程工具](features/native-tools/coder.zh-CN.md)
* [浏览器与电脑控制](features/native-tools/browser-and-computer.zh-CN.md)
* [定时任务与调度](features/native-tools/cron.zh-CN.md)
* [语音](features/native-tools/voice.zh-CN.md)
* [记忆工具](features/native-tools/memory-tools.zh-CN.md)
* [工具级记忆](features/native-tools/tool-memory.zh-CN.md)
* [第三方集成](features/native-tools/integrations.zh-CN.md)
* [智能体协调](features/native-tools/agent-coordination.zh-CN.md)
* [系统与工具](features/native-tools/system-and-utilities.zh-CN.md)
* [潜意识循环](features/subconscious.zh-CN.md)
* [隐私与安全](features/privacy-and-security.zh-CN.md)
* [平台与可用性](features/platform.zh-CN.md)
* [云部署](features/cloud-deploy.zh-CN.md)
## 开发
* [概述](developing/README.zh-CN.md)
* [环境搭建](developing/getting-set-up.zh-CN.md)
* [构建 Rust 核心](developing/building-rust-core.zh-CN.md)
* [测试策略](developing/testing-strategy.zh-CN.md)
* [E2E 测试](developing/e2e-testing.zh-CN.md)
* [发布策略](developing/release-policy.zh-CN.md)
* [Polymarket 集成(v1 读取 + 交易)](developing/integrations/polymarket.zh-CN.md)
* [Chromium Embedded Framework](developing/cef.zh-CN.md)
* [智能体可观测性](developing/agent-observability.zh-CN.md)
* [架构](developing/architecture/README.zh-CN.md)
* [Agent Harness](developing/architecture/agent-harness.zh-CN.md)
* [前端(app/src/](developing/architecture/frontend.zh-CN.md)
* [Tauri 壳层(app/src-tauri/](developing/architecture/tauri-shell.zh-CN.md)
## 法律
* [使用条款](legal/terms-of-use.zh-CN.md)
* [隐私政策](legal/privacy-policy.zh-CN.md)
-75
View File
@@ -1,75 +0,0 @@
---
description: 从源码构建、运行、测试和发布 OpenHuman。
icon: code-branch
lang: zh-CN
---
# 概览
OpenHuman 在 [github.com/tinyhumansai/openhuman](https://github.com/tinyhumansai/openhuman) 以 GPLv3 协议开源。本节面向贡献者和所有从源码运行 OpenHuman 的人。
如果你只是想使用应用,请前往[快速开始](../overview/getting-started.zh-CN.md)。如果你来这里是为了阅读架构文档、hack 一个新特性,或者提交一个 PR,那你来对地方了。
***
## 代码结构
| 路径 | 内容 |
| ---- | ---- |
| `app/` | pnpm workspace `openhuman-app`。Vite + React 前端(`app/src/`)和 Tauri 桌面宿主(`app/src-tauri/`)。 |
| `src/` | Rust 库 crate `openhuman`,并包含 `openhuman-core` CLI 二进制文件。领域逻辑、JSON-RPC、MCP 路由。 |
| `gitbooks/` | 本站(面向公众的文档)。 |
| `docs/` | 尚未迁移到 GitBook 的深层参考资料(记忆流水线图、智能体流程等)。 |
仓库根目录的 `CLAUDE.md` 是给在该代码库上工作的 AI 智能体的权威参考。人类也适用同样的规则。
***
## 从这里开始
如果你是第一次拉取仓库:
1. [**环境搭建**](getting-set-up.zh-CN.md)。工具链、依赖、vendored Tauri CLI、sidecar staging —— 让 `pnpm dev` 真正跑起来所需的一切。
2. [**构建 Rust 核心**](building-rust-core.zh-CN.md)。仅针对仓库根目录 Rust crate 的新机搭建:固定工具链、OS 包,以及精确的 `cargo` 命令。
3. [**架构**](architecture.zh-CN.md)。桌面应用、Rust 核心 sidecar、JSON-RPC 桥接,以及双 socket 如何协同工作。在做非平凡改动之前先读这个。
4. [**前端**](architecture/frontend.zh-CN.md) 和 [**Tauri 壳层**](architecture/tauri-shell.zh-CN.md)。React 应用,以及包裹它的桌面宿主。
5. [**MCP 服务器**](mcp-server.zh-CN.md)。可选的 stdio MCP 模式,将只读的 OpenHuman 记忆工具暴露给本地客户端。
***
## 测试
OpenHuman 有三层测试。知道你的改动属于哪一层:
* [**测试策略**](testing-strategy.zh-CN.md)。什么时候写 Vitest、什么时候写 cargo tests、什么时候写 WDIO。
* [**E2E 测试**](e2e-testing.zh-CN.md)。WDIO/Appium spec、双平台设置(Linux tauri-driver、macOS Appium Mac2),以及如何在本地运行单个 spec。
* [**智能体可观测性**](agent-observability.zh-CN.md)。让 E2E 和智能体运行事后可调试的工件捕获层。
PR 必须通过 **变更行覆盖率 ≥ 80%** 的门禁。为新行为添加测试,不要只测 happy path。
***
## 发布
* [**发布策略**](release-policy.zh-CN.md)。版本策略、发布节奏、OAuth + 安装包规则。
* [**云端部署**](../features/cloud-deploy.md)。当变更跨越桌面边界时,后端/云端侧的部署。
***
## 深入探索
* [**Agent Harness**](architecture/agent-harness.zh-CN.md)。智能体面向代码的工具表面,以及如何扩展它。
* [**Chromium Embedded Framework**](cef.zh-CN.md)。嵌入式提供商 webview 如何工作、为什么不运行注入的 JS,以及各提供商 scanner 实际上做了什么。
对于仍在构建中的特性,[Subconscious Loop](../features/subconscious.zh-CN.md) 页面从头到尾涵盖了后台任务评估系统。
***
## 贡献
* 在 [tinyhumansai/openhuman](https://github.com/tinyhumansai/openhuman) 提交 issue 和 PR。
* PR 目标分支为 `main`。推送到你的 fork,不要推 upstream。
* 遵循 [`CONTRIBUTING.md`](../../CONTRIBUTING.md) 和 issue/PR 模板。
* 保持改动聚焦。一个 bug fix 不需要附带周边清理;一个一次性操作不需要 helper。
帮助构建 AGI 并不意味着一定要提交内核代码 —— bug 修复、文档、集成和测试都在推动进展。
@@ -1,81 +0,0 @@
---
description: 使 E2E 测试可调试的工件捕获层。日志、跟踪、截图。
icon: eye
---
# E2E 的 Agent 可观测性
本文档描述了使桌面应用可通过现有 WDIO/Appium/tauri-driver harness 被编码智能体(Codex、Claude Code、Cursor)检查的工件捕获层。
它有意保持精简:一个规范的 onboarding + 隐私流程,包含磁盘截图、页面源码 dump 和 mock 后端请求日志。更广泛的计划见仓库根目录的 `AGENT_OBSERVABILITY_PLAN.md`
## TL;DR
```bash
bash app/scripts/e2e-agent-review.sh
```
工件落在:
```text
app/test/e2e/artifacts/<ISO-timestamp>-agent-review/
01-welcome.png
01-welcome.source.xml
02-post-welcome.png
02-post-welcome.source.xml
03-post-onboarding.png
03-post-onboarding.source.xml
04-privacy-panel.png
04-privacy-panel.source.xml
mock-requests-after-welcome.json
mock-requests-after-onboarding.json
mock-requests-after-privacy.json
failure-<test>.png # 仅在失败时
failure-<test>.source.xml # 仅在失败时
meta.json # 运行元数据 + 检查点索引
```
脚本最后会打印解析后的工件目录。
## 组成部分
| 组件 | 路径 | 作用 |
|-------|------|------|
| 辅助函数 | `app/test/e2e/helpers/artifacts.ts` | 运行目录、`captureCheckpoint``captureFailureArtifacts``saveMockRequestLog` |
| WDIO hook | `app/test/wdio.conf.ts` (`afterTest`) | 任何失败测试都会 dump 截图 + 源码 |
| 规范 spec | `app/test/e2e/specs/agent-review.spec.ts` | Welcome → onboarding → 隐私面板,带命名检查点 |
| Wrapper 脚本 | `app/scripts/e2e-agent-review.sh` | 构建 + 运行 + 打印工件目录 |
| 稳定选择器 | `OnboardingNextButton``Onboarding` 遮罩层 + 跳过按钮、`WelcomeStep``PrivacyPanel` 上的 `data-testid` | 智能体可靠的导航锚点 |
## 环境覆盖
| 变量 | 效果 |
|----------|--------|
| `E2E_ARTIFACT_DIR` | 强制指定运行目录(跳过自动时间戳命名) |
| `E2E_ARTIFACT_ROOT` | 自动生成运行目录的父目录(默认:`app/test/e2e/artifacts` |
| `E2E_ARTIFACT_LABEL` | 自动生成的运行目录名中使用的标签(默认:`run`wrapper 设为 `agent-review` |
## 在新 spec 中使用辅助函数
```ts
import {
captureCheckpoint,
saveMockRequestLog,
} from '../helpers/artifacts';
import { getRequestLog } from '../mock-server';
await captureCheckpoint('after-connect-click');
saveMockRequestLog('after-connect-click', getRequestLog());
```
`captureCheckpoint` 会对捕获进行编号,使运行目录按时间顺序阅读。
`captureFailureArtifacts` 已接入 `wdio.conf.ts`,在任何失败测试中自动触发,spec 不应直接调用它。
## 有意排除的范围
- 跨每个组件状态的视觉基线 / 图像差异。
- 每次点击都截图(太吵)。
- 实时集成(Gmail、Notion、Telegram);仅 mock 服务器。
- 新测试框架 / reporter。
仅在证明此循环有效后才扩展到更多流程。
-353
View File
@@ -1,353 +0,0 @@
---
description: OpenHuman 代码库的深度架构参考 —— 仓库布局、运行时范围、双 socket 同步、RPC 流程。
icon: code-branch
lang: zh-CN
---
# OpenHuman 架构
**基于 Rust 构建的加密社区 AI 超级助手。**
OpenHuman 是一款为加密货币生态系统量身打造的跨平台通信与自动化平台。单一的 React + Rust(Tauri)代码库可以面向多个平台;**我们目前为用户文档和发布的仅是桌面端** —— **Windows、macOS 和 Linux**。Android、iOS 和 Web **尚未**在当前文档或发布中支持。技术栈包括一个托管的 Node.js 运行时,用于支持工具能力的技能;持久化的 Rust 原生 WebSocket 基础设施;以及一个 AI 工具协议,让语言模型实时调用任何已连接的服务。
---
## 仓库布局(monorepo
| 路径 | 内容 |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **`app/`** | Yarn workspace **`openhuman-app`**Vite/React UI`app/src/`)、Tauri 壳层(`app/src-tauri/`)、Vitest 测试 |
| **仓库根目录 `src/`** | Rust **`openhuman_core`** 库 + **`openhuman-core`** CLI 二进制文件 —— 核心服务器、JSON-RPC、一等 JavaScript 运行时(`src/openhuman/javascript/`),由托管的 Node.js 实现驱动、频道、内存等 |
| **`Cargo.toml`**(根目录) | 构建 `openhuman-core` 二进制文件(`cargo build --bin openhuman-core`),staging 到 `app/src-tauri/binaries/` 以供桌面打包 |
| **`skills/`** | 运行时消耗的技能包 |
| **`docs/`** | 本书 + 每棵树指南(`docs/src/``docs/src-tauri/` |
桌面应用 **WebView**`app/` 加载 UI;繁重的 RPC 和技能在 **`openhuman-core`** 进程中运行,可通过 HTTP 从 Tauri 主机访问(`core_rpc_relay`)。
---
## 平台覆盖范围
**今天支持的(终端用户):** 桌面端。Windows、macOS、Linux(原生安装包)。
**尚未支持:** Android、iOS、独立 Web 客户端(仓库中可能以实验性目标存在;不要视为产品就绪)。
```text
OpenHuman(已发布)
|
Desktop
/ | \
Windows macOS Linux
x64 x64 x64
ARM64 ARM64 ARM64
```
Tauri v2 将 Rust 核心编译为每个平台的原生二进制文件,将 React 前端作为轻量级 WebView 嵌入。桌面构建产出 `.dmg``.msi``.AppImage``.deb` 安装包。额外目标(移动端、Web)在明确文档化支持之前均超出范围。
---
## 高层架构
```text
+------------------------------------------------------------------+
| React 前端 |
| Redux Toolkit | Socket.io 客户端 | MCP 传输层 | UI |
+------------------------------------------------------------------+
| Tauri IPC 桥接 |
+------------------------------------------------------------------+
| Rust 核心引擎 |
| |
| +------------------+ +------------------+ +-----------------+ |
| | QuickJS 技能 | | Socket 管理器 | | AI 加密 | |
| | 运行时引擎 | | (持久化 WS) | | & 内存存储 | |
| +------------------+ +------------------+ +-----------------+ |
| |
| +------------------+ +------------------+ +-----------------+ |
| | 技能注册表 | | Cron 调度器 | | 会话 & 认证 | |
| | & 桥接 API | | (5s tick 循环) | | 管理 | |
| +------------------+ +------------------+ +-----------------+ |
| |
| +------------------+ +------------------+ +-----------------+ |
| | Telegram | | SQLite 存储 | | OS 钥匙串 | |
| | 集成 | | (rusqlite) | | 集成 | |
| +------------------+ +------------------+ +-----------------+ |
+------------------------------------------------------------------+
|
+-----------+-----------+
| |
后端服务 外部 API
(Socket.io 服务器) (Telegram 等)
```
前端通过两种方式与 **openhuman** Rust 核心通信:用于一小部分壳层命令的 **Tauri IPC**(窗口、AI 文件辅助函数、**`core_rpc_relay`**),以及用于业务逻辑和技能的 **HTTP JSON-RPC**。核心拥有持久连接(如适用)、内存/功能的加密工作,以及 **QuickJS** 沙盒化技能执行。
---
## Rust 驱动的性能
OpenHuman 选择 Tauri + Rust 而非 Electron,基于根本的性能和安全原因:
| 指标 | OpenHumanTauri + Rust | 典型 Electron 应用 |
| ------------------------- | -------------------------------------------------------- | ---------------------------- |
| 二进制体积 | 取决于功能(CEF 运行时 + 技能包占主导) | ~150 MB+ |
| 每技能上下文内存 | ~1-2 MBQuickJS | ~150 MB+Chromium 渲染器) |
| 冷启动 | 亚 500ms | 2-5 秒 |
| 垃圾回收暂停 | 无(Rust 所有权模型) | V8 GC 暂停 |
| 内存安全 | 编译期保证 | 运行时异常 |
| TLS 实现 | rustls(无 OpenSSL 依赖) | Chromium 的 BoringSSL |
**这对加密平台为何重要**:交易员和分析师在运行 OpenHuman 的同时,还会运行资源密集型工具、图表软件、多个浏览器标签、交易终端。原生二进制文件加上亚 500ms 启动意味着应用感觉像原生应用,不会碍事。零 GC 暂停意味着实时价格推送和警报永远不会因内存管理而延迟。
**Tokio 异步运行时**驱动所有 I/O。WebSocket 连接、HTTP 请求、文件操作和技能间通信,都是线程池上的非阻塞任务。数千个并发操作(技能执行、cron job、socket 事件)共享一小套固定的 OS 线程。
---
## 实时 Socket 基础设施
OpenHuman 实现了**双 socket 架构**:桌面端使用 Rust 原生 WebSocket 客户端,Web 端使用 JavaScript Socket.io 客户端。Rust 实现能在应用后台存活,独立于 WebView 运行,并通过 rustls 处理 TLS。
```text
桌面模式: Web 模式:
+-------------+ +-------------+
| React UI | | React UI |
+------+------+ +------+------+
| Tauri IPC | Direct
+------+------+ +------+------+
| Rust Socket | | JS Socket |
| Manager | | .io Client |
+------+------+ +------+------+
| tokio-tungstenite | Socket.io
| + rustls TLS | (websocket/polling)
+------+------+ +------+------+
| Backend | | Backend |
+-------------+ +-------------+
```
**Rust Socket 管理器**通过原始 WebSocket 实现 Engine.IO v4 + Socket.IO v4 帧:
- **握手**WebSocket 连接、Engine.IO OPEN(提取 `sid``pingInterval``pingTimeout`)、带 JWT 认证的 Socket.IO CONNECT、CONNECT ACK
- **保活**:响应 Engine.IO PING 以 PONG;超时阈值 = `pingInterval + pingTimeout + 5s`(默认:50 秒)
- **重连**:指数退避,从 1 秒到最大 30 秒。成功连接丢失后重置为 1s;如果连接从未建立则持续增长
- **CORS 绕过**Rust `reqwest` HTTP 客户端直接发起外部 API 调用,不受浏览器 CORS 限制
socket 连接在所有技能间**共享**。当事件到达时,socket 管理器通过异步消息通道将它们路由到相应的技能。这完全消除了每个技能的连接开销。
**`tool:sync` 协议**:每次 socket 连接和技能生命周期变化时,客户端都会发出一个 `tool:sync` 事件,包含可用工具的完整列表及其连接状态。这使后端 AI 系统能实时感知所有能力。
---
## 技能运行时引擎
OpenHuman 的决定性能力是其运行在 Rust 进程内部的**沙盒化 JavaScript 执行引擎**。技能是轻量级自动化脚本,通过自定义工具、集成和定时任务扩展平台。
```text
+---------------------------------------------------------------+
| RuntimeEngine |
| |
| +-------------------+ +-------------------+ |
| | SkillRegistry | | CronScheduler | |
| | (HashMap + MPSC) | | (5s tick loop) | |
| +--------+----------+ +--------+----------+ |
| | | |
| +--------v----------+ +--------v----------+ +----------+ |
| | JavaScript Layer | | runtime_node | | Bridge | |
| | skill metadata | | managed Node.js | | APIs | |
| | + prompt context | | system/bundled | +----+-----+ |
| | + tool discovery | | tool execution | | |
| +-------------------+ +-------------------+ | |
| | |
| +---------------------------------------------------v-----+ |
| | net | db | store | cron | log | tauri | | |
| | HTTP SQLite KV Schedule Log Platform| | |
| +------------------------------------------------------+ | |
+---------------------------------------------------------------+
```
**Node.js 运行时**:核心尽可能解析兼容的系统 `node`,否则将托管发行版安装到 OpenHuman 缓存中。技能主要暴露工具元数据,并使用运行时桥接来列出和执行工具,而非在核心内运行隔离的 QuickJS VM。
| 参数 | 值 |
| ---------------------- | ----- |
| 公共语言槽位 | `javascript` |
| 当前 JS 后端 | `runtime_node` |
| 托管 Node 版本 | 默认 `v22.11.0` |
| 运行时来源 | 系统 `node` 或托管安装 |
| 完整性验证 | 针对 `SHASUMS256.txt` 的 SHA-256 |
**工具桥架构**`SKILL.md` 包提供元数据、指令和可选的捆绑 JS 辅助函数。Rust 核心拥有权威的工具注册表,JavaScript 运行时桥接列出工具并将具名工具调用分派到核心或 Node-backed 辅助函数中。
**桥接 API** 向运行时桥接和 Node-backed 辅助函数暴露平台能力:
| 桥接 | 能力 |
| --------- | ----------------------------------------------------------- |
| **net** | 通过 `reqwest` 的 HTTP fetch(默认 30s 超时,所有方法) |
| **db** | 通过 `rusqlite` 的每个技能 SQLite 数据库 |
| **store** | 键值持久化 |
| **cron** | 定时注册(6 字段 cron 表达式) |
| **log** | 通过 Rust `log` crate 的结构化日志 |
| **tauri** | 平台检测、通知、白名单环境变量 |
**技能发现** 使用 `SKILL.md` 加上可选的捆绑资源:
| 字段 | 用途 |
| ------------------ | ------- |
| `name` | 人类可读的显示名称 |
| `description` | 触发/选择摘要 |
| `metadata.id` | 存在时的稳定技能 slug |
| `allowed-tools` | 工具允许列表指引 |
| 捆绑资源 | 脚本、参考、资源 |
技能从 GitHub 仓库同步并在运行时发现。执行不再建模为每个技能一个嵌入式 QuickJS VMJavaScript 行为通过共享运行时桥接流动。
**Cron 调度器**:一个 5 秒 tick 循环对照 UTC 时间检查所有已注册的调度,使用 `cron` crate 进行表达式解析。当调度触发时,调度器向技能的通道发送 `CronTrigger` 消息,调用技能的 `onCronTrigger()` 处理程序。
---
## AI & 工具协议(MCP
OpenHuman 实现了**模型上下文协议**,一个基于 Socket.io 的 JSON-RPC 2.0 层,让 AI 模型发现并由技能暴露的工具。
```text
用户提示
|
v
AI 模型(后端)
|
| 1. mcp:listTools --> 前端/Rust 聚合所有技能工具
| <-- 工具目录
|
| 2. 决定调用哪个工具
|
| 3. mcp:toolCall { skillId__toolName, arguments }
| |
| v
| Socket 管理器路由到技能注册表
| |
| v
| QuickJS 技能实例执行工具
| |
| v
| 桥接 API 调用(HTTP、DB 等)
| |
| <-- mcp:toolCallResponse { result }
|
v
AI 对用户的响应
```
**传输**:每次请求 30 秒超时,`mcp:` 事件前缀,请求 ID 在待处理响应映射中跟踪。工具名称以 `skillId__toolName` 命名空间化,以实现明确路由。
**工具同步**`tool:sync` 事件在每次 socket 连接和技能状态变化时广播完整的工具清单、技能 ID、名称、连接状态和工具列表。后端 AI 系统始终拥有可用能力的最新视图。
**AI 记忆系统**
| 功能 | 实现 |
| ------------------ | ------------------------------------------------------ |
| 静态加密 | 带 Argon2id 密钥派生的 AES-256-GCM |
| 分块 | 每块 512 token64 token 重叠 |
| 搜索 | 混合:70% 向量相似度 + 30% FTS5 全文 |
| 嵌入 | OpenAI `text-embedding-3-small` |
| 知识图谱 | 通过 REST API 的 Neo4j,用于实体关系 |
| 会话 | 带压缩和工具压缩的 JSONL 转录 |
记忆加密密钥通过 Argon2id 从用户凭证派生,确保记忆文件在未经认证的情况下不可读。混合搜索结合语义理解(向量相似度)和关键词精确度(SQLite FTS5)以实现可靠的召回。
---
## 安全架构
```text
+-------------------------------------------------------------------+
| 安全层 |
| |
| +------------------+ +------------------+ +------------------+ |
| | OS 钥匙串 | | AES-256-GCM | | 沙盒化 | |
| | (macOS/Win/Lin) | | 内存加密 | | QuickJS 每 | |
| | 用于凭证 | | + Argon2id KDF | | 技能 (64 MB) | |
| +------------------+ +------------------+ +------------------+ |
| |
| +------------------+ +------------------+ +------------------+ |
| | 一次性 | | rustls TLS | | 无 localStorage | |
| | 登录 token | | 用于所有网络 | | 存储敏感数据 | |
| | (5-min TTL) | | 连接 | | | |
| +------------------+ +------------------+ +------------------+ |
+-------------------------------------------------------------------+
```
- **凭证存储**:通过 `keyring` crate 的 OS 钥匙串集成(macOS Keychain、Windows Credential Manager、Linux Secret Service),仅限桌面端
- **内存加密**:带 Argon2id 密钥派生的 AES-256-GCM。所有 AI 内存静态加密
- **技能沙盒化**:每个 QuickJS 实例都有强制内存限制(默认 64 MB)和栈限制(512 KB)。禁止跨技能内存访问
- **认证交接**:Web 到桌面认证使用 5 分钟 TTL 的一次性登录 token,通过 Rust HTTP 客户端交换(绕过 CORS
- **网络 TLS**:所有 WebSocket 和 HTTP 连接使用 rustls,不依赖平台 OpenSSL
- **状态管理**:敏感数据保存在 Redux(内存)和 OS 钥匙串(持久化)中。凭证或 token 不使用 localStorage
- **提示注入防护**:用户提示在模型/工具执行前经过规范化/评分,并在服务器端强制执行(`allow | review | block`)。详见 [`docs/PROMPT_INJECTION_GUARD.md`](../../docs/PROMPT_INJECTION_GUARD.md)
---
## 端到端数据流
从用户操作到外部服务再返回的完整流程:
```text
用户在聊天 UI 中输入命令
|
v
React 前端分派到 AI 提供商
|
v
AI 模型接收提示 + 工具目录(通过 tool:sync
|
v
AI 决定调用技能工具(例如,发送 Telegram 消息)
|
v
通过 Socket.io 发送 mcp:toolCall 事件
|
v
Socket 管理器(Rust)接收事件,解析 skillId__toolName
|
v
技能注册表通过 MPSC 通道将消息路由到正确的 QuickJS 实例
|
v
QuickJS 技能执行工具处理程序
|
v
桥接 APInet.rs 通过 reqwest 发起 HTTP 请求(无 CORSrustls TLS
|
v
外部服务响应(例如,Telegram API
|
v
结果回流:桥接 -> QuickJS -> 注册表 -> Socket -> MCP -> AI -> UI
|
v
用户在聊天界面中看到结果
```
每一层都是异步且非阻塞的。Rust 核心在固定的 Tokio 线程池上处理数千个并发的技能执行、cron 触发和 socket 事件。
---
## 技术栈
| 层 | 技术 | 原因 |
| -------------- | ------------------------------- | -------------------------------------------------------- |
| **前端** | React 19, TypeScript 5.8 | 现代组件模型,类型安全 |
| **状态** | Redux Toolkit + Persist | 可预测状态,支持离线持久化 |
| **构建** | Vite 7 | 亚秒级 HMR,优化的生产构建 |
| **样式** | Tailwind CSS | 工具优先,一致的设计系统 |
| **框架** | Tauri v2 | 原生跨平台,开销最小 |
| **语言** | Rust (2021 edition) | 内存安全,零成本抽象 |
| **异步** | Tokio | 高性能异步 I/O 运行时 |
| **JS 运行时** | Node.js | 用于工具辅助函数和技能相关 JS 的托管 V8 运行时 |
| **数据库** | SQLite (rusqlite) | 嵌入式,零配置,每技能隔离 |
| **WebSocket** | tokio-tungstenite + rustls | 持久连接,原生 TLS |
| **HTTP** | reqwest | 异步 HTTP,支持 rustls + native-tLS 双栈 |
| **加密** | aes-gcm + argon2 | AES-256-GCM 加密,Argon2id 密钥派生 |
| **调度** | cron crate + 自定义调度器 | 标准 cron 表达式,5 秒精度 |
| **Telegram** | 已移除 | Telegram 集成已移除 |
| **实时** | Socket.io(客户端) | 双向基于事件的通信 |
| **AI** | MCPJSON-RPC 2.0 | LLM 集成的标准化工具协议 |
| **搜索** | OpenAI 嵌入 + SQLite FTS5 | 混合语义 + 关键词搜索 |
| **图谱** | Neo4j | 实体关系知识图谱 |
@@ -1,81 +0,0 @@
---
description: >-
OpenHuman 系统的高层轮廓(桌面壳层、Rust 核心、Memory Tree、Agent 循环)。指向仓库中的深度开发者架构文档。
icon: code-branch
lang: zh-CN
---
# 架构
OpenHuman 基于 GNU GPL3 开源。本页是系统的高层轮廓;深度开发者架构参考位于仓库中的 [深度架构文档](../architecture.zh-CN.md)。
## 系统形态
OpenHuman 是一款 **React + Tauri v2 桌面应用**,搭配一个承担重活的 **Rust 核心**
```text
┌──────────────────────────────────────────────────┐
│ Tauri 壳层 (app/src-tauri/) │
│ • 窗口管理、OS 集成、sidecar 生命周期 │
│ • 用于集成提供商的 CEF 子 WebView │
└──────────────────────────────────────────────────┘
│ JSON-RPC (HTTP) ↕
┌──────────────────────────────────────────────────┐
│ Rust 核心 (openhuman 二进制, src/) │
│ • Memory Tree 流水线 │
│ • 集成适配器 + 自动获取调度器 │
│ • 提供商路由器(模型路由) │
│ • TokenJuice 压缩 │
│ • 原生工具(搜索、获取、文件系统、git…) │
│ • 语音(STT 输入、TTS 输出、Meet Agent
└──────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────┐
│ React 前端 (app/src/) │
│ • 页面、导航 │
│ • 通过 coreRpcClient 与核心通信 │
│ • 无业务逻辑 —— 仅负责展示 │
└──────────────────────────────────────────────────┘
```
**逻辑归属:**
* **Rust 核心**。所有业务逻辑。Memory Tree、集成、模型路由、工具、语音。具有权威性。
* **Tauri 壳层**。窗口管理、进程生命周期、IPC。是交付载体,不是功能的栖身之所。
* **React 前端**。UI 与编排。通过 JSON-RPC 调用核心。
## 数据流
1. **连接**。通过 OAuth 接入[集成](../../features/integrations/README.zh-CN.md)。后端保存 token;核心永远不会以明文形式看到它。
2. **自动获取**。每二十分钟,[调度器](../../features/obsidian-wiki/auto-fetch.zh-CN.md)会遍历每个活跃连接,并要求每个原生提供商进行同步。
3. **规范化**。提供商输出(邮件页面、GitHub diff、Slack 频道转储)被归一化为带来源标签的 Markdown。
4. **分块**。Markdown 被拆分为 ≤3k token 的确定性块。
5. **存储**。块存入 SQLite (`<workspace>/memory_tree/chunks.db`),并以 `.md` 文件形式存入 `<workspace>/wiki/`
6. **评分**。后台工作线程运行嵌入、实体提取、热度评分。
7. **摘要**。从块池中构建并刷新来源 / 主题 / 全局摘要树。
8. **检索**。当你提问时,Agent 查询 Memory Tree(搜索 / 钻取 / 主题 / 全局 / 获取)。
9. **压缩**。工具输出和大型源数据在进入 LLM 上下文前经过 [TokenJuice](../../features/token-compression.zh-CN.md) 处理。
10. **路由**。[路由器](../../features/model-routing/) 根据任务提示选择合适的提供商 + 模型。
## 隐私边界
留在你机器上的数据:
* Memory Tree SQLite 数据库。
* Obsidian Markdown 仓库。
* 音频捕获缓冲区和任何本地模型状态。
经过 OpenHuman 后端的数据(在一个订阅下):
* LLM 调用(模型提供商)。
* 网页搜索智能体。
* 集成 OAuth 和工具智能体。
* TTS 流。
完整图景请参阅 [隐私与安全](../../features/privacy-and-security.zh-CN.md)。
## 开源
* **仓库:** [github.com/tinyhumansai/openhuman](https://github.com/tinyhumansai/openhuman)。GNU GPL3。
* 欢迎提交 **Issue 和 PR**。项目处于早期测试阶段。
* 对于贡献者,权威开发者指南是[深度架构文档](../architecture.zh-CN.md)。
@@ -1,311 +0,0 @@
---
description: >-
智能体轮次实际如何运行 —— 工具调用循环、子智能体分派、原型、分类、hook,以及围绕它们的成本/预算机制。
icon: layer-group
---
# Agent Harness
Agent Harness 是将用户消息(或 webhook 触发、cron tick)转变为完整的、使用工具的 LLM 交互的运行时。它拥有工具调用循环、子智能体分派、触发器-分类流水线和围绕它们的 hook 表面。它**不**拥有提供商 HTTP 传输、工具实现、提示部分组装或记忆存储 —— 那些是 harness 组合起来的独立领域。
本页先走过一个轮次中发生了什么,然后放大每个活动部件。
## 轮次的形态
每个轮次 —— 无论是用户刚输入消息、Telegram webhook 刚触发,还是 9am cron 刚 tick —— 都流经相同的生命周期:
```text
┌─ 入站 ─────────────────────────────────────────────────────────┐
│ 用户消息 · 渠道入站 · webhook · cron · composio 事件 │
└──────────────────────────┬────────────────────────────────────────┘
▼ (仅外部触发器)
┌──────────────────────┐
│ 触发器分类 │ 分类 → 丢弃 / 通知 /
│ (小型本地 LLM) │ 生成 reactor / 生成 orchestrator
└──────────┬───────────┘
┌──────────────────────────────┐
│ Agent::turn() │
│ 1. 恢复转录 │
│ 2. 构建系统提示* │
│ 3. 注入记忆上下文 │
│ 4. 进入工具调用循环 ────┼──► 提供商调用
│ 5. 分派工具调用 ────┼──► 工具执行 / 子智能体生成
│ 6. 上下文守卫 / 压缩 │
│ 7. 停止 hook 检查 │
│ 8. 最终助手文本 │
└──────────┬───────────────────┘
│ 异步,在用户看到回复后
┌─────────────────┐
│ 轮次后 │ archivist · learning · 成本日志 ·
│ hook │ 情景记忆索引
└─────────────────┘
* 系统提示仅在第一轮构建 —— 后续轮次逐字复用渲染后的提示,
以便推理后端的 KV-cache 前缀保持有效。
```
本页其余部分就是同一个图表,展开版。
## 会话和 `Agent::turn`
**会话**是 `Agent` 实例正在运行的实时对话。`Agent` 结构体拥有:
* 对话历史(系统 + 用户 + 助手 + 工具消息)。
* 要调用的提供商客户端(由[模型路由器](../../features/model-routing/)解析模型)。
* 模型可见的工具注册表。
* 在每条用户消息前为相关记忆补水的记忆加载器。
* 每轮预算 —— 最大工具迭代次数、最大 payload 大小、最大 USD 成本。
`Agent::turn(user_message)` 是热路径。在一个轮次中它:
1. **恢复会话转录**,如果这是一个新进程 —— 从磁盘重新加载精确的提供商消息,以便推理后端的 KV-cache 前缀仍然命中。
2. **构建系统提示**(仅在第一轮)。这拉入身份、soul、profile、记忆、已连接集成、可用工具、安全前言 —— 由提示部分构建器组装。
3. **注入记忆上下文**,通过记忆加载器为新用户消息注入:[记忆树](../../features/obsidian-wiki/memory-tree.zh-CN.md) 中的相关块,附带引用,使 UI 可以展示来源。
4. **进入工具调用循环**(下一节)。
5. **在后台生成轮次后 hook** —— 用户在 archivist / learning / 成本日志完成前就得到答案。
系统提示在后续轮次中**不**重建。即使是微小的字节变化也会使 KV-cache 前缀失效并强制完整重新 prefill,因此动态每轮上下文(记忆召回、新学习片段)作为用户可见的消息内容追加,而非拼接到系统提示中。
## 工具调用循环
`Agent::turn` 内部,工具调用循环是内部引擎。它最多运行 `max_tool_iterations` 轮(默认 10):
```text
loop {
1. 上下文守卫 - 如果历史太长,microcompact / autocompact
2. 停止 hook 检查 - 预算上限、最大迭代次数、自定义 kill switch
3. 提供商调用 - 发送消息 + 工具 spec,流式响应
4. 解析响应 - 将助手文本与工具调用分离
5. 如果没有工具调用 - 返回最终文本
6. 执行工具调用 - 分派每个(下一节)
7. 总结超大结果 - 将巨大工具输出路由到 summarizer 智能体
8. 追加结果 - 将工具结果推入历史,再次循环
}
```
每次迭代都会发出实时 `AgentProgress` 事件,以便 UI 可以逐 token 渲染流式传输、"正在调用工具 X" 状态和每轮成本更新。
### 工具分派和工具调用方言
不同的 LLM 说不同的工具调用方言。harness 通过 `ToolDispatcher` trait 抽象了这一点,它有三个具体实现:
* **Native** —— 拥有一等工具调用 API 的提供商(Anthropic、OpenAI)。工具调用以结构化字段返回,不在文本体中。
* **XML** —— 未原生训练工具调用但可遵循指令的模型的 fallback。工具被包装在助手文本中的 `<tool_call>{...}</tool_call>` 标签内。
* **P-Format** —— 某些较小模型使用的紧凑文本格式。
dispatcher 按提供商选择,使循环本身方言无关。相同的循环代码驱动 Claude、GPT、Gemini 和本地 Ollama 模型。
### 循环中的上下文管理
长工具调用链可能超出上下文窗口。两层处理:
* **工具结果预算** —— 每个工具结果都对照每调用字节预算检查。任何超出的内容都会被硬截断,并附带解释性标记,以便模型知道它没有看到完整输出。
* **Microcompact / autocompact** —— 当总历史接近上下文窗口时,harness 在下次提供商调用前将旧轮次压缩为摘要。压缩后的历史保持系统提示和最近轮次不变(KV-cache 稳定性),并重写中间部分。
### 超大工具结果 —— summarizer 绕道
某些工具调用返回巨大的 payload —— Composio action dump 200 KB JSON、网页抓取返回 50 KB markdown、跨越数千行的日志上的 `file_read`。在 payload 中间硬截断会丢弃恰好落在截断点之后的任何内容。
当工具结果超过 summarizer 阈值时,它在进入父历史之前通过专用的 `summarizer` 子智能体路由。summarizer 按照保留标识符和关键事实的提取合约压缩 payload,父智能体只看到压缩后的摘要。当 summarization 失败或 payload 大到在其上支付 LLM 调用在经济上没有意义时,硬截断仍是下游的备用方案。
### 缺失命令的自愈
当代码执行器子智能体运行 shell 命令且运行时回答 "command not found" 时,自愈拦截器捕获错误,生成一个 `ToolMaker` 子智能体为缺失命令编写 polyfill 脚本,然后重试原始调用。每个命令有尝试上限,因此真正不可能的命令不会无限循环。
## 子智能体 —— orchestrator 模式
OpenHuman 是**多智能体**的。与用户聊天的智能体是 **Orchestrator** —— 一个高级别的、策略层面的智能体,决定何时直接回答、何时使用直接工具、何时生成专家子智能体。
### 为什么多智能体
一个知道一切的单个智能体也有一个小书大小的系统提示。将工作拆分到专家意味着:
* 每个子智能体获得一个**窄系统提示**,只有它需要的部分(可以剥离身份 / 记忆 / 安全前言)。
* 每个子智能体获得一个**过滤后的工具注册表** —— 集成智能体不需要文件系统工具,coder 不需要 Composio 目录。
* 子智能体历史永远不会泄露回父级 —— 父级看到一个紧凑的工具结果,而非内部对话。
* 更便宜的模型可以做叶子工作。Orchestrator 使用强推理模型;研究子智能体可能使用更快、更便宜的模型。
### 内置原型
每个原型位于 `agents/<name>/` 下,带一个 `agent.toml`(元数据、工具范围、模型提示)和一个提示:
| 原型 | Orchestrator 何时选择它 |
| ------------------- | --------------------------------------------------------------------------------------- |
| `orchestrator` | 顶层智能体。永远不会被另一个 orchestrator 生成。 |
| `planner` | 多步分解 —— 将复杂请求分解为有序子任务。 |
| `researcher` | 网页/文档查找、引用搜寻。 |
| `code_executor` | 在工作区中编写、运行和调试代码。 |
| `critic` | 代码审查、对另一个智能体输出的质量检查。 |
| `summarizer` | 压缩超大工具结果(由 harness 调用,通常不是模型调用)。 |
| `archivist` | 记忆蒸馏 —— 持久化什么、遗忘什么。 |
| `tool_maker` | 自愈 —— 为缺失的 shell 命令编写 polyfill。 |
| `tools_agent` | 任意工具绑定任务的通用专家。 |
| `integrations_agent`| 绑定到特定 Composio 工具包(Gmail、GitHub、Slack…)以执行该工具包的动作。|
| `trigger_triage` | 将传入的外部事件分类为丢弃 / 通知 / 生成 reactor / 生成智能体。 |
| `trigger_reactor` | 对分类后的触发器的轻量级反应,不需要完整的 orchestrator 轮次。 |
| `morning_briefing` | 由 cron 运行的精选每日摘要。 |
| `welcome` / `help` | Onboarding 流程。 |
自定义原型作为 TOML 文件发布在 `$OPENHUMAN_WORKSPACE/agents/*.toml`(或 `~/.openhuman/agents/*.toml` 用于用户全局专家)。自定义定义在 id 冲突时覆盖内置定义。
### 运行子智能体
当 orchestrator 调用 `spawn_subagent`(或 `delegate_*` 便捷工具之一)时,runner
1. 从 task-local 读取父执行上下文 —— 父提供商、sandbox 模式、取消围栏、转录根。
2. 解析子智能体的模型 —— 继承父级、遵循提示(`fast` / `reasoning` / `summarization`),或固定到精确模型。
3. 按定义的 `tools``disallowed_tools``skill_filter` 过滤父级的工具注册表。在 `fork` 模式下,父级的完整注册表逐字继承。
4. 构建窄系统提示,省略定义要求剥离的部分。
5. 使用与父级相同的机制运行内部工具调用循环。
6. 返回一个紧凑的文本结果。子智能体内部历史永远不会拼接到父级中 —— orchestrator 看到一个单一的工具结果并继续。
对于不需要阻塞 orchestrator 轮次的任务,`spawn_worker_thread` 在后台运行子智能体,orchestrator 立即继续。
### 生成层级和 tiers
并非每个智能体都被允许生成每个其他智能体。harness 建模了一个三层层级,镜像模型之间的成本 / 延迟 / 思考深度拆分:
```text
Chat (快速,UX 聚焦 —— 例如 orchestrator 使用 `chat` 提示)
├─► Worker ◄─── 快速路径:一次委托,叶子做工作
└─► Reasoning (慢速,深度思考 —— 例如 planner 使用 `reasoning` 提示)
└─► Worker ◄─── 深度路径:reasoning 分解,workers 执行
```
每个 `AgentDefinition` 携带一个 `agent_tier` 字段(`chat` / `reasoning` / `worker`,默认 `worker`)。契约:
| Tier | 可以生成 | 禁止生成 | 典型成员 |
| ------------ | ----------------- | ---------------------------- | -------------------------------------------------------- |
| `chat` | `reasoning`, `worker` | 另一个 `chat` | `orchestrator` |
| `reasoning` | `worker` | 另一个 `reasoning`、任何 `chat` | `planner`(当今的规范代表) |
| `worker` | nothing[^1] | 任何东西 | researcher、code_executor、critic、archivist、tool_maker、integrations_agent、… |
[^1]: Skill-wildcard 条目(`{ skills = "*" }`)被豁免,因为它们坍缩为单个 `delegate_to_integrations_agent` 工具,其目标是 worker —— 它们是扇出委托表面,不是递归生成。
**为什么有这些规则。**
- *Chat → chat 毫无意义。* Chat tier 存在是为了 snappy UX。Chat 智能体生成另一个 chat 智能体只是加倍 TTFT 并燃烧 token 而不购买任何新能力。
- *Reasoning → reasoning 会爆炸深度。* Reasoning tier 很昂贵。Reasoning 智能体链倾向于重新分解相同问题并创建失控的层级。
- *Worker → anything 混合执行和编排。* Workers 是叶子,因此父级总是看到一个紧凑结果,而非嵌套委托的转录。
**强制执行。** 两层:
1. **加载时(静态)。** [`agents::loader::validate_tier_hierarchy`](../../../src/openhuman/agent/agents/loader.rs) 在合并的注册表(内置 + workspace TOML)上运行,并拒绝启动列出同级或 worker-with-subagents 条目的注册表。内置原型在编译测试时检查;用户发布的 TOML 在 workspace 加载时检查。
2. **运行时深度门禁(动态)。** 独立于 tier,子智能体 runner 通过 task-local 计数器将总生成链深度限制为 `MAX_SPAWN_DEPTH = 3`,该计数器在 `run_subagent` 之间递增,作为 `SpawnDepthExceeded` 智能体错误展示。这使得一个删除了 tier 注释的用户发布 TOML 仍然无法递归超过三跳。
> **状态:** 加载时 tier 检查、`agent_tier` 字段和运行时深度计数器 task-local 已上线。深度由静态加载器契约和运行时 `MAX_SPAWN_DEPTH = 3` 守卫共同限制。
### 工具包特定专家
对于具有数百个动作的 Composio 工具包(仅 GitHub 就有 500+),将每个动作加载到子智能体的工具集中会膨胀提示大小。harness 通过廉价的纯 CPU 过滤器(动词检测、token 重叠、动词对齐提升)将工具包的动作与父级精炼的任务提示进行排名,并仅将排名靠前的子集加载到子智能体中。无需模型调用,纯启发式 —— 快速且可解释。
## 分类 —— 处理外部触发器
当 webhook 触发、cron tick 或 Composio 事件到达时,系统不能直接将它们交给 orchestrator。大多数触发器是噪音;有些值得通知;只有少数值得完整的智能体轮次。**触发器-分类流水线**是门禁。
```text
TriggerEnvelope ──► run_triage ──► TriageDecision ──► apply_decision
│ │
│ ├─► 丢弃 (噪音)
│ ├─► 仅通知
│ ├─► 生成 trigger_reactor
│ └─► 生成 orchestrator
└── 小型本地 LLM(云端 LLM 重试 fallback
```
evaluator 有意保持廉价 —— 在可用时使用小型本地模型,重试时 fallback 到远程模型。决策被缓存,因此相同的触发器不会重新分类。只有升级到"生成 orchestrator"的触发器才会通过完整的 `Agent::turn` 机制。
## Hook —— 可观测性和策略杠杆
两个 hook 表面包裹循环,位于两端:
### 停止 hook(轮次中)
停止 hook 在工具调用循环的**迭代之间**触发。它们是预算上限、速率限制和自定义 kill switch 的策略杠杆。内置 hook
* **预算停止 hook** —— 使用每轮成本累加器限制轮次的累计 USD 成本。
* **最大迭代次数停止 hook** —— 从智能体持久配置外部限制迭代次数。
返回 `Stop` 的 hook 会以清晰的原因中止循环,调用者可以将该原因展示给用户。停止 hook 与中断(下一节)不同:它们是策略驱动的,不是用户驱动的。
### 轮次后 hook
轮次后 hook 在轮次**完成后**触发,在后台。它们获得 `TurnContext` 快照 —— 用户消息、助手响应、每个工具调用及其参数和结果、总 wall-clock、迭代次数、会话 ID。内置消费者:
* **Archivist** —— 蒸馏轮次中哪些事实值得持久化到长期记忆。
* **Learning** —— 为 reflection、工具跟踪器和用户 profile 更新提供输入。
* **成本日志** —— 最终每轮成本行。
* **情景记忆索引** —— 将轮次作为块写入[记忆树](../../features/obsidian-wiki/memory-tree.zh-CN.md)以供未来召回。
Hook 通过 `tokio::spawn` 运行,因此用户在它们完成前就得到了答案。
## 中断 —— 优雅取消
`InterruptFence` 在循环的固定安全点检查 —— 每次工具执行前、每次子智能体生成前、每次提供商调用前。当用户按下 Ctrl+C 或发送 `/stop`
* 围栏翻转。
* 每个正在运行的子智能体看到相同的 flag(通过 `Arc` 共享)并在其下一个检查点退出。
* 进行中的提供商流被丢弃。
* Archivist 仍然使用任何存在的部分上下文触发,因此对话不会丢失。
中断是用户驱动的;停止 hook 是策略驱动的。它们共享底层的"干净停止循环"管道,但从不同侧面进入。
## 成本核算
每个提供商响应携带一个 `UsageInfo` 块 —— 输入 token、输出 token、缓存输入 token,以及由 OpenHuman 后端填充的权威 `charged_amount_usd``TurnCost` 在一个轮次内对每个提供商调用求和,以便 harness 可以:
* 通过进度通道发出每轮成本遥测。
* 为预算停止 hook 提供输入,使失控的轮次在循环中自我切断。
* 记录精确的轮次结束成本行。
当后端不展示收费金额时(旧构建、不通过它计费的提供商),一个小的每 tier 费率表提供 token 费率 floor 估计。后端直接成本在可用时总是优先。
## Fork 上下文 —— 跨 harness 的 KV-cache 复用
harness 使用 task-local `ParentExecutionContext` 将父状态线程化到子智能体中,而不会爆炸每个函数签名。相同的模式携带当前 sandbox 模式、中断围栏和停止 hook 列表。继承父级提供商、模型和提示前缀的子智能体可以在推理后端上**共享父级的 KV-cache 前缀** —— 比从头重新 prefill 明显更便宜。
## 自愈回顾
几个小型自适应系统位于主循环之上:
* **缺失命令的自愈** —— `ToolMaker` polyfill,有上限的重试尝试。
* **Payload summarizer 断路器** —— 会话中连续三次子智能体失败会禁用 summarizationfallback 到截断。
* **分类本地-vs-远程重试** —— 本地 LLM 优先;解析失败时远程 fallback。
这些都不会改变循环的形状 —— 它们只是让常见故障模式无需用户干预即可恢复。
## 代码中该看哪里
harness 完全位于 `src/openhuman/agent/` 下。该目录中的 README 枚举了公共表面;负载最重的文件是:
| 文件 / 目录 | 里面有什么 |
| ----------------------------- | ----------------------------------------------------------------- |
| `harness/session/turn.rs` | `Agent::turn` —— 上述生命周期。 |
| `harness/tool_loop.rs` | 内部工具调用循环。 |
| `harness/subagent_runner/` | `run_subagent`、fork 模式、超大结果交接。 |
| `harness/definition.rs` | `AgentDefinition` —— 原型声明的内容。 |
| `harness/tool_filter.rs` | 集成子智能体的工具包动作排名。 |
| `harness/payload_summarizer.rs` | 超大工具结果绕道。 |
| `harness/self_healing.rs` | 缺失命令拦截器。 |
| `harness/interrupt.rs` | 取消围栏。 |
| `dispatcher.rs` | 工具调用方言抽象。 |
| `triage/` | 外部触发器分类 + 升级。 |
| `agents/` | 内置原型 —— 每个智能体一个子目录。 |
| `hooks.rs` / `stop_hooks.rs` | 轮次后和轮次中 hook 表面。 |
| `cost.rs` | 每轮 USD/token 核算。 |
| `progress.rs` | 到 UI 的实时进度事件。 |
| `memory_loader.rs` | 每条用户消息的记忆树上下文注入。 |
## 另请参阅
* [架构概览](README.zh-CN.md) —— harness 在更大图景中的位置。
* [记忆树](../../features/obsidian-wiki/memory-tree.zh-CN.md) —— 记忆加载器从中读取、轮次后 hook 写入的内容。
* [自动模型路由](../../features/model-routing/README.zh-CN.md) —— `model: "hint:reasoning"` 如何解析为具体的提供商+模型。
* [原生工具 —— 智能体协调](../../features/native-tools/agent-coordination.zh-CN.md) —— `spawn_subagent``delegate_*``todo_write` 的用户可见表面。
@@ -1,129 +0,0 @@
---
description: Desktop Companion 领域 —— Clicky 风格的交互循环,将热键、语音、屏幕智能、LLM、TTS 和视觉指向整合为单一产品体验。
icon: robot
---
# Desktop Companion (`src/openhuman/desktop_companion/`)
Desktop Companion 编排一个 Clicky 风格的交互循环:热键激活、麦克风捕获、屏幕上下文、LLM 推理、语音合成和视觉指向。它复用现有构建块,而非重新实现它们。
## 构建块
| 模块 | 提供的能力 | 路径 |
|--------|-----------------|------|
| **screen_intelligence** | 权限门控的捕获会话、`capture_now()``VisionSummary``AppContextInfo` | `src/openhuman/screen_intelligence/` |
| **voice** | 热键监听器(push/tap)、音频捕获、云端 STTWhisper)、TTS (`reply_speech`) | `src/openhuman/voice/` |
| **meet_agent** | LLM 编排模式(STT -> LLM -> TTS)、WAV 打包 | `src/openhuman/meet_agent/` |
| **overlay** | 浮动 UI 表面、注意力事件、打字机气泡 | `src/openhuman/overlay/` |
| **provider_surfaces** | 连接应用事件队列 (`ingest_event`, `list_queue`) | `src/openhuman/provider_surfaces/` |
| **accessibility** | 前台应用上下文 (`foreground_context()`) | `src/openhuman/accessibility/` |
## 模块布局
```text
src/openhuman/desktop_companion/
mod.rs — 模块导出(轻量)
types.rs — CompanionState enum、CompanionConfig、ConversationTurn、会话 param/result 类型
session.rs — 单例会话生命周期、状态机、TTL、对话历史
pipeline.rs — STT -> 屏幕上下文 -> LLM -> TTS -> 指向编排
pointing.rs — [POINT:x,y:label:screenN] 标签解析器、多显示器坐标映射
handoff.rs — 连接应用动作的 provider-surface 队列匹配
bus.rs — CompanionStateChangedEvent 的广播通道
schemas.rs — RPC 控制器 (companion_start_session, companion_stop_session 等)
```
## 状态机
```text
Idle -> Listening -> Thinking -> Speaking -> Pointing -> Idle
| |
v v
Listening Listening (中断)
任何状态 -> Error -> Idle (重置)
```
有效转换由 `session::is_valid_transition()` 强制执行。关键路径:
- **Happy path**Idle -> Listening -> Thinking -> Speaking -> Pointing -> Idle
- **无指向**Thinking -> Speaking -> Idle(响应中没有 POINT 标签)
- **中断**Speaking/Pointing -> Listening(用户重新激活热键)
- **取消**Thinking -> Idle(用户在思考中途取消)
- **错误恢复**Any -> Error -> Idle
## 交互流水线
`pipeline.rs` 编排单个轮次:
1. **激活** —— 状态转换为 Listening(将由 Tauri 壳层热键桥接驱动,见 PR 2)
2. **STT** —— 通过 `voice::cloud_transcribe`Whisper)转录音频样本
3. **屏幕上下文** —— `accessibility::foreground_context()` 获取应用名称 + 窗口标题
4. **LLM** —— 通过 `BackendOAuthClient` 进行聊天补全,携带系统提示、屏幕上下文和滚动对话历史(最近 20 轮作为上下文)
5. **解析响应** —— 通过 `pointing::parse_and_map()` 提取 `[POINT:x,y:label:screenN]` 标签
6. **Handoff 检查** —— 扫描响应中的提供商关键词,与 `provider_surfaces` 队列匹配
7. **TTS** —— 通过 `voice::reply_speech`ElevenLabs)合成语音
8. **指向** —— 为 overlay 动画发射指向目标
9. **返回 Idle**
流水线通过 `CancellationToken` 支持取消 —— Tauri 壳层可以在任何检查点取消(STT、LLM、TTS 阶段之间)。
文本输入也通过 `run_text_turn()` 支持,跳过 STT。
## 会话生命周期
- **一次一个会话** —— 由进程级 `Mutex<Option<CompanionSessionInner>>` 强制执行
- **需要同意** —— `start_session` 拒绝 `consent=false`
- **TTL 强制执行** —— 当 `status()` 检测到 TTL 已过时,会话自动过期
- **对话历史** —— 上限 50 轮,溢出时最旧的被丢弃
## RPC 表面
命名空间:`companion`。所有方法都通过标准控制器注册表。
| 方法 | 说明 |
|--------|-------------|
| `companion_start_session` | 以显式同意 + 可选 TTL 启动会话 |
| `companion_stop_session` | 结束活跃会话 |
| `companion_status` | 当前状态、会话信息、剩余 TTL |
| `companion_config_get` | 读取 companion 配置 |
| `companion_config_set` | 更新 companion 配置 |
## 事件总线
`CompanionStateChangedEvent` 通过 `tokio::sync::broadcast` 通道广播(与 `overlay::bus` 相同模式)。三个 `DomainEvent` 变体路由到 `"companion"` 领域:
- `CompanionSessionStarted { session_id }`
- `CompanionStateChanged { session_id, state, previous_state }`
- `CompanionSessionEnded { session_id, reason }`
## 指向系统
LLM 响应可以嵌入 `[POINT:x,y:label:screenN]` 标签。`pointing.rs`
- 通过正则解析标签
- 使用 `ScreenGeometry` 将屏幕相对坐标映射为绝对桌面坐标
- 将坐标钳制到屏幕边界
- 索引越界时回退到 screen 0
- 从显示文本中剥离标签
## Provider-surface handoff
`handoff.rs` 扫描清理后的 LLM 响应文本中的提供商关键词(slack、discord、telegram 等),并将它们与 `provider_surfaces` 队列中的条目匹配。当找到匹配时,`HandoffEvent` 被包含在 `TurnResult` 中,供 Tauri 壳层 / overlay 展示。
## 平台范围
- **macOS**:完整支持 —— 热键、屏幕捕获、指向、TTS、overlay
- **Windows/Linux**:部分 —— 热键可用(rdev),屏幕上下文 stub,无指向
平台特定代码通过 `#[cfg(target_os = "macos")]` 门控。
## 测试
| 文件 | 覆盖范围 |
|------|----------|
| `session_tests.rs` | 会话 CRUD、状态机转换、TTL、同意、对话历史 |
| `pipeline_tests.rs` | 轮次编排、取消、输入验证、系统提示 |
| `pointing_tests.rs` | 标签解析、坐标映射、多显示器、边界情况 |
| `handoff.rs` (inline) | 关键词匹配、空队列、提供商覆盖 |
| `schemas.rs` (inline) | 控制器计数、schema 字段验证 |
| `tests/json_rpc_e2e.rs` | 完整 RPC 往返:start -> status -> config -> stop |
File diff suppressed because it is too large Load Diff
@@ -1,209 +0,0 @@
---
description: 桌面宿主 (`app/src-tauri/`) —— Tauri v2 + WebView、IPC、嵌入式核心生命周期、核心桥接。
icon: desktop
---
# Tauri Shell (`app/src-tauri/`)
OpenHuman 的桌面宿主:Tauri v2 + WebView、IPC 命令、窗口管理,以及桥接到嵌入式 `openhuman-core` Rust 运行时(核心 JSON-RPC)。它**不会**重复完整的领域栈;那部分存在于仓库根目录的 Rust crate 中(`openhuman_core``src/main.rs`)。
## 职责
1. **Web UI**。从 `app/dist` 加载 Vite 构建(或开发服务器,端口 1420)。
2. **IPC**。暴露一小套明确的 Tauri 命令(见 [Commands](#tauri-ipc-commands-app-src-tauri))。
3. **核心生命周期**。启动进程内核心服务器,并通过 `core_rpc_relay` 代理 JSON-RPC。
4. **磁盘上的 AI 提示**。从资源 / 开发 cwd 解析捆绑的 `src/openhuman/agent/prompts`,用于 `ai_get_config` / `write_ai_config_file`
5. **窗口 + 托盘**。桌面窗口行为和系统托盘(见 `lib.rs`)。
## 核心进程模型
`app/package.json``core:stage` 现在有意保持为 no-op,仅用于脚本兼容性。桌面应用会在进程内链接核心,因此本地构建不再需要在 `app/src-tauri/binaries/` 下 staging `openhuman-core-*` sidecar。
## 卡死进程恢复
正常应用退出从 `RunEvent::ExitRequested` 运行 teardownCEF 关闭前先关闭子 webview,触发嵌入式核心的 cancellation token,最终进程扫描在短暂的宽限期后向直接子进程发送 `SIGTERM`,然后升级使用 `SIGKILL` 处理顽固进程。扫描摘要记录为 `[app] sweep: term=N kill=M total=K`;任何非零 `kill` 计数都是警告,意味着子进程忽略了优雅关闭。
在 macOS 上,硬退出(强制退出、`SIGKILL`、渲染器崩溃)可能跳过正常的 teardown。下一次启动在 CEF 缓存 preflight 之前运行启动恢复:它列出可执行路径属于正在启动的 `.app/Contents` 的 OpenHuman 进程,跳过当前进程,发送 `SIGTERM`,短暂等待,然后对仍然匹配相同 pid+command 的顽固进程发送 `SIGKILL`。日志使用 `[startup-recovery]` 前缀。
当设置了 `OPENHUMAN_CORE_REUSE_EXISTING=1` 时(以便手动 CLI-core 复用仍然有效),以及当 CEF `SingletonLock` 被实时进程持有时(以便正常的 second-instance 路径可以在不杀死已运行应用的情况下失败),启动恢复跳过。Tauri 命令 `process_diagnostics_list_owned` 返回当前拥有的进程列表;macOS 实现是 bundle 作用域的,Linux/Windows 目前返回空。
## Tauri Shell 架构 (`app/src-tauri/`)
### 概述
**`app/src-tauri`** crateRust 包 **`OpenHuman`**,二进制文件 **`OpenHuman`**)是一个**仅限桌面**的宿主。它嵌入 React UI,注册插件(深度链接、打开器、OS、通知、自动启动、更新器),管理主窗口和托盘,并**中继 JSON-RPC** 到嵌入式核心服务器。
非桌面目标在编译时失败(`lib.rs` 中的 `compile_error!`)。
### 目录布局(实际)
```text
app/src-tauri/src/
├── lib.rs # `run()`、托盘/菜单动作、插件、`generate_handler!`、核心启动
├── main.rs # 二进制入口
├── core_process.rs # CoreProcessHandle、嵌入式核心服务器任务
├── core_rpc.rs # 核心 JSON-RPC 的 HTTP 客户端
├── commands/
│ ├── mod.rs # 重新导出
│ ├── core_relay.rs # `core_rpc_relay`、服务管理的核心引导
│ ├── openhuman.rs # Daemon 宿主配置、systemd 风格服务辅助函数
│ └── window.rs # 显示/隐藏/最小化/关闭窗口
└── utils/
├── mod.rs
└── dev_paths.rs # 解析捆绑的 AI 提示路径
```
此树中**没有** `src-tauri/src/services/session_service.rs`;会话语义在 Web 层 + 后端 + 核心中按适用情况处理。
### 数据流:UI → 核心
```text
React (invoke)
→ core_rpc_relay { method, params, serviceManaged? }
→ core_rpc::call HTTP POST 到 OPENHUMAN_CORE_RPC_URL
→ 嵌入式 openhuman 核心服务器
```
`core_process.rs` 中的 `CoreProcessHandle` 拥有嵌入式服务器任务;`commands/core_relay.rs` 可选地在 relay 之前确保**服务管理**的核心正在运行。
### 窗口和托盘行为
- 壳层在启动时创建托盘图标,并将动作连接到打开主窗口或退出。
- 在 daemon 模式(`daemon` / `--daemon`)下,主窗口在启动时隐藏,可以从托盘动作重新打开。
- 在 macOS 上,`RunEvent::Reopen` 也会恢复并聚焦主窗口。
- Windows 和 Linux 使用相同的托盘动作(`Open OpenHuman``Quit`),某些 Linux 设置上有桌面环境特定的托盘渲染差异。
### 捆绑资源
`tauri.conf.json` 捆绑 **`../../skills/skills`** 和 **`../../src/openhuman/agent/prompts`**,使技能和提示 markdown 随应用一起发布。
### 相关
- IPC 表面:见下方的 [Commands](#tauri-ipc-commands-app-src-tauri) 部分
- HTTP 桥接:见下方的 [Core bridge & helpers](#core-bridge-helpers-app-src-tauri) 部分
- Rust 领域(实现):仓库根目录 `src/openhuman/``src/core_server/`
## Tauri IPC 命令 (`app/src-tauri`) {#tauri-ipc-commands-app-src-tauri}
所有命令都在 **`app/src-tauri/src/lib.rs`** 中的 `tauri::generate_handler![...]` 内注册(桌面构建)。下方名称是 **Rust** 命令名称(在 JS 中通过 serde 应用 camelCase)。
### Demo / 诊断
| 命令 | 用途 |
| ------- | ------------------------------------------ |
| `greet` | Demo 字符串(生产中可安全移除) |
### AI 配置(捆绑提示)
| 命令 | 用途 |
| ---------------------- | -------------------------------------------------------------------------------------------- |
| `ai_get_config` | 从捆绑或开发 `src/openhuman/agent/prompts` 下解析的 `SOUL.md` / `TOOLS.md` 构建 `AIPreview` |
| `ai_refresh_config` | 与 `ai_get_config` 相同的读取路径(刷新 hook) |
| `write_ai_config_file` | 在仓库 `src/openhuman/agent/prompts` 下写入单个 `.md`(开发 / 安全文件名检查) |
### 核心 JSON-RPC 中继
| 命令 | 用途 |
| ---------------- | -------------------------------------------------------------------------------------------------------------- |
| `core_rpc_relay` | Body: `{ method, params?, serviceManaged? }` → 转发到本地 **`openhuman-core`** HTTP JSON-RPC (`core_rpc.rs`) |
从前端使用 **`app/src/services/coreRpcClient.ts`** (`callCoreRpc`)。
### 窗口管理
来自 **`commands/window.rs`**(名称可能略有不同;见 `lib.rs`):
| 命令 | 用途 |
| ------------------- | ----------------- |
| `show_window` | 显示主窗口 |
| `hide_window` | 隐藏主窗口 |
| `toggle_window` | 切换可见性 |
| `is_window_visible` | 查询可见性 |
| `minimize_window` | 最小化 |
| `maximize_window` | 最大化 |
| `close_window` | 关闭 |
| `set_window_title` | 设置标题字符串 |
### OpenHuman daemon / 服务辅助函数
来自 **`commands/openhuman.rs`**(见源码获取精确 payload):
| 命令 | 用途 |
| ---------------------------------- | ---------------------------------------------- |
| `openhuman_get_daemon_host_config` | 读取 daemon 宿主偏好设置(例如托盘) |
| `openhuman_set_daemon_host_config` | 持久化 daemon 宿主偏好设置 |
| `openhuman_service_install` | 安装后台服务(平台特定) |
| `openhuman_service_start` | 启动服务 |
| `openhuman_service_stop` | 停止服务 |
| `openhuman_service_status` | 查询状态 |
| `openhuman_service_uninstall` | 卸载服务 |
### 屏幕共享选择器(CEF / macOS
来自 **`screen_capture/mod.rs`**。支持 `webview_accounts/runtime.js` 中的页面内 `getDisplayMedia` shim。会话门控:shim 必须在成功枚举/缩略图捕获之前用实时用户手势打开会话。见 issue #713(选择器 UX+ #812(会话门控)。
| 命令 | 用途 |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| `screen_share_begin_session` | 从账户 webview 打开 30s 会话,在 `navigator.userActivation.isActive` 手势之后。返回 `{ token, sources }`。每个账户限速 10/分钟。 |
| `screen_share_thumbnail` | 将单个来源的缩略图捕获为 base64 PNG。需要 live token 和会话颁发的 `id`。仅 macOS;其他平台返回错误。 |
| `screen_share_finalize_session` | 关闭会话。由 shim 在 Share 或 Cancel 时调用;使用未知/过期 token 安全调用(no-op)。 |
### 已移除 / 不存在
以下命令**不**存在于当前的 `generate_handler!` 列表中:`exchange_token``get_auth_state``socket_connect``start_telegram_login`。认证和 socket 在 **React** 应用和 **核心** 进程中处理,而非通过这些 IPC 名称。
### 示例:核心 RPC
```typescript
import { invoke } from "@tauri-apps/api/core";
const result = await invoke("core_rpc_relay", {
request: {
method: "your.rpc.method",
params: { foo: "bar" },
serviceManaged: false,
},
});
```
---
_见 `app/src-tauri/src/lib.rs` 获取权威列表。_
## Core bridge & helpers (`app/src-tauri`) {#core-bridge-helpers-app-src-tauri}
本文档替代了旧的 "SessionService / SocketService" 拆分。Tauri crate **不**嵌入重复的 Socket.io 服务器或 Telegram 客户端;相反,它专注于对 **`openhuman-core`** 二进制文件的**进程管理**和 **HTTP JSON-RPC**
### `CoreProcessHandle` (`core_process.rs`)
- 解析 **`openhuman-core`** 可执行文件(staging 在 `binaries/` 下或 `PATH` / 开发布局中)。
- 启动或附加到核心进程并暴露其 RPC URL (`OPENHUMAN_CORE_RPC_URL`)。
-`lib.rs` 的应用设置期间使用 (`app.manage(core_handle)`)。
### `core_rpc` (`core_rpc.rs`)
- 核心 JSON-RPC 表面的 HTTP 客户端(localhost)。
-**`core_rpc_relay`** 使用,以转发前端的 `method` + `params`
### `commands/core_relay.rs`
- **`core_rpc_relay`**。确保核心正在运行(进程内句柄或**服务管理**路径),然后调用 `core_rpc`
- **`ensure_service_managed_core_running`**。当 RPC 不可用时引导 systemd/launchd 风格服务(核心 CLI 内的平台特定行为)。
### `commands/openhuman.rs`
- Daemon 宿主 JSON 配置(例如托盘可见性),位于应用数据目录下。
-**openhuman** 后台服务提供 install/start/stop/status/uninstall 辅助函数。
### `utils/dev_paths.rs`
- 解析 AI preview 的开发和捆绑资源路径下的 **`src/openhuman/agent/prompts`**。
### `utils/tauriSocket.ts`(前端)
不在 `src-tauri` 中,但与 shell **配对**React 应用监听镜像 Rust 端客户端 socket 活动的 Tauri 事件。见 `app/src/utils/tauriSocket.ts` 和 [前端服务](frontend.zh-CN.md#services-layer) 章节。
---
@@ -1,190 +0,0 @@
---
description: 在全新机器上从头构建 Rust 核心。
icon: terminal
lang: zh-CN
---
# 构建 Rust 核心
本页面向贡献者,是在全新机器上编译 Rust 核心的参考文档。
它仅涵盖**仓库根目录的 crate**:
- Cargo 包:`openhuman`
- 二进制文件:`openhuman-core`
- 库:`openhuman_core`
如果你需要完整的桌面应用(`pnpm dev`、Tauri、CEF、前端工具链),请使用[环境搭建](getting-set-up.zh-CN.md)。该路径有额外的 JavaScript、子模块和桌面运行时依赖,**不**需要用于纯核心的 `cargo` 工作流。
## 1. 安装指定版本的 Rust 工具链
仓库在 [`rust-toolchain.toml`](../../rust-toolchain.toml) 中固定了 Rust 版本:
- Channel`1.93.0`
- Components`rustfmt``clippy`
推荐安装方式:
```bash
rustup toolchain install 1.93.0 --component rustfmt --component clippy
rustup default 1.93.0
```
你也可以在安装 `rustup` 后,让 `cargo``rust-toolchain.toml` 自动安装。
## 2. 克隆仓库
仅核心开发:
```bash
git clone https://github.com/tinyhumansai/openhuman.git
cd openhuman
```
这对根目录 crate 来说已足够。
桌面/Tauri 开发则不同:
- 只有在构建桌面壳层或 CEF 感知的 Tauri 工具链时,才需要 `app/src-tauri/vendor/` 子模块。
- 该流程请遵循[环境搭建](getting-set-up.zh-CN.md)并运行 `git submodule update --init --recursive`
## 3. 构建命令
从仓库根目录运行:
```bash
# 快速依赖 + 类型检查
cargo check --manifest-path Cargo.toml
# 实际 CLI / RPC 二进制文件的 Debug 构建
cargo build --manifest-path Cargo.toml --bin openhuman-core
# Release 构建
cargo build --manifest-path Cargo.toml --release --bin openhuman-core
# Rust 测试
cargo test --manifest-path Cargo.toml
```
注意:
- **包**名是 `openhuman`,但可运行的二进制文件是 **`openhuman-core`**。
- 如果你更喜欢面向包的 cargo 命令用于打包脚本,请使用 `-p openhuman`
- 构建好的二进制文件位于 `target/debug/openhuman-core``target/release/openhuman-core`
## 4. macOS 前置条件
安装:
- Xcode Command Line Tools`xcode-select --install`
原因:
- `whisper-rs` 在构建期间编译原生代码。
- 在 macOS 上,该 crate 在 [`Cargo.toml`](../../Cargo.toml) 中以 `metal` 特性启用构建,因此需要 Apple 工具链和 SDK 头文件。
安装 Xcode CLT 后,核心应该能用上述 cargo 命令构建。
## 5. Linux 前置条件
### 仅核心包集合
在全新 Linux 机器上运行 `cargo` 前,先安装这些包。
**Ubuntu / Debian**
```bash
sudo apt-get update
sudo apt-get install -y \
build-essential cmake pkg-config clang libssl-dev libclang-dev \
libasound2-dev libxi-dev libxtst-dev libxdo-dev libudev-dev \
libstdc++-14-dev
```
**Arch Linux**
```bash
sudo pacman -S --needed base-devel cmake pkgconf clang openssl \
alsa-lib libxi libxtst xdotool libevdev
```
> 在 Arch 上,`clang` 包含 `libclang``base-devel` 包含 `gcc`(提供 `libstdc++`),因此不需要单独的 `-dev` 包。
这些包的重要性:
- `build-essential` / `base-devel``cmake``pkg-config` / `pkgconf`:传递性 Rust 依赖使用的原生构建。
- `clang``libclang-dev`bindgen / C 和 C++ 编译路径,被原生 crate 使用。
- `libssl-dev` / `openssl`:某些网络依赖需要的 OpenSSL 头文件。
- `libasound2-dev` / `alsa-lib``libxi-dev` / `libxi``libxtst-dev` / `libxtst``libxdo-dev` / `xdotool``libudev-dev`Arch 中已包含在 `systemd-libs` 内)、`libevdev`:被核心构建引入的音频/输入/设备 crate 所需。
### `whisper-rs` + `clang` 注意事项
`whisper-rs-sys``clang` 下可能会失败并提示:
```text
fatal error: 'array' file not found
```
这就是为什么文档特别指出 `libstdc++-14-dev``clang` 在 Ubuntu runner 上可能会选择 GCC 14 的 C++ 头文件。
如果你的发行版布局仍然导致构建无法解析 `libstdc++.so`,请使用 [`AGENTS.md`](../../AGENTS.md) 中记录的相同变通方案:
```bash
# Ubuntu/Debian —— 按需调整 GCC 版本
sudo ln -sf /usr/lib/gcc/x86_64-linux-gnu/13/libstdc++.so /usr/lib/x86_64-linux-gnu/libstdc++.so
```
Arch Linux 通常不需要此变通方案,因为 `gcc-libs``libstdc++.so` 放在了默认库搜索路径上。
### Linux 桌面/Tauri 包集合
如果你构建的是桌面壳层而非仅核心 crate,请安装更广泛的依赖集合。
**Ubuntu / Debian**(镜像自 [`.github/workflows/build-desktop.yml`](../../.github/workflows/build-desktop.yml)):
```bash
sudo apt-get update
sudo apt-get install -y \
libgtk-3-dev libwebkit2gtk-4.1-dev libayatana-appindicator3-dev librsvg2-dev \
patchelf cmake libasound2-dev libxdo-dev libxtst-dev libx11-dev libxi-dev \
libevdev-dev libssl-dev libclang-dev \
libnss3 libnspr4 libatk1.0-0 libatk-bridge2.0-0 libcups2 libdrm2 \
libxkbcommon0 libxcomposite1 libxdamage1 libxfixes3 libxrandr2 \
libgbm1 libpango-1.0-0 libcairo2 libatspi2.0-0 libxshmfence1 libu2f-udev
```
**Arch Linux**
```bash
sudo pacman -S --needed gtk3 webkit2gtk-4.1 libayatana-appindicator \
librsvg patchelf nss nspr at-spi2-core libcups libdrm \
libxkbcommon libxcomposite libxdamage libxfixes libxrandr \
mesa pango cairo libxshmfence
```
仅在需要 `app/src-tauri/` 时使用桌面列表;对于根 crate 工作,上面较小的仅核心列表是相关的基线。
## 6. Windows 前置条件
安装:
- 通过 `rustup` 安装 Rust
- Visual Studio Build Tools 2022 或带 **使用 C++ 的桌面开发** 工作负载的 Visual Studio
- CI 和发布构建使用的 MSVC 目标:`x86_64-pc-windows-msvc`
安装 Microsoft 工具链后推荐的命令:
```powershell
rustup toolchain install 1.93.0 --component rustfmt --component clippy
rustup target add x86_64-pc-windows-msvc
cargo build --manifest-path Cargo.toml --bin openhuman-core
```
Windows 注意事项:
- 仓库对 `whisper-rs-sys` 打补丁以强制使用静态 MSVC CRT,并避免 [`Cargo.toml`](../../Cargo.toml) 中提到的 `LNK2038` / `LNK1169` 不匹配。请使用 MSVC 工具链,而非 MinGW。
## 7. 相关路径
- [环境搭建](getting-set-up.zh-CN.md):完整的桌面贡献者设置,含 `pnpm`、Tauri、子模块和 sidecar staging。
- [OpenHuman 架构](architecture/README.zh-CN.md):核心在桌面应用和 RPC 流程中的位置。
-172
View File
@@ -1,172 +0,0 @@
---
description: >-
为什么 OpenHuman 自带 Chromium 运行时,我们今天用它做什么,以及同样的 CDP 表面接下来能解锁什么。
icon: chrome
---
# Chromium Embedded Framework
OpenHuman 不运行在平台内置的 webview 上。它通过 `tauri-runtime` 的一个 fork 自带 **Chromium Embedded Framework (CEF) 运行时**,而这一个决策对产品几乎所有 "OpenHuman 知道你的工具里发生了什么" 的功能都是 load-bearing 的。
本页解释为什么 CEF 在 bundle 中,代码库今天用它做什么,以及同样的表面可以去哪里。
## 为什么用 CEF 而不是 stock webview
Stock Tauri 使用每个平台的原生 webview。macOS 上的 WKWebView、Windows 上的 WebView2、Linux 上的 WebKitGTK。这些用于渲染 OpenHuman 应用本身都能正常工作。它们对我们的用例有一个致命的局限性:**没有一个暴露 Chrome DevTools Protocol (CDP)**。
CDP 是 load-bearing 的原语。OpenHuman 中每个 "观察 Slack / WhatsApp / Telegram / Discord / Meet 内部发生了什么" 的功能都通过 CDP 与这些嵌入应用对话,而非通过注入的 JavaScript。CDP 提供:
* `Target.getTargets` 用于发现每个页面和服务 worker。
* `IndexedDB.requestDatabaseNames` / `requestDatabase` / `requestData` 用于遍历第三方应用的本地存储。
* `DOMSnapshot.captureSnapshot` 用于不会触发框架反应性的只读 DOM 检查。
* `Runtime.evaluate` 用于短暂的一次性读取(单个固定的 JSON 序列化器,从来不是持久桥接)。
* `Page.addScriptToEvaluateOnNewDocument` 用于极少数我们真正需要在页面 JS 运行前渲染器端 shim 的情况。
Stock webview 不能给我们任何这些。所以我们 vendor CEF。
Vendored 运行时位于 [`app/src-tauri/vendor/tauri-cef/`](https://github.com/tinyhumansai/openhuman/tree/main/app/src-tauri/vendor/tauri-cef)(从上游 `tauri-cef` 分支 fork 到 `tinyhumansai/tauri-cef:feat/cef-notification-intercept`,当前 CEF 146.4.1)。每个 Tauri crate 在 `app/src-tauri/Cargo.toml` 中通过 `[patch.crates-io]` 指向此 fork。Vendored `cargo-tauri` CLI 将 Chromium 正确捆绑到 `Contents/Frameworks/`stock `@tauri-apps/cli` 会产生一个损坏的 bundle,在 `cef::library_loader::LibraryLoader::new` 中 panic。[`scripts/ensure-tauri-cli.sh`](../../scripts/ensure-tauri-cli.sh) 在 fork 比安装的二进制文件更新时重新安装 vendored CLI。
## CEF 今天用于什么
### 嵌入的第三方 webview
每个作为托管 Web 应用运行的已连接提供商都有自己的子 CEF webview
* WhatsApp Web
* Telegram Web
* Slack
* Discord
* Google Meet
* LinkedIn
* Gmail
* Zoom
* browserscan
每个账户的存储隔离到 `{app_local_data_dir}/webview_accounts/{id}/`。两个 Slack workspace,两个浏览器配置文件。代码:[`app/src-tauri/src/webview_accounts/mod.rs`](../../app/src-tauri/src/webview_accounts/mod.rs)。
### CDP 驱动的扫描器
每个提供商在 [`app/src-tauri/src/`](https://github.com/tinyhumansai/openhuman/tree/main/app/src-tauri/src) 中都有一个**扫描器模块**。每个扫描器持有到 CEF 的 `--remote-debugging-port=19222` 的长期 WebSocket,并按固定节奏 tick
| 扫描器 | 节奏 | 做什么 |
| ------------------ | ------------------------------- | -------------------------------------------------------------------- |
| `whatsapp_scanner` | 2s DOM tick + 30s 完整 IDB 遍历 | 读取消息存储、拉取媒体元数据 |
| `telegram_scanner` | 相同 | 额外加上 QR 登录 hand-off 到原生 Telegram Desktop |
| `slack_scanner` | 30s IDB 遍历 | 纯 IDB —— 无需 DOM 抓取 |
| `discord_scanner` | 定期 | 通过 CDP 的频道 + DM 状态 |
| `meet_scanner` | 定期 | 通话期间的实时字幕 + 参与者状态 |
| `imessage_scanner` | 定期 | **无 webview。** 在 macOS 上直接读取 `~/Library/Messages/chat.db` |
每次扫描都会发出 `webview:event` payload,并直接向核心 RPC POST `openhuman.memory_doc_ingest`,因此无论 UI 窗口是否打开或后台运行,记忆都会增长。
### Google Meet mascot 摄像头
最炫的 CEF 技巧。Meet Agent 不只是"参加会议",它还**将自己广播为摄像头**。之所以能工作,是因为 CEF 允许我们:
1. 在任何 Meet 代码运行前通过 `Page.addScriptToEvaluateOnNewDocument` 注入一个微小桥接 (`camera_bridge.js`)。
2. 覆盖 `navigator.mediaDevices.getUserMedia`,使其从隐藏的 640×480 canvas 返回 `MediaStream`,而非真实摄像头。
3. 在该 canvas 上渲染 mascot SVG,通过 Rust 经 CDP 驱动的 `window.__openhumanSetMood(...)` 交换情绪状态(idle、thinking、talking)。
还有一个构建时路径,将 mascot SVG 栅格化为 Y4M,并使用 CEF 的原生 `--use-file-for-fake-video-capture` flag,一个完全原生的 fake-camera 来源,完全不使用 JS。
代码:[`app/src-tauri/src/meet_video/`](https://github.com/tinyhumansai/openhuman/tree/main/app/src-tauri/src/meet_video)。
### 原生通知拦截
`feat/cef-notification-intercept` 上的 fork 为 `Notification.permission``Notification.requestPermission()``navigator.permissions.query({name: "notifications"})` 添加了渲染器端 shim。这些现在在每条运行时代码路径上都安装在真正的 `tauri-runtime-cef` 路径中,因此当 Slack 检查它是否可以显示通知时,答案与 CEF 的权限回调已经授予的内容一致。
这是 `docs/TAURI_CEF_FINDINGS_AND_CHANGES.md` 的大部分内容。这就是 Slack 在一次会话中不再五次询问相同权限的原因。
## "不注入新 JS" 规则
规则记录在 [`CLAUDE.md`](../../CLAUDE.md) 中:**迁移的提供商以零注入 JavaScript 加载**。所有抓取都通过扫描器侧的 CDP 原生进行。
这很重要,因为任何在第三方来源内部运行的宿主控制代码都是攻击面责任。Slack 内部的持久 JS 桥接离失效只有一个 Slack 更新之遥,离通过攻击者控制的 JS 泄露桥接只有一个错误之遥。从渲染器外部的 CDP 严格更好。
| 提供商 | 已迁移? | 启动时加载什么 |
| ----------- | ------------- | -------------------------------- |
| WhatsApp | ✅ | 零 JS |
| Telegram | ✅ | 零 JS |
| Slack | ✅ | 零 JS |
| Discord | ✅ | 零 JS |
| browserscan | ✅ | 零 JS |
| Gmail | grandfathered | 遗留 `runtime.js` 桥接 |
| LinkedIn | grandfathered | 遗留 `LINKEDIN_RECIPE_JS` |
| Google Meet | grandfathered | 摄像头 + 音频 + 字幕桥接 |
遗留注入应该缩小,永远不要增长。新提供商直接走 CDP-only 路径。
## CEF 预热
一个隐藏的 CEF webview (`cef-prewarm`) 在应用启动时启动浏览器,因此当用户点击时第一个子 webview 立即生成。它在 `cef::shutdown()` 前被拆除以避免退出时的竞争。见 `app/src-tauri/src/lib.rs` 中 prewarm + 关闭生命周期附近的代码。
## Windows 启动诊断
CEF 在 onboarding UI 能够从渲染器故障中恢复之前初始化。如果 Windows 用户报告静默退出、永久的 "Connecting..." 转圈,或在第一个交互窗口出现前的 `tauri-runtime-cef` 断言,请在 issue 中询问这些细节:
* Windows 版本和完整构建号,特别是 Insider 构建。
* OpenHuman 版本和安装包类型(`.msi``.exe`)。
* 重试前是否将 `%LOCALAPPDATA%\com.openhuman.app` 移到了一边。
* `[startup]``[cef-profile]``[cef-startup]` 的启动日志行。
* 任何命名 `tauri-runtime-cef/src/lib.rs` 的 panic 文本。
对于 Windows Insider 构建,还要确认相同的安装包是否在当前稳定版 Windows 发布上启动。这会将 profile/缓存问题与 CEF 启动中的 OS/运行时兼容性回归分开。
## Linux shell fallbackCEF 启动崩溃时)
在某些 Linux 桌面上,特别是 NVIDIA 专有驱动设置下的 Wayland/XWaylandTauri/CEF shell 可能在 React 应用变得可用之前的原生窗口配置期间失败。一个已知症状是 CEF 报告主浏览器上下文后的 X11 `BadWindow` 错误。
当核心本身健康时,你可以通过分别运行核心和前端来继续开发:
```bash
cargo build --bin openhuman-core
./target/debug/openhuman-core run --port 7788
```
在另一个终端:
```bash
cd app
pnpm dev
```
在常规浏览器中打开 Vite URL,选择 **Advanced** / remote core 模式,将 RPC URL 设置为 `http://127.0.0.1:7788/rpc`,并使用核心写入的 bearer token。这会绕过原生专属功能,如托盘、自动更新和嵌入提供商 webview,但保持智能体、记忆、技能和 RPC 表面可用于调试。
## 插件审计
添加到 `app/src-tauri/src/lib.rs` 的任何新内容都必须审计 `js_init_script` 调用。`tauri-plugin-opener` 默认附带一个 init 脚本 (`init-iife.js`),添加了一个全局点击监听器;我们将其配置为 `.open_js_links_on_click(false)`,使其不在第三方 webview 内运行。`tauri-plugin-notification` 的 init 脚本同样从 vendored 副本中删除。
## 这里可以如何演进
CDP 表面是通用的。今天它为固定列表的提供商提供记忆摄入;同样的原语可以做更多。
### 浏览器自动化作为一等智能体工具
今天智能体有[原生工具](../features/native-tools/README.zh-CN.md)用于文件系统、git、网页搜索和网页获取。下一个明显的工具是**"驱动真实浏览器会话"**:登录用户已认证过的 SaaS,填写表单,抓取分页表格,下载导出。
plumbing 已经存在。`@openhuman/browser_task` 技能可以启动一个专用 CEF webview,通过 CDP 从核心驱动它,并将结果作为工具调用展示。用户现有的每账户配置文件意味着无需重新认证。
### Headless CEF 用于服务端回放
同样的扫描器模式(长期 WebSocket → IDB 遍历 + DOM snapshot)无需 UI 即可工作。核心 sidecar 中的 Headless CEF 可以按计划回放会话,适用于在云端托管核心并希望从不暴露干净 OAuth API 的来源自动获取的用户。
### 浏览器进程层的隐私 hook
CEF 的 `CefRequestHandler` 已经允许我们拦截网络请求。从"拦截并记录"到"拦截并重写"只有一小步:广告拦截、跟踪器拦截、每个提供商的 DNS 固定、请求重写。隐私作为一等浏览器功能,而非每个来源内泄漏的 JS shim。
### CDP 驱动的测试框架
扫描器模式、生成 webview、遍历 IDB、snapshot DOM、评估一个短暂表达式,在结构上与 E2E 测试编排相同。我们可以将 `@openhuman/web_test` 作为公共技能发布:`connect_cef → snapshot → evaluate → assert`。用纯 Rust 针对任何 Web 应用编写的测试,无需 Selenium / Playwright 依赖。
### 渲染器 ↔ Rust 消息通道
今天每个 CDP `Runtime.evaluate` 都是 fire-and-forget。从渲染器到 Rust 的长期双向通道(Tauri 为主机应用做 IPC 的方式)将解锁流式用例:实时打字检测、实时选择/高亮跟踪、主动推送。设计它时不违反"第三方来源中不允许持久 JS 桥接"规则是有趣的约束。
### 多账户合并
每个连接账户都有自己的配置文件和自己的 IDB。CDP 可以 snapshot 一个账户的 IDB,与另一个账户的解密合并,并 upsert 到共享的记忆文档中,例如跨三个 workspace 的统一 Slack 记忆。
## 另请参阅
* [`docs/TAURI_CEF_FINDINGS_AND_CHANGES.md`](../../docs/TAURI_CEF_FINDINGS_AND_CHANGES.md)。通知权限深度解析。
* [`CLAUDE.md`](../../CLAUDE.md)。权威的"不注入新 JS"规则。
-256
View File
@@ -1,256 +0,0 @@
---
description: 使用 WDIO + Appium 进行端到端测试。CI 和本地设置。
icon: vials
lang: zh-CN
---
# E2E 测试指南
## 概述
桌面 E2E 测试使用 **WebDriverIO (WDIO)** 通过 Appium 驱动 Tauri 应用:
| 平台 | 驱动 | 端口 | 应用格式 | 选择器 |
| --------------------------- | --------------- | ---- | ---------------- | --------- |
| **Linux / Appium Chromium** | Appium Chromium | 4723 | Debug 二进制文件 | CSS / DOM |
| **macOS / Appium Chromium** | Appium Chromium | 4723 | `.app` 包 | CSS / DOM |
OpenHuman 桌面应用目前使用 CEF 运行时(`tauri-runtime-cef`)。CI 通过 Appium Chromium driver 驱动 Linux debug 二进制;手动 macOS 和 Windows E2E 使用同一个 Chromium-driver backend。
---
## 快速开始
### Linux / Appium Chromium
```bash
# 安装 Appium 和 Chromium driver(一次性)
npm install -g appium@3
appium driver install --source=npm appium-chromium-driver
# 构建 E2E 应用
pnpm --filter openhuman-app test:e2e:build
# 运行所有流程
pnpm --filter openhuman-app test:e2e:all:flows
# 运行单个 spec
bash app/scripts/e2e-run-spec.sh test/e2e/specs/smoke.spec.ts smoke
```
在无头 Linux 上,harness 在 **Xvfb** 虚拟显示下运行。
### macOS / Appium Chromium
```bash
# 安装 Appium + Chromium driver(一次性,需要 Node 24+
npm install -g appium@3
appium driver install --source=npm appium-chromium-driver
# 构建 .app 包
pnpm --filter openhuman-app test:e2e:build
# 运行所有流程
pnpm --filter openhuman-app test:e2e:all:flows
```
### macOS 上的 Docker(本地运行 Linux harness
使用 Docker 从 macOS 运行相同的基于 Linux 的 harness。
```bash
# 构建 + 运行所有 E2E 流程
docker compose -f e2e/docker-compose.yml run --rm e2e
# 先构建应用(如需要)
docker compose -f e2e/docker-compose.yml run --rm e2e \
pnpm --filter openhuman-app test:e2e:build
# 运行单个 spec
docker compose -f e2e/docker-compose.yml run --rm e2e \
bash app/scripts/e2e-run-spec.sh test/e2e/specs/smoke.spec.ts smoke
```
需要 Docker Desktop 或 Colima。仓库通过 bind mount 挂载,因此构建在运行之间持久化。
---
## 架构
### 平台检测
`app/test/e2e/helpers/platform.ts` 导出:
- `isTauriDriver()`legacy shim,现在对支持 DOM 的 Chromium session 始终返回 `true`
- `isMac2()`legacy shim,现在始终返回 `false`
- `supportsExecuteScript()``true`,因为 Chromium driver 在所有平台都支持 `browser.execute()`
### 元素辅助函数
`app/test/e2e/helpers/element-helpers.ts` 提供统一 API
| 辅助函数 | Appium Chromium |
| ------------------------- | -------------------------------------------- |
| `waitForText(text)` | DOM 文本内容上的 XPath |
| `waitForButton(text)` | `button` / `[role="button"]` XPath |
| `clickText(text)` | 标准 `el.click()` |
| `clickNativeButton(text)` | button 上的标准 `el.click()` |
| `clickToggle()` | `[role="switch"]` / `input[type="checkbox"]` |
| `waitForWindowVisible()` | 窗口句柄检查 |
| `waitForWebView()` | `document.readyState` 检查 |
| `hasAppChrome()` | 窗口句柄检查 |
| `dumpAccessibilityTree()` | HTML 页面源码 |
### 稳定的测试 ID
优先为 E2E spec 点击或轮询的 UI affordance 使用稳定的 `data-testid` hook。使用分类法 `<surface>-<element>-<id?>`,例如:
- `cron-jobs-panel``cron-refresh`
- `cron-job-row-<jobId>``cron-job-toggle-<jobId>``cron-job-run-<jobId>``cron-job-view-runs-<jobId>``cron-job-remove-<jobId>`
- `settings-nav-<routeId>`
- `skill-row-<skillId>``skill-install-<skillId>``skill-uninstall-<skillId>`
- `thread-row-<threadId>``new-thread-button``send-message-button`
- `onboarding-next-button`
当 spec 瞄准这些 hook 之一时,使用 `element-helpers.ts` 中的 `waitForTestId(testId)``clickTestId(testId)`。对行/动作发现保留文本选择器,对用户可见文案断言也保留文本选择器。
### 深度链接辅助函数
`app/test/e2e/helpers/deep-link-helpers.ts` 处理 auth 深度链接:
- **Appium Chromium**:所有平台都使用 `browser.execute(window.__simulateDeepLink(url))`
- **macOS fallback**`macos: deepLink` 扩展命令,然后 `open -a ...`
对于发布候选版,在触碰 CEF preflight、单实例或深度链接启动代码时,还要在 Linux 或 macOS 上运行一次手动 secondary-instance 冒烟测试:
1. 正常启动 OpenHuman 并保持运行。
2. 通过 OS opener 触发 `openhuman://auth?token=e2e-token&key=auth`
3. 确认已运行的窗口接收到回调,且不会启动第二个完整的 CEF 实例。
4. 确认 secondary 进程干净退出,没有 CEF 缓存锁错误。
这捕捉了一类回归:secondary 进程在 Tauri 的深度链接转发路径安装之前,于 CEF 缓存 preflight 期间退出。
### 编写跨平台 spec
1. 在 spec 中使用 `element-helpers.ts` 中的**辅助函数**,永远不要使用原始的 `XCUIElementType*` 选择器
2. 使用 **`clickNativeButton(text)`** 代替内联 button-clicking 代码
3. 使用 **`hasAppChrome()`** 代替检查 `XCUIElementTypeMenuBar`
4. 使用 **`waitForWebView()`** 代替检查 `XCUIElementTypeWebView`
5. 对于仅 macOS 的测试,使用 `process.platform` 守卫或单独的 spec 文件
6. 对 hash 路由使用 `navigateViaHash(route)`;它等待 hash、`document.readyState` 和挂载的 React root 后返回。在 onboarding 之后,`walkOnboarding()` 也等待 `#/home` 加上 Home 页面标记,然后 spec 才会导航到别处。
---
## 环境变量
| 变量 | 默认值 | 说明 |
| --------------------------- | ---------- | ----------------------------------------------- |
| `APPIUM_PORT` | `4723` | Appium 服务器端口 |
| `E2E_MOCK_PORT` | `18473` | Mock 后端服务器端口 |
| `OPENHUMAN_WORKSPACE` | (临时目录) | 应用工作区目录 |
| `OPENHUMAN_SERVICE_MOCK` | `0` | 启用服务 mock 模式 |
| `OPENHUMAN_E2E_MODE` | 未设置 | 启用破坏性测试支持 RPC;E2E runner 将其设为 `1` |
| `OPENHUMAN_E2E_AUTH_BYPASS` | 未设置 | 启用 JWT 绕过认证 |
| `DEBUG_E2E_DEEPLINK` | (verbose) | 设为 `0` 以静默深度链接日志 |
| `E2E_FORCE_CARGO_CLEAN` | 未设置 | E2E 构建前强制 cargo clean |
---
## CI 工作流
### Push / PR 检查
默认的 pull-request 门禁是 `.github/workflows/pr-ci.yml`。它先构建一个 Linux E2E 兼容的桌面 artifact,然后并行运行 Linux Appium/Chromium `mega-flow` lane、Playwright web lane、Rust 和 coverage jobs。
macOS 和 Windows 桌面 E2E 不会在每个 PR 上运行。需要跨平台桌面信号时,请使用手动触发的 E2E workflow 或 release pretest workflow。
### macOS / Appium Chromium
macOS/Appium Chromium 可用于本地运行,也可以通过手动触发的 E2E workflow 运行:
1. 安装 Appium + Chromium driver
2. 构建 `.app`
3. 运行所有 E2E 流程
---
## 故障排除
### Linux"WebView not ready" 超时
对于默认 CEF 运行时,这通常意味着旧的本地 runner 正试图通过 WebKitWebDriver 驱动 CEF-backed WebView。当前 CI 在 Linux 上使用 Appium Chromium driver;请使用 `app/scripts/e2e-run-session.sh` 或 PR CI workflow 作为受支持的 Linux 路径。
确保 `DISPLAY` 已设置且 Xvfb 正在运行:
```bash
export DISPLAY=:99
Xvfb :99 -screen 0 1280x1024x24 &
```
还要确保 dbus 已启动(webkit2gtk 需要):
```bash
eval $(dbus-launch --sh-syntax)
```
### Linux:找不到 Appium Chromium driver
```bash
npm install -g appium@3
appium driver install --source=npm appium-chromium-driver
```
### macOS:深度链接在 `tauri dev` 中不工作
深度链接需要 `.app` 包。请改用 `pnpm tauri build --debug --bundles app`
### Docker:首次运行构建很慢
首次 Docker 构建会编译 Rust 并安装 E2E harness 依赖。后续运行使用缓存层。Cargo registry 和 git 源通过 Docker volume 缓存。
## SpecNotifications
**文件**`app/test/e2e/specs/notifications.spec.ts`
通过实时 core sidecar 和 Notifications UI 页面测试 notification RPC 方法:
- `notification_ingest`,通过 core RPC 创建新通知
- `notification_list`,验证摄入的通知被返回
- `notification_mark_read`,将通知标记为已读
- `notification_stats`,检查聚合统计形状
- UINotifications 页面渲染集成通知部分(`[data-testid="integration-notifications-section"]`
- UINotifications 页面显示 System Events 部分(`[data-testid="system-events-section"]`
**运行**
```bash
bash app/scripts/e2e-run-spec.sh test/e2e/specs/notifications.spec.ts notifications
```
**平台说明**RPC 测试(`notification_ingest``notification_list``notification_mark_read``notification_stats`)通过统一的 Appium Chromium backend 运行。UI 断言需要 `browser.execute()` 支持,当前 backend 在所有平台都支持。
---
## Agent 可观测的工件流
对于一种规范的、可检查的 run,将截图、页面源码 dump 和 mock 请求日志写入磁盘:
```bash
bash app/scripts/e2e-agent-review.sh
```
工件落在 `app/test/e2e/artifacts/<timestamp>-agent-review/`。完整详情 + 辅助 API[`AGENT-OBSERVABILITY.md`](agent-observability.zh-CN.md)。任何失败的测试都会触发 `wdio.conf.ts``afterTest` hook,将 `failure-*.png` + `failure-*.source.xml` 写入同一运行目录。
---
## Rust 推理提供商 E2E
这些测试(`tests/inference_provider_e2e.rs`)使用 **wiremock** 模拟 HTTP upstream,不需要实时 LLM API 调用。它们覆盖 OpenAI 兼容聊天、Anthropic 认证风格、每模型温度抑制、Ollama 本地提供商和 `/v1` HTTP 端点认证层。
```bash
# 本地:
bash scripts/test-rust-inference-e2e.sh
# 通过 DockerLinux,与 CI 相同镜像):
docker compose -f e2e/docker-compose.yml run --rm inference-e2e
```
-240
View File
@@ -1,240 +0,0 @@
---
description: 如何从源码构建 OpenHuman —— 工具链、vendored Tauri CLI 和本地桌面构建。
icon: wrench
lang: zh-CN
---
# 构建与安装 OpenHuman
本指南涵盖完整的桌面/源码安装路径和发布安装包。
如果你只需要在新机器上运行仓库根目录的 Rust crate,请使用[构建 Rust 核心](building-rust-core.zh-CN.md)。该页面记录了固定的 Rust 工具链、OS 包前置条件以及 `openhuman-core` 的精确 `cargo` 命令。
本指南涵盖两条路径:
1. 从源码构建并编译 OpenHuman
2. 安装最新的稳定发布二进制文件
## 前置条件
- `git`
- Node.js 24 或更高版本(见 `app/package.json`
- `pnpm@10.10.0`(见根目录 `package.json``packageManager` 字段)
- 通过 `rustup` 安装的 Rust 1.93.0,含 `rustfmt``clippy`(见 `rust-toolchain.toml`
- CMake,原生 Rust 依赖所需
- `app/src-tauri/vendor/` 下的 Git 子模块,vendored CEF-aware Tauri CLI 所需
- 平台桌面构建工具:macOS 上的 Xcode Command Line Tools,或 Linux 上的 Tauri GTK/WebKit/AppIndicator 包集合
macOS Homebrew 快速开始:
```bash
brew install node@24 pnpm rustup-init cmake
rustup toolchain install 1.93.0 --profile minimal
rustup component add rustfmt clippy --toolchain 1.93.0
```
Arch Linux 快速开始:
```bash
sudo pacman -S --needed nodejs npm rustup cmake base-devel clang openssl \
alsa-lib xdotool libxtst libxi libevdev gtk3 webkit2gtk-4.1 \
libayatana-appindicator librsvg patchelf nss nspr at-spi2-core \
libcups libdrm libxkbcommon libxcomposite libxdamage libxfixes \
libxrandr mesa pango cairo libxshmfence
npm install -g pnpm@10.10.0
rustup toolchain install 1.93.0 --profile minimal
rustup component add rustfmt clippy --toolchain 1.93.0
```
## 从源码构建(本地编译)
从仓库根目录运行:
```bash
# 1) 克隆并进入仓库
git clone https://github.com/tinyhumansai/openhuman.git
cd openhuman
# 2) 获取 vendored Tauri/CEF 源码
git submodule update --init --recursive
# 3) 安装 JS 依赖(workspace
pnpm install
# 4) 构建桌面应用产物
pnpm build
```
本地开发(而非生产构建):
```bash
# 仅 Web UI 开发
pnpm dev
# 使用 vendored Tauri/CEF CLI 的桌面应用开发:从 workspace 根目录运行
pnpm --filter openhuman-app dev:app
```
## 安装最新稳定版(macOS/Linux x64
主要安装命令:
```bash
curl -fsSL https://raw.githubusercontent.com/tinyhumansai/openhuman/main/scripts/install.sh | bash
```
安装器行为:
- 解析你平台的最新稳定 OpenHuman 发布版本
- 可用时验证产物摘要
- 本地安装(默认不需要 sudo
- macOS:将 `OpenHuman.app` 安装到 `~/Applications`
- Linux x64:将 AppImage 安装为 `~/.local/bin/openhuman` 并写入桌面入口
实用 flag
```bash
# 预览操作而不写入文件
curl -fsSL https://raw.githubusercontent.com/tinyhumansai/openhuman/main/scripts/install.sh | bash -s -- --dry-run
```
## Windows(最新稳定版)
使用 PowerShell
```powershell
irm https://raw.githubusercontent.com/tinyhumansai/openhuman/main/scripts/install.ps1 | iex
```
Windows 安装器行为:
- 解析最新稳定版
- 下载 x64 的 MSI/EXE
- 可用时验证摘要
- 在安装包支持的情况下执行按用户安装
## ARM Linux 构建(aarch64
ARM Linux 构建由于 CEF 和 GTK 依赖需要特殊处理。
### 前置条件
```bash
# 安装 xvfb 用于 headless 构建/测试
sudo apt install xvfb
```
### 构建
```bash
cd app
pnpm tauri build --target aarch64-unknown-linux-gnu
```
### 运行 ARM 二进制文件
该二进制文件需要设置 CEF 库路径:
### 选项 1 —— 直接调用
```bash
REL_DIR=app/src-tauri/target/aarch64-unknown-linux-gnu/release
CEF_DIR=$(ls -d "$REL_DIR"/build/cef-dll-sys-*/out/cef_linux_aarch64 2>/dev/null | head -n1)
export LD_LIBRARY_PATH="$CEF_DIR:$REL_DIR/deps:$REL_DIR${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
"$REL_DIR/OpenHuman" --no-sandbox
```
### 选项 2 —— Wrapper 脚本(推荐)
保存到 `~/bin/openhuman` 并赋予可执行权限(`chmod +x ~/bin/openhuman`):
```bash
#!/bin/bash
REL_DIR=/path/to/app/src-tauri/target/aarch64-unknown-linux-gnu/release
CEF_DIR=$(ls -d "$REL_DIR"/build/cef-dll-sys-*/out/cef_linux_aarch64 2>/dev/null | head -n1)
export LD_LIBRARY_PATH="$CEF_DIR:$REL_DIR/deps:$REL_DIR${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
exec "$REL_DIR/OpenHuman" --no-sandbox "$@"
```
### DEB 包安装
```bash
DEB_FILE=$(ls app/src-tauri/target/aarch64-unknown-linux-gnu/release/bundle/deb/OpenHuman_*_arm64.deb | head -n1)
sudo dpkg -i "$DEB_FILE"
```
### GTK 初始化修复
ARM 构建需要 GTK 在 Tauri 创建系统托盘之前初始化。这在 `vendor/tauri-cef/crates/tauri-runtime-cef/src/lib.rs` 中处理:
```rust
// CEF 初始化后,添加:
#[cfg(target_os = "linux")]
{
gtk::init().ok();
}
```
如果托盘初始化失败并提示 "GTK has not been initialized",请确保此修复已到位后重新构建。
全平台手动下载链接:
- 网站:https://tinyhuman.ai/openhuman
- 最新发布:https://github.com/tinyhumansai/openhuman/releases/latest
## 故障排除
### macOS`pnpm dev:app` 退出并提示 "CEF cache is held by another OpenHuman instance"
**症状**
`pnpm dev:app`(或 Tauri 壳层的任何 debug 构建)在窗口出现前退出,提示类似:
```text
[openhuman] CEF cache at /Users/<you>/Library/Caches/com.openhuman.app/cef is held by another OpenHuman instance (host <hostname>, pid 12345).
Quit the running instance and try again.
Workaround:
pkill -f "OpenHuman.app/Contents"
pkill -f "openhuman-core"
```
**原因**
CEFChromium Embedded Framework)通过 `~/Library/Caches/com.openhuman.app/cef` 下的 `SingletonLock` 符号链接对其用户数据目录持有独占锁。已安装的 `.app` 包和开发二进制文件使用相同的标识符(`com.openhuman.app`),因此它们无法并排运行。如果没有 preflight,`cef::initialize` 会返回失败,而 vendored `tauri-runtime-cef` 会以 Rust 回溯和无可操作消息的方式 panic(这是 preflight 落地前的 issue #864)。
**修复**
退出另一个 OpenHuman 实例并重新运行。最快路径:
```bash
pkill -f "OpenHuman.app/Contents"
pkill -f "openhuman-core"
pnpm dev:app
```
如果锁是由崩溃进程留下的(PID 已不存在),preflight 会自动移除陈旧的 `SingletonLock`,开发启动将继续,无需手动清理。
**已知限制**
开发和发布构建仍然共享 `com.openhuman.app` 作为缓存标识符。将开发隔离到单独的 `com.openhuman.app.dev` 缓存需要修改 vendored `tauri-runtime-cef`(缓存路径在运行时内部从 bundle 标识符构建,未暴露给 openhuman 壳层)。作为 #864 的后续跟踪。
### 核心端口上的陈旧 `openhuman` RPC 进程
**症状**
之前的 Tauri 构建或 `openhuman-core run` harness 在 `OPENHUMAN_CORE_PORT`(默认 `7788`)上留下了一个监听进程。在 issue #1130 之前,新的 Tauri 构建会静默附加到该监听器,导致版本漂移,以及新构建的 `OPENHUMAN_CORE_TOKEN` 不匹配时出现 401。
**当前行为(issue #1130**
`core_process::ensure_running` 现在在启动时探测端口:
- 如果 `GET /` 将监听器识别为 OpenHuman 核心(JSON body 含 `"name": "openhuman"`),则将其视为之前运行的陈旧进程并主动终止(Unix 上 `SIGTERM`750ms 后 `SIGKILL`Windows 上 `taskkill /F /T /PID`)。Tauri 主机随后会生成自己的全新嵌入式核心。
- 如果监听器是其他东西(或不讲 HTTP),启动会大声失败,并在日志中显示冲突,而非静默附加。
- 设置 `OPENHUMAN_CORE_REUSE_EXISTING=1` 以选择回到遗留的 attach-to-anything 行为,在将 `openhuman-core run` 作为手动调试 harness 运行时很有用。
**手动清理(仍然有效)**
```bash
pkill -f "OpenHuman.app/Contents"
pkill -f "openhuman-core"
```
@@ -1,128 +0,0 @@
---
lang: zh-CN
---
# Polymarket 集成(读取 + 交易)
本文档描述 issue #1398 的 Polymarket 集成。
## 范围
`polymarket` 工具现在支持以下 API 上的市场浏览和交易工作流:
- Gamma API (`https://gamma-api.polymarket.com`)
- CLOB API (`https://clob.polymarket.com`)
支持的读取操作:
- `list_markets`
- `get_market`
- `list_events`
- `get_orderbook`
- `get_price`
- `get_positions`
- `get_balance`
- `get_open_orders`
- `get_usdc_allowance`
支持的写入操作:
- `place_order`
- `cancel_order`
## 架构
实现位于 `src/openhuman/tools/impl/network/polymarket.rs`,辅助模块包括:
- `clob_auth.rs`L1 凭据派生 + L2 HMAC 头
- `polymarket_orders.rs`EIP-712 订单类型数据签名
关键运行时行为:
- Layer-2 API 凭据在首次认证调用时派生并缓存。
- 派生凭据持久化到 `integrations.polymarket.derived_clob_credentials`(在 secret-store 迁移落地前使用明文配置 fallback)。
- 下单前获取 `GET /nonce?user=<eoa>` 以避免重放/nonce 不匹配。
- USDC.e 授权通过 Polygon `eth_call` 对 ERC-20 `allowance(owner, spender)` 进行读取。
## 认证与签名流程
### L1 握手(一次性引导)
- 使用 Polygon chain id `137` 签署 CLOB `ClobAuth` EIP-712 payload。
- 调用 `POST /auth/api-key`;如需,fallback 到 `GET /auth/derive-api-key`
- 持久化返回的 `{ apiKey, secret, passphrase }` 以供 L2 使用。
### L2 认证请求
每个认证的 CLOB 请求签署:
- `timestamp + method + request_path (+ POST 的 body)`
Headers
- `POLY_ADDRESS`
- `POLY_SIGNATURE`
- `POLY_TIMESTAMP`
- `POLY_NONCE: 0`
- `POLY_API_KEY`
- `POLY_PASSPHRASE`
### 订单签名
`place_order` 使用以下 domain 签署 EIP-712 订单:
- name: `Polymarket CTF Exchange`
- version: `1`
- chain id: `137`
- verifying contract: `integrations.polymarket.clob_exchange_contract`
## 权限
写入操作目前由显式的临时审批 flag 保护。
- `place_order``cancel_order` 需要 `approved=true`
- 如果省略或 `false`,工具返回:
- `Polymarket write requires explicit user approval. Re-invoke with arguments.approved = true after confirming with the user.`
这是临时的,直到 #1339 的共享审批门禁集成进来。
## 配置
配置路径:`integrations.polymarket`
字段:
- `enabled`(默认 `false`
- `gamma_base_url`(默认 `https://gamma-api.polymarket.com`
- `clob_base_url`(默认 `https://clob.polymarket.com`
- `timeout_secs`(默认 `15`
- `eoa_address`(可选默认用户地址)
- `polygon_rpc_url`(默认 `https://polygon-rpc.com`
- `usdc_contract`(默认 `0x2791Bca1f2de4661ED88A30C99A7a9449Aa84174`
- `clob_exchange_contract`(默认 `0x4bFb41d5B3570DeFd03C39a9A4D8dE6Bd8B8982E`
- `derived_clob_credentials`(可选缓存的 L2 凭据)
## USDC Allowance 合约
`get_usdc_allowance` 仅报告授权状态;不改变链上状态。
- TokenPolygon 上的 USDC.e (`0x2791Bca1f2de4661ED88A30C99A7a9449Aa84174`)
- SpenderPolymarket exchange (`0x4bFb41d5B3570DeFd03C39a9A4D8dE6Bd8B8982E`)
如果授权不足,必须单独执行审批(wallet 工具 / 显式用户审批流程)。
## 错误与重试行为
- 4xx 错误视为客户端错误,不重试。
- 429 和 5xx 错误视为瞬态错误,最多重试 3 次。
- 退避固定为每次重试间隔 500ms。
- 超时表现为显式的 deadline 错误。
## 测试策略
单元测试位于 `src/openhuman/tools/impl/network/polymarket_tests.rs` 及辅助模块测试中。
- 现有读取路径和重试行为测试保持覆盖。
- 新增认证读取操作、写入审批门禁和 Polygon 授权读取的覆盖。
- `clob_auth.rs` 测试覆盖 HMAC/头 fixture 行为。
- `polymarket_orders.rs` 测试覆盖 domain 和确定性签名 fixture 行为。
-152
View File
@@ -1,152 +0,0 @@
---
description: 将 OpenHuman Core 作为只读 stdio Model Context Protocol 服务器运行。
icon: plug
lang: zh-CN
---
# MCP 服务器
OpenHuman Core 可以作为可选的 stdio MCP 服务器运行,供 Claude Desktop、Cursor 或 Zed 等本地 MCP 客户端使用。
```bash
openhuman-core mcp
```
该命令不会启动 HTTP JSON-RPC 服务器。它从 stdin 读取换行分隔的 JSON-RPC 2.0 消息,并将 MCP 响应写入 stdout。日志输出到 stderr;添加 `--verbose` 以获得调试输出。
## 客户端来源
`initialize` 期间,MCP 服务器捕获 stdio 会话的 `params.clientInfo.name`。名称通过以下方式规范化:修剪首尾空白,转换为小写,将每个非 ASCII 字母数字字符序列替换为单个连字符,然后修剪首尾连字符。例如,`Claude Desktop` 变为 `claude-desktop``Cursor` 变为 `cursor``Windsurf` 变为 `windsurf`
如果客户端省略了 `clientInfo.name`、发送空值,或发送一个规范化后结果为空的名称,会话会回退到裸的 `mcp` 来源标签。可写的 MCP 工具应使用此会话来源标签作为记忆来源,以便旧客户端保持现有的 `mcp` 行为,而可识别客户端可以作为 `mcp:<client>` 写入。
## 工具
MCP 表面经过精心设计为只读,并通过现有的控制器注册表以及核心安全策略的读取门禁:
| MCP 工具 | 背后的 RPC | 用途 |
| --- | --- | --- |
| `searxng_search`* | `openhuman.tools_searxng_search` | 搜索配置的自托管 SearXNG 实例。 |
| `memory.search` | `openhuman.memory_tree_search` | 对记忆树块进行关键词搜索。 |
| `memory.recall` | `openhuman.memory_tree_recall` | 对记忆树摘要/块进行语义召回。 |
| `tree.read_chunk` | `openhuman.memory_tree_get_chunk` | 读取搜索或召回返回的一个块。 |
| `tree.browse` | `openhuman.memory_tree_list_chunks` | 分页块列表,支持来源/实体/时间过滤。 |
| `tree.top_entities` | `openhuman.memory_tree_top_entities` | 引用最多的规范化实体,可选按类型过滤。 |
| `tree.list_sources` | `openhuman.memory_tree_list_sources` | 不同的摄入来源及其块计数和最后活动时间戳。 |
* 仅在启用 SearXNG 时存在 `searxng_search`
`searxng_search` 在启用 SearXNG 时加入 MCP 目录。它接受 `query`、可选的 `categories``web``news``images`)、可选的 `language`,以及可选的 `max_results`1-50)。
`memory.search``memory.recall` 接受 `query` 加可选的 `k`(默认 10,上限 50)。`tree.read_chunk` 接受 `chunk_id``tree.browse` 接受可选的 `source_kinds``source_ids``entity_ids``since_ms``until_ms``query``k``offset``tree.top_entities` 接受可选的 `kind``k``tree.list_sources` 接受可选的 `user_email_hint`
`config.toml` 或通过环境变量启用 SearXNG:
```toml
[searxng]
enabled = true
base_url = "http://localhost:8080"
max_results = 10
default_language = "en"
timeout_seconds = 10
```
```bash
OPENHUMAN_SEARXNG_ENABLED=true
OPENHUMAN_SEARXNG_BASE_URL=http://localhost:8080
OPENHUMAN_SEARXNG_MAX_RESULTS=10
OPENHUMAN_SEARXNG_DEFAULT_LANGUAGE=en
OPENHUMAN_SEARXNG_TIMEOUT_SECONDS=10
```
## 资源
MCP 服务器将内置提示词资产作为静态资源暴露出来。支持 `resources/list``resources/read` 的客户端可以在不执行任何工具调用的情况下,直接查看完整的智能体个性定义和子智能体提示词模板。
### 能力声明
`initialize` 响应包含以下内容:
```json
{
"capabilities": {
"tools": {},
"resources": { "subscribe": false, "listChanged": false }
}
}
```
### URI 方案
| URI | 内容 |
| --- | --- |
| `openhuman://prompts/identity` | `IDENTITY.md` — 核心智能体身份定义 |
| `openhuman://prompts/soul` | `SOUL.md` — 核心智能体个性与价值观 |
| `openhuman://prompts/user` | `USER.md` — 用户档案上下文 |
| `openhuman://prompts/agents/<id>` | 18 个内置子智能体各自的 `<id>/prompt.md` |
所有资源的 `mimeType` 均为 `"text/markdown"`
### 目录一致性
单元测试 `catalog_mirrors_builtins` 会将资源目录与 `loader.rs` 中的 `BUILTINS` 切片进行交叉验证。若新增内置子智能体而未在目录中添加对应条目,该测试将失败,从而阻断 CI。
### 资源模板
由于目录完全静态——所有 URI 均为具体地址,不存在模板化——`resources/templates/list` 始终返回空的 `resourceTemplates` 数组。该处理器是为了 MCP 规范合规性:当客户端在看到 `resources` 能力后探测 `resources/templates/list` 时,会获得格式良好的结果而非 `-32601 Method not found`
### 冒烟测试
```bash
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"smoke","version":"0"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"resources/list"}' \
'{"jsonrpc":"2.0","id":3,"method":"resources/templates/list"}' \
'{"jsonrpc":"2.0","id":4,"method":"resources/read","params":{"uri":"openhuman://prompts/identity"}}' \
| openhuman-core mcp
```
## 工具注册表
HTTP JSON-RPC 服务器还暴露一个只读的全局工具注册表,供需要发现元数据而不打开 MCP stdio 会话的智能体和仪表板使用:
| RPC 方法 | 用途 |
| --- | --- |
| `openhuman.tool_registry_list` | 列出 MCP stdio 工具和控制器支持的工具,包含稳定的 `tool_id`、路由、版本、输入/输出 schema、允许的智能体、标签、启用状态和健康状况。 |
| `openhuman.tool_registry_get` | 通过 `tool_id` 返回一个注册表条目,例如 `memory.search``tools.web_search`。 |
| `openhuman.tool_registry_diagnostics` | 返回脱敏的清单统计、疑似写入面、策略面以及外部能力提供方诊断信息。 |
注册表仅用于发现。它不改变工具分派或权限检查;MCP 调用仍通过 `tools/call`,控制器支持的工具仍通过其现有的 JSON-RPC 方法路由。
### 外部能力提供方
OpenHuman 可以在 `config.toml` 中记录可信外部能力提供方。这只是治理元数据:不会安装包、执行远程代码,也不会绕过现有 MCP/控制器分派路径。
```toml
[[capability_providers]]
id = "Acme Tools"
display_name = "Acme Tools"
source_uri = "https://example.com/openhuman/acme-tools"
source_digest = "sha256:abc123"
trust_state = "trusted"
enabled = true
```
Provider id 会在策略检查前规范化。例如 `Acme Tools` 会变成 `acme-tools`;规范化后重复的 id 会被拒绝。只有同时满足 `enabled = true``trust_state = "trusted"` 的提供方,才会被后续准入检查视为可用。没有 provider 配置时保持旧行为:provider 注册表为空,现有工具不会被隐藏。
## 冒烟测试
```bash
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"smoke","version":"0"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
| openhuman-core mcp
```
响应应包含来自 `initialize``capabilities.tools` 和来自 `tools/list` 的精选工具名称。成功的运行向 stdout 写入恰好两行紧凑的 JSON 响应;`notifications/initialized` 消息是通知,没有响应。
```json
{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{},"resources":{"subscribe":false,"listChanged":false}},"serverInfo":{"name":"openhuman-core","version":"<crate version>"},"instructions":"..."}}
{"jsonrpc":"2.0","id":2,"result":{"tools":[{"name":"memory.search",...},{"name":"memory.recall",...},{"name":"tree.read_chunk",...},{"name":"tree.browse",...},{"name":"tree.top_entities",...},{"name":"tree.list_sources",...}]}}
```
@@ -1,81 +0,0 @@
---
description: 发布节奏、版本策略、OAuth 与安装包规则。发布是如何运作的。
icon: ship
lang: zh-CN
---
# 发布策略:最新桌面构建与 OAuth
本 runbook 描述了我们如何避免用户在**过时的桌面安装包**上完成 **OAuth**(包括 **Gmail**),而规范流程始终要求**最新**发布版本。
## 分发
- [tinyhumansai/openhuman](https://github.com/tinyhumansai/openhuman/releases) 的 **GitHub Releases** 是桌面构建的主要来源。
- **Tauri 更新器**端点(见 `scripts/prepareTauriConfig.js` 和发布工作流)应将用户指向当前发布产物。
- **淘汰旧稳定版产物:** 当弃用一条发布线时,在 **GitHub Releases** 上移除或隐藏过时的安装包资源,将 **网站 / CDN** 下载链接更新为 **releases/latest**(或当前版本),刷新**更新器 manifest**(例如 Gist / `latest.json`)使其不再指向已弃用的构建,并抽查旧直接 URL 在适当位置是否被**重定向、返回 404 或 410**。验证方式:尝试从文档或书签中已知的旧资源 URL,确认它们不再提供主要安装路径。
## OAuth 最低应用版本
生产 Web 构建在**构建时**嵌入一个**最低支持的应用 semver**,使 OAuth 深度链接无法在已弃用的二进制文件上完成。每个安装包携带构建时设定的 floor;对于从不升级的用户,提高 floor 需要他们安装一个**新**的发布版本(或通过应用内更新)。可选的未来工作:仅通过**运行时** API 强制执行移动的最低版本,捆绑值仅作为 fallback。
| 变量 | 用途 |
| ------------------------------------ | --------------------------------------------------------------------------------------------------------------------- |
| `VITE_MINIMUM_SUPPORTED_APP_VERSION` | 例如 `0.51.0` —— 桌面应用必须 **≥** 此版本才能完成 `openhuman://oauth/success`。 |
| `VITE_LATEST_APP_DOWNLOAD_URL` | 可选;默认为 `https://github.com/tinyhumansai/openhuman/releases/latest`。当门禁阻止 OAuth 时打开。 |
将这些配置为 **GitHub Actions 变量**。它们必须同时存在于独立的 **`pnpm build`** 步骤和 **`.github/workflows/build-desktop.yml`** 中的 **`tauri-apps/tauri-action`** 步骤环境变量中(由 `release-production.yml` / `release-staging.yml` 调用的可重用矩阵),以便嵌入已发布安装包的 Vite bundle 包含该门禁。本地开发时保持 `VITE_MINIMUM_SUPPORTED_APP_VERSION` **未设置**(门禁禁用)。
实现:`app/src/utils/oauthAppVersionGate.ts``app/src/utils/desktopDeepLinkListener.ts`
## Gmail / Google Cloud OAuth
- Google Cloud Console 中的 **Redirect URIs** 必须匹配**当前**后端 + 隧道回调路径。
- 桌面 scheme`openhuman://`)是稳定的;当 `VITE_MINIMUM_SUPPORTED_APP_VERSION` 设置时,**已安装的二进制文件**必须满足最低版本。
## 发布清单(避免回归)
1. 按照现有版本工作流提升 `app/package.json``app/src-tauri/tauri.conf.json`(以及根目录 `Cargo.toml` / core)的版本。
2. 当弃用对旧安装包的支持时,在该发布**之前**或**同时**将 **`VITE_MINIMUM_SUPPORTED_APP_VERSION`** 设置为新的 floor(仓库 Actions 变量 + 上述两个工作流步骤)。
3. 从用户可见表面(GitHub Release 资源、网站、CDN、更新器 feed)移除、重定向或淘汰旧稳定版安装包和陈旧**更新器**条目。确认已弃用的资源无法从默认安装/更新流程中访问。
4.**releases/latest** 的全新安装上冒烟测试 **Gmail 连接**
5. 完成[手动冒烟清单](../../docs/RELEASE-MANUAL-SMOKE.md),然后将完成的签字块(逐字复制,每个已勾选项目保持勾选)粘贴到发布 PR 描述中,然后再打 tag。
## 工作流:staging vs. production
两个一等 GitHub Actions 工作流,每个环境一个。按意图选择,而非切换 flag。
| 工作流 | 分支 | 提升 | 推送的 Tags | 并发组 | 使用场景 |
| ------------------------------------------------------- | --------- | ------- | -------------------------- | ----------------------- | --------------------------------------------------------------------- |
| [`release-staging.yml`](../../.github/workflows/release-staging.yml) | `main` | 仅 `patch` | `v<version>-staging` | `release-staging` | 为 QA 切割 staging 构建。运行频繁;semver 移动范围窄。 |
| [`release-production.yml`](../../.github/workflows/release-production.yml) | `main` | `patch` / `minor` / `major`(仅在 `main_head` 上) | `v<version>` | `release-production` | 提升已验证的 staging tag,或从 `main` HEAD 热修。 |
两个流程使用的矩阵构建 / 签名 / Sentry-DIF / 产物上传流水线位于 [`.github/workflows/build-desktop.yml`](../../.github/workflows/build-desktop.yml) 中,作为 `workflow_call` 可重用工作流。上述两个顶层工作流拥有 ref 解析、版本提升、tagging 和发布/清理;构建本身是共享的。
### 切割 staging 构建
1. 通过 `workflow_dispatch``main` 运行 **Release (Staging)**
2. 工作流在 `main` 上提升 `patch`commit `chore(staging): vX.Y.Z`,推送分支,并在该 commit 上创建不可变的 `vX.Y.Z-staging` tag。
3. 构建矩阵从 **tag**(而非 main HEAD)运行,因此即使 `main` 已经前进,rerun 也会重建字节相同的内容。
4. 失败时 staging tag 会被自动删除;`main` 上的提升 commit 保留,因此下一次切割从 `vX.Y.(Z+1)` 继续。
没有单独的 `staging` 分支,staging 切割和 production 提升都存在于 `main` 上。两者仅通过 tag 后缀(`-staging` vs 无)和创建工作流来区分。
### 提升为 production(默认流程)
1. 通过 `workflow_dispatch``release_source = staging_tag`(默认)运行 **Release Production**
2. 留空 `staging_tag` 以提升最新的 `v*-staging`,或传入显式 tag(例如 `v1.2.4-staging`)以固定版本。
3. 工作流去除 `-staging` 后缀,在同一 commit 上创建 `v<version>`,并从该 tag 运行 production 构建矩阵。**不再提升版本**,产物复用 staging 已验证的内容。
### 从 `main` HEAD 热修
1. 通过 `workflow_dispatch``release_source = main_head` 和所需的 `release_type``patch` / `minor` / `major`)运行 **Release Production**
2. 工作流运行遗留的提升-and-tag 路径:在 `main` 上提升,commit `chore(release): vX.Y.Z`,推送,tag `vX.Y.Z`,构建。
3. 仅当需要不经过 staging 的 production-only 修复时才使用此路径。
### Tag 策略与回滚
- **命名。** Staging tag 使用 SemVer 预发布后缀 `-staging``v1.2.4-staging`),因此它们在排序上位于匹配的 production tag *之前*。提升到 production 时逐字去除后缀;两个 tag 之间捆绑安装包中嵌入的版本是相同的。
- **冲突。** 如果目标 tag 已存在于本地或 `origin` 上,两个工作流都会快速失败。通过删除陈旧 tag(仅限组织维护者)或跳过它来解决。
- **回滚(production)。** 失败的构建矩阵会触发 `cleanup-failed-release`,删除草稿 GitHub Release 和 `v<version>` tag。它从中提升的 staging tag 保持不变,修复后可以重新提升。
- **回滚(staging)。** 失败的 staging 构建会删除 `v<version>-staging` tag。`main` 上的提升 commit 保留;下一次 staging 切割从新的 patch 号继续,而不是重新使用它(我们接受 patch 号中的一个小"缺口",而不是与并发合并竞争)。
- **谁可以删除 tag。** 与 `main` 相同的写入权限。工作流驱动的清理通过工作流的 token 使用 `actions/github-script` 运行删除(GitHub App token 仅由 `prepare-build` 用于提升 commit + tag 推送);手动删除(`git push --delete origin <tag>`)需要同等的维护者权限。
@@ -1,157 +0,0 @@
---
description: OpenHuman 如何测试其产品 —— Vitest、cargo test、WDIO E2E。每种测试该放哪里。
icon: vial
lang: zh-CN
---
# 测试策略
OpenHuman 如何测试其产品。"我的测试该放哪里?"的权威答案。 companion 文档为 [`TEST-COVERAGE-MATRIX.md`](../../docs/TEST-COVERAGE-MATRIX.md)。
---
## 测试层级
| 层级 | 存放位置 | 测试内容 | 驱动方式 |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| **Rust 单元测试** | 同一 `*.rs` 文件内的 `#[cfg(test)] mod tests`,或同级 `tests.rs`,或域名下的 `tests/` 子目录(例如 `src/openhuman/channels/tests/` | 纯领域逻辑、schema、RPC handler 形态、内存状态机 | `cargo test` |
| **Rust 集成测试** | 仓库根目录的 `tests/*.rs` | 完整领域接线,含真实 Tokio 运行时、模拟外部服务、JSON-RPC 端到端(`tests/json_rpc_e2e.rs`)、领域 × 领域交互 | `pnpm test:rust`(调用 `bash scripts/test-rust-with-mock.sh` |
| **Vitest 单元测试** | 与源码共存于 `app/src/**` 下的 `*.test.ts(x)`,或 `app/src/**/__tests__/` 下 | React 组件、hook、store slice、纯工具函数、service 层适配器 | `pnpm test:unit` |
| **WDIO E2E** | `app/test/e2e/specs/*.spec.ts` | 完整桌面流程:UI → Tauri → core sidecar → JSON-RPC;用户可见行为 | Linux CI: `tauri-driver`(端口 4444)。macOS 本地: Appium Mac2(端口 4723)。详见 [E2E 测试](e2e-testing.zh-CN.md)。 |
| **手动冒烟测试** | [`docs/RELEASE-MANUAL-SMOKE.md`](../../docs/RELEASE-MANUAL-SMOKE.md) | 驱动程序无法断言的 OS 级表面:TCC 权限弹窗、Gatekeeper、代码签名、DMG 安装、OS 原生通知 | 发布切割时由人工执行,在发布 PR 中签字确认 |
---
## 决策树 —— 我的测试该放哪里?
```text
变更是否在 JSON-RPC 边界之后(在 src/ 中)?
├─ 是 —— 是否跨领域或与外部服务通信?
│ ├─ 是 → Rust 集成测试 (tests/*.rs)
│ └─ 否 → Rust 单元测试(源码旁)
└─ 否 —— 变更在 app/ 中
├─ 是纯函数、hook、slice 或独立组件?
│ └─ 是 → Vitest 单元测试 (*.test.tsx 与源码共存)
└─ 是否用户可见 且 跨越 UI ⇄ Tauri ⇄ sidecar ⇄ JSON-RPC
├─ 是 → WDIO E2E (app/test/e2e/specs/*.spec.ts)
└─ 是否 OS 级(TCC、Gatekeeper、安装、OS 通知)?
└─ 是 → 手动冒烟清单
```
如果一项变更触及多个层级,在**每个**触及的层级都写测试。不要用一层替代另一层。
---
## 失败路径要求
覆盖矩阵中的每个功能叶子节点,除了 happy path 外,**至少**还要有一个**失败 / 边界**断言。例如:
- 文件写入工具:happy = 写入了字节;failure = 路径限制拒绝。
- OAuth 流程:happy = 签发了 tokenedge = 过期刷新 token 恢复。
- 记忆存储:happy = 存储并召回;edge = 遗忘后再召回返回空。
只断言 happy path 的 spec 是不完整的。
---
## Mock 策略
- **单元 / 集成 / E2E 中禁止真实网络。** 使用共享 mock 后端(`scripts/mock-api-core.mjs``scripts/mock-api-server.mjs``app/test/e2e/mock-server.ts`)。
- 测试用 admin 端点:`GET /__admin/health``POST /__admin/reset``POST /__admin/behavior``GET /__admin/requests`
- **外部服务**Telegram、Slack、Gmail、Notion、Ollama、OpenAI 等)在 mock 后端层面被 stub;测试通过 `getRequestLog()` 断言请求形态。
- 唯一可接受的例外是记录在案的发布切割手动冒烟步骤。
---
## 确定性规则
- 禁止 wall-clock 等待,使用 `waitForApp``waitForAppReady``waitForWebView` 辅助函数,或显式的元素就绪谓词。
- 禁止共享文件系统状态,每个 E2E spec 在隔离的 `OPENHUMAN_WORKSPACE` 中运行(由 `app/scripts/e2e-run-spec.sh` 创建/清理)。
- 禁止顺序依赖的 spec,每个 spec 必须能独立通过。
- 禁止依赖绝对坐标或动画时序。
- 禁止在 tauri-driver 上通过 `browser.keys()` 使用真实键盘,通过 `browser.execute(...)` 合成(参见 `command-palette.spec.ts` 中的模式)。
---
## 现有 harness 提供的能力
- **Mock 后端引导**`app/test/e2e/mock-server.ts` 中的 `startMockServer` / `stopMockServer`
- **Auth 捷径**`helpers/deep-link-helpers.ts` 中的 `triggerAuthDeepLink` / `triggerAuthDeepLinkBypass` 跳过真实 OAuth。
- **元素辅助函数**`helpers/element-helpers.ts` 中的 `clickNativeButton``waitForWebView``clickToggle`,在 spec 中使用这些代替原始的 `XCUIElementType*` 选择器。
- **共享流程**`helpers/shared-flows.ts` 中的 `completeOnboardingIfVisible``navigateViaHash``navigateToSkills``walkOnboarding`
- **从 spec 调用 Core RPC**`helpers/core-rpc.ts` 中的 `callOpenhumanRpc`,当 UI 步骤可能脆弱时直接驱动 sidecar。
- **平台守卫**`helpers/platform.ts` 中的 `isTauriDriver``isMac2``supportsExecuteScript`
- **失败时捕获工件**`captureFailureArtifacts``wdio.conf.ts` 运行,截图 + DOM dump 输出到 `app/test/e2e/artifacts/`
---
## 命名与结构规范
- WDIO spec:端到端产品流用 `<feature-area>-flow.spec.ts`;更窄的表面用 `<feature>.spec.ts`
- Vitest 同位置:优先 `Component.tsx` + `Component.test.tsx` 同级;仅在组合多个相关测试时使用 `__tests__/`
- Rust 集成测试:文件名用 snake_case 匹配表面,JSON-RPC 驱动流用 `<feature>_e2e.rs`,跨领域用 `<feature>_integration.rs`
- 每个 `describe` / `mod tests` 块对应一个功能列表 ID 范围,如果映射不明显,在注释中链接矩阵行。
---
## 合并前门禁
开 PR 前运行。CI 会跑同一套,但本地更快:
```bash
# Rust 核心
cargo fmt --check
cargo check --manifest-path Cargo.toml
cargo clippy --manifest-path Cargo.toml -- -D warnings
cargo test --manifest-path Cargo.toml
# Tauri 壳层
cargo check --manifest-path app/src-tauri/Cargo.toml
# 前端
pnpm typecheck
pnpm lint
pnpm format:check
pnpm test:unit
# 带 mock 后端的 Rust 集成测试
pnpm test:rust
# E2E(慢 —— 仅在行为用户可见变更时运行)
pnpm test:e2e:build
bash app/scripts/e2e-run-spec.sh test/e2e/specs/<your-spec>.spec.ts <id>
```
---
## 无法被驱动程序自动化的 —— 需要手动冒烟
某些表面无法被 WDIO / Appium 驱动,因为它们跨越 OS 级信任边界或硬件路径。完整的清单 + 签字块位于 [`docs/RELEASE-MANUAL-SMOKE.md`](../../docs/RELEASE-MANUAL-SMOKE.md),该文件是每次发布必须验证内容的权威来源。涵盖示例:
- macOS TCC 权限弹窗(辅助功能、输入监控、屏幕录制、麦克风)
- Gatekeeper 首次启动签名验证
- 代码签名完整性(`codesign --verify --deep --strict`
- DMG 安装 / 拖入 Applications 流程
- 自动更新下载 + 重启
- Linux OS 原生通知 toast(无显示服务器的 driver 无法看见 Xvfb 之外的 Linux
如果一项功能没有自动化覆盖,也不在手动冒烟清单上,视为未测试,开一个覆盖缺口。
---
## 覆盖矩阵即契约
[覆盖矩阵](../../docs/TEST-COVERAGE-MATRIX.md) 中的每个功能叶子节点映射到:
1. 一个或多个测试路径,**或**
2. 一个合理的 `🚫` 并附手动冒烟条目。
当你添加 / 删除 / 重命名功能时,**在同一 PR 中更新矩阵行**。CI 将在 #965 落地后守卫此契约。
---
## 不确定时
- 尽可能把测试推到层级栈的**底层**(Rust 单元 > Rust 集成 > Vitest > WDIO)。更低层级更快、更确定、运行成本更低。
- WDIO 用于真正跨越 UI ⇄ Tauri ⇄ sidecar ⇄ JSON-RPC 的行为。不要仅仅因为 UI 存在就通过 WDIO 驱动一个可单元测试的关注点。
- 失败的 happy path 是回归。缺失的失败路径测试是缺口。两者都是 bug。
+121
View File
@@ -0,0 +1,121 @@
---
description: >-
Human-in-the-loop consent for side-effecting tool calls — the agent parks any
risky action until you approve it, and fails closed if you don't.
icon: shield-check
---
# Approval Gate
The Approval Gate is the checkpoint between the agent and the outside world. Whenever the agent wants to run a tool that has a real-world effect — post to Slack, send an email, create a calendar event, run a shell command, install a package — the gate intercepts the call, shows you exactly what's about to happen, and waits for your decision before anything runs.
It's on by default. Nothing with an external effect leaves your machine in an interactive chat without you saying yes.
***
## What triggers a prompt
Every acting tool call is classified into a **command class**, and your **autonomy tier** decides whether that class runs silently, prompts, or is blocked.
| Command class | What it covers |
| ------------- | ---------------------------------------------------------- |
| Read | Provably read-only / observational (curated allowlist) |
| Write | State-changing; the fail-closed default for anything unrecognized |
| Network | Reaches the network (curl, wget, ssh, scp, …) |
| Install | Installs an OS or global language package |
| Destructive | Catastrophic / irreversible / privilege-escalating |
The tier comes from **Settings → Agent access** (`[autonomy].level`):
| Tier | Read | Write | Network / Install / Destructive |
| -------------- | ----- | ------ | ------------------------------- |
| Read-only | Allow | Block | Block |
| Supervised *(default)* | Allow | Prompt | Prompt |
| Full | Allow | Allow | Prompt |
Anything that lands on **Prompt** is parked at the gate. `Block` is refused outright — no in-tier approval can authorize it. Classification is fail-closed: a command that isn't provably read-only is treated as at least `Write`, and across a piped command the highest class wins (so `ls | curl …` is `Network`).
***
## The flow
```
agent wants to act
classify command ──► Block ──► refused
Prompt
on "Always allow" list? ──► yes ──► run immediately
│ no
park call · persist pending row · emit approval_request
┌──────────────┬───────────────┬────────────┐
▼ ▼ ▼ ▼
Approve Always allow Deny 10-min TTL
(once) (+ allowlist) │
│ │ │ ▼
▼ ▼ ▼ Deny
run run refused (fail closed)
```
When a call is parked, an **Approval Request card** appears above the chat composer. It shows the tool name, a safe one-line summary of the action, and the (redacted) command. Three choices:
* **Approve** — run this one call.
* **Always allow** — run it, and add the tool to your `auto_approve` list so it skips the prompt next time.
* **Deny** — refuse this call.
You can also just type **yes** / **no** in chat — the reply is routed back to the parked request.
***
## Always allow
Approving with **Always allow** persists the tool name onto `[autonomy].auto_approve` (config save + live-policy reload), so the gate short-circuits to *allow* for that tool on future turns. The list ships with safe read-only tools pre-approved (`file_read`, `memory_search`, `memory_list`, `get_time`, `list_dir`, `glob`, `grep`) and is editable in **Settings → Agent access**. Remove an entry there to start being prompted again.
***
## Fail-closed behavior
Every non-approve path resolves to **Deny**:
* **Timeout** — a parked request lives for 10 minutes; if undecided it is transitioned to a terminal `deny`.
* **Persist failure** or a dropped channel — denied.
* The timeout path re-reads the stored decision first, so an approval that committed in the race still wins.
Pending requests are stored in SQLite (`{workspace_dir}/approval/approval.db`) and **survive a core restart**. After an approved tool finishes, the gate records a write-once execution outcome (success / error, error text sanitized and capped) as a durable audit trail. Everything persisted or broadcast is redacted first — PII and chat content are scrubbed and home paths stripped.
***
## Background and cron bypass
The gate is **interactive-only**. Background, triage, and cron turns carry no chat context, so there's nobody to answer a prompt — these turns are pre-authorized and pass straight through (no row, no event). Approval is only enforced for live chat turns. (The Subconscious loop has its own, separate escalation-card approval for *unsolicited* writes — see below.)
***
## Configuration & RPC
* **`OPENHUMAN_APPROVAL_GATE`** — set to `0` / `false` to skip installing the gate entirely. With no gate, `Prompt`-class calls run unprompted. On by default.
* **`[autonomy].level`** and **`[autonomy].auto_approve`** — tier and allowlist, via the `config.update_autonomy_settings` RPC or Settings → Agent access.
The `approval` controller exposes three JSON-RPC methods:
| Method | Purpose |
| ----------------------------------- | ------------------------------------------------------------------- |
| `openhuman.approval_list_pending` | The live queue of parked requests. |
| `openhuman.approval_list_recent_decisions` | Decided/executed audit rows (`limit` 1500, default 50). Surfaced in **Settings → Approval history**. |
| `openhuman.approval_decide` | Apply a decision (`approve_once` / `approve_always_for_tool` / `deny`). |
`list_pending` / `list_recent_decisions` return empty (not an error) when no gate is installed; `decide` errors when the gate is absent or the request is unknown or already decided.
***
## See also
* [Privacy & Security](privacy-and-security.md) — autonomy tiers, trusted roots, and path hardening.
* [Subconscious Loop](subconscious.md) — the background loop and its separate escalation approvals.
* [Security architecture](../developing/architecture/security.md) — the command-classification and policy internals.
+109
View File
@@ -0,0 +1,109 @@
---
description: >-
Plans, credits and saved cards over Stripe and Coinbase, plus a local
real-time dashboard for token usage, cost and budget enforcement.
icon: credit-card
---
# Billing, Cost & Usage
OpenHuman keeps two related but separate ledgers. **Billing** is what you pay the hosted backend — plans, credit top-ups, saved cards and coupons, all settled through Stripe or Coinbase. **Cost & Usage** is what the agent spends on your behalf, tracked locally per provider call so you can see (and cap) real token spend before the bill ever lands.
The first lives in the cloud; the second never leaves your workspace.
***
## Part 1 — Billing & Payments
The `billing` domain is a thin RPC adapter. It holds **no payment logic or state of its own** — every operation forwards an authenticated HTTPS call to the hosted backend (`/payments/*`, `/coupons/*`) using your stored app-session JWT, and surfaces the JSON response verbatim. Authorization, plan ownership and payment policy are all enforced backend-side. A missing or invalid session yields the backend's `401`/`403` directly; JWTs and card data are never logged.
Pre-HTTP, the adapter does light input validation only: non-empty plan/coupon/payment-method ids, a finite positive `amountUsd`, and a gateway whitelist of `stripe` / `coinbase`.
### Plans
Three tiers are offered, each with a monthly and annual interval:
| Tier | Monthly | Annual | Per-call discount vs pay-as-you-go |
| --- | --- | --- | --- |
| **Free** | $0 | $0 | None — pay-as-you-go baseline |
| **Basic** | $19.99 | $199 | 50% cheaper per call |
| **Pro** | $199.99 | $1,799.99 | 90% cheaper per call |
Higher tiers do not unlock features so much as lower the **per-call margin** over the pay-as-you-go baseline. All tiers have "access to everything"; you are buying cheaper inference, not gated capabilities.
### Payment providers
Two gateways are wired, and only two:
- **Stripe** — plan purchases (Checkout sessions), the customer billing portal, credit top-ups, saved-card management (SetupIntents) and auto-recharge.
- **Coinbase Commerce** — crypto charges, used for credit top-ups and annual billing.
`top_up_credits` and `create_coinbase_charge` default to the `stripe` gateway and `annual` interval; an empty or whitespace gateway normalises to Stripe.
### Credits, top-ups & auto-recharge
Beyond a subscription you hold a **USD credit balance**. You can read the balance, page through transaction history, and top up via either gateway. **Auto-recharge** (Stripe only) re-fills credits from a saved card when the balance runs low; you can read and update its settings, and list / add / update / delete saved cards. Adding a card creates a Stripe SetupIntent; deleting one is treated as a dangerous operation.
### Coupons
Coupon codes are redeemed against the backend (`POST /coupons/redeem`), and you can list the coupons currently redeemed on your account (`GET /coupons/me`).
### Where billing lives in the app
The desktop **Settings → Billing** panel intentionally has no embedded payment UI — it links out to the hosted web **billing dashboard**, which is the single place to manage plans, cards and invoices. The agent can also read billing state through default-ON tools (plan, balance, transactions, cards, coupons, the Stripe portal link); every money-moving or payment-method mutator ships **default-OFF** behind a `billing_writes` toggle, and card deletion is flagged dangerous.
### RPC surface
Namespace `billing`, exposed as `openhuman.billing_*` (15 methods), e.g. `billing_get_current_plan`, `billing_get_balance`, `billing_get_transactions`, `billing_purchase_plan`, `billing_top_up`, `billing_create_coinbase_charge`, `billing_get_cards`, `billing_create_setup_intent`, `billing_update_auto_recharge`, `billing_redeem_coupon`.
***
## Part 2 — Cost & Usage Dashboard
The `cost` domain is entirely local. It records every provider call's token usage and computed USD cost to an append-only JSONL file (`<workspace>/state/costs.jsonl`), keeps in-memory daily/monthly aggregates, enforces budgets, and serves a 7-day dashboard over JSON-RPC. A process-global singleton tracker is shared by the agent turn loop (which logs telemetry after each provider call) and the dashboard handlers, so each call is persisted exactly once.
### Real-time token & cost tracking
For each call, per-call cost is computed from token counts and per-million-token prices (clamping non-finite or negative prices to `0.0`). When the provider echoes an authoritative `charged_amount_usd` that value wins; otherwise OpenHuman falls back to a static pricing catalog of known models. Usage is bucketed in UTC, keyed by model, with the **provider** derived from the `provider/model` prefix. All-zero usage payloads are skipped so providers that don't report usage don't inflate the request count.
### Budgets & enforcement
Budget enforcement is configured under the `[cost]` config block:
| Setting | Default | Role |
| --- | --- | --- |
| `enabled` | `true` | Gates **enforcement only**, not telemetry |
| `daily_limit_usd` | `10.00` | Hard daily cap |
| `monthly_limit_usd` | `100.00` | Hard monthly cap |
| `warn_at_percent` | `80` | Warn threshold for `check_budget` |
`check_budget` returns `Allowed`, `Warning` (warn threshold reached) or `Exceeded` (over the daily or monthly cap). A crucial detail: **`enabled` controls enforcement, not capture.** When it is `false`, `check_budget` always returns `Allowed` and hard caps are off — but the agent still records usage unconditionally, so your spend history accumulates and you can review it *before* opting into hard caps. To hide the panel set `dashboard.enabled = false`; to clear history delete the JSONL file (it is local and never leaves the workspace).
### The 7-day dashboard
Settings → **Usage & Limits** hosts the cost dashboard (alongside background-activity controls). It renders a 7-day daily history (gap days zero-filled, oldest first), a token-usage chart, a monthly-pace projection, budget utilisation and a per-model cost breakdown. Dashboard colour-coding uses fractions of the monthly budget: bars flip to amber at the `warn_threshold` (default `0.8`) and red at the `alert_threshold` (default `0.95`). `budget_utilization` is clamped to `1.0` for display, while status is computed from the raw value. The panel polls roughly every 10 seconds and shows an "Updated Ns ago" freshness pill. A read-only fallback tracker (sharing the same JSONL file) serves the UI when the global tracker isn't yet initialised.
### RPC surface
Namespace `cost`, exposed as `openhuman.cost_*`:
| Method | Inputs | Output |
| --- | --- | --- |
| `cost_get_dashboard` | none | 7-day buckets, summary metrics, budget utilisation/status, per-model breakdown |
| `cost_get_daily_history` | `days?` (default 7, clamped 1366) | Ordered daily entries, oldest first, gaps zero-filled |
| `cost_get_summary` | none | Live session / daily / monthly cost summary |
These are also exposed as read-only, default-ON agent tools so the agent can inspect its own spend.
***
## Cost & token compression
Because cost tracks **real token counts**, anything that shrinks the prompt directly lowers spend. OpenHuman's [TokenJuice token compression](token-compression.md) reduces the tokens sent on each call, and [model routing](model-routing/README.md) sends work to the cheapest model that can handle it — both of which show up as lower bars in the dashboard and slower budget burn.
***
## See also
- [Token compression (TokenJuice)](token-compression.md)
- [Model routing](model-routing/README.md)
+81
View File
@@ -0,0 +1,81 @@
---
description: >-
Messaging platforms OpenHuman talks back to you on — inbound dispatch into the
agent loop, outbound replies and proactive delivery, and per-channel
credentials.
icon: messages-square
---
# Messaging Channels
A **channel** is a messaging platform OpenHuman uses to _talk back_ to you. This is the mirror image of an [integration](integrations/README.md): an integration is mostly a source the agent _reads from_ (your inbox, your calendar, your CRM), while a channel is a two-way conversation surface — you message the agent on a platform you already use, and the agent replies there.
Under the hood every channel implements one small Rust contract — a `send` path for outbound messages and a `listen` path for inbound ones — so the same agent loop serves Telegram, Discord, the built-in web chat, and a dozen others without per-platform branching in the core.
***
## What a channel does
Each channel does two things:
* **Inbound** — when a message arrives, the channel normalizes it into a `ChannelMessage` (sender, reply target, content, optional thread id) and hands it to the dispatch loop. Dispatch spawns or resumes an agent run, scopes its tools, and the agent works the request. Some platforms support a `/models` and `/model` command to switch the model for that sender's session; Telegram additionally supports remote-control commands.
* **Outbound** — the agent's response is sent back through the same channel to your `reply_target`, threaded when the platform supports it. Channels can also deliver **proactively** (no incoming message to reply to) when fired by a [trigger](integrations/triggers.md), a cron job, or the [subconscious loop](subconscious.md). A channel only receives proactive sends if it advertises a default delivery target; channels without one are skipped rather than posted to an empty recipient.
Channels that support it can show a typing indicator, stream progressive **draft updates**, post **threaded replies**, and add **emoji reactions** — capabilities are declared per channel, not assumed.
***
## Supported channels
OpenHuman ships **18 channel provider modules** (16 built by default plus two behind Cargo feature flags), of which **17 are real messaging platforms**`presentation` is an internal response-rendering helper for the web chat, not a platform you connect to. A separate `cli` channel serves the `openhuman-core` terminal binary. Seven channels are exposed in the Settings UI; the rest are enabled through `config.toml`.
| Channel | Direction | Inbound transport | Credential mode | In Settings UI |
| --- | --- | --- | --- | --- |
| **Telegram** | Two-way | Bot API long-poll | Connect via OpenHuman (managed DM) **or** your own BotFather token | Yes |
| **Discord** | Two-way | Gateway | Your own bot token, OAuth install, **or** managed account link | Yes |
| **Web** | Two-way | In-app | Built-in, no setup (local) | Yes |
| **iMessage** | Two-way | macOS Messages (AppleScript) | Local-only, no credentials (needs Full Disk Access) | Yes |
| **Lark / Feishu** | Two-way | WebSocket or webhook | Your own app id + secret | Yes |
| **DingTalk** | Two-way | Stream Mode WebSocket | Your own client id + secret | Yes |
| **元宝 (Yuanbao)** | Two-way | WebSocket | Your own AppID + AppSecret | Yes |
| **Slack** | Two-way | Events/socket | Your own bot token | `config.toml` |
| **WhatsApp** | Two-way | Meta Cloud webhook | Your own access token | `config.toml` |
| **IRC** | Two-way | Persistent socket | Your own server/nick | `config.toml` |
| **Matrix** | Two-way | Sync loop | Your own homeserver creds (feature `channel-matrix`) | `config.toml` |
| **Signal** | Two-way | signal-cli REST events | Your own linked signal-cli account | `config.toml` |
| **Mattermost** | Two-way | WebSocket | Your own bot token | `config.toml` |
| **QQ** | Two-way | WebSocket | Your own bot credentials | `config.toml` |
| **Linq** | Two-way (SMS) | Webhook | Your own API token | `config.toml` |
| **Email** | Two-way | IMAP IDLE + SMTP | Your own mailbox credentials | `config.toml` |
WhatsApp also has an experimental peer-to-peer variant behind the `whatsapp-web` feature flag. Channels marked "webhook" keep a live connection alive but receive inbound messages by HTTP push, so they need a reachable HTTPS endpoint configured on the provider's side.
Telegram is the most fully featured channel — it supports typing indicators and live draft updates, and is currently the only channel wired to a per-channel approval surface, so `Prompt`-class tool calls can be answered inline rather than parked. Discord adds native threaded replies; Lark also threads. Web supports rich text and stays entirely local.
***
## Credential modes
Channels authenticate one of a few ways:
* **Connect via OpenHuman (managed)** — a one-click, encrypted connection brokered through the OpenHuman backend. Today this covers Telegram (message the managed bot directly) and Discord (link your account or install via OAuth). No tokens live on your machine.
* **Your own credentials** — you supply a bot token, API key/secret, or app credentials. Telegram (BotFather token), Discord (bot token), Slack, WhatsApp, Lark/Feishu, DingTalk, Yuanbao, Matrix, Signal, Mattermost, QQ, Linq, IRC, and Email all support this. Maximum control; you own the platform account, rate limits, and any webhook endpoint.
* **Local, no credentials** — the **Web** chat and **iMessage** need no tokens at all. Web runs inside the desktop app; iMessage drives the local macOS Messages app over an AppleScript bridge (grant Full Disk Access). Both keep messages on your machine.
Secrets supplied for any mode are stored through OpenHuman's credential layer and protected at rest by the [encryption layer](privacy-and-security.md) — never written to `config.toml` in plaintext for the UI-managed channels.
***
## Choosing the default channel
Open **Settings → Automation & Channels → Messaging Channels** to pick which channel is the **active route** — the one OpenHuman uses for proactive, recipient-less delivery (cron, triggers, subconscious). The default is the in-app **Web** chat until you change it. Setting a new default takes effect immediately, without restarting the channel runtime, and the panel shows which channel is currently active. Inbound messages always get answered on whatever channel they arrived on, regardless of the default route.
***
## See also
* [Integrations](integrations/README.md) — the read-side catalog the agent pulls context from.
* [Triggers](integrations/triggers.md) — live events that fire proactive channel delivery.
* [Subconscious Loop](subconscious.md) — the background loop that can reach you through the active channel.
* [Privacy & Security](privacy-and-security.md) — where credentials live and the backend boundary.
* [OS Keyring & Secret Storage](os-keyring-and-secret-storage.md) — at-rest protection for channel secrets.
-515
View File
@@ -1,515 +0,0 @@
---
description: 在云端托管 headless openhuman-core——DigitalOcean App Platform、Fly.io 或任何 VPS 上的 Docker Compose。
icon: cloud
lang: zh-CN
---
# 云端部署
OpenHuman 是一个桌面应用,但它的 **Rust 核心**`openhuman-core`)是一个可以托管在云端的 headless JSON-RPC 服务器。单独部署核心的用途包括:
- 多设备访问,让多个桌面客户端指向同一个托管核心
- 没有本地 Rust 工具链的内部测试人员
- 应该比笔记本 session 更长寿的长运行 cron job / webhook
本指南涵盖四条部署路径,由易到难:
1. [DigitalOcean App Platform:一键部署](#1-digitalocean-app-platform-一键部署)
2. [DigitalOcean App Platform:通过 doctl 手动部署](#2-digitalocean-app-platform-通过-doctl-手动部署)
3. [任何 VPS 通过 Docker Compose](#3-任何-vps-通过-docker-compose)
4. [Fly.io](#4-flyio)
每条路径部署的内容相同:一个运行 `openhuman-core serve` 在端口 `7788` 上的单一容器。公共主机应位于提供商的 TLS 之后,例如 `https://core.example.com/rpc`。仅限私有的主机——localhost、RFC1918 网络或 Tailscale 等 tailnet——可以使用纯 HTTP,例如 `http://100.x.x.x:7788/rpc`,当核心无法从公共互联网访问时。桌面应用已经知道如何与远程核心通信;在 `app/.env.local` 中设置 `OPENHUMAN_CORE_RPC_URL``OPENHUMAN_CORE_TOKEN=...`,然后启动即可。
---
## Bearer token 的单一事实来源
每次 `/rpc` 调用都携带 `Authorization: Bearer <token>`。核心在启动时通过两种方式加载该 token([`src/core/auth.rs`](../../src/core/auth.rs)):
1. **`OPENHUMAN_CORE_TOKEN` 环境变量**——由调用方预置(Tauri 壳层、Docker、App Platform、systemd unit 等)。核心原样使用此值,**绝不**写入文件。
2. **`{workspace}/core.token` 文件**——仅在 `OPENHUMAN_CORE_TOKEN` 未设置时由核心在首次启动时生成。独立运行的 `openhuman core run` 使用此方式,以便 CLI 客户端可以 `cat` 该文件。
**任何远程 / Docker 化部署的经验法则:始终设置 `OPENHUMAN_CORE_TOKEN`。** 不要在容器中依赖 `core.token`——临时文件系统会在重新部署时丢失它,任何试图从容器外部读取该文件的客户端都会得到过期或空值。这两条路径在启动时故意互斥;混合使用是"重新部署后 dashboard 报 401"的最常见原因。
要检查*运行中*的核心在使用什么,在主机上运行 [`scripts/print-core-token.sh`](../../scripts/print-core-token.sh)(或在容器内使用 `docker compose exec`):
```bash
scripts/print-core-token.sh --where # 打印 'env' 或 'file:/path'
scripts/print-core-token.sh --redact # 前 8 个十六进制字符 + '…'(适合日志)
scripts/print-core-token.sh # 完整值(直接管道到客户端)
```
桌面应用的首次运行选择器也在 Core RPC URL + token 字段旁暴露了一个**测试连接**按钮,它向该 URL 和输入的 token 触发 `core.ping`,并在持久化配置之前内联报告 `Connected ✓` / `Auth failed` / `Unreachable`
---
## 开始前的准备工作
| 设置 | 必需 | 说明 |
| ---- | ---- | ---- |
| `OPENHUMAN_CORE_TOKEN` | 是 | 客户端发送给 `/rpc` 的 Bearer token。用 `openssl rand -hex 32` 生成。**任何持有此 token 的人都可以驱动核心。** |
| `BACKEND_URL` | 是 | 核心通信的 Tinyhumans 后端(生产环境为 `https://api.tinyhumans.ai`)。 |
| `OPENHUMAN_APP_ENV` | 否 | `production``staging`。默认 `production`。 |
| `OPENHUMAN_CORE_HOST` | 否 | 容器中默认 `0.0.0.0`。 |
| `OPENHUMAN_CORE_PORT` | 否 | 默认 `7788`。 |
| `RUST_LOG` | 否 | `info` 足够;`debug` 用于排查。 |
运行中的容器暴露的端点:
- `GET /health`,公共存活探针。每条部署路径的健康检查都使用它。
- `POST /rpc`,受 bearer 保护的 JSON-RPC 入口。
- `GET /events``GET /ws/dictation`,公共流式通道。
`OPENHUMAN_WORKSPACE` 目录(容器内为 `/home/openhuman/.openhuman`)保存核心的配置、sqlite 数据库和技能状态。**在每个生产部署中将其挂载到持久卷**,否则重启时会丢失数据。
---
## 1. DigitalOcean App Platform:一键部署
点击下面的按钮,从本仓库的 [`.do/app.yaml`](../../.do/app.yaml) 创建一个新的 App Platform 应用:
[![Deploy to DO](https://www.deploytodo.com/do-btn-blue.svg)](https://cloud.digitalocean.com/apps/new?repo=https://github.com/tinyhumansai/openhuman/tree/main)
然后,在 App Platform UI 中,**在首次部署完成之前**:
1. 打开 **Settings → App-Level Environment Variables** 标签页。
2. 将占位符 `OPENHUMAN_CORE_TOKEN` 值替换为强 secret`openssl rand -hex 32`)。标记为 encrypted。
3. 如果部署的是 staging,将 `OPENHUMAN_APP_ENV` 改为 `staging``BACKEND_URL` 改为 `https://staging-api.tinyhumans.ai`
4. 点击 **Save**。App Platform 用新 secret 重新部署。
App Platform 处理 TLS、崩溃重启、日志流式传输和 `git push` 时的滚动重新部署(在 `.do/app.yaml` 中设置 `deploy_on_push: true` 以选择加入)。
> **持久化说明:** App Platform Basic 不提供块存储。核心的工作区位于容器的临时文件系统中,重新部署时会丢失。如需持久存储,请附加托管数据库或升级到支持卷的套餐。参见 [Compose 路径](#3-任何-vps-通过-docker-compose)获取开箱即用持久卷的自助替代方案。
---
## 2. DigitalOcean App Platform:通过 doctl 手动部署
如果你不想点击 UI
```bash
# 一次性:安装 doctl 并认证。
doctl auth init
# 编辑 .do/app.yaml - 将 OPENHUMAN_CORE_TOKEN 设置为真实值(或通过 --spec 配合 envsubst 在创建时传入)。然后:
doctl apps create --spec .do/app.yaml
# 观察构建:
doctl apps list
doctl apps logs <app-id> --type build --follow
```
编辑 spec 后更新现有应用:
```bash
doctl apps update <app-id> --spec .do/app.yaml
```
---
## 3. 任何 VPS 通过 Docker Compose
适用于任何安装了 Docker Engine ≥ 24 和 Compose 插件的主机。DigitalOcean Droplet、Hetzner、Linode、EC2、家用服务器。
每个生产版本都会向 GHCR 发布多标签镜像:
```bash
docker pull ghcr.io/tinyhumansai/openhuman-core:latest # 追踪最新生产切版
docker pull ghcr.io/tinyhumansai/openhuman-core:v1.2.4 # 按 GitHub Release tag 固定
docker pull ghcr.io/tinyhumansai/openhuman-core:1.2.4 # 按 SemVer 固定
```
镜像是 `linux/amd64`。arm64 主机拉取独立 tarball,该 tarball 附在同一 GitHub Release 中(`openhuman-core-<version>-aarch64-unknown-linux-gnu.tar.gz`),或在 arm64 构建器上从源码构建镜像。
使用已发布镜像快速运行:
```bash
docker run -d --name openhuman-core -p 7788:7788 \
-e OPENHUMAN_CORE_TOKEN="$(openssl rand -hex 32)" \
-e BACKEND_URL=https://api.tinyhumans.ai \
-e OPENHUMAN_APP_ENV=production \
-v openhuman-workspace:/home/openhuman/.openhuman \
ghcr.io/tinyhumansai/openhuman-core:latest
```
或使用仓库内的 Compose 文件(仍然从 `Dockerfile` 本地构建镜像;在 `docker-compose.yml` 中将 `image:` 字段切换为 `ghcr.io/tinyhumansai/openhuman-core:latest` 以改用已发布镜像):
```bash
# 在服务器上:
git clone https://github.com/tinyhumansai/openhuman.git
cd openhuman
# 配置 secrets
cp .env.example .env
# 编辑 .env - 至少填写:
# BACKEND_URL=https://api.tinyhumans.ai
# OPENHUMAN_CORE_TOKEN=<openssl rand -hex 32>
# OPENHUMAN_APP_ENV=production
# 构建并启动:
docker compose up -d
# 验证:
docker compose ps
curl -fsS http://localhost:7788/health
```
### 无 Docker 的 Headless 安装
如果主机无法运行 Docker,抓取附在最新 [GitHub Release](https://github.com/tinyhumansai/openhuman/releases/latest) 中的独立 CLI tarball
```bash
# 选择匹配你主机架构的 tarball。
ARCH="$(uname -m)"
case "$ARCH" in
x86_64) TARGET=x86_64-unknown-linux-gnu ;;
aarch64) TARGET=aarch64-unknown-linux-gnu ;;
*) echo "Unsupported arch: $ARCH"; exit 1 ;;
esac
VERSION=1.2.4 # 设置为你想要的版本
curl -fsSL "https://github.com/tinyhumansai/openhuman/releases/download/v${VERSION}/openhuman-core-${VERSION}-${TARGET}.tar.gz" \
| tar -xz -C /usr/local/bin
openhuman-core --version
```
然后在你选择的 service manager 下运行 `openhuman-core serve`systemd、supervisord 等),使用上述相同的环境变量。
### Headless 自更新契约
Headless 部署应将 `openhuman.update_apply` 视为安全原语:它下载 release asset,将其原子地写入当前二进制文件旁边,然后返回。不会自动退出。
`openhuman.update_run` 遵循 `config.update.restart_strategy`
- `self_replace`(默认):stage 二进制文件,发布一个进程内重启请求,让运行中的核心自行 respawn。
- `supervisor`stage 二进制文件并返回 `restart_requested=false`。你的外部 service manager 必须重启进程。
对于长运行的 Linux 服务,设置:
```toml
[update]
restart_strategy = "supervisor"
rpc_mutations_enabled = false
```
或等效的环境变量:
```bash
OPENHUMAN_AUTO_UPDATE_RESTART_STRATEGY=supervisor
OPENHUMAN_AUTO_UPDATE_RPC_MUTATIONS_ENABLED=false
```
推荐的 `systemd` 配置:
```ini
Restart=always
ExecReload=/bin/kill -HUP $MAINPID
```
运维流程:
1. 调用 `openhuman.update_check` 发现 release。
2. 在你的 `update.toml` 中配置 `restart_strategy = "supervisor"`(或设置 `OPENHUMAN_AUTO_UPDATE_RESTART_STRATEGY=supervisor`),以便核心 stage 新二进制文件而不尝试自行 re-exec,然后调用 `openhuman.update_apply``openhuman.update_run``restart_strategy` 是配置设置,不是 RPC 参数。
3. 显式重启 unit`systemctl restart openhuman`
如果下载或 staging 失败,运行中的二进制文件会保持原位,不会请求重启。如果 staged 二进制文件在重启后被证明有问题,通过你的包管理器、镜像标签或 release artifact 恢复之前的二进制文件,然后再次重启 supervisor。
Compose 文件([`docker-compose.yml`](../../docker-compose.yml))将核心映射到 `:7788`,挂载命名卷 `openhuman-workspace` 用于持久化,并设置 `restart: unless-stopped` 以便主机重启后核心自动恢复。
### 更新
```bash
git pull
docker compose build
docker compose up -d
```
对于暴露 RPC 的生产部署,建议禁用可变的 update RPC(`OPENHUMAN_AUTO_UPDATE_RPC_MUTATIONS_ENABLED=false`),并通过你现有的镜像标签或包管理流程执行 rollout。
### 日志
```bash
docker compose logs -f openhuman-core
```
### 轮换 bearer token
`OPENHUMAN_CORE_TOKEN` 是公共互联网与完整 RPC 访问之间的唯一屏障。按时间表轮换它,并在任何疑似泄露后轮换:
```bash
# 1. 生成新 token 并更新服务器端 .env。
openssl rand -hex 32 > /tmp/new-token
sed -i.bak "s|^OPENHUMAN_CORE_TOKEN=.*|OPENHUMAN_CORE_TOKEN=$(cat /tmp/new-token)|" .env
rm /tmp/new-token .env.bak
# 2. 重启容器,让新值到达核心进程。
docker compose up -d --force-recreate openhuman-core
# 3. 确认运行中的容器正在使用新 token(脱敏)。
docker compose exec openhuman-core /bin/sh -c \
'echo -n "$OPENHUMAN_CORE_TOKEN" | head -c 8; echo "…"'
# 4. 更新每个桌面客户端(切换模式 → 在选择器中重新粘贴,或编辑 app/.env.local 中的 OPENHUMAN_CORE_TOKEN 然后重新启动)。仍然持有旧 token 的客户端会在下一次 /rpc 调用时收到 HTTP 401——这是预期行为,不是回归。
```
对于 App Platform,在 **Settings → App-Level Environment Variables** 中执行相同操作:编辑 `OPENHUMAN_CORE_TOKEN` secret,让 App Platform 重新部署。没有单独的 token 文件需要删除;环境变量是唯一的状态。
### 置于 TLS 之后
使用 Caddy、nginx 或 Traefik 作为 `:7788` 的反向代理。最小 `Caddyfile`
```caddy
core.example.com {
reverse_proxy localhost:7788
}
```
---
## 将桌面应用指向托管核心
在桌面应用的环境文件(`app/.env.local`)中:
```bash
# 使用托管核心而不是生成本地 sidecar。
OPENHUMAN_CORE_RUN_MODE=external
OPENHUMAN_CORE_RPC_URL=https://core.example.com/rpc
OPENHUMAN_CORE_TOKEN=<你在服务器上设置的相同 token>
```
对于没有公共 IP 的私有 tailnet-only VM,改用 tailnet URL
```bash
OPENHUMAN_CORE_RUN_MODE=external
OPENHUMAN_CORE_RPC_URL=http://100.x.x.x:7788/rpc
OPENHUMAN_CORE_TOKEN=<你在服务器上设置的相同 token>
```
重启桌面应用。`App.tsx` 中的 provider 链会将所有 RPC 调用路由到远程核心;其他一切不变。公共 `http://` 主机被应用选择器拒绝;对任何可从公共网络访问的核心使用 HTTPS。
---
## 命名卷所有权与 Docker entrypoint
Docker 默认创建 `root:root` 拥有的命名卷。因为核心以非 root `openhuman` 用户(UID 10001)运行,banner 之后的首次写入——`init_rpc_token → write_token_file``$OPENHUMAN_WORKSPACE`——如果没有先修复所有权,会抛出 `Permission denied (os error 13)`
镜像在 `/usr/local/bin/docker-entrypoint-core.sh` 附带一个专用 entrypoint,它:
1.`root` 启动。
2.`$OPENHUMAN_WORKSPACE``$HOME/.openhuman``OPENHUMAN_CORE_TOKEN` 未设置时 `core.token` 被写入的目录)运行 `mkdir -p` + `chown openhuman:openhuman`
3. 调用 `exec gosu openhuman openhuman-core "$@"` 来降权并移交控制权给二进制文件。
这是**幂等的**:在新创建的卷上,chown 会修复 root 拥有的目录;在已经修复过的卷上,chown 是 no-op。从早于该修复的镜像升级时,不需要手动执行 `docker volume rm`
该 entrypoint 名为 `docker-entrypoint-core.sh`,仅接入根 `Dockerfile`。E2E 镜像(`e2e/docker-entrypoint.sh`)不受影响。
---
## 4. Fly.io
[Fly.io](https://fly.io) 非常适合 `openhuman-core`:它自动处理 TLS,在所有套餐上支持持久卷,并且可以自动停止空闲机器以削减成本。
### 前置条件
- 安装并认证 [flyctl](https://fly.io/docs/flyctl/install/)`fly auth login`
- Fly.io 账户
### 步骤 1 —— 启动应用
```bash
fly launch --no-deploy --config .fly/fly.toml
```
Fly.io 自动检测 `Dockerfile`。选择靠近你用户的区域,并在提示时跳过首次部署。这会生成一个配置文件。
### 步骤 2 —— 配置 `.fly/fly.toml`
仓库在 [`.fly/fly.toml`](../../.fly/fly.toml) 附带一个模板。用 `fly launch` 期间选择的值填充 `<your-app-name>``<your-region>`
```toml
app = '<your-app-name>'
primary_region = '<your-region>'
[build]
dockerfile = "Dockerfile"
[env]
OPENHUMAN_CORE_HOST = "0.0.0.0"
OPENHUMAN_CORE_PORT = "7788"
OPENHUMAN_WORKSPACE = "/home/openhuman/.openhuman"
RUST_LOG = "info"
[[mounts]]
source = "openhuman_workspace"
destination = "/home/openhuman/.openhuman"
[http_service]
internal_port = 7788
force_https = true
auto_stop_machines = 'stop'
auto_start_machines = true
# min_machines_running = 0 在空闲时完全停止机器(最便宜),但
# 空闲后的第一个请求需要支付冷启动惩罚(容器启动 +
# Rust 二进制文件初始化——数秒)。设为 1 可保持一台机器热备。
min_machines_running = 0
processes = ['app']
[[http_service.checks]]
interval = "30s"
timeout = "5s"
grace_period = "10s"
method = "GET"
path = "/health"
[[vm]]
memory = '1gb'
cpus = 1
```
### 步骤 3 —— 创建持久卷
```bash
fly volumes create openhuman_workspace --size 5 --region <your-region> --config .fly/fly.toml
```
**将工作区挂载到持久卷**,否则每次重新部署都会丢失数据。
### 步骤 4 —— 设置 secrets
```bash
# 必需
fly secrets set OPENHUMAN_CORE_TOKEN="$(openssl rand -hex 32)"
fly secrets set BACKEND_URL="https://api.tinyhumans.ai"
fly secrets set OPENHUMAN_APP_ENV="production"
# 建议——任何可公开访问的部署:
fly secrets set OPENHUMAN_AUTO_UPDATE_RPC_MUTATIONS_ENABLED="false"
fly secrets set OPENHUMAN_AUTO_UPDATE_RESTART_STRATEGY="supervisor"
# 可选——错误报告和 analytics:
fly secrets set OPENHUMAN_CORE_SENTRY_DSN="https://<key>@o<org>.ingest.sentry.io/<project>"
fly secrets set OPENHUMAN_ANALYTICS_ENABLED="true"
```
保存 `OPENHUMAN_CORE_TOKEN` 的值——你稍后需要它来连接桌面应用。**任何持有此 token 的人都可以驱动核心**;像密码一样对待它,并在任何疑似泄露后通过 `fly secrets set OPENHUMAN_CORE_TOKEN="$(openssl rand -hex 32)"` 轮换。
### 步骤 5 —— 部署
```bash
fly deploy --config .fly/fly.toml
```
验证核心健康:
```bash
curl -fsS https://<your-app-name>.fly.dev/health
```
### 步骤 6 —— 将桌面应用指向托管核心
`app/.env.local` 中:
```bash
OPENHUMAN_CORE_RUN_MODE=external
OPENHUMAN_CORE_RPC_URL=https://<your-app-name>.fly.dev/rpc
OPENHUMAN_CORE_TOKEN=<你在步骤 4 中设置的 token>
```
或使用桌面应用中的**首次运行选择器**Core RPC URL + token 字段,带**测试连接**按钮)来配置,无需编辑文件。
### 持续部署
要在每次推送到 `main` 时自动重新部署,在 `.github/workflows/fly-deploy.yml` 添加 workflow 文件:
```yaml
name: Fly Deploy
on:
push:
branches:
- main
paths:
- 'src/**'
- 'Cargo.toml'
- 'Cargo.lock'
- 'Dockerfile'
- '.fly/fly.toml'
- 'scripts/docker-entrypoint-core.sh'
jobs:
deploy:
name: Deploy openhuman-core
runs-on: ubuntu-latest
concurrency: deploy-group
steps:
- uses: actions/checkout@v4
# 将 Fly action 固定到带标签的 release(或完整 commit SHA),而不是 @master——
# 追踪移动分支意味着信任未来推送到那里的每个 commit,包括被入侵的维护者账户所做的任何 commit。
- uses: superfly/flyctl-actions/setup-flyctl@1.5
- run: flyctl deploy --remote-only --config .fly/fly.toml
env:
FLY_API_TOKEN: ${{ secrets.FLY_API_TOKEN }}
```
`fly tokens create deploy` 生成部署 token,并将其作为名为 `FLY_API_TOKEN` 的仓库 secret 添加。
### 更新
```bash
fly deploy --config .fly/fly.toml
```
对于固定版本的部署,在 `.fly/fly.toml` 中更新镜像标签并重新部署:
```toml
[build]
image = "ghcr.io/tinyhumansai/openhuman-core:v1.2.4"
```
### 日志
```bash
fly logs --config .fly/fly.toml
```
### 已知陷阱 —— 卷上的 UID 不匹配
如果你在从 `Dockerfile` 构建(创建 UID 10001 的 `openhuman` 用户)和拉取预构建的 GHCR 镜像(使用 UID 1000)之间切换,已写入持久卷的文件会被旧 UID 拥有,并在启动时产生 `Permission denied (os error 13)`
通过 SSH 进入并重新拥有工作区来修复:
```bash
fly ssh console --config .fly/fly.toml
chown -R openhuman:openhuman /home/openhuman/.openhuman/
exit
fly machine restart --config .fly/fly.toml
```
---
## 冒烟测试
云部署路径有两种需要防范的失败模式:
- **`docker-image`**——设置 `OPENHUMAN_CORE_TOKEN` 且不挂载卷。保护 DigitalOcean App Platform 路径(`.do/app.yaml`),其中 token 始终预置且不挂载持久卷。
- **`docker-volume-permissions`**——省略 `OPENHUMAN_CORE_TOKEN` 并在 `/home/openhuman/.openhuman` 挂载一个新的匿名卷。复现 issue #2065 的确切失败模式,并断言 `/health` 返回 200 且日志中不存在 `Permission denied (os error 13)`
要在本地运行相同的检查:
```bash
docker build -t openhuman-core:smoke .
# Token-set 路径(App Platform):
docker run -d --name oh-smoke -p 7788:7788 \
-e OPENHUMAN_CORE_TOKEN=smoke-test-token \
openhuman-core:smoke
curl -fsS http://localhost:7788/health
docker rm -f oh-smoke
# Fresh-volume / no-token 路径(Docker Compose、VPS):
docker volume create oh-vol-test
docker run -d --name oh-vol-smoke -p 7789:7788 \
-v oh-vol-test:/home/openhuman/.openhuman \
openhuman-core:smoke
curl -fsS http://localhost:7789/health
docker rm -f oh-vol-smoke
docker volume rm oh-vol-test
```
+79
View File
@@ -0,0 +1,79 @@
---
description: >-
Long-term goals, per-thread objectives with budgets, and a kanban task board -
how OpenHuman stays pointed at what matters.
icon: target
---
# Goals & Todos
OpenHuman keeps the agent aligned with what you actually care about through three complementary layers: durable **long-term goals**, a single **thread goal** per conversation, and a collaborative **task board** of todos. Each one is editable by both you and the agent, and all of them survive restarts.
***
## Long-term goals
A short, human-readable list of your durable objectives — things like _"Ship the desktop app"_ or _"Grow the community to 10k."_ It lives as a plain Markdown file (`MEMORY_GOALS.md`) in your workspace, so you can open and edit it directly.
The list is deliberately tiny — capped at roughly **8 items / 500 tokens** — so it's cheap for the agent to read on every relevant turn and easy for you to keep honest. Each goal gets a stable short id (`g1`, `g2`, …) so it can be edited or deleted without depending on order.
* **Goals Panel** (Intelligence → Goals) shows the list with add / edit / delete actions.
* **Reflect** runs a background `goals_agent` that reviews your goals against recent memory and conversation, then makes minimal, justified changes — adding what you've clearly started pursuing, retiring what you've dropped. On first run it bootstraps an initial set from your context.
* The agent reads and updates the same list mid-conversation via its `goals_list` / `goals_add` / edit tools, so your edits and the agent's stay in lock-step.
RPC surface: `openhuman.memory_goals_list` / `_add` / `_edit` / `_delete` / `_reflect`.
***
## Thread goals
Each conversation can carry **one** thread goal — a durable "completion contract" the agent works across turns, interrupts, resumes, and budget boundaries. A thread goal has an objective, a status, and an optional **token budget** so you can cap how much work a thread is allowed to consume.
| Status | Meaning |
| ----------------- | ------------------------------------------------------------------- |
| `active` | The agent may keep working the objective. |
| `paused` | Suspended (e.g. you interrupted); reactivates when the thread resumes. |
| `budget_limited` | Tokens spent ≥ budget; substantive work halts until you raise it. |
| `complete` | Objective satisfied. |
The orchestrator sets a goal with `goal_set`, reads it with `goal_get`, and finishes it with `goal_complete`. Updates emit `thread/goal/updated` events so the UI stays live.
**Autonomous idle continuation.** If a thread has an active goal and goes idle — no in-flight turn, no activity for a configured interval (e.g. 10 minutes) — the [heartbeat](subconscious.md) can inject a single continuation turn that resumes the transcript and keeps working the objective. It's opt-in (`heartbeat.goal_continuation_enabled`) and guarded by a one-shot suppression flag per idle period, so the agent never self-drives into a loop.
***
## Task board (todos)
Every conversation also hosts a **kanban-style task board** — a list of discrete work cards that you and the agent build together. Unlike a thread goal (one durable objective), the board is a collection of concrete items, each with rich structure:
* Title / description, and a **status**: `todo`, `in_progress`, `awaiting_approval`, `ready`, `blocked`, `done`, `rejected`.
* Optional objective and desired outcome.
* An ordered **execution plan**, **acceptance criteria** checklist, assigned agent, and an **approval mode** (required / not required).
* Notes, blocker reason, and evidence / links.
The agent reads the board with `todo_list`, appends with `todo_add`, edits with `todo_edit`, and advances status with `todo_update_status`. Destructive operations (clear / remove / replace) are disabled by default. You and the agent share the same persistence, so edits stay consistent.
Two reserved boards back special views:
* **`user-tasks`** — your personal task list, not attached to any conversation. Create and manage these from the **User Task Composer** (Intelligence → Tasks). You can optionally attach a task to a conversation and assign it to the orchestrator with `approvalMode: not_required`, so the background dispatcher auto-picks and runs it.
* **`task-sources`** — an inbox for tasks ingested from external sources before they're promoted to an agent workstream.
RPC surface: `openhuman.todos_list` / `_add` / `_edit` / `_update_status` / `_set_session_thread`. Responses include a rendered `markdown` field so the board renders identically in the UI and in agent transcripts.
***
## How the three relate
| Layer | Scope | Count | Who drives it |
| ------------------ | -------------------- | --------------- | ----------------------------------- |
| Long-term goals | Your whole account | ~8 max | You + periodic `goals_agent` reflect |
| Thread goal | One conversation | 1 per thread | Orchestrator, with optional budget |
| Task board (todos) | One conversation | Many cards | You + agent, collaboratively |
***
## See also
* [Subconscious Loop](subconscious.md) — the background loop that powers idle continuation and task evaluation.
* [Memory Tree](obsidian-wiki/memory-tree.md) — what goal reflection reads from.
* [SuperContext](super-context.md) — first-turn grounding that complements goal-driven work.
+7 -2
View File
@@ -59,9 +59,14 @@ Three integrations are special. OpenHuman uses them to _talk back_ to you, not j
Set your default under **Settings → Automation & Channels → Messaging Channels**. The active route status shows which channel is currently in use. Telegram offers two credential modes: connect via OpenHuman (one-click, encrypted) or provide your own credentials for maximum control.
## Skills
## Beyond the curated catalog: MCP & Skills
Beyond third-party services, OpenHuman has **skills**, small sandboxed modules that run inside the app, fetch external data, run on a schedule, transform information, and respond to events. Each runs with enforced resource limits. Skills install from the Skills tab and integrate with the same Memory Tree as everything else.
The 118+ OAuth connectors are the curated path. Beyond them, OpenHuman opens up the wider open-tooling ecosystem:
* **MCP servers** — a built-in registry browses thousands of [Model Context Protocol](https://modelcontextprotocol.io) servers (Smithery + the official registry) that install locally as new agent tools.
* **Skills** — a browsable, ~90,000-entry catalog of `SKILL.md` capability bundles aggregated from HermesHub, ClawHub, LobeHub and more. (Note: the old in-app skills runtime has been removed; Skills are now a metadata catalog you install from the Skills tab.)
See [MCP Servers & Skills](mcp-and-skills.md) for the full picture.
## Native voice and tools
@@ -1,85 +0,0 @@
---
description: >-
118+ 第三方集成——Gmail、Notion、GitHub、Slack、Stripe、日历等,
一键 OAuth 连接,无需 API 密钥。
icon: plug
---
# 第三方集成(118+
OpenHuman 搭载对 **118+ 第三方服务**的后端代理访问。任意服务通过托管路径连接都只需在应用内一键 OAuth,无需手动接入 API 密钥,也无需穿梭于插件市场。
底层连接器层由 [Composio](https://composio.dev) 驱动。默认托管模式下,OpenHuman 后端拥有 Composio API 密钥、OAuth token 经纪、速率限制和触发器 webhook 分发。如果你切换到直连模式,core 用你自己的 Composio API 密钥与 Composio 通信;同步工具调用可以工作,但实时触发器 webhook 必须配置在你自己的 webhook 基础设施上。
服务连接后,会同时出现在四个位置:
1. 作为**智能体工具**,模型可以直接调用。
2. 作为**记忆源**[自动拉取](../obsidian-wiki/auto-fetch.zh-CN.md)每二十分钟将其同步到[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)。
3. 作为**个人化信号**,你在各服务上的活动为你的偏好模型提供数据。
4. 作为**触发器源**,实时事件(新邮件、新 charge、入站 DM)流入[触发器](triggers.zh-CN.md)流水线,可以自动触发智能体操作。
## 目录中的部分服务
目录涵盖生产力、商业、社交、消息和 Google 类目。不完全示例:
| 类别 | 示例 |
| ----------------------- | ---------------------------------------------------- |
| **邮件与日历** | Gmail、Outlook、Google Calendar、Apple Calendar |
| **文档与存储** | Google Docs、Google Drive、Notion、Dropbox、Airtable |
| **代码与开发** | GitHub、Linear、Jira、Figma |
| **通讯** | Slack、Discord、Microsoft Teams、Telegram、WhatsApp |
| **CRM 与销售** | Salesforce、HubSpot |
| **商业与支付** | Stripe、Shopify |
| **项目管理** | Asana、Trello |
| **社交** | Twitter / X、Spotify、YouTube |
## 原生 vs 代理
部分服务有**原生 provider**。Rust 模块知道如何直接将服务摄入记忆树(例如 Gmail 的原生摄入路径)。其他仅暴露为**代理工具**:智能体可以调用,但没有自动摄入。新的原生 provider 随着功能落地陆续添加。
## 连接如何工作
点击任意集成的**连接**。浏览器窗口打开进行 OAuth。登录后,连接变为活跃状态,OpenHuman 在下一个 20 分钟 tick 开始同步。
每个集成显示其当前状态:
* **未连接**。集成尚未设置。
* **已连接**。集成活跃并正在同步。
* **管理**。活跃集成,可重新配置或断开。
你可以随时从 Skills 标签页撤销任何连接。
## 消息渠道
三个集成是特殊的。OpenHuman 用它们*回复*你,而不只是读取:
* **Telegram**。主要消息渠道。双向:发送和接收消息、管理聊天、搜索历史、创建群组、代表你执行 80+ 操作。所有操作通过你自己加密的凭据运行。
* **Discord**。通过 Discord 发送和接收消息。连接你的账户以接收 OpenHuman 消息。
* **Web**。桌面应用内的浏览器聊天界面。消息完全保留在本地。
在**设置 → 自动化与渠道 → 消息渠道**中设置你的默认值。活跃路由状态显示当前使用的渠道。Telegram 提供两种凭据模式:通过 OpenHuman 连接(一键,加密)或提供你自己的凭据以获得最大控制权。
## 技能
除了第三方服务,OpenHuman 还有**技能**——运行在应用内的小型沙盒模块,获取外部数据、按计划运行、转换信息、响应事件。每个技能都强制执行资源限制。技能从 Skills 标签页安装,与其他所有内容一样集成到同一个记忆树。
## 原生语音和工具
有两个功能作为原生功能搭载,而非集成,因为它们对桌面体验是基础性的:
* [**语音**](../native-tools/voice.zh-CN.md)。语音转文字输入、文字转语音输出,加上实时 Google Meet 智能体——加入会议、转录到记忆树、在通话中说话。
* [**原生工具**](../native-tools/README.zh-CN.md)。内置网络搜索、网络抓取,以及完整的文件系统/git/lint/test/grep 编码工具集,智能体开箱即用。
## 隐私边界
OpenHuman core 从不直接调用任何第三方 API。所有请求都通过 OpenHuman 后端,该后端处理 OAuth token 和速率限制。你的 token 永不以明文形式存储在电脑磁盘上,智能体只看到工具调用的*结果*,而不是凭据。
如果你选择直连 Composio 模式,该边界会改变:你本地的 core 使用你自己的 Composio API 密钥,你负责 Composio 账户、速率限制、计费关系,以及触发器投递所需的任何 webhook 端点。
完整边界见[隐私与安全](../privacy-and-security.zh-CN.md)。
## 另见
* [触发器](triggers.zh-CN.md),已连接集成的实时事件以及它们如何触发智能体操作。
* [从集成自动拉取](../obsidian-wiki/auto-fetch.zh-CN.md)
* [记忆树](../obsidian-wiki/memory-tree.zh-CN.md)
@@ -0,0 +1,46 @@
---
description: >-
Beyond the curated OAuth connectors - browse thousands of MCP servers and a
90,000-entry Skills catalog, and let OpenHuman act as an MCP server itself.
icon: blocks
---
# MCP Servers & Skills
The [one-click OAuth integrations](README.md) are the curated path. Beyond them, OpenHuman opens up the wider open-tooling ecosystem in two ways: the **Model Context Protocol (MCP)** registry and the **Skills** catalog. OpenHuman can also expose _itself_ as an MCP server to other clients.
***
## MCP servers (thousands)
OpenHuman has a built-in **MCP registry** that browses the open MCP ecosystem and lets you install servers locally as new typed tools for the agent.
* **Two upstream registries, merged.** Discovery fans out in parallel to [Smithery.ai](https://smithery.ai) and the official [`registry.modelcontextprotocol.io`](https://registry.modelcontextprotocol.io), then merges the results — thousands of servers across both.
* **Search → install → connect.** Search the catalog, view a server's details, and install it. A local install spawns the server as a stdio subprocess; deployed servers connect over HTTP.
* **Supervised connections.** Installed servers are persisted in a local SQLite store (`mcp_clients/mcp_clients.db`) with their command, args, and transport. A supervisor loop keeps enabled servers connected, probing every ~60s with per-server exponential backoff.
Once connected, an MCP server's tools are available to the agent exactly like native tools.
> The catalogs are open-ended and grow on their own — the "5,000+" figure reflects the combined Smithery + official ecosystem size, not a fixed list baked into OpenHuman.
### OpenHuman as an MCP server
OpenHuman can run the other way around, too. `openhuman-core mcp` exposes OpenHuman over stdio as an MCP server, offering read-only tools — memory search / recall, Memory Tree browsing, and optional web search — to clients like Claude Desktop. See [MCP Server](../../developing/mcp-server.md) for setup.
***
## Skills (90,000-entry catalog)
**Skills** are a large, browsable catalog of agent skills — `SKILL.md`-style capability bundles — aggregated from multiple upstream sources (HermesHub, ClawHub, LobeHub, and more).
* **One aggregated catalog.** Sourced from HermesHub (configurable via `OPENHUMAN_SKILL_REGISTRY_CATALOG_URL`), the catalog runs to roughly **90,000 entries**. Each entry carries id, name, description, source, author, version, tags, platforms, a download URL, and license.
* **Cached and fast.** The catalog is fetched on boot in the background (without blocking startup), cached locally at `~/.openhuman/skill-registry/cache.json` with a ~1-hour TTL and served stale-while-revalidate. A single-flight gate prevents duplicate downloads of the large catalog.
* **Metadata-first.** OpenHuman's in-app skills runtime (the old QuickJS sandbox) has been **removed** — Skills are now a metadata catalog you browse and install from the Skills tab, not code executing inside the app. Availability varies per entry: some expose a direct `SKILL.md` download, others point to external hosting.
***
## See also
* [Third-party Integrations](README.md) — the curated 100+ OAuth connectors.
* [MCP Server](../../developing/mcp-server.md) — running OpenHuman as an MCP server.
* [Available Tools](../native-tools/README.md) — the native tools that ship by default.
@@ -1,138 +0,0 @@
---
description: >-
已连接集成(Gmail 新邮件、Notion 编辑、Stripe charge)的实时事件
作为触发器到达,被分类器分类,并可自动触发智能体操作。
icon: bolt
---
# 触发器
已连接的集成不仅仅是智能体可以按需读取的地方。它也是**实时事件源**。当有人给你发邮件、编辑 Notion 页面、在你的某个仓库打开 GitHub Issue、在 Stripe 上给你的卡收费、或在 Slack 上给你发 DM 时,OpenHuman 几乎实时接收该事件,并可以决定是否要对其采取行动。
本页关于这条流水线:触发器如何到达、如何分类、以及触发器如何无需你输入一个字就变成完整的智能体操作。
## 什么是触发器
触发器是你所连接集成发布的外部事件。常见形态:
| 集成 | 示例触发器 |
| --- | --- |
| **Gmail** | `GMAIL_NEW_GMAIL_MESSAGE`,收件箱中的新邮件 |
| **Slack** | `SLACK_NEW_MESSAGE`,你被提及的频道/DM 消息 |
| **Notion** | `NOTION_PAGE_UPDATED`,被跟踪的页面有变化 |
| **GitHub** | `GITHUB_ISSUE_OPENED``GITHUB_PULL_REQUEST_OPENED`,你的仓库上 |
| **Stripe** | `STRIPE_CHARGE_SUCCEEDED`,你账户上的一笔成功 charge |
| **日历** | `GOOGLE_CALENDAR_EVENT_CREATED`,你日历上的新事件 |
完整集合来自为[第三方集成](README.zh-CN.md)提供支持的 [Composio](https://composio.dev) 连接器层。当连接活跃时,相关的触发器订阅会自动接入。
### Gmail OAuth 作用域
Gmail 触发器订阅需要所连接 Google 账户的邮件读取权限。新鲜的 OpenHuman Gmail 授权请求 `https://www.googleapis.com/auth/gmail.readonly`,这样 `GMAIL_NEW_GMAIL_MESSAGE` 可以启用,原生 Gmail 同步路径可以读取新邮件元数据。
如果旧 Gmail 连接在此作用域被请求之前创建,请从设置中重新连接 Gmail 然后再启用 Gmail 触发器。
## 触发器从哪里来,从头到尾
```text
┌────────────────────┐
│ third-party API │ Gmail / Slack / Notion / GitHub / ...
└─────────┬──────────┘
│ webhook
┌────────────────────┐
│ OpenHuman backend │ HMAC 验证 webhook,规范 payload
└─────────┬──────────┘
│ Socket.IO 事件("composio:trigger"
┌────────────────────┐
│ Rust core │ 在进程内事件总线上发布 DomainEvent::ComposioTriggerReceived
│(你的笔记本)│
└─────────┬──────────┘
┌────────────────────┐
│ Trigger Triage │ 分类:drop / acknowledge / react / escalate
└─────────┬──────────┘
┌────────────────────┐
│ 以下之一: │
│ - nothing │ ← drop
│ - memory note │ ← acknowledge
│ - Trigger Reactor │ ← react1-2 个工具调用)
│ - Orchestrator │ ← escalate(完整多步规划)
└────────────────────┘
```
Webhook 永远不会被原始地到达你的机器。后端持有 OAuth token 并直接从第三方接收 webhook。它进行 HMAC 验证、规范 payload,并通过已认证的 socket 将其转发给你的 Rust core。你的笔记本在总线上看到一个干净的、经过验证的 `ComposioTriggerReceived` 事件,没有别的。
## 分类步骤
在任何操作运行之前,每个触发器都经过 [`trigger_triage`](https://github.com/tinyhumansai/openhuman/tree/main/src/openhuman/agent/agents/trigger_triage) 智能体。它的唯一工作是决定系统其余部分应该做什么。
它精确选择四种操作之一:
| 操作 | 发生什么 | 何时使用 |
| --- | --- | --- |
| **`drop`** | 什么也不做。触发器被静默记录并丢弃。 | 垃圾邮件、重复、不相关的噪音。默认用于你不在乎的东西。 |
| **`acknowledge`** | 持久化一条短期记忆笔记,不运行智能体。 | 值得记住的被动通知("档案中创建了一个新页面")。 |
| **`react`** | 使用一到两个工具调用运行 [`trigger_reactor`](https://github.com/tinyhumansai/openhuman/tree/main/src/openhuman/agent/agents/trigger_reactor) 智能体。 | 一个小的、单步的副作用:存储一条记忆条目、发布快速确认、将线程标记为已读。 |
| **`escalate`** | 全权交给带规划能力的 **orchestrator** 智能体。 | 需要推理、多步、或多技能的任何东西:起草回复、更新多个 Notion 页面、决定如何分类入站 issue。 |
分类智能体拥有与智能体其余部分相同的记忆和工作区上下文。它可以判断触发器是否与你现在正在做的事情相关、涉及哪些人、以及是否是你之前要求 OpenHuman 采取行动的那类事情。
## 触发器何时变成智能体操作
这就是区分"OpenHuman 有 Gmail 集成"和"OpenHuman 在值班你的收件箱"的部分:
- **`react`** 是廉价路径。Trigger Reactor 是一个有严格预算的窄专家,只有几个工具调用。它非常适合:写一条简短的记忆笔记说"看到 Stripe 新增一笔 $84 charge,客户 X,商户 Y"、静默将同一自动提醒标记为已处理因为你本周已经分类过两次、或存储用户以后可能想查找的事件的结构化记录。
- **`escalate`** 是重型路径。当分类智能体决定触发器需要真正的工作时,它将自包含的任务描述交给 Orchestrator。orchestrator 可以访问你完整的技能表面、工具、记忆和[潜意识循环](../subconscious.zh-CN.md)输出。从那里它可能:
- 起草一封重要邮件的回复并排队等待你批准。
- 为入站 issue 拉取相关的 Notion / Linear / Drive 上下文并写一条结构化评论。
- 基于单个入站事件更新三个已连接系统("这个客户的计划在 Stripe 变了,更新 HubSpot,在 #revenue 发帖,并在他们的 Notion 文件中添加一条笔记")。
- 判断触发器意味着一个会议刚刚被预定并为该通话预加载[会议智能体](../mascot/meeting-agents.zh-CN.md)。
两种情况下操作都在你的机器上运行,针对你的本地记忆树,使用与智能体其余部分相同的模型路由和工具表面。
## 为什么要一个分类步骤
跳过分类器并将每个触发器直接管道到 orchestrator 很有诱惑力。这是一个坏主意,有两个原因:
1. **大多数触发器是噪音。** 一个已连接的 Gmail 账户每小时触发数十个触发器,其中绝大多数是用户不在乎的。在每个上运行 orchestrator 会消耗预算并产生持续的后台活动流。
2. **不同的触发器值得不同的上限。** 一个自动 Stripe 收据和个人 Slack DM 不应该花相同的 token 数来处理。分类让廉价路径保持廉价,并将 orchestrator 保留给值得它的东西。
分类在快速模型层运行(参见[自动模型路由](../model-routing/README.zh-CN.md)),所以分类本身在亚秒级完成。
## 配置和退出
- **默认开启。** 一旦集成被连接,其触发器自动进入流水线。
- **退出。** 分类路径由 `OPENHUMAN_TRIGGER_TRIAGE_DISABLED` 环境变量控制。设为 `1` / `true` / `yes` 关闭智能体分类并退回到仅被动日志记录。集成本身保持连接;只有自动操作行为被抑制。
- **每触发器设置。** 触发器设置(哪些集成和事件类型应该被评估)在**设置**下管理;底层 RPC 方法是 `update_composio_trigger_settings` / `get_composio_trigger_settings`
- **审计日志。** 每个触发器,无论决策如何,都被写入触发器历史,这样你可以看到什么到达了、分类器决定了什么、以及(如果有的话)运行了什么。决策和升级也作为进程内总线上的 `TriggerEvaluated` / `TriggerEscalated` 事件发布,这意味着核心内部的任何东西都可以订阅它们。
## 隐私边界
触发器遵循与产品其余部分相同的边界(参见[隐私与安全](../privacy-and-security.zh-CN.md)):
- 第三方 token 位于后端,永不在你的笔记本上。
- Webhook 在到达你的机器之前由后端进行 HMAC 验证。
- 触发器 payload 由你的本地 core 处理;分类和任何反应在你机器上运行,针对你的本地记忆树。
- `acknowledge` / `react` / `escalate` 路径写入的记忆笔记存储在你本地 SQLite 记忆树和 Markdown 存储库中,与任何其他来源相同。
## 开发者实现指针
- 分类智能体:`src/openhuman/agent/agents/trigger_triage/`
- Reactor 智能体:`src/openhuman/agent/agents/trigger_reactor/`
- Composio 总线订阅器:`src/openhuman/composio/bus.rs``ComposioTriggerSubscriber`
- 触发器历史持久化:`src/openhuman/composio/trigger_history.rs`
- 领域事件:`DomainEvent::ComposioTriggerReceived``DomainEvent::TriggerEscalated``src/core/event_bus/events.rs`
- 触发器设置 RPC`src/openhuman/config/` 中的 `update_composio_trigger_settings` / `get_composio_trigger_settings`
## 另见
* [第三方集成](README.zh-CN.md),触发器来源的服务目录。
* [从集成自动拉取](../obsidian-wiki/auto-fetch.zh-CN.md),轮询对应部分,定期将源数据摄入记忆树。
* [潜意识循环](../subconscious.zh-CN.md),使用触发器上下文和记忆提前规划的背景循环。
* [会议智能体](../mascot/meeting-agents.zh-CN.md),升级触发器可以落地的地方之一(日历事件有 Meet 链接)。
+94
View File
@@ -0,0 +1,94 @@
---
description: >-
Pair an iOS companion app to your desktop OpenHuman over an end-to-end
encrypted tunnel, scanned from a QR code.
icon: smartphone
---
# iOS Companion
The iOS Companion lets you reach your desktop OpenHuman from your phone: you scan a QR code shown on the desktop, the two devices agree on a shared key, and from then on the phone talks to the desktop core over an encrypted channel.
{% hint style="warning" %}
**Experimental / non-shipping.** The iOS client is in-progress and is **not** part of the shipped desktop product. APIs, wire formats, and the pairing flow can change without notice, and an upgrade may force you to re-pair. Treat everything below as a developer preview.
{% endhint %}
The desktop core is always the source of truth. The phone is a thin client — it does not run its own agent, it relays requests to the core and renders the results.
***
## What it is
Pairing is brokered by the Rust `devices` domain in the core. The core registers a pairing channel with the tinyhumans backend's `tunnel:*` Socket.IO relay, generates a fresh X25519 keypair, and renders a QR code. The phone scans it, generates **its own** X25519 keypair, and connects back over the same relay. The backend is a **blind forwarder** — it relays opaque frames and never sees plaintext.
Once paired, the device shows up in **Settings → Devices** on the desktop with an online/offline dot, and can be revoked at any time.
***
## Pairing via QR code
```
Desktop core Backend relay iOS app
| | |
|-- devices_create_pairing RPC | |
|-- tunnel:register ----------------->| |
|<-- channel_id, expires_at ----------| |
|-- generate X25519 keypair | |
|-- tunnel:connect (role: core) ----->| |
| | |
| shows QR: | |
| cid, pt, cpk, rpc?, exp | |
|.................. scan QR ......................> |
| | generate device |
| | X25519 keypair |
| |<-- tunnel:connect ------|
| | (role: client) |
|<------ tunnel:frame (handshake) ----|------------------------|
|-- X25519 DH + derive session keys | |
|-- persist PairedDevice | |
|-- publish DevicePaired event | |
| device appears in Devices list | |
```
The QR payload (carried as an `openhuman://pair?...` deep link) contains the channel id (`cid`), a single-use pairing token (`pt`), the core's public key (`cpk`), an optional LAN URL (`rpc`), and an expiry (`exp`). The pairing token is single-use, hashed at rest on the backend, and the QR is rejected client-side once `exp` has passed (the backend enforces the real ~10 minute TTL).
***
## The end-to-end tunnel
Confidentiality and integrity live entirely on the two endpoints. The exact primitives, from `src/openhuman/devices/crypto.rs`:
* **Key agreement:** X25519 Diffie-Hellman. Each side has a long-term static keypair (the core's is in the QR; the device's is minted at scan time) plus an ephemeral keypair minted per session for forward secrecy.
* **Session-key derivation:** HKDF-SHA256 over `ikm = static_dh || eph_dh`, salted with `client_eph_pub || server_eph_pub`. Two **directional** 32-byte subkeys are expanded with distinct info tags — `openhuman-tunnel/v1/c2s` and `openhuman-tunnel/v1/s2c` — so a frame one side seals can never decrypt under its own opener (closes the cross-direction reflection attack class).
* **Frame cipher:** XChaCha20-Poly1305 (AEAD, 192-bit nonce). Wire format is `version(0x02) || nonce(24) || ciphertext+tag`, with a random nonce per frame.
* **Replay protection:** a sliding window over the last 128 nonces seen per opener.
Static DH authenticates the peer via the QR-code provenance; ephemeral DH means a later static-key leak cannot decrypt past traffic. The legacy single-key `version=0x01` frame shape is rejected with an explicit "re-pair required" error — peers must re-pair after an upgrade. Outbound frames are capped at 64 KB.
***
## Transport strategies
The phone may reach the core three ways. `TransportManager` (`app/src/services/transport/`) picks one from the saved `ConnectionProfile`; for a paired device it **races LAN against the tunnel** (2 s LAN timeout) and uses whichever answers `openhuman.ping` first.
| Strategy | Class | When it's used | Trade-offs |
| --- | --- | --- | --- |
| **LAN HTTP** (`LanHttpTransport`) | Direct HTTP to the core's LAN `rpc_url` | Phone and desktop on the same network | Fastest, lowest latency. Requires same LAN; not encrypted by this layer (relies on local network trust). |
| **Tunnel** (`TunnelTransport`) | E2E encrypted frames over the backend Socket.IO relay | Anywhere with internet; default fallback | Works across networks; X25519 + XChaCha20-Poly1305 end to end. Higher latency (relayed); depends on backend availability. |
| **Cloud HTTP** (`CloudHttpTransport`) | HTTP to a cloud-hosted core endpoint | Profile `kind: "cloud"`, when LAN and tunnel are unreachable | Reachable from anywhere; depends on a hosted core and its own auth. |
***
## Device management & revocation
Paired devices are persisted by the core in SQLite (`{workspace_dir}/devices/devices.db`, table `paired_devices`): channel id, label, the device's public key, a SHA-256 hash of the core session token, and timestamps. The core's X25519 private key is stored encrypted at rest (via the OS keyring `SecretStore`) so handshakes survive a restart.
* **List** — `devices_list` returns non-revoked devices, overlaying a live `peer_online` flag sourced from `tunnel:peer-status` (online status is never persisted).
* **Revoke** — `devices_revoke` soft-deletes the device, tears down all in-memory and tunnel state for the channel, and publishes a `DeviceRevoked` event. Today revocation is local-side: the backend channel is left to expire via its pairing-token TTL (a backend revoke endpoint is a follow-up).
***
## See also
* [Privacy & Security](privacy-and-security.md) — how OpenHuman handles your data and keys.
* [Voice](native-tools/voice.md) — push-to-talk and dictation, the headline use case for a phone companion.
-73
View File
@@ -1,73 +0,0 @@
---
description: >-
OpenHuman 的屏幕之脸——一个桌面吉祥物,能说话、能反应、
能加入你的会议、能在你不看的时候在后台思考。
icon: face-smile
---
# 吉祥物
OpenHuman 有一张脸。吉祥物是一个生活在你桌面上的动画角色,作为智能体的可见表面——它在说什么、它在思考什么、它何时空闲、何时忙碌、何时有话要告诉你。
它不是装饰性镀层。吉祥物接入智能体同一套组件:语音、记忆、[潜意识循环](../subconscious.zh-CN.md)和 [Google Meet 集成](../native-tools/voice.zh-CN.md)。智能体说话时,吉祥物就是说话的那个;智能体思考时,吉祥物就是思考的那个。
## 它做什么
### 它说话,并与自己的声音口型同步
智能体回复时,音频通过托管 TTS 模型生成并流式传输到你的扬声器。同时,吉祥物驱动一个 viseme 贴图与音频对齐,这样它的嘴型与说出的词语相匹配。没有单独的"说话头像"视频,你听到的同一音频流驱动着动画。
吉祥物所依赖的语音转文字、文字转语音、会议管道见[原生语音](../native-tools/voice.zh-CN.md)。
### 它加入你的会议,作为真实参与者
吉祥物是 OpenHuman 的旗舰语音集成。它可以作为真实参与者加入 Google Meet 会议:它听到每个人、将笔记记入你的[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)、当它有话要说时在通话中说话,并将其自己的动画脸作为摄像头画面管道输送到会议中。
这是头条用例,有专门页面,见[会议智能体](meeting-agents.zh-CN.md)。
### 它移动并对周围环境做出反应
吉祥物有情绪状态(空闲、思考、倾听、说话、惊讶、做梦),它根据智能体的行为在状态间转换。当你开始打字时它切换到倾听姿势。当模型在推理时它显示出来。当工具调用返回值得注意的内容时它做出反应。当你停止交互一段时间后,它进入空闲状态。
它应该让人感觉是活的,而非轨道动画。
### 它记得你
吉祥物是拥有[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)的智能体的可见部分。它记得你们聊过什么、你生活中的人是谁、你盘子上有什么、已决定了什么、什么还悬而未决,跨越你连接的所有来源。当它早上问候你时,它不是从零开始。
这种记忆使性格在数周和数月间保持一致。今天和你说话的吉祥物知道上周二和你说话的吉祥物知道的东西。
### 它在后台思考——潜意识
即使你已经停止打字,吉祥物也在继续思考。[潜意识循环](../subconscious.zh-CN.md)是一个后台 tick
* 加载你的待办任务和背景目标。
* 读取你工作区和最近记忆的当前状态。
* 决定对每项做什么(自主执行、保留、或升级给你审批)。
* 将结果写入你可以审计的活动日志。
所以当你回到桌前,吉祥物可能已经起草了邮件、刷新了仪表板、或排队了它需要问你的问题。屏幕上的那张脸就是做了工作的那张。
### 它做梦
当你离开得足够久,吉祥物进入做梦状态。做梦是智能体的离线整合过程,将一天的块凝练为更长期限的摘要、刷新已升温实体的主题树、浮现不符合任何单一来源的模式。吉祥物在做梦时动画不同,这样你可以一眼看出:它不是空闲,它在处理。
当你回来时,做梦已经折叠到记忆树中。吉祥物醒来时比睡前更聪明。
## 为什么要有一个吉祥物?
大多数助手只是一个闪烁的文本输入。对工具来说这没问题。对于要整天陪伴你、对你生活有持久记忆、代表你执行操作的东西来说,这还不够。
吉祥物的存在是因为:
* **存在胜过面板。** 你可以扫一眼的脸在一帧中告诉你智能体是忙碌、空闲、做梦还是在试图引起你注意。
* **它让语音通话感觉像对话。** 一个与自己的语音口型同步的动画角色的摄像头画面,与黑色方块的机器人声音是截然不同的体验。
* **性格是一个 UX 表面。** 屏幕上始终如一的角色比无脸的 API 更值得信任、更容易交谈、更容易原谅错误。
## 另见
* [会议智能体](meeting-agents.zh-CN.md),吉祥物在 Google Meet 中:倾听、说话、动画、使用工具。
* [原生语音](../native-tools/voice.zh-CN.md),吉祥物所依赖的 STT / TTS 管道。
* [记忆树](../obsidian-wiki/memory-tree.zh-CN.md),吉祥物记住什么以及如何记住。
* [潜意识循环](../subconscious.zh-CN.md),你离开时它在思考什么。
* [Chromium Embedded Framework](../../developing/cef.zh-CN.md),摄像头进入 Meet 的管道(开发者参考)。
@@ -1,94 +0,0 @@
---
description: >-
吉祥物作为真实参与者加入会议:倾听、记笔记、在通话中说话、
将动画脸管道到摄像头网格,并在会议中间使用工具。不只是笔记工具。
icon: video
---
# 会议智能体
吉祥物的旗舰集成是**会议智能体**:你在桌面上对话的同一角色可以代表你加入 Google Meet,坐在参与者网格中作为动画脸,听到房间里的每个人,用自己的声音在通话中说话,并在会议进行时使用工具。
它不是笔记工具。笔记工具安静地坐着产生转录。会议智能体参与——它回答问题、实时查找、在与同一个人之前的会议中记住事情,并在你(或它)决定有用的内容要补充时做出贡献。
## 它在通话中实际做什么
### 1. 作为真实参与者加入
吉祥物通过嵌入式 webview 加入会议,与一个人从浏览器加入的方式相同。网格中有一个名字、一张脸和一个瓦片。其他参与者像看到任何其他与会者一样看到和听到它——没有日历 bot、没有拨入号码、没有"此会议正被……录制"横幅。
在底层,会议大脑位于 `src/openhuman/meet_agent/brain.rs`webview 端是 OpenHuman 用于其他嵌入式 provider 的相同 CEF 子窗口。
### 2. 它倾听房间里的每个人
会议的入站音频被捕获并实时推送通过流式语音转文字。转录按说话者分离,经过与桌面听写相同的幻觉过滤和后处理,并在会议展开时折叠到[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)中——在正确的人、正确的主题、正确的项目下,带有吉祥物以后可以使用的反向链接。
因为转录正在实时结构化,吉祥物可以在会议仍在进行时回答关于_这个_会议(或与同一个人任何之前的会议)的问题。
### 3. 它互动——回答、提问、跟进
智能体没有静音。当你是指向它("Ghosty,你能拉出上个季度的数字吗?"),或者当它决定有用的内容要补充时,它使用项目正常 LLM 堆栈实时生成回复并在会议中说话。
对话轮次通过快速模型层路由(参见[自动模型路由](../model-routing/README.zh-CN.md)),这样延迟感觉像在和一个正在倾听的人说话,而不是等待聊天机器人。
### 4. 它说话——自己的 TTS 音频播放回通话
回复由项目 TTS 堆栈生成并直接作为出站麦克风 feed 流式传输到会议中。它不是通过你的本地扬声器播放并被你的麦克风重新捕获——它直接作为智能体的音频注入,所以它干净地到达其他每个人,不会通过你的房间回声。
### 5. 它动画——吉祥物的脸就是摄像头 feed
吉祥物的画布被管道到 Meet 通话作为出站摄像头流(commit `b6d05cb4` 引入的工作,Mascot 帧流水线在 `f5dce783` 中进一步打磨)。当智能体在说话时,吉祥物在摄像头瓦片上说话——嘴型与所有其他人听到的同一 TTS 音频口型同步。当它在倾听时,它显示倾听姿势。当它在说话前推理时,你看到思考姿势。
其他参与者在网格中看不到黑色瓦片或静态头像。他们看到一个与正在说的话实时反应的动画角色,这使得通话感觉像与活着的东西对话,而不是声音从无处传来。
### 6. 它在会议中间使用工具——这是笔记工具做不到的部分
这就是转录 bot 和会议_智能体_之间的区别。
当通话发生时,吉祥物可以访问它在桌面上相同的工具表面:
- [**记忆树**](../obsidian-wiki/memory-tree.zh-CN.md)——召回之前的会议、决策、开放线程、谁上次说了什么、承诺了什么。
- [**从集成自动拉取**](../obsidian-wiki/auto-fetch.zh-CN.md)和[**第三方集成**](../integrations/README.zh-CN.md)——从 Slack 拉取线程、一封邮件、一个 Linear ticket、一个 Notion 文档、一个日历条目、一个 Drive 文件。
- [**原生工具**](../native-tools/README.zh-CN.md)——搜索网络、抓取页面、运行快速代码/数据查询,全部不离开通话。
- [**潜意识循环**](../subconscious.zh-CN.md)输出——任何它在后台一直在工作的东西都随手可得。
所以当通话中有人问"等等,我们不是上个月决定放弃 Q3 发布的吗?",吉祥物不只是转录问题。它回答它——用实际的决策、做出它的会议、以及谁同意了。
这将它从_笔记工具_移到_房间里信息最丰富的参与者_。
## 为什么它感觉是活的
只转录的会议智能体是工具。参与的会议智能体是一种存在。Meet 集成刻意构建为让吉祥物感觉像一个真正的与会者,而不是录音设备:
- 它在摄像头网格上有**一张脸**,会口型同步和反应,不是黑色方块或标志。
- 它有**自己的声音**,播放到通话中,而不是你的扬声器。
- 它有**持久记忆**房间里的、项目、之前的决策——所以它可以被命名并上下文回答。
- 它有**工具**,所以它可以行动于所说的话,而不只是记录。
- 它在会议之间运行**潜意识循环**——所以当它加入你的下一个通话时,它已经做完了在上一个会议中承诺的功课。
实际结果是,参与者不再把它当作 bot 开始对待,而是开始把它当作一个恰好非常快速查找东西的同事。
## 设置、控制、隐私
- **加入通话。** 你可以从桌面 app 给吉祥物一个 Google Meet 链接;它将打开嵌入式 Meet webview,用配置的显示名称加入,并将其摄像头瓦片切换到吉祥物画布。
- **麦克风和摄像头控制。** 智能体的麦克风是 TTS 注入流,不是你真正的麦克风。智能体的摄像头是吉祥物帧生成器,不是你真正的网络摄像头。你可以随时从 app 中将智能体的麦克风静音,就像在 Meet 中静音自己一样。
- **转录和记忆。** 实时转录以与任何其他来源相同的方式落在[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)中——在通话中的人、项目和出现的主题下。它们是本地优先的,遵循项目的[隐私与安全](../privacy-and-security.zh-CN.md)规则。
- **无秘密录制。** 智能体在网格中作为正常参与者出现;通话中的每个人都可以看到它,并在它说话时看到。
## 开发者实现指针
好奇这是如何连接的:
- 大脑 - `src/openhuman/meet_agent/brain.rs`(LLM 轮次、说话/不说话决定、工具调用)。
- 语音管道 - `src/openhuman/voice/`STT in、TTS out、幻觉过滤、后处理)。参见[原生语音](../native-tools/voice.zh-CN.md)。
- 作为出站摄像头的吉祥物画布 - `app/src/features/meet/MascotFrameProducer.tsx` 和 Tauri 端 `mascot_native_window.rs` 窗口。
- 嵌入式 Meet webview - 参见 [Chromium Embedded Framework](../../developing/cef.zh-CN.md)。与六个零注入提供商不同,Google Meet 使用 CDP 驱动 join 自动化:CDP 通过 `Page.addScriptToEvaluateOnNewDocument``Runtime.evaluate` 注入 bridging scripts / recipes`GOOGLE_MEET_RECIPE_JS` / `provider_recipe_js("google-meet")`),而不是真正的零注入交付。
- 要阅读的上下文的重要 commit - `0bc74575`(实时记笔记)、`f1203479`(真实 LLM 轮次 + 调优 TTS)、`b6d05cb4`(吉祥物画布作为出站摄像头)、`f5dce783`(吉祥物帧流水线 + 屏外会议窗口)。
## 另见
- [吉祥物](./)——屏幕上的角色本身,会议之外。
- [原生语音](../native-tools/voice.zh-CN.md)——会议智能体所依赖的 STT / TTS。
- [记忆树](../obsidian-wiki/memory-tree.zh-CN.md)——转录和决策落地的地方。
- [原生工具](../native-tools/README.zh-CN.md)——吉祥物在通话中可以伸手拿什么。
- [自动模型路由](../model-routing/README.zh-CN.md)——对话轮次为什么感觉低延迟。
@@ -1,63 +0,0 @@
---
description: >-
一个订阅,多个模型。任务通过 hint 前缀选择模型:
推理发给强模型,快速路径发给快模型,视觉发给视觉模型。
icon: route
---
# 自动模型路由
智能体的不同部分需要不同的模型。长推理需要前沿模型。快速的"修这个拼写错误"需要又快又便宜的模型。视觉需要视觉模型。OpenHuman 通过内置**路由 provider**处理这一切,所以你永远不需要考虑它。
## 请求如何被路由
任何聊天调用上的 model 参数可以取两种形式:
- **具体模型名**。例如 `anthropic/claude-sonnet-4`。路由到带该精确模型的默认 provider。
- **Hint 前缀**。例如 `hint:reasoning`。在路由表中查找 hint 并解析为 `(provider, model)` 对。
```rust
// src/openhuman/providers/router.rs
fn resolve(&self, model: &str) -> (usize, String) {
if let Some(hint) = model.strip_prefix("hint:") {
if let Some((idx, resolved_model)) = self.routes.get(hint) {
return (*idx, resolved_model.clone());
}
}
(self.default_index, model.to_string())
}
```
路由器包装了多个预创建的 providersAnthropic、OpenAI、Google、Groq 等),每次请求选择正确的一个。Hint 可以在运行时重新映射而无需重启 core。
## 常见 hint
| Hint | 典型目标 | 使用场景 |
| --- | --- | --- |
| `hint:reasoning` | 强推理模型 | 多步规划、数学、重度代码轮次 |
| `hint:fast` | 快速/便宜模型 | UI 助手、自动补全、小型分类调用 |
| `hint:vision` | 有视觉能力的模型 | 截图、图像附件、OCR |
| `hint:summarize` | 擅长压缩的模型 | 记忆树摘要构建器 |
| `hint:code` | 代码调优的模型 | 原生编码器轮次 |
精确映射可配置;默认值提供每个 provider 的合理路由。
## 一个订阅
路由在单一 OpenHuman 订阅背后发生。你不需要分别为 Anthropic、OpenAI、Google 等持有单独的 API 密钥,后端经纪访问,路由器为每个任务选择正确的一个。这就是 README 中"一个订阅,多个 provider"的承诺,具体化了。
## 覆盖路由
- **全局**。配置 TOML`src/openhuman/config/schema/types.rs` 中的 `Config` 结构体)可以在启动时提供自定义路由表。
- **每次调用**。传递具体模型名(无 `hint:` 前缀),路由器回退到带该精确模型的默认 provider。
- **对于技能**。技能可以在其 manifest 中固定一个 hint 或模型。
## 为什么这不是简单的"模型切换器"
路由不是 UI 下拉菜单。智能体循环本身根据它要做什么发出 hint。你不选择模型;*任务*选择。这就是"多模型"和"智能路由"的区别。
## 另见
- [智能 Token 压缩](../token-compression.zh-CN.md)。什么使大型推理调用负担得起。
- [原生工具](../native-tools/README.zh-CN.md)。不同的工具调用暗示不同的路由。
- [本地 AI(可选)](local-ai.zh-CN.md)。轻量聊天 hint 可以在设备上运行。
@@ -1,99 +0,0 @@
---
description: >-
可选、自愿开启的本地 AI,通过 Ollama 或 LM Studio 提供。
为记忆嵌入向量、摘要树构建和后台推理循环提供端侧支持。聊天/视觉/语音走云端。
icon: microchip
---
# 本地 AI(可选)
OpenHuman 可以为以下工作负载在你机器上运行本地模型:当本地保留数据最为重要时:**记忆嵌入向量、摘要树构建和后台推理循环**。它是**自愿开启**的,默认**关闭**。
这是一个刻意的范围界定。之前的设计尝试将聊天、视觉、STT 和 TTS 全部放在 Gemma 3 的设备上,结果是对硬件较敏感的资源占用,与产品其余部分所需的东西冲突。如今,本地最有价值的东西(循环、低延迟、隐私敏感的内存工作)走本地;最有价值于前沿模型的东西(默认聊天、推理、视觉)走云端。
## 开启后什么在本地运行
| 工作负载 | 默认模型 | 实现 |
| ------------------------- | --------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| **记忆嵌入向量** | `all-minilm:latest` | `src/openhuman/embeddings/ollama.rs`——用于[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)向量搜索。 |
| **摘要树构建** | `gemma3:1b-it-qat`(可配置) | `src/openhuman/tree_summarizer/ops.rs`——记忆树的源/主题/全局摘要构建器。 |
| **心跳循环** | 小型聊天模型 | `src/openhuman/heartbeat/`——周期性后台反思。 |
| **学习 / 反思** | 小型聊天模型 | `src/openhuman/learning/reflection.rs`——巩固所学内容的通过。 |
| **潜意识** | 小型聊天模型 | `src/openhuman/subconscious/executor.rs`——后台评估循环。 |
每个都是**按功能开启的 opt-in flag**。开启本地 AI 不会静默将所有内容路由到它,你选择工作负载。
## 什么留在云端
| 工作负载 | 为什么走云端 |
| ------------------ | --------------------------------------------------------------------------------------------------- |
| **聊天(默认)** | 前沿推理质量。通过[模型路由器](README.zh-CN.md)在单一订阅下路由。 |
| **视觉** | 同上。 |
| **STT** | 后端代理转录(`src/openhuman/voice/cloud_transcribe.rs`)。 |
| **TTS** | 底层托管[文字转语音](../native-tools/voice.zh-CN.md)`reply_speech.rs`)。 |
| **网络搜索** | 后端代理(你的机器上没有 API key)。 |
对于**轻量级或中等聊天 hint**(`hint:reaction``hint:classify``hint:format``hint:sentiment``hint:summarize``hint:medium``hint:tool_lite`),当本地 AI 开启且 Ollama 可达时,[路由器](README.zh-CN.md)会优先使用本地 provider。重型 hint`hint:reasoning``hint:agentic``hint:coding`)走云端。
## 工作原理
在底层,OpenHuman 支持两种本地 provider 路径:
* [Ollama](https://ollama.com),用于捆绑模型生命周期、嵌入向量和现有模型资产流。
* [LM Studio](https://lmstudio.ai),通过其本地 OpenAI 兼容服务器用于聊天风格本地推理。
对于 OllamaOpenHuman 在可能的情况下与其 OpenAI 兼容的 `/v1` 端点对话。这意味着:
* `OpenAiCompatibleProvider``src/openhuman/providers/compatible.rs`)与 Ollama 的包装方式与与远程 OpenAI 风格 provider 完全相同。没有特殊案例代码路径。
* Provider 路由器在启动时创建一个_健康门控_的本地 provider。如果 Ollama 不可达,请求透明地回退到远程 provider,没有破碎状态。
* 模型按需由 Ollama 拉取并缓存在其自己的存储中。OpenHuman 自己不附带权重。
对于 LM Studio,设置 `local_ai.provider = "lm_studio"` 并确保 LM Studio 本地服务器正在运行。OpenHuman 默认为 `http://localhost:1234/v1`,探测 `GET /v1/models`,并将聊天请求发送到 `POST /v1/chat/completions`。你可以用 `local_ai.base_url``OPENHUMAN_LM_STUDIO_BASE_URL``LM_STUDIO_BASE_URL` 覆盖端点。
## 选择加入
本地 AI 由 core 配置中的两个 flag 门控(`src/openhuman/config/schema/local_ai.rs`):
| Flag | 默认 | 含义 |
| ------------------------------------ | ------- | ------------------------------------------------------------------- |
| `local_ai.runtime_enabled` | `false` | 主开关。`false` ⇒ 根本不创建本地 provider。 |
| `local_ai.opt_in_confirmed` | `false` | 明确的 opt-in 标记。除非你重新 opt-in,否则 Bootstrap 强制为 `false`。 |
| `local_ai.provider` | `ollama` | 本地 provider`ollama``lm_studio`。 |
| `local_ai.base_url` | 未设置 | 可选的 provider URL。LM Studio 默认为 `http://localhost:1234/v1`。 |
| `local_ai.usage.embeddings` | `false` | 使用本地进行记忆嵌入向量。 |
| `local_ai.usage.heartbeat` | `false` | 使用本地进行心跳循环。 |
| `local_ai.usage.learning_reflection` | `false` | 使用本地进行学习通过。 |
| `local_ai.usage.subconscious` | `false` | 使用本地进行潜意识循环。 |
在桌面 app 中,**设置 → AI 与技能 → 本地 AI** 暴露预设,选择一个("仅嵌入向量"、"记忆 + 反思"、"全部本地"),正确的 flag 组合会为你设置。状态(Ollama 可达性、模型可用性、每个子系统启用)通过 `openhuman.inference_status` 实时暴露。
## 何时开启
如果以下任一为真,开启本地 AI 是值得的:
* 你摄入大量邮件 / 聊天并希望**嵌入向量永不离开机器**。
* 你希望**摘要树构建**离线工作。
* 你对后台反思("潜意识")循环隐私敏感。
如果你的连接源很少,云端路径更快,隐私收益很小,则**不值得**开启。也有硬件成本:Ollama 和一个小型 Gemma 模型需要几 GB 的 RAM 并拉取几 GB 的权重。
## 你需要什么
* 安装并运行本地的 [**Ollama**](https://ollama.com),或启用本地服务器的 [**LM Studio**](https://lmstudio.ai)。
* 模型有足够的磁盘(`gemma3:1b-it-qat` \~700 MB`all-minilm:latest` \~23 MB)。
* 有足够的 RAM 保持模型驻留(建议 8 GB+,理想 16 GB+)。
OpenHuman 处理其余:生命周期(`src/openhuman/inference/local/service/`)、API 客户端、健康检查,以及当本地 provider 消失时优雅地回退到远程。
### LM Studio 故障排除
* 确认 LM Studio 本地服务器已启用并在 `http://localhost:1234/v1` 可达。
* 在调用 OpenHuman 之前在 LM Studio 中加载所选模型。当配置的 `local_ai.chat_model_id` 不在 `/v1/models` 中时,诊断报告 `load_lm_studio_model`
* 如果 LM Studio 使用不同端口,设置 `local_ai.base_url``OPENHUMAN_LM_STUDIO_BASE_URL`
* LM Studio 模型下载在 LM Studio 内部管理。OpenHuman 不会从本地资产下载控制中拉取 LM Studio 模型。
## 另见
* [记忆树](../obsidian-wiki/memory-tree.zh-CN.md)。本地嵌入向量 + 摘要 powering 什么。
* [自动模型路由](README.zh-CN.md)。轻量聊天 hint 如何优先使用本地 provider。
* [隐私与安全](../privacy-and-security.zh-CN.md)。当你 opt-in 时什么移至端侧。
@@ -1,42 +0,0 @@
---
description: >-
OpenHuman 智能体开箱即用的完整工具集——研究、编码、
控制你的机器、安排任务、回复你,以及调用 118+ 第三方服务。
icon: toolbox
---
# 原生工具
OpenHuman 的智能体并非空载交付。智能体背后的每个模型在安装瞬间就有一套精选工具可用——无需插件市场、无需接入 API 密钥、无需注册 MCP 服务器。整个工具带都在盒子里。
本页是索引。每个子页面覆盖一个工具族。
## 为什么原生提供这些工具
纯插件模式意味着工具跑在不同进程里,通过 RPC 交互,各自维护认证和打包逻辑。这对于开放式扩展性没问题,但对于每个智能体都需要的**核心**工具(读文件、搜索网页、编辑代码、设提醒、加入会议),以内置方式提供意味着:
* 一致的错误处理。
* 零安装门槛。
* 所有输出自动经过[智能 Token 压缩](../token-compression.zh-CN.md)。
* 可预测的安全边界——文件系统工具遵守工作区作用域,网络工具通过 OpenHuman 代理。
## 工具带
| 类别 | 包含内容 |
| ------ | -------------- |
| [网络搜索](web-search.zh-CN.md) | 无需自带 API key 搜索实时网页。 |
| [网页抓取](web-scraper.zh-CN.md) | 从任意 URL 拉取干净文本——文章、文档、README。 |
| [编码器](coder.zh-CN.md) | 读/写/编辑/补丁文件,globgrepgitlinttest。 |
| [浏览器与计算机控制](browser-and-computer.zh-CN.md) | 打开 URL、截图、点击、输入、移动鼠标。 |
| [定时任务与调度](cron.zh-CN.md) | 循环任务、一次性提醒、定时智能体运行。 |
| [语音](voice.zh-CN.md) | 语音转文字输入、文字转语音输出、实时 Google Meet 智能体。 |
| [记忆工具](memory-tools.zh-CN.md) | 在[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)中召回、存储、遗忘和搜索。 |
| [第三方集成](../integrations/README.zh-CN.md) | 智能体视角中的 [118+ 已连接服务](../integrations/README.zh-CN.md)。 |
| [智能体协作](agent-coordination.zh-CN.md) | 生成子智能体、委托给技能、规划、询问用户。 |
| [系统与工具](system-and-utilities.zh-CN.md) | Shell、node、SQL、当前时间、推送通知、LSP。 |
## 另见
* [智能 Token 压缩](../token-compression.zh-CN.md) —— 保持工具输出成本有界的机制。
* [第三方集成](../integrations/README.zh-CN.md) —— 118+ 目录的面向用户介绍和 OAuth 流程。
* [隐私与安全](../privacy-and-security.zh-CN.md) —— 每个工具运行所在的安全边界。
@@ -1,37 +0,0 @@
---
description: 智能体用来规划、委托和求助的工具。
icon: sitemap
---
# 智能体协作
除了做工作,智能体还有用于*组织*工作的工具——规划多步任务、委托给专家、生成子智能体,以及当某些东西真正模糊时暂停询问用户。
## 系列中的工具
| 工具 | 功能 |
| ----------------------- | --------------------------------------------------------------------------------------------- |
| `todo_write` | 在长任务中维护结构化 TODO 列表。随着工作进展标记完成。 |
| `spawn_subagent` | 为独立子任务启动具有自己上下文窗口的新智能体。 |
| `spawn_worker_thread` | 不需要阻塞主对话的后台工作。 |
| `delegate` | 将任务交给专家(例如具有不同提示/工具/权限的原型)。 |
| `archetype_delegation` | 路由到命名原型——coder、researcher、planner 等。 |
| `skill_delegation` | 交接给工作区中安装的[技能](../integrations/README.zh-CN.md#skills)。 |
| `ask_clarification` | 暂停并向用户提出精确问题,而不是猜测。 |
| `plan_exit` | 退出规划阶段并开始执行。 |
| `check_onboarding_status` / `complete_onboarding` | 根据用户是否完成入门进行门控。 |
## 为什么这些是工具,不是隐式行为
长任务在智能体试图将所有东西保存在一个头脑中时会崩溃。通过 TODO 和子智能体拆分工作意味着:
* 每个子智能体获得干净的上下文——更少 token、更少干扰。
* 主线程保持高级别进度视图。
* 一个分支中的失败不会污染其余。
询问澄清也是一个工具,是刻意的:这使得"我应该问用户"成为一个*可见的*决定,智能体可以被引导,而不是紧急出现的行为。
## 另见
* [编码器](coder.zh-CN.md)——coder-archetype 子智能体通常使用什么。
* [潜意识循环](../subconscious.zh-CN.md)——始终开启的后台智能体线程。
@@ -1,33 +0,0 @@
---
description: 原生打开 URL、截图、点击、输入、移动鼠标。
icon: display
---
# 浏览器与计算机控制
当智能体需要像人一样*使用*你的机器时——打开页面、截图、点击按钮、输入短语——这些工具就是它做这些事的方式。
## 浏览器
* **打开**一个 URL,进入智能体可以回读的嵌入式 webview。
* **截图**当前页面。
* **检查**图像输出和元数据,以便智能体描述它看到的内容。
浏览器界面通过 CEFChromium Embedded Framework)运行,并包含一个安全层,限制页面能做什么。参见 [Chromium Embedded Framework](../../developing/cef.zh-CN.md) 了解平台详情。
## 计算机(鼠标 + 键盘)
* **鼠标**——移动、点击、拖拽。
* **键盘**——输入文本、发送快捷键。
* **类人路径**——移动和点击遵循类人轨迹,而非瞬移,因此不会触发简单的机器人检测。
## 适用于
* 驱动没有 API 或没有[原生集成](../integrations/README.zh-CN.md)的网站。
* 单次截图不够的多步骤 UI 流程。
* 在聊天中自动化本地应用。
## 另见
* [网页抓取](web-scraper.zh-CN.md) —— 当你只需要文章而非整个页面时。
* [Chromium Embedded Framework](../../developing/cef.zh-CN.md) —— 运行时浏览器层。
@@ -1,43 +0,0 @@
---
description: 一个完整的工具集,用于处理真实代码库——读、写、编辑、搜索、git、lint、test。
icon: code
---
# 编码器
编码器系列使 OpenHuman 成为可行的编码伙伴,而不是一个*假装*了解代码库的聊天窗口。
## 系列中的工具
| 工具 | 功能 |
| ---------------- | ----------------------------------------------------------------- |
| `file_read` | 读文件(带行号,像 `cat -n`)。 |
| `file_write` | 写一个新文件。 |
| `edit_file` | 定向编辑——严格唯一性检查的匹配替换。 |
| `apply_patch` | 应用统一 diff。 |
| `glob_search` | 按 glob 模式查找文件。 |
| `grep` | 跨树 ripgrep 风格搜索。 |
| `list_files` | 遍历目录树。 |
| `read_diff` | 两个文件或版本之间的 diff。 |
| `git_operations` | Status、diff、log、blame、branch、commit。 |
| `run_linter` | 运行项目的 linter。 |
| `run_tests` | 运行项目的 test 命令。 |
| `csv_export` | 将查询结果导出为 CSV。 |
## 为什么这些是原生的,而非纯 shell
Shell 工具加 `cat`/`sed`/`awk` 技术上可以完成所有这些。原生工具存在是因为:
* 编辑通过唯一性检查,所以智能体不会意外覆盖错误的行。
* 读取返回智能体可以在后续中引用的行号。
* Git 操作将输出解析为结构化数据,而不是让智能体刮擦 porcelain。
* Lint 和 test 运行连接到项目的实际命令,而非通用猜测。
## 工作区作用域
文件系统工具遵守工作区边界——智能体未经明确许可不能在其外部读写。边界与应用的其余部分用于 `OPENHUMAN_WORKSPACE` 的相同。
## 另见
* [系统与工具](system-and-utilities.zh-CN.md) —— `shell``node_exec``npm_exec` 用于开发循环的其余部分。
* [智能体协作](agent-coordination.zh-CN.md) —— `todo_write``spawn_subagent` 用于更大的重构。
@@ -1,37 +0,0 @@
---
description: 循环任务、一次性提醒和定时智能体运行——一等公民。
icon: clock
---
# 定时任务与调度
调度是一等公民能力,而非权宜之计。智能体可以设置循环任务("每个工作日早上 9 点,总结我的收件箱")、一次性提醒("三小时后提醒我这件事")以及按 cron 时间表运行的任意智能体任务。
## 系列中的工具
| 工具 | 功能 |
| ------------- | ------------------------------------------------------------------ |
| `cron_add` | 创建新计划任务——cron 表达式 + 智能体提示。 |
| `cron_list` | 列出现有任务及其下次运行时间。 |
| `cron_update` | 编辑现有任务——更改时间表、提示或启用状态。 |
| `cron_remove` | 删除任务。 |
| `cron_run` | 立即运行一次任务,无论其时间表如何。 |
| `cron_runs` | 检查最近运行历史——何时、多久、产生了什么。 |
[系统与工具](system-and-utilities.zh-CN.md)中还有一个一次性 `schedule` 工具,用于"在时间 T 做一次"而不需要循环条目的情况。
## 适用于
* 按你选择的消息渠道发送的每日/每周摘要。
* 轮询没有推送事件的慢速集成。
* 智能体自己拥有的提醒("周四提醒我跟进 Alice")。
* 循环研究——"每周一,检查这个话题有什么新内容,给我写个简报"。
## 如何与其余部分关联
每次 cron 运行都是一次正常的智能体调用,所以它可以使用任何其他工具——搜索网页、查询[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)、调用[第三方集成](../integrations/README.zh-CN.md)、发消息。运行历史被记录,这样你可以看到每个 tick 产生了什么。
## 另见
* [系统与工具](system-and-utilities.zh-CN.md) —— 一次性 `schedule` 工具。
* [智能体协作](agent-coordination.zh-CN.md) —— 向子智能体扇出的任务。
@@ -1,33 +0,0 @@
---
description: 智能体对 118+ 已连接第三方服务的视图。
icon: plug
---
# 第三方集成
OpenHuman 的智能体可以通过单一代理工具接口调用 [118+ 第三方服务](../integrations/README.zh-CN.md)——Gmail、Notion、GitHub、Slack、Stripe、日历,以及长长的尾部的服务。
## 它在智能体看来如何
一旦你通过 OAuth 连接了服务,其操作就变为可调用工具。智能体不需要知道工具是与 Gmail 还是与本地文件对话——它只调用工具,代理用你的 token 通过 OpenHuman 后端路由请求,结果像任何其他工具输出一样返回。
一些变为可用的例子:
* "在 Slack 上向 #engineering 发送消息。"
* "在 openhuman 仓库中创建一个 issue。"
* "我日历上明天有什么?"
* "拉取过去 20 笔超过 $1000 的 Stripe charge。"
## 原生 vs 代理
部分服务有**原生 provider**——Rust 模块知道如何直接将服务摄入[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)(例如 Gmail 的原生摄入路径)。其他仅暴露为**代理工具**:智能体可以调用,但没有自动摄入。新的原生 provider 随着功能落地陆续添加。
## 隐私边界
OpenHuman core 从不直接调用任何第三方 API。所有请求都通过 OpenHuman 后端,该后端处理 OAuth token 和速率限制。你的 token 永不以明文形式存储在你机器的磁盘上,智能体只看到工具调用的*结果*,而不是凭据。
## 另见
* [第三方集成(目录)](../integrations/README.zh-CN.md)——面向用户的介绍、OAuth 流程和连接管理。
* [自动拉取](../obsidian-wiki/auto-fetch.zh-CN.md)——已连接服务如何流入记忆树。
* [隐私与安全](../privacy-and-security.zh-CN.md)——完整边界。
@@ -1,27 +0,0 @@
---
description: 智能体如何在对话期间读取、写入和搜索自己的长期记忆。
icon: brain
---
# 记忆工具
[记忆树](../obsidian-wiki/memory-tree.zh-CN.md) 是 OpenHuman 的知识库。记忆工具是智能体在对话期间如何与其对话的。
## 系列中的工具
| 工具 | 功能 |
| -------- | ----------------------------------------------------------------------------------------------------------- |
| `recall` | 按查询搜索记忆树——源作用域、主题作用域或全局。返回带来源的块。 |
| `store` | 写入智能体认为值得保留的新记忆(事实、偏好、一段上下文)。 |
| `forget` | 按 ID 删除记忆——当某些东西出错、过时或用户明确要求忘记时使用。 |
还有一个树感知的检索表面(深入主题、获取一天的全局摘要)——智能体根据问题选择正确的一个。
## 为什么这些是工具,不是隐式上下文
记忆树太大了,无法倾倒到每个对话中。工具让模型*询问*——"我对 Alice 知道什么?""昨天发生了什么?""提醒我上周 Stripe webhook 说了什么"——检索层只返回相关块,并附带你 Obsidian 存储库中源文件的来源追溯。
## 另见
* [记忆树](../obsidian-wiki/memory-tree.zh-CN.md)——这些工具从什么读取和写入什么。
* [自动拉取](../obsidian-wiki/auto-fetch.zh-CN.md)——树如何首先被填充。
@@ -1,36 +0,0 @@
---
description: Shell、node、SQL、当前时间、推送通知——完善工具带的小工具。
icon: gear
---
# 系统与工具
兜底系列。智能体伸手拿来完成任务的小巧、锋利工具。
## 系列中的工具
| 工具 | 功能 |
| ------------------- | ----------------------------------------------------------------------------- |
| `shell` | 运行 shell 命令。有界输出,捕获退出码。 |
| `node_exec` | 运行 Node.js 片段——用于临时脚本。 |
| `npm_exec` | 运行 `npm`/`pnpm`/`yarn` 脚本。 |
| `current_time` | 获取任意时区的当前时间,带格式化选项。 |
| `schedule` | 一次性"在时间 T 做这个"——循环任务见 [Cron](cron.zh-CN.md)。 |
| `pushover` | 向你的设备发送推送通知。 |
| `insert_sql_record` | 向智能体的结构化工作区 SQL 存储追加一行。 |
| `lsp` | 查询语言服务器(定义、引用、诊断)。 |
| `workspace_state` | 检查当前工作区——打开的文件、最近的编辑、环境。 |
| `proxy_config` | 读取或更改出站请求的代理配置。 |
| `tool_stats` | 自我反思——本会话中使用了哪些工具以及频率。 |
## 适用于
* 不适合更丰富工具家族的工组流部分。
* "就跑这个命令,告诉我它打印了什么"。
* 时间感知行为("用户现在几点?")而不是将时区假设烘焙到提示中。
* 让智能体在完成长时间运行的任务后*通知你*。
## 另见
* [编码器](coder.zh-CN.md) —— 对于文件系统重的工作,优先使用专用工具而非 `shell`
* [定时任务与调度](cron.zh-CN.md) —— 对于任何循环性的任务。
@@ -1,83 +0,0 @@
---
description: 持久化的、工具作用域规则,用于安全关键型指引和学习成果。
icon: shield-check
lang: zh-CN
---
# 工具级记忆
工具级记忆层捕获关于智能体应如何使用特定工具的**可执行指引**——它与[记忆工具](memory-tools.zh-CN.md)的通用召回不同,也与 `tool_effectiveness` 统计命名空间相区别。它是把"永远不要给 Sarah 发邮件"转化为智能体在每一轮后续中都必须遵守的硬约束的表面。
它实现了 [issue #1400](https://github.com/tinyhumansai/openhuman/issues/1400)——一个用于持久化学习成果和高优先级规则的一流存储与检索系统。
## 存储内容
每个工具都有自己的命名空间 **`tool-{tool_name}`**,与 `global``skill-{id}` 以及仅用于统计的 `tool_effectiveness` 命名空间相区分。在其内部,每条记录都是一个 `ToolMemoryRule`
| 字段 | 用途 |
| ---- | ---- |
| `id` | 每条规则的稳定 UUID。Upsert 会复用相同 id。 |
| `tool_name` | 规则适用的工具(例如 `send_email``shell`)。 |
| `rule` | 智能体必须遵循的自然语言指引。 |
| `priority` | `critical``high``normal`。驱动检索 + 压缩策略。 |
| `source` | `user_explicit``post_turn``programmatic`——来源。 |
| `tags` | 自由标签(`safety``permission`……)。 |
| `created_at` / `updated_at` | RFC3339 时间戳。 |
统计(`tool_effectiveness/tool/{name}`)和规则(`tool-{name}/rule/{id}`)按设计位于*不同*的命名空间——一个追踪"发生了什么",另一个追踪"对此该做什么"。
## 优先级层级
| 优先级 | 存储位置 | 抗压缩? |
| ------ | -------- | -------- |
| `critical` | 通过 `ToolMemoryRulesSection` 钉入**系统提示**。 | **是**——系统提示按 session 冻结,不会被 mid-session 压缩器重写。 |
| `high` | 同一块系统提示中,排在 critical 之后。 | **是**——机制相同。 |
| `normal` | 存储在命名空间中;通过 `memory_recall` 按需检索。 | 否——与任何其他命名空间记忆一样可被压缩。 |
抗压缩属性是结构性的:critical 和 high 规则 riding 在*系统提示*中,而推理后端的 prefix cache 会在整个 session 期间保持其冻结。没有任何方式能让 token 压缩静默丢弃一条 `critical` 规则。
## 捕获流水线
每轮之后有两条自动捕获路径触发(通过 `ToolMemoryCaptureHook`):
1. **用户指令**——用户消息中的 `never <verb> <noun>``don't <verb> ...``do not <verb> ...``stop <verb>ing ...` 等句子会被提升为匹配工具的 **Critical** 规则。通用名词别名将 `"email"` 映射到名为 `send_email` 的工具,`"shell"` 映射到 `bash`/`exec` 等;当没有别名匹配时,规则会落在该轮次中第一个运行的工具上,使其保持在相关调用现场附近。
2. **重复工具失败**——在一轮中失败两次或以上的工具会获得一条 **Normal** 优先级的观察记录,失败类别被内联摘要,以便智能体下次考虑该工具时有上下文。
当学习子系统开启时,该 hook 默认启用。用 `OPENHUMAN_LEARNING_TOOL_MEMORY_CAPTURE_ENABLED=0` 选择性禁用。
## 工具选择时的检索
在 session 开始时,harness 通过 `ToolMemoryStore::rules_for_prompt` 预取每条 Critical 和 High 规则,将它们渲染到 `## Tool-scoped rules` 块中,并把该块钉入系统提示。因为提示在 session 生命周期内被冻结,这些规则在每一轮的工具选择时——以及任何实际工具执行之前——都是可见的。
低优先级指引不占用提示预算;智能体通过针对 `tool-{name}` 命名空间调用 `memory_recall` 按需获取它们。
## RPC 表面
`memory` 命名空间下暴露六个方法:
| 方法 | 用途 |
| ---- | ---- |
| `memory.tool_rule_put` | Upsert 一条规则。对安全关键型条目使用 `priority='critical'`。 |
| `memory.tool_rule_get` | 通过 `(tool_name, id)` 获取一条规则。 |
| `memory.tool_rule_list` | 列出某工具的所有规则,按优先级 + 新鲜度排序。 |
| `memory.tool_rule_delete` | 删除一条规则。 |
| `memory.tool_rules_for_prompt` | 返回渲染后的 Markdown 块 + 结构化快照——session builder 所钉入的内容。 |
| `memory.tool_rules_json` | 原始 JSON 列表(供信封消费者使用)。 |
JSON payload 使用 snake_case`priority: "critical"``source: "user_explicit"`)。每个方法都经过与其他记忆 RPC 相同的 `active_memory_client` 管道。
## 端到端安全场景
"永远不要给 Sarah 发邮件"路径已被回归测试覆盖:
1. 用户在调用了 `send_email` 的轮次中说 *"Never email Sarah at sarah@example.com."*
2. `ToolMemoryCaptureHook` 提取该指令,将 `email` 别名映射到 `send_email` 工具,并在 `tool-send_email/rule/{uuid}` 下写入一条 Critical 规则。
3. 在下一个 session 中,`prefetch_tool_memory_rules_blocking` 拉取每条 Critical 和 High 规则,session builder 将 `ToolMemoryRulesSection` 追加到系统提示。
4. 智能体在选择工具之前就看到 `### \`send_email\`` 后跟 `- **[critical]** Never email Sarah at sarah@example.com.`,并且该规则在任何 mid-session token 压缩中都能存活。
覆盖率与集成测试位于 `src/openhuman/memory/tool_memory/`
## 另请参阅
- [记忆工具](memory-tools.zh-CN.md)——通用 `recall``store``forget`
- [智能 Token 压缩](../token-compression.zh-CN.md)——系统提示被保护免受的内容。
@@ -1,43 +0,0 @@
---
description: >-
原生语音——语音转文字输入、文字转语音输出、吉祥物口型同步,
以及一个实时 Google Meet 智能体,能听会说。
icon: microphone
---
# 语音
OpenHuman 在你需要时是语音优先的。STT、TTS 和实时 Google Meet 智能体是核心的一部分,而非第三方插件。
## 语音转文字
* **热键**——按键说话和切换模式。
* **音频捕获**——跨平台麦克风捕获,带语音活动检测。
* **流式转录**——你说话时词语即时出现。
* **幻觉过滤器**——剥离已知人工产物("感谢观看"、静默诱导短语)。
* **后处理**——标点、大写、听写清理。
听写可以替换你桌面上活动的文本输入,或直接发送到与智能体的聊天中。
## 文字转语音
回复语音通过托管 TTS 模型路由。智能体的回复可以用你选择的嗓音说出来,带自然的时机和韵律。语音选择可按用户配置,吉祥物头像通过 viseme 贴图与音频流口型同步。
## 实时 Google Meet 智能体
OpenHuman 的旗舰语音集成:
* 通过嵌入式 webview 加入 Google Meet。
* 实时流式输出音频到 STT,转录通话中的每个人,并在会议进行时将结构化笔记写入[记忆树](../obsidian-wiki/memory-tree.zh-CN.md)。
* 当你让它说话(或它觉得有需要补充的有用内容时),它通过 TTS 模型生成音频并**作为出站摄像头/麦克风流播放回会议**,这样其他参与者真的能听到它。
## 隐私
* 音频捕获是本地的。流式 STT 通过 OpenHuman 后端;除实时转录外不保留任何录音。
* TTS 音频流式传输后丢弃——不存储。
* 会议转录内容会像其他来源一样落入你的本地记忆树中。
## 另见
* [记忆树](../obsidian-wiki/memory-tree.zh-CN.md) —— Meet 转录和笔记存放的地方。
* [自动模型路由](../model-routing/) —— Meet 的大脑使用 `hint:fast` 实现低延迟对话轮次。
@@ -1,31 +0,0 @@
---
description: 一个专门的"获取并阅读"工具,返回干净的文本而非原始 HTML。
icon: globe
---
# 网页抓取
一个专门构建的获取工具,区别于通用的 `http_request` / `curl`。它的存在是因为智能体不需要原始 HTML——它需要的是*文章*。
## 功能
* 获取一个 URL。
* 剥离 Boilerplate(导航、广告、页脚、脚本)。
* 返回智能体可以推理的干净文本。
## 护栏
* 响应上限 1 MB——大页面被截断,而非静默丢弃。
* 20 秒超时——慢速服务器不会阻塞对话。
* 遵守与其他网络工具相同的代理和 URL 防护规则。
## 适用于
* 阅读文章、博客文章、文档页面、GitHub README,去除噪音。
* 跟进[网络搜索](web-search.zh-CN.md)的结果。
* 按需摘要单个页面。
## 另见
* [网络搜索](web-search.zh-CN.md) —— 找到要输入抓取器的 URL。
* [智能 Token 压缩](../token-compression.zh-CN.md) —— 在长页面到达模型之前对其进行修剪。
@@ -1,23 +0,0 @@
---
description: 智能体可直接调用的原生搜索工具——无需 API key。
icon: magnifying-glass
---
# 网络搜索
智能体可以自行搜索实时网页。由服务器端代理(Parallel)支持,所以你无需携带搜索 API key,该工具返回标题、摘要片段和 URL,供后续跟进。
## 适用于
* 研究——"X 的最新动态是什么"。
* 引用追踪——"为我找到 Y 的三个来源"。
* 回答前的事实核查——如果智能体不够自信,会快速搜索。
## 与通用 HTTP 的区别
一个纯粹的 `http_request` 工具可以获取 URL 但无法*找到* URL。网络搜索是发现层:它为智能体挑选正确的 URL,然后交给[网页抓取](web-scraper.zh-CN.md)进行实际阅读。
## 另见
* [网页抓取](web-scraper.zh-CN.md) —— 获取并清理特定 URL。
* [智能 Token 压缩](../token-compression.zh-CN.md) —— 搜索摘要片段在进入模型之前被压缩。
@@ -0,0 +1,98 @@
---
description: >-
The notification center, the Activity transparency hub, and the Routines
scheduler — everything OpenHuman tells you about, and everything it does in
the background.
icon: bell
---
# Notifications & Activity
OpenHuman surfaces two kinds of "what's happening" in one place: **notifications** (things you should look at — an important Slack message, a failed webhook, a high-priority email) and **activity** (a transparent ledger of what the agent did on its own while you weren't watching). This page covers the notification center, the Activity hub that fronts it, and the Routines screen for managing scheduled automations.
***
## Notification Center
The notification center is fed by two independent streams that render side by side under **Activity → Alerts**.
### Integration notifications
Notifications captured from connected accounts (Gmail, Slack, WhatsApp, Discord, …) are ingested through the `notification.ingest` RPC, persisted to a per-workspace SQLite store, and then **triaged by a local LLM in the background**. Ingest returns immediately; triage runs in a spawned task and back-fills the score a moment later, so a freshly arrived item can briefly show as unscored.
Triage assigns each notification an **action**, which maps to a fixed 0.01.0 importance score:
| Triage action | Score | What it means |
| ------------- | ----- | ------------------------------------------------ |
| `drop` | 0.10 | Noise — not worth surfacing |
| `acknowledge` | 0.35 | Low value, informational |
| `react` | 0.65 | Worth a follow-up |
| `escalate` | 0.90 | High priority — hand to the agent |
Only `react` and `escalate` are considered "routed" actions; `drop` and `acknowledge` stay quiet. Each ingested item carries a one-sentence `triage_reason` justifying the classifier's call, plus a lifecycle status: **unread → read → acted → dismissed**. Duplicate content received within a 60-second window collapses to a single entry.
### System (core-bridge) notifications
The second stream translates selected internal events into compact, user-facing alerts and pushes them over the socket bridge as they happen. These are persisted before broadcast, so anything fired while the app was closed syncs down on the next open. Each carries a **category** and an in-app deep link:
| Source event | Category | Surfaces when |
| ---------------------- | ---------- | ----------------------------------------------- |
| Cron job completed | Agents | Always (success or failure) |
| Webhook processed | System | **Only on failure** — successes are silent |
| Sub-agent finished | Agents | Always |
| Sub-agent failed | Agents | Always |
| Notification triaged | Agents | Only when routed (`escalate`/`react`) |
| API key rejected | System | Always — links to the LLM settings tab |
The category set the notification center understands is **messages, agents, skills, system, meetings, reminders, important**. The Alerts view shows a filter chip row, but only for categories that actually appear in the current feed, plus **Mark all read** and **Clear**. Clicking a notification marks it read and follows its deep link. Some core notifications carry **action buttons** (e.g. a meeting auto-join prompt) and are pinned to the top of the center.
### Per-provider routing & thresholds
Every provider has its own settings (`notification.settings_set`), letting you tune the noise per source:
| Setting | Effect |
| ---------------------- | --------------------------------------------------------------------------- |
| `enabled` | When off, that provider's notifications are not ingested at all |
| `importance_threshold` | Minimum score (0.01.0) to display; `0.0` shows everything |
| `route_to_orchestrator`| When on, high-importance (`react`/`escalate`) items are forwarded to the agent |
Auto-routing re-reads the provider's settings the moment before escalating, so toggling a setting mid-flight takes effect immediately. A notification is only routed to the agent when its score clears the provider threshold **and** `route_to_orchestrator` is enabled.
***
## Activity hub
The Activity surface (`/activity`) is the transparency layer over everything the agent does without you in the loop. It has three tabs:
| Tab | What it shows |
| ----------------------- | --------------------------------------------------------------------------------------------- |
| **Automations** | Workflows the agent runs on your behalf (the workflows panel) |
| **Background Activity** | The subconscious engine: status bar, active tasks, approval cards, and the evaluation ledger |
| **Alerts** | The notification center described above (integration + system streams) |
The **Background Activity** tab embeds the subconscious loop's controls and activity log — its tick interval, mode, a manual **Run Now** trigger, and a chronological feed of every background task evaluation with a colored status dot. That loop is documented in full on the [Subconscious Loop](subconscious.md) page; the Activity hub is just its front door.
Older deep links (`?tab=memory`, `?tab=agents`, `?tab=tasks`, …) now live under Settings → Developer & Diagnostics and fall back to the Automations tab.
***
## Routines
Routines (`/routines`) is the user-facing management UI for scheduled automations — the desktop face of the cron system. Jobs are sorted by next-run time, each rendered as a card showing:
* The schedule, rendered human-readable (e.g. "every day at 9am") from its cron expression.
* The job **type** badge — *agent* (runs a prompt through the agent) or *command*.
* The **next run** time (when enabled) and the **last run status** dot — sage for success, coral for failure, neutral when it has not run yet.
* A toggle to **enable/disable** the routine, a **Run Now** button for manual triggering (it polls until the run lands), and an expandable **run history**.
Routines surface and manage the scheduled jobs; the underlying scheduling engine, cron syntax, and the agent tools for creating jobs programmatically are covered on the [Cron / scheduled tasks](native-tools/cron.md) page. Completed and failed runs also emit Agents-category notifications into the center described above.
***
## See also
* [Subconscious Loop](subconscious.md) — the background engine behind the Background Activity tab.
* [Cron / scheduled tasks](native-tools/cron.md) — the scheduling engine and agent tools behind Routines.
* [Triggers](integrations/triggers.md) — webhooks and inbound events that can raise notifications.
</content>
</invoke>
+16 -2
View File
@@ -5,14 +5,28 @@ description: >-
icon: book-open
---
# Obsidian-Style Memory
# Memory
<figure><img src="../../.gitbook/assets/image (1).png" alt=""><figcaption><p>A preview of the OpenHuman memory in Obsidian. Data from various sources (GMail, Slack, Whatsapp etc..) is organized as a memory tree.</p></figcaption></figure>
OpenHuman's memory is not a black box. The same chunks the agent reasons over are written as plain `.md` files into a vault inside your workspace. You can open it in [Obsidian](https://obsidian.md), browse it, edit it, and link notes by hand, and the agent will see your edits.
OpenHuman's memory is not a black box. The same chunks the agent reasons over are written as plain `.md` files into an Obsidian-compatible vault inside your workspace. You can open it in [Obsidian](https://obsidian.md), browse it, edit it, and link notes by hand, and the agent will see your edits.
The design is directly inspired by [Andrej Karpathy's obsidian-wiki workflow](https://x.com/karpathy/status/2039805659525644595): a personal wiki where every interesting thing in your life ends up as a linkable note.
## In this section
The memory system spans several layers — this page covers the on-disk vault; the rest of the section goes deeper:
| Page | What it covers |
| ---- | -------------- |
| [Memory Tree](memory-tree.md) | The hierarchical summary forest (L0 buffers → summaries → digests) that produces the vault. |
| [Memory Sources & Scoping](sources.md) | The typed registry of connectors that feed memory, and per-agent source allowlisting. |
| [Auto-fetch from Integrations](auto-fetch.md) | The 20-minute sync loop that keeps memory fresh on its own. |
| [Scoring & Ranking](scoring.md) | How chunks are admitted, enriched with entities, and indexed for recall. |
| [Retrieval & Recall](retrieval.md) | The `memory_tree` tool modes the agent uses to read memory back. |
| [Memory Diff (Git-Backed)](memory-diff.md) | A git ledger of how memory changes over time — "what's new since I last looked." |
| [agentmemory backend](agentmemory-backend.md) | Optional shared `agentmemory` store across other coding agents. |
## Where the vault lives
```
@@ -1,53 +0,0 @@
---
description: >-
每个记忆块也作为 Markdown 文件存在于与你 Obsidian 兼容的存储库中,
你可以打开和编辑。灵感来自 Karpathy 的 obsidian-wiki 工作流。
icon: book-open
---
# Obsidian 风格的记忆
<figure><img src="../../.gitbook/assets/image (1).png" alt=""><figcaption><p>OpenHuman 记忆在 Obsidian 中的预览。来自各种来源(GMail、Slack、Whatsapp 等)的数据被组织成一棵记忆树。</p></figcaption></figure>
OpenHuman 的记忆不是一个黑箱。智能体在其上推理的相同块作为普通的 `.md` 文件写入你工作区内的存储库中。你可以在 [Obsidian](https://obsidian.md) 中打开它,浏览、编辑、手动链接笔记,智能体都会看到你的改动。
设计直接灵感来自 [Andrej Karpathy 的 obsidian-wiki 工作流](https://x.com/karpathy/status/2039805659525644595):一个个人 wiki,你生活中每个有趣的事物最终都成为一个可链接的笔记。
## 存储库在哪里
```text
<workspace>/
└── wiki/
├── summaries/ # 自动生成的源 / 主题 / 全局摘要
├── notes/ # 你的手写笔记(自由格式)
└── … # 每个已连接工具包的文件夹
```
`summaries/` 文件夹按层级布局:全局树按日期,源树按源,主题树按实体。每个文件的前置元数据携带来源(源 id、时间范围、作用域),以便智能体可以将任何声明追溯到产生它的块。
## 打开存储库
在桌面 app 中,**记忆**标签页有一个**"在 Obsidian 中查看存储库"**按钮。它使用 `obsidian://open?path=...` 深度链接,所以你需要已安装 Obsidian。
你也可以在任何编辑器中打开该文件夹,它其实就是 Markdown。文件之间的链接使用标准的 `[[wiki-link]]` 语法,因此 Obsidian 的图谱视图、反向链接和标签浏览器开箱即用。
## 手动编辑笔记
`wiki/notes/` 中的任何内容都会被纳入摄取范围。处理 Gmail 和 Slack 的相同流水线会获取你的手写笔记,对它们进行分块、评分,并与其他所有内容一起折叠到主题树和全局树中。
这意味着你可以:
* 将会议笔记放入 `wiki/notes/2026-05-08-board-call.md`,智能体明天就会知道背景。
* 按项目、人物、股票代码维护一个文件,主题树将你的手动笔记视为另一个数据源。
* 批量导入现有 Obsidian 存储库:将 `.md` 文件放入并触发摄入。
## 为什么这很重要
你无法信任你无法读取的记忆。大多数"AI 记忆"系统将状态隐藏在不透明的嵌入中;OpenHuman 的存储库则相反,智能体的记忆**确确实实**就是一个你拥有的 Markdown 文件夹。如果智能体弄错了什么,你可以找到文件,修复它,下一次检索就是正确的。
这也是最干净的导出方式:即使明天不再使用 OpenHuman,你仍然保留一个完整的个人 wiki。
## 另见
* [记忆树](memory-tree.zh-CN.md)。产生存储库的流水线。
* [从集成自动拉取](auto-fetch.zh-CN.md)。存储库如何自行增长。
@@ -1,166 +0,0 @@
---
description: >-
可选的 `Memory` trait 后端,委托给本地运行的 agentmemory REST 服务器,
适用于在 Claude Code、Cursor、Codex、OpenCode 和 OpenHuman 间
自托管 agentmemory 的用户。
icon: database
---
# agentmemory 后端
OpenHuman 默认的 `Memory` trait 后端是 `sqlite`——即 [记忆树](memory-tree.zh-CN.md) 中记录的统一存储。对于已经在本地运行 [agentmemory](https://github.com/rohitg00/agentmemory) 的用户——通常是因为他们希望在 Claude Code、Cursor、Codex、OpenCode 和 OpenHuman 之间共享单一持久化存储——OpenHuman 暴露了一个可选后端,将每个 trait 调用代理到 agentmemory 的 REST 层面。
选择 `backend = "agentmemory"` 会跳过 OpenHuman 的 SQLite + 嵌入器路径。agentmemory 拥有存储、嵌入和检索层。OpenHuman 成为一个精简的 REST 客户端。
## 何时使用
在以下情况下使用 agentmemory 后端:
- 你已经为一个或多个编码智能体运行 `npx -y @agentmemory/agentmemory`,并希望 OpenHuman 共享相同的持久化存储。
- 你希望混合 BM25 + 向量 + 图检索,而无需在 OpenHuman 端配置单独的嵌入器。
- 你偏好 agentmemory 的生命周期(整合、保留评分、自动遗忘、图提取)而不是 OpenHuman 的统一存储。
在以下情况下保持默认的 `sqlite` 后端:
- 你想要完全自包含的单进程操作,无外部守护进程依赖。
- 你依赖 OpenHuman 特定的记忆树功能(分块、密封、摘要树),这些功能在 SQLite 存储之上运行。记忆树流水线不受 trait 后端影响——它在主机的文档存储上操作,正交——但 agentmemory 后端在你已经在其他智能体上标准化使用 agentmemory 时最有价值。
## 快速开始
1. **安装 + 启动 agentmemory**(一个终端):
```bash
npx -y @agentmemory/agentmemory
```
默认为 `http://localhost:3111`REST+ `ws://localhost:49134`(引擎)。首次启动在 `~/.agentmemory/.hmac` 生成 HMAC 密钥并打印一次。
2. **在 `config.toml` 中将 OpenHuman 指向它**
```toml
[memory]
backend = "agentmemory"
# 以下为默认值——仅在覆盖时设置。
# agentmemory_url = "http://localhost:3111"
# agentmemory_secret = "" # HMAC bearer token,可选
# agentmemory_timeout_ms = 5000
```
3. **重启 OpenHuman**。Factory 会跳过 SQLite 路径并记录 `[memory::factory] using agentmemory backend at <url>`。
就这样。现有的 OpenHuman 调用点(`store`、`recall`、`get`、`list`、`forget`、`namespace_summaries`、`count`、`health_check`)保持不变。
## 配置 keys
| 字段 | 默认值 | 用途 |
| --- | --- | --- |
| `agentmemory_url` | `http://localhost:3111` | agentmemory REST 服务器的基础 URL |
| `agentmemory_secret` | 无 | 可选的 HMAC bearer token。作为 `Authorization: Bearer <secret>` 发送 |
| `agentmemory_timeout_ms` | `5000` | 每个请求的 reqwest 超时 |
当 `backend == "agentmemory"` 时,以下现有 `MemoryConfig` 字段被**忽略**——agentmemory 通过 `~/.agentmemory/.env` 管理自己的嵌入堆栈:
- `embedding_provider`
- `embedding_model`
- `embedding_dimensions`
- `sqlite_open_timeout_secs`
在此路径上设置它们是空操作。本地 AI Ollama 健康检查也不在此路径上运行——agentmemory 的守护进程管理自己的嵌入器生命周期。
## 字段映射
OpenHuman 的 `MemoryEntry` ↔ agentmemory 传输行:
| OpenHuman 字段 | agentmemory 字段 | 备注 |
| --- | --- | --- |
| `namespace` | `project` | 空时默认为 `"default"` |
| `key` | `title` | |
| `content` | `content` | |
| `id` | `id` | agentmemory 生成的(`mem_<rand>` |
| `category: Core` | `type: "fact"` | |
| `category: Daily` | `type: "conversation"` | |
| `category: Conversation` | `type: "conversation"` | |
| `category: Custom(s)` | `type: "fact"` + `concepts: [s]` | 自定义标签滚入 concepts 数组以保持可查询性 |
| `session_id` | `sessionIds: [...]` | OpenHuman 暴露单个 idagentmemory 持久化一个数组 |
| `timestamp` | `updatedAt`RFC3339 | 如果 `updatedAt` 缺失则回退到 `createdAt` |
| `score`(仅召回命中) | smart-search `score` | 在 `recall` 响应中填充,`get` / `list` 时为 `None` |
agentmemory 携带额外字段——`concepts`(自动提取)、`files`(路径标签)、`strength`(保留评分)、`version`、`supersedes`(生命周期链)——此后端保留为默认值。它们是 agentmemory 生命周期层的内部字段,不需要通过 OpenHuman 的 trait 进行往返。
## Trait 方法 → 端点
| `Memory` 方法 | agentmemory REST | 备注 |
| --- | --- | --- |
| `store` | `POST /agentmemory/remember` | `{project, title, content, type, concepts, sessionIds}` |
| `recall` | `POST /agentmemory/smart-search` | 混合 BM25 + 向量 + 图 |
| `get` | `POST /agentmemory/smart-search` | + 客户端精确 title 过滤 |
| `list` | `GET /agentmemory/memories?latest=true&project=<ns>` | |
| `forget` | `get(ns, key)` → `POST /agentmemory/forget` | 两步:先解析 id 再 forget |
| `namespace_summaries` | `GET /agentmemory/projects` | 返回 `[{name, count, lastUpdated}]` |
| `count` | `GET /agentmemory/health` | 读取 `memories` 字段 |
| `health_check` | `GET /agentmemory/livez` | |
`RecallOpts.category`、`RecallOpts.session_id` 和 `RecallOpts.min_score` 作为**客户端过滤**应用于 smart-search 响应。agentmemory 的 REST 面今天不将它们作为服务器端过滤器暴露。对于非常大的召回窗口(limit > 100),建议发出更严格的查询字符串以减少服务器端工作,而不是依赖客户端后过滤。
## 安全性
当 `agentmemory_secret` 被设置时,客户端遵守 agentmemory 的 v0.9.12 明文 Bearer 守卫约定:
- **环回主机**`localhost`、`127.0.0.1`、`::1`)上的 `http://` —— 允许。本地开发路径。
- **`https://`** 到任何主机 —— 允许。
- **到非环回主机的明文 HTTP** —— 在构造时发出一次性 stderr 警告。Bearer 在线路上是可观察的。
- **`AGENTMEMORY_REQUIRE_HTTPS=1`**(进程环境,ASCII 大小写不敏感,匹配 `1` 或 `true`)—— 将警告升级为构造时的硬性拒绝。后端启动失败而不是静默泄露 bearer。
生产部署应设置 `AGENTMEMORY_REQUIRE_HTTPS=1`,这样配置错误的 TLS 终结器会明显报错,而不是静默泄露。
明文 bearer guard 镜像了 agentmemory [PR #315](https://github.com/rohitg00/agentmemory/pull/315) 中的集成插件 guard,因此在 Hermes / OpenClaw / pi 上看到过相同警告的操作员会在 OpenHuman 上认出它。
## 故障模式
| 故障 | 后端行为 |
| --- | --- |
| 启动时守护进程不可达 | `from_config` 成功(URL 解析),但首次调用时 `health_check()` 返回 false。Trait 方法向上冒泡 `reqwest` 传输错误 |
| 网络超时 | 按 trait 约定返回 `anyhow::Error`;浮出到调用者 |
| 4xx / 5xx 响应 | 带状态 + body 片段的 `anyhow::Error` |
| Bearer 通过明文非环回(无环境变量) | 一次性 stderr 警告,请求继续 |
| Bearer 通过明文非环回 + `AGENTMEMORY_REQUIRE_HTTPS=1` | 构造时硬性拒绝 |
| 空的 `agentmemory_url` | 构造时硬性拒绝并提示留空以使用默认值 |
| 无效的 URL 语法 | 构造时硬性拒绝并附带解析器错误 |
**不会自动回退到 SQLite。** 如果守护进程在启动时未运行,后端会明显抛出传输错误。操作员在 `config.toml` 中切回 `backend = "sqlite"` 以恢复。理由:静默的 SQLite 回退会隐藏配置错误的守护进程——"私密、简单、可预测"胜过"神奇容忍"。
## 性能说明
后端是一个精简的 REST 代理——每个 trait 调用增加一个 HTTP 往返。实际影响:
- `store` 和 `forget` 是单 RTT。
- `recall`、`get`、`list` 是单 RTT。
- 对未知 key 的 `forget` 是两个 RTT(隐式 `get` 查找 + 一个空操作确认)。调用者可以通过检查先前 `list` 的返回值来短路这个。
- agentmemory 的 REST 默认是 `127.0.0.1` —— 同主机延迟低于一毫秒。通过 HTTPS 终结的管理部署,预期每个 RTT 约 10–30ms。
- 默认每请求超时为 5 秒。如果在 iii 引擎冷启动时看到间歇性超时,增加 `agentmemory_timeout_ms`agentmemory 长时间空闲后的第一次请求延迟可达 3–5 秒,取决于持久化状态。
## 迁移:从 SQLite 到 agentmemory
目前没有原地迁移。建议路径:
1. 通过 OpenHuman 现有的导出 RPC(或直接 SQL)从 SQLite 存储导出你现有的记忆。
2. 遍历导出,将每一行 POST 到 `/agentmemory/remember`,使用相同的 `project` + `title` + `content`。agentmemory 将分配新 idOpenHuman 端在首次 `list` 时获取它们。
3. 设置 `backend = "agentmemory"` 并重启。
专门的批量导入路径作为后续跟进。
## 实现参考
仓库内文件:
- [`store/agentmemory/mod.rs`](https://github.com/tinyhumansai/openhuman/tree/main/src/openhuman/memory/store/agentmemory/mod.rs) —— 模块表面
- [`store/agentmemory/backend.rs`](https://github.com/tinyhumansai/openhuman/tree/main/src/openhuman/memory/store/agentmemory/backend.rs) —— `impl Memory for AgentMemoryBackend`
- [`store/agentmemory/client.rs`](https://github.com/tinyhumansai/openhuman/tree/main/src/openhuman/memory/store/agentmemory/client.rs) —— reqwest 包装器 + 明文 bearer guard
- [`store/agentmemory/mapping.rs`](https://github.com/tinyhumansai/openhuman/tree/main/src/openhuman/memory/store/agentmemory/mapping.rs) —— `MemoryEntry` ↔ agentmemory JSON
- [`tests/agentmemory_backend.rs`](https://github.com/tinyhumansai/openhuman/tree/main/tests/agentmemory_backend.rs) —— 12 个 axum-mock 集成测试
相关的上游:
- agentmemory 仓库 —— <https://github.com/rohitg00/agentmemory>
- agentmemory REST 约定 —— `~/.agentmemory/.env` keys + 端点列表在 agentmemory README 中
- v0.9.12 明文 bearer guard —— agentmemory PR #315
@@ -1,60 +0,0 @@
---
description: >-
每隔二十分钟,OpenHuman 遍历每个活跃集成,将新数据整合进你的记忆树。
无需提示词,无需编写轮询循环。
icon: arrows-rotate
---
# 自动拉取集成
大多数"AI 助手"是被动的:你提问,它们思考,它们回答。OpenHuman 则相反。它持续从你的技术栈中拉取数据,所以当你问"昨晚我的收件箱收到了什么?"时,答案已经在[记忆树](memory-tree.zh-CN.md)里了。
## 工作原理
一个单一的周期性调度器每二十分钟触发一次。每次触发时,它遍历每个活跃的[集成](../integrations/README.zh-CN.md),查找匹配的原生 provider,如果该连接的距上次同步的时间足够长,就调用 `provider.sync(ctx, SyncReason::Periodic)`
```text
每 20 分钟
|
v
遍历每个活跃连接(Gmail、Notion、GitHub……)
|
+--> 检查 sync_statetoolkit, connection_id
| - 上次同步时间戳
| - 每日预算
| - 去重集合
| - 游标
|
+--> 如果间隔已过 -> provider.sync()
|
+--> 成功 -> record_sync_success(ts)
```
这里有几个关键点:
* **一个全局触发,而不是每个连接一个任务。** 每个用户的连接数很少;一个 20 分钟的触发周期足够了,而且 bookkeeping 很简单。
* **状态按 `(toolkit, connection_id)` 划分。** 每个连接有自己的游标、上次同步时间戳、去重集合和每日预算。重启时从中重建;即使重启后错过了一次周期性同步也无害,因为下一个触发周期会重新拾取。
* **原生同步与事件驱动路径共享。** 当 webhook 或 `on_connection_created` 事件触发非周期性同步时,它们在同一个 sync_state 上盖戳,所以调度器不会冗余地重新触发。
* **错误被记录并静默处理。** 调度器绝不能在其循环中 panic,否则周期性同步会在进程剩余生命周期内静默停止。
## 什么进入记忆树
每个 provider 负责定义自己的摄入逻辑。例如 Gmail provider 获取一页新消息,运行邮件规范化器,通过相同的手动 UI 摄入路径传输结果,块进入 SQLite,摘要 bucket 被填充,任何被触及的实体都会将主题树标记为脏。
其他 providersGitHub、Slack、Notion……)遵循相同的形状:从游标后获取新项目 → 规范化 → 摄入到[记忆树](memory-tree.zh-CN.md)。
## 为什么是 20 分钟触发周期
最初设计每 60 秒运行一次。当连接了多个 provider 时,这意味着持续不断的 HTTP 获取和数据库写入,在笔记本上明显繁忙。二十分钟用一点延迟换取明显更少的前台负载。每个 provider 的 `sync_interval_secs` 仍然限制实际同步之间的**最小**延迟;全局触发周期只放宽上限。
## 调优和可见性
* **每个 provider 的间隔。** 每个原生 provider 声明自己的 `sync_interval_secs`,所以高流量工具包(Gmail)可以比低流量工具包(Stripe)更频繁地同步。
* **每日预算。** 每个连接有每日请求预算,以保持 API 成本和速率限制合理。
* **日志。** 同步活动以 debug 级别记录在 core 日志中。
## 另见
* [第三方集成](../integrations/README.zh-CN.md)。自动拉取运行的连接器层。
* [记忆树](memory-tree.zh-CN.md)。一切最终到达的地方。
* [智能 Token 压缩](../token-compression.zh-CN.md)。使"获取一切"保持低成本的原因。
@@ -0,0 +1,121 @@
---
description: >-
Git-backed change tracking for memory. Every sync is committed to a snapshot
ledger, so the agent can ask "what changed since I last looked?"
icon: git-compare
---
# Memory Diff
The [Memory Tree](memory-tree.md) tells the agent what it knows. **Memory Diff** tells it what _changed_. It is a derived ledger that records the state of every memory source over time, so any agent (or you) can ask: what's new, what was edited, what disappeared - since the last sync, since I last read it, or since a named baseline.
The chunk store (`mem_tree_chunks`) stays authoritative. The diff ledger is a read-only view built _from_ already-ingested data - so snapshots cost zero API calls. Source: `src/openhuman/memory_diff/`.
***
## It's a git repository
The whole thing is a real [libgit2](https://libgit2.org/) repository living at `<workspace>/memory_diff/repo` (`git_store.rs`). Rather than invent a snapshot format, OpenHuman maps memory-change tracking straight onto git's native primitives:
| Memory-diff concept | Git primitive |
| ------------------- | --------------------------------------------------- |
| Snapshot | Commit (the snapshot id **is** the commit SHA) |
| Item | One flat blob, named by the (encoded) item id |
| Source | A subtree under `<source_id>/` in the root tree |
| Checkpoint | Annotated tag `ckpt_<uuid>` at HEAD |
| Read marker | Ref `refs/openhuman/read/<source_id>` → commit SHA |
| Diff | A git tree-to-tree diff scoped to one source's path |
Snapshot metadata that has no natural git home - source kind, label, trigger (`auto` / `manual`), item count, millisecond timestamp - rides along in the **commit message as trailers** (`Source-Id:`, `Trigger:`, `Item-Count:`, `Taken-At-Ms:`, …) and is parsed back out on read.
***
## The snapshot model
After each successful sync, `auto_snapshot_after_sync()` reads the current chunks for that one source out of `mem_tree_chunks`, groups them into one blob per item, and commits them under `<source_id>/`. Crucially, **every other source is carried forward** from the parent commit - so each commit's tree reflects the whole world, even though only one source actually changed.
```
mem_tree_chunks (authoritative)
|
| sync finishes for source A
v
take_snapshot(A) items grouped, one blob per item
|
v
┌──────────────────────────────────────────────┐
│ commit_snapshot │
│ │
│ root tree = parent tree │
│ ├── src_A/ ← rebuilt from new items │
│ ├── src_B/ ← carried forward unchanged │
│ └── src_C/ ← carried forward unchanged │
│ │
│ message trailers: Source-Id, Trigger, … │
└───────────────────────┬───────────────────────┘
v
HEAD ─► commit (= snapshot id / SHA)
```
A diff for source A is then just `git diff <from-tree>..<to-tree>` with the pathspec pinned to `src_A/`. Added / Removed / Modified fall straight out of git's delta status; **Unchanged** is computed as `to_item_count - added - modified`. Item identity is the blob name, so editing an item's content keeps the name (→ `Modified`) while changing its id is `Removed` + `Added`.
All writes serialise through a process-global lock, because git's HEAD/parent bookkeeping is read-modify-write and concurrent commits could otherwise fork history.
***
## What the agent uses it for
The headline use case is **"what changed since I last looked?"** During a conversation the agent calls the `memory_diff` tool (`tools.rs`). Its parameters:
| Param | Effect |
| -------------------- | ----------------------------------------------------------------------------------------------- |
| _(none)_ | Lists enabled sources with their snapshot counts. |
| `source_id` | Diffs one source. |
| `checkpoint_id` | Cross-source diff: everything that changed since that named checkpoint. |
| `since_read` | When diffing a source, show changes since the **read marker** rather than since the last sync. Default `true`. |
| `commit` | When using `since_read`, advance the read marker to head after reading. Default `true`; set `false` to preview without acknowledging. |
| `include_text_diff` | Include line-level unified diffs for modified items (truncated to ~2000 chars). Default `false`. |
The read-marker mechanic is the turn-to-turn primitive (`diff_since_read` in `ops.rs`): the first call returns the full delta and moves `refs/openhuman/read/<source_id>` up to the current head; the next call returns _only_ what arrived since. So an agent that polls a source repeatedly never re-reads the same news twice. The tool is `ReadOnly` with respect to your data - the only write it performs is advancing that internal marker, never anything in `action_dir`.
Output is concise markdown, e.g.:
```
## Memory Changes (Inbox)
**2 added, 1 modified, 0 removed** (47 unchanged)
### Added
- Invoice #4021 from Acme
- Re: Q3 planning
### Modified
- Standup notes
```
***
## Checkpoints and cross-source diffs
A **checkpoint** is a named baseline across _all_ enabled sources - an annotated git tag at HEAD that records the latest snapshot id per source (`create_checkpoint`). It will even take a fresh snapshot for any source lacking one, so the baseline is complete. Later, `diff_since_checkpoint` walks each recorded snapshot to its current head and aggregates per-source changes into one `CrossSourceDiff` - "everything that's happened across my whole memory since this morning's baseline."
Checkpoints are cheap to prune: `cleanup` deletes tags older than N days, but **snapshot commits are never deleted** - git history _is_ the ledger, and git's delta compression keeps it compact.
***
## Why this matters to you
Because the ledger is real git history, Memory Diff gives the agent's knowledge an **audit trail**:
* **Traceability.** Every change to what the agent knows is a commit with a timestamp, a trigger (`auto` vs `manual`), and an item count.
* **No surprises.** The agent acts on _deltas_, not the whole world each turn - so a single new email gets noticed without re-reading your entire inbox.
* **Recoverable history.** Snapshots are kept indefinitely; you can always reconstruct what a source looked like at any past point.
* **Cheap.** It's built from data already ingested by the Memory Tree, so tracking change costs no extra model or API calls.
***
## See also
* [Memory Tree](memory-tree.md) - the authoritative knowledge base that snapshots are derived from.
* [Auto-fetch from Integrations](auto-fetch.md) - what triggers the syncs that produce new snapshots.
* [Obsidian Wiki](README.md) - the Markdown vault these sources ingest into.
* [Subconscious Loop](../subconscious.md) - the background loop that reviews new memory changes for actionable items.
@@ -1,172 +0,0 @@
---
description: >-
OpenHuman 的本地优先存储库。从工具中摄入数据,规范化为 Markdown,
分块,评分,并折叠为层级化的摘要树。
icon: tree
---
# 记忆树
<figure><img src="../../.gitbook/assets/image.png" alt=""><figcaption><p>记忆树。所有文档的高度压缩视图。</p></figcaption></figure>
记忆树是 OpenHuman 的存储库。它不是一个披着"记忆"外衣的向量数据库,而是一个确定性的、bucket-sealed(桶密封)处理流水线,将你一天中杂乱的数据流——聊天、邮件、文档、集成同步结果——转化为你机器上结构化的、可查询的、带摘要支撑的 Markdown。
## 它做什么
每个连接的源都走同样的流水线:
```text
源适配器(聊天 / 邮件 / 文档)
|
v
规范化 规范化的 Markdown + 来源元数据
|
v
分块器 确定性的 ID,≤3k token 的有界片段
|
v
内容存储 原子 .md 文件(正文 + 标签)
|
v
存储 持久化(块、评分、摘要、任务、热度)
|
v
评分 信号 + 向量 + 实体提取
|
v
源 / 主题 / 全局树 按作用域的摘要树
|
v
检索 搜索 / 深入 / 主题 / 全局 / 获取
```
热路径(规范化 → 分块 → 快速评分 → 持久化 → 入队后续工作)很快。重型工作——向量生成、实体提取、密封摘要 bucket、每日摘要——在后台 workers 中运行,UI 永远不会阻塞。
如果你开启了[本地 AI](../model-routing/local-ai.zh-CN.md),嵌入向量和摘要树的构建可以在**设备上通过 Ollama** 运行;否则它们像其他模型调用一样通过 OpenHuman 后端处理。
## 三棵树,三个作用域
* **源树**,每个源一个滚动缓冲区(L0),填满后密封为 L1 → L2 → …。每个 Gmail 标签、每个 Slack 频道、每个上传的文档各一棵。
* **主题树**,按实体懒加载的摘要,由**热度**驱动。某个实体(人、项目、股票代码、仓库)出现得越频繁,其主题树就越积极地被构建和刷新。
* **全局树**,一个跨当天摄入的所有内容的每日全局摘要。
检索可以针对任何作用域:搜索单个源,深入某个主题,或拉取全局摘要。
## 它在磁盘上的位置
位于你的工作区内(默认 `~/.openhuman`,或 `OPENHUMAN_WORKSPACE` 指向的路径):
| 路径 | 内容 |
| ------------------------- | ---------------------------------------------- |
| `memory_tree/chunks.db` | 块、评分、摘要、实体索引、任务、热度 |
| `wiki/` | Markdown 存储库 —— 见 [Obsidian Wiki](./README.zh-CN.md) |
一切都是本地的。除非你明确发送包含原始数据的聊天消息,否则你的原始数据不会离开你的机器。
## 为什么是树,而不是向量存储
向量存储回答"与这个查询相似的是什么?"记忆需要回答更多:
* **今天发生了什么?**(全局摘要)
* **这个人的最新情况是什么?**(主题树,热度驱动)
* **上周二下午 3 点 Stripe webhook 说了什么?**(源树 + 来源追溯)
树给你压缩**和**导航。嵌入向量仍然存在于内部,所以语义搜索继续工作,但上面的结构才是让记忆感觉像大脑而不是一堆碎片的原因。
## 流水线如何工作?
用户看到的功能很简单:连接一个源,智能体就获得了对其的持久记忆。实现这一功能的流水线横跨一条 HTTP 触发的摄入路径、一个持久化的任务队列、一组后台 workers、三个独立的摘要树,以及一个每日 UTC 调度器。
### 1. 摄入
新的聊天 / 邮件 / 文档到达。热路径将其规范化为 Markdown,用确定性 ID 分块,运行廉价的快速评分,在单个事务中持久化所有内容,将每个块标记为 `pending_extraction`,并为 workers 入队后续工作。
这里有三个重要属性:
* **确定性的。** 块 ID 是内容寻址的,所以对相同输入重新运行摄入永远不会产生重复。
* **快速的。** 这条路径中没有 LLM 调用——只有廉价的启发式方法。
* **写入有界。** 所有操作在一个事务中完成,所以部分摄入不会留下悬空的行。
### 2. 队列
后续工作进入持久化的任务队列(与块在同一个磁盘存储中)。每个任务携带一种类型、一个 payload、一个去重 key、重试记录和一个调度窗口。类型如下:
| 类型 | 功能 |
| --------------- | -------------------------------------------------------------------------------------- |
| `extract_chunk` | 深度评分 + 实体提取。决定 `admitted` 还是 `dropped`。 |
| `append_buffer` | 将一个 admitted 的叶子添加到源的(或主题的)树的 L0 缓冲区。可能触发密封。 |
| `seal` | 将 L0 缓冲区压缩为 L1 摘要;如果父缓冲区已满,则向上级联。 |
| `topic_route` | 将叶子路由到每个实体的主题树,由热度检查控制。 |
| `digest_daily` | 构建全局每日摘要节点。 |
| `flush_stale` | 强制密封停留太久的缓冲区。 |
### 3. Workers
一个小型的后台 workers 池(默认 3 个)从队列中取出任务并运行。池被摄入路径立即唤醒,有一个短轮询后备方案,所以错过的唤醒不会搁置工作。共享信号量限制并发 LLM 调用,这样新源的突发不会意外地扇出到数十个并发嵌入向量。
启动时,任何 worker 租约已过期的任务(因为崩溃或 kill)会被返还到队列。崩溃不会丢失已 admitted 但尚未密封的工作。
### 4. 树状态
三棵独立的树从同一个叶子流构建。
* **源树** —— 每个源一个。新叶子进入 L0 缓冲区;当缓冲区填满(或 stale-flush 触发),一个 `seal` 写入 L1 摘要,级联继续向上。
* **主题树** —— 每个高热度实体一个。路由器检查实体是否足够热以值得拥有自己的树,如果是,则追加到其缓冲区。
* **全局树** —— 一棵树,每天增长一个节点,随着天数累积向上行走。
### 5. 调度器
调度器循环独立于摄入路径运行。每天 00:00 UTC 它为昨天入队一个全局每日摘要,并为今天入队一个 stale-flush。调度器**不自己运行**摘要器——一切通过队列,所以重试、去重和 stale-lock 恢复保持集中。
### 6. 叶子生命周期
每个块经历一个小型状态机:
```text
pending_extraction --> admitted --> buffered --> sealed
\
--> dropped
```
* 提取根据深度评分决定 `admitted` 还是 `dropped`
* Admitted 的叶子移入缓冲区(`buffered`)。
* 当缓冲区密封时,里面的每个叶子被标记为 `sealed`
* `dropped` 的叶子停在这里。它们的块行保留用于来源追溯,但没有缓冲区或摘要引用它们。
这就是为什么检索可以显示来源追溯而无需重新运行流水线:块行及其终端生命周期状态就够了。
## 触发摄入
* **自动的** —— 每个活跃的集成每 20 分钟自动拉取一次;见 [自动拉取](auto-fetch.zh-CN.md)。
* **手动的** —— 桌面 app 的"记忆"标签页暴露了每个源的"运行摄入"触发器。
* **RPC** —— `openhuman.memory_tree_ingest`,用于高级工作流。
## 在桌面 app 中 —— 智能标签页
从底部导航栏打开。
**系统状态。** 页面顶部显示当前状态(空闲、摄入中、摘要中)和一个**运行摄入**按钮,用于手动触发对任何连接源的同步。
**记忆指标:**
| 指标 | 显示内容 |
| ---------------------- | ------------------------------------------------------------------------------------------ |
| **存储** | `<workspace>/memory_tree/chunks.db` 和 Obsidian 存储库总大小。 |
| **源** | 已摄入的不同源数量(每个 Gmail 标签、Slack 频道、文档等各算一个)。 |
| **块** | 存储中 ≤3k token 的块总数。 |
| **主题** | 目前已实例化的主题树数量(从"热"实体构建的每个实体摘要)。 |
| **最早 / 最新记忆** | 最旧和最新块的时间戳。 |
**记忆图谱。** 一个实体及其关系的力导向可视化,从实体索引绘制。图谱随着自动拉取获取更多数据而增长——早期稀疏,几天内变得密集。
**Obsidian 存储库。** 一个**"在 Obsidian 中查看存储库"**按钮通过 `obsidian://open?path=...` 深度链接直接打开 `<workspace>/wiki/`。你也可以在任何文件浏览器中打开该文件夹。
**摄入活动。** 一个显示摄入事件随时间分布的热力图,类似于 GitHub 的贡献图。可用于发现自动拉取空闲的时期(例如连接中断导致同步停止)。
**搜索与检索。** 记忆树上的搜索栏。支持源作用域、主题作用域或全局查询,任何结果都可以链接回底层块文件(在你的 Obsidian 存储库中)以获取完整来源追溯。
**路由。** 智能标签页还显示智能体每个任务使用的模型——见[自动模型路由](../model-routing/README.zh-CN.md)。
## 交换后端
记忆树流水线(分块 → 评分 → 密封 → 摘要)是默认的。在多个智能体间自托管 [agentmemory](https://github.com/rohitg00/agentmemory) 且希望 OpenHuman 共享相同持久化存储的操作员可以通过 `MemoryConfig.backend = "agentmemory"` 选择外部后端——参见 [agentmemory 后端](agentmemory-backend.zh-CN.md) 了解配置 keys、字段映射、端点表、安全措施和故障模式。
@@ -0,0 +1,112 @@
---
description: >-
How OpenHuman reads back from the memory tree. A handful of deterministic
retrieval primitives, canonical entity resolution, a co-occurrence graph, and
a specialist memory sub-agent that combines them.
icon: search
---
# Memory Retrieval & Recall
The [Memory Tree](memory-tree.md) is the write path: it folds the stream of your day into chunks, scores, and hierarchical summary trees on disk. **Retrieval** is the read path - how the agent finds the right node, hydrates the right raw chunk, and resolves "Alice" to a stable id before answering you.
There is deliberately **no classifier, gate, or composer** in the retrieval layer. The primitives are deterministic and scope-specific; deciding *which* primitive to call and *how* to combine results is left to the calling agent (or, for the deterministic `walk`, to a pure routing algorithm). Source: `src/openhuman/memory_tree/retrieval/mod.rs`.
***
## The `memory_tree` tool: one mode dispatcher
The agent-facing surface is a single multi-mode tool named `memory_tree` (`src/openhuman/memory/query/mod.rs`). Its `mode` field routes to one underlying implementation. Every retrieval mode returns the same `RetrievalHit` shape so the model sees a uniform schema regardless of which mode ran.
| Mode | What it's for | Typical use |
| --- | --- | --- |
| `search_entities` | Fuzzy `LIKE` lookup over the canonical entity index; resolves a surface name to a canonical id. | Call **first** when the user mentions someone by name ("what did Alice say?"). |
| `query_source` | Per-source summary retrieval filtered by source kind + time window, with optional semantic rerank. | "Summarise my Slack #eng from last week." |
| `drill_down` | BFS walk of a summary node's `child_ids`, one or more levels down, with optional rerank. | Expand a coarse summary into its finer-grained children. |
| `cover_window` | Minimum set of nodes covering a `[since_ms, until_ms]` time window. | "Last 24h" recaps and other time-bounded catch-ups. |
| `fetch_leaves` | Batch hydration of raw leaf chunks by id (cap 20). | Pull exact source text for citation after a summary hit. |
| `ingest_document` | Write a document into the tree for future retrieval (the one **write** mode). | Persist a fetched web page / GitHub file; re-ingesting the same `source_id` replaces old chunks. |
| `walk` / `smart_walk` | Deterministic E2GraphRAG retrieval - extracts query entities, routes between the entity graph and dense summaries with **no LLM**, returns ranked evidence. | Answer a natural-language question in one shot without an agent loop. |
The historical `query_global` and `query_topic` modes were **removed**: source trees hold all the content, and walking the source hierarchy plus the entity index reconstructs both the time and topic projections (`mod.rs` dispatcher tests assert their absence).
***
## The `RetrievalHit` shape
Every primitive emits `RetrievalHit` (`src/openhuman/memory_tree/retrieval/types.rs`). The important fields:
- `node_id`, `node_kind` - `leaf` (a raw `mem_tree_chunks` row) or `summary` (a sealed `mem_tree_summaries` row). Consumers branch on this (e.g. "only `drill_down` on summaries").
- `tree_id` / `tree_kind` / `tree_scope` / `level` - provenance, so a UI can say "from Slack #eng".
- `content` - the snippet (summary text or raw chunk body).
- `entities` / `topics` - canonical ids and tags carried on the node.
- `time_range_start` / `time_range_end` - RFC3339, so hits from different tools sort on a common axis.
- `score` - relevance.
- `child_ids` - next level down (empty on leaves); the cursor for `drill_down`.
- `source_ref` - back-pointer to the original source (populated on leaves).
Query-style modes wrap hits in a `QueryResponse { hits, total, truncated }` where `total` is the pre-truncation match count, so the agent can tell whether a higher-limit follow-up would surface more.
***
## Entity resolution and canonical ids
Names are messy; ids are not. Before answering a question about a person, the agent resolves the surface form to a **canonical id** like `person:alice` or `email:alice@example.com`.
- `search_entities` does the fuzzy lookup over the entity index that the tree summariser maintains.
- The canonical registry lives in `src/openhuman/memory_entities/` - one Markdown file per entity at `<content_root>/entities/<kind>/<canonical_id>.md`, with YAML frontmatter (`id`, `kind`, `display_name`, `aliases`, `emails`, `handles`) plus a free-form notes body the user can edit in Obsidian. `lookup_alias` matches by alias / email / handle / display name, case-insensitively.
- `kind` matches `memory_tree::score::extract::EntityKind`, so the ids the scorer emits round-trip through the registry unchanged. The vault is the source of truth - Obsidian, grep, and vector search all see the same data without a separate store.
***
## The entity graph (read-only, derived)
`src/openhuman/memory_graph/` exposes entity relationships **without** a parallel triple-store table. The premise: *the graph is the tree mapped out*. Two entities that co-occur on the same tree node form an edge; the weight is the count of distinct shared nodes.
- `co_occurring_entities(config, subject, limit)` - `GraphEdge { subject, object, weight }` sorted by weight.
- `neighbors(config, subject, limit)` - neighbour ids only.
It is a pure read-only SELF-JOIN over `mem_tree_entity_index` - no new tables, no new schema. This graph is exactly what powers the deterministic `walk` routing below.
***
## Deterministic walk (`walk` / `smart_walk`, no LLM)
`walk` and `smart_walk` both route through `fast_retrieve` (`src/openhuman/memory_tree/retrieval/fast.rs`), an **E2GraphRAG-style** algorithm that replaces the old agentic turn-by-turn loops. It never invokes an LLM. Routing is decided purely by query entities and co-occurrence-graph hop distance:
1. Extract query entities `Eq` (spaCy NLP, regex fallback).
2. `Eq` empty -> **global**: dense rerank over the summary tree.
3. Otherwise compute related entity pairs within `h` hops:
- none related -> **global with occurrence ranking**: dense top-2k, re-ranked by how many `Eq` entities each summary mentions.
- related pairs found -> **local**: intersect the entity-index node sets of each pair, tightening `h` while candidates exceed `k`, then rank survivors by entity coverage and recency.
Tunables (`FastRetrieveOptions`): `limit` (`k`, default 10, cap 100), `max_hops` (`h`, default 2, cap 4), and an optional `time_window_days` look-back on the dense branch. Output is a structured `QueryResponse` of hits - no synthesised prose - for a higher-level context agent to consume.
***
## Time-windowed recall
For "what happened in the last 24h" style questions, `cover_window` computes the **minimum set of nodes** that covers `[since_ms, until_ms]` (epoch-millis). Because summary nodes carry `time_range_start` / `time_range_end`, a single high-level summary can cover a whole window without fanning out to every leaf - the agent only drills down or fetches leaves when it needs detail or a citation.
***
## `memory_recall` - legacy key-value search
Distinct from the tree, `memory_recall` (`src/openhuman/memory/tools/recall.rs`) searches the older namespaced key-value memory: `memory_recall { namespace, query, limit }` over namespaces like `global`, `background`, `autocomplete`, or `skill-{id}`. It returns scored results and is best for exact preference / fact lookups ("does the user prefer dark mode?") that predate the tree.
***
## The memory agent (specialist sub-agent)
`src/openhuman/agent_memory/` owns a specialist retrieval sub-agent invoked via the `call_memory_agent` tool. It navigates the memory tree to answer a question by combining strategies the primitives expose: vector search, keyword search over raw files, entity search and relationship following, hierarchical tree browse, direct content reads, and source listing.
Its tool allowlist (`src/openhuman/agent_memory/agent/agent.toml`) is the full retrieval surface: `memory_tree` (with all the modes above, including deterministic `walk` / `smart_walk`), `memory_recall`, and `query_memory`. The prompt and iteration cap live alongside in `agent/prompt.md` + `agent/prompt.rs`; performance is tracked by the benchmark harness in `ops.rs` (`scripts/bench-memory-walk.sh`).
***
## See also
- [memory-tree.md](memory-tree.md) - the write path that builds the trees retrieval reads.
- [memory-diff.md](memory-diff.md) - how memory changes are tracked over time.
- [README.md](README.md) - feature index for the Obsidian-backed wiki.
- [../subconscious.md](../subconscious.md) - the background loop that consumes recalled context.
+118
View File
@@ -0,0 +1,118 @@
---
description: >-
How OpenHuman scores every chunk before it enters the Memory Tree - weighted
signals gate admission, entity extraction enriches it, and an inverted index
plus vector embeddings make it recallable.
icon: scale
---
# Memory Scoring & Ranking
Not every chunk deserves a place in the Memory Tree. A "thanks!" reply, an email footer, or a calendar auto-notification carry almost no signal, and folding them into summary trees only dilutes the result and burns LLM tokens. Scoring is the gate: a per-chunk pass that runs **after chunking and before the chunk is appended to the L0 buffer**, deciding whether the chunk is worth keeping, enriching it with extracted entities, and indexing it for retrieval.
The entry point is `score_chunk` in [`src/openhuman/memory_tree/score/mod.rs`](../../../src/openhuman/memory_tree/score/mod.rs). It is a pure function - it computes a result but does not touch the store; callers persist based on `ScoreResult::kept`.
***
## Why scoring exists
Two goals, both in service of a dense, relevant tree:
* **Keep the tree signal-rich.** Trees compress _and_ navigate. Admitting noise makes both worse - summaries get vaguer and retrieval surfaces junk.
* **Control cost.** The deep path can call an LLM for entity extraction and importance rating. Scoring is structured so that obvious keeps and obvious drops never pay that cost - only genuinely borderline chunks consult the model.
***
## The signals
`score_chunk` computes a bag of independent signals, each normalised to `[0.0, 1.0]`, defined in [`score/signals/`](../../../src/openhuman/memory_tree/score/signals/). They are combined into a single weighted total and stored alongside it in `mem_tree_score` so every admit/drop decision stays auditable.
| Signal | What it measures | Default weight |
| ----------------- | ------------------------------------------------------------------------------------------------- | -------------- |
| `token_count` | Plateau over chunk size: 0 below `TOKEN_MIN` (10), ramps to 1 by `TOKEN_RAMP_LOW` (30), eases back to 0.5 toward `TOKEN_MAX` (8000). | 1.0 |
| `unique_words` | Type-token ratio noise detector; low lexical diversity scores low, very short messages return a neutral 0.5. | 1.0 |
| `metadata_weight` | Base weight per `SourceKind` (Email > Document > Chat). | 1.5 |
| `source_weight` | Per-`DataSource` weight inferred from `provider:<name>` tags, with `SourceKind` defaults. | 1.5 |
| `interaction` | Engagement-tag bonus (`sent`, `reply`, `dm`, `mention`); absent tags return 0.5 so silent content isn't penalised. | 3.0 |
| `entity_density` | Distinct entities per token, capped at ~1 entity / 100 tokens. More entities → more substantive. | 1.0 |
| `llm_importance` | LLM-derived importance rating in `[0.0, 1.0]`. Off by default; weight `2.0` once an LLM extractor is wired in. | 0.0 |
`interaction` is deliberately the strongest signal - direct user engagement is the clearest proxy for "this mattered to a human." Weights live in `SignalWeights` ([`signals/types.rs`](../../../src/openhuman/memory_tree/score/signals/types.rs)); `combine` / `combine_cheap_only` in [`signals/ops.rs`](../../../src/openhuman/memory_tree/score/signals/ops.rs) produce the normalised total (the cheap variant excludes the `llm_importance` term).
***
## The admission gate
```
chunk --> regex extract --> cheap signals --> combine_cheap_only
|
+--------------------------------------+--------------------------------------+
| | |
total >= DEFINITE_KEEP (0.85) DEFINITE_DROP < total < KEEP total <= DEFINITE_DROP (0.15)
admit, skip LLM borderline -> LLM extract, drop, skip LLM
merge, recompute, recombine
| | |
+--------------------------------------+--------------------------------------+
v
final total >= DROP_THRESHOLD (0.3) ? --> admit / drop
v
extract entities --> canonicalise --> index + embed
```
The three band constants are defined in `mod.rs`:
* `DEFAULT_DEFINITE_KEEP = 0.85` - cheap total at or above this is admitted without the LLM.
* `DEFAULT_DEFINITE_DROP = 0.15` - cheap total at or below this is dropped without the LLM.
* `DEFAULT_DROP_THRESHOLD = 0.3` - the final admission cutoff applied after any LLM augmentation.
Only chunks whose cheap total lands **strictly between** the two definite bands pay for an LLM call - that is where the importance signal is most informative. A short-circuit on either side skips it.
Two refinements sit on top. Chunks tagged `priority_high` at ingest (GitHub commit messages, closed/merged issues and PRs) get a `PRIORITY_BOOST` of `+0.25` (clamped to 1.0) so high-signal source material clears the gate and ranks higher. And a guard drops "tiny, entity-free" chatter - content under `TOKEN_MIN` tokens with no extracted entities - so it can't squeak through on metadata priors alone (priority-tagged chunks bypass this guard).
Dropped chunks still get a score row written for diagnostics, with a `drop_reason`; their chunk row survives for provenance, but no buffer or summary references them.
***
## Entity extraction
Extraction enriches a chunk and feeds both the `entity_density` / `llm_importance` signals and the index. It is pluggable via the `EntityExtractor` trait in [`score/extract/`](../../../src/openhuman/memory_tree/score/extract/), and runs in two stages:
* **`RegexEntityExtractor`** - always on, deterministic, cheap. Once-compiled patterns pull mechanical identifiers: email, URL, handle (`@alice` and Discord-style `alice#1234`), and hashtag. UTF-8 safe (spans are char offsets).
* **`LlmEntityExtractor`** - consulted only on borderline chunks. A single structured-JSON call asks the model for semantic NER (Person / Organization / Location / Topic / …) plus an importance rating, with span recovery and a soft warn-and-empty fallback on transport failure.
The two are chained by **`CompositeExtractor`**, which runs a sequence of extractors and tolerates per-extractor failures. Outputs are merged (`ExtractedEntities::merge` deduplicates entities and takes the max importance), then **canonicalised** by [`resolver.rs`](../../../src/openhuman/memory_tree/score/resolver.rs) - lowercasing emails, stripping leading `@`/`#`, and assigning stable `canonical_id` strings - so the same person or topic resolves to one identity across chunks.
***
## The entity index & graph
Canonical entities for each kept chunk are written to **`mem_tree_entity_index`**, an inverted index mapping `entity_id → node_id` ([`store.rs`](../../../src/openhuman/memory_tree/score/store.rs)). This is the connective tissue the rest of the Memory Tree reads from:
* **Retrieval** resolves a query's entities against the index to find candidate nodes.
* **Topic routing** uses entity hotness to decide which entities deserve their own topic tree.
* **The Memory graph** (the force-directed visualization in the Intelligence tab) is drawn from co-occurrence edges - two entities mentioned in the same chunk get an undirected edge (`graph::pairs_from_entities`), written in the same transaction as the index so the two never diverge.
***
## Embeddings for semantic recall
Scoring also produces vectors. The embedder in [`score/embed/`](../../../src/openhuman/memory_tree/score/embed/) turns each chunk (and later, summary) into a fixed `EMBEDDING_DIM = 1024`-float `Vec<f32>`, packed into a SQLite BLOB, so retrieval can rerank candidates by cosine similarity rather than relying on the entity index alone.
The active embedder is selected by `build_embedder_from_config` ([`embed/factory.rs`](../../../src/openhuman/memory_tree/score/embed/factory.rs)) walking a resolution ladder, identical for read and write paths:
1. **Explicit Ollama override** (`memory_tree.embedding_endpoint` + `embedding_model`) - power users / E2E rigs.
2. **Local Ollama** via the unified `embeddings` workload setting - the "Memory embeddings" checkbox in [Local AI](../model-routing/local-ai.md) Settings.
3. **User-configured OpenAI-compatible** endpoint (`OpenAiCompatEmbedder`, e.g. LM Studio).
4. **Managed cloud** (`CloudEmbedder`, OpenHuman backend / Voyage) - the default once logged in.
5. **No provider** - the read path falls back to `InertEmbedder` (zero vectors) so retrieval still runs; the write path returns `None`, skips embedding, and flags `semantic_recall` degraded so the chunk can be re-embedded later.
Embeddings run on the background workers, not the ingest hot path, so a burst of new sources never blocks the UI. Trees give compression and navigation; embeddings keep similarity search working underneath them.
***
## See also
* [Memory Trees](memory-tree.md) - the pipeline scoring sits inside.
* [Retrieval](retrieval.md) - how the index and embeddings are queried.
* [Obsidian Wiki](README.md) - the Markdown vault scored chunks land in.
* [Token Compression](../token-compression.md) - why keeping the tree dense matters.
+109
View File
@@ -0,0 +1,109 @@
---
description: >-
The typed registry of connectors that feed your Memory Tree — local folders,
GitHub repos, RSS feeds, web pages, and Composio OAuth integrations — plus
per-agent source scoping for privacy and focus.
icon: database
---
# Memory Sources & Scoping
A **memory source** is a configured connector that feeds the [Memory Tree](memory-tree.md). Where the tree owns _"how do I store and summarize?"_, the `memory_sources` domain (`src/openhuman/memory_sources/`) owns the upstream question: **"what feeds my memory?"** It is a typed registry of connectors — persisted in `config.toml` under `[[memory_sources]]` — with CRUD at runtime, a uniform reader abstraction, per-source sync status, and the `openhuman.memory_sources_*` RPC surface.
The domain only _defines connectors and reads from them_. The ingestion engine and sync scheduling live in `memory` / `memory_sync`; sources dispatch work to the right backend.
***
## Source kinds
Every source is a single flat `MemorySourceEntry` (`src/openhuman/memory_sources/types.rs`) whose `kind` discriminator (the `SourceKind` enum) decides which fields are required — validation is enforced at add/update time by `validate()`, not the type system. The kinds:
| Kind | `SourceKind` | What it ingests |
| ---------------- | -------------- | -------------------------------------------------------------------------------------------- |
| **Composio** | `Composio` | An OAuth-connected SaaS integration (Gmail, Slack, Notion, …) — sync is provider-driven. |
| **Conversation** | `Conversation` | The agent's own conversation transcripts. |
| **Folder** | `Folder` | A local directory, globbed (default `**/*.md`, 10 MB/file cap) with a path-traversal guard. |
| **GitHub repo** | `GithubRepo` | Project activity — commits, issues, PRs — via the `gh` CLI or a public REST fallback. |
| **RSS feed** | `RssFeed` | RSS/Atom feed items. |
| **Web page** | `WebPage` | A fetched web page, optionally narrowed by a CSS `selector`. |
| **Twitter query**| `TwitterQuery` | A saved Twitter query — reader scaffolded, sync intentionally unimplemented pending creds. |
Each entry also carries optional per-sync budgets — `max_tokens_per_sync`, `max_cost_per_sync_usd`, `sync_depth_days` — so a chatty source can't blow up your token spend on one run.
***
## Adding and configuring sources
Sources are CRUD-ed through the `memory_sources` controllers (`src/openhuman/memory_sources/schemas.rs``rpc.rs`), namespace `openhuman.memory_sources_*`:
| RPC | Purpose |
| ------------- | ------------------------------------------------------------------ |
| `list` | List configured sources (lazily reconciles Composio first). |
| `get` | Fetch one source by `id`. |
| `add` | Add a source; kind-specific fields are flat on the request. |
| `update` | Partial update via `MemorySourcePatch`. |
| `remove` | Delete a source by `id`. |
| `list_items` | List readable items from a source via its reader. |
| `read_item` | Read one item's content. |
| `sync` | Queue a manual sync (returns immediately; progress via events). |
| `status_list` | Per-source sync status. |
All mutations reload the live `Config`, apply the change, and `config.save()` atomically (`registry.rs`). In the desktop app these surface in the Intelligence / Memory tab alongside the [Auto-fetch](auto-fetch.md) cadence.
***
## The reader abstraction
Every kind implements one async trait, `SourceReader` (`src/openhuman/memory_sources/readers/mod.rs`):
```rust
#[async_trait]
pub trait SourceReader: Send + Sync {
fn kind(&self) -> SourceKind;
async fn list_items(&self, source, config) -> Result<Vec<SourceItem>, String>;
async fn read_item(&self, source, item_id, config) -> Result<SourceContent, String>;
}
```
A `reader_for(kind)` dispatcher hands back the right implementation (`FolderReader`, `GithubReader`, `RssReader`, `WebPageReader`, etc.). On a manual `sync`, reader-backed kinds walk `list_items` and ingest each item through `memory::ingest_pipeline::ingest_document` (`sync.rs`); Composio sources delegate wholesale to `memory_sync::composio::run_connection_sync` rather than reading item-by-item, so `ComposioReader::read_item` is an explanatory placeholder.
***
## Sync status & freshness
`status.rs` computes a `SourceStatus` per source by querying `mem_tree_chunks` (chunks synced/pending, last-chunk timestamp) using a `source_id LIKE` prefix — `mem_src:{id}:%` for reader kinds, `{toolkit}:%` for Composio. Each source gets a `FreshnessLabel`:
- **Active** — last chunk ≤ 30 s ago.
- **Recent** — last chunk ≤ 5 min ago.
- **Idle** — older, or no chunks yet.
Sync progress streams as `MemorySyncStageChanged` events (Requested → Fetching → Stored → Ingesting → Completed/Failed), tagged with `connection_id = Some(source.id)`, so the UI can show live progress without polling. `status_list` degrades a per-source query failure to an `Idle` zero-row entry rather than failing the whole call.
**Composio auto-upsert.** When an OAuth connection is created, `memory_sync::composio::bus` calls `upsert_composio_source`, so freshly-connected integrations appear as sources with no restart. `list_rpc` also performs a lazy reconciliation (`reconcile::ensure_composio_sources`) on every list, catching connections made before this hook existed.
***
## Source scoping for agent profiles
By default an agent recalls from **every** source. Source scoping lets an agent profile restrict recall to a whitelist of source ids — so a customer-support flavour never surfaces your personal Gmail, and a research flavour stays focused on the repos and feeds that matter. This is a privacy and focus control, not just a relevance tweak.
The mechanism lives in `src/openhuman/memory/source_scope.rs`. Threading an allowlist through every memory tool and the deep `select_trees` retrieval layer would touch dozens of call sites, so — mirroring `thread_context` — the channel sets a `tokio::task_local!` around the agent turn and the retrieval layer reads it ambiently, with no explicit plumbing:
- **`None`** (outside any scope, or `with_source_scope(None, …)`) means **unrestricted** — the default for cron, sub-agents, the CLI, and any profile that left `memory_sources` unset.
- **`Some(set)`** restricts recall to source scopes in the set. An **empty** set surfaces nothing (the profile selected no sources).
The gate is **tag-discriminated and fail-open** for everything that is not a memory-source chunk. Every source-ingested chunk carries the `memory_sources` tag; the gate (`chunk_source_allowed`) only touches tagged chunks:
- A chunk **without** the `memory_sources` tag — working memory, conversation transcripts, internal chunks — **always passes**, even under an empty allowlist.
- A **tagged** memory-source chunk passes only if its source id is allowed — matched against either the raw `source_id` (Composio / channel scopes like `slack:#eng`) or the registry id extracted from a `mem_src:<id>:<item>` composite (reader-based sources).
So tightening a profile's scope hides its connected sources without ever starving it of its own conversation context.
***
## See also
- [Auto-fetch](auto-fetch.md) — the 20-minute cadence that keeps active sources fresh.
- [Memory Trees](memory-tree.md) — the pipeline every source feeds into.
- [Obsidian Wiki](README.md) — the Markdown vault sources land in.
- [Integrations](../integrations/README.md) — connecting the OAuth providers behind Composio sources.
+123
View File
@@ -0,0 +1,123 @@
---
description: >-
How OpenHuman learns your communication style, identity, tooling, vetoes, and
goals from everyday use, then surfaces them as ambient defaults in every reply.
icon: brain
---
# Personalization & Self-Learning
OpenHuman gets to know you the way a good assistant would: not by asking you to fill in a settings form, but by paying attention. As you chat, connect accounts, and correct it, it quietly collects evidence about how you like to work, scores that evidence for stability, and promotes the durable signals into your **`PROFILE.md`** and into the system prompt of every future turn.
Nothing is locked in from a single message. A preference has to keep showing up before it earns a place in your profile, and anything that fades stops being injected. You stay in control: the profile is a plain Markdown file you can edit, and you can pin or forget any learned fact.
***
## What gets learned
Learning is organized into six **facet classes**. Each class has its own decay rate and its own budget so one noisy category can't crowd out the others.
| Class | What it captures | Examples |
| --- | --- | --- |
| **Style** | How you like replies written | `verbosity=terse`, `format=bullets`, `emoji=skip` |
| **Identity** | Stable facts about you | `name=Alice`, `timezone=PST`, `role=engineer` |
| **Tooling** | Your developer toolchain | `package_manager=pnpm`, `editor=neovim`, `lang=rust` |
| **Veto** | Things you've explicitly rejected | `avoid em dashes`, `no nested bullets` |
| **Goal** | Active goals and ongoing projects | free-form goal sentences |
| **Channel** | Your preferred place to talk | `primary=desktop-chat` |
Recurring people, topics, and past threads are **not** stored here — those live in the [memory tree](obsidian-wiki/memory-tree.md) and are pulled in per-turn by `memory_recall`.
***
## The learning pipeline
Evidence flows through four stages: **capture → score → render → inject**.
```
your activity candidate buffer stability detector
──────────── ──────────────── ──────────────────
chat turns ──┐
corrections ──┤
email signatures ──┼──→ LearningCandidate ──→ rebuild every 30 min
connected accounts──┤ (class, key, value, + event-driven (~60s
LinkedIn (opt-in) ──┘ cue family, evidence) after new data)
│ score each (class, key)
│ resolve value conflicts
│ apply per-class budgets
user_profile facets
(Active / Provisional /
Candidate / Dropped)
CacheRebuilt ───┤
┌───────────────────────┴───────────────┐
▼ ▼
PROFILE.md system prompt
(managed blocks) ("Your standing preferences")
```
**Capture.** Producers watch your activity and push a `LearningCandidate` into a bounded buffer. Each candidate records the `(class, key, value)` it asserts, a pointer back to the evidence, and a **cue family** describing how strong the signal is — `Explicit` (you said it outright), `Structural` (from account data or a file), `Behavioral` (inferred from how you act), or `Recurrence` (a statistical pattern).
**Score.** A background **stability detector** rebuilds the cache roughly every 30 minutes, and sooner when new email or documents arrive. It aggregates every candidate for a given fact and computes a stability score: stronger cue families count for more, recent evidence counts for more than stale evidence (each class has a half-life), and an explicit statement from you doubles the weight.
| Class | Evidence half-life |
| --- | --- |
| Identity | 90 days |
| Veto | 60 days |
| Tooling / Goal | 30 days |
| Style | 14 days |
| Channel | 7 days |
The score decides each fact's lifecycle state:
| State | Meaning |
| --- | --- |
| **Active** | Strong enough to render in `PROFILE.md` and inject into the prompt |
| **Provisional** | Stored and tracked, but not yet shown |
| **Candidate** | Still gathering evidence |
| **Dropped** | Faded below the floor; removed |
When two values compete for the same fact (e.g. `verbosity=terse` vs `verbosity=detailed`), the higher-stability value wins. The loser is dropped, and if it ever becomes true again it re-earns its place naturally.
***
## Where it's stored: `PROFILE.md`
The learned profile is materialized into **`PROFILE.md`** in your workspace — a real, editable Markdown file. Each facet class owns a managed block (`## Style`, `## Identity`, `## Tooling`, `## Vetoes`, `## Goals`), and only **Active** facets are written, sorted by stability. Pinned entries are marked `*(pinned)*`.
The renderer only touches its own managed blocks. Anything you write by hand outside those blocks — and the separate `## Connected Accounts` block owned by the integrations layer — is left untouched. Empty classes show a `*(no entries yet)*` placeholder rather than disappearing.
> **Per-session freeze.** `PROFILE.md` is folded into the agent's system prompt at the start of a session and held stable for that session's lifetime, which keeps prompt caching fast. Edits you make mid-session are picked up on the next rebuild and the next session — not retroactively in the current one.
***
## How it surfaces in replies
On every turn the agent reads the Active facets and injects them as a compact **"Your standing preferences"** section in the system prompt, alongside a standing instruction to call `memory_recall` before answering anything that leans on past sessions. The result is that the agent defaults to your verbosity, your tools, and your vetoes without being reminded — and reaches into memory for the specifics.
***
## Optional LinkedIn enrichment
During onboarding you can let OpenHuman bootstrap your identity from LinkedIn. The flow searches your connected Gmail for a `linkedin.com/in/...` profile URL, and (when available) scrapes the public profile, then compresses what it finds into `PROFILE.md` via the `learning_save_profile` step. It runs once, as a fire-and-forget pass — entirely opt-in, and skipped cleanly if no profile is found.
***
## Reviewing and controlling what's learned
Everything learned is inspectable and reversible:
- **Edit `PROFILE.md` directly.** It's your file. Correct, add, or delete anything; the next rebuild respects your edits.
- **The Brain page** (raised center button in the bottom bar, `/brain`) is the home for memory and intelligence — the knowledge graph, goals, sources, and sync status all live here.
- **Pin** a fact to lock it Active and shield it from decay, or **forget** a fact to drop it and block it from coming back. Under the hood these are the `learning_pin_facet`, `learning_unpin_facet`, and `learning_forget_facet` operations over the `openhuman.learning_*` RPC surface, alongside `learning_list_facets` and `learning_rebuild_cache`.
***
## See also
* [Memory Tree](obsidian-wiki/memory-tree.md), where recurring people, topics, and threads live and are recalled per-turn.
* [Goals & To-dos](goals-and-todos.md), the goal-tracking surface that pairs with learned `goal/*` facets.
* [Subconscious Loop](subconscious.md), the background engine that keeps thinking about your workspace between turns.
-75
View File
@@ -1,75 +0,0 @@
---
description: >-
OpenHuman 以什么形式交付(原生 React + Tauri v2 桌面应用,Rust core)、
支持的平台,以及当前范围内的功能。
icon: layer-plus
---
# 平台与可用性
OpenHuman 是一个原生桌面应用,不是浏览器扩展,也不是 Electron 包装器。基于 **React + Tauri v2** 构建,搭载 **Rust core**,它体积小、启动快、不干扰你的工作流。
***
## 支持的平台
| 平台 | 架构 | 分发方式 |
| ---------- | ---------------------- | -------------------------- |
| **macOS** | Intel、Apple Silicon | `.dmg` 安装包、Homebrew |
| **Windows**| x64、ARM64 | `.msi` 安装包 |
| **Linux** | x64 | AppImage、`.deb` |
***
## 为什么是原生应用
OpenHuman 作为原生应用构建而非 Web 包装器,有三个原因:
**体积小。** 只有典型通信工具的几分之一。不到一秒启动,内存占用极少。
**启动快。** 无需初始化浏览器引擎。立即就绪接受请求。
**操作系统级安全。** 凭据保存在你平台的安全密钥链中:macOS Keychain、Windows Credential Manager、Linux Secret Service。敏感数据永不放在浏览器存储或明文文件中。本地记忆树的 SQLite 数据库位于你的工作区文件夹中,由你拥有。
***
## 架构概览
```text
┌──────────────────────────────────────────────────┐
│ Tauri shell - windowing, OS integration │
└──────────────────────────────────────────────────┘
│ JSON-RPC ↕
┌──────────────────────────────────────────────────┐
│ Rust core(进程内 `openhuman` core)│
│ • Memory Tree, integrations, auto-fetch │
│ • Model router, TokenJuice, native tools │
│ • Voice (STT in, TTS out, Meet agent) │
└──────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────┐
│ React frontend - screens, navigation │
└──────────────────────────────────────────────────┘
```
Shell 是载体(负责窗口化、进程生命周期、IPC)。所有产品逻辑都在 Rust core 中。React 前端通过 JSON-RPC 与 core 通信。参见[架构](../developing/architecture/)获取完整图景。
***
## 实时通信
桌面应用与 OpenHuman 后端保持持久连接。响应在生成时流式输出;输出渐进出现,而非等待后的最终结果。如果网络断开,应用会自动重连,使用渐进退避。
***
## 离线行为
你的本地状态保存在设备上。偏好设置、设置和连接的源配置在离线时仍然可用。本地记忆树完全可访问,你可以浏览 [Obsidian 存储库](obsidian-wiki/),在无网络连接的情况下阅读你现有的笔记。
自动拉取和实时 LLM 调用需要网络连接。网络恢复时,下一个 20 分钟触发周期会从上次停止的地方继续。
***
## 自动更新
桌面 shell 通过 Tauri 的更新插件自动更新,针对 GitHub Releases 上发布的一份清单。进程内 OpenHuman core 打包在同一 bundle 中,所以 shell 更新会同时升级两者。
@@ -1,97 +0,0 @@
---
icon: shield
---
# 隐私与安全
OpenHuman 的设计使得**你生活的记忆活在你的机器上**。本地 SQLite 记忆树、Markdown Obsidian 存储库、你的音频缓冲区,所有这些都在你的控制之下。OpenHuman 后端处理必须经纪的事(LLM 调用、OAuth token、搜索代理),别无其他。
***
## 隐私设计
**记忆树是本地的。** SQLite 数据库(`<workspace>/memory_tree/chunks.db`)和 Markdown 存储库(`<workspace>/wiki/`)位于你的机器上。智能体在本地读取;你的原始源数据没有任何内容存在于 OpenHuman 后端。
**集成 token 由后端持有,不在你的笔记本上。** OAuth token 永不以明文形式写入磁盘。OpenHuman 后端经纪每个集成请求,core 从不直接与任何第三方 API 通信。
**操作系统级凭据存储。** 敏感 token 存储在你平台的安全密钥链中:macOS Keychain、Windows Credential Manager、Linux Secret Service。
**不使用你的数据训练。** 你的对话、记忆树和个人信息永不用于训练 AI 模型或改进系统。
**可选**[本地 AI](model-routing/local-ai.zh-CN.md)**。** 如果你想让嵌入向量和摘要树构建保留在你的机器上,可以选择加入。心跳/学习/潜意识循环同样可以移至端侧。
***
## 保留在你机器上的内容
| | |
| ------------------------------- | --------------------------------------------------------------- |
| **记忆树 SQLite 数据库** | 本地 - `<workspace>/memory_tree/chunks.db`。 |
| **Obsidian Markdown 存储库** | 本地 - `<workspace>/wiki/`。你可以阅读、编辑、复制、删除。 |
| **音频捕获缓冲区** | 本地。STT 后丢弃。 |
| **本地模型状态** | 本地。 |
## OpenHuman 后端处理的内容
| | |
| ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **LLM 调用** | 通过一个订阅由后端代理,然后按[模型路由器](model-routing/)转发至底层提供商(Anthropic / OpenAI / Google 等)。 |
| **网络搜索代理** | 原生[网络搜索工具](native-tools/web-search.zh-CN.md)调用后端代理,这样你无需携带搜索 API key。 |
| **集成 OAuth 和工具代理** | [118+ 集成](integrations/README.zh-CN.md)的 token 存储和限速请求经纪。 |
| **TTS 流媒体** | 托管[文字转语音](native-tools/voice.zh-CN.md)音频流。音频生成后丢弃——不保留。 |
***
## 权限和访问控制
OpenHuman 仅在你完成 OAuth 流程后才访问集成。每个连接有自己的作用域;你可以随时从 Skills 标签页撤销它们。
[自动拉取](obsidian-wiki/auto-fetch.zh-CN.md)确实在连接活跃时持续运行,这正是它的意义所在。但它受以下约束:
* 你授予该集成的 **OAuth 作用域**
* 每个 provider 的**同步间隔**(例如 Gmail 默认每 15 分钟)。
* 每个连接的**每日预算**,限制 API 使用。
如果你撤销一个连接,下一个 tick 停止同步;已在你本地记忆树中的块保留在那里,因为它们是你的。
***
## 为什么本地记忆是隐私的
大多数 AI 助手面临权衡:更多上下文意味着更多原始数据发送到云端。记忆树消除了这一权衡。
因为规范化、分块、评分和摘要树都在**你本地 Rust core 内部**运行,你的原始源数据永不离开你的机器。LLM 看到的唯一东西是智能体在当前轮次从你本地记忆树中检索的内容,而该检索由你的提示管理,而非后台上传。
压缩和本地性共同构成隐私架构。
<figure><img src="../.gitbook/assets/V17 — Privacy Shield@2x.png" alt=""><figcaption></figcaption></figure>
## 安全性
**传输加密。** 应用与 OpenHuman 后端之间的所有通信使用 TLS。没有数据以明文传输。
**沙盒技能。** 每个技能在自己的隔离执行环境中运行,强制执行内存和资源限制。技能无法访问彼此的数据、主机系统的文件系统或你的凭据。
**工作区作用域工具。** 原生[文件系统工具](native-tools/coder.zh-CN.md)在用户打开的工作区内操作;它们没有对磁盘其余部分的ambient访问权限。
**短生命周期 token。** 应用与后端之间的认证 token 是有时间限制的。
***
## 信任与风险智能
OpenHuman 包含一个智能层,旨在帮助你推理已连接来源的可信度、信息质量和潜在风险。
**诈骗和冒充信号。** 与诈骗、冒充或协调滥用相关的行为模式可以浮现为警告。信号来自模式,而非来自共享个人消息内容。
**上下文动态信任。** 信任是上下文相关的,一个领域的可信度不会自动转移到另一个。OpenHuman 通过聚合数据和历史准确性而非静态分数来呈现信任。
**建议性,而非执行性。** 信任和风险输出是为你判断提供信息的建议性信号。OpenHuman 不会封禁用户、删除消息或执行审核决策。
***
## 共享环境
在团队或社区环境中,隐私仍以用户为中心。每个用户的已连接源都作用域到其账户;管理员无法通过后门访问其他用户的记忆树。
社区级智能从聚合和匿名化的信号中推导,绝不从直接访问个人消息内容获取。
+109
View File
@@ -0,0 +1,109 @@
---
description: >-
Invite friends, earn referral credit, redeem promo codes, and unlock Discord
community roles - the in-app Rewards & Referrals surface.
icon: gift
---
# Rewards & Referrals
OpenHuman bundles three loosely related growth mechanics behind one surface: a **referral program** (share a code, earn credit when friends convert), **promo coupons** (redeem a code for promotional credit), and a **community rewards** track (link Discord, unlock roles as you hit usage milestones). Invite-code management lives on its own screen.
All of this requires a signed-in backend session. On a local-only session the Rewards page shows an empty state prompting you to sign in - none of these features work offline.
***
## The Rewards screen
Lives at `/rewards` with three chip tabs. The middle **Rewards** (community) tab is selected by default.
| Tab | What it does |
| ------------- | ----------------------------------------------------------------- |
| **Referrals** | Your referral code, earnings, and referred-user activity |
| **Rewards** | Discord connection, progress ring, and unlockable community roles |
| **Coupons** | Redeem promo codes for promotional credit + redemption history |
***
## Referrals
Each account has a single **referral code**. Copy it, or use **Share** (native share sheet, falling back to clipboard) to send a prefilled message with your code and the app download link.
The tab shows four tiles: your code, **total earned** (USD), **pending referrals**, and **completed** referrals. Below them, an **activity table** lists each referred user (masked identity, e.g. `j***@gmail.com`), a status badge, the reward amount, and a timestamp.
| Referral status | Meaning |
| --------------- | -------------------------------------------------- |
| Joined | Referred user signed up but hasn't converted yet |
| Completed | Referred user converted - referral reward credited |
| Expired | Relationship lapsed (reserved; backend-driven) |
If you were referred by someone else and are still eligible, an **apply** form lets you enter their code. Eligibility (`canApplyReferral`) is decided by the backend - typically only users who haven't already subscribed or already applied a code can claim one. Once applied, the form is replaced by a confirmation of the linked code.
Reward amounts, conversion rules, and eligibility are all enforced **server-side**. The desktop core is a thin adapter here.
### Under the hood
The referral domain (`src/openhuman/referral/`) is a stateless RPC adapter, not business logic. It exists because the desktop WebView `fetch` can fail with a generic "Load failed" (CORS/TLS/WebKit), so these calls reuse the same server-side `reqwest` path as billing.
| RPC | Backend call | Purpose |
| -------------------- | ----------------------- | ----------------------------------------- |
| `referral.get_stats` | `GET /referral/stats` | Code, totals, and referred-user rows |
| `referral.claim` | `POST /referral/claim` | Apply a referral code (optional device fingerprint for abuse signals) |
Both fail closed with `no backend session token` when no session is stored.
***
## Coupons (Redeem)
The Coupons tab redeems **promo codes** for promotional credit - separate from referral rewards. Two tiles show your **promo credit balance** (USD) and the **count of redeemed codes**. Enter a code and redeem; redemption is either applied immediately or accepted as **pending** when it's conditional on a later action.
A **recent redemptions** table lists each code, its reward amount, status, and when it was redeemed.
| Coupon status | Meaning |
| --------------- | ------------------------------------------------------ |
| Applied | Fulfilled - credit is on your account |
| Pending action | Conditional coupon awaiting a triggering action |
| Redeemed | Accepted, not yet fulfilled |
***
## Community rewards & Discord
The Rewards tab gamifies usage. A **progress ring** shows how many achievements you've unlocked out of the total, and a list of **roles & rewards** describes each milestone (some carry an optional USD credit). Status badges (current streak, cumulative tokens) summarize your activity at the bottom.
Rewards are delivered as **Discord roles**, so the tab is built around linking your Discord account:
1. **Connect Discord** runs an OAuth consent flow (`openhuman.auth.oauth_connect` with provider `discord`); on success the snapshot refreshes and shows your Discord username.
2. **Join Discord** opens the community server invite.
3. **Disconnect** unlinks the account (clears the stored Discord ID, idempotent).
Once linked, each unlocked achievement shows its Discord role-assignment state:
| Role status | Meaning |
| --------------- | ----------------------------------------------------------- |
| Assigned | Role granted on the server |
| Pending | Unlocked but the role hasn't been assigned yet |
| Join to claim | Linked but not in the server - join to receive the role |
If you've unlocked a role-bearing achievement but haven't joined the server, a **claim banner** prompts you to join. Membership status is one of `member`, `not_in_guild`, `not_linked`, or `unavailable`.
> GitHub-based contributor rewards are a **separate** mechanism: a GitHub Actions workflow (`.github/workflows/contributor-rewards.yml`) that posts a Discord/merch invite comment when a contributor's first PR merges. It is not part of the in-app Rewards screen and uses no in-app GitHub OAuth.
***
## Invite codes
The **Invites** screen (`/invites`) is distinct from referral codes. It manages personal **invite codes** that gate new-user signup:
- **Redeem** - if you haven't been invited yet, enter an invite code to claim your spot.
- **Your invite codes** - a list of the codes issued to you. Each row shows the code (monospace), a copy button, and an **enabled/disabled** state. A code flips to disabled once its uses are exhausted (`currentUses >= maxUses`), and shows who claimed it.
Invite codes carry a `type` (`USER` or `CAMPAIGN`), `maxUses`/`currentUses` counters, and a `usageHistory` of who redeemed them and when.
***
## See also
* [Billing & usage](billing-and-usage.md) - where referral, coupon, and achievement credit gets spent.
* [Welcome](../README.md) - the documentation home.
+105
View File
@@ -0,0 +1,105 @@
---
description: >-
On-device screen capture, OCR + vision summarization, input automation, and
inline autocomplete — gated behind explicit macOS privacy permissions.
icon: scan-eye
---
# Screen Intelligence
Screen Intelligence lets the agent see what you're working on. When you start a consent-gated session, OpenHuman periodically screenshots your **active window**, runs it through on-device OCR and a local vision model, and synthesizes a short "what the user is doing right now" note into memory. On top of that capture loop it offers **input automation** (text staging, keyboard actions, a panic stop) and a separate **system-wide autocomplete** that suggests inline text completions in any focused field.
Everything runs locally. Capture, OCR, and summarization happen on your machine using the local model (Ollama); nothing is sent to the cloud as part of this feature.
> **Platform:** Screen Intelligence is **macOS-only** in V1. On Windows and Linux the engine reports `platform_supported: false`, `start_session` returns a `macOS-only` error, and the embedded capture server does not autostart. Microphone permission detection is the only cross-platform piece.
***
## What it does
A session runs two cooperating background workers:
| Worker | Role |
| --- | --- |
| **Capture worker** | Polls the foreground window at `baseline_fps` (default 1 fps), screenshots the **active window only** via `screencapture -l <windowID>` (no fullscreen fallback), applies the allow/deny policy, optionally saves a PNG, and enqueues the frame. |
| **Processing worker** | Drains to the **latest** frame, compresses it (PNG → JPEG, longest edge ≤ 1024 px, quality 72), runs Apple Vision OCR, then a vision LLM, then a synthesis LLM, and persists a `VisionSummary` to memory. |
Capture only proceeds when the foreground context exposes a `window_id`. Stale queued frames are discarded — only the most recent frame is analyzed, deduped by capture timestamp.
### Sessions
Sessions are explicitly consent-gated and time-boxed:
* You start a session from Settings (or the `screen_intelligence.start_session` RPC) with `consent: true`.
* TTL is clamped to **303600 seconds** (the Settings panel defaults to 300 s). When the TTL expires the session stops on its own.
* **Analyze now** flushes the pending frame through the vision pipeline immediately.
* **Stop** ends the session; the `panic_stop` input action force-stops it instantly.
***
## Input automation
While a session is active, `input_action` lets the agent stage text and signal keyboard intent into the session's autocomplete context. Actions are blocked when no session is active or when the foreground app is denylisted by the active policy. The special `panic_stop` action immediately tears down the session regardless of state — a hard kill switch.
The capture-side autocomplete helpers (`autocomplete_suggest` / `autocomplete_commit`) maintain a small in-memory context buffer (capped at 256 chars) and return heuristic suggestions; they're gated on `autocomplete_enabled`.
***
## Autocomplete
Separate from the capture session, OpenHuman ships a **system-wide inline autocomplete** engine (the `autocomplete` domain). It is also **macOS-only** at runtime.
* It captures your currently-focused text field through the macOS accessibility (AX) layer, runs **local** inference to generate a short single-line continuation, and renders it in a floating overlay badge.
* Press **Tab** to accept (it inserts the text and cleans up any stray indentation the app added), or **Escape** to reject.
* Accepted completions are saved as personalization examples in a local KV store and a local memory-doc namespace, and feed back into later suggestions.
* It special-cases terminals (extracting just the input line), skips blocked/disabled apps and OpenHuman's own window, and filters low-quality suggestions (too short, no alphanumerics, or an echo of what you just typed).
* Debounce is clamped 502000 ms; the displayed/applied suggestion is capped at 64 characters. After 5 consecutive inference failures the engine auto-stops to avoid notification floods.
There is also an in-app path: the OpenHuman composer passes an explicit `context` to the engine, bypassing AX capture entirely.
***
## Permissions
Capture and automation require macOS privacy grants. OpenHuman detects each one and can open the relevant **System Settings → Privacy & Security** pane and trigger the system prompt.
| Permission | Why it's needed | Detected via |
| --- | --- | --- |
| **Screen Recording** | Screenshot the active window for OCR + vision. | `CGPreflightScreenCaptureAccess` |
| **Accessibility** | Read the foreground window/element, capture focused text, and insert text (required to start a session and to run autocomplete). | `AXIsProcessTrusted` |
| **Input Monitoring** | Listen for accept/reject key edges (Tab/Escape) and the Globe/Fn hotkey. | `IOHIDCheckAccess` |
| **Microphone** | Voice features (cross-platform; the only permission detected off macOS). | CPAL device probe |
> **Restart after granting.** macOS TCC grants are per-executable **and per-process** — a running core never sees a freshly granted permission. After you grant a permission you must restart the core for it to take effect. The status payload carries `permission_check_process_path` and the core process pid/start time so the UI can confirm a restart actually happened. The panel exposes a "refresh with restart" action for this.
***
## Privacy considerations
* **Local-only processing.** OCR (Apple Vision) and both LLM passes run on-device via the local model. The vision pipeline requires `local_ai.runtime_enabled = true` with the `ollama` provider; without it, analysis errors out rather than falling back to a cloud call.
* **What's stored.** Each summary is written to **unified memory** in the `background` namespace (source type `screenshot`, tagged `screen_intelligence`) as a small markdown doc with the app name, window title, capture timestamp, and confidence. Raw frames are **not** persisted to memory.
* **Screenshots on disk are opt-in.** PNGs are only saved to `{workspace}/screenshots/` when **Keep screenshots** is enabled; otherwise a temp file is written for OCR and deleted immediately.
* **Allow/deny policy.** A policy mode (`all_except_blacklist` or `whitelist_only`) plus allow/deny app lists control which windows are ever captured.
* **Time-boxed + consent-gated.** Nothing captures until you start a session, and every session expires on its TTL.
***
## Enabling & configuring
Configure it under **Settings → Screen awareness** (the `ScreenIntelligencePanel`):
* **Permissions section** — per-permission status (granted / denied / unknown), request buttons, and the restart-to-refresh flow.
* **Enabled** — master toggle for the feature.
* **Mode** — `All except blacklist` or `Whitelist only`, backed by the allow/deny lists.
* **Screen monitoring** — toggle the capture loop for a session.
* **Session controls** — Start / Stop / **Analyze now**, with a live remaining-time countdown. Start is disabled until Accessibility is granted and the platform is supported.
Under the hood these map to the `[screen_intelligence]` config block (`enabled`, `vision_enabled`, `use_vision_model`, `keep_screenshots`, `baseline_fps`, `session_ttl_secs`, `policy_mode`, allow/deny lists, `autocomplete_enabled`) and the `screen_intelligence.*` JSON-RPC surface. The same capabilities are available from the `openhuman screen-intelligence` CLI (`status`, `start`, `stop`, `capture`, `doctor`, `run`).
***
## See also
* [Memory Tree](obsidian-wiki/memory-tree.md) — where vision summaries land.
* [Privacy & Security](privacy-and-security.md) — the broader permission and data model.
* [Voice](native-tools/voice.md) — the other feature that uses the Microphone permission.
-186
View File
@@ -1,186 +0,0 @@
---
description: >-
后台循环,评估用户/系统任务相对于工作区的状态,
并决定做什么。
icon: loader
---
# 潜意识循环
一个后台任务评估和执行系统。在每个周期 tick 上,它加载用户定义和系统任务列表,读取你工作区的当前状态,决定对每项做什么,然后要么自主行动,要么升级给你审批。
把它想象成智能体的空闲线程:你停止打字后仍在继续思考的部分。
***
## 一个 tick 如何工作
```text
┌─────────────────────────────────────────────────────────┐
│ Heartbeat │
│ (tick 之间睡眠几分钟) │
└──────────────────────┬──────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ Subconscious Engine │
│ │
│ 1. 加载到期任务 │
│ 2. 将每项标记为进行中 │
│ 3. 构建情况报告(记忆 + 工作区) │
│ 4. 用本地模型评估每个任务 │
│ 5. 执行决定(act / noop / escalate
│ 6. 将结果写回活动日志 │
└─────────────────────────────────────────────────────────┘
┌───────────┼───────────┐
▼ ▼ ▼
noop act escalate
(skip) (execute) (deeper agent)
```
每个 tick 是独立的。如果一个 tick 在下一个开始时仍在运行(慢速模型调用、网络抖动),新的 tick 接管,旧的在进行中的条目被标记为已取消。Tick 永远不会堆叠。
***
## 任务类型
### 系统任务
引擎启动时自动播种。不能删除,只能禁用。默认覆盖你希望任何助手监视的事情:
* 检查已连接技能的错误或断开
* 审查新记忆更新中的可操作项目
* 监控系统健康(本地模型、记忆、连接)
你可以通过在 workspace 的 `HEARTBEAT.md` 文件中列出额外的系统任务来扩展,每行一个任务。
### 用户任务
你从 UI 手动添加的任何内容。切换开/关、编辑、删除。例如:
* "检查紧急邮件"(只读)
* "发送每日摘要到 Slack"(写意图)
* "总结 Notion 更新"(只读)
***
## 决策
对于每个到期任务,本地模型返回三个决策之一:
| 决策 | 含义 |
| -------- | --------------------------------------------------- |
| Skip | 现在没什么相关的 |
| Act | 发现了相关的东西,执行任务 |
| Escalate | 需要更深入的推理,交给云端智能体 |
决定如何执行取决于任务是否有**写意图**(它要求智能体执行一个操作)还是**只读**(它要求智能体查看和报告):
```text
Decision: Skip
→ 记录"没什么新东西",调度下一次运行
Decision: Act
→ 在本地模型上执行(读或写)
Decision: Escalate
├─ 写意图任务
│ → 用完整权限运行云端智能体
│ → 不需要批准(你明确要求了该操作)
└─ 只读任务
→ 用仅分析模式运行云端智能体
→ 如果智能体浮出未经请求的推荐操作
│ → 为你审批创建升级卡片
│ → 批准后 → 用完整权限重新运行
└─ 否则 → 记录结果,完成
```
每个任务评估都带着彩色点和简短状态落在活动日志中:
| 状态 | 颜色 | 文本 |
| ----------------- | -------------- | ---------------------- |
| 进行中 | 蓝色(脉冲) | "评估中…" |
| 已行动 | 绿色 | 结果文本 |
| 已跳过 | 灰色 | "没什么新东西" |
| 等待批准 | 琥珀色 | "等待批准" |
| 失败 | 珊瑚色 | 错误消息 |
| 已取消 | 灰色 | "已取消" |
| 已忽略 | 灰色 | "已跳过" |
***
## 两个模型,一个循环
| 阶段 | 运行位置 | 为什么 |
| -------------------------------------- | ----------------------- | -------------------------------------------- |
| 每个任务评估(每个 tick) | 本地模型(Ollama) | 免费,无速率限制,适合端侧 |
| 仅文本执行(摘要、检查) | 本地模型 | 相同 |
| 工具使用执行(发送、发布…) | 云端智能体 | 工具、更大上下文、速率限制重试 |
| 升级读取的分析模式 | 云端智能体(只读) | 本地模型 defer 时更深入的推理 |
这种分割保持了循环便宜:只有当任务真正需要时你才为云端调用付费。
***
## 审批门
只有当智能体想要采取**你没有明确要求的写操作**时才需要审批。
| 任务意图 | 智能体想要写 | 需要审批? |
| ------------------------------ | -------------------- | -------------------------- |
| "发送摘要到 Slack"(写) | 是 | 否,你要求的 |
| "检查紧急邮件"(读) | 否 | 否,只读结果 |
| "检查紧急邮件"(读) | 是(转发它们) | **是**,未经请求的写 |
审批流程:
1. 云端智能体以仅分析模式运行。
2. 它浮出一个推荐,例如 _"将 3 封紧急邮件转发到 #team-alerts。"_
3. 升级卡片出现在 UI 的**需要审批**下。
4. **继续**用完整权限重新运行。
5. **跳过**什么都不做。
与技能相关的升级(断开的集成、过期的 OAuth、缺失的范围)显示一个**在技能中修复**按钮,直接带你到技能页面而不是。
***
## 失败处理
失败计数器跟踪连续 tick 全评估步骤失败(本地模型宕机、网络断开)。任何成功 tick 将其重置为零,并在 UI 状态栏中以珊瑚色显示(当非零时)。
每任务失败不会触发此计数器,tick 本身仍被认为成功。
如果一个 tick 失败或被取消,引擎不会推进其"上次看到"时间戳,所以下一次成功 tick 覆盖相同的窗口。你工作区中的任何内容都不会被跳过。
***
## 配置
循环可在桌面 app 中配置:
* **启用 / 禁用。** 打开或关闭整个后台循环。
* **Tick 间隔。** tick 触发的频率。默认为 5 分钟;这也是最小值。
* **推理。** 本地模型是否在每个 tick 评估任务。如果你想仅通过手动**立即运行**按钮运行,则禁用。
* **上下文预算。** 情况报告一次可以传入多少工作区。默认值是合理的;为更丰富的上下文调高,为更紧密的成本调低。
***
## 在 UI 中
位于**智能 → 潜意识**。
* **状态栏。** 任务数、总 tick 数、上次 tick 时间、失败计数器(如果有)。
* **进行中的任务。** 系统任务(只读,带"默认"badge)和你自己的任务(切换 + 删除)。
* **需要审批。** 待处理升级的琥珀色卡片。每张有标题、描述和优先级。按钮:**继续**、**在技能中修复**(当相关时)、或**跳过**。
* **活动日志。** 每个任务评估的按时间顺序的 feed,彩色点 + 结果。当任何内容进行中时自动刷新。
* **立即运行。** 手动触发一个 tick。立即返回;UI 轮询结果。
***
## 另见
* [记忆树](obsidian-wiki/memory-tree.zh-CN.md),情况报告从中读取。
* [从集成自动拉取](obsidian-wiki/auto-fetch.zh-CN.md),tick 之间工作区如何保持新鲜。
* [本地 AI(可选)](model-routing/local-ai.zh-CN.md),为评估提供支持的端侧模型。
+70
View File
@@ -0,0 +1,70 @@
---
description: >-
Deterministic, harness-driven context preparation on the first turn of every
new conversation - so the agent never starts cold.
icon: sparkles
---
# SuperContext
A fresh chat shouldn't start cold. **SuperContext** makes the agent gather relevant background _before_ it reads your first message — automatically, on every new thread, without you asking and without waiting on a tool call.
Most agents start a conversation blank and only fetch context if the model decides to call a "look things up" tool. That adds a round-trip, costs tokens, and depends on the model choosing well. SuperContext flips it: the harness itself prepares context up front, deterministically, so the very first reply already knows the relevant memories, files, and connected data.
***
## How it works
On the **first turn of a new thread**, if SuperContext is enabled, the harness:
1. Spawns a read-only `context_scout` sub-agent.
2. The scout sweeps your available data — the [Memory Tree](obsidian-wiki/memory-tree.md), workspace files, and connected integrations — and assembles a bounded **context bundle**.
3. The bundle is validated, then prepended to your message under a `Prepared context (super context)` header before the orchestrator model ever sees the turn.
4. The model answers your message already grounded in that context.
```
New thread, first message
┌──────────────────────────────┐
│ Harness gate (deterministic) │
│ super_context_enabled? │
└───────────────┬──────────────┘
│ yes
context_scout (read-only)
sweeps memory + files + data
[context_bundle] … [/context_bundle]
│ validated & extracted
Prepended to your message → orchestrator
```
Because the scout is **read-only**, it can never take an action on a fresh thread — it only reads and summarizes. And because it runs in the harness rather than as an optional tool, the redundant `agent_prepare_context` tool is suppressed for that turn, so the agent doesn't do the work twice.
***
## Safety and robustness
The scout returns its findings wrapped in `[context_bundle] … [/context_bundle]` tags. Only the bracketed envelope is ever injected — any surrounding prose the model emits ("sure, here's what I found…") is stripped out. If the bundle is missing, malformed (unterminated, reversed, or duplicated tags), or empty, the turn proceeds **gracefully without augmentation** rather than injecting garbage. A cold start is always preferable to a broken one.
***
## Turning it on or off
SuperContext is **on by default**.
* **From the composer.** A **Super Context** toggle appears below the chat input on a fresh thread. The flag is read when a thread is constructed, so toggling it affects **newly started threads**, not the one you're already in.
* **Config.** `context.super_context_enabled` (boolean, default `true`).
* **Environment.** `OPENHUMAN_SUPER_CONTEXT` (or `OPENHUMAN_CONTEXT_SUPER_CONTEXT_ENABLED`).
* **RPC.** `get_super_context_enabled()` reads the flag; `set_super_context_enabled(value)` sets and persists it.
***
## See also
* [Memory Tree](obsidian-wiki/memory-tree.md) — the primary source the scout reads from.
* [Auto-fetch from Integrations](obsidian-wiki/auto-fetch.md) — keeps that source fresh between conversations.
* [Subconscious Loop](subconscious.md) — the other side of "keeps thinking when you've stopped typing."
+63
View File
@@ -0,0 +1,63 @@
---
description: >-
Five built-in theme families, light/dark/auto variants, and a full visual
Theme Studio - make OpenHuman look the way you want.
icon: palette
---
# Themes & Theme Studio
OpenHuman is fully re-skinnable at runtime. Pick from built-in themes, switch light/dark/auto, or open the **Theme Studio** to design your own — every change applies instantly and persists locally, no restart required.
***
## Built-in themes
Five theme families ship out of the box, each with light and dark variants:
| Family | Feel |
| ------------ | ------------------------------------- |
| **Classic** | The default OpenHuman look. |
| **Ocean** | Cool blues around the `#4A83DD` primary. |
| **Sepia** | Warm, paper-like, easy on the eyes. |
| **Matrix** | High-contrast green-on-black. |
| **HAL 9000** | Deep black with a red accent. |
Each can be applied as **Light**, **Dark**, or **Auto**. Auto follows your operating system's `prefers-color-scheme` setting and re-applies live the moment you flip your OS between light and dark — no reload.
***
## Theme Studio
**Settings → Theme Studio** is a full visual editor. From it you can:
* **Pick a family** from a gallery of theme tiles (built-ins plus your own custom themes).
* **Adjust every colour token** with colour pickers — surfaces, text, borders, and accent ramps. A live **contrast warning** flags when text-on-background luminance drops below a readable threshold.
* **Swap fonts per role** — title, heading, body, mono, and serif can each use a different family.
* **Configure the backdrop** — an animated WebGL mesh, a flat colour, or a custom image, with an optional dotted overlay.
* **Manage custom themes** — create, edit, reset, delete, and **export / import as JSON** to share a theme with someone else.
### Editing a preset auto-forks
Changing any token on a built-in preset transparently creates a **new custom theme** — the original preset stays pristine. So you can start from Ocean, tweak it, and keep both.
***
## Where it's stored
Theme state — the active theme, the light/dark/auto variant, and all your custom themes — lives in Redux and persists to `localStorage` via `redux-persist`, so it survives app restarts and is scoped to your user. Sharing a theme is just exporting the JSON and having someone import it.
***
## Under the hood
Themes are driven by CSS custom-property **tokens** (space-separated RGB channel triples, so Tailwind opacity modifiers like `bg-surface/50` keep working). The `ThemeProvider` writes the active theme's overrides onto the `<html>` root; unspecified tokens fall through to the light/dark defaults.
For the full token taxonomy, Tailwind wiring, and component-authoring best practices, see the contributor reference: [Theming (developing)](../developing/theming.md).
***
## See also
* [Theming (contributor reference)](../developing/theming.md) — token system, Tailwind wiring, migration codemod.
* [Realtime Mascot](mascot/README.md) — the other big piece of OpenHuman's personality.
+130 -32
View File
@@ -1,51 +1,149 @@
---
description: >-
TokenJuice - a rule overlay that compacts verbose tool output before it ever
enters LLM context. Sweeping through thousands of emails stays cheap.
TokenJuice - a multi-stage compression router that compacts verbose tool
output and ingested data before it ever enters LLM context.
icon: file-zipper
---
# Smart Token Compression
LLM tokens are expensive, and verbose tool output is where most of them go to die. A `git status` in a busy repo, a `cargo build` log, a 600-message email thread, a `docker ps -a` against a real cluster, each of these can balloon a context window for almost no information gain.
LLM tokens are expensive, and verbose tool output is where most of them go to die. A `git status` in a busy repo, a `cargo build` log, a 600-message email thread, a `docker ps -a` against a real cluster each can balloon a context window for almost no information gain.
OpenHuman ships with **TokenJuice**, a port of [vincentkoc/tokenjuice](https://github.com/vincentkoc/tokenjuice) integrated directly into the tool-execution path. Before any tool result reaches the model, TokenJuice runs the output through a rule overlay that strips the noise and keeps the signal.
OpenHuman ships with **TokenJuice**, a compression router wired directly into the tool-execution and memory-ingestion paths. Before any tool result or ingested payload reaches a model, TokenJuice classifies it, routes it to a specialized compressor, optionally offloads the full original to a recoverable cache, and records how many tokens (and dollars) it saved.
## Three-layer rule overlay
It began as a port of [vincentkoc/tokenjuice](https://github.com/vincentkoc/tokenjuice) — that JSON rule overlay is still in here as the log/command compressor — but it has since grown into a multi-stage, content-aware pipeline.
Rules are JSON, and they merge in this order, later layers override earlier ones:
***
<table><thead><tr><th width="134.41796875">Layer</th><th>Path</th><th>Purpose</th></tr></thead><tbody><tr><td><strong>Builtin</strong></td><td>shipped with the binary</td><td>sensible defaults for git, npm, cargo, docker, kubectl, ls, etc.</td></tr><tr><td><strong>User</strong></td><td><code>~/.config/tokenjuice/rules/</code></td><td>your personal overrides, apply across every project</td></tr><tr><td><strong>Project</strong></td><td><code>.tokenjuice/rules/</code></td><td>repo-specific overrides, checked in, shared with the team</td></tr></tbody></table>
## The pipeline, step by step
Each rule names a tool/command pattern and a reduction strategy (truncate, dedup lines, fold whitespace, drop matching regexes, summarize sections, …). New rules are just JSON files; no recompile required.
Every blob that flows through `compact_tool_output(...)` takes the same path (`src/openhuman/tokenjuice/compress.rs`):
```
raw tool result / ingested payload
1. Size gate router enabled? input ≥ min_bytes_to_compress (2 KB)?
│ yes
2. Detect kind Json · Diff · Html · Search · Code · Log · PlainText
3. Select compressor one specialized compressor per kind (+ per-kind toggles)
4. Compress run it; if it declines or grows the output, fall back / pass through
5. CCR eligibility lossy AND ≥ ccr_min_tokens (≈500)? → offload original to cache
6. Append marker ⟦tj:<hash>⟧ footer so the agent can retrieve the full original
7. Record savings tokens + cost saved, by model and by compressor
compact text → LLM context
```
1. **Size gate.** If the router is disabled or the input is below `min_bytes_to_compress` (default **2048 bytes**), it passes through untouched — tiny outputs aren't worth compressing.
2. **Content detection** (`detect/kind.rs`). The blob is classified into one of seven `ContentKind`s. Precedence: an explicit hint → MIME/extension tag → a per-tool prior (e.g. `grep` → Search, `git_operations` → Diff, `run_tests` → Log) → cheap structural heuristics (JSON → Diff → HTML → Search → Code → Log → PlainText). No regex on the hot path.
3. **Compressor selection.** Each kind routes to a dedicated compressor, honoring per-kind toggles (`search_enabled`, `code_enabled`, `html_enabled`, `ml_compression_enabled`).
4. **Compression.** The compressor runs. If it declines or its output is no smaller than the input, TokenJuice falls back to the generic compressor or passes the original through — it never makes things bigger.
5. **CCR offload.** For **lossy** compressions where the original is large enough (`ccr_min_tokens`, default ~500 tokens), the full original is stowed in the **Compress-Cache-Retrieve** store so nothing is permanently lost.
6. **Recovery marker.** A footer carrying the canonical marker `⟦tj:<hash>⟧` is appended, telling the agent it's looking at a partial view and how to fetch the rest.
7. **Savings accounting.** Tokens saved and estimated cost saved are recorded, attributed by model and by compressor.
***
## The compressors
Each content kind has a purpose-built compressor (`src/openhuman/tokenjuice/compressors/`):
| Compressor | Kind | What it does |
| ---------------- | ----------- | -------------------------------------------------------------------------------------------------------- |
| **SmartCrusher** | JSON | Re-renders arrays of objects as a compact table; past ~40 rows keeps head + tail + error rows + numeric outliers. |
| **Code** | Code | Keeps signatures and imports, collapses deep function bodies to `{ … N lines … }` (tree-sitter when available, brace-depth heuristic otherwise). Preserves `TODO`/`FIXME`/`error`/`panic`/`unsafe` markers. |
| **Log** | Log | For **command output**, delegates to the JSON rule engine (below). For other logs, keeps errors / warnings / stack traces / summaries and drops the noise. |
| **Search** | Search | Groups grep/ripgrep `path:line:body` hits by file, ranks by query-term density, keeps the top matches per file, and tallies `[+N more]`. |
| **Diff** | Diff | Keeps changed lines and hunk headers, collapses long unchanged runs to an anchor; lockfile hunks shrink to a one-line `+A/-B` summary. |
| **Html** | HTML | Strips markup to readable text with sensible block-boundary newlines and entity decoding — allocation-light, no DOM. |
| **MlText** | PlainText | Opt-in ML salience compression (see below). |
| **Generic** | fallback | Head/tail summariser for command output that no specific rule matched; declines on structured blobs so they're preserved. |
Multi-byte text (CJK, emoji, combining marks) is handled grapheme-by-grapheme throughout — never split mid-character.
***
## ML compression (opt-in)
Beyond the deterministic compressors, TokenJuice can route plain text through a **ModernBERT** token-salience model that scores and drops low-information spans (`src/openhuman/tokenjuice/ml/`).
* **Off by default.** Enable with `ml_compression_enabled = true` in `[tokenjuice]`.
* **Runs locally** as the `kompress` backend of the shared Python runtime sidecar — no data leaves your machine.
* **Tunable:** `ml_model_id` (default `answerdotai/ModernBERT-base`), `ml_target_ratio` (default `0.5`), `ml_max_input_chars` (default `200000`), `ml_device` (`cpu`/`auto`), `ml_sidecar_idle_timeout_secs`.
* **Graceful:** if the sidecar is unavailable or an input exceeds the char cap, it degrades to the native compressors without ever failing the agent loop.
***
## Nothing is lost: CCR cache & retrieval
Lossy compression would normally mean throwing data away. TokenJuice instead **offloads** the full original into the **Compress-Cache-Retrieve (CCR)** store and leaves a breadcrumb (`src/openhuman/tokenjuice/cache/`).
* **In-memory tier** (always on): a process-global store keyed by SHA-256 hash, bounded by entry count (`max_cache_entries`, default 256) and total bytes (`max_cache_bytes`, default 64 MiB), FIFO eviction.
* **On-disk tier** (optional): `<workspace>/.tokenjuice/ccr/`, enabled with `ccr_disk_enabled`, survives memory eviction. Optional TTL via `ccr_ttl_secs`.
* **The marker:** compacted output ends with a footer like `[compacted tool output — PARTIAL view; full original available via tokenjuice_retrieve with token "…"]` carrying the `⟦tj:<hash>⟧` token.
* **Retrieval tool:** the agent calls the read-only **`tokenjuice_retrieve`** tool with that token (optionally a byte/line `range`) to pull back the full original or a slice. The token is an unguessable SHA-256 digest.
So the agent gets the cheap compacted view by default, and can transparently "zoom in" on the full text only when it actually needs it.
***
## Savings tracking
Every compression is metered (`src/openhuman/tokenjuice/savings.rs`). TokenJuice tracks events, original vs. compacted tokens, tokens saved, and **estimated cost saved in USD** (using per-model input pricing) — aggregated `total`, `by_model`, and `by_compressor`. Stats persist to `<workspace>/state/tokenjuice_savings.json` and survive restarts.
Read them over RPC with `tokenjuice.savings_stats`; clear them with `tokenjuice.savings_reset`.
***
## The rule overlay (command & log output)
The original three-layer JSON rule overlay still powers the Log/command compressor. Rules merge in order, later layers overriding earlier ones:
| Layer | Path | Purpose |
| ------------ | ----------------------------- | ------------------------------------------------------------- |
| **Builtin** | shipped with the binary | ~96 vendored rules for git, npm, cargo, docker, kubectl, ls… |
| **User** | `~/.config/tokenjuice/rules/` | personal overrides, apply everywhere |
| **Project** | `.tokenjuice/rules/` | repo-specific overrides, checked in and shared with the team |
Each rule names a command/tool pattern and a reduction strategy (skip/keep filters, transforms like strip-ANSI and dedupe, head/tail summarize, named counters, canned messages). Rules are JSON — add one and it applies with no recompile.
***
## Configuration, RPC & tools
Everything lives under the `[tokenjuice]` config block (`src/openhuman/config/schema/tokenjuice.rs`) and can be changed live.
* **Master switch:** `router_enabled` (default `true`).
* **Thresholds:** `min_bytes_to_compress`, `ccr_min_tokens`.
* **CCR:** `ccr_enabled`, `ccr_disk_enabled`, `max_cache_entries`, `max_cache_bytes`, `ccr_ttl_secs`.
* **Per-kind:** `search_enabled`, `code_enabled`, `html_enabled`, plus the `ml_*` keys.
* **RPC** (`tokenjuice.*`): `detect`, `compress` (dry-run the pipeline), `settings_get` / `settings_update` (live partial patch), `cache_stats`, `retrieve`, `savings_stats`, `savings_reset`.
* **Agent tool:** `tokenjuice_retrieve` (read-only) recovers offloaded originals.
* **Debugging:** start the core with `RUST_LOG=openhuman_core::openhuman::tokenjuice=debug` to watch detection, matching, and how much each blob is trimmed.
***
## Why this matters for memory
TokenJuice is what makes [auto-fetch](obsidian-wiki/auto-fetch.md) economically viable. When the Gmail provider syncs a page of 200 messages, TokenJuice compacts each canonicalized email _before_ it enters the model that builds summaries. The same applies to GitHub diffs, Slack channel dumps, and any other firehose source.
TokenJuice is what makes [auto-fetch](obsidian-wiki/auto-fetch.md) economically viable. When the Gmail provider syncs a page of 200 messages, TokenJuice compacts each canonicalized email _before_ it enters the model that builds [Memory Tree](obsidian-wiki/memory-tree.md) summaries. The same applies to GitHub diffs, Slack dumps, and any other firehose source. Concretely: ingesting six months of email through a frontier model costs single-digit dollars instead of hundreds.
Concretely: ingesting your last six months of email through a frontier model costs single-digit dollars instead of hundreds.
## Where it lives in the pipeline
```
tool call result
TokenJuice (classify → match rule → reduce)
LLM context
```
Implementation: `src/openhuman/tokenjuice/` (`classify.rs`, `reduce.rs`, `rules/compiler.rs`, `tool_integration.rs`).
## Inspecting and overriding
* Drop a JSON file in `~/.config/tokenjuice/rules/` to add or override a rule globally.
* Drop one in `.tokenjuice/rules/` inside a repo to do the same per-project.
* Start the core with `RUST_LOG=openhuman_core::openhuman::tokenjuice=debug` to see what's matching and how much output is being trimmed.
***
## See also
* [Native Tools](native-tools/README.md). most heavy tool output flows through TokenJuice.
* [Memory Tree](obsidian-wiki/memory-tree.md). the downstream consumer of compressed output.
* [Available Tools](native-tools/README.md) most heavy tool output flows through TokenJuice.
* [Memory Tree](obsidian-wiki/memory-tree.md) the downstream consumer of compressed output.
* [Billing, Cost & Usage](billing-and-usage.md) — where token savings show up as real money.
@@ -1,51 +0,0 @@
---
description: >-
TokenJuice - 一层规则叠加,在工具输出进入 LLM 上下文之前将其压缩。
处理成千上万封邮件依然成本低廉。
icon: file-zipper
---
# 智能 Token 压缩
LLM Token 价格不菲,而冗长的工具输出是消耗大多数 Token 的地方。繁忙仓库里的 `git status`、一次 `cargo build` 日志、一个 600 条消息的邮件串,或者针对真实集群的 `docker ps -a`,这些都可能把上下文窗口撑得很大,却几乎不带多少有效信息。
OpenHuman 搭载 **TokenJuice**,这是 [vincentkoc/tokenjuice](https://github.com/vincentkoc/tokenjuice) 的移植版本,直接集成到工具执行路径中。在任何工具结果到达模型之前,TokenJuice 会将输出通过一层规则叠加进行处理,去除噪音、保留信号。
## 三层规则叠加
规则是 JSON,按以下顺序合并,后面的层级覆盖前面的:
<table><thead><tr><th width="134.41796875">层级</th><th>路径</th><th>用途</th></tr></thead><tbody><tr><td><strong>内置</strong></td><td>随二进制文件发布</td><td>为 git、npm、cargo、docker、kubectl、ls 等提供的合理默认值</td></tr><tr><td><strong>用户</strong></td><td><code>~/.config/tokenjuice/rules/</code></td><td>你的个人覆盖,应用于所有项目</td></tr><tr><td><strong>项目</strong></td><td><code>.tokenjuice/rules/</code></td><td>仓库特定的覆盖,纳入版本控制,与团队共享</td></tr></tbody></table>
每条规则命名一个工具/命令模式和一个压缩策略(截断、行去重、折叠空白、删除匹配的正则表达式、摘要分段等)。新规则就是 JSON 文件,无需重新编译。
## 为什么这和记忆有关
TokenJuice 是使[自动拉取](obsidian-wiki/auto-fetch.zh-CN.md)在经济上可行的原因。当 Gmail provider 同步一页 200 条消息时,TokenJuice 在每个规范化的邮件进入构建摘要的模型**之前**就将其压缩。GitHub diff、Slack 频道转储以及其他任何高流量来源同理。
具体来说:通过前沿模型摄入你最近六个月的邮件费用从数百美元降到个位数美元。
## 它在流水线中的位置
```text
工具调用结果
TokenJuice(分类 → 匹配规则 → 压缩)
LLM 上下文
```
实现:`src/openhuman/tokenjuice/``classify.rs``reduce.rs``rules/compiler.rs``tool_integration.rs`)。
## 检查和覆盖
*`~/.config/tokenjuice/rules/` 中放入一个 JSON 文件来全局添加或覆写规则。
* 在仓库内的 `.tokenjuice/rules/` 中放入一个来做同样的项目级设置。
* 使用 `RUST_LOG=openhuman_core::openhuman::tokenjuice=debug` 启动 core,可以查看匹配了什么以及多少输出被裁剪了。
## 另见
* [原生工具](native-tools/README.zh-CN.md)。大多数重型工具输出都经过 TokenJuice。
* [记忆树](obsidian-wiki/memory-tree.zh-CN.md)。压缩输出的下游消费者。
+114
View File
@@ -0,0 +1,114 @@
---
description: >-
Local, non-custodial multi-chain crypto wallet the agent can read balances
from and send transfers with — keys stay in-core and never cross the wire.
icon: wallet
---
# Wallet
A deliberately **basic**, **non-custodial** multi-chain crypto wallet owned by the Rust core. It manages one account per supported chain derived from a single recovery phrase, reads balances, and runs a strict **prepare → confirm → execute** flow for native sends and a small set of token transfers.
It is intentionally minimal: key/account management plus the primitive on-chain operations. Higher-level DeFi (swaps, bridges, generic contract/dapp calls) lives in a separate `web3` module and is **not** part of the wallet's agent or RPC surface.
The most important property to understand: **signing and broadcast happen entirely in-core from the decrypted recovery phrase. No private keys ever leave the device or cross the network.** This is your money — the wallet is conservative by design.
***
## Supported chains and token standards
Setup derives **exactly one account per chain**. EVM is a single account reused across six networks. Only the standards listed below are supported for transfers — anything else (swaps, arbitrary contract calls) is out of scope for the wallet.
| Chain | Networks | Native | Token standard | Notes |
| --- | --- | --- | --- | --- |
| EVM | Ethereum, Base, Arbitrum, Optimism, Polygon, BNB Chain | ETH / BNB / etc. | ERC-20 (BEP-20 on BNB Chain) | One `Evm` account across all six; network selected per request, defaults to Ethereum mainnet. EIP-1559 / typed-tx signing. |
| Bitcoin | Mainnet | BTC | — | P2WPKH (native SegWit). **Rejects token transfers.** Esplora REST for balance/broadcast. |
| Solana | Mainnet / devnet (per RPC) | SOL | SPL | ed25519 signing; native + SPL token transfers. |
| Tron | Mainnet | TRX | TRC-20 | TronGrid REST for native + TRC-20 transfers. |
Built-in asset catalogs include the native asset plus common stablecoins per chain (e.g. USDC/USDT as ERC-20/BEP-20, USDC as SPL on Solana, USDT as TRC-20 on Tron).
***
## Onboarding and the recovery phrase
Setup is a single, all-or-nothing operation: it persists a consent flag, the mnemonic word count, the setup source, exactly one derived account per supported chain, and the encrypted recovery phrase. Valid BIP-39 mnemonic word counts are **12, 15, 18, 21, or 24**.
The recovery phrase is the only secret. The wallet stores the per-chain account **addresses** (safe to surface) separately from the secret material, and only the encrypted phrase can reconstruct private keys.
***
## Key custody and security
The wallet is **non-custodial and local** — there is no server-side key escrow.
- **The recovery phrase is always encrypted at rest.** It is encrypted via the core `encryption` domain before being persisted anywhere.
- **Preferred home: the OS keychain.** The encrypted phrase lives in the operating system keychain under the key `wallet.mnemonic`, scoped by a workspace-derived user id. Access is gated by the keyring consent policy.
- **Fallback: workspace JSON.** When the keychain is unavailable (e.g. headless), the encrypted phrase falls back to `{workspace_dir}/state/wallet-state.json`. On load, any JSON-resident secret is transparently **migrated into the keychain** when one becomes available, and stripped from the JSON.
- **Atomic, guarded writes.** `wallet-state.json` is written atomically (temp file + fsync + persist) under a process-wide lock. Corrupt or invalid state files are quarantined rather than trusted.
- **Decrypt only at signing time.** Chain signers decrypt the phrase in-core only when deriving a key to sign a confirmed transaction. Plaintext keys are never persisted and never serialized over the wire.
See [OS keyring & secret storage](os-keyring-and-secret-storage.md) for how secrets are stored across platforms, and [Privacy & security](privacy-and-security.md) for the broader model.
***
## Reading balances and chain info
Read-only surfaces require no confirmation:
- **Status** — onboarding state plus the safe per-chain account addresses.
- **Balances** — native-asset balances per account. Note: only **EVM balances read live** today (Ethereum mainnet); BTC, Solana, and Tron call their providers but fall back to a zero balance with a "provider missing" status on error.
- **Network defaults / supported assets** — per-chain RPC and explorer URLs, capability flags, and the built-in asset catalog.
- **Chain status** — per-chain readiness and the active RPC URL.
RPC endpoints are overridable per chain/network via `OPENHUMAN_WALLET_RPC_*` environment variables. URLs are redacted to scheme + host in logs.
***
## Sending transfers: prepare → confirm → execute
Every write is a two-step, intentional flow — the wallet never sends in one shot.
1. **Prepare** (`prepare_transfer`) — validates the amount, destination address, and (for tokens) calldata, estimates fees, and returns a **prepared quote** with a `quoteId`. Quotes are held in an in-memory store with a **5-minute TTL**, capped at 64, and are **not** persisted across restarts.
2. **Confirm + execute** (`execute_prepared`) — requires `confirmed: true` and a valid `quoteId`. The quote is consumed atomically before broadcast so concurrent confirmations can't double-submit; on failure it's restored with a refreshed TTL so it stays retryable.
**Quote-owner binding.** Each quote is bound to the chat thread that prepared it. A quote can only be executed by the same owner that prepared it; a `quoteId` leaked into a shared channel returns an indistinguishable "not found" error rather than letting another session hijack it.
Transfers are limited to **native sends and the token standards in the table above**. Bitcoin rejects token transfers. Swaps, bridges, and generic contract calls are not available here.
***
## Transaction status tracking
After broadcast, three read-only inspectors let the agent follow a transaction by hash:
- **`tx_status`** — lifecycle state: pending / confirmed / failed / not found.
- **`tx_receipt`** — receipt details: success, fee, block.
- **`lookup_tx`** — the raw transaction payload.
***
## Agent tools and approval safety
The agent reaches the wallet through six tools:
| Tool | Purpose |
| --- | --- |
| `wallet_status` | Onboarding status + account addresses. |
| `wallet_chain_status` | Per-chain readiness + active RPC. |
| `wallet_prepare_transfer` | Build a validated, fee-estimated quote. |
| `wallet_tx_status` | Transaction lifecycle state by hash. |
| `wallet_tx_receipt` | Transaction receipt by hash. |
| `wallet_lookup_tx` | Raw transaction lookup by hash. |
Notice there is **no agent tool that executes a transfer.** The agent can prepare a quote, but actually moving funds (`execute_prepared`) goes through the RPC surface, where it must be explicitly confirmed and pass the owner-binding check. Combined with the prepare-then-confirm flow and the per-thread quote binding, this keeps the agent from silently spending funds.
Because these are financial actions, they should be surfaced through the [approval gate](approval-gate.md) so a human confirms before money moves. Treat every transfer as a high-stakes action.
***
## See also
- [Approval gate](approval-gate.md)
- [Privacy & security](privacy-and-security.md)
- [OS keyring & secret storage](os-keyring-and-secret-storage.md)
-75
View File
@@ -1,75 +0,0 @@
---
description: >-
OpenHuman 如何在使用服务时收集、使用、处理、存储和保护信息。
icon: key
---
# 隐私政策
最后更新日期:2026/02/02
本隐私政策描述了我们如何在您使用我们的系统级 AI 助手(以下简称"服务")时收集、使用、处理、存储和保护信息。该服务旨在作为用户设备(如笔记本电脑或台式电脑)上的通用辅助智能体运行。我们致力于保护用户隐私,并最小化数据收集和保留。
本隐私政策旨在遵守适用的全球数据保护法律,包括欧盟《通用数据保护条例》(GDPR)、印度《数字个人数据保护法》(DPDP)、《加州消费者隐私法》及《隐私权法》(CCPA/CPRA)、巴西《通用数据保护法》(LGPD)以及加拿大《个人信息保护和电子文件法》(PIPEDA)。
## 我们收集的信息
我们仅在提供服务所必需的范围内收集和处理信息,且仅响应用户的明确操作。
1.1 用户内容与系统数据 当您使用服务时,它仅在您明确指示的情况下处理您设备上的文件、应用数据、系统上下文或其他信息,且仅处理完成所请求任务所需的范围。
1.2 最小化用户元数据 我们收集有限的用户元数据,用于基本账户或服务功能。仅限于:• 名字 • 姓氏 • 用户提供的个人资料信息(如适用,例如个人简介)
除非用户请求的任务明确要求且允许,否则我们不会收集电话号码、联系人列表、精确位置数据、设备标识符、浏览历史、行为分析或后台系统活动。
## 我们如何使用信息
我们仅将信息用于运营、维护并提供服务,包括:• 执行用户明确请求的任务
• 跨文件、应用或工作流提供上下文辅助
• 提高服务的可靠性和安全性
我们不会将个人数据用于广告、营销、画像或行为追踪。我们不会出售、出租或交易个人数据。用户数据不会被用于训练共享的、公开的或第三方的 AI 模型。
## 数据保留与删除
服务采用默认零保留设计。
• 系统数据和内容被临时处理
• 不长期保留系统活动、文件或操作的日志
• 完成任务所需的临时数据在任务完成后立即删除,或在最长 30 天内删除
用户可随时请求删除其数据。收到此类请求后,所有关联数据将被永久删除且无法恢复。撤销权限或卸载服务将立即停止进一步处理。
## 处理的法律依据
在法律要求的范围内,我们基于以下一项或多项法律依据处理个人数据:
• 用户同意
• 履行合同
• 合法权益,在适用情况下并与用户权利相平衡
## 安全措施
我们实施合理且适当的技术和组织措施,旨在保护信息免受未经授权的访问、丢失、滥用或篡改。这些措施包括传输中和静止状态的加密、最小权限访问控制、隔离处理环境以及内部监控和审计机制。人员对用户数据的访问仅限于运营、安全或法律必要性。
## 服务提供商
我们可能会聘请第三方服务提供商来支持服务的运营,例如基础设施托管或 AI 推理提供商。这些提供商仅代表我们行事,受保密义务约束,且被限制将数据用于自身目的。
## 国际数据传输
信息可能会在您的居住国以外的地点进行处理。在适用情况下,我们实施适当的保障措施,以确保符合适用的数据保护法律的充分保护。
## 您的权利
根据您的所在地,您可能有权访问、更正、删除、限制处理或撤回同意您的个人数据,以及请求数据可携带性。请求可随时在仪表板上完成,或发送至 privacy@tinyhumans.ai。
## 本政策的变更
我们可能会不时更新本隐私政策。重大变更将按照法律要求进行通知。变更不会以降低现有隐私保护的方式追溯适用。
-60
View File
@@ -1,60 +0,0 @@
---
description: 使用 OpenHuman 服务的条款和条件。
icon: file-contract
---
# 使用条款
最后更新日期:2026/02/02
本使用条款(以下简称"条款")管辖您对服务的使用。通过安装、访问或使用服务,您同意受这些条款的约束。如果您不同意,则不得使用服务。
## 服务
服务是一款系统级 AI 助手,旨在帮助用户完成任务、自动化工作流,并基于明确的用户指令与文件、应用和系统资源进行交互。服务仅根据用户请求行事,不会自主运行。
## 用户责任
您同意:
• 在遵守适用法律和法规的前提下使用服务
• 仅对您有权访问的内容和系统使用服务
• 不将服务用于监视、骚扰或非法活动
• 不尝试逆向工程、干扰或滥用服务
您有责任在根据任何输出采取行动之前对其进行审查和验证。
## 内容与数据所有权
您保留对您的内容、文件和数据的所有权。我们不声称拥有用户内容或 AI 生成输出的所有权。我们对数据的访问是有限的、基于权限的,且仅为提供服务之目的。
## AI 输出免责声明
服务使用以概率方式生成输出的人工智能系统。输出可能不准确、不完整或过时。服务不保证任何目的的准确性、可靠性或适用性,也不替代专业判断。您对您如何使用 AI 生成的输出负有全部责任。
## 可用性与修改
我们可能会随时修改、暂停或终止服务或其任何部分,包括改进功能、解决安全问题或遵守法律要求。
## 责任限制
在法律允许的最大范围内,服务按"原样"和"可用"基础提供。我们否认所有明示或默示的保证。我们对间接、附带、后果性、特殊或惩罚性损害不承担责任,包括数据丢失、利润损失或依赖 AI 生成输出造成的损失。我们对第三方软件、操作系统或服务导致的中断或故障不承担责任。
## 赔偿
您同意赔偿并使服务及其关联方免受因您使用服务、违反这些条款或侵犯第三方权利而引起的索赔、损害、损失或费用的损害。
## 终止
您可以随时停止使用服务。如果您违反这些条款或法律要求,我们可能会暂停或终止访问。终止后,数据将按照隐私政策处理。
## 适用法律
这些条款受 [插入司法管辖区] 的法律管辖,不考虑法律冲突原则。
## 安装时显示的同意声明
通过安装或使用本 AI 助手,您同意它仅为了执行您明确请求的任务而访问系统信息、文件和应用上下文。助手不会持续监控您的系统,不会在没有指令的情况下行动,也不会在完成任务所需范围之外保留系统数据。您始终控制权限,并可以随时撤销访问或删除您的数据。您的数据不会被用于广告或训练共享的 AI 模型。
@@ -1,78 +0,0 @@
---
description: >-
安装 OpenHuman,完成应用内入门引导(登录、连接 Gmail、
选择 AI 运行方式),然后对你的记忆树发出第一个请求。
icon: play
---
# 快速入门
本文将引导你完成安装 OpenHuman、完成应用内入门引导,以及发出第一个请求。
OpenHuman 遵循 GNU GPL3 开源许可证,代码库位于 [github.com/tinyhumansai/openhuman](https://github.com/tinyhumansai/openhuman)。
***
## 系统要求
OpenHuman 支持 **macOS、Windows 和 Linux** 桌面端。建议 4 GB 以上内存;如果要摄入超大型邮箱或仓库,或在同一台机器上运行[本地模型](../features/model-routing/local-ai.zh-CN.md),建议 16 GB 以上。
### 权限
首次启动 OpenHuman 时,操作系统会提示授予应用所需的权限(macOS 上的 Accessibility、语音热键的 Input Monitoring,以及计划使用[会议智能体](../features/mascot/meeting-agents.zh-CN.md)时的相机/麦克风)。你随时可以在 **设置 → 自动化与渠道** 中查看和调整这些权限。
***
## 1. 下载并安装
从 [https://tinyhumans.ai/openhuman](https://tinyhumans.ai/openhuman) 或通过你平台的软件包管理器获取 OpenHuman 桌面应用。安装后打开应用。
## 2. 登录
第一个屏幕是**"登录!让我们开始吧"**。提供多种登录方式,包括社交登录。如果你要将应用指向自定义 core RPC URL(自建后端的情况),还有一个**高级**面板;大多数用户可以忽略它。
{% hint style="info" %}
**无永久锁定。** 登录不会授予 OpenHuman 对任何内容的持续访问权。所有第三方访问都需要在以下步骤中每个集成单独进行明确的 OAuth 批准。
{% endhint %}
## 3. 发出你的第一个请求
一旦 Gmail 完成摄入(首次自动拉取会在二十分钟内触发),可以尝试以下提示:
**简报**
* "过去 12 小时我需要了解什么?"
* "有什么在等着我?"
**跨源查询**
* "总结我今天错过了什么。"
* "这周有哪些关键决策?"
* "从我最近的对话中提取行动项。"
* "Sarah 在邮件和聊天中对这个项目说了什么?"
OpenHuman 自动为每个任务选择合适的模型。参见[自动模型路由](../features/model-routing/)。
***
## 4. 打开 Obsidian 存储库
"记忆"标签页有一个**"在 Obsidian 中查看存储库"**按钮。点击它可以在 [Obsidian](https://obsidian.md) 中打开 `<workspace>/wiki/`。你可以浏览智能体的摘要、放入你自己的笔记,甚至构建手动链接——智能体会在下一次摄入时获取你的编辑。参见 [Obsidian 风格的记忆](../features/obsidian-wiki/)。
***
## 5. 让吉祥物做更多
现在智能体有了记忆和一个模型,产品的其余部分就是给它更多发挥空间:
* [**会议智能体**](../features/mascot/meeting-agents.zh-CN.md) —— 放入一个 Google Meet 链接,吉祥物作为真实参与者加入:它倾听、将笔记记入记忆树、在通话中说话,并实时使用工具。
* [**从集成自动拉取**](../features/obsidian-wiki/auto-fetch.zh-CN.md) —— 从**设置**中连接更多源;每二十分钟调度器将新数据拉入你的树。
* [**原生语音**](../features/native-tools/voice.zh-CN.md) —— 按键说话输入和 TTS 回复,这样你可以和 OpenHuman 对话而不是打字。
* [**潜意识循环**](../features/subconscious.zh-CN.md) —— 让你离开时吉祥物继续处理待办任务。
## 加入社区
OpenHuman 处于早期测试阶段。在这个阶段,反馈和贡献能带来真正的改变。
* **GitHub** [github.com/tinyhumansai/openhuman](https://github.com/tinyhumansai/openhuman)
* **Discord** [discord.tinyhumans.ai](https://discord.tinyhumans.ai)
@@ -1,62 +0,0 @@
---
description: >-
诊断登录失败、OAuth 回调无法完成以及远程核心 RPC 认证问题。
icon: key
lang: zh-CN
---
# 登录故障排查
当社交登录卡住、返回欢迎界面,或核心日志中出现未授权的 `/auth` 请求时,使用此 checklist。
## 检查后端可达性
从桌面应用所在的同一网络,验证公共 OpenHuman 端点:
```bash
curl -I https://tinyhumans.ai/
curl -I https://api.tinyhumans.ai/health
```
如果网站能加载但 API 端点失败,桌面应用可能无法将 OAuth 回调兑换为 session。在 issue 报告中记录 HTTP 状态码、区域和 DNS 结果。
## 检查所选核心
如果你使用**高级**远程核心模式,在开始 OAuth 之前确认 RPC URL 和 bearer token
```bash
curl -sS https://your-core.example/rpc \
-H "Content-Type: application/json" \
-H "Authorization: Bearer CORE_TOKEN" \
-d '{"jsonrpc":"2.0","id":1,"method":"core.ping","params":{}}'
```
`401` 响应表示桌面 token 与远程核心 token 不匹配。在重试 Google 或 GitHub 登录之前先修复这个问题。
## 检查深度链接回调
成功的桌面 OAuth 以 `openhuman://auth?...` 回调结束。如果浏览器显示了该 URL 但应用仍停留在欢迎界面:
1. 确保只运行了一个 OpenHuman 桌面实例。
2. 重启应用,保持相同的远程核心设置,并重试登录。
3. 如果使用远程核心,检查核心是否收到 `openhuman.auth_store_session`
对于远程核心,临时手动注入可以确认核心本身是正常的:
```bash
curl -sS https://your-core.example/rpc \
-H "Content-Type: application/json" \
-H "Authorization: Bearer CORE_TOKEN" \
-d '{"jsonrpc":"2.0","id":1,"method":"openhuman.auth_store_session","params":{"token":"JWT_FROM_CALLBACK"} }'
```
不要把真实 JWT 粘贴到公共 GitHub issue 中。对 token 进行脱敏处理,只附加状态码、主机名、应用版本、操作系统和相关日志行。
## Bug 报告中应包含的内容
* 应用版本和操作系统。
* 核心模式是本地还是远程。
* RPC URL 主机、脱敏后的 token 状态和 `core.ping` 结果。
* 使用的 OAuth 提供商。
* 浏览器中是否出现了 `openhuman://auth` URL。
* 如果存在,第一条未授权日志行。