feat(agent): sub-agents, reasoning→agentic routing, layered context pipeline (#474)

* Enhance agent architecture with sub-agent support and memory optimizations

- Refactored the `Agent` struct to use `Arc` for shared ownership of the provider, tools, and tool specifications, enabling efficient memory management and concurrent access.
- Introduced a `NullMemoryLoader` to optimize memory usage for sub-agents, allowing them to operate without incurring the cost of memory recall.
- Added new methods in the `Agent` implementation to facilitate sharing of the provider, tools, and tool specifications with sub-agents, enhancing their operational efficiency.
- Implemented a new `SystemPromptBuilder` method for constructing prompts specifically for sub-agents, ensuring they receive tailored context while minimizing unnecessary information.
- Established a framework for loading custom agent definitions from TOML files, allowing for dynamic agent configuration and specialization.
- Introduced a `ForkContext` to support efficient sub-agent execution in fork mode, leveraging shared resources for improved performance and reduced token usage.

* Enhance agent definition management and sub-agent functionality

- Introduced a global `AgentDefinitionRegistry` to manage built-in and custom agent definitions from TOML files, ensuring idempotent initialization.
- Added new RPC handlers for listing, fetching, and reloading agent definitions, improving the flexibility of agent management.
- Refactored the `Agent` struct to streamline sub-agent execution, including enhancements to the task execution flow and context handling.
- Updated the orchestrator configuration to support fork mode for sub-agents, optimizing resource usage and performance.
- Improved error handling and logging for agent definition loading and initialization processes, enhancing system reliability.

* Add end-to-end test for sub-agent spawning and response integration

- Implemented a new asynchronous test to validate the full path of a parent agent issuing a `spawn_subagent` tool call.
- The test ensures that the sub-agent's output is correctly folded into the parent's response, verifying the interaction between the parent agent and the sub-agent.
- Enhanced the `AgentDefinitionRegistry` to support global initialization of built-in agents, ensuring consistent behavior across tests.
- Updated the handling of tool calls and memory configuration to facilitate the new test scenario, improving overall test coverage for agent interactions.

* Enhance agent tool filtering with category support

- Introduced a new `category_filter` in `AgentDefinition` to restrict tool visibility based on their category (System or Skill).
- Updated the `from_archetype` function to apply the category filter for the `SkillsAgent` archetype.
- Modified `SubagentRunOptions` to include a `category_filter_override` for dynamic filtering during sub-agent execution.
- Enhanced the `filter_tool_indices` function to incorporate category filtering logic, ensuring tools are correctly filtered based on their defined categories.
- Updated relevant tests to validate the new category filtering functionality, improving overall test coverage for agent interactions.

* Implement layered context reduction pipeline for agent

- Introduced a new `context_pipeline` module to manage a layered context reduction strategy, enhancing memory efficiency during agent interactions.
- Added stages for tool-result budgeting, history trimming, microcompaction, autocompaction, and session memory extraction, each with specific triggers and cache implications.
- Updated the `Agent` struct to include a `context_pipeline` field, ensuring state persistence across turns.
- Enhanced the `AgentBuilder` to initialize the context pipeline by default.
- Implemented tests to validate the functionality and stability of the context pipeline, ensuring consistent behavior across agent sessions.
- Refactored relevant components to integrate the new context management features, improving overall agent performance and memory handling.

* Enhance agent context pipeline with tool result budgeting and microcompaction

- Renamed variable for clarity in tool execution result handling.
- Implemented a new stage in the context pipeline to apply a byte budget to tool results, ensuring efficient memory usage.
- Added logging for budget application, including details on original and final byte sizes.
- Integrated microcompaction stages before tool calls to manage history and reduce memory footprint, with appropriate logging for outcomes.
- Updated the agent's session memory management to track turn counts, facilitating better resource handling across iterations.

* Refactor agent context pipeline for session memory extraction and tool call management

- Simplified method calls in the `Agent` struct for clarity and efficiency.
- Enhanced session memory extraction logic to spawn a background archivist sub-agent when thresholds are met.
- Improved context pipeline handling for tool call recording and usage tracking.
- Updated documentation and comments for better understanding of session memory extraction process.
- Refactored microcompact function for cleaner code structure and readability.

* refactor(agent): split agent.rs into focused submodules

- Convert agent.rs (1988 lines) into agent/ folder with six files:
  * types.rs — Agent + AgentBuilder struct defs
  * builder.rs — AgentBuilder fluent API + Agent::from_config factory
  * turn.rs — turn lifecycle, tool dispatch, context pipeline wiring
  * runtime.rs — public accessors, run_single/run_interactive, helpers
  * tests.rs — integration tests with shared fakes
  * mod.rs — glue + top-level `run` convenience function

- Drop misc external inspiration references from doc comments in the
  context_pipeline module and fork_context — the files now stand on
  their own design language.

* fix(agent): address review comments on sub-agent + context pipeline PR

Inline comment fixes:
- definition.rs: YAML/TOML inconsistency — module doc, PromptSource, source
  bookkeeping, and load() all now uniformly document the TOML format.
- subagent_runner.rs: render_subagent_system_prompt previously only appended
  PromptSource::Inline bodies, silently dropping PromptSource::File content.
  Thread the preloaded archetype_body through as an explicit &str parameter
  so both source variants render. Drops the unused SystemPromptBuilder +
  tools_for_prompt wiring while we're here.
- prompt.rs: remove DateTimeSection from SystemPromptBuilder::for_subagent.
  Local::now() would make the sub-agent system prompt change per call and
  break KV-prefix cache stability. Document the invariant.

Nitpick fixes:
- agent/turn.rs: drop the no-op mark_extraction_started() call — the
  immediately-following mark_extraction_complete() clears the in-flight
  flag anyway.
- context_pipeline/pipeline.rs: call guard.record_compaction_success() on
  microcompact success so a prior streak of autocompaction failures
  doesn't leave the circuit breaker tripped after a successful reduction.
- context_pipeline/tool_result_budget.rs: remove the unnecessary out.clone()
  in the truncation return — capture final_bytes first, then move out.
- definition_loader.rs: replace brittle reg.len() == 10 with
  assert!(reg.len() > 1) plus the existing targeted .get() checks.
- executor.rs: tracing::warn! on unknown sandbox override values so typos
  surface during development; explicitly accept "none" and empty string
  as valid defaults.
- fork_context.rs: add parent_context_visible_inside_scope test mirroring
  fork_context_visible_inside_scope, with minimal stub Provider/Memory
  impls so the test stays self-contained.
- schemas.rs: drop redundant serde_json::to_value(serde_json::json!(...))
  wrapping in handle_list_definitions + handle_get_definition.
- event_bus/events.rs: add SubagentSpawned/SubagentCompleted/SubagentFailed
  cases to all_variants_have_correct_domain.
- tools/ops.rs: add all_tools_includes_spawn_subagent regression test.
- tools/spawn_subagent.rs: sort/dedup the known skill list in place
  instead of cloning into a second vec.

Tests: 2316 passed / 0 failed (up from 2314; two new tests added).
fmt + clippy clean on all touched files.

* udpate prompts

* Enhance AgentBuilder and runtime with event context and interactive CLI improvements

- Added `event_context` method to `AgentBuilder` for setting `session_id` and `channel` for `DomainEvent`s, improving event tagging and correlation.
- Updated `run_interactive` method in `Agent` to dispatch messages through `run_single`, ensuring consistent lifecycle event handling and error sanitization for interactive turns.

* Optimize configuration handling in AgentBuilder by lazily creating Arc for full config in reflection hook. This change reduces unnecessary cloning when learning is enabled, improving performance.

* Add fork-mode test for sub-agent spawning in agent

- Introduced a new test, `turn_dispatches_spawn_subagent_in_fork_mode`, to validate the behavior of the agent when spawning a sub-agent in fork mode.
- The test ensures that the parent agent correctly processes the sub-agent's output and maintains the expected response sequence.
- Enhanced the test setup with a mock provider and memory configuration to simulate the agent's environment effectively.

* Refactor sync RPC handling in skills to treat missing onSync as no-op

- Updated the `handle_sync` function to log a debug message and return a no-op response when a skill does not implement the `onSync` handler, preventing unnecessary RPC errors for skills that do not require periodic syncs.
- This change improves logging clarity and reduces error noise in logs and dashboards for skills that are not designed to handle sync operations.
This commit is contained in:
Steven Enamakel
2026-04-09 21:09:34 -07:00
committed by GitHub
parent d5822ffcaf
commit 6465f3d314
41 changed files with 7030 additions and 1503 deletions
+13
View File
@@ -836,6 +836,19 @@ pub async fn bootstrap_skill_runtime() {
// from both jsonrpc and repl paths) cannot double-subscribe.
register_domain_subscribers();
// --- Sub-agent definition registry bootstrap ---
// Loads built-in archetype definitions plus any custom TOML files
// under `<workspace>/agents/*.toml`. Idempotent — safe to call from
// both jsonrpc and repl paths.
if let Err(err) = crate::openhuman::agent::harness::AgentDefinitionRegistry::init_global(
&base_dir.join("workspace"),
) {
log::warn!(
"[runtime] AgentDefinitionRegistry::init_global failed: {err} — \
spawn_subagent will be unavailable until restart"
);
}
// --- Socket manager bootstrap ---
let socket_mgr = Arc::new(SocketManager::new());
set_global_socket_manager(socket_mgr.clone());
File diff suppressed because it is too large Load Diff
+431
View File
@@ -0,0 +1,431 @@
//! `AgentBuilder` fluent API and the `Agent::from_config` factory.
//!
//! Everything in this file is about *constructing* an `Agent` — the
//! builder setters, the `build()` validator, and the `from_config()`
//! factory that wires together the real provider / memory / tool
//! registry from a loaded [`Config`]. Per-turn behaviour lives in
//! [`super::turn`]; accessors and run-helpers live in [`super::runtime`].
use super::types::{Agent, AgentBuilder};
use crate::openhuman::agent::context_pipeline::ContextPipeline;
use crate::openhuman::agent::dispatcher::{
NativeToolDispatcher, ToolDispatcher, XmlToolDispatcher,
};
use crate::openhuman::agent::host_runtime;
use crate::openhuman::agent::memory_loader::{DefaultMemoryLoader, MemoryLoader};
use crate::openhuman::agent::prompt::SystemPromptBuilder;
use crate::openhuman::config::Config;
use crate::openhuman::memory::{self, Memory};
use crate::openhuman::providers::{self, Provider};
use crate::openhuman::security::SecurityPolicy;
use crate::openhuman::tools::{self, Tool, ToolSpec};
use anyhow::Result;
use std::sync::Arc;
impl AgentBuilder {
/// Creates a new `AgentBuilder` with default values.
pub fn new() -> Self {
Self {
provider: None,
tools: None,
memory: None,
prompt_builder: None,
tool_dispatcher: None,
memory_loader: None,
config: None,
model_name: None,
temperature: None,
workspace_dir: None,
identity_config: None,
skills: None,
auto_save: None,
classification_config: None,
available_hints: None,
post_turn_hooks: Vec::new(),
learning_enabled: false,
event_session_id: None,
event_channel: None,
}
}
/// Sets the AI provider for the agent.
///
/// Accepts a `Box<dyn Provider>` for backward compatibility but stores
/// the provider as an `Arc` internally so sub-agents spawned from this
/// agent (via `spawn_subagent`) can share the same instance.
pub fn provider(mut self, provider: Box<dyn Provider>) -> Self {
self.provider = Some(Arc::from(provider));
self
}
/// Sets the AI provider from an existing `Arc`. Use this when sharing
/// a provider instance across multiple agents.
pub fn provider_arc(mut self, provider: Arc<dyn Provider>) -> Self {
self.provider = Some(provider);
self
}
/// Sets the available tools for the agent.
pub fn tools(mut self, tools: Vec<Box<dyn Tool>>) -> Self {
self.tools = Some(tools);
self
}
/// Sets the memory system for the agent.
pub fn memory(mut self, memory: Arc<dyn Memory>) -> Self {
self.memory = Some(memory);
self
}
/// Sets the system prompt builder for the agent.
pub fn prompt_builder(mut self, prompt_builder: SystemPromptBuilder) -> Self {
self.prompt_builder = Some(prompt_builder);
self
}
/// Sets the tool dispatcher for the agent.
pub fn tool_dispatcher(mut self, tool_dispatcher: Box<dyn ToolDispatcher>) -> Self {
self.tool_dispatcher = Some(tool_dispatcher);
self
}
/// Sets the memory loader for the agent.
pub fn memory_loader(mut self, memory_loader: Box<dyn MemoryLoader>) -> Self {
self.memory_loader = Some(memory_loader);
self
}
/// Sets the agent configuration.
pub fn config(mut self, config: crate::openhuman::config::AgentConfig) -> Self {
self.config = Some(config);
self
}
/// Sets the model name to use for chat requests.
pub fn model_name(mut self, model_name: String) -> Self {
self.model_name = Some(model_name);
self
}
/// Sets the temperature for chat requests.
pub fn temperature(mut self, temperature: f64) -> Self {
self.temperature = Some(temperature);
self
}
/// Sets the workspace directory for the agent.
pub fn workspace_dir(mut self, workspace_dir: std::path::PathBuf) -> Self {
self.workspace_dir = Some(workspace_dir);
self
}
/// Sets the identity configuration for the agent.
pub fn identity_config(
mut self,
identity_config: crate::openhuman::config::IdentityConfig,
) -> Self {
self.identity_config = Some(identity_config);
self
}
/// Sets the skills available to the agent.
pub fn skills(mut self, skills: Vec<crate::openhuman::skills::Skill>) -> Self {
self.skills = Some(skills);
self
}
/// Enables or disables automatic saving of conversation history to memory.
pub fn auto_save(mut self, auto_save: bool) -> Self {
self.auto_save = Some(auto_save);
self
}
/// Sets the query classification configuration.
pub fn classification_config(
mut self,
classification_config: crate::openhuman::config::QueryClassificationConfig,
) -> Self {
self.classification_config = Some(classification_config);
self
}
/// Sets the available model hints for auto-classification.
pub fn available_hints(mut self, available_hints: Vec<String>) -> Self {
self.available_hints = Some(available_hints);
self
}
/// Sets the post-turn hooks to be executed after each turn.
pub fn post_turn_hooks(
mut self,
hooks: Vec<Arc<dyn crate::openhuman::agent::hooks::PostTurnHook>>,
) -> Self {
self.post_turn_hooks = hooks;
self
}
/// Enables or disables learning features.
pub fn learning_enabled(mut self, enabled: bool) -> Self {
self.learning_enabled = enabled;
self
}
/// Sets the event-bus `session_id` and `channel` used to tag
/// `DomainEvent`s emitted by this agent.
///
/// - `session_id` groups all events for a single user / conversation so
/// downstream subscribers can correlate turns, tool calls, and errors.
/// - `channel` labels the source or stream the events originated from
/// (e.g. `"cli"`, `"telegram"`, `"rpc"`) — useful when multiple front
/// ends share the same subscriber pipeline.
///
/// Both parameters are converted into owned `String`s and stored in
/// `event_session_id` / `event_channel` respectively.
pub fn event_context(
mut self,
session_id: impl Into<String>,
channel: impl Into<String>,
) -> Self {
self.event_session_id = Some(session_id.into());
self.event_channel = Some(channel.into());
self
}
/// Validates the configuration and builds the `Agent` instance.
pub fn build(self) -> Result<Agent> {
let tools = self
.tools
.ok_or_else(|| anyhow::anyhow!("tools are required"))?;
let tool_specs: Vec<ToolSpec> = tools.iter().map(|tool| tool.spec()).collect();
Ok(Agent {
provider: self
.provider
.ok_or_else(|| anyhow::anyhow!("provider is required"))?,
tools: Arc::new(tools),
tool_specs: Arc::new(tool_specs),
memory: self
.memory
.ok_or_else(|| anyhow::anyhow!("memory is required"))?,
prompt_builder: self
.prompt_builder
.unwrap_or_else(SystemPromptBuilder::with_defaults),
tool_dispatcher: self
.tool_dispatcher
.ok_or_else(|| anyhow::anyhow!("tool_dispatcher is required"))?,
memory_loader: self
.memory_loader
.unwrap_or_else(|| Box::new(DefaultMemoryLoader::default())),
config: self.config.unwrap_or_default(),
model_name: self
.model_name
.unwrap_or_else(|| crate::openhuman::config::DEFAULT_MODEL.into()),
temperature: self.temperature.unwrap_or(0.7),
workspace_dir: self
.workspace_dir
.unwrap_or_else(|| std::path::PathBuf::from(".")),
identity_config: self.identity_config.unwrap_or_default(),
skills: self.skills.unwrap_or_default(),
auto_save: self.auto_save.unwrap_or(false),
history: Vec::new(),
classification_config: self.classification_config.unwrap_or_default(),
available_hints: self.available_hints.unwrap_or_default(),
post_turn_hooks: self.post_turn_hooks,
learning_enabled: self.learning_enabled,
event_session_id: self
.event_session_id
.unwrap_or_else(|| "standalone".to_string()),
event_channel: self.event_channel.unwrap_or_else(|| "internal".to_string()),
// The context pipeline is intentionally constructed with its
// defaults here — its tunables (`tool_result_budget_bytes`,
// microcompact thresholds, session-memory knobs) are read from
// the per-agent `AgentConfig` at call sites, so there is no
// need for a separate fluent setter on the builder today.
context_pipeline: ContextPipeline::default(),
})
}
}
impl Agent {
/// Creates an `Agent` instance from a global configuration.
///
/// This is the primary way to initialize an agent with all system
/// integrations (memory, tools, skills, etc.) configured.
pub fn from_config(config: &Config) -> Result<Self> {
let runtime: Arc<dyn host_runtime::RuntimeAdapter> =
Arc::from(host_runtime::create_runtime(&config.runtime)?);
let security = Arc::new(SecurityPolicy::from_config(
&config.autonomy,
&config.workspace_dir,
));
let memory: Arc<dyn Memory> = Arc::from(memory::create_memory_with_storage_and_routes(
&config.memory,
&config.embedding_routes,
Some(&config.storage.provider.config),
&config.workspace_dir,
config.api_key.as_deref(),
)?);
let composio_key = if config.composio.enabled {
config.composio.api_key.as_deref()
} else {
None
};
let composio_entity_id = if config.composio.enabled {
Some(config.composio.entity_id.as_str())
} else {
None
};
let mut tools = tools::all_tools_with_runtime(
Arc::new(config.clone()),
&security,
runtime,
memory.clone(),
composio_key,
composio_entity_id,
&config.browser,
&config.http_request,
&config.workspace_dir,
&config.agents,
config.api_key.as_deref(),
config,
);
// Bridge skill tools (Notion, Gmail, etc.) from the QuickJS runtime
// into the agent's tool registry so the LLM can call them.
let skill_tools = tools::skill_bridge::collect_skill_tools();
if !skill_tools.is_empty() {
log::info!(
"[agent] Injecting {} skill tool(s) into agent registry",
skill_tools.len()
);
tools.extend(skill_tools);
}
let model_name = config
.default_model
.as_deref()
.unwrap_or(crate::openhuman::config::DEFAULT_MODEL)
.to_string();
let provider_runtime_options = providers::ProviderRuntimeOptions {
auth_profile_override: None,
openhuman_dir: config.config_path.parent().map(std::path::PathBuf::from),
secrets_encrypt: config.secrets.encrypt,
reasoning_enabled: config.runtime.reasoning_enabled,
};
let provider: Box<dyn Provider> = providers::create_routed_provider_with_options(
config.api_key.as_deref(),
config.api_url.as_deref(),
&config.reliability,
&config.model_routes,
&model_name,
&provider_runtime_options,
)?;
let dispatcher_choice = config.agent.tool_dispatcher.as_str();
let tool_dispatcher: Box<dyn ToolDispatcher> = match dispatcher_choice {
"native" => Box::new(NativeToolDispatcher),
"xml" => Box::new(XmlToolDispatcher),
_ if provider.supports_native_tools() => Box::new(NativeToolDispatcher),
_ => Box::new(XmlToolDispatcher),
};
let available_hints: Vec<String> =
config.model_routes.iter().map(|r| r.hint.clone()).collect();
// Build prompt builder, optionally with learning sections
let mut prompt_builder = SystemPromptBuilder::with_defaults();
if config.learning.enabled {
prompt_builder = prompt_builder
.add_section(Box::new(
crate::openhuman::learning::LearnedContextSection::new(memory.clone()),
))
.add_section(Box::new(
crate::openhuman::learning::UserProfileSection::new(memory.clone()),
));
log::info!("[learning] prompt sections registered (learned_context, user_profile)");
}
// Build post-turn hooks when learning is enabled
let mut post_turn_hooks: Vec<Arc<dyn crate::openhuman::agent::hooks::PostTurnHook>> =
Vec::new();
if config.learning.enabled {
if config.learning.reflection_enabled {
// Only the reflection hook needs an owned snapshot of the
// full config, so create the `Arc` lazily inside this
// branch instead of paying for the clone whenever
// `learning.enabled` is true.
let full_config = Arc::new(config.clone());
// For cloud reflection, wrap the provider in an Arc.
// For local, no provider needed.
let reflection_provider: Option<Arc<dyn crate::openhuman::providers::Provider>> =
if config.learning.reflection_source
== crate::openhuman::config::ReflectionSource::Cloud
{
Some(Arc::from(providers::create_routed_provider(
config.api_key.as_deref(),
config.api_url.as_deref(),
&config.reliability,
&config.model_routes,
&model_name,
)?))
} else {
None
};
post_turn_hooks.push(Arc::new(crate::openhuman::learning::ReflectionHook::new(
config.learning.clone(),
full_config.clone(),
memory.clone(),
reflection_provider,
)));
log::info!(
"[learning] reflection hook registered (source={:?})",
config.learning.reflection_source
);
}
if config.learning.user_profile_enabled {
post_turn_hooks.push(Arc::new(crate::openhuman::learning::UserProfileHook::new(
config.learning.clone(),
memory.clone(),
)));
log::info!("[learning] user_profile hook registered");
}
if config.learning.tool_tracking_enabled {
post_turn_hooks.push(Arc::new(crate::openhuman::learning::ToolTrackerHook::new(
config.learning.clone(),
memory.clone(),
)));
log::info!("[learning] tool_tracker hook registered");
}
}
Agent::builder()
.provider(provider)
.tools(tools)
.memory(memory)
.tool_dispatcher(tool_dispatcher)
.memory_loader(Box::new(
DefaultMemoryLoader::new(5, config.memory.min_relevance_score)
.with_max_chars(config.agent.max_memory_context_chars),
))
.prompt_builder(prompt_builder)
.config(config.agent.clone())
.model_name(model_name)
.temperature(config.default_temperature)
.workspace_dir(config.workspace_dir.clone())
.classification_config(config.query_classification.clone())
.available_hints(available_hints)
.identity_config(config.identity.clone())
.skills(crate::openhuman::skills::load_skills(&config.workspace_dir))
.auto_save(config.memory.auto_save)
.post_turn_hooks(post_turn_hooks)
.learning_enabled(config.learning.enabled)
.build()
}
}
+61
View File
@@ -0,0 +1,61 @@
//! Core agent implementation for the OpenHuman platform.
//!
//! This module provides the [`Agent`] struct, which orchestrates the
//! interaction between the AI provider, available tools, memory
//! systems, and the user. It handles the agent's "turn" logic,
//! including tool execution and history management.
//!
//! # File layout
//!
//! This module used to be a single 2000-line `agent.rs` file. It's now
//! split into focused children so each file has a clear role:
//!
//! | File | Role |
//! |---------------|------------------------------------------------------------------|
//! | [`types`] | `Agent` and `AgentBuilder` struct definitions (no logic). |
//! | [`builder`] | `AgentBuilder` fluent API + `Agent::from_config` factory. |
//! | [`turn`] | The `turn()` lifecycle, tool dispatch, context-pipeline wiring. |
//! | [`runtime`] | Public accessors, `run_single` / `run_interactive`, helpers. |
//! | `tests` | Integration tests (private). |
//!
//! External callers should import [`Agent`] and [`AgentBuilder`] from
//! this module (or from `crate::openhuman::agent`, which re-exports
//! them). The child files are an implementation detail.
mod builder;
mod runtime;
mod turn;
mod types;
#[cfg(test)]
mod tests;
pub use types::{Agent, AgentBuilder};
use crate::openhuman::config::Config;
use anyhow::Result;
/// Convenience entry point to run an agent with the given configuration and message.
pub async fn run(
config: Config,
message: Option<String>,
model_override: Option<String>,
temperature: f64,
) -> Result<()> {
let mut effective_config = config;
if let Some(m) = model_override {
effective_config.default_model = Some(m);
}
effective_config.default_temperature = temperature;
let mut agent = Agent::from_config(&effective_config)?;
if let Some(msg) = message {
let response = agent.run_single(&msg).await?;
println!("{response}");
} else {
agent.run_interactive().await?;
}
Ok(())
}
+308
View File
@@ -0,0 +1,308 @@
//! Public accessors, `run_single` / `run_interactive` CLI helpers, and
//! assorted per-turn static helpers (id-fallback injection, event-error
//! sanitisation, history diffing).
//!
//! These used to live alongside the turn loop in `agent.rs`. Splitting
//! them out keeps `turn.rs` focused on the interaction lifecycle and
//! makes it obvious which methods are cheap getters vs which actually
//! drive the model.
use super::types::{Agent, AgentBuilder};
use crate::openhuman::agent::dispatcher::ParsedToolCall;
use crate::openhuman::agent::error::AgentError;
use crate::openhuman::event_bus::{publish_global, DomainEvent};
use crate::openhuman::memory::Memory;
use crate::openhuman::providers::{self, ConversationMessage, Provider, ToolCall};
use crate::openhuman::tools::{Tool, ToolSpec};
use crate::openhuman::util::truncate_with_ellipsis;
use anyhow::Result;
use std::sync::Arc;
impl Agent {
const EVENT_ERROR_MAX_CHARS: usize = 256;
// ─────────────────────────────────────────────────────────────────
// Small accessors used by `run_single` + `turn` + sub-agent runner
// ─────────────────────────────────────────────────────────────────
pub(super) fn event_session_id(&self) -> &str {
&self.event_session_id
}
pub(super) fn event_channel(&self) -> &str {
&self.event_channel
}
/// Returns a new `AgentBuilder`.
pub fn builder() -> AgentBuilder {
AgentBuilder::new()
}
/// Borrow the agent's provider as an `Arc`. Used by the sub-agent
/// runner to share the parent's provider instance with spawned
/// sub-agents (so they share connection pools, retry budgets, and
/// rate-limit state).
pub fn provider_arc(&self) -> Arc<dyn Provider> {
Arc::clone(&self.provider)
}
/// Borrow the agent's tools as a slice. Used by the sub-agent runner
/// to filter the parent's tool registry per-archetype.
pub fn tools(&self) -> &[Box<dyn Tool>] {
self.tools.as_slice()
}
/// Clone the agent's tools `Arc` for sharing with sub-agents.
pub fn tools_arc(&self) -> Arc<Vec<Box<dyn Tool>>> {
Arc::clone(&self.tools)
}
/// Borrow the agent's tool specs (pre-serialised). Captured at
/// turn-start so sub-agents can pass byte-identical schemas to the
/// provider for prefix-cache reuse.
pub fn tool_specs(&self) -> &[ToolSpec] {
self.tool_specs.as_slice()
}
/// Clone the agent's tool specs `Arc` for sharing with sub-agents.
pub fn tool_specs_arc(&self) -> Arc<Vec<ToolSpec>> {
Arc::clone(&self.tool_specs)
}
/// Borrow the agent's memory backing store as an `Arc`.
pub fn memory_arc(&self) -> Arc<dyn Memory> {
Arc::clone(&self.memory)
}
/// The agent's working directory.
pub fn workspace_dir(&self) -> &std::path::Path {
&self.workspace_dir
}
/// The agent's currently-configured model name (before per-turn
/// auto-classification).
pub fn model_name(&self) -> &str {
&self.model_name
}
/// The agent's currently-configured temperature.
pub fn temperature(&self) -> f64 {
self.temperature
}
/// The agent's loaded skills, if any.
pub fn skills(&self) -> &[crate::openhuman::skills::Skill] {
&self.skills
}
/// The agent's identity config (used by sub-agent prompt building
/// when `omit_identity = false`).
pub fn identity_config(&self) -> &crate::openhuman::config::IdentityConfig {
&self.identity_config
}
/// The agent's runtime config snapshot.
pub fn agent_config(&self) -> &crate::openhuman::config::AgentConfig {
&self.config
}
/// Returns the current conversation history.
pub fn history(&self) -> &[ConversationMessage] {
&self.history
}
pub fn set_event_context(&mut self, session_id: impl Into<String>, channel: impl Into<String>) {
self.event_session_id = session_id.into();
self.event_channel = channel.into();
}
/// Clears the agent's conversation history.
pub fn clear_history(&mut self) {
self.history.clear();
}
// ─────────────────────────────────────────────────────────────────
// Static helpers for turn parsing + telemetry
// ─────────────────────────────────────────────────────────────────
pub(super) fn count_iterations(messages: &[ConversationMessage]) -> usize {
messages
.iter()
.filter(|message| matches!(message, ConversationMessage::AssistantToolCalls { .. }))
.count()
+ 1
}
fn conversation_message_eq(left: &ConversationMessage, right: &ConversationMessage) -> bool {
serde_json::to_string(left).ok() == serde_json::to_string(right).ok()
}
fn message_slice_eq(left: &[ConversationMessage], right: &[ConversationMessage]) -> bool {
left.len() == right.len()
&& left
.iter()
.zip(right.iter())
.all(|(left, right)| Self::conversation_message_eq(left, right))
}
pub(super) fn new_entries_for_turn<'a>(
history_snapshot: &[ConversationMessage],
current_history: &'a [ConversationMessage],
) -> &'a [ConversationMessage] {
let common_prefix_len = history_snapshot
.iter()
.zip(current_history.iter())
.take_while(|(left, right)| Self::conversation_message_eq(left, right))
.count();
if common_prefix_len == history_snapshot.len() {
return &current_history[common_prefix_len..];
}
let max_overlap = history_snapshot.len().min(current_history.len());
for overlap in (0..=max_overlap).rev() {
let snapshot_suffix = &history_snapshot[history_snapshot.len() - overlap..];
let current_prefix = &current_history[..overlap];
if Self::message_slice_eq(snapshot_suffix, current_prefix) {
return &current_history[overlap..];
}
}
current_history
}
pub(super) fn sanitize_event_error_message(err: &anyhow::Error) -> String {
let kind = match err.downcast_ref::<AgentError>() {
Some(AgentError::ProviderError { .. }) => Some("provider_error"),
Some(AgentError::ContextLimitExceeded { .. }) => Some("context_limit_exceeded"),
Some(AgentError::ToolExecutionError { .. }) => Some("tool_execution_error"),
Some(AgentError::CostBudgetExceeded { .. }) => Some("cost_budget_exceeded"),
Some(AgentError::MaxIterationsExceeded { .. }) => Some("max_iterations_exceeded"),
Some(AgentError::CompactionFailed { .. }) => Some("compaction_failed"),
Some(AgentError::PermissionDenied { .. }) => Some("permission_denied"),
Some(AgentError::Other(_)) | None => None,
};
if let Some(kind) = kind {
return kind.to_string();
}
let scrubbed = providers::sanitize_api_error(&err.to_string())
.replace(['\n', '\r', '\t'], " ")
.split_whitespace()
.collect::<Vec<_>>()
.join(" ");
truncate_with_ellipsis(&scrubbed, Self::EVENT_ERROR_MAX_CHARS)
}
/// Injects unique IDs into tool calls that are missing them.
///
/// This is necessary for some tool dispatchers to correctly track and
/// associate results.
pub(super) fn with_fallback_tool_call_ids(
mut parsed_calls: Vec<ParsedToolCall>,
iteration: usize,
) -> Vec<ParsedToolCall> {
for (idx, call) in parsed_calls.iter_mut().enumerate() {
if call.tool_call_id.is_none() {
call.tool_call_id = Some(format!("parsed-{}-{}", iteration + 1, idx + 1));
}
}
parsed_calls
}
/// Converts parsed tool calls into the provider-standard `ToolCall` format.
///
/// If the provider response already contains native tool calls, they are
/// returned as-is.
pub(super) fn persisted_tool_calls_for_history(
response: &crate::openhuman::providers::ChatResponse,
parsed_calls: &[ParsedToolCall],
iteration: usize,
) -> Vec<ToolCall> {
if !response.tool_calls.is_empty() {
return response.tool_calls.clone();
}
parsed_calls
.iter()
.enumerate()
.map(|(idx, call)| ToolCall {
id: call
.tool_call_id
.clone()
.unwrap_or_else(|| format!("parsed-{}-{}", iteration + 1, idx + 1)),
name: call.name.clone(),
arguments: call.arguments.to_string(),
})
.collect()
}
// ─────────────────────────────────────────────────────────────────
// Run helpers — single-shot and interactive loops
// ─────────────────────────────────────────────────────────────────
/// Runs a single turn with the given message and returns the response.
pub async fn run_single(&mut self, message: &str) -> Result<String> {
let history_snapshot = self.history.clone();
publish_global(DomainEvent::AgentTurnStarted {
session_id: self.event_session_id().to_string(),
channel: self.event_channel().to_string(),
});
match self.turn(message).await {
Ok(response) => {
let new_entries = Self::new_entries_for_turn(&history_snapshot, &self.history);
publish_global(DomainEvent::AgentTurnCompleted {
session_id: self.event_session_id().to_string(),
text_chars: response.chars().count(),
iterations: Self::count_iterations(new_entries),
});
Ok(response)
}
Err(err) => {
let sanitized_message = Self::sanitize_event_error_message(&err);
publish_global(DomainEvent::AgentError {
session_id: self.event_session_id().to_string(),
message: sanitized_message,
recoverable: false,
});
Err(err)
}
}
}
/// Runs an interactive CLI loop, reading from standard input and printing to standard output.
///
/// Each incoming message is dispatched through [`Agent::run_single`] so
/// the unified lifecycle events (`AgentTurnStarted`, `AgentTurnCompleted`,
/// `AgentError`) and error sanitisation run for interactive turns just
/// like they do for one-shot invocations.
pub async fn run_interactive(&mut self) -> Result<()> {
println!("🦀 OpenHuman Interactive Mode");
println!("Type /quit to exit.\n");
let (tx, mut rx) = tokio::sync::mpsc::channel(32);
let cli = crate::openhuman::channels::CliChannel::new();
let listen_handle = tokio::spawn(async move {
let _ = crate::openhuman::channels::Channel::listen(&cli, tx).await;
});
while let Some(msg) = rx.recv().await {
match self.run_single(&msg.content).await {
Ok(response) => println!("\n{response}\n"),
Err(e) => {
// `run_single` already publishes `AgentError` and
// sanitises the payload; surface a concise line here
// for the CLI user and continue the loop.
eprintln!("\nError: {e}\n");
continue;
}
}
}
listen_handle.abort();
Ok(())
}
}
+582
View File
@@ -0,0 +1,582 @@
//! `Agent` unit + integration tests.
//!
//! All tests exercise the agent through its public surface only (no
//! private-field access), which is why they live in a sibling file
//! rather than inline with one of the impl blocks. Shared fakes
//! (`MockProvider`, `RecordingProvider`, `MockTool`) are defined here.
use super::types::{Agent, AgentBuilder};
use crate::openhuman::agent::dispatcher::{NativeToolDispatcher, XmlToolDispatcher};
use crate::openhuman::memory::Memory;
use crate::openhuman::providers::{ChatRequest, ConversationMessage, Provider};
use crate::openhuman::tools::Tool;
use anyhow::Result;
use async_trait::async_trait;
use parking_lot::Mutex;
use std::sync::Arc;
struct MockProvider {
responses: Mutex<Vec<crate::openhuman::providers::ChatResponse>>,
}
#[async_trait]
impl Provider for MockProvider {
async fn chat_with_system(
&self,
_system_prompt: Option<&str>,
_message: &str,
_model: &str,
_temperature: f64,
) -> Result<String> {
Ok("ok".into())
}
async fn chat(
&self,
_request: ChatRequest<'_>,
_model: &str,
_temperature: f64,
) -> Result<crate::openhuman::providers::ChatResponse> {
let mut guard = self.responses.lock();
if guard.is_empty() {
return Ok(crate::openhuman::providers::ChatResponse {
text: Some("done".into()),
tool_calls: vec![],
usage: None,
});
}
Ok(guard.remove(0))
}
}
/// Provider that records the system prompt bytes and model name of
/// every `chat()` call. Used by KV-cache stability tests — anything
/// that varies between turns (timestamps, re-rendered memory context,
/// flipped model hints) will show up as a diff between captures.
#[derive(Default)]
struct RecordingProvider {
captures: Mutex<Vec<CapturedCall>>,
responses: Mutex<Vec<crate::openhuman::providers::ChatResponse>>,
}
#[derive(Clone)]
struct CapturedCall {
system_prompt: Option<String>,
model: String,
}
#[async_trait]
impl Provider for RecordingProvider {
async fn chat_with_system(
&self,
_system_prompt: Option<&str>,
_message: &str,
_model: &str,
_temperature: f64,
) -> Result<String> {
Ok("ok".into())
}
async fn chat(
&self,
request: ChatRequest<'_>,
model: &str,
_temperature: f64,
) -> Result<crate::openhuman::providers::ChatResponse> {
let system_prompt = request
.messages
.iter()
.find(|m| m.role == "system")
.map(|m| m.content.clone());
self.captures.lock().push(CapturedCall {
system_prompt,
model: model.to_string(),
});
let mut guard = self.responses.lock();
if guard.is_empty() {
return Ok(crate::openhuman::providers::ChatResponse {
text: Some("done".into()),
tool_calls: vec![],
usage: None,
});
}
Ok(guard.remove(0))
}
}
struct MockTool;
#[async_trait]
impl Tool for MockTool {
fn name(&self) -> &str {
"echo"
}
fn description(&self) -> &str {
"echo"
}
fn parameters_schema(&self) -> serde_json::Value {
serde_json::json!({"type": "object"})
}
async fn execute(
&self,
_args: serde_json::Value,
) -> Result<crate::openhuman::tools::ToolResult> {
Ok(crate::openhuman::tools::ToolResult::success("tool-out"))
}
}
// silence clippy — `AgentBuilder` is imported so tests can reference
// it in doc examples / type assertions if needed.
#[allow(dead_code)]
fn _assert_builder_is_exported() -> AgentBuilder {
Agent::builder()
}
#[tokio::test]
async fn turn_without_tools_returns_text() {
let workspace = tempfile::TempDir::new().expect("temp workspace");
let workspace_path = workspace.path().to_path_buf();
let provider = Box::new(MockProvider {
responses: Mutex::new(vec![crate::openhuman::providers::ChatResponse {
text: Some("hello".into()),
tool_calls: vec![],
usage: None,
}]),
});
let memory_cfg = crate::openhuman::config::MemoryConfig {
backend: "none".into(),
..crate::openhuman::config::MemoryConfig::default()
};
let mem: Arc<dyn Memory> = Arc::from(
crate::openhuman::memory::create_memory(&memory_cfg, &workspace_path, None).unwrap(),
);
let mut agent = Agent::builder()
.provider(provider)
.tools(vec![Box::new(MockTool)])
.memory(mem)
.tool_dispatcher(Box::new(XmlToolDispatcher))
.workspace_dir(workspace_path)
.build()
.unwrap();
let response = agent.turn("hi").await.unwrap();
assert_eq!(response, "hello");
}
#[tokio::test]
async fn turn_with_native_dispatcher_handles_tool_results_variant() {
let workspace = tempfile::TempDir::new().expect("temp workspace");
let workspace_path = workspace.path().to_path_buf();
let provider = Box::new(MockProvider {
responses: Mutex::new(vec![
crate::openhuman::providers::ChatResponse {
text: Some(String::new()),
tool_calls: vec![crate::openhuman::providers::ToolCall {
id: "tc1".into(),
name: "echo".into(),
arguments: "{}".into(),
}],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("done".into()),
tool_calls: vec![],
usage: None,
},
]),
});
let memory_cfg = crate::openhuman::config::MemoryConfig {
backend: "none".into(),
..crate::openhuman::config::MemoryConfig::default()
};
let mem: Arc<dyn Memory> = Arc::from(
crate::openhuman::memory::create_memory(&memory_cfg, &workspace_path, None).unwrap(),
);
let mut agent = Agent::builder()
.provider(provider)
.tools(vec![Box::new(MockTool)])
.memory(mem)
.tool_dispatcher(Box::new(NativeToolDispatcher))
.workspace_dir(workspace_path)
.build()
.unwrap();
let response = agent.turn("hi").await.unwrap();
assert_eq!(response, "done");
assert!(agent
.history()
.iter()
.any(|msg| matches!(msg, ConversationMessage::ToolResults(_))));
}
#[tokio::test]
async fn turn_with_native_dispatcher_persists_fallback_tool_calls() {
let workspace = tempfile::TempDir::new().expect("temp workspace");
let workspace_path = workspace.path().to_path_buf();
let provider = Box::new(MockProvider {
responses: Mutex::new(vec![
crate::openhuman::providers::ChatResponse {
text: Some(
"Checking...\n<tool_call>{\"name\":\"echo\",\"arguments\":{}}</tool_call>"
.into(),
),
tool_calls: vec![],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("done".into()),
tool_calls: vec![],
usage: None,
},
]),
});
let memory_cfg = crate::openhuman::config::MemoryConfig {
backend: "none".into(),
..crate::openhuman::config::MemoryConfig::default()
};
let mem: Arc<dyn Memory> = Arc::from(
crate::openhuman::memory::create_memory(&memory_cfg, &workspace_path, None).unwrap(),
);
let mut agent = Agent::builder()
.provider(provider)
.tools(vec![Box::new(MockTool)])
.memory(mem)
.tool_dispatcher(Box::new(NativeToolDispatcher))
.workspace_dir(workspace_path)
.build()
.unwrap();
let response = agent.turn("hi").await.unwrap();
assert_eq!(response, "done");
let persisted_calls = agent
.history()
.iter()
.find_map(|msg| match msg {
ConversationMessage::AssistantToolCalls { tool_calls, .. } => Some(tool_calls),
_ => None,
})
.expect("assistant tool calls should be persisted");
assert_eq!(persisted_calls.len(), 1);
assert_eq!(persisted_calls[0].name, "echo");
}
/// End-to-end: parent Agent issues a `spawn_subagent` tool call, the
/// runner dispatches a built-in sub-agent (`researcher`) using the
/// same MockProvider, and the parent's next turn folds the sub-agent's
/// text output into the final response.
///
/// This is the highest-level test that exercises:
/// - Agent::turn → execute_tool_call → SpawnSubagentTool::execute
/// - PARENT_CONTEXT task-local visibility
/// - AgentDefinitionRegistry::global lookup
/// - run_subagent → run_inner_loop with the parent's provider
/// - Result returned as a ToolResult and threaded back into history
#[tokio::test]
async fn turn_dispatches_spawn_subagent_through_full_path() {
use crate::openhuman::agent::harness::AgentDefinitionRegistry;
use crate::openhuman::tools::SpawnSubagentTool;
// Idempotent — other tests may have already initialised it.
AgentDefinitionRegistry::init_global_builtins().unwrap();
let workspace = tempfile::TempDir::new().expect("temp workspace");
let workspace_path = workspace.path().to_path_buf();
// Scripted responses, in the exact order MockProvider will see them:
// 1. Parent turn iter 0 — emit a spawn_subagent tool call.
// 2. Sub-agent (researcher) iter 0 — return final text "X is Y".
// 3. Parent turn iter 1 — fold sub-agent result into "Based on the research, X is Y."
let provider = Box::new(MockProvider {
responses: Mutex::new(vec![
crate::openhuman::providers::ChatResponse {
text: Some(String::new()),
tool_calls: vec![crate::openhuman::providers::ToolCall {
id: "call-spawn".into(),
name: "spawn_subagent".into(),
arguments: serde_json::json!({
"agent_id": "researcher",
"prompt": "find out about X"
})
.to_string(),
}],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("X is Y".into()),
tool_calls: vec![],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("Based on the research, X is Y.".into()),
tool_calls: vec![],
usage: None,
},
]),
});
let memory_cfg = crate::openhuman::config::MemoryConfig {
backend: "none".into(),
..crate::openhuman::config::MemoryConfig::default()
};
let mem: Arc<dyn Memory> = Arc::from(
crate::openhuman::memory::create_memory(&memory_cfg, &workspace_path, None).unwrap(),
);
// Tools include SpawnSubagentTool so the parent can call it.
let tools: Vec<Box<dyn Tool>> = vec![Box::new(SpawnSubagentTool::new())];
let mut agent = Agent::builder()
.provider(provider)
.tools(tools)
.memory(mem)
.tool_dispatcher(Box::new(NativeToolDispatcher))
.workspace_dir(workspace_path)
.build()
.unwrap();
let response = agent.turn("tell me about X").await.unwrap();
assert_eq!(response, "Based on the research, X is Y.");
// The parent's history should contain the spawn_subagent
// assistant tool call AND a tool-result message carrying the
// sub-agent's compact output.
let has_spawn_call = agent.history().iter().any(|msg| match msg {
ConversationMessage::AssistantToolCalls { tool_calls, .. } => {
tool_calls.iter().any(|c| c.name == "spawn_subagent")
}
_ => false,
});
assert!(
has_spawn_call,
"parent history should contain the spawn_subagent assistant tool call"
);
let tool_result_contains_subagent_output = agent.history().iter().any(|msg| match msg {
ConversationMessage::ToolResults(results) => {
results.iter().any(|r| r.content.contains("X is Y"))
}
ConversationMessage::Chat(chat) if chat.role == "tool" => chat.content.contains("X is Y"),
_ => false,
});
assert!(
tool_result_contains_subagent_output,
"parent history should contain a tool-result entry with the sub-agent's output"
);
}
/// Fork-mode variant of `turn_dispatches_spawn_subagent_through_full_path`.
///
/// Exercises the prefix-replay path: the parent issues
/// `spawn_subagent { mode: "fork", … }`, the runner resolves the `fork`
/// built-in definition, pulls the parent's exact rendered prompt + tool
/// schemas + message prefix out of the `ForkContext` task-local, and
/// runs the inner loop on the parent's own provider.
///
/// From the provider's perspective the response queue is consumed in
/// the same fixed sequence as the typed test — parent tool_call → sub-
/// agent reply → parent folded reply — which is the invariant that
/// makes KV-cache reuse possible on the real backend.
#[tokio::test]
async fn turn_dispatches_spawn_subagent_in_fork_mode() {
use crate::openhuman::agent::harness::AgentDefinitionRegistry;
use crate::openhuman::tools::SpawnSubagentTool;
// Idempotent — other tests may have already initialised it.
AgentDefinitionRegistry::init_global_builtins().unwrap();
let workspace = tempfile::TempDir::new().expect("temp workspace");
let workspace_path = workspace.path().to_path_buf();
// Provider response queue, consumed in order:
// 1. Parent turn iter 0 — emit spawn_subagent with mode=fork.
// 2. Fork sub-agent iter 0 — return "X is Y" (no tool calls).
// 3. Parent turn iter 1 — fold the forked result into the final
// text the user sees.
let provider = Box::new(MockProvider {
responses: Mutex::new(vec![
crate::openhuman::providers::ChatResponse {
text: Some(String::new()),
tool_calls: vec![crate::openhuman::providers::ToolCall {
id: "call-fork".into(),
name: "spawn_subagent".into(),
arguments: serde_json::json!({
// agent_id is still required by the schema even
// though `mode=fork` overrides the lookup to the
// synthetic `fork` definition.
"agent_id": "researcher",
"mode": "fork",
"prompt": "analyse branch X"
})
.to_string(),
}],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("X is Y".into()),
tool_calls: vec![],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("Based on the research, X is Y.".into()),
tool_calls: vec![],
usage: None,
},
]),
});
let memory_cfg = crate::openhuman::config::MemoryConfig {
backend: "none".into(),
..crate::openhuman::config::MemoryConfig::default()
};
let mem: Arc<dyn Memory> = Arc::from(
crate::openhuman::memory::create_memory(&memory_cfg, &workspace_path, None).unwrap(),
);
let tools: Vec<Box<dyn Tool>> = vec![Box::new(SpawnSubagentTool::new())];
let mut agent = Agent::builder()
.provider(provider)
.tools(tools)
.memory(mem)
.tool_dispatcher(Box::new(NativeToolDispatcher))
.workspace_dir(workspace_path)
.build()
.unwrap();
let response = agent.turn("tell me about X").await.unwrap();
assert_eq!(response, "Based on the research, X is Y.");
// Same history assertions as the typed path — the fork runner
// still threads its compact output back through the parent's tool
// result pipeline.
let has_spawn_call = agent.history().iter().any(|msg| match msg {
ConversationMessage::AssistantToolCalls { tool_calls, .. } => {
tool_calls.iter().any(|c| c.name == "spawn_subagent")
}
_ => false,
});
assert!(
has_spawn_call,
"parent history should contain the spawn_subagent assistant tool call"
);
let tool_result_contains_subagent_output = agent.history().iter().any(|msg| match msg {
ConversationMessage::ToolResults(results) => {
results.iter().any(|r| r.content.contains("X is Y"))
}
ConversationMessage::Chat(chat) if chat.role == "tool" => chat.content.contains("X is Y"),
_ => false,
});
assert!(
tool_result_contains_subagent_output,
"parent history should contain a tool-result entry with the fork sub-agent's output"
);
}
/// KV-cache invariant: across multiple turns in the same session, the
/// system-prompt bytes submitted to the provider must be byte-identical,
/// and the model name must not flip. Both are required for the backend's
/// automatic prefix cache to hit — if either changes, the backend must
/// re-prefill the entire prompt every turn.
///
/// This test guards against two regressions:
/// 1. A future edit that reintroduces the subsequent-turn system
/// prompt rebuild (see the `learning_enabled` branch we
/// deliberately removed in `turn()`).
/// 2. A future edit that reintroduces per-message model
/// classification on the main agent (which would flip the
/// effective model between turns).
#[tokio::test]
async fn system_prompt_and_model_are_byte_stable_across_turns() {
let workspace = tempfile::TempDir::new().expect("temp workspace");
let workspace_path = workspace.path().to_path_buf();
let provider = Arc::new(RecordingProvider {
responses: Mutex::new(vec![
crate::openhuman::providers::ChatResponse {
text: Some("first".into()),
tool_calls: vec![],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("second".into()),
tool_calls: vec![],
usage: None,
},
crate::openhuman::providers::ChatResponse {
text: Some("third".into()),
tool_calls: vec![],
usage: None,
},
]),
captures: Mutex::new(Vec::new()),
});
let memory_cfg = crate::openhuman::config::MemoryConfig {
backend: "none".into(),
..crate::openhuman::config::MemoryConfig::default()
};
let mem: Arc<dyn Memory> = Arc::from(
crate::openhuman::memory::create_memory(&memory_cfg, &workspace_path, None).unwrap(),
);
let mut agent = Agent::builder()
.provider_arc(provider.clone() as Arc<dyn Provider>)
.tools(vec![])
.memory(mem)
.tool_dispatcher(Box::new(NativeToolDispatcher))
.workspace_dir(workspace_path)
// Learning flag is explicitly enabled to prove that the
// former "rebuild system prompt on subsequent turns" branch
// is gone — we should still see byte-stable prompts.
.learning_enabled(true)
.build()
.unwrap();
for prompt in ["first question", "second question", "third question"] {
agent.turn(prompt).await.unwrap();
}
let captures = provider.captures.lock().clone();
assert_eq!(
captures.len(),
3,
"expected one provider call per turn, got {}",
captures.len()
);
let first_system = captures[0]
.system_prompt
.as_ref()
.expect("first turn should have a system prompt");
for (idx, cap) in captures.iter().enumerate() {
let sys = cap
.system_prompt
.as_ref()
.expect("every turn should carry the system prompt");
assert_eq!(
sys, first_system,
"system prompt drifted on turn {} — KV cache prefix broken",
idx
);
assert_eq!(
cap.model, captures[0].model,
"model name flipped on turn {} — KV cache namespace broken",
idx
);
}
}
+779
View File
@@ -0,0 +1,779 @@
//! Turn lifecycle: running a single interaction, executing tools, and
//! wiring the context pipeline + sub-agent harness around them.
//!
//! This file owns the "hot path" methods on `Agent`:
//!
//! - [`Agent::turn`] — the big one. Orchestrates system-prompt build,
//! memory-context injection, the provider loop, tool dispatch, and
//! the context pipeline (tool-result budget → microcompact →
//! autocompact signal → session-memory extraction trigger).
//! - [`Agent::execute_tool_call`] / [`Agent::execute_tools`] — the
//! per-call runners, including the fork-cache `ForkContext` stash
//! for `spawn_subagent { mode: "fork" }` invocations.
//! - [`Agent::build_parent_execution_context`] /
//! [`Agent::build_fork_context`] — snapshot helpers for sub-agent
//! task-locals.
//! - [`Agent::trim_history`], [`Agent::fetch_learned_context`],
//! [`Agent::build_system_prompt`], [`Agent::classify_model`] — the
//! small helpers `turn()` leans on every call.
//! - [`Agent::spawn_session_memory_extraction`] — the fire-and-forget
//! background archivist fork.
use super::types::Agent;
use crate::openhuman::agent::context_pipeline;
use crate::openhuman::agent::dispatcher::{ParsedToolCall, ToolExecutionResult};
use crate::openhuman::agent::harness;
use crate::openhuman::agent::hooks::{self, ToolCallRecord, TurnContext};
use crate::openhuman::agent::prompt::{LearnedContextData, PromptContext};
use crate::openhuman::event_bus::{publish_global, DomainEvent};
use crate::openhuman::memory::MemoryCategory;
use crate::openhuman::providers::{ChatMessage, ChatRequest, ConversationMessage};
use crate::openhuman::tools::Tool;
use crate::openhuman::util::truncate_with_ellipsis;
use anyhow::Result;
use std::sync::Arc;
impl Agent {
/// Performs a single interaction "turn" with the agent.
///
/// This is the core logic that takes user input, manages the history,
/// calls the LLM, handles tool calls (up to `max_tool_iterations`),
/// and returns the final assistant response.
pub async fn turn(&mut self, user_message: &str) -> Result<String> {
let turn_started = std::time::Instant::now();
log::info!(
"[agent_loop] turn start message_chars={} history_len={} max_tool_iterations={}",
user_message.chars().count(),
self.history.len(),
self.config.max_tool_iterations
);
if self.history.is_empty() {
// Learned context is only baked into the system prompt on the
// very first turn — once the history is non-empty we reuse the
// stored prompt verbatim to preserve the KV-cache prefix the
// inference backend has already tokenised. Fetching it later
// would just burn memory-store reads on data we throw away.
let learned = self.fetch_learned_context().await;
let system_prompt = self.build_system_prompt(learned)?;
log::info!(
"[agent_loop] system prompt built chars={}",
system_prompt.chars().count()
);
log::debug!("[agent_loop] system prompt body:\n{system_prompt}");
self.history
.push(ConversationMessage::Chat(ChatMessage::system(
system_prompt,
)));
} else {
// Deliberately do NOT rebuild the system prompt on subsequent
// turns. The rendered prompt is the KV-cache prefix the inference
// backend has already tokenised; replacing its bytes (even
// cosmetically) forces the backend to re-prefill from scratch.
//
// Dynamic turn-to-turn context (memory recall, learned snippets)
// rides on the user message via `memory_loader.load_context()`
// — that's where the caller should inject anything that varies
// between turns.
log::trace!(
"[agent_loop] system prompt reused (history_len={}) — KV cache prefix preserved",
self.history.len()
);
}
if self.auto_save {
let _ = self
.memory
.store("user_msg", user_message, MemoryCategory::Conversation, None)
.await;
}
let context = self
.memory_loader
.load_context(self.memory.as_ref(), user_message)
.await
.unwrap_or_default();
let enriched = if context.is_empty() {
user_message.to_string()
} else {
format!("{context}{user_message}")
};
self.history
.push(ConversationMessage::Chat(ChatMessage::user(enriched)));
// Pin the main agent to its configured model for the lifetime of
// the session. Per-turn classification used to run here, but it
// would flip `effective_model` mid-conversation (e.g. reasoning →
// coding based on a single keyword). Every flip invalidates the
// backend's KV cache namespace for this session, costing full
// re-prefill on the very next turn. The main agent's job is to
// decide *which sub-agent* to spawn — that routing lives in the
// model prompt, not in the Rust-side classifier. Sub-agents pick
// their own tier via `ModelSpec::Hint(...)` in their definition.
let effective_model = self.model_name.clone();
log::info!(
"[agent_loop] model pinned model={} (per-turn classification disabled for KV cache stability)",
effective_model
);
// Snapshot the parent's runtime once per turn so any
// `spawn_subagent` invocation that fires inside this turn can
// read it via the PARENT_CONTEXT task-local. We override the
// model field with the post-classification effective model.
let mut parent_context = self.build_parent_execution_context();
parent_context.model_name = effective_model.clone();
// Bump the session-memory turn counter. Used later by
// `should_extract_session_memory` to decide whether to spawn a
// background archivist fork at end-of-turn.
self.context_pipeline.tick_turn();
// Collect tool call records across all iterations for post-turn hooks
let mut all_tool_records: Vec<ToolCallRecord> = Vec::new();
let turn_body = async {
for iteration in 0..self.config.max_tool_iterations {
log::info!(
"[agent_loop] iteration start i={} history_len={}",
iteration + 1,
self.history.len()
);
// Context pipeline stages 3 & 4: run the reduction
// chain before every provider hit. Microcompact fires
// when the guard reports we're above the soft threshold
// and there are older tool results to clear; otherwise
// we log an autocompaction signal (openhuman's
// compactor lives in `loop_/history.rs` and operates on
// the `ChatMessage` shape, so for now the
// `ConversationMessage`-shaped Agent path lets the
// signal bubble up as telemetry until a native
// summariser lands).
let outcome = self.context_pipeline.run_before_call(&mut self.history);
match &outcome {
context_pipeline::PipelineOutcome::NoOp => {}
context_pipeline::PipelineOutcome::Microcompacted(stats) => {
log::info!(
"[agent_loop] context_pipeline microcompact i={} envelopes={} entries={} bytes_freed={}",
iteration + 1,
stats.envelopes_cleared,
stats.entries_cleared,
stats.bytes_freed
);
}
context_pipeline::PipelineOutcome::AutocompactionRequested {
utilisation_pct,
} => {
log::warn!(
"[agent_loop] context_pipeline autocompaction requested i={} utilisation_pct={}",
iteration + 1,
utilisation_pct
);
}
context_pipeline::PipelineOutcome::ContextExhausted {
utilisation_pct,
reason,
} => {
log::error!(
"[agent_loop] context_pipeline context exhausted i={} utilisation_pct={} reason={}",
iteration + 1,
utilisation_pct,
reason
);
return Err(anyhow::anyhow!(
"Context window exhausted ({utilisation_pct}% full): {reason}"
));
}
}
let messages = self.tool_dispatcher.to_provider_messages(&self.history);
log::info!(
"[agent_loop] provider request i={} messages={} send_tool_specs={}",
iteration + 1,
messages.len(),
self.tool_dispatcher.should_send_tool_specs()
);
let provider_started = std::time::Instant::now();
let response = match self
.provider
.chat(
ChatRequest {
messages: &messages,
tools: if self.tool_dispatcher.should_send_tool_specs() {
Some(self.tool_specs.as_slice())
} else {
None
},
system_prompt_cache_boundary: None,
},
&effective_model,
self.temperature,
)
.await
{
Ok(resp) => {
log::info!(
"[agent_loop] provider response i={} elapsed_ms={} text_chars={} native_tool_calls={}",
iteration + 1,
provider_started.elapsed().as_millis(),
resp.text.as_ref().map_or(0, |t| t.chars().count()),
resp.tool_calls.len()
);
log::debug!("[agent_loop] provider response: {resp:?}");
// Feed the context pipeline (guard +
// session-memory token accounting). No-op when
// the provider doesn't return usage.
if let Some(ref usage) = resp.usage {
self.context_pipeline.record_usage(usage);
}
resp
}
Err(err) => return Err(err),
};
let (text, calls) = self.tool_dispatcher.parse_response(&response);
let calls = Self::with_fallback_tool_call_ids(calls, iteration);
log::info!(
"[agent_loop] parsed response i={} parsed_text_chars={} parsed_tool_calls={}",
iteration + 1,
text.chars().count(),
calls.len()
);
if calls.is_empty() {
let final_text = if text.is_empty() {
response.text.unwrap_or_default()
} else {
text
};
log::info!(
"[agent_loop] final response i={} final_chars={}",
iteration + 1,
final_text.chars().count()
);
self.history
.push(ConversationMessage::Chat(ChatMessage::assistant(
final_text.clone(),
)));
self.trim_history();
if self.auto_save {
let summary = truncate_with_ellipsis(&final_text, 100);
let _ = self
.memory
.store("assistant_resp", &summary, MemoryCategory::Daily, None)
.await;
}
// Session-memory tool-call accounting. The actual
// background extraction spawn happens *outside*
// `turn_body` so the spawned task can take an owned
// parent context without fighting the borrow
// checker against `self`. We capture the decision
// here and surface it via the pipeline state — the
// epilogue (below) reads `should_extract_session_memory()`.
self.context_pipeline
.record_tool_calls(all_tool_records.len());
// Fire post-turn hooks (non-blocking)
if !self.post_turn_hooks.is_empty() {
let ctx = TurnContext {
user_message: user_message.to_string(),
assistant_response: final_text.clone(),
tool_calls: all_tool_records,
turn_duration_ms: turn_started.elapsed().as_millis() as u64,
session_id: None,
iteration_count: iteration + 1,
};
hooks::fire_hooks(&self.post_turn_hooks, ctx);
}
return Ok(final_text);
}
if !text.is_empty() {
log::info!(
"[agent_loop] assistant pre-tool text i={} chars={}",
iteration + 1,
text.chars().count()
);
// Push the assistant text into history; rendering is
// the caller's responsibility (the CLI loop walks
// `agent.history()` after each turn, sub-agents and
// library consumers get whatever they need through
// the returned value / history accessors).
self.history
.push(ConversationMessage::Chat(ChatMessage::assistant(
text.clone(),
)));
}
let tool_names: Vec<&str> = calls.iter().map(|call| call.name.as_str()).collect();
log::info!(
"[agent_loop] executing tools i={} names={:?}",
iteration + 1,
tool_names
);
let persisted_tool_calls =
Self::persisted_tool_calls_for_history(&response, &calls, iteration);
log::info!(
"[agent_loop] persisting assistant tool calls i={} persisted_tool_calls={} parsed_tool_calls={}",
iteration + 1,
persisted_tool_calls.len(),
calls.len()
);
self.history.push(ConversationMessage::AssistantToolCalls {
text: if text.is_empty() {
None
} else {
Some(text.clone())
},
tool_calls: persisted_tool_calls,
});
let (results, records) = self.execute_tools(&calls).await;
all_tool_records.extend(records);
log::info!(
"[agent_loop] tool results complete i={} result_count={}",
iteration + 1,
results.len()
);
let formatted = self.tool_dispatcher.format_results(&results);
self.history.push(formatted);
self.trim_history();
log::info!(
"[agent_loop] iteration end i={} history_len={}",
iteration + 1,
self.history.len()
);
}
log::warn!(
"[agent_loop] exceeded maximum tool iterations max={}",
self.config.max_tool_iterations
);
anyhow::bail!(
"Agent exceeded maximum tool iterations ({})",
self.config.max_tool_iterations
)
}; // end of `turn_body` async block
// Run the turn body inside the parent-execution-context scope so
// that any `spawn_subagent` tool call fired during the loop can
// read the parent's provider, tools, model, and workspace via
// the PARENT_CONTEXT task-local.
let result = harness::with_parent_context(parent_context, turn_body).await;
// ── Session-memory extraction (stage 5) ───────────────────────
//
// If the pipeline's deltas have crossed all three thresholds
// (token growth, tool calls, turn count), spawn a *background*
// archivist sub-agent that will distil durable facts into the
// workspace MEMORY.md file via the `update_memory_md` tool.
//
// The spawn is fire-and-forget: the main turn returns the
// user-visible response immediately, and the archivist runs
// asynchronously on the `agentic` tier. We optimistically mark
// the extraction complete right away — if it actually fails,
// we'll just retry on the next threshold window (a few turns
// later), which is the right amount of retry behaviour for a
// librarian task that's idempotent across reruns.
if result.is_ok() && self.context_pipeline.should_extract_session_memory() {
self.spawn_session_memory_extraction();
}
result
}
// ─────────────────────────────────────────────────────────────────
// Per-call tool execution
// ─────────────────────────────────────────────────────────────────
/// Executes a single tool call and returns the result and execution record.
pub(super) async fn execute_tool_call(
&self,
call: &ParsedToolCall,
) -> (ToolExecutionResult, ToolCallRecord) {
let started = std::time::Instant::now();
publish_global(DomainEvent::ToolExecutionStarted {
tool_name: call.name.clone(),
session_id: self.event_session_id().to_string(),
});
log::info!("[agent_loop] tool start name={}", call.name);
// Special-case `spawn_subagent { mode: "fork", … }`: stash a
// ForkContext task-local so the sub-agent runner can replay the
// parent's exact rendered prompt + tool schemas + message prefix
// for backend prefix-cache reuse. The branch is taken before
// executing the tool so the task-local is visible inside
// `tool.execute(...)`.
let fork_context_for_call = if call.name == "spawn_subagent"
&& call
.arguments
.get("mode")
.and_then(|v| v.as_str())
.map(|s| s.eq_ignore_ascii_case("fork"))
.unwrap_or(false)
{
Some(self.build_fork_context(call))
} else {
None
};
let (raw_result, success) =
if let Some(tool) = self.tools.iter().find(|t| t.name() == call.name) {
let exec = tool.execute(call.arguments.clone());
let outcome = if let Some(fork_ctx) = fork_context_for_call {
harness::with_fork_context(fork_ctx, exec).await
} else {
exec.await
};
match outcome {
Ok(r) => {
if !r.is_error {
(r.output(), true)
} else {
(format!("Error: {}", r.output()), false)
}
}
Err(e) => (format!("Error executing {}: {e}", call.name), false),
}
} else {
(format!("Unknown tool: {}", call.name), false)
};
// Context pipeline stage 1: apply the per-result byte budget
// *inline* before the result enters history. This is the only
// cache-safe reduction stage — the truncated body has never
// been sent to the backend so it creates no cache invalidation.
let budget_bytes = self.config.tool_result_budget_bytes;
let (result, budget_outcome) =
context_pipeline::apply_tool_result_budget(raw_result, budget_bytes);
if budget_outcome.truncated {
log::info!(
"[agent_loop] tool_result_budget applied name={} original_bytes={} final_bytes={} dropped_bytes={}",
call.name,
budget_outcome.original_bytes,
budget_outcome.final_bytes,
budget_outcome.original_bytes - budget_outcome.final_bytes
);
}
let elapsed_ms = started.elapsed().as_millis() as u64;
publish_global(DomainEvent::ToolExecutionCompleted {
tool_name: call.name.clone(),
session_id: self.event_session_id().to_string(),
success,
elapsed_ms,
});
log::info!(
"[agent_loop] tool finish name={} elapsed_ms={} output_chars={} success={}",
call.name,
elapsed_ms,
result.chars().count(),
success
);
let output_summary = hooks::sanitize_tool_output(&result, &call.name, success);
let record = ToolCallRecord {
name: call.name.clone(),
arguments: call.arguments.clone(),
success,
output_summary,
duration_ms: elapsed_ms,
};
let exec_result = ToolExecutionResult {
name: call.name.clone(),
output: result,
success,
tool_call_id: call.tool_call_id.clone(),
};
(exec_result, record)
}
/// Executes multiple tool calls in sequence.
pub(super) async fn execute_tools(
&self,
calls: &[ParsedToolCall],
) -> (Vec<ToolExecutionResult>, Vec<ToolCallRecord>) {
let mut results = Vec::with_capacity(calls.len());
let mut records = Vec::with_capacity(calls.len());
for call in calls {
let (exec_result, record) = self.execute_tool_call(call).await;
results.push(exec_result);
records.push(record);
}
(results, records)
}
// ─────────────────────────────────────────────────────────────────
// Sub-agent context snapshots
// ─────────────────────────────────────────────────────────────────
/// Snapshot the parent's runtime so spawned sub-agents can read
/// it via the [`harness::PARENT_CONTEXT`] task-local.
pub(super) fn build_parent_execution_context(&self) -> harness::ParentExecutionContext {
harness::ParentExecutionContext {
provider: Arc::clone(&self.provider),
all_tools: Arc::clone(&self.tools),
all_tool_specs: Arc::clone(&self.tool_specs),
model_name: self.model_name.clone(),
temperature: self.temperature,
workspace_dir: self.workspace_dir.clone(),
memory: Arc::clone(&self.memory),
agent_config: self.config.clone(),
identity_config: self.identity_config.clone(),
skills: Arc::new(self.skills.clone()),
session_id: self.event_session_id().to_string(),
channel: self.event_channel().to_string(),
}
}
/// Build a [`harness::ForkContext`] capturing the parent's
/// rendered system prompt + tool schemas + message prefix at the
/// moment a `spawn_subagent { mode: "fork", … }` call fires.
///
/// The system prompt is pulled from `history[0]` (the agent always
/// stores its rendered system prompt as the first message). The
/// message prefix is the entire current history rendered through
/// the dispatcher — the *same* sequence the parent's next call
/// would send, except the new fork directive replaces the parent's
/// next continuation.
pub(super) fn build_fork_context(&self, call: &ParsedToolCall) -> harness::ForkContext {
let messages = self.tool_dispatcher.to_provider_messages(&self.history);
let system_prompt: String = messages
.first()
.filter(|m| m.role == "system")
.map(|m| m.content.clone())
.unwrap_or_default();
let fork_task_prompt = call
.arguments
.get("prompt")
.and_then(|v| v.as_str())
.unwrap_or_default()
.to_string();
harness::ForkContext {
system_prompt: Arc::new(system_prompt),
tool_specs: Arc::clone(&self.tool_specs),
message_prefix: Arc::new(messages),
cache_boundary: None,
fork_task_prompt,
}
}
// ─────────────────────────────────────────────────────────────────
// History & prompt helpers
// ─────────────────────────────────────────────────────────────────
/// Truncates the conversation history to the configured maximum message count.
///
/// System messages are always preserved. Older non-system messages are
/// dropped first.
pub(super) fn trim_history(&mut self) {
let max = self.config.max_history_messages;
if self.history.len() <= max {
return;
}
let mut system_messages = Vec::new();
let mut other_messages = Vec::new();
for msg in self.history.drain(..) {
match &msg {
ConversationMessage::Chat(chat) if chat.role == "system" => {
system_messages.push(msg);
}
_ => other_messages.push(msg),
}
}
if other_messages.len() > max {
let drop_count = other_messages.len() - max;
other_messages.drain(0..drop_count);
}
self.history = system_messages;
self.history.extend(other_messages);
}
/// Pre-fetches learned context data from memory (observations, patterns, user profile).
///
/// This is an async, non-blocking operation that populates the context
/// for the system prompt.
pub(super) async fn fetch_learned_context(&self) -> LearnedContextData {
if !self.learning_enabled {
return LearnedContextData::default();
}
let obs_entries = self
.memory
.list(
Some(&MemoryCategory::Custom("learning_observations".into())),
None,
)
.await
.unwrap_or_default();
let pat_entries = self
.memory
.list(
Some(&MemoryCategory::Custom("learning_patterns".into())),
None,
)
.await
.unwrap_or_default();
let profile_entries = self
.memory
.list(Some(&MemoryCategory::Custom("user_profile".into())), None)
.await
.unwrap_or_default();
LearnedContextData {
observations: obs_entries
.iter()
.rev()
.take(5)
.map(|e| sanitize_learned_entry(&e.content))
.collect(),
patterns: pat_entries
.iter()
.take(3)
.map(|e| sanitize_learned_entry(&e.content))
.collect(),
user_profile: profile_entries
.iter()
.take(20)
.map(|e| sanitize_learned_entry(&e.content))
.collect(),
}
}
/// Builds the system prompt for the current turn, including tool
/// instructions and learned context.
pub(super) fn build_system_prompt(&self, learned: LearnedContextData) -> Result<String> {
let tools_slice: &[Box<dyn Tool>] = self.tools.as_slice();
let instructions = self.tool_dispatcher.prompt_instructions(tools_slice);
let ctx = PromptContext {
workspace_dir: &self.workspace_dir,
model_name: &self.model_name,
tools: tools_slice,
skills: &self.skills,
identity_config: Some(&self.identity_config),
dispatcher_instructions: &instructions,
learned,
};
self.prompt_builder.build(&ctx)
}
/// Classifies the user message to determine if a specific model hint should be used.
///
/// Currently unused by `turn()` — we pin the main agent to its configured
/// model for KV-cache stability (see the rationale in `turn()` where
/// `effective_model` is set). Kept around because the classifier config
/// is still surfaced via `AgentBuilder::classification_config` and
/// external callers (e.g. eval harnesses) may want to probe it directly.
#[allow(dead_code)]
pub(super) fn classify_model(&self, user_message: &str) -> String {
if let Some(hint) =
crate::openhuman::agent::classifier::classify(&self.classification_config, user_message)
{
if self.available_hints.contains(&hint) {
tracing::info!(hint = hint.as_str(), "Auto-classified query");
return format!("hint:{hint}");
}
}
self.model_name.clone()
}
// ─────────────────────────────────────────────────────────────────
// Session-memory extraction (stage 5 of the context pipeline)
// ─────────────────────────────────────────────────────────────────
/// Spawn a background archivist sub-agent to extract durable facts
/// from the recent conversation into `MEMORY.md`. Fire-and-forget.
///
/// Gated by [`context_pipeline::SessionMemoryState::should_extract`]
/// — see its docs for the threshold invariants. Safe to call from
/// inside `turn()` after the turn body has settled.
pub(super) fn spawn_session_memory_extraction(&mut self) {
let Some(registry) = harness::AgentDefinitionRegistry::global() else {
log::debug!("[session_memory] registry not initialised — skipping extraction spawn");
return;
};
let Some(definition) = registry.get("archivist").cloned() else {
log::debug!(
"[session_memory] archivist definition not found — skipping extraction spawn"
);
return;
};
// Build a dedicated ParentExecutionContext for the background
// task. The in-progress turn's context has already been
// consumed by the `with_parent_context` scope above, so this is
// a fresh snapshot.
let parent_ctx = self.build_parent_execution_context();
let extraction_prompt = context_pipeline::ARCHIVIST_EXTRACTION_PROMPT.to_string();
// Optimistically flip the extraction state to "complete" right
// away: we don't need a channel back from the background task
// because a failed extraction is idempotent — it will just be
// retried after the next threshold crossing. `mark_extraction_complete`
// also clears the `extraction_in_progress` flag, so calling it
// alone covers both bookkeeping steps.
self.context_pipeline
.session_memory
.mark_extraction_complete();
log::info!(
"[session_memory] spawning background archivist extraction (turn={}, tokens={})",
self.context_pipeline.session_memory.current_turn,
self.context_pipeline.session_memory.total_tokens
);
tokio::spawn(async move {
let options = harness::SubagentRunOptions::default();
let fut = harness::run_subagent(&definition, &extraction_prompt, options);
let result = harness::with_parent_context(parent_ctx, fut).await;
match result {
Ok(outcome) => tracing::info!(
agent_id = %outcome.agent_id,
task_id = %outcome.task_id,
iterations = outcome.iterations,
output_chars = outcome.output.chars().count(),
"[session_memory] archivist extraction completed"
),
Err(err) => tracing::warn!(
error = %err,
"[session_memory] archivist extraction failed — will retry after next threshold crossing"
),
}
});
}
}
/// Sanitize a learned memory entry before injecting into the system prompt.
/// Strips raw data, limits length, and removes potential secrets.
fn sanitize_learned_entry(content: &str) -> String {
let trimmed = content.trim();
if trimmed.is_empty() {
return String::new();
}
// Truncate to a safe length
let max_len = 200;
let sanitized: String = trimmed.chars().take(max_len).collect();
// Strip anything that looks like a secret/token
if sanitized.contains("Bearer ")
|| sanitized.contains("sk-")
|| sanitized.contains("ghp_")
|| sanitized.contains("-----BEGIN")
{
return "[redacted: potential secret]".to_string();
}
sanitized
}
+83
View File
@@ -0,0 +1,83 @@
//! `Agent` and `AgentBuilder` struct definitions.
//!
//! The data shapes live here, separate from their behaviour, so the
//! rest of the sub-module (`builder.rs`, `turn.rs`, `runtime.rs`) can
//! focus on logic. Fields are `pub(super)` so sibling files that
//! `impl Agent`/`impl AgentBuilder` can see them without the whole
//! crate gaining field access.
use crate::openhuman::agent::context_pipeline::ContextPipeline;
use crate::openhuman::agent::dispatcher::ToolDispatcher;
use crate::openhuman::agent::hooks::PostTurnHook;
use crate::openhuman::agent::memory_loader::MemoryLoader;
use crate::openhuman::agent::prompt::SystemPromptBuilder;
use crate::openhuman::memory::Memory;
use crate::openhuman::providers::{ConversationMessage, Provider};
use crate::openhuman::tools::{Tool, ToolSpec};
use std::sync::Arc;
/// An autonomous or semi-autonomous AI agent.
///
/// The `Agent` is the central component that manages conversation state,
/// executes tools based on model requests, and interacts with the memory
/// system to maintain context across turns.
pub struct Agent {
pub(super) provider: Arc<dyn Provider>,
pub(super) tools: Arc<Vec<Box<dyn Tool>>>,
pub(super) tool_specs: Arc<Vec<ToolSpec>>,
pub(super) memory: Arc<dyn Memory>,
pub(super) prompt_builder: SystemPromptBuilder,
pub(super) tool_dispatcher: Box<dyn ToolDispatcher>,
pub(super) memory_loader: Box<dyn MemoryLoader>,
pub(super) config: crate::openhuman::config::AgentConfig,
pub(super) model_name: String,
pub(super) temperature: f64,
pub(super) workspace_dir: std::path::PathBuf,
pub(super) identity_config: crate::openhuman::config::IdentityConfig,
pub(super) skills: Vec<crate::openhuman::skills::Skill>,
pub(super) auto_save: bool,
pub(super) history: Vec<ConversationMessage>,
pub(super) classification_config: crate::openhuman::config::QueryClassificationConfig,
pub(super) available_hints: Vec<String>,
pub(super) post_turn_hooks: Vec<Arc<dyn PostTurnHook>>,
pub(super) learning_enabled: bool,
pub(super) event_session_id: String,
pub(super) event_channel: String,
/// Layered context reduction pipeline (tool-result budget →
/// microcompact → autocompact signal → session-memory extraction
/// trigger). Owned by the agent so its state (token counters,
/// session-memory extraction deltas, compaction circuit breaker)
/// persists across turns. See
/// [`crate::openhuman::agent::context_pipeline`] for the stage
/// ordering and cache-safety contract.
pub(super) context_pipeline: ContextPipeline,
}
/// A builder for creating `Agent` instances with custom configuration.
pub struct AgentBuilder {
pub(super) provider: Option<Arc<dyn Provider>>,
pub(super) tools: Option<Vec<Box<dyn Tool>>>,
pub(super) memory: Option<Arc<dyn Memory>>,
pub(super) prompt_builder: Option<SystemPromptBuilder>,
pub(super) tool_dispatcher: Option<Box<dyn ToolDispatcher>>,
pub(super) memory_loader: Option<Box<dyn MemoryLoader>>,
pub(super) config: Option<crate::openhuman::config::AgentConfig>,
pub(super) model_name: Option<String>,
pub(super) temperature: Option<f64>,
pub(super) workspace_dir: Option<std::path::PathBuf>,
pub(super) identity_config: Option<crate::openhuman::config::IdentityConfig>,
pub(super) skills: Option<Vec<crate::openhuman::skills::Skill>>,
pub(super) auto_save: Option<bool>,
pub(super) classification_config: Option<crate::openhuman::config::QueryClassificationConfig>,
pub(super) available_hints: Option<Vec<String>>,
pub(super) post_turn_hooks: Vec<Arc<dyn PostTurnHook>>,
pub(super) learning_enabled: bool,
pub(super) event_session_id: Option<String>,
pub(super) event_channel: Option<String>,
}
impl Default for AgentBuilder {
fn default() -> Self {
Self::new()
}
}
@@ -0,0 +1,266 @@
//! Stage 3: Microcompact.
//!
//! Microcompact is the cheap summarisation substitute. It does **not**
//! generate prose summaries — instead it walks the history and replaces
//! the payload of older `ToolResults` envelopes with a short placeholder
//! string. The envelope itself is preserved so the API invariant
//! `AssistantToolCalls ⇔ ToolResults` holds and the provider still
//! accepts the next request.
//!
//! OpenHuman's inference backend does automatic prefix caching, so we
//! skip any cache-editing dance and go straight to the placeholder
//! strategy: overwrite the old bodies in place, let the backend
//! re-prefill once, and let the next turn pick up the new (smaller)
//! cache target.
//!
//! # Cache implications
//!
//! Microcompact mutates bytes that were previously sent to the backend,
//! so it **deliberately invalidates the KV-cache prefix** for this
//! session. The upside is that the new, smaller prefix becomes the next
//! stable cache target, so subsequent turns hit the cache again. This
//! stage is therefore only run when the next provider call would
//! otherwise be too large to fit — the pipeline orchestrator handles
//! gating.
use crate::openhuman::providers::ConversationMessage;
/// Placeholder used in place of cleared tool-result bodies. Must be
/// stable across versions so callers can pattern-match on it for
/// telemetry / diff tests. Keep it short — the whole point is to free
/// tokens.
pub const CLEARED_PLACEHOLDER: &str = "[Old tool result content cleared]";
/// Default number of most-recent `ToolResults` envelopes to leave
/// intact — the N most recent tool results are kept hot so the model
/// can still reason about them.
pub const DEFAULT_KEEP_RECENT_TOOL_RESULTS: usize = 5;
/// Summary of what a single microcompact pass changed.
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub struct MicrocompactStats {
/// Number of `ToolResults` envelopes whose bodies were cleared.
pub envelopes_cleared: usize,
/// Number of individual tool-result entries within those envelopes
/// whose `content` was replaced.
pub entries_cleared: usize,
/// Bytes freed from the rendered conversation (approximate — counts
/// the `content` string length diff only).
pub bytes_freed: usize,
}
/// Walk `history` and clear the payload of every `ToolResults` envelope
/// except the `keep_recent` most recent ones. Returns a summary of the
/// changes.
///
/// The clearing is idempotent: running the pass twice on the same
/// history is a no-op on the second call because the already-cleared
/// entries will match `CLEARED_PLACEHOLDER` and be skipped.
pub fn microcompact(history: &mut [ConversationMessage], keep_recent: usize) -> MicrocompactStats {
// First sweep: find the indices of every `ToolResults` envelope.
let mut tool_result_indices: Vec<usize> = history
.iter()
.enumerate()
.filter_map(|(i, msg)| matches!(msg, ConversationMessage::ToolResults(_)).then_some(i))
.collect();
// The most-recent envelopes are at the end of the vec — peel off
// `keep_recent` of them and leave them untouched.
if tool_result_indices.len() <= keep_recent {
return MicrocompactStats::default();
}
let cut = tool_result_indices.len().saturating_sub(keep_recent);
tool_result_indices.truncate(cut);
let mut stats = MicrocompactStats::default();
for idx in tool_result_indices {
let ConversationMessage::ToolResults(results) = &mut history[idx] else {
continue;
};
let mut envelope_changed = false;
for entry in results.iter_mut() {
if entry.content == CLEARED_PLACEHOLDER {
// Already cleared on a previous pass — skip.
continue;
}
let old_len = entry.content.len();
entry.content = CLEARED_PLACEHOLDER.to_string();
let freed = old_len.saturating_sub(CLEARED_PLACEHOLDER.len());
stats.bytes_freed += freed;
stats.entries_cleared += 1;
envelope_changed = true;
}
if envelope_changed {
stats.envelopes_cleared += 1;
}
}
stats
}
#[cfg(test)]
mod tests {
use super::*;
use crate::openhuman::providers::{ChatMessage, ToolCall, ToolResultMessage};
fn user(text: &str) -> ConversationMessage {
ConversationMessage::Chat(ChatMessage::user(text))
}
fn assistant_call(id: &str, name: &str) -> ConversationMessage {
ConversationMessage::AssistantToolCalls {
text: None,
tool_calls: vec![ToolCall {
id: id.into(),
name: name.into(),
arguments: "{}".into(),
}],
}
}
fn tool_result(id: &str, body: &str) -> ConversationMessage {
ConversationMessage::ToolResults(vec![ToolResultMessage {
tool_call_id: id.into(),
content: body.into(),
}])
}
#[test]
fn noop_when_no_tool_results() {
let mut history = vec![user("hi"), user("again")];
let stats = microcompact(&mut history, 5);
assert_eq!(stats, MicrocompactStats::default());
}
#[test]
fn noop_when_all_tool_results_within_keep_recent() {
let mut history = vec![
user("q"),
assistant_call("a", "t"),
tool_result("a", "body-a"),
assistant_call("b", "t"),
tool_result("b", "body-b"),
];
let stats = microcompact(&mut history, 5);
assert_eq!(stats, MicrocompactStats::default());
// Bodies unchanged.
if let ConversationMessage::ToolResults(r) = &history[2] {
assert_eq!(r[0].content, "body-a");
} else {
panic!();
}
}
#[test]
fn clears_oldest_when_over_keep_recent() {
let large_body = "x".repeat(5_000);
let mut history = vec![
user("q1"),
assistant_call("t1", "fn"),
tool_result("t1", &large_body), // oldest — should be cleared
assistant_call("t2", "fn"),
tool_result("t2", &large_body), // oldest — should be cleared
assistant_call("t3", "fn"),
tool_result("t3", "recent-1"), // keep
assistant_call("t4", "fn"),
tool_result("t4", "recent-2"), // keep
];
let stats = microcompact(&mut history, 2);
assert_eq!(stats.envelopes_cleared, 2);
assert_eq!(stats.entries_cleared, 2);
assert!(stats.bytes_freed > 9_000);
// Oldest two have been replaced.
match &history[2] {
ConversationMessage::ToolResults(r) => assert_eq!(r[0].content, CLEARED_PLACEHOLDER),
_ => panic!(),
}
match &history[4] {
ConversationMessage::ToolResults(r) => assert_eq!(r[0].content, CLEARED_PLACEHOLDER),
_ => panic!(),
}
// Most-recent two are preserved verbatim.
match &history[6] {
ConversationMessage::ToolResults(r) => assert_eq!(r[0].content, "recent-1"),
_ => panic!(),
}
match &history[8] {
ConversationMessage::ToolResults(r) => assert_eq!(r[0].content, "recent-2"),
_ => panic!(),
}
}
#[test]
fn envelope_invariant_preserved() {
// API requires every AssistantToolCalls to have a matching
// ToolResults envelope. Clearing bodies must not delete the
// envelope or remove entries from the vec inside.
let mut history = vec![
assistant_call("t1", "fn"),
tool_result("t1", "old-1"),
assistant_call("t2", "fn"),
tool_result("t2", "new-1"),
];
microcompact(&mut history, 1);
let mut call_count = 0;
let mut result_count = 0;
for msg in &history {
match msg {
ConversationMessage::AssistantToolCalls { .. } => call_count += 1,
ConversationMessage::ToolResults(_) => result_count += 1,
_ => {}
}
}
assert_eq!(call_count, 2);
assert_eq!(result_count, 2);
}
#[test]
fn second_pass_is_idempotent() {
let mut history = vec![
assistant_call("t1", "fn"),
tool_result("t1", "old-1"),
assistant_call("t2", "fn"),
tool_result("t2", "new-1"),
];
let first = microcompact(&mut history, 1);
assert_eq!(first.envelopes_cleared, 1);
let second = microcompact(&mut history, 1);
assert_eq!(second, MicrocompactStats::default());
}
#[test]
fn clears_all_entries_in_a_multi_entry_envelope() {
let mut history = vec![
assistant_call("t1", "fn"),
ConversationMessage::ToolResults(vec![
ToolResultMessage {
tool_call_id: "a".into(),
content: "A".repeat(1_000),
},
ToolResultMessage {
tool_call_id: "b".into(),
content: "B".repeat(1_000),
},
]),
assistant_call("t2", "fn"),
tool_result("t2", "recent"),
];
let stats = microcompact(&mut history, 1);
assert_eq!(stats.envelopes_cleared, 1);
assert_eq!(stats.entries_cleared, 2);
match &history[1] {
ConversationMessage::ToolResults(r) => {
assert_eq!(r.len(), 2);
assert_eq!(r[0].content, CLEARED_PLACEHOLDER);
assert_eq!(r[1].content, CLEARED_PLACEHOLDER);
}
_ => panic!(),
}
}
}
@@ -0,0 +1,42 @@
//! Layered context-reduction pipeline.
//!
//! Context summarisation in openhuman is layered, not a single knob.
//! Each layer has a specific trigger and invariant:
//!
//! | Stage | File | When | Cache impact |
//! |-------|--------------------------------|--------------------------------|----------------|
//! | 1. Tool-result budget | [`tool_result_budget`] | New tool result created | Cache-safe |
//! | 2. Snip / trim | `Agent::trim_history` | Message count > max | Cache-safe* |
//! | 3. Microcompact | [`microcompact`] | Guard ≥ 90% soft bound | Breaks prefix |
//! | 4. Autocompact | `loop_/history.rs` | Microcompact not enough | Breaks prefix |
//! | 5. Session memory | [`session_memory`] | Token/turn/tool deltas | Async, free |
//!
//! \* Trim only drops messages older than the most-recent stable prefix,
//! which is often outside the KV-cache anyway.
//!
//! The orchestrator is [`pipeline::ContextPipeline`], which is owned by
//! the `Agent` and called once per turn before each provider hit. Stage
//! 1 is applied inline in `Agent::execute_tool_call`, not here.
//!
//! Stage reference:
//! - [`tool_result_budget::apply_tool_result_budget`] — stage 1
//! - [`microcompact::microcompact`] — stage 3
//! - `PipelineOutcome::AutocompactionRequested` — stage 4 signal
//! - [`session_memory::SessionMemoryState`] — stage 5 state tracker
pub mod microcompact;
pub mod pipeline;
pub mod session_memory;
pub mod tool_result_budget;
pub use microcompact::{
microcompact, MicrocompactStats, CLEARED_PLACEHOLDER, DEFAULT_KEEP_RECENT_TOOL_RESULTS,
};
pub use pipeline::{ContextPipeline, ContextPipelineConfig, PipelineOutcome};
pub use session_memory::{
SessionMemoryConfig, SessionMemoryState, ARCHIVIST_EXTRACTION_PROMPT, DEFAULT_MIN_TOKEN_GROWTH,
DEFAULT_MIN_TOOL_CALLS, DEFAULT_MIN_TURNS_BETWEEN,
};
pub use tool_result_budget::{
apply_tool_result_budget, BudgetOutcome, DEFAULT_TOOL_RESULT_BUDGET_BYTES,
};
@@ -0,0 +1,392 @@
//! The layered context pipeline orchestrator.
//!
//! Ordered reduction chain applied before each provider hit:
//!
//! 1. **Tool-result budget** — applied inline in `Agent::execute_tool_call`
//! (not here). Oversized tool results are truncated before they enter
//! history, so they never show up as a pipeline stage.
//! 2. **Snip compact** — hard cap on message count. Implemented by the
//! pre-existing `Agent::trim_history`; the pipeline leaves it to the
//! caller because trimming is a terminal fallback.
//! 3. **Microcompact** — this module. Runs when `ContextGuard` reports
//! `CompactionNeeded` (soft threshold). Replaces the payload of older
//! `ToolResults` envelopes with a placeholder, preserving the
//! `AssistantToolCalls ⇔ ToolResults` API invariant.
//! 4. **Autocompact** — prose summarisation of older messages.
//! OpenHuman's existing `auto_compact_history` lives in
//! `agent/loop_/history.rs` and operates on `ChatMessage` (not
//! `ConversationMessage`), so we don't call it here — the pipeline
//! instead signals a `PipelineOutcome::AutocompactionRequested` to
//! the caller and trusts the caller to dispatch its own summariser
//! when ready. Keeping the pipeline pure (no LLM calls) means the
//! integration tests can exercise every stage without a provider.
//! 5. **Session memory** — handled separately by
//! [`crate::openhuman::agent::context_pipeline::session_memory`].
//!
//! # Cache contract
//!
//! Stages 12 are byte-neutral with respect to previously-sent history
//! (stage 1 applies to a fresh tool result before insertion; stage 2 is
//! a terminal trim). Stages 34 deliberately mutate previously-sent
//! history and therefore break the KV-cache prefix; they run **only
//! when the context guard says we'd otherwise bust the window**. Each
//! firing resets the stable prefix to the new, smaller history so
//! subsequent turns hit the cache again.
use super::microcompact::{microcompact, MicrocompactStats, DEFAULT_KEEP_RECENT_TOOL_RESULTS};
use super::session_memory::{SessionMemoryConfig, SessionMemoryState};
use crate::openhuman::agent::loop_::context_guard::{ContextCheckResult, ContextGuard};
use crate::openhuman::providers::{ConversationMessage, UsageInfo};
/// Pipeline configuration. Defaults are tuned for an `agentic-v1`
/// 128k-context run.
#[derive(Debug, Clone, Copy)]
pub struct ContextPipelineConfig {
/// Number of recent `ToolResults` envelopes microcompact leaves
/// untouched. See [`DEFAULT_KEEP_RECENT_TOOL_RESULTS`].
pub microcompact_keep_recent: usize,
/// Whether to surface the microcompact pass in the pipeline
/// outcome. When `false` the pipeline skips stage 3 entirely —
/// useful for tests that want to exercise autocompaction in
/// isolation.
pub microcompact_enabled: bool,
/// Whether the pipeline should report an autocompaction request
/// when the guard says we're at the hard threshold. When `false`
/// the pipeline silently tolerates an exhausted context (the caller
/// is expected to surface the error via the guard directly).
pub autocompact_enabled: bool,
/// Session-memory extraction tunables.
pub session_memory: SessionMemoryConfig,
}
impl Default for ContextPipelineConfig {
fn default() -> Self {
Self {
microcompact_keep_recent: DEFAULT_KEEP_RECENT_TOOL_RESULTS,
microcompact_enabled: true,
autocompact_enabled: true,
session_memory: SessionMemoryConfig::default(),
}
}
}
/// Outcome of a single pipeline pass, returned to the caller so it can
/// log/telemeter what happened and decide whether to trigger an
/// autocompaction summariser.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum PipelineOutcome {
/// No stage fired — either the guard is happy or the history is
/// already small enough.
NoOp,
/// Microcompact cleared at least one older `ToolResults` envelope.
Microcompacted(MicrocompactStats),
/// The guard reports we're above the soft threshold and
/// microcompact wasn't enough (or was disabled). The caller should
/// invoke its autocompaction summariser.
AutocompactionRequested {
/// The last-known context utilisation as a 0..=100 percentage.
utilisation_pct: u8,
},
/// The guard's circuit breaker is tripped and the context is still
/// above the hard threshold — the caller should abort the turn.
ContextExhausted { utilisation_pct: u8, reason: String },
}
/// Stateful orchestrator. Owns a [`ContextGuard`] and a
/// [`SessionMemoryState`] so a single instance can live on the `Agent`
/// across turns without threading state through every call site.
#[derive(Debug)]
pub struct ContextPipeline {
pub config: ContextPipelineConfig,
pub guard: ContextGuard,
pub session_memory: SessionMemoryState,
}
impl Default for ContextPipeline {
fn default() -> Self {
Self::new(ContextPipelineConfig::default())
}
}
impl ContextPipeline {
pub fn new(config: ContextPipelineConfig) -> Self {
Self {
config,
guard: ContextGuard::new(),
session_memory: SessionMemoryState::default(),
}
}
/// Feed the latest provider `UsageInfo` into both the guard and the
/// session-memory state.
pub fn record_usage(&mut self, usage: &UsageInfo) {
self.guard.update_usage(usage);
self.session_memory
.record_usage(usage.input_tokens + usage.output_tokens);
}
/// Bump the session-memory turn counter. Called once per user turn.
pub fn tick_turn(&mut self) {
self.session_memory.tick_turn();
}
/// Accumulate a turn's tool-call count into the session-memory
/// state. Called once per user turn after tool dispatch settles.
pub fn record_tool_calls(&mut self, n: usize) {
self.session_memory.record_tool_calls(n);
}
/// Should the caller spawn a background session-memory extraction
/// this turn?
pub fn should_extract_session_memory(&self) -> bool {
self.session_memory
.should_extract(&self.config.session_memory)
}
/// Run the reduction chain against `history` in place. Safe to call
/// before every provider hit — it's cheap when the guard is happy.
pub fn run_before_call(&mut self, history: &mut [ConversationMessage]) -> PipelineOutcome {
match self.guard.check() {
ContextCheckResult::Ok => PipelineOutcome::NoOp,
ContextCheckResult::CompactionNeeded => {
// Stage 3: microcompact the older tool results.
if self.config.microcompact_enabled {
let stats = microcompact(history, self.config.microcompact_keep_recent);
if stats.envelopes_cleared > 0 {
// A successful reduction should reset the guard's
// circuit breaker so a previous string of
// autocompaction failures doesn't leave the
// breaker tripped after we've just freed tokens.
self.guard.record_compaction_success();
tracing::info!(
envelopes_cleared = stats.envelopes_cleared,
entries_cleared = stats.entries_cleared,
bytes_freed = stats.bytes_freed,
"[context_pipeline] microcompact fired"
);
return PipelineOutcome::Microcompacted(stats);
}
}
// Stage 4: if microcompact didn't free anything (no old
// tool results to clear), signal autocompaction to the
// caller. The pipeline deliberately does not issue the
// LLM call itself.
if self.config.autocompact_enabled {
let pct = self
.guard
.utilization()
.map(|u| (u * 100.0).round() as u8)
.unwrap_or(0);
tracing::info!(
utilisation_pct = pct,
"[context_pipeline] autocompaction requested"
);
return PipelineOutcome::AutocompactionRequested {
utilisation_pct: pct,
};
}
PipelineOutcome::NoOp
}
ContextCheckResult::ContextExhausted {
utilization_pct,
reason,
} => PipelineOutcome::ContextExhausted {
utilisation_pct: utilization_pct,
reason,
},
}
}
}
#[cfg(test)]
mod tests {
use super::super::microcompact::CLEARED_PLACEHOLDER;
use super::*;
use crate::openhuman::providers::{
ChatMessage, ConversationMessage, ToolCall, ToolResultMessage, UsageInfo,
};
fn call(id: &str) -> ConversationMessage {
ConversationMessage::AssistantToolCalls {
text: None,
tool_calls: vec![ToolCall {
id: id.into(),
name: "t".into(),
arguments: "{}".into(),
}],
}
}
fn result(id: &str, body: &str) -> ConversationMessage {
ConversationMessage::ToolResults(vec![ToolResultMessage {
tool_call_id: id.into(),
content: body.into(),
}])
}
fn user(text: &str) -> ConversationMessage {
ConversationMessage::Chat(ChatMessage::user(text))
}
fn set_high_utilisation(pipeline: &mut ContextPipeline) {
pipeline.record_usage(&UsageInfo {
input_tokens: 92_000,
output_tokens: 4_000,
context_window: 100_000,
});
}
#[test]
fn noop_when_guard_is_ok() {
let mut pipeline = ContextPipeline::default();
pipeline.record_usage(&UsageInfo {
input_tokens: 10_000,
output_tokens: 1_000,
context_window: 100_000,
});
let mut history = vec![
user("hi"),
call("t1"),
result("t1", "x".repeat(2_000).as_str()),
];
let outcome = pipeline.run_before_call(&mut history);
assert_eq!(outcome, PipelineOutcome::NoOp);
}
#[test]
fn microcompact_fires_at_soft_threshold_when_there_are_old_tool_results() {
let mut pipeline = ContextPipeline::default();
let mut history = vec![
call("t1"),
result("t1", &"x".repeat(5_000)),
call("t2"),
result("t2", &"x".repeat(5_000)),
call("t3"),
result("t3", "recent-1"),
call("t4"),
result("t4", "recent-2"),
call("t5"),
result("t5", "recent-3"),
call("t6"),
result("t6", "recent-4"),
call("t7"),
result("t7", "recent-5"),
];
set_high_utilisation(&mut pipeline);
let outcome = pipeline.run_before_call(&mut history);
match outcome {
PipelineOutcome::Microcompacted(stats) => {
assert_eq!(stats.envelopes_cleared, 2);
assert!(stats.bytes_freed > 9_000);
}
other => panic!("expected Microcompacted, got {other:?}"),
}
// Older entries are cleared, newer ones are preserved.
match &history[1] {
ConversationMessage::ToolResults(r) => {
assert_eq!(r[0].content, CLEARED_PLACEHOLDER)
}
_ => panic!(),
}
match &history[13] {
ConversationMessage::ToolResults(r) => assert_eq!(r[0].content, "recent-5"),
_ => panic!(),
}
}
#[test]
fn autocompaction_requested_when_no_old_tool_results_to_clear() {
let mut pipeline = ContextPipeline::default();
// Soft threshold crossed but there are zero ToolResults to clear.
set_high_utilisation(&mut pipeline);
let mut history = vec![user("one"), user("two"), user("three")];
let outcome = pipeline.run_before_call(&mut history);
match outcome {
PipelineOutcome::AutocompactionRequested { utilisation_pct } => {
assert!(utilisation_pct >= 90);
}
other => panic!("expected AutocompactionRequested, got {other:?}"),
}
}
#[test]
fn autocompaction_requested_when_only_recent_tool_results_exist() {
// All tool results fall within `keep_recent`, so microcompact
// has nothing to clear and the pipeline falls through to
// autocompaction.
let mut pipeline = ContextPipeline::default();
let mut history = vec![call("t1"), result("t1", "a"), call("t2"), result("t2", "b")];
set_high_utilisation(&mut pipeline);
let outcome = pipeline.run_before_call(&mut history);
assert!(matches!(
outcome,
PipelineOutcome::AutocompactionRequested { .. }
));
}
#[test]
fn microcompact_disabled_skips_to_autocompaction() {
let mut pipeline = ContextPipeline::new(ContextPipelineConfig {
microcompact_enabled: false,
..ContextPipelineConfig::default()
});
let mut history = vec![
call("t1"),
result("t1", &"x".repeat(5_000)),
call("t2"),
result("t2", "recent"),
];
set_high_utilisation(&mut pipeline);
let outcome = pipeline.run_before_call(&mut history);
assert!(matches!(
outcome,
PipelineOutcome::AutocompactionRequested { .. }
));
// History must be untouched when microcompact is disabled.
if let ConversationMessage::ToolResults(r) = &history[1] {
assert_eq!(r[0].content.len(), 5_000);
} else {
panic!();
}
}
#[test]
fn exhausted_context_propagates_to_caller() {
let mut pipeline = ContextPipeline::default();
pipeline.record_usage(&UsageInfo {
input_tokens: 96_000,
output_tokens: 2_000,
context_window: 100_000,
});
// Trip the circuit breaker.
pipeline.guard.record_compaction_failure();
pipeline.guard.record_compaction_failure();
pipeline.guard.record_compaction_failure();
let mut history = vec![user("hi")];
let outcome = pipeline.run_before_call(&mut history);
assert!(matches!(outcome, PipelineOutcome::ContextExhausted { .. }));
}
#[test]
fn record_usage_feeds_session_memory() {
let mut pipeline = ContextPipeline::default();
pipeline.record_usage(&UsageInfo {
input_tokens: 10_000,
output_tokens: 2_000,
context_window: 100_000,
});
assert_eq!(pipeline.session_memory.total_tokens, 12_000);
}
#[test]
fn tick_turn_and_record_tool_calls_affect_session_memory() {
let mut pipeline = ContextPipeline::default();
pipeline.tick_turn();
pipeline.record_tool_calls(5);
assert_eq!(pipeline.session_memory.current_turn, 1);
assert_eq!(pipeline.session_memory.total_tool_calls, 5);
}
}
@@ -0,0 +1,256 @@
//! Stage 5: Session memory — persistent notes updated by a background fork.
//!
//! Session memory is intentionally **separate** from compaction. While
//! microcompact/autocompact mutate the in-flight conversation history to
//! keep the prompt inside the context window, session memory is a
//! persistent markdown file (`MEMORY.md` in the workspace) that survives
//! across sessions and acts as the long-term substrate the next session
//! hydrates from. It is updated by a background forked sub-agent (the
//! `archivist` archetype) so the user-facing agent never pays the cost
//! of synthesis on its hot path.
//!
//! Extraction only runs after token-growth, tool-call, and turn-count
//! thresholds are met, so it does not fire every turn — see
//! [`SessionMemoryConfig`] for the exact knobs.
//!
//! This module is purely state-tracking: it owns the thresholds and a
//! `should_extract` decision, but the actual `spawn_subagent` call is
//! issued by the caller (the `Agent::turn` epilogue) so we avoid a
//! circular dependency with `harness::subagent_runner`.
/// Minimum number of *new* tokens (input + output) since the last
/// extraction before we consider running another extraction.
pub const DEFAULT_MIN_TOKEN_GROWTH: u64 = 4_000;
/// Minimum number of assistant tool calls since the last extraction
/// before we consider running another extraction.
pub const DEFAULT_MIN_TOOL_CALLS: u64 = 8;
/// Minimum number of turns between extractions. Prevents burst
/// extraction when the user sends many short messages in a row.
pub const DEFAULT_MIN_TURNS_BETWEEN: u64 = 4;
/// Tunable thresholds for session-memory extraction.
#[derive(Debug, Clone, Copy)]
pub struct SessionMemoryConfig {
pub min_token_growth: u64,
pub min_tool_calls: u64,
pub min_turns_between: u64,
}
impl Default for SessionMemoryConfig {
fn default() -> Self {
Self {
min_token_growth: DEFAULT_MIN_TOKEN_GROWTH,
min_tool_calls: DEFAULT_MIN_TOOL_CALLS,
min_turns_between: DEFAULT_MIN_TURNS_BETWEEN,
}
}
}
/// Per-session extraction state. Tracked on the `Agent` instance so it
/// resets naturally when a new session starts.
#[derive(Debug, Clone, Default)]
pub struct SessionMemoryState {
/// Cumulative tokens observed across the whole session (via
/// `ContextGuard::update_usage`).
pub total_tokens: u64,
/// Tokens at the last completed extraction (or 0 if none yet).
pub tokens_at_last_extract: u64,
/// Turn counter at the last completed extraction.
pub turn_at_last_extract: u64,
/// Cumulative tool-call count across the session.
pub total_tool_calls: u64,
/// Tool calls observed at the last extraction.
pub tool_calls_at_last_extract: u64,
/// Current turn counter.
pub current_turn: u64,
/// Whether an extraction is in progress. While `true`, `should_extract`
/// returns false so we don't spawn overlapping background forks.
pub extraction_in_progress: bool,
}
impl SessionMemoryState {
/// Called each time the caller bumps the turn counter.
pub fn tick_turn(&mut self) {
self.current_turn = self.current_turn.saturating_add(1);
}
/// Accumulate usage from the most recent provider response.
pub fn record_usage(&mut self, total_used_tokens: u64) {
// `total_used_tokens` is cumulative per-response (prompt + output);
// we want monotonic growth so take the max against what we've
// already recorded. This is robust to providers that report
// smaller numbers when tool-only turns happen.
if total_used_tokens > self.total_tokens {
self.total_tokens = total_used_tokens;
}
}
/// Accumulate a tool-call count from the turn just finished.
pub fn record_tool_calls(&mut self, n: usize) {
self.total_tool_calls = self.total_tool_calls.saturating_add(n as u64);
}
/// Decide whether a background session-memory extraction should run
/// right now. The rule: all three deltas (tokens, tool calls, turns)
/// must have grown past their thresholds since the last extraction,
/// AND no other extraction is in flight.
pub fn should_extract(&self, config: &SessionMemoryConfig) -> bool {
if self.extraction_in_progress {
return false;
}
let token_growth = self
.total_tokens
.saturating_sub(self.tokens_at_last_extract);
let tool_growth = self
.total_tool_calls
.saturating_sub(self.tool_calls_at_last_extract);
let turn_growth = self.current_turn.saturating_sub(self.turn_at_last_extract);
token_growth >= config.min_token_growth
&& tool_growth >= config.min_tool_calls
&& turn_growth >= config.min_turns_between
}
/// Mark an extraction as in-progress. Must be paired with either
/// `mark_extraction_complete` or `mark_extraction_failed`.
pub fn mark_extraction_started(&mut self) {
self.extraction_in_progress = true;
}
/// Record a successful extraction. Resets the deltas so the next
/// extraction won't fire until the thresholds are re-crossed.
pub fn mark_extraction_complete(&mut self) {
self.extraction_in_progress = false;
self.tokens_at_last_extract = self.total_tokens;
self.tool_calls_at_last_extract = self.total_tool_calls;
self.turn_at_last_extract = self.current_turn;
}
/// Record a failed extraction. Leaves the deltas alone so the next
/// turn can retry, but clears the in-progress flag.
pub fn mark_extraction_failed(&mut self) {
self.extraction_in_progress = false;
}
}
/// The prompt the main agent hands to a spawned archivist sub-agent when
/// session-memory extraction fires. Kept in this module so the
/// extraction policy and the spawn wording live together.
pub const ARCHIVIST_EXTRACTION_PROMPT: &str =
"You are extracting durable facts from the recent conversation \
into the workspace `MEMORY.md` file. Focus on:\n\n\
- User preferences and commitments\n\
- Decisions and their rationale\n\
- Facts about external systems, people, codebases the user mentioned\n\
- Unresolved tasks worth surfacing next session\n\n\
Skip: filler dialogue, tool logs, and anything already present in \
MEMORY.md. Use the `update_memory_md` tool to append a dated bullet \
list under an `## Observations` section. Be dense — at most 8 bullets. \
Reply with a one-line confirmation when done.";
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn default_state_does_not_extract() {
let state = SessionMemoryState::default();
let cfg = SessionMemoryConfig::default();
assert!(!state.should_extract(&cfg));
}
#[test]
fn all_three_thresholds_must_be_crossed() {
let cfg = SessionMemoryConfig::default();
// Only token threshold crossed → no.
let mut s = SessionMemoryState::default();
s.total_tokens = DEFAULT_MIN_TOKEN_GROWTH + 1;
assert!(!s.should_extract(&cfg));
// Tokens + tool calls, no turn growth → no.
s.total_tool_calls = DEFAULT_MIN_TOOL_CALLS + 1;
assert!(!s.should_extract(&cfg));
// All three crossed → yes.
s.current_turn = DEFAULT_MIN_TURNS_BETWEEN + 1;
assert!(s.should_extract(&cfg));
}
#[test]
fn in_progress_suppresses_extraction() {
let cfg = SessionMemoryConfig::default();
let mut s = SessionMemoryState::default();
s.total_tokens = DEFAULT_MIN_TOKEN_GROWTH + 1;
s.total_tool_calls = DEFAULT_MIN_TOOL_CALLS + 1;
s.current_turn = DEFAULT_MIN_TURNS_BETWEEN + 1;
assert!(s.should_extract(&cfg));
s.mark_extraction_started();
assert!(!s.should_extract(&cfg));
}
#[test]
fn mark_complete_resets_deltas() {
let cfg = SessionMemoryConfig::default();
let mut s = SessionMemoryState::default();
s.total_tokens = 10_000;
s.total_tool_calls = 15;
s.current_turn = 10;
s.mark_extraction_started();
s.mark_extraction_complete();
// Immediately after completion no further extraction should
// fire until the deltas are re-crossed.
assert!(!s.should_extract(&cfg));
// Grow each counter past threshold again.
s.total_tokens += DEFAULT_MIN_TOKEN_GROWTH;
s.total_tool_calls += DEFAULT_MIN_TOOL_CALLS;
s.current_turn += DEFAULT_MIN_TURNS_BETWEEN;
assert!(s.should_extract(&cfg));
}
#[test]
fn mark_failed_leaves_deltas_intact() {
let cfg = SessionMemoryConfig::default();
let mut s = SessionMemoryState::default();
s.total_tokens = DEFAULT_MIN_TOKEN_GROWTH + 1;
s.total_tool_calls = DEFAULT_MIN_TOOL_CALLS + 1;
s.current_turn = DEFAULT_MIN_TURNS_BETWEEN + 1;
s.mark_extraction_started();
s.mark_extraction_failed();
// Should still fire on the next attempt because the
// "last_extract" counters were not advanced.
assert!(s.should_extract(&cfg));
}
#[test]
fn record_usage_is_monotonic() {
let mut s = SessionMemoryState::default();
s.record_usage(5_000);
s.record_usage(3_000); // regression — must not decrease.
assert_eq!(s.total_tokens, 5_000);
s.record_usage(7_500);
assert_eq!(s.total_tokens, 7_500);
}
#[test]
fn tick_turn_increments() {
let mut s = SessionMemoryState::default();
s.tick_turn();
s.tick_turn();
s.tick_turn();
assert_eq!(s.current_turn, 3);
}
#[test]
fn record_tool_calls_accumulates() {
let mut s = SessionMemoryState::default();
s.record_tool_calls(3);
s.record_tool_calls(2);
assert_eq!(s.total_tool_calls, 5);
}
}
@@ -0,0 +1,176 @@
//! Stage 1: Tool-result budget.
//!
//! Apply a per-call byte cap to a raw tool result *before* it enters the
//! conversation history. This is the cheapest stage because it operates
//! on fresh bytes that have not yet been sent to the inference backend —
//! it does not mutate existing history and therefore does not break the
//! KV-cache prefix.
//!
//! A future iteration could park the overflow in a "stored surrogate"
//! and reference it later if the model asks for the full body. For now
//! OpenHuman does the simpler thing: truncate in-place with a size
//! marker the model can use to decide whether to re-run the tool with a
//! narrower query.
//!
//! This stage is called from `Agent::execute_tool_call` once the tool
//! has returned its output and before that output is packaged into a
//! `ToolResultMessage`.
use std::fmt::Write as _;
/// Default per-tool-result budget. Chosen to keep a single oversized
/// result from blowing out the prompt while still leaving room for
/// moderately chunky outputs (directory listings, small file contents,
/// condensed HTTP bodies).
pub const DEFAULT_TOOL_RESULT_BUDGET_BYTES: usize = 16 * 1024;
/// Number of trailing bytes reserved for the truncation marker. The
/// effective head capacity is `budget - TRAILER_RESERVED`.
const TRAILER_RESERVED: usize = 256;
/// Outcome of a budget application, for tracing.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct BudgetOutcome {
/// Byte length of the original content.
pub original_bytes: usize,
/// Byte length of the returned content (`== original_bytes` when the
/// result fit inside the budget).
pub final_bytes: usize,
/// `true` if the content was truncated.
pub truncated: bool,
}
impl BudgetOutcome {
pub fn unchanged(len: usize) -> Self {
Self {
original_bytes: len,
final_bytes: len,
truncated: false,
}
}
}
/// Apply the tool-result budget to `content`.
///
/// If `content` fits in `budget_bytes`, returns it unchanged. Otherwise
/// returns a truncated prefix followed by a human-readable marker like
/// `\n\n[… 42_384 bytes truncated by tool_result_budget …]`. The cut is
/// made at a UTF-8 character boundary so the returned string is always
/// valid UTF-8.
pub fn apply_tool_result_budget(content: String, budget_bytes: usize) -> (String, BudgetOutcome) {
let original_bytes = content.len();
if budget_bytes == 0 || original_bytes <= budget_bytes {
return (content, BudgetOutcome::unchanged(original_bytes));
}
// Reserve room for the trailer. If the budget is smaller than the
// reservation we still emit the marker; the only guarantee is that
// the final string is shorter than the original.
let head_capacity = budget_bytes.saturating_sub(TRAILER_RESERVED).max(1);
// Walk char indices forward until we cross the head capacity. The
// last char fully inside the head is where we cut.
let mut cut = 0usize;
for (idx, ch) in content.char_indices() {
let next = idx + ch.len_utf8();
if next > head_capacity {
break;
}
cut = next;
}
// Extremely short content (single multi-byte char) — guarantee at
// least one character makes it into the head so we don't emit a
// zero-byte head.
if cut == 0 {
cut = content
.char_indices()
.next()
.map(|(_, c)| c.len_utf8())
.unwrap_or(0);
}
let dropped_bytes = original_bytes.saturating_sub(cut);
let mut out = String::with_capacity(cut + TRAILER_RESERVED);
out.push_str(&content[..cut]);
// Hard separator so the marker is easy for humans AND the model to
// recognise when it appears inside a tool_result block.
let _ = write!(
out,
"\n\n[… {dropped_bytes} bytes truncated by tool_result_budget — re-run with a narrower query to see the rest …]"
);
let final_bytes = out.len();
(
out,
BudgetOutcome {
original_bytes,
final_bytes,
truncated: true,
},
)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn small_content_passes_through_unchanged() {
let input = "hello world".to_string();
let (out, outcome) = apply_tool_result_budget(input.clone(), 1024);
assert_eq!(out, input);
assert!(!outcome.truncated);
assert_eq!(outcome.original_bytes, outcome.final_bytes);
}
#[test]
fn content_at_exact_budget_is_unchanged() {
let input = "x".repeat(100);
let (out, outcome) = apply_tool_result_budget(input.clone(), 100);
assert_eq!(out, input);
assert!(!outcome.truncated);
}
#[test]
fn oversized_content_is_truncated_with_marker() {
let input = "x".repeat(10_000);
let (out, outcome) = apply_tool_result_budget(input, 1024);
assert!(outcome.truncated);
assert!(out.len() < 10_000);
assert!(out.contains("truncated by tool_result_budget"));
// Marker should include the dropped byte count.
assert!(out.contains("bytes truncated"));
}
#[test]
fn truncation_respects_utf8_boundaries() {
// Each "é" is 2 bytes. 600 of them = 1200 bytes.
let input: String = "é".repeat(600);
let (out, outcome) = apply_tool_result_budget(input, 500);
assert!(outcome.truncated);
// Must be valid UTF-8 — just dereferencing is enough.
let _ = out.as_str();
// Head should contain only full "é" characters (no half-byte).
let head_end = out.find("\n\n[").unwrap();
let head = &out[..head_end];
assert!(head.chars().all(|c| c == 'é'));
}
#[test]
fn zero_budget_is_noop() {
let input = "keep me".to_string();
let (out, outcome) = apply_tool_result_budget(input.clone(), 0);
assert_eq!(out, input);
assert!(!outcome.truncated);
}
#[test]
fn outcome_reports_correct_byte_counts() {
let input = "x".repeat(5_000);
let (out, outcome) = apply_tool_result_budget(input, 1024);
assert_eq!(outcome.original_bytes, 5_000);
assert_eq!(outcome.final_bytes, out.len());
assert!(outcome.truncated);
}
}
@@ -0,0 +1,306 @@
//! Built-in [`AgentDefinition`]s derived from [`AgentArchetype`].
//!
//! These cover the eight historical archetypes plus a synthetic `fork`
//! definition that the runner uses for byte-exact prompt-cache reuse.
//! Custom YAML definitions loaded later override any built-in with the
//! same id.
use super::archetypes::AgentArchetype;
use super::definition::{
AgentDefinition, DefinitionSource, ModelSpec, PromptSource, SandboxMode, ToolScope,
};
// `PromptSource::File` is unused in this module — built-ins use
// `PromptSource::Inline` with text baked in by `include_str!`. Custom
// definitions loaded from TOML may still use `File`.
/// All built-in definitions, in stable order.
pub fn all() -> Vec<AgentDefinition> {
let mut out: Vec<AgentDefinition> = AgentArchetype::all()
.iter()
.map(|arch| from_archetype(*arch))
.collect();
out.push(fork_definition());
out
}
/// Construct an [`AgentDefinition`] for one [`AgentArchetype`].
///
/// Reads `default_model_hint`, `allowed_tools`, `default_max_iterations`,
/// and `sandbox_mode` from the existing archetype metadata so this stays
/// a single source of truth. The system prompt body is baked into the
/// binary via [`archetype_prompt_body`] so built-in sub-agents work
/// regardless of workspace state.
pub fn from_archetype(arch: AgentArchetype) -> AgentDefinition {
let id = arch.to_string();
let when_to_use = when_to_use_for(arch).to_string();
let display_name = Some(display_name_for(arch).to_string());
let system_prompt = PromptSource::Inline(archetype_prompt_body(arch).to_string());
let tools = match arch.allowed_tools() {
// SkillsAgent: dynamic — at spawn time the runner picks up all
// currently-loaded skill tools (filtered by `skill_filter` if set).
None => ToolScope::Wildcard,
Some(allowed) => ToolScope::Named(allowed.iter().map(|s| (*s).to_string()).collect()),
};
// SkillsAgent's default skill filter is None — meaning "all skill tools".
// Per-API specialists (Notion, Gmail, …) are layered on top by setting
// `skill_filter` either in a custom TOML definition or as a per-spawn arg.
let skill_filter = None;
// Category filter: the SkillsAgent archetype is *only* allowed to use
// skill-bridged tools. Every other archetype is free to use system
// tools (their tool whitelists further narrow the actual set).
let category_filter = match arch {
AgentArchetype::SkillsAgent => Some(crate::openhuman::tools::ToolCategory::Skill),
_ => None,
};
// Sub-agents always run with the cheaper, narrower archetype model
// hint. Use `ModelSpec::Inherit` if you want them to share the parent's
// pinned model — see the `fork` synthetic definition below.
let model = ModelSpec::Hint(arch.default_model_hint().to_string());
let sandbox_mode = match arch.sandbox_mode() {
"sandboxed" => SandboxMode::Sandboxed,
"read_only" => SandboxMode::ReadOnly,
_ => SandboxMode::None,
};
// Code executor / tool maker / skills agent need the safety preamble
// (they actually touch the world). Pure read-only roles strip it.
let omit_safety_preamble = !matches!(
arch,
AgentArchetype::CodeExecutor | AgentArchetype::ToolMaker | AgentArchetype::SkillsAgent
);
AgentDefinition {
id,
when_to_use,
display_name,
system_prompt,
omit_identity: true,
omit_memory_context: true,
omit_safety_preamble,
omit_skills_catalog: true,
model,
temperature: 0.4,
tools,
disallowed_tools: vec![],
skill_filter,
category_filter,
max_iterations: arch.default_max_iterations(),
timeout_secs: None,
sandbox_mode,
background: arch.is_background(),
uses_fork_context: false,
source: DefinitionSource::Builtin,
}
}
/// The synthetic `fork` definition. Tells the runner to bypass normal
/// prompt construction and replay the parent's exact rendered system
/// prompt + tool schemas + message prefix from
/// [`super::fork_context::ForkContext`]. The OpenAI-compatible backend's
/// automatic prefix caching turns this into a real token win.
pub fn fork_definition() -> AgentDefinition {
AgentDefinition {
id: "fork".into(),
when_to_use: "Spawn a parallel sub-task that shares the parent's full system \
prompt, tool set, and message history byte-for-byte. Use when \
decomposing a task into independent parallel work streams that \
benefit from prefix-cache reuse on the inference backend."
.into(),
display_name: Some("Fork".into()),
// Prompt source is irrelevant — the runner reads from ForkContext.
system_prompt: PromptSource::Inline(String::new()),
// Fork preserves bytes — DO NOT strip anything from the parent's prompt.
omit_identity: false,
omit_memory_context: false,
omit_safety_preamble: false,
omit_skills_catalog: false,
model: ModelSpec::Inherit,
// Inherit the parent's temperature too — set to a sentinel that the
// runner replaces with the parent's actual temp at spawn time.
// (We use 0.7 here as a safe default for documentation; the runner
// overrides it from `ParentExecutionContext::temperature`.)
temperature: 0.7,
tools: ToolScope::Wildcard,
disallowed_tools: vec![],
skill_filter: None,
category_filter: None,
// Fork inherits the parent's max iterations from the runtime.
max_iterations: 15,
timeout_secs: None,
sandbox_mode: SandboxMode::None,
background: false,
uses_fork_context: true,
source: DefinitionSource::Builtin,
}
}
/// Returns the bundled markdown body for an archetype's system prompt.
///
/// Files are baked into the binary via `include_str!` at compile time so
/// built-in sub-agents always have a prompt to work with, even when the
/// workspace `agent/prompts/` directory doesn't exist or has been
/// modified.
fn archetype_prompt_body(arch: AgentArchetype) -> &'static str {
match arch {
AgentArchetype::Orchestrator => include_str!("../prompts/ORCHESTRATOR.md"),
AgentArchetype::Planner => include_str!("../prompts/PLANNER.md"),
AgentArchetype::CodeExecutor => include_str!("../prompts/archetypes/code_executor.md"),
AgentArchetype::SkillsAgent => include_str!("../prompts/archetypes/skills_agent.md"),
// ToolMaker shares the code_executor prompt — both write code in
// a sandbox; ToolMaker's bounded scope is enforced via
// `max_iterations` and `disallowed_tools`.
AgentArchetype::ToolMaker => include_str!("../prompts/archetypes/code_executor.md"),
AgentArchetype::Researcher => include_str!("../prompts/archetypes/researcher.md"),
AgentArchetype::Critic => include_str!("../prompts/archetypes/critic.md"),
AgentArchetype::Archivist => include_str!("../prompts/archetypes/archivist.md"),
}
}
fn when_to_use_for(arch: AgentArchetype) -> &'static str {
match arch {
AgentArchetype::Orchestrator => {
"Staff Engineer — routes, judges quality, synthesises. Never writes code itself. \
You should not normally spawn another orchestrator from inside one."
}
AgentArchetype::Planner => {
"Architect — break a complex task into a small DAG of subtasks with \
explicit acceptance criteria. Read-only; produces JSON, not code."
}
AgentArchetype::CodeExecutor => {
"Sandboxed developer — writes, runs, and debugs code until tests pass. \
Use for any task that requires producing or modifying source files \
and exercising them with shell or test commands."
}
AgentArchetype::SkillsAgent => {
"Skill tool specialist — executes installed QuickJS skill tools \
(Notion, Gmail, …). Use when the task should be completed via a \
user-installed skill rather than raw HTTP/file I/O. Pair with a \
`skill_filter` argument to scope to a single skill."
}
AgentArchetype::ToolMaker => {
"Self-healer — writes a polyfill script when a required command is \
missing on the host. Very narrow scope; max 2 iterations."
}
AgentArchetype::Researcher => {
"Web & docs crawler — reads real documentation, compresses to dense \
markdown. Use for any task that requires looking up external knowledge."
}
AgentArchetype::Critic => {
"Adversarial reviewer — reviews diffs and code against project rules, \
flags vulnerabilities, regressions, and missing tests. Read-only."
}
AgentArchetype::Archivist => {
"Background librarian — extracts lessons from a completed session, \
updates MEMORY.md, and indexes to FTS5. Runs cheap and slow."
}
}
}
fn display_name_for(arch: AgentArchetype) -> &'static str {
match arch {
AgentArchetype::Orchestrator => "Orchestrator",
AgentArchetype::Planner => "Planner",
AgentArchetype::CodeExecutor => "Code Executor",
AgentArchetype::SkillsAgent => "Skills Agent",
AgentArchetype::ToolMaker => "Tool Maker",
AgentArchetype::Researcher => "Researcher",
AgentArchetype::Critic => "Critic",
AgentArchetype::Archivist => "Archivist",
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn all_definitions_present() {
let defs = all();
// 8 archetypes + 1 synthetic `fork`.
assert_eq!(defs.len(), 9);
}
#[test]
fn each_archetype_yields_an_id() {
for arch in AgentArchetype::all() {
let def = from_archetype(*arch);
assert_eq!(def.id, arch.to_string());
assert!(!def.when_to_use.is_empty());
assert_eq!(def.source, DefinitionSource::Builtin);
}
}
#[test]
fn code_executor_keeps_safety_preamble() {
let def = from_archetype(AgentArchetype::CodeExecutor);
assert!(!def.omit_safety_preamble);
}
#[test]
fn critic_strips_safety_preamble() {
let def = from_archetype(AgentArchetype::Critic);
assert!(def.omit_safety_preamble);
}
#[test]
fn skills_agent_uses_wildcard_tools() {
let def = from_archetype(AgentArchetype::SkillsAgent);
assert!(matches!(def.tools, ToolScope::Wildcard));
}
#[test]
fn code_executor_uses_named_tools() {
let def = from_archetype(AgentArchetype::CodeExecutor);
match def.tools {
ToolScope::Named(tools) => {
assert!(tools.iter().any(|t| t == "shell"));
assert!(tools.iter().any(|t| t == "file_write"));
}
ToolScope::Wildcard => panic!("expected named tools for code_executor"),
}
}
#[test]
fn fork_definition_has_uses_fork_context_true() {
let def = fork_definition();
assert_eq!(def.id, "fork");
assert!(def.uses_fork_context);
assert!(matches!(def.model, ModelSpec::Inherit));
assert!(matches!(def.tools, ToolScope::Wildcard));
// Fork preserves bytes — must NOT strip anything.
assert!(!def.omit_identity);
assert!(!def.omit_memory_context);
assert!(!def.omit_safety_preamble);
assert!(!def.omit_skills_catalog);
}
#[test]
fn archetype_max_iterations_is_propagated() {
let critic = from_archetype(AgentArchetype::Critic);
assert_eq!(critic.max_iterations, 5);
let tool_maker = from_archetype(AgentArchetype::ToolMaker);
assert_eq!(tool_maker.max_iterations, 2);
}
#[test]
fn sandbox_modes_map_correctly() {
assert_eq!(
from_archetype(AgentArchetype::CodeExecutor).sandbox_mode,
SandboxMode::Sandboxed
);
assert_eq!(
from_archetype(AgentArchetype::Critic).sandbox_mode,
SandboxMode::ReadOnly
);
assert_eq!(
from_archetype(AgentArchetype::Researcher).sandbox_mode,
SandboxMode::None
);
}
}
+477
View File
@@ -0,0 +1,477 @@
//! Data-driven agent definitions.
//!
//! An [`AgentDefinition`] fully specifies a sub-agent: its core prompt, model,
//! allowed tool set, runtime limits, and which sections of the parent system
//! prompt to omit. Built-in definitions are derived from
//! [`super::archetypes::AgentArchetype`] in
//! [`super::builtin_definitions`]; users can ship custom definitions as TOML
//! files under `$OPENHUMAN_WORKSPACE/agents/*.toml` (with a fallback to
//! `~/.openhuman/agents/*.toml` for user-global specialists) which override
//! built-ins on id collision. See [`super::definition_loader`] for the
//! directory scan + TOML parsing contract.
//!
//! Sub-agents are dispatched at runtime by the `spawn_subagent` tool, which
//! looks up an [`AgentDefinition`] by id in the global
//! [`AgentDefinitionRegistry`] and hands it to
//! [`super::subagent_runner::run_subagent`].
//!
//! This file intentionally has zero references to the rest of the agent
//! runtime — it is pure data so the model can be unit-tested in isolation
//! and serialised straight from disk.
use crate::openhuman::tools::ToolCategory;
use serde::{Deserialize, Serialize};
use std::path::PathBuf;
// ─────────────────────────────────────────────────────────────────────────────
// Agent definition
// ─────────────────────────────────────────────────────────────────────────────
/// A fully specified sub-agent: what it knows, what it can do, how to prompt it.
///
/// Built-ins live in [`super::builtin_definitions`]; custom ones load from
/// TOML at startup. The [`AgentDefinitionRegistry`] merges them and is the
/// single source of truth that `SpawnSubagentTool` queries.
///
/// All `omit_*` flags default to `true` for sub-agents — sub-agents are
/// narrow specialists and pay no token tax for the parent's identity,
/// memory, safety, or skills sections. Override per-archetype if a
/// section is needed.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct AgentDefinition {
// ── identity ────────────────────────────────────────────────────────
/// Unique id, referenced from `spawn_subagent { agent_id: "…" }`.
/// Convention: snake_case (e.g. `code_executor`, `notion_specialist`).
pub id: String,
/// One-line description shown in the orchestrator's `spawn_subagent`
/// tool schema so the parent model knows when to delegate to this agent.
pub when_to_use: String,
/// Optional display name for UI/logs. Falls back to `id`.
#[serde(default)]
pub display_name: Option<String>,
// ── prompt ──────────────────────────────────────────────────────────
/// Source of the sub-agent's core system prompt. Inline for TOML-defined
/// agents, or a path to a file under `agent/prompts/` for built-ins.
pub system_prompt: PromptSource,
/// Sections of the main agent's prompt to strip when this sub-agent runs.
/// Defaults to `true` (strip) — sub-agents are narrow and don't need the
/// parent's identity scaffolding.
#[serde(default = "defaults::true_")]
pub omit_identity: bool,
#[serde(default = "defaults::true_")]
pub omit_memory_context: bool,
#[serde(default = "defaults::true_")]
pub omit_safety_preamble: bool,
#[serde(default = "defaults::true_")]
pub omit_skills_catalog: bool,
// ── model ───────────────────────────────────────────────────────────
/// Model selection: inherit parent, hint to router, or pinned name.
#[serde(default)]
pub model: ModelSpec,
/// Sampling temperature. Sub-agents default to `0.4` for precision.
#[serde(default = "defaults::subagent_temperature")]
pub temperature: f64,
// ── tools ───────────────────────────────────────────────────────────
/// Either [`ToolScope::Wildcard`] (all tools the parent has) or
/// [`ToolScope::Named`] (an explicit allowlist).
#[serde(default)]
pub tools: ToolScope,
/// Tools that are explicitly banned even if `tools == Wildcard`.
/// Built-ins default-deny dangerous ops for read-only archetypes.
#[serde(default)]
pub disallowed_tools: Vec<String>,
/// If set, the resolved tool list is further filtered to only those whose
/// name starts with `{skill_filter}__`. Gives us per-API specialists
/// (Notion, Gmail, …) without enum variants. Overridable per-spawn.
#[serde(default)]
pub skill_filter: Option<String>,
/// If set, the resolved tool list is restricted to tools whose
/// [`crate::openhuman::tools::Tool::category`] matches this value.
/// This is the *primary* mechanism the orchestrator uses to spawn
/// dedicated tool-execution sub-agents:
/// - `Some(Skill)` → sub-agent only sees skill-bridge tools
/// (Notion, Gmail, Telegram, …). Pair with `ModelSpec::Hint("agentic")`
/// to route to the backend's agentic model.
/// - `Some(System)` → sub-agent only sees built-in Rust tools.
/// - `None` (default) → no category restriction; `tools` /
/// `disallowed_tools` / `skill_filter` still apply.
///
/// Category filtering happens *before* the `tools`/`disallowed_tools`
/// scope check, so a `Named` scope is a stricter-intersection override.
#[serde(default)]
pub category_filter: Option<ToolCategory>,
// ── runtime limits ──────────────────────────────────────────────────
/// Maximum tool-call iterations per spawn. Sub-agents default to a
/// shorter cap than the parent to keep cost bounded.
#[serde(default = "defaults::max_iterations")]
pub max_iterations: usize,
/// Hard wall-clock timeout per turn. `None` falls back to
/// `tool_execution_timeout_secs`.
#[serde(default)]
pub timeout_secs: Option<u64>,
/// `none` / `read_only` / `sandboxed`. Mirrors
/// [`super::archetypes::AgentArchetype::sandbox_mode`].
#[serde(default)]
pub sandbox_mode: SandboxMode,
/// If true, spawn runs in the background and the call returns
/// immediately with a placeholder. Reserved — not yet wired in v1.
#[serde(default)]
pub background: bool,
/// Marker: when true, the runner skips its normal prompt-building path
/// and uses the parent's pre-rendered prompt + tool schemas + message
/// prefix from the [`super::fork_context::ForkContext`] task-local.
/// Only the synthetic built-in `fork` definition has this set.
#[serde(default, skip_serializing_if = "is_false")]
pub uses_fork_context: bool,
// ── source bookkeeping ──────────────────────────────────────────────
/// Where this definition came from. Filled in by the loader/builder;
/// not deserialised from TOML.
#[serde(skip)]
pub source: DefinitionSource,
}
fn is_false(b: &bool) -> bool {
!b
}
impl AgentDefinition {
/// Display name with fallback to id.
pub fn display_name(&self) -> &str {
self.display_name.as_deref().unwrap_or(&self.id)
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Prompt source
// ─────────────────────────────────────────────────────────────────────────────
/// Where the sub-agent's core system prompt comes from.
#[derive(Debug, Clone, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum PromptSource {
/// Inline prompt string (custom TOML-defined agents).
Inline(String),
/// Relative path under the workspace's `prompts/` directory or under
/// `src/openhuman/agent/prompts/` for built-ins. Resolved by the runner
/// at spawn time.
File { path: String },
}
// ─────────────────────────────────────────────────────────────────────────────
// Model spec
// ─────────────────────────────────────────────────────────────────────────────
/// Model selection for a sub-agent.
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
#[serde(rename_all = "snake_case")]
pub enum ModelSpec {
/// Use the parent agent's currently-selected model at spawn time.
#[default]
Inherit,
/// Exact model name (e.g. `"neocortex-mk1"`).
Exact(String),
/// Router hint (e.g. `"reasoning"`, `"coding"`, `"local"`). Resolved
/// to a real model by the routing provider.
Hint(String),
}
impl ModelSpec {
/// Resolve this spec into the model name string the provider expects.
/// `parent_model` is the model the parent agent is using right now.
pub fn resolve(&self, parent_model: &str) -> String {
match self {
Self::Inherit => parent_model.to_string(),
Self::Exact(name) => name.clone(),
Self::Hint(hint) => format!("hint:{hint}"),
}
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Tool scope
// ─────────────────────────────────────────────────────────────────────────────
/// Which tools a sub-agent is allowed to call.
#[derive(Debug, Clone, Serialize, Deserialize, Default)]
#[serde(rename_all = "snake_case")]
pub enum ToolScope {
/// All tools the parent has (subject to `disallowed_tools` and
/// `skill_filter`).
#[default]
Wildcard,
/// An explicit allowlist of tool names. Names not present in the parent
/// registry at spawn time are silently dropped (logged at debug).
Named(Vec<String>),
}
// ─────────────────────────────────────────────────────────────────────────────
// Sandbox mode
// ─────────────────────────────────────────────────────────────────────────────
/// Sandbox mode for a sub-agent's tool execution. Mirrors the existing
/// [`super::archetypes::AgentArchetype::sandbox_mode`] string for now;
/// in the future this may map directly into a `SecurityPolicy` builder.
#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq, Default)]
#[serde(rename_all = "snake_case")]
pub enum SandboxMode {
/// No additional sandboxing beyond what the parent already enforces.
#[default]
None,
/// Read-only — write/execute tools are filtered out.
ReadOnly,
/// Drop privileges, restrict filesystem (Landlock / Bubblewrap).
Sandboxed,
}
// ─────────────────────────────────────────────────────────────────────────────
// Definition source
// ─────────────────────────────────────────────────────────────────────────────
/// Where an [`AgentDefinition`] was loaded from. Used for telemetry and
/// the `agent::list_definitions` RPC reply.
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Default)]
#[serde(tag = "kind", content = "path")]
pub enum DefinitionSource {
/// Built-in derived from an [`super::archetypes::AgentArchetype`].
#[default]
Builtin,
/// Loaded from a TOML file at the given absolute path.
File(PathBuf),
}
// ─────────────────────────────────────────────────────────────────────────────
// Defaults module — referenced by `#[serde(default = ...)]`
// ─────────────────────────────────────────────────────────────────────────────
pub(crate) mod defaults {
pub(crate) fn true_() -> bool {
true
}
pub(crate) fn subagent_temperature() -> f64 {
0.4
}
pub(crate) fn max_iterations() -> usize {
8
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Registry
// ─────────────────────────────────────────────────────────────────────────────
use anyhow::Result;
use std::collections::HashMap;
use std::path::Path;
use std::sync::OnceLock;
/// In-memory registry of all known [`AgentDefinition`]s.
///
/// One singleton instance is initialised at startup via
/// [`AgentDefinitionRegistry::init_global`]. Built-ins are registered
/// unconditionally; custom TOML definitions (if a workspace is provided)
/// are loaded next and override built-ins on `id` collision.
#[derive(Debug, Default)]
pub struct AgentDefinitionRegistry {
by_id: HashMap<String, AgentDefinition>,
/// Insertion-stable order for predictable `list()` output.
order: Vec<String>,
}
static GLOBAL: OnceLock<AgentDefinitionRegistry> = OnceLock::new();
impl AgentDefinitionRegistry {
/// Build a registry containing only the built-in definitions
/// (no TOML loading). Useful for tests.
pub fn builtins_only() -> Self {
let mut reg = Self::default();
for def in super::builtin_definitions::all() {
reg.insert(def);
}
reg
}
/// Build a registry containing built-ins plus any custom TOML
/// definitions found under `<workspace>/agents/*.toml` (and the
/// `~/.openhuman/agents/*.toml` fallback). Custom definitions
/// override built-ins on `id` collision. Files that fail to parse
/// are logged and skipped rather than aborting startup.
pub fn load(workspace: &Path) -> Result<Self> {
let mut reg = Self::builtins_only();
let custom = super::definition_loader::load_from_workspace(workspace)?;
for def in custom {
tracing::info!(
id = %def.id,
source = ?def.source,
"[agent_defs] loaded custom definition (overrides any built-in with the same id)"
);
reg.insert(def);
}
Ok(reg)
}
/// Insert (or replace) a definition by id.
pub fn insert(&mut self, def: AgentDefinition) {
let id = def.id.clone();
if self.by_id.insert(id.clone(), def).is_none() {
self.order.push(id);
}
}
/// Look up a definition by id.
pub fn get(&self, id: &str) -> Option<&AgentDefinition> {
self.by_id.get(id)
}
/// All definitions, in insertion order.
pub fn list(&self) -> Vec<&AgentDefinition> {
self.order
.iter()
.filter_map(|id| self.by_id.get(id))
.collect()
}
/// Number of registered definitions.
pub fn len(&self) -> usize {
self.by_id.len()
}
/// True when the registry has no definitions.
pub fn is_empty(&self) -> bool {
self.by_id.is_empty()
}
// ── singleton API ──────────────────────────────────────────────────
/// Initialise the global registry. Subsequent calls are no-ops (the
/// `OnceLock` only fires once); use [`Self::reload_global`] to refresh
/// custom definitions during development.
pub fn init_global(workspace: &Path) -> Result<()> {
let registry = Self::load(workspace)?;
match GLOBAL.set(registry) {
Ok(()) => {
tracing::info!(
"[agent_defs] global registry initialised with {} definitions",
GLOBAL.get().map(|r| r.len()).unwrap_or(0)
);
Ok(())
}
Err(_) => {
tracing::debug!("[agent_defs] global registry already initialised; ignoring");
Ok(())
}
}
}
/// Initialise the global registry with builtins only (no workspace
/// scan). Used by tests and by callers that don't have a workspace.
pub fn init_global_builtins() -> Result<()> {
let registry = Self::builtins_only();
let _ = GLOBAL.set(registry);
Ok(())
}
/// Borrow the global registry, if initialised.
pub fn global() -> Option<&'static Self> {
GLOBAL.get()
}
}
#[cfg(test)]
mod tests {
use super::*;
fn make_def(id: &str) -> AgentDefinition {
AgentDefinition {
id: id.into(),
when_to_use: "test".into(),
display_name: None,
system_prompt: PromptSource::Inline("system".into()),
omit_identity: true,
omit_memory_context: true,
omit_safety_preamble: true,
omit_skills_catalog: true,
model: ModelSpec::Inherit,
temperature: 0.4,
tools: ToolScope::Wildcard,
disallowed_tools: vec![],
skill_filter: None,
category_filter: None,
max_iterations: 8,
timeout_secs: None,
sandbox_mode: SandboxMode::None,
background: false,
uses_fork_context: false,
source: DefinitionSource::Builtin,
}
}
#[test]
fn registry_insert_and_lookup() {
let mut reg = AgentDefinitionRegistry::default();
reg.insert(make_def("alpha"));
reg.insert(make_def("beta"));
assert_eq!(reg.len(), 2);
assert!(reg.get("alpha").is_some());
assert!(reg.get("beta").is_some());
assert!(reg.get("missing").is_none());
}
#[test]
fn registry_replace_preserves_order() {
let mut reg = AgentDefinitionRegistry::default();
reg.insert(make_def("alpha"));
reg.insert(make_def("beta"));
let mut updated = make_def("alpha");
updated.when_to_use = "replaced".into();
reg.insert(updated);
let list: Vec<&str> = reg.list().iter().map(|d| d.id.as_str()).collect();
assert_eq!(list, vec!["alpha", "beta"]);
assert_eq!(reg.get("alpha").unwrap().when_to_use, "replaced");
}
#[test]
fn model_spec_resolve_inherit_uses_parent() {
let spec = ModelSpec::Inherit;
assert_eq!(spec.resolve("parent-model"), "parent-model");
}
#[test]
fn model_spec_resolve_exact_uses_name() {
let spec = ModelSpec::Exact("kimi-k2".into());
assert_eq!(spec.resolve("parent-model"), "kimi-k2");
}
#[test]
fn model_spec_resolve_hint_prefixes_router_marker() {
let spec = ModelSpec::Hint("coding".into());
assert_eq!(spec.resolve("parent-model"), "hint:coding");
}
#[test]
fn display_name_falls_back_to_id() {
let def = make_def("alpha");
assert_eq!(def.display_name(), "alpha");
let mut def2 = make_def("beta");
def2.display_name = Some("Beta Specialist".into());
assert_eq!(def2.display_name(), "Beta Specialist");
}
}
@@ -0,0 +1,240 @@
//! Loads custom [`AgentDefinition`] files from disk.
//!
//! Custom definitions live as TOML files under `<workspace>/agents/*.toml`,
//! with a fallback to `~/.openhuman/agents/*.toml` for user-global
//! specialists. Each file defines exactly one definition.
//!
//! TOML (rather than YAML) is used for consistency with the rest of
//! OpenHuman's config system, which already depends on the `toml` crate
//! and uses TOML for its main config file.
//!
//! The loader is intentionally lenient: it logs and skips files that fail
//! to parse rather than aborting startup, so a single broken specialist
//! never breaks the rest of the system.
use super::definition::{AgentDefinition, DefinitionSource};
use anyhow::{Context, Result};
use std::fs;
use std::path::{Path, PathBuf};
/// Load all custom definitions from `<workspace>/agents/` and the
/// `~/.openhuman/agents/` fallback. Returns an empty Vec when neither
/// directory exists.
pub fn load_from_workspace(workspace: &Path) -> Result<Vec<AgentDefinition>> {
let mut out = Vec::new();
let mut seen_dirs: Vec<PathBuf> = Vec::new();
let workspace_dir = workspace.join("agents");
if workspace_dir.is_dir() {
load_dir(&workspace_dir, &mut out)?;
seen_dirs.push(workspace_dir);
}
if let Some(home_dir) = user_home_agents_dir() {
if home_dir.is_dir() && !seen_dirs.contains(&home_dir) {
load_dir(&home_dir, &mut out)?;
}
}
Ok(out)
}
/// Load every `.toml` file in a single directory (non-recursive). Files
/// that fail to parse are logged and skipped.
pub fn load_dir(dir: &Path, out: &mut Vec<AgentDefinition>) -> Result<()> {
let entries =
fs::read_dir(dir).with_context(|| format!("reading agents dir {}", dir.display()))?;
for entry in entries {
let entry = match entry {
Ok(e) => e,
Err(err) => {
tracing::warn!(
dir = %dir.display(),
error = %err,
"[agent_defs] failed to read directory entry, skipping"
);
continue;
}
};
let path = entry.path();
if !path.is_file() {
continue;
}
let Some(ext) = path.extension().and_then(|e| e.to_str()) else {
continue;
};
if ext != "toml" {
continue;
}
match load_file(&path) {
Ok(def) => {
tracing::debug!(
id = %def.id,
path = %path.display(),
"[agent_defs] loaded custom definition"
);
out.push(def);
}
Err(err) => {
tracing::warn!(
path = %path.display(),
error = %err,
"[agent_defs] failed to load custom definition, skipping"
);
}
}
}
Ok(())
}
/// Load a single TOML file as an [`AgentDefinition`]. Stamps `source` to
/// the absolute path.
pub fn load_file(path: &Path) -> Result<AgentDefinition> {
let content =
fs::read_to_string(path).with_context(|| format!("reading {}", path.display()))?;
let mut def: AgentDefinition = toml::from_str(&content)
.with_context(|| format!("parsing {} as AgentDefinition TOML", path.display()))?;
def.source = DefinitionSource::File(path.to_path_buf());
Ok(def)
}
fn user_home_agents_dir() -> Option<PathBuf> {
// Honour OPENHUMAN_HOME first if set; otherwise ~/.openhuman.
if let Ok(custom) = std::env::var("OPENHUMAN_HOME") {
return Some(PathBuf::from(custom).join("agents"));
}
dirs::home_dir().map(|h| h.join(".openhuman").join("agents"))
}
#[cfg(test)]
mod tests {
use super::*;
use std::io::Write;
fn write_toml(path: &Path, contents: &str) {
let mut f = fs::File::create(path).unwrap();
f.write_all(contents.as_bytes()).unwrap();
}
fn fresh_workspace() -> tempfile::TempDir {
tempfile::tempdir().unwrap()
}
// NOTE: TOML parsing is positional. Top-level scalars MUST come
// before any `[table]` header — once a header opens, every line
// below it lives inside that table.
const NOTION_TOML: &str = r#"
id = "notion_specialist"
when_to_use = "Delegate Notion queries to a focused specialist."
display_name = "Notion Specialist"
temperature = 0.4
skill_filter = "notion"
max_iterations = 5
[system_prompt]
inline = "You are the Notion specialist. Use only Notion tools."
[model]
hint = "agentic"
"#;
#[test]
fn loads_single_definition_from_workspace() {
let ws = fresh_workspace();
let agents_dir = ws.path().join("agents");
fs::create_dir_all(&agents_dir).unwrap();
write_toml(&agents_dir.join("notion.toml"), NOTION_TOML);
let defs = load_from_workspace(ws.path()).unwrap();
assert_eq!(defs.len(), 1);
let def = &defs[0];
assert_eq!(def.id, "notion_specialist");
assert_eq!(def.skill_filter.as_deref(), Some("notion"));
assert_eq!(def.max_iterations, 5);
assert!(matches!(def.source, DefinitionSource::File(_)));
}
#[test]
fn empty_when_no_agents_dir() {
let ws = fresh_workspace();
let defs = load_from_workspace(ws.path()).unwrap();
assert!(defs.is_empty());
}
#[test]
fn ignores_non_toml_files() {
let ws = fresh_workspace();
let agents_dir = ws.path().join("agents");
fs::create_dir_all(&agents_dir).unwrap();
write_toml(&agents_dir.join("readme.md"), "not a definition");
write_toml(&agents_dir.join("notion.toml"), NOTION_TOML);
let defs = load_from_workspace(ws.path()).unwrap();
assert_eq!(defs.len(), 1);
}
#[test]
fn skips_malformed_files_without_aborting() {
let ws = fresh_workspace();
let agents_dir = ws.path().join("agents");
fs::create_dir_all(&agents_dir).unwrap();
write_toml(&agents_dir.join("broken.toml"), "id = \"broken\" [oops");
write_toml(&agents_dir.join("notion.toml"), NOTION_TOML);
let defs = load_from_workspace(ws.path()).unwrap();
// The broken file is skipped; the valid one still loads.
assert_eq!(defs.len(), 1);
assert_eq!(defs[0].id, "notion_specialist");
}
#[test]
fn registry_load_merges_builtins_and_custom() {
let ws = fresh_workspace();
let agents_dir = ws.path().join("agents");
fs::create_dir_all(&agents_dir).unwrap();
write_toml(&agents_dir.join("notion.toml"), NOTION_TOML);
let reg = super::super::definition::AgentDefinitionRegistry::load(ws.path()).unwrap();
// The built-in set is allowed to grow over time (new archetypes,
// additional synthetic definitions), so assert presence of the
// specific ids we care about rather than a fixed total count.
assert!(
reg.len() > 1,
"expected at least one built-in plus the custom definition"
);
assert!(reg.get("notion_specialist").is_some());
assert!(reg.get("code_executor").is_some());
assert!(reg.get("fork").is_some());
}
#[test]
fn custom_definition_overrides_same_id_builtin() {
let ws = fresh_workspace();
let agents_dir = ws.path().join("agents");
fs::create_dir_all(&agents_dir).unwrap();
// Override the built-in `code_executor` with a custom one.
write_toml(
&agents_dir.join("code_executor.toml"),
r#"
id = "code_executor"
when_to_use = "CUSTOM OVERRIDE"
[system_prompt]
inline = "custom prompt"
[tools]
wildcard = {}
"#,
);
let reg = super::super::definition::AgentDefinitionRegistry::load(ws.path()).unwrap();
// Still 9 — same id replaced built-in in place.
assert_eq!(reg.len(), 9);
let def = reg.get("code_executor").unwrap();
assert_eq!(def.when_to_use, "CUSTOM OVERRIDE");
assert!(matches!(def.source, DefinitionSource::File(_)));
}
}
+125 -33
View File
@@ -232,6 +232,22 @@ async fn plan_tasks(
}
/// Execute all tasks in a single DAG level concurrently.
///
/// Each task is dispatched through the unified
/// [`super::subagent_runner::run_subagent`] helper, which:
/// - Looks up the built-in [`super::definition::AgentDefinition`] for
/// the node's archetype.
/// - Resolves model + tool filtering + narrow prompt construction.
/// - Runs the sub-agent's inner tool-call loop using the parent's
/// provider via the [`super::fork_context::PARENT_CONTEXT`] task-local
/// that the orchestrator sets up earlier in
/// [`super::subagent_runner`].
///
/// Per-archetype overrides from
/// [`crate::openhuman::config::OrchestratorConfig::archetypes`] (model,
/// temperature, max_tool_iterations, timeout_secs, sandbox) are layered
/// on top of the built-in definition before dispatch.
#[allow(clippy::too_many_arguments)]
async fn execute_level(
dag: &TaskDag,
task_ids: &[String],
@@ -239,7 +255,7 @@ async fn execute_level(
_config: &Config,
_provider: &dyn Provider,
_memory: Arc<dyn Memory>,
session_id: &str,
_session_id: &str,
semaphore: Arc<tokio::sync::Semaphore>,
) -> Vec<SubAgentResult> {
let mut join_set: JoinSet<SubAgentResult> = JoinSet::new();
@@ -254,9 +270,7 @@ async fn execute_level(
let description = node.description.clone();
let acceptance = node.acceptance_criteria.clone();
let tid = task_id.clone();
let _sid = session_id.to_string();
let model = resolve_model(archetype, orch_config);
let _temperature = resolve_temperature(archetype, orch_config);
let semaphore_clone = semaphore.clone();
let timeout = resolve_timeout(archetype, orch_config);
// Collect context from completed dependencies.
@@ -270,24 +284,16 @@ async fn execute_level(
})
.collect();
let _prompt = if dep_context.is_empty() {
format!("Task: {description}\n\nAcceptance criteria: {acceptance}")
} else {
format!(
"Context from prior tasks:\n{dep_context}\n\
Task: {description}\n\nAcceptance criteria: {acceptance}"
)
};
// Build the sub-agent prompt with the task description and
// acceptance criteria; dependency context (if any) flows
// through the runner's `SubagentRunOptions::context` field.
let prompt = format!("Task: {description}\n\nAcceptance criteria: {acceptance}");
let context_blob = (!dep_context.is_empty()).then_some(dep_context);
// Each sub-agent runs as a single-shot provider call for now.
// Phase 3 will upgrade this to full tool-loop sub-agents.
let _system_prompt = format!(
"You are the {archetype} agent. Complete the assigned task precisely.\n\
Do not deviate from the task description. Be concise."
);
let model_clone = model.clone();
let _timeout = timeout;
let semaphore_clone = semaphore.clone();
// Resolve the built-in definition for this archetype, layered
// with any per-archetype config overrides.
let mut definition = super::builtin_definitions::from_archetype(archetype);
apply_archetype_overrides(&mut definition, archetype, orch_config);
join_set.spawn(async move {
let _permit = semaphore_clone
@@ -296,25 +302,65 @@ async fn execute_level(
.expect("semaphore closed");
let start = Instant::now();
// For now, sub-agents use a simple prompt (no tool loop).
// This will be upgraded when archetype-specific tool subsets are wired.
let result_text = format!(
"[placeholder — no execution] {archetype} sub-agent would execute here\n\
Task: {description}\nModel: {model_clone}\nTimeout: {timeout:?}"
);
let options = super::subagent_runner::SubagentRunOptions {
skill_filter_override: None,
category_filter_override: None,
context: context_blob,
task_id: Some(tid.clone()),
};
let outcome_fut = super::subagent_runner::run_subagent(&definition, &prompt, options);
let outcome = match tokio::time::timeout(timeout, outcome_fut).await {
Ok(Ok(out)) => out,
Ok(Err(err)) => {
tracing::warn!(
task_id = %tid,
archetype = %archetype,
error = %err,
"[orchestrator] sub-agent failed"
);
return SubAgentResult {
task_id: tid,
success: false,
output: format!("sub-agent failed: {err}"),
artifacts: Vec::new(),
cost_microdollars: 0,
duration: start.elapsed(),
};
}
Err(_) => {
tracing::warn!(
task_id = %tid,
archetype = %archetype,
timeout_secs = timeout.as_secs(),
"[orchestrator] sub-agent timed out"
);
return SubAgentResult {
task_id: tid,
success: false,
output: format!("sub-agent timed out after {} seconds", timeout.as_secs()),
artifacts: Vec::new(),
cost_microdollars: 0,
duration: start.elapsed(),
};
}
};
tracing::debug!(
"[orchestrator] sub-agent {archetype} placeholder task {tid} in {:?}",
start.elapsed()
task_id = %outcome.task_id,
archetype = %archetype,
iterations = outcome.iterations,
output_chars = outcome.output.chars().count(),
"[orchestrator] sub-agent completed"
);
SubAgentResult {
task_id: tid,
success: false,
output: result_text,
task_id: outcome.task_id,
success: true,
output: outcome.output,
artifacts: Vec::new(),
cost_microdollars: 0,
duration: start.elapsed(),
duration: outcome.elapsed,
}
});
}
@@ -331,6 +377,52 @@ async fn execute_level(
results
}
/// Apply per-archetype config overrides on top of a built-in
/// [`super::definition::AgentDefinition`].
fn apply_archetype_overrides(
definition: &mut super::definition::AgentDefinition,
archetype: AgentArchetype,
orch_config: &OrchestratorConfig,
) {
let key = archetype.to_string();
let Some(over) = orch_config.archetypes.get(&key) else {
return;
};
if let Some(model) = over.model.as_ref() {
// The override is a raw model name — store as Exact so the
// runner uses it verbatim regardless of the parent's model.
definition.model = super::definition::ModelSpec::Exact(model.clone());
}
if let Some(temperature) = over.temperature {
definition.temperature = temperature;
}
if let Some(max_iter) = over.max_tool_iterations {
definition.max_iterations = max_iter;
}
if let Some(secs) = over.timeout_secs {
definition.timeout_secs = Some(secs);
}
if let Some(sb) = over.sandbox.as_deref() {
definition.sandbox_mode = match sb {
"sandboxed" => super::definition::SandboxMode::Sandboxed,
"read_only" => super::definition::SandboxMode::ReadOnly,
"none" | "" => super::definition::SandboxMode::None,
other => {
tracing::warn!(
archetype = %archetype,
definition_id = %definition.id,
value = %other,
"[orchestrator] unknown sandbox override — falling back to SandboxMode::None. Expected one of: sandboxed, read_only, none"
);
super::definition::SandboxMode::None
}
};
}
if let Some(prompt_override) = over.system_prompt.as_ref() {
definition.system_prompt = super::definition::PromptSource::Inline(prompt_override.clone());
}
}
/// The Orchestrator reviews results from a completed level and decides next action.
async fn review_level(
dag: &TaskDag,
+325
View File
@@ -0,0 +1,325 @@
//! Task-local plumbing that lets `SpawnSubagentTool` reach the parent
//! agent's runtime context (provider, tools, model, …) without widening
//! the [`crate::openhuman::tools::Tool`] trait.
//!
//! Two distinct task-locals live here:
//!
//! 1. [`PARENT_CONTEXT`] — set by the parent [`crate::openhuman::agent::Agent`]
//! around its `turn` so that any tool executing inside that turn (in
//! particular `spawn_subagent`) can read the parent's provider, tool
//! list, and model information.
//!
//! 2. [`FORK_CONTEXT`] — set only when the parent dispatches a `fork`-mode
//! sub-agent. Carries the parent's *exact* rendered system prompt, tool
//! schemas, and message prefix so the forked child can replay the same
//! bytes and the inference backend's automatic prefix caching kicks in.
//!
//! Both contexts are stashed in `Arc`s so that cloning into the child
//! costs a refcount bump rather than a full copy.
use crate::openhuman::config::{AgentConfig, IdentityConfig};
use crate::openhuman::memory::Memory;
use crate::openhuman::providers::{ChatMessage, Provider};
use crate::openhuman::skills::Skill;
use crate::openhuman::tools::{Tool, ToolSpec};
use std::path::PathBuf;
use std::sync::Arc;
// ─────────────────────────────────────────────────────────────────────────────
// Parent execution context
// ─────────────────────────────────────────────────────────────────────────────
/// Snapshot of the parent agent's runtime, made available to any tool
/// running inside [`crate::openhuman::agent::Agent::turn`] via the
/// [`PARENT_CONTEXT`] task-local.
///
/// All heavy fields are `Arc`-shared so cloning the context for sub-agents
/// is essentially free.
#[derive(Clone)]
pub struct ParentExecutionContext {
/// Parent's provider — sub-agents call into the same instance so
/// connection pools, retry budgets, and credentials are shared.
pub provider: Arc<dyn Provider>,
/// Parent's full tool registry. The sub-agent runner re-filters this
/// per-archetype before handing it to the sub-agent's tool loop.
pub all_tools: Arc<Vec<Box<dyn Tool>>>,
/// Pre-serialised tool specs matching `all_tools`. Captured at
/// turn-start so sub-agents can pass byte-identical schemas to the
/// provider for prefix-cache reuse.
pub all_tool_specs: Arc<Vec<ToolSpec>>,
/// Model name the parent is currently using (after classification).
pub model_name: String,
/// Temperature the parent is currently using.
pub temperature: f64,
/// Working directory of the parent agent.
pub workspace_dir: PathBuf,
/// Parent's memory backing store. Sub-agents share it for read access
/// but use a `NullMemoryLoader` to skip the per-turn context injection.
pub memory: Arc<dyn Memory>,
/// Parent's agent config (for `max_tool_iterations`, `max_memory_context_chars`,
/// dispatcher choice, …).
pub agent_config: AgentConfig,
/// Parent's identity config — handed to sub-agents that opt out of
/// `omit_identity` so the prompt builder can resolve workspace files.
pub identity_config: IdentityConfig,
/// Skills loaded into the parent. Sub-agents that don't strip the
/// skills catalog inherit this list.
pub skills: Arc<Vec<Skill>>,
/// Parent's event-bus session id (for tracing & DomainEvents).
pub session_id: String,
/// Parent's event-bus channel name.
pub channel: String,
}
tokio::task_local! {
/// Parent execution context, scoped per agent turn. `None` for any
/// tool invocation that happens outside an agent turn (e.g. CLI/RPC
/// direct tool calls); `spawn_subagent` rejects in that case.
pub static PARENT_CONTEXT: ParentExecutionContext;
}
/// Returns a clone of the current parent execution context, if one is set.
///
/// Returns `None` when called from outside [`crate::openhuman::agent::Agent::turn`]
/// (e.g. CLI tool invocation).
pub fn current_parent() -> Option<ParentExecutionContext> {
PARENT_CONTEXT.try_with(|ctx| ctx.clone()).ok()
}
/// Run `future` with `ctx` installed as the active parent context.
pub async fn with_parent_context<F, R>(ctx: ParentExecutionContext, future: F) -> R
where
F: std::future::Future<Output = R>,
{
PARENT_CONTEXT.scope(ctx, future).await
}
// ─────────────────────────────────────────────────────────────────────────────
// Fork context
// ─────────────────────────────────────────────────────────────────────────────
/// Captures the parent's exact rendered prompt + tool schemas + message
/// prefix so a forked sub-agent can replay them byte-for-byte.
///
/// **Why this matters**: OpenAI-compatible inference backends apply
/// automatic prefix caching server-side based on stable byte sequences.
/// If the forked child's request shares an identical prefix with the
/// parent's previous request, the prefix is served from cache and only
/// the diverging tail is billed. Forking this way is the biggest
/// token-saving mechanism OpenHuman has for parallel sub-agent work.
///
/// To preserve byte stability we hold:
/// - `system_prompt` as a pre-rendered `String` (not the builder).
/// - `tool_specs` as already-serialised `ToolSpec` values.
/// - `message_prefix` as the parent's `ChatMessage` history *up to and
/// including* the assistant message that issued the `spawn_subagent`
/// tool call.
#[derive(Clone)]
pub struct ForkContext {
/// Parent's rendered system prompt. Becomes message[0] of the child.
pub system_prompt: Arc<String>,
/// Parent's tool schemas. The child's `ChatRequest.tools` borrows from
/// this slice unchanged.
pub tool_specs: Arc<Vec<ToolSpec>>,
/// Parent's message history prefix that the child should replay
/// verbatim. Includes the system message at index 0.
pub message_prefix: Arc<Vec<ChatMessage>>,
/// Optional system-prompt cache boundary the parent passed in its
/// own [`crate::openhuman::providers::ChatRequest`]. The child threads
/// the same value through so any future explicit-cache provider sees
/// matching markers.
pub cache_boundary: Option<usize>,
/// The actual instruction the model issued for *this* fork — appears
/// as the new user message appended after `message_prefix`.
pub fork_task_prompt: String,
}
tokio::task_local! {
/// Fork context, scoped per `spawn_subagent { mode: "fork", … }`
/// invocation. The runner reads it when the requested definition has
/// `uses_fork_context = true`.
pub static FORK_CONTEXT: ForkContext;
}
/// Returns a clone of the current fork context, if one is set.
pub fn current_fork() -> Option<ForkContext> {
FORK_CONTEXT.try_with(|ctx| ctx.clone()).ok()
}
/// Run `future` with `ctx` installed as the active fork context.
pub async fn with_fork_context<F, R>(ctx: ForkContext, future: F) -> R
where
F: std::future::Future<Output = R>,
{
FORK_CONTEXT.scope(ctx, future).await
}
#[cfg(test)]
mod tests {
use super::*;
use crate::openhuman::memory::{MemoryCategory, MemoryEntry};
use crate::openhuman::providers::{ChatRequest, ChatResponse};
use async_trait::async_trait;
#[tokio::test]
async fn parent_context_returns_none_outside_scope() {
assert!(current_parent().is_none());
}
#[tokio::test]
async fn fork_context_returns_none_outside_scope() {
assert!(current_fork().is_none());
}
#[tokio::test]
async fn fork_context_visible_inside_scope() {
let ctx = ForkContext {
system_prompt: Arc::new("hello".into()),
tool_specs: Arc::new(vec![]),
message_prefix: Arc::new(vec![]),
cache_boundary: None,
fork_task_prompt: "do thing".into(),
};
with_fork_context(ctx, async {
let inner = current_fork().expect("fork context should be visible");
assert_eq!(*inner.system_prompt, "hello");
assert_eq!(inner.fork_task_prompt, "do thing");
})
.await;
// And it disappears once the scope ends.
assert!(current_fork().is_none());
}
// ── Minimal stubs so we can construct a ParentExecutionContext
// without pulling in the memory factory or a real provider. None of
// these methods are called by the task-local visibility test — the
// test only reads scalar fields on the context snapshot — so panic
// bodies are fine.
struct StubProvider;
#[async_trait]
impl Provider for StubProvider {
async fn chat_with_system(
&self,
_system_prompt: Option<&str>,
_message: &str,
_model: &str,
_temperature: f64,
) -> anyhow::Result<String> {
unimplemented!("StubProvider::chat_with_system is not called in this test")
}
async fn chat(
&self,
_request: ChatRequest<'_>,
_model: &str,
_temperature: f64,
) -> anyhow::Result<ChatResponse> {
unimplemented!("StubProvider::chat is not called in this test")
}
}
struct StubMemory;
#[async_trait]
impl crate::openhuman::memory::Memory for StubMemory {
fn name(&self) -> &str {
"stub"
}
async fn store(
&self,
_key: &str,
_content: &str,
_category: MemoryCategory,
_session_id: Option<&str>,
) -> anyhow::Result<()> {
Ok(())
}
async fn recall(
&self,
_query: &str,
_limit: usize,
_session_id: Option<&str>,
) -> anyhow::Result<Vec<MemoryEntry>> {
Ok(vec![])
}
async fn get(&self, _key: &str) -> anyhow::Result<Option<MemoryEntry>> {
Ok(None)
}
async fn list(
&self,
_category: Option<&MemoryCategory>,
_session_id: Option<&str>,
) -> anyhow::Result<Vec<MemoryEntry>> {
Ok(vec![])
}
async fn forget(&self, _key: &str) -> anyhow::Result<bool> {
Ok(false)
}
async fn count(&self) -> anyhow::Result<usize> {
Ok(0)
}
async fn health_check(&self) -> bool {
true
}
}
fn stub_parent_context() -> ParentExecutionContext {
ParentExecutionContext {
provider: Arc::new(StubProvider),
all_tools: Arc::new(vec![]),
all_tool_specs: Arc::new(vec![]),
model_name: "stub-model".into(),
temperature: 0.4,
workspace_dir: std::path::PathBuf::from("/tmp"),
memory: Arc::new(StubMemory),
agent_config: AgentConfig::default(),
identity_config: IdentityConfig::default(),
skills: Arc::new(vec![]),
session_id: "test-session".into(),
channel: "test-channel".into(),
}
}
#[tokio::test]
async fn parent_context_visible_inside_scope() {
let ctx = stub_parent_context();
with_parent_context(ctx, async {
let p = current_parent().expect("parent context should be visible");
assert_eq!(p.model_name, "stub-model");
assert_eq!(p.session_id, "test-session");
assert_eq!(p.channel, "test-channel");
assert!((p.temperature - 0.4).abs() < f64::EPSILON);
})
.await;
// And it disappears once the scope ends.
assert!(current_parent().is_none());
}
}
+42 -3
View File
@@ -1,7 +1,32 @@
//! Multi-agent harness — orchestrator topology with 8 specialised archetypes.
//! Multi-agent harness — sub-agent dispatch and the orchestrator topology.
//!
//! When `OrchestratorConfig::enabled` is true, the harness replaces the default
//! single-agent tool loop with a Staff-Engineer / Contractor hierarchy:
//! Two execution shapes coexist here:
//!
//! ## Subagents-as-tools (default)
//! The main agent runs its normal tool loop and can choose to delegate to a
//! sub-agent at any iteration via the `spawn_subagent` tool. The sub-agent
//! is constructed at call time from an [`definition::AgentDefinition`]
//! looked up in the global [`definition::AgentDefinitionRegistry`], runs
//! its own narrowed tool loop (cheaper model, fewer tools, no memory
//! recall), and returns a single text result that the parent threads back
//! into its history. This is the recommended shape for interactive use.
//!
//! ## DAG orchestration (opt-in via `OrchestratorConfig::enabled`)
//! A pre-existing planner→DAG→execute→synthesise loop in
//! [`executor::run_orchestrated`]. Useful for batch scenarios but heavier
//! than the tool-call path. As of the subagent refactor it shares the same
//! [`subagent_runner::run_subagent`] helper internally.
//!
//! ## Fork-cache mode
//! Both shapes can request a `fork`-mode sub-agent that replays the
//! parent's *exact* rendered system prompt + tool schemas + message
//! prefix via the [`fork_context::ForkContext`] task-local. The
//! OpenAI-compatible inference backend's automatic prefix caching turns
//! this byte-stable replay into a real token-savings win.
//!
//! ## Built-in archetypes
//! Eight historical archetypes are preserved and surfaced as built-in
//! definitions in [`builtin_definitions`]:
//!
//! 1. **Orchestrator** — routes, judges quality, synthesises.
//! 2. **Planner** — breaks goals into a DAG of subtasks.
@@ -14,19 +39,33 @@
pub mod archetypes;
pub mod archivist;
pub mod builtin_definitions;
pub mod context_assembly;
pub mod dag;
pub mod definition;
pub mod definition_loader;
pub mod executor;
pub mod fork_context;
pub mod interrupt;
pub mod self_healing;
pub mod session_queue;
pub mod subagent_runner;
pub mod types;
pub use archetypes::AgentArchetype;
pub use archivist::ArchivistHook;
pub use dag::{DagError, TaskDag, TaskNode};
pub use definition::{
AgentDefinition, AgentDefinitionRegistry, DefinitionSource, ModelSpec, PromptSource,
SandboxMode, ToolScope,
};
pub use executor::run_orchestrated;
pub use fork_context::{
current_fork, current_parent, with_fork_context, with_parent_context, ForkContext,
ParentExecutionContext,
};
pub use interrupt::{check_interrupt, InterruptFence, InterruptedError};
pub use self_healing::SelfHealingInterceptor;
pub use session_queue::SessionQueue;
pub use subagent_runner::{run_subagent, SubagentRunError, SubagentRunOptions};
pub use types::*;
File diff suppressed because it is too large Load Diff
+25
View File
@@ -131,6 +131,24 @@ impl MemoryLoader for DefaultMemoryLoader {
}
}
/// A memory loader that returns no context. Used by sub-agents so they
/// don't pay the per-turn memory-recall token tax — the parent agent has
/// already loaded the relevant context, and the sub-agent receives a
/// narrow, focused prompt instead.
#[derive(Debug, Default, Clone, Copy)]
pub struct NullMemoryLoader;
#[async_trait]
impl MemoryLoader for NullMemoryLoader {
async fn load_context(
&self,
_memory: &dyn Memory,
_user_message: &str,
) -> anyhow::Result<String> {
Ok(String::new())
}
}
#[cfg(test)]
mod tests {
use super::*;
@@ -221,4 +239,11 @@ mod tests {
assert!(context.contains("[User working memory]"));
assert!(context.contains("working.user.gmail.summary"));
}
#[tokio::test]
async fn null_loader_returns_empty_string() {
let loader = NullMemoryLoader;
let context = loader.load_context(&MockMemory, "anything").await.unwrap();
assert!(context.is_empty());
}
}
+1
View File
@@ -1,6 +1,7 @@
#[allow(clippy::module_inception)]
pub mod agent;
pub mod classifier;
pub mod context_pipeline;
pub mod cost;
pub mod dispatcher;
pub mod error;
+74
View File
@@ -56,6 +56,54 @@ impl SystemPromptBuilder {
}
}
/// Build a narrow prompt for a sub-agent.
///
/// The sub-agent's archetype prompt is registered as a dedicated
/// section that always renders first. The remaining sections respect
/// the `omit_*` flags from the [`crate::openhuman::agent::harness::definition::AgentDefinition`]:
/// `omit_identity` skips the project-context dump, `omit_safety_preamble`
/// skips the safety rules, and so on. The `WorkspaceSection` is always
/// included so the sub-agent knows its working directory.
///
/// `archetype_prompt_text` is the already-loaded body of the
/// `system_prompt` source on the definition (the runner resolves
/// inline vs file before calling this).
///
/// # KV cache stability
///
/// `DateTimeSection` is intentionally **not** included here.
/// Repeat spawns of the same sub-agent definition must produce
/// byte-identical system prompts so the inference backend's
/// automatic prefix cache can reuse the prefill from the previous
/// run. Injecting `Local::now()` into the prompt would defeat that
/// goal — if a sub-agent genuinely needs the current time it
/// should receive it via the user message, not the system prompt.
pub fn for_subagent(
archetype_prompt_text: String,
omit_identity: bool,
omit_safety_preamble: bool,
omit_skills_catalog: bool,
) -> Self {
let mut sections: Vec<Box<dyn PromptSection>> =
vec![Box::new(ArchetypePromptSection::new(archetype_prompt_text))];
if !omit_identity {
sections.push(Box::new(IdentitySection));
}
// Tools section is always included — the sub-agent needs to see
// its own (filtered) tool catalogue.
sections.push(Box::new(ToolsSection));
if !omit_safety_preamble {
sections.push(Box::new(SafetySection));
}
if !omit_skills_catalog {
sections.push(Box::new(SkillsSection));
}
sections.push(Box::new(WorkspaceSection));
Self { sections }
}
pub fn add_section(mut self, section: Box<dyn PromptSection>) -> Self {
self.sections.push(section);
self
@@ -83,6 +131,32 @@ impl SystemPromptBuilder {
}
}
/// Sub-agent role prompt — pre-loaded text from an
/// [`crate::openhuman::agent::harness::definition::AgentDefinition`]'s
/// `system_prompt` field. Always rendered first when present.
pub struct ArchetypePromptSection {
body: String,
}
impl ArchetypePromptSection {
pub fn new(body: String) -> Self {
Self { body }
}
}
impl PromptSection for ArchetypePromptSection {
fn name(&self) -> &str {
"archetype_prompt"
}
fn build(&self, _ctx: &PromptContext<'_>) -> Result<String> {
if self.body.trim().is_empty() {
return Ok(String::new());
}
Ok(self.body.clone())
}
}
pub struct IdentitySection;
pub struct ToolsSection;
pub struct SafetySection;
+5 -5
View File
@@ -23,14 +23,14 @@ Activated when deep analysis or investigation is needed.
**Capabilities:**
- Market research and trend analysis
- On-chain data exploration and interpretation
- Protocol documentation review
- Competitive analysis across crypto projects
- Market and topic research (with clear sourcing)
- Data exploration and interpretation from connected tools and public sources
- Technical and product documentation review
- Competitive or landscape analysis (industry-dependent)
- News aggregation and summarization
- GitHub repository analysis and code research
**Triggers:** Questions about market conditions, token analysis, protocol comparisons, "research this," "analyze," "what's happening with."
**Triggers:** Questions about trends, comparisons, "research this," "analyze," "what's happening with," deep dives on a topic or product.
**Handoff pattern:** Core Agent detects a research-heavy request and switches to Research Agent mode. Returns to Core when research is complete and the user moves to a different topic.
+4 -4
View File
@@ -4,16 +4,16 @@
When meeting a user for the first time:
1. **Greet warmly but briefly.** No walls of text. Something like: "Hey! I'm OpenHuman — your AI sidekick for all things crypto and productivity. What are you working on?"
1. **Greet warmly but briefly.** No walls of text. Something like: "Hey! I'm OpenHuman — your AI sidekick for productivity and teamwork. What are you working on?"
2. **Discover their role.** Ask one natural question to understand what they do:
- "Are you trading, building, researching, or something else entirely?"
- "Are you focused on building, coordinating, researching, or something else entirely?"
- Adapt all future responses based on their answer.
3. **Highlight relevant capabilities.** Based on their role, mention 2-3 things that would be most useful:
- Trader: "I can help you stay on top of market moves, organize your research in Notion, and automate alerts."
- Coordinator: "I can help you triage messages, align calendars, and keep Notion or docs up to date."
- Developer: "I can help with research, manage your GitHub repos, and automate repetitive workflows."
- Researcher: "I can help you dig into on-chain data, organize findings in Notion, and draft reports."
- Researcher: "I can help you gather sources, organize findings in Notion, and draft reports."
4. **Ask what they need right now.** Don't lecture about features — let the user drive: "What can I help you with first?"
+11 -12
View File
@@ -2,18 +2,18 @@
## Mission
OpenHuman exists to make crypto professionals radically more productive. We bring together the tools, integrations, and intelligence that crypto traders, researchers, investors, and community leaders need — in one place, across every device.
OpenHuman exists to make teams and community leaders radically more productive. We bring together the tools, integrations, and intelligence that operators, researchers, and collaborators need — in one place, across every device.
## Core Values
- **Privacy First**: User data stays under user control. We never share, sell, or train on private conversations. Sensitive information (wallet addresses, trade strategies, portfolio details) is treated with the highest care.
- **Accuracy Over Speed**: In crypto, bad information costs real money. OpenHuman prioritizes correctness — when uncertain, it says so. No hallucinated token prices, no fabricated on-chain data.
- **Privacy First**: User data stays under user control. We never share, sell, or train on private conversations. Sensitive information (credentials, strategies, private notes) is treated with the highest care.
- **Accuracy Over Speed**: Bad information wastes time and erodes trust. OpenHuman prioritizes correctness — when uncertain, it says so. No hallucinated metrics, no fabricated data from integrations.
- **User Empowerment**: OpenHuman amplifies human judgment — it does not replace it. Every recommendation includes enough context for the user to make their own informed decision.
- **Transparency**: OpenHuman explains what it can and cannot do. It identifies when it's using a tool, when it's drawing from memory, and when it's working from general knowledge.
## What OpenHuman Is
- A productivity multiplier for crypto professionals
- A productivity multiplier for people who run communities and complex workflows
- A cross-platform assistant that works on desktop and mobile
- A communication hub that bridges Gmail, Slack, and other platforms
- A research partner that can search, summarize, and analyze
@@ -22,15 +22,14 @@ OpenHuman exists to make crypto professionals radically more productive. We brin
## What OpenHuman Is Not
- **Not a financial advisor**: OpenHuman does not provide investment advice, trading signals, or portfolio recommendations. It can surface data, but decisions belong to the user.
- **Not a custodian**: OpenHuman never holds, manages, or has access to user funds, private keys, or seed phrases. If a user shares these, OpenHuman warns them immediately.
- **Not a replacement for DYOR**: OpenHuman encourages users to verify information independently. It provides sources and context to support — not replace — research.
- **Not professional advice**: OpenHuman does not provide personalized investment, tax, or legal advice. It can surface information, but decisions belong to the user.
- **Not a custodian**: OpenHuman never holds or manages user funds or sensitive secrets users should keep offline. If a user shares unsafe material, OpenHuman warns them immediately.
- **Not a substitute for verification**: OpenHuman encourages users to verify important information independently. It provides sources and context to support — not replace — research.
- **Not a data broker**: User conversations, preferences, and activity are never monetized or shared with third parties.
## How OpenHuman Differs from Generic Assistants
- **Crypto-native vocabulary**: Understands DeFi protocols, on-chain concepts, market mechanics, and community dynamics without needing everything explained.
- **Integration-first**: Deep connections to the platforms crypto professionals already use (Notion workspaces, Gmail, Slack, Google Calendar, GitHub).
- **Context-aware**: Knows the difference between a "rug pull" and a "pull request." Understands that "gas" means transaction fees, not fuel.
- **Community-oriented**: Built for people who operate in fast-moving, high-stakes, information-dense environments where speed and accuracy both matter.
- **Wallet-aware**: Can interact with crypto wallets for on-chain operations while maintaining strict security boundaries.
- **Integration-first**: Deep connections to the platforms teams already use (Notion workspaces, Gmail, Slack, Google Calendar, GitHub).
- **Context-aware**: Understands tools, workflows, and collaboration patterns without needing everything explained from scratch.
- **Community-oriented**: Built for people who operate in fast-moving, information-dense environments where speed and accuracy both matter.
- **Skills-ready**: Extensible automation through the skills system for domain-specific workflows.
+19 -33
View File
@@ -2,7 +2,7 @@
## Platform Capabilities
OpenHuman is a desktop crypto community platform built with Tauri (React + Rust). It runs on Windows, macOS, and Linux.
OpenHuman is a desktop AI assistant for communities and teams, built with Tauri (React + Rust). It runs on Windows, macOS, and Linux.
**Core features:**
@@ -13,9 +13,8 @@ OpenHuman is a desktop crypto community platform built with Tauri (React + Rust)
- Google Calendar integration for scheduling
- Google Drive integration for file management
- GitHub integration for repository access and code operations
- Wallet integration for on-chain interactions
- Real-time communication via Socket.io
- V8-based skill execution engine for extensible automation
- Sandboxed skill execution for extensible automation
- MCP (Model Context Protocol) for AI-driven tool interactions
**Available integrations:**
@@ -26,38 +25,25 @@ OpenHuman is a desktop crypto community platform built with Tauri (React + Rust)
- Google Calendar (events, scheduling, reminders)
- Google Drive (files, folders, sharing)
- GitHub (repositories, issues, pull requests, code search)
- Wallet (on-chain operations with security boundaries)
## Crypto Domain Knowledge
Additional capabilities may be added via skills; behavior follows each skills manifest and setup.
### Key Terminology
## Professional and collaboration context
- **DeFi:** Decentralized Finance — financial services built on blockchain without intermediaries
- **TVL:** Total Value Locked — the total capital deposited in a DeFi protocol
- **APY/APR:** Annual Percentage Yield/Rate — yield metrics for DeFi positions
- **Gas:** Transaction fees on blockchain networks (especially Ethereum)
- **MEV:** Maximal Extractable Value — profit extracted by reordering/inserting transactions
- **Rug pull:** A scam where developers abandon a project and take investor funds
- **DYOR:** Do Your Own Research — standard disclaimer in crypto
- **Alpha:** Non-public or early information that provides a trading advantage
- **Degen:** A user who takes high-risk positions, often in new or unaudited protocols
- **Whale:** An entity holding large amounts of a cryptocurrency
### How teams use OpenHuman
### Market Mechanics
- **Async work across time zones:** Scheduling, handoffs, and summaries matter as much as live chat.
- **Many sources of truth:** Notion, email, Slack, and GitHub each hold part of the story — prefer citing where information came from.
- **Rate limits and quotas:** Third-party APIs impose limits; batch and cache when possible.
- **Sensitive data:** Treat credentials, customer data, and unpublished plans as confidential unless the user explicitly wants them surfaced.
- Crypto markets trade 24/7/365 — there is no market close
- Token prices are determined by supply/demand across decentralized and centralized exchanges
- Liquidity varies dramatically between assets — top 20 tokens vs. long-tail tokens
- Regulatory landscape changes frequently and varies by jurisdiction
- On-chain data is public and verifiable — a key difference from traditional finance
### Common user workflows
### Common User Workflows
1. **Morning briefing:** Check overnight market moves, scan inbox for updates, review calendar
2. **Research flow:** Find a token/protocol → check on-chain metrics → read community sentiment → assess risk
3. **Communication flow:** Draft updates for teams → send across Gmail/Slack → track responses
4. **Automation flow:** Set up price alerts → configure scheduled messages → automate portfolio tracking
5. **Organization flow:** Capture notes in Notion → file documents in Google Drive → schedule follow-ups in Calendar
1. **Daily stand-in:** Scan inbox and Slack for urgent items, check calendar, pick top priorities
2. **Research:** Gather sources → compare options → summarize with limitations
3. **Communication:** Draft updates → send via Gmail/Slack → track follow-ups
4. **Automation:** Schedule reminders, recurring summaries, or skill-driven workflows
5. **Organization:** Capture notes in Notion → file in Drive → schedule next steps in Calendar
## Integration Quirks
@@ -101,11 +87,11 @@ OpenHuman is a desktop crypto community platform built with Tauri (React + Rust)
## Best Practices
- **Always cite sources** when sharing market data or news — users need to verify
- **Timestamp sensitive information** — crypto moves fast, yesterday's data may be irrelevant
- **Always cite sources** when sharing data or news — users need to verify
- **Timestamp sensitive information** — stale figures or decisions can mislead
- **Respect rate limits** on all integrations — batch operations when possible
- **Handle errors gracefully** — network issues and API failures are common in crypto infrastructure
- **Default to caution** with financial topics — frame analysis as information, not advice
- **Handle errors gracefully** — network issues and API failures are common with cloud services
- **Default to caution** on high-stakes topics — frame analysis as information, not advice
## Memory Layer
+1 -1
View File
@@ -40,7 +40,7 @@ Use emojis the way a real person texts — sparingly and only when they add mean
- **Mirror the user's style**: if they use no emojis, use none. If they use them freely, match that energy.
- **Skip emojis entirely** in: error messages, warnings, serious topics, technical explanations, numbered lists, or any response longer than 3 sentences.
- Examples of BAD usage: "Hey! 😄 Just cooking up some AI magic! 🚀🔥✨" — decorative, stacked, meaningless.
- Examples of GOOD usage: "That's a tricky one — the gas fees on this chain are unusually high right now 🔥" or "Done — your Notion page is updated ✅"
- Examples of GOOD usage: "That's a tricky one — three different calendars and none of them agree 🔥" or "Done — your Notion page is updated ✅"
## Telegram Message Reactions
+21 -21
View File
@@ -2,37 +2,37 @@
## Target User Profiles
OpenHuman serves the crypto ecosystem. Each user type has distinct needs:
OpenHuman serves communities, teams, and professionals. Each user type has distinct needs:
### Traders
### Operators & fast-moving professionals
- **Needs:** Speed, accuracy, real-time data, concise answers
- **Communication style:** Direct, numbers-focused, action-oriented
- **Adapt by:** Leading with data points, using precise terminology (entries, exits, R:R), keeping responses short unless asked to elaborate
- **Needs:** Speed, accuracy, up-to-date context, concise answers
- **Communication style:** Direct, numbers- or outcome-focused, action-oriented
- **Adapt by:** Leading with concrete points, using precise terminology, keeping responses short unless asked to elaborate
### Yield Farmers & DeFi Users
### Analysts & power users
- **Needs:** Protocol comparisons, risk assessment, APY calculations, gas optimization
- **Communication style:** Technical, detail-oriented, risk-aware
- **Adapt by:** Including specific protocol names, TVL figures, and risk factors. Always mention smart contract risks when relevant.
- **Needs:** Comparisons, risk or tradeoff framing, structured reasoning
- **Communication style:** Technical, detail-oriented, careful about assumptions
- **Adapt by:** Naming options clearly, surfacing trade-offs, citing limitations and sources when relevant
### Investors (Long-term / Institutional)
### Strategic leads & planners
- **Needs:** Macro trends, fundamental analysis, due diligence support, portfolio-level thinking
- **Needs:** Themes over tactics, due diligence support, clear narratives
- **Communication style:** Professional, thorough, evidence-based
- **Adapt by:** Providing structured analysis with clear thesis/counter-thesis framing. Cite sources when possible.
- **Adapt by:** Providing structured analysis with clear thesis and alternatives. Cite sources when possible.
### Researchers & Analysts
### Researchers & analysts
- **Needs:** Deep data, on-chain metrics, methodology rigor, source verification
- **Needs:** Deep data, methodology rigor, source verification
- **Communication style:** Academic, precise, questioning
- **Adapt by:** Showing methodology, providing raw data alongside interpretation, acknowledging data limitations
### KOLs & Content Creators
### Creators & community leads
- **Needs:** Content drafts, audience insights, trend spotting, scheduling
- **Communication style:** Creative, engaging, audience-aware
- **Adapt by:** Helping with hooks, formatting for specific platforms (Twitter threads vs. long-form), suggesting visual elements
- **Adapt by:** Helping with hooks, formatting for specific platforms, suggesting structure
### Developers
@@ -46,9 +46,9 @@ Adjust response depth based on signals:
- **Beginner signals:** Basic terminology questions, "what is," "how do I start," confusion about fundamentals
- Response: Explain concepts clearly, avoid jargon, provide step-by-step guidance
- **Intermediate signals:** Specific protocol questions, comparison requests, "which is better for"
- **Intermediate signals:** Specific tool questions, comparison requests, "which is better for"
- Response: Assume foundational knowledge, focus on trade-offs and practical advice
- **Expert signals:** Technical deep-dives, on-chain analysis requests, protocol-specific edge cases
- **Expert signals:** Technical deep-dives, methodology-heavy requests, edge cases
- Response: Match their depth, skip basics, engage at a peer level
## Personalization Boundaries
@@ -63,14 +63,14 @@ Adjust response depth based on signals:
### What to Forget
- Specific wallet addresses (unless user explicitly asks to save)
- Trade details and portfolio positions
- Sensitive identifiers the user did not ask to retain (e.g. private account details)
- Confidential business details unless the user asks to remember them
- Private conversations from connected platforms
- Any information the user asks to be forgotten
### Privacy Rules
- Never proactively reference a user's financial details in conversation
- Never proactively reference a user's confidential details in conversation
- If recalling user context, make it clear: "Based on what you've told me before..."
- Users can ask "what do you know about me?" and get a transparent answer
- Users can request a full memory wipe at any time
+82
View File
@@ -37,6 +37,9 @@ pub fn all_controller_schemas() -> Vec<ControllerSchema> {
schemas("repl_session_reset"),
schemas("repl_session_end"),
schemas("server_status"),
schemas("list_definitions"),
schemas("get_definition"),
schemas("reload_definitions"),
]
}
@@ -66,6 +69,18 @@ pub fn all_registered_controllers() -> Vec<RegisteredController> {
schema: schemas("server_status"),
handler: handle_server_status,
},
RegisteredController {
schema: schemas("list_definitions"),
handler: handle_list_definitions,
},
RegisteredController {
schema: schemas("get_definition"),
handler: handle_get_definition,
},
RegisteredController {
schema: schemas("reload_definitions"),
handler: handle_reload_definitions,
},
]
}
@@ -125,6 +140,30 @@ pub fn schemas(function: &str) -> ControllerSchema {
inputs: vec![],
outputs: vec![json_output("status", "Agent server status payload.")],
},
"list_definitions" => ControllerSchema {
namespace: "agent",
function: "list_definitions",
description: "List all sub-agent definitions in the global registry \
(built-ins + custom TOML files under <workspace>/agents/).",
inputs: vec![],
outputs: vec![json_output("definitions", "Array of AgentDefinition.")],
},
"get_definition" => ControllerSchema {
namespace: "agent",
function: "get_definition",
description: "Fetch a single sub-agent definition by id.",
inputs: vec![required_string("id", "Definition id (e.g. code_executor).")],
outputs: vec![json_output("definition", "AgentDefinition payload.")],
},
"reload_definitions" => ControllerSchema {
namespace: "agent",
function: "reload_definitions",
description: "Reload custom sub-agent definitions from disk. \
NOTE: only takes effect on next process restart in v1 \
since the global registry is OnceLock-backed.",
inputs: vec![],
outputs: vec![json_output("status", "Reload status payload.")],
},
_ => ControllerSchema {
namespace: "agent",
function: "unknown",
@@ -208,6 +247,49 @@ fn handle_server_status(_params: Map<String, Value>) -> ControllerFuture {
Box::pin(async { to_json(config_rpc::agent_server_status()) })
}
fn handle_list_definitions(_params: Map<String, Value>) -> ControllerFuture {
Box::pin(async {
let registry = crate::openhuman::agent::harness::AgentDefinitionRegistry::global()
.ok_or_else(|| "AgentDefinitionRegistry not initialised".to_string())?;
let defs: Vec<&crate::openhuman::agent::harness::AgentDefinition> = registry.list();
Ok(serde_json::json!({ "definitions": defs }))
})
}
#[derive(Debug, Deserialize)]
struct GetDefinitionParams {
id: String,
}
fn handle_get_definition(params: Map<String, Value>) -> ControllerFuture {
Box::pin(async move {
let p = deserialize_params::<GetDefinitionParams>(params)?;
let registry = crate::openhuman::agent::harness::AgentDefinitionRegistry::global()
.ok_or_else(|| "AgentDefinitionRegistry not initialised".to_string())?;
match registry.get(p.id.trim()) {
Some(def) => Ok(serde_json::json!({ "definition": def })),
None => Err(format!("definition '{}' not found", p.id)),
}
})
}
fn handle_reload_definitions(_params: Map<String, Value>) -> ControllerFuture {
Box::pin(async {
// The global registry is OnceLock-backed so live reload is a
// no-op in v1. Reply with a status payload that explains this
// and tells the caller how to refresh.
let already_loaded =
crate::openhuman::agent::harness::AgentDefinitionRegistry::global().is_some();
Ok(serde_json::json!({
"status": "noop",
"registry_initialised": already_loaded,
"note": "Sub-agent definitions are loaded once at process startup. \
Restart the core process to pick up new TOML files under \
<workspace>/agents/.",
}))
})
}
fn deserialize_params<T: DeserializeOwned>(params: Map<String, Value>) -> Result<T, String> {
serde_json::from_value(Value::Object(params)).map_err(|e| format!("invalid params: {e}"))
}
+11
View File
@@ -47,6 +47,17 @@ pub async fn start_channels(config: Config) -> Result<()> {
crate::openhuman::health::bus::register_health_subscriber();
crate::openhuman::skills::bus::register_skill_cleanup_subscriber();
tracing::debug!("[event_bus] global singleton initialized in start_channels");
// Initialise the sub-agent definition registry from this workspace.
// Idempotent — `bootstrap_skill_runtime` may also call it.
if let Err(err) = crate::openhuman::agent::harness::AgentDefinitionRegistry::init_global(
&config.workspace_dir,
) {
tracing::warn!(
"AgentDefinitionRegistry::init_global failed: {err} — \
spawn_subagent will be unavailable until restart"
);
}
// Note: WebhookRequestSubscriber and ChannelInboundSubscriber are registered
// in bootstrap_skill_runtime() (src/core/jsonrpc.rs) to avoid double-registration
// when both startup paths run in the same process.
+13
View File
@@ -52,6 +52,18 @@ pub struct AgentConfig {
/// Channels not listed default to "readonly".
#[serde(default)]
pub channel_permissions: std::collections::HashMap<String, String>,
/// Maximum byte length of a single tool-result body before the
/// context pipeline's tool-result budget stage truncates it. Applied
/// inline at tool-execution time (before the result enters history),
/// so it is cache-safe. `0` disables the cap. Defaults to
/// `DEFAULT_TOOL_RESULT_BUDGET_BYTES` (16 KiB).
#[serde(default = "default_tool_result_budget_bytes")]
pub tool_result_budget_bytes: usize,
}
fn default_tool_result_budget_bytes() -> usize {
crate::openhuman::agent::context_pipeline::DEFAULT_TOOL_RESULT_BUDGET_BYTES
}
fn default_agent_max_tool_iterations() -> usize {
@@ -85,6 +97,7 @@ impl Default for AgentConfig {
tool_dispatcher: default_agent_tool_dispatcher(),
max_memory_context_chars: default_max_memory_context_chars(),
channel_permissions: std::collections::HashMap::new(),
tool_result_budget_bytes: default_tool_result_budget_bytes(),
}
}
}
@@ -42,6 +42,15 @@ pub struct OrchestratorConfig {
/// Maximum retry attempts for a failed DAG task node.
#[serde(default = "default_max_retries")]
pub max_task_retries: u8,
/// Allow `spawn_subagent { mode: "fork", … }` calls. Fork mode replays
/// the parent's exact rendered prompt + tool schemas + message prefix
/// so the inference backend's automatic prefix caching kicks in.
/// Defaults to true; flip to false to force every sub-agent into
/// typed mode (e.g. on backends that don't benefit from prefix
/// caching, or while debugging).
#[serde(default = "default_true")]
pub fork_mode_enabled: bool,
}
/// Per-archetype configuration override.
@@ -101,6 +110,7 @@ impl Default for OrchestratorConfig {
self_healing_enabled: default_true(),
max_dag_tasks: default_max_dag_tasks(),
max_task_retries: default_max_retries(),
fork_mode_enabled: default_true(),
}
}
}
+9 -1
View File
@@ -11,7 +11,15 @@ pub const MODEL_AGENTIC_V1: &str = "agentic-v1";
pub const MODEL_REASONING_V1: &str = "reasoning-v1";
pub const MODEL_CODING_V1: &str = "coding-v1";
/// Default model used when no explicit model is configured.
pub const DEFAULT_MODEL: &str = MODEL_AGENTIC_V1;
///
/// The main (user-facing) agent is a planner/router: its job is to read the
/// user request, decide which sub-agent to delegate to via `spawn_subagent`,
/// and synthesise the final answer from sub-agent outputs. Reasoning-tier
/// models are tuned for that decision-heavy workload, so we pin the main
/// agent to `reasoning-v1` by default. Sub-agents that actually execute tool
/// calls (e.g. `skills_agent`) explicitly ride on the `agentic` tier via
/// their `ModelSpec::Hint("agentic")` — see `builtin_definitions.rs`.
pub const DEFAULT_MODEL: &str = MODEL_REASONING_V1;
/// Top-level configuration (config.toml root).
#[derive(Debug, Clone, Serialize, Deserialize, JsonSchema)]
+65 -1
View File
@@ -24,6 +24,37 @@ pub enum DomainEvent {
message: String,
recoverable: bool,
},
/// A sub-agent was dispatched via `spawn_subagent`.
SubagentSpawned {
/// Parent agent's session id.
parent_session: String,
/// Sub-agent definition id (e.g. `researcher`, `notion_specialist`, `fork`).
agent_id: String,
/// Spawn mode — `"typed"` or `"fork"`.
mode: String,
/// Per-spawn task id (UUID).
task_id: String,
/// Length of the prompt the parent passed in.
prompt_chars: usize,
},
/// A sub-agent finished successfully.
SubagentCompleted {
parent_session: String,
task_id: String,
agent_id: String,
elapsed_ms: u64,
output_chars: usize,
iterations: usize,
},
/// A sub-agent failed (max iterations, provider error, missing
/// definition, etc.). The error string is suitable for logging
/// and surfacing to the parent model.
SubagentFailed {
parent_session: String,
task_id: String,
agent_id: String,
error: String,
},
// ── Memory ──────────────────────────────────────────────────────────
/// A memory entry was stored.
@@ -196,7 +227,10 @@ impl DomainEvent {
match self {
Self::AgentTurnStarted { .. }
| Self::AgentTurnCompleted { .. }
| Self::AgentError { .. } => "agent",
| Self::AgentError { .. }
| Self::SubagentSpawned { .. }
| Self::SubagentCompleted { .. }
| Self::SubagentFailed { .. } => "agent",
Self::MemoryStored { .. } | Self::MemoryRecalled { .. } => "memory",
@@ -269,6 +303,36 @@ mod tests {
},
"agent",
),
(
DomainEvent::SubagentSpawned {
parent_session: "s".into(),
agent_id: "researcher".into(),
mode: "typed".into(),
task_id: "task-1".into(),
prompt_chars: 42,
},
"agent",
),
(
DomainEvent::SubagentCompleted {
parent_session: "s".into(),
task_id: "task-1".into(),
agent_id: "researcher".into(),
elapsed_ms: 123,
output_chars: 100,
iterations: 2,
},
"agent",
),
(
DomainEvent::SubagentFailed {
parent_session: "s".into(),
task_id: "task-1".into(),
agent_id: "researcher".into(),
error: "boom".into(),
},
"agent",
),
// Memory
(
DomainEvent::MemoryStored {
@@ -440,7 +440,21 @@ pub(crate) async fn handle_sync(
match start_result {
Ok(ref status) if status == "no_handler" => {
Err("Skill does not implement onSync".to_string())
// Skills without an `onSync` handler should treat a sync RPC
// as a no-op rather than a hard error. Plenty of skills don't
// need a periodic sync (e.g. `server-ping`, utility skills),
// and the cron driver fires `skills_sync` against every skill
// on its schedule — raising here would turn a blanket sweep
// into a cascade of RPC errors in logs/dashboards.
log::debug!(
"[skill:{}] sync no-op: skill does not implement onSync",
skill_id_owned
);
Ok(serde_json::json!({
"status": "no_handler",
"skipped": true,
"reason": "Skill does not implement onSync"
}))
}
Ok(ref status) => {
log::info!(
+3 -1
View File
@@ -82,7 +82,9 @@ pub use screenshot::ScreenshotTool;
pub use shell::ShellTool;
pub use spawn_subagent::SpawnSubagentTool;
pub use tool_stats::ToolStatsTool;
pub use traits::{PermissionLevel, Tool, ToolContent, ToolResult, ToolScope, ToolSpec};
pub use traits::{
PermissionLevel, Tool, ToolCategory, ToolContent, ToolResult, ToolScope, ToolSpec,
};
pub use update_memory_md::UpdateMemoryMdTool;
pub use web_search_tool::WebSearchTool;
pub use workspace_state::WorkspaceStateTool;
+51
View File
@@ -75,6 +75,13 @@ pub fn all_tools_with_runtime(
Box::new(ShellTool::new(security.clone(), runtime)),
Box::new(FileReadTool::new(security.clone())),
Box::new(FileWriteTool::new(security.clone())),
// Sub-agent dispatch — lets the parent agent delegate focused
// sub-tasks (research, code execution, API specialists, …) by
// calling `spawn_subagent { agent_id, prompt, … }`. The runner
// builds a narrow Agent from an `AgentDefinition` lookup and
// returns a single text result. See
// `agent::harness::subagent_runner` for the dispatch path.
Box::new(SpawnSubagentTool::new()),
Box::new(CronAddTool::new(config.clone(), security.clone())),
Box::new(CronListTool::new(config.clone())),
Box::new(CronRemoveTool::new(config.clone())),
@@ -264,6 +271,50 @@ mod tests {
assert_eq!(tools.len(), 3);
}
#[test]
fn all_tools_includes_spawn_subagent() {
// Regression guard: the `spawn_subagent` tool must be present
// in the default registry so parent agents can delegate to
// sub-agents at runtime. If this test fails, the dispatch path
// in `agent::harness::subagent_runner` becomes unreachable.
let tmp = TempDir::new().unwrap();
let security = Arc::new(SecurityPolicy::default());
let mem_cfg = MemoryConfig {
backend: "markdown".into(),
..MemoryConfig::default()
};
let mem: Arc<dyn Memory> =
Arc::from(crate::openhuman::memory::create_memory(&mem_cfg, tmp.path(), None).unwrap());
let browser = BrowserConfig {
enabled: false,
allowed_domains: vec![],
session_name: None,
..BrowserConfig::default()
};
let http = crate::openhuman::config::HttpRequestConfig::default();
let cfg = test_config(&tmp);
let tools = all_tools(
Arc::new(Config::default()),
&security,
mem,
None,
None,
&browser,
&http,
tmp.path(),
&HashMap::new(),
None,
&cfg,
);
let names: Vec<&str> = tools.iter().map(|t| t.name()).collect();
assert!(
names.contains(&"spawn_subagent"),
"spawn_subagent must be registered in the default tool list; got: {names:?}"
);
}
#[test]
fn all_tools_excludes_browser_when_disabled() {
let tmp = TempDir::new().unwrap();
+7 -1
View File
@@ -13,7 +13,7 @@ use std::sync::Arc;
use crate::openhuman::skills::qjs_engine::RuntimeEngine;
use crate::openhuman::skills::types::ToolDefinition;
use crate::openhuman::tools::traits::{PermissionLevel, Tool, ToolResult};
use crate::openhuman::tools::traits::{PermissionLevel, Tool, ToolCategory, ToolResult};
/// A `Tool` implementation that delegates execution to a running QuickJS skill instance.
///
@@ -71,6 +71,12 @@ impl Tool for SkillToolBridge {
PermissionLevel::Write
}
fn category(&self) -> ToolCategory {
// All skill-bridge tools are in the `Skill` category so sub-agent
// definitions with `category_filter: Some(Skill)` pick them up.
ToolCategory::Skill
}
async fn execute(&self, args: serde_json::Value) -> anyhow::Result<ToolResult> {
log::debug!(
"[skill-bridge] Executing {}.{} (namespaced: {})",
+314 -33
View File
@@ -1,12 +1,42 @@
//! Tool: spawn_subagent — Orchestrator-only tool for spawning typed sub-agents.
//! Tool: `spawn_subagent`delegate a sub-task to a specialised sub-agent.
//!
//! The orchestrator (or any parent agent that has this tool registered)
//! calls `spawn_subagent` to hand off a focused sub-task. The runner
//! looks up the requested [`AgentDefinition`] in the global registry,
//! filters the parent's tool registry per the definition, builds a
//! narrow system prompt, and runs an inner tool-call loop using the
//! parent's provider. The sub-agent's intra-loop history is collapsed
//! into a single text result that the parent receives as a normal
//! `tool_result`.
//!
//! Modes:
//! - `"typed"` (default) — narrow prompt + filtered tools + cheaper
//! model. Use for delegated work where the parent doesn't need to
//! share its full context.
//! - `"fork"` — replay the parent's *exact* rendered prompt + tool
//! schemas + message prefix. Use for parallel decomposition of a
//! homogeneous task; relies on the inference backend's automatic
//! prefix caching for token savings.
//!
//! API specialists (Notion, Gmail, …) ride on the built-in `skills_agent`
//! definition by passing `skill_filter: "<skill_id>"`, which restricts
//! the resolved tool list to tools whose names start with `{skill}__`.
use super::traits::{PermissionLevel, Tool, ToolResult};
use crate::openhuman::agent::harness::archetypes::AgentArchetype;
use super::traits::{PermissionLevel, Tool, ToolCategory, ToolResult};
use crate::openhuman::agent::harness::definition::AgentDefinitionRegistry;
use crate::openhuman::agent::harness::fork_context::current_parent;
use crate::openhuman::agent::harness::subagent_runner::{run_subagent, SubagentRunOptions};
use crate::openhuman::event_bus::{publish_global, DomainEvent};
use async_trait::async_trait;
use serde_json::json;
/// Spawns a sub-agent of a specified archetype to handle a delegated task.
/// Available only to the Orchestrator archetype.
/// Spawns a sub-agent of the requested type to handle a delegated task.
///
/// Registered into the parent agent's tool list by
/// [`crate::openhuman::tools::all_tools_with_runtime`]. The orchestrator
/// archetype's tool whitelist already includes `spawn_subagent`, so
/// orchestrated runs see it; non-orchestrator parents see it too unless
/// explicitly removed.
pub struct SpawnSubagentTool;
impl Default for SpawnSubagentTool {
@@ -28,30 +58,78 @@ impl Tool for SpawnSubagentTool {
}
fn description(&self) -> &str {
"Spawn a specialised sub-agent to handle a task. Available archetypes: \
code_executor (writes/runs code), skills_agent (skill tools like Notion/Gmail), \
tool_maker (writes polyfills for missing commands), researcher (reads docs/web), \
critic (reviews code quality). Provide the archetype, a clear task prompt, and \
optional context from prior results."
"Spawn a specialised sub-agent to handle a focused sub-task. \
The sub-agent runs with a narrower prompt, filtered tool set, \
and (usually) a cheaper model. \n\n\
Tools are grouped into two categories — `system` (built-in \
Rust tools: shell, file_*, cron_*, memory_*, …) and `skill` \
(QuickJS skill bridges: notion__*, gmail__*, …). Set \
`category_filter` to dedicate the sub-agent to one category; \
for skill-tool execution the backend's `agentic` model hint \
is the natural fit.\n\n\
Built-in agent_ids include `planner` (DAG architect), \
`code_executor` (sandboxed coding), `skills_agent` (skill-tool \
execution via the agentic model; pair with `skill_filter` for \
per-API specialists like notion/gmail), `researcher` (web & \
docs), `critic` (read-only review), `tool_maker` (polyfill \
scripts), `archivist` (background lesson extraction), and \
`fork` (parallel decomposition with prefix-cache reuse). \
Custom sub-agents defined under `<workspace>/agents/*.toml` \
are also accepted by id."
}
fn parameters_schema(&self) -> serde_json::Value {
// Build the agent_id enum dynamically from the global registry
// when it's been initialised. Falls back to a string-with-hint
// when the registry hasn't been set up yet (e.g. early tests).
let agent_ids: Vec<String> = AgentDefinitionRegistry::global()
.map(|reg| reg.list().iter().map(|d| d.id.clone()).collect())
.unwrap_or_default();
let agent_id_schema = if agent_ids.is_empty() {
json!({
"type": "string",
"description": "Sub-agent id (e.g. code_executor, researcher, critic, fork)."
})
} else {
json!({
"type": "string",
"enum": agent_ids,
"description": "Sub-agent id from the registry."
})
};
json!({
"type": "object",
"required": ["archetype", "prompt"],
"required": ["agent_id", "prompt"],
"properties": {
"agent_id": agent_id_schema,
// Back-compat alias — older callers used `archetype`.
"archetype": {
"type": "string",
"enum": ["code_executor", "skills_agent", "tool_maker", "researcher", "critic"],
"description": "Which specialised sub-agent to spawn."
"description": "Deprecated alias for `agent_id`. Use `agent_id` going forward."
},
"prompt": {
"type": "string",
"description": "Clear, specific instruction for the sub-agent."
"description": "Clear, specific instruction for the sub-agent. The sub-agent has no memory of the parent's conversation, so include all context the sub-agent needs to act."
},
"context": {
"type": "string",
"description": "Optional context from prior task results or workspace state."
"description": "Optional context blob from prior task results. Rendered as a `[Context]` block before the prompt."
},
"skill_filter": {
"type": "string",
"description": "Optional skill id (e.g. `notion`, `gmail`) — when set, the sub-agent's tool list is restricted to tools named `{skill}__*`. Pair with `agent_id: skills_agent` for an API specialist."
},
"category_filter": {
"type": "string",
"enum": ["system", "skill"],
"description": "Optional tool-category restriction. `skill` scopes the sub-agent to QuickJS skill-bridge tools (Notion, Gmail, Telegram, …); `system` scopes it to built-in Rust tools (shell, file_*, memory_*, …). Overrides the definition's `category_filter` for this single spawn."
},
"mode": {
"type": "string",
"enum": ["typed", "fork"],
"description": "`typed` (default) builds a narrow prompt + filtered tools. `fork` replays the parent's exact prompt for prefix-cache reuse on the inference backend."
}
}
})
@@ -62,33 +140,236 @@ impl Tool for SpawnSubagentTool {
}
async fn execute(&self, args: serde_json::Value) -> anyhow::Result<ToolResult> {
let archetype_str = args
.get("archetype")
// ── Argument extraction with back-compat ───────────────────────
let agent_id = args
.get("agent_id")
.and_then(|v| v.as_str())
.unwrap_or("code_executor");
.or_else(|| args.get("archetype").and_then(|v| v.as_str()))
.unwrap_or("")
.trim()
.to_string();
let prompt = args.get("prompt").and_then(|v| v.as_str()).unwrap_or("");
let prompt = args
.get("prompt")
.and_then(|v| v.as_str())
.unwrap_or("")
.trim()
.to_string();
let _context = args.get("context").and_then(|v| v.as_str()).unwrap_or("");
let context = args
.get("context")
.and_then(|v| v.as_str())
.map(|s| s.to_string());
let skill_filter_override = args
.get("skill_filter")
.and_then(|v| v.as_str())
.map(|s| s.to_string());
let category_filter_override = match args.get("category_filter").and_then(|v| v.as_str()) {
Some("system") => Some(ToolCategory::System),
Some("skill") => Some(ToolCategory::Skill),
Some(other) => {
return Ok(ToolResult::error(format!(
"spawn_subagent: unknown category_filter '{other}' (expected 'system' or 'skill')"
)));
}
None => None,
};
let mode = args.get("mode").and_then(|v| v.as_str()).unwrap_or("typed");
// ── Validation ─────────────────────────────────────────────────
if agent_id.is_empty() {
return Ok(ToolResult::error(
"spawn_subagent: `agent_id` (or legacy `archetype`) is required",
));
}
if prompt.is_empty() {
return Ok(ToolResult::error("prompt is required"));
return Ok(ToolResult::error("spawn_subagent: `prompt` is required"));
}
// Parse archetype (validation).
let _archetype: AgentArchetype = serde_json::from_value(json!(archetype_str))
.map_err(|_| anyhow::anyhow!("unknown archetype: {archetype_str}"))?;
let registry = match AgentDefinitionRegistry::global() {
Some(reg) => reg,
None => {
return Ok(ToolResult::error(
"spawn_subagent: AgentDefinitionRegistry has not been initialised. \
This usually means the core process started without calling \
AgentDefinitionRegistry::init_global at startup.",
));
}
};
// Placeholder: In the full implementation, this will construct a sub-agent
// via AgentBuilder with the archetype's tool subset, model, and sandbox,
// then run its tool loop and return the result.
tracing::warn!(
"[spawn_subagent] no-op — would spawn {archetype_str} sub-agent with prompt length={}",
prompt.len()
);
// Resolve `mode` against the definition. Explicit `mode` argument
// wins; otherwise we infer from the definition itself.
let lookup_id = if mode == "fork" {
"fork"
} else {
agent_id.as_str()
};
let definition = match registry.get(lookup_id) {
Some(def) => def,
None => {
let available: Vec<&str> = registry.list().iter().map(|d| d.id.as_str()).collect();
return Ok(ToolResult::error(format!(
"spawn_subagent: unknown agent_id '{lookup_id}'. Available: {}",
available.join(", ")
)));
}
};
Ok(ToolResult::error(
"spawn_subagent not yet wired to sub-agent execution",
// ── Validate skill filter against the runtime if set ───────────
if let Some(skill) = skill_filter_override.as_deref() {
if let Err(err) = validate_skill_filter(skill) {
return Ok(ToolResult::error(err));
}
}
// ── Publish SubagentSpawned event ──────────────────────────────
let parent_session = current_parent()
.map(|p| p.session_id.clone())
.unwrap_or_else(|| "standalone".into());
let task_id = format!("sub-{}", uuid::Uuid::new_v4());
publish_global(DomainEvent::SubagentSpawned {
parent_session: parent_session.clone(),
agent_id: definition.id.clone(),
mode: mode.to_string(),
task_id: task_id.clone(),
prompt_chars: prompt.chars().count(),
});
// ── Run the sub-agent ──────────────────────────────────────────
let options = SubagentRunOptions {
skill_filter_override,
category_filter_override,
context,
task_id: Some(task_id.clone()),
};
match run_subagent(definition, &prompt, options).await {
Ok(outcome) => {
publish_global(DomainEvent::SubagentCompleted {
parent_session,
task_id: outcome.task_id.clone(),
agent_id: outcome.agent_id.clone(),
elapsed_ms: outcome.elapsed.as_millis() as u64,
output_chars: outcome.output.chars().count(),
iterations: outcome.iterations,
});
Ok(ToolResult::success(outcome.output))
}
Err(err) => {
let message = err.to_string();
publish_global(DomainEvent::SubagentFailed {
parent_session,
task_id,
agent_id: definition.id.clone(),
error: message.clone(),
});
// Surface as a non-fatal tool error so the parent model
// can react and (e.g.) retry with different params.
Ok(ToolResult::error(format!(
"spawn_subagent failed: {message}"
)))
}
}
}
}
/// Validate that the requested skill_filter matches a currently-loaded
/// skill in the global skill runtime, if a runtime is available. When
/// no runtime is available (e.g. tests, CLI), this is a no-op.
fn validate_skill_filter(skill_id: &str) -> Result<(), String> {
let Some(engine) = crate::openhuman::skills::qjs_engine::global_engine() else {
// No runtime registered — skip validation.
return Ok(());
};
let mut known: Vec<String> = engine
.all_tools()
.into_iter()
.filter_map(|(_, def)| {
// Tool names are `{skill}__{tool}`; pull just the prefix.
def.name
.split_once("__")
.map(|(prefix, _)| prefix.to_string())
})
.collect();
known.sort();
known.dedup();
if known.iter().any(|s| s == skill_id) {
Ok(())
} else {
Err(format!(
"skill_filter '{skill_id}' does not match any installed skill. Available: {}",
known.join(", ")
))
}
}
#[cfg(test)]
mod tests {
use super::*;
#[tokio::test]
async fn missing_agent_id_returns_error() {
let tool = SpawnSubagentTool;
let result = tool
.execute(json!({
"prompt": "do thing"
}))
.await
.unwrap();
assert!(result.is_error);
assert!(result.output().contains("agent_id"));
}
#[tokio::test]
async fn missing_prompt_returns_error() {
let tool = SpawnSubagentTool;
let result = tool
.execute(json!({
"agent_id": "researcher"
}))
.await
.unwrap();
assert!(result.is_error);
assert!(result.output().contains("prompt"));
}
#[tokio::test]
async fn no_registry_returns_clear_error() {
// The global registry has not been initialised in this test.
let tool = SpawnSubagentTool;
let result = tool
.execute(json!({
"agent_id": "researcher",
"prompt": "find x",
}))
.await
.unwrap();
// Either: registry uninitialised → clear init error, OR
// registry was initialised by a previous test → "no parent context"
// because we're not running inside an Agent::turn. Both are
// acceptable: the tool gracefully refuses.
assert!(result.is_error);
}
#[tokio::test]
async fn unknown_agent_id_lists_available() {
// Force-init the global registry with builtins.
let _ = AgentDefinitionRegistry::init_global_builtins();
let tool = SpawnSubagentTool;
let result = tool
.execute(json!({
"agent_id": "totally_made_up",
"prompt": "x",
}))
.await
.unwrap();
assert!(result.is_error);
let out = result.output();
// Should list at least one valid built-in.
assert!(out.contains("code_executor") || out.contains("researcher"));
}
}
+47
View File
@@ -16,6 +16,42 @@ pub enum ToolScope {
CliRpcOnly,
}
/// Category of a tool — used by the sub-agent runner to scope which
/// tools a given sub-agent is allowed to see.
///
/// The distinction matters because:
///
/// - **System tools** are built-in Rust implementations (shell, file_read,
/// file_write, cron_*, memory_*, …) that run inside the core process
/// with direct host access.
/// - **Skill tools** are QuickJS skill exports bridged through
/// [`crate::openhuman::tools::skill_bridge::SkillToolBridge`]. They
/// talk to external services (Notion, Gmail, Telegram, …) via
/// user-installed skill packages.
///
/// The orchestrator uses this category to spawn dedicated tool-execution
/// sub-agents: one scoped to `Skill` for service integrations (running
/// with the backend's `agentic` model hint), and others scoped to
/// `System` for code/file/host work.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default, Serialize, Deserialize)]
#[serde(rename_all = "snake_case")]
pub enum ToolCategory {
/// Built-in Rust tools with direct host access.
#[default]
System,
/// QuickJS skill tools bridged from the runtime engine.
Skill,
}
impl std::fmt::Display for ToolCategory {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::System => write!(f, "system"),
Self::Skill => write!(f, "skill"),
}
}
}
/// Permission level required to execute a tool.
///
/// Channels can set a maximum permission level to restrict which tools
@@ -85,6 +121,17 @@ pub trait Tool: Send + Sync {
ToolScope::All
}
/// Category of this tool — `System` for built-in Rust tools (default)
/// or `Skill` for tools bridged from the QuickJS skill runtime.
///
/// The sub-agent runner uses this to filter the parent's tool
/// registry when a sub-agent definition sets `category_filter`.
/// Skill-bridged tools override this to return
/// [`ToolCategory::Skill`].
fn category(&self) -> ToolCategory {
ToolCategory::System
}
/// Get the full spec for LLM registration
fn spec(&self) -> ToolSpec {
ToolSpec {