Files
openhuman/docs/tinyagents-migration-spec.md
T

45 KiB

TinyAgents Migration Spec

Status: draft migration backlog

TinyAgents source reviewed: tinyhumansai/tinyagents origin/main at 8f226f1, crate version 1.1.0. Refreshed against tinyhumansai/tinyagents main at 348a0e7dc71a1f9039f3d523a2a384661a7a9acd after the SDK/docs update. Current OpenHuman dependency in this checkout is tinyagents = { version = "1.5.0", features = ["sqlite"] }.

OpenHuman already depends on TinyAgents and already routes the live agent turn through src/openhuman/tinyagents/. This spec is not a proposal to add TinyAgents. It is a todo list for moving the rest of OpenHuman's generic agent runtime behavior onto TinyAgents primitives while keeping OpenHuman-owned product semantics in OpenHuman.

Current inventory snapshot: tinyagents-harness-migration-audit.md.

Goal

Use TinyAgents as the generic runtime for:

  • model/provider abstraction and model selection
  • tools and tool schemas
  • middleware around model/tool calls
  • streaming, events, traces, and replayable run status
  • session transcript storage and migration targets
  • prompt/response cache layout protection
  • token usage and cost rollups
  • state graphs, fanout, reducers, checkpoints, and interrupts
  • sub-agent recursion, steering, cancellation, and reusable sessions
  • deterministic testkit coverage for the runtime seams

OpenHuman should continue to own:

  • desktop product UX and Tauri/RPC boundaries
  • user/workspace config, credentials, keychain, and approval records
  • OpenHuman memory stores, thread transcripts, run ledgers, and controllers
  • security policy, sandboxing, tool permission tiers, and workspace roots
  • product-specific built-in agents, prompts, MCP setup, Composio, channels, and native tools
  • compatibility with existing JSON-RPC method names and persisted state

Sources Reviewed

TinyAgents SDK:

  • src/lib.rs
  • src/harness/*
  • src/graph/*
  • src/registry/*
  • src/language/*
  • src/repl/*
  • docs/modules/harness/*.md
  • docs/modules/graph/*.md
  • docs/modules/registry/*.md
  • docs/modules/expressive-language/README.md
  • docs/modules/repl-language/README.md
  • examples: agent_loop_tools, orchestrator_subagents, durable_graph, openai_graph_agent, openai_self_blueprint

OpenHuman Rust core:

  • src/openhuman/tinyagents/*
  • src/openhuman/agent/**
  • src/openhuman/agent_orchestration/**
  • src/openhuman/agent_registry/**
  • src/openhuman/tools/**
  • src/openhuman/inference/**
  • src/openhuman/cost/**
  • src/openhuman/context/**
  • src/openhuman/approval/**
  • src/openhuman/security/**
  • src/openhuman/mcp_registry/**
  • src/core/event_bus/**
  • src/core/all.rs
  • gitbooks/developing/architecture/agent-harness.md

Current Adoption Inventory

Already done or partially done:

  • Cargo.toml pins tinyagents = { version = "1.5.0", features = ["sqlite"] }.
  • src/openhuman/tinyagents/mod.rs registers OpenHuman Provider and Tool adapters on tinyagents::harness::runtime::AgentHarness.
  • ProviderModel maps OpenHuman ChatRequest/ChatResponse into tinyagents::harness::model::{ModelRequest, ModelResponse, ModelStream}.
  • ToolAdapter and SharedToolAdapter map OpenHuman tools into tinyagents::harness::tool::Tool.
  • OpenhumanEventBridge maps TinyAgents AgentEvent into AgentProgress and the global cost tracker.
  • StopHookMiddleware, ContextCompressionMiddleware, and MessageTrimMiddleware are already used on the TinyAgents path.
  • SqlRunLedgerCheckpointer implements TinyAgents Checkpointer on top of the OpenHuman session DB while the migration re-points checkpoint rows to the crate checkpointer.
  • spawn_parallel_graph uses GraphBuilder and graph::parallel::map_reduce for reusable concurrent fanout.
  • model_council, workflow_runs, agent_teams, and tinyagents/delegation.rs already use TinyAgents graphs.
  • Built-in agents already have graph.rs selectors, but most return AgentGraph::Default.

Important current gaps:

  • OpenHuman still has separate registries for agents, tools, MCP tools, model providers, controllers, cost, and event bus projections.
  • Tool safety metadata exists in OpenHuman traits but is not fully expressed as TinyAgents tool safety/runtime metadata.
  • Cost/usage is still converted through a bridge, not an end-to-end TinyAgents usage/cost journal.
  • Event streams are mirrored into AgentProgress, but TinyAgents event journals and status stores are not the canonical durable inspection surface.
  • Provider model profiles, model resolution, and fallback remain primarily in OpenHuman inference/router logic.
  • Sub-agent lifecycle, durable state, worker threads, and wait/abort controls still live in OpenHuman orchestration stores.

Some todos below are local adapter work. Others still require upstream TinyAgents SDK extensions. In particular, the SDK has a strong tool/runtime boundary today (ToolSchema, ToolExecutionContext, middleware hooks), but OpenHuman's full tool safety metadata is richer than the current SDK schema. After the TinyAgents main refresh, durable session/cache/journal primitives are no longer the broad SDK gap they were in the original baseline; the remaining work is to design OpenHuman's compatibility adapter, migrate old transcripts and run-ledger rows, and prove restart/resume parity.

Migration Rules

  • Keep every product-facing JSON-RPC contract stable unless a migration plan is written next to the code change.
  • Do not bypass OpenHuman approval, security policy, sandbox, workspace root, or credential boundaries by adopting a generic TinyAgents tool API.
  • Prefer adapters first, then flip ownership once tests prove parity.
  • Preserve existing transcript and run-ledger compatibility. TinyAgents may become the internal runtime without changing persisted public records in the same PR.
  • Old OpenHuman session JSONL/Markdown data should move through a one-time, idempotent migration script into TinyAgents-compatible store/journal records before the old readers are deleted.
  • Every migration task needs unit coverage plus at least one JSON-RPC or harness-level e2e when behavior crosses controller, tool, provider, or graph boundaries.

Phase 0 - Baseline And Drift Control

  • Add a version/feature compatibility note to the OpenHuman architecture doc.
    • OpenHuman files: gitbooks/developing/architecture/agent-harness.md, Cargo.toml.
    • TinyAgents components: crate features default, openai, sqlite, repl.
    • Acceptance: document why default features are used, why TinyAgents sqlite is enabled through the aligned rusqlite stack, and which OpenHuman adapters still replace SDK-owned providers.
    • Done: added "TinyAgents crate: features & compatibility" section to agent-harness.md (default-only, openai/sqlite/repl rationale, adapter map) + fixed stale council_graph.rs/member_graph.rs links.

Phase 1 - Tools

  • [~] Make OpenHuman tool metadata round-trip into TinyAgents tool metadata.

    • OpenHuman files: src/openhuman/tools/traits.rs, src/openhuman/tinyagents/tools.rs, src/openhuman/tinyagents/convert.rs.
    • TinyAgents components: harness::tool::{ToolSchema, ToolFormat, ToolExecutionContext, ToolResult}.
    • Migrate: permission level, external effect, generated runtime context, timeout policy, concurrency safety, result-size cap, display label/detail, markdown support.
    • Acceptance: a TinyAgents tool call has enough metadata for middleware to enforce approval, security, timeout, concurrency, truncation, and display behavior without re-querying OpenHuman trait methods ad hoc.
    • 1.3 update: crate ToolPolicy, Tool::policy(), ToolRegistry::policies(), and ToolPolicyMiddleware now provide the SDK-owned safety/runtime/access projection. ToolSchema still carries only name/description/parameters/format — it has no model-visible metadata/extension map — so OpenHuman should map enforcement fields into ToolPolicy while keeping display/schema annotations as app-side metadata. Existing side-lookup middleware (ApprovalSecurityMiddleware, ToolOutputMiddleware) can shrink as those policy snapshots become the source of truth.
  • Move unknown-tool recovery into a reusable middleware or tool policy layer.

    • Current path: run_policy_for sets UnknownToolPolicy::ReturnToolError; no sentinel tool is registered.
    • TinyAgents components: ToolRegistry, ToolMiddleware, AgentEvent::ToolStarted/ToolCompleted, repairable tool results.
    • Acceptance: hallucinated tool names remain recoverable, sub-agent wording is preserved, and TinyAgents event stream records the original requested tool name without exposing the sentinel as a model-visible tool.
    • 1.3 update: crate RunPolicy.unknown_tool now has UnknownToolPolicy::{Fail, ReturnToolError, Rewrite} and emits AgentEvent::UnknownToolCall with the original requested name/arguments. OpenHuman now uses that policy directly; UNKNOWN_TOOL_SENTINEL and UnknownToolRewriteMiddleware are gone from source.
  • Route approval and security through TinyAgents middleware.

    • Current OpenHuman files: src/openhuman/approval/*, src/openhuman/security/*, src/openhuman/tinyagents/tools.rs.
    • TinyAgents components: ToolMiddleware, ToolExecutionContext, tool safety metadata.
    • Acceptance: approval checks happen in before_tool/wrap_tool, emit typed events, preserve audit rows, and return model-consumable denial results.
    • Done: ApprovalSecurityMiddleware (tinyagents/middleware.rs, a wrap_tool middleware) replaces the inline approval block in execute_openhuman_tool. Denials short-circuit with a model-consumable result; approved external-effect calls now record a terminal audit row (record_execution) the old path dropped. Typed approval events still ride DomainEvent (the crate AgentEvent enum has no approval variant — SDK gap). Tool-internal security (path/command live_policy) stays per-tool by design. Follow-ups: channel permission-ceiling threading; per-tool metadata side-lookup (Task C).
  • Use TinyAgents bounded-concurrent tool execution where safe.

    • OpenHuman files: src/openhuman/tools/traits.rs, src/openhuman/tinyagents/tools.rs.
    • TinyAgents components: graph Send, graph fanout, or harness tool execution policy.
    • Acceptance: read-only/concurrency-safe tool batches can run in parallel with deterministic result ordering and identical transcript semantics.

Phase 2 - Models And Providers

  • Register OpenHuman inference providers as TinyAgents model registry entries.

    • OpenHuman files: src/openhuman/inference/provider/*, src/openhuman/inference/model_ids.rs, src/openhuman/tinyagents/model.rs.
    • TinyAgents components: harness::model::{ChatModel, ModelRegistry, ModelProfile, CapabilitySet, ModelRequest, ModelResponse}.
    • Acceptance: every workload route (agentic, reasoning, coding, memory, subconscious, etc.) can resolve to a TinyAgents model entry while retaining OpenHuman provider strings and config compatibility.
  • [~] Translate OpenHuman provider capability data into TinyAgents model profiles.

    • OpenHuman files: src/openhuman/inference/provider/traits.rs, src/openhuman/inference/provider/factory.rs, docs/inference-provider-catalog.md.
    • TinyAgents components: ModelProfile, CapabilitySet, registry model catalog.
    • Acceptance: context window, tool calling, streaming, vision, structured output, reasoning, local/cloud source, and provider-family metadata are available before dispatch.
    • Partial: every ProviderModel registered by the shared runner now carries a crate ModelProfile built at construction from the provider's canonical capability accessors — tool calling (+parallel), vision (modalities.image_in), streaming, local/remote source — plus the runner-threaded token limits (with_context_windowmax_input_tokens, output cap → max_output_tokens). ChatModel::profile() returns it, so the crate's pre-dispatch validation and structured-output strategy see real capabilities. Remaining: structured-output/JSON-schema/reasoning flags (no OpenHuman capability source yet), release/status metadata, and a registry-level model catalog (ties into the workload-route registry item above).
  • Move model fallback and retry policy to TinyAgents policy/middleware.

    • OpenHuman files: src/openhuman/inference/provider/reliable.rs, src/openhuman/inference/provider/router.rs, src/openhuman/tinyagents/mod.rs.
    • TinyAgents components: RunPolicy, RetryPolicy, FallbackPolicy, ModelFallbackMiddleware, AgentEvent::RetryScheduled, AgentEvent::FallbackSelected.
    • Acceptance: OpenHuman provider retry does not double-retry under TinyAgents, fallback events are typed, and tests cover transient 429/5xx, config rejection, and billing exhaustion.
  • Preserve provider-specific metadata in the TinyAgents message model.

    • OpenHuman files: src/openhuman/inference/provider/traits.rs, src/openhuman/tinyagents/convert.rs, src/openhuman/tinyagents/model.rs.
    • TinyAgents components: ContentBlock, AssistantMessage, provider metadata, tool-call ids.
    • Acceptance: Gemini thought signatures, reasoning content, native tool-call ids, cached tokens, and raw provider metadata survive multi-turn history.

Phase 3 - Middleware

  • Convert OpenHuman turn cross-cuts into named TinyAgents middleware.

    • Current OpenHuman surfaces: stop hooks, approval gate, security policy, output caps, context compression, memory injection, tool allowlists, cost/usage, prompt cache stability, event bridge.
    • TinyAgents components: Middleware, ModelMiddleware, ToolMiddleware, MiddlewareStack, RunContext.
    • Acceptance: each cross-cut has a stable middleware name, tests for ordering, emitted events, and explicit interaction with streaming/retry/fallback.
  • Add OpenHuman policy middleware for dynamic tool exposure.

    • OpenHuman files: src/openhuman/agent_registry/*, src/openhuman/agent/harness/subagent_runner/**, src/openhuman/tools/user_filter.rs.
    • TinyAgents components: before_model, before_tool, ToolRegistry.
    • Acceptance: agent tool_allowlist, tool_denylist, sub-agent tool scope, MCP tool visibility, and channel permission ceilings are enforced through middleware rather than scattered pre-filtering.
  • Add prompt/cache-layout middleware tests.

    • OpenHuman files: src/openhuman/context/*, src/openhuman/agent/harness/session/turn/core.rs, src/openhuman/tinyagents/summarize.rs.
    • TinyAgents components: harness::cache::{CachePolicy, CacheLayoutEvent, PromptCacheLayout, ResponseCache}, ContextCompressionMiddleware, MessageTrimMiddleware.
    • Acceptance: system prompt prefix remains stable across later turns; volatile memory, timestamps, tool results, and steering messages land in the tail.
  • Move OpenHuman response-cache and provider KV-cache protection onto TinyAgents cache primitives.

    • OpenHuman files: src/openhuman/agent/harness/session/turn/core.rs, src/openhuman/context/*, src/openhuman/tinyagents/middleware.rs.
    • TinyAgents components: harness::cache::{ResponseCache, PromptCacheLayout, CachePolicy}, AgentEvent::CacheHit, AgentEvent::CacheMiss.
    • Acceptance: repeated deterministic model requests can be served by the TinyAgents response cache where safe, prompt-prefix stability is asserted by PromptCacheLayout, and OpenHuman cache-align warnings become TinyAgents cache-layout events.

Phase 4 - Events, Status, And Observability

  • Make TinyAgents event journals the canonical internal run event stream.

    • OpenHuman files: src/openhuman/tinyagents/observability.rs, src/core/event_bus/*, src/openhuman/notifications/*, src/openhuman/session_db/run_ledger/*.
    • TinyAgents components: HarnessEventJournal, HarnessStatusStore, GraphEventJournal, GraphStatusStore, harness::store::AppendStore, harness::store::JsonlAppendStore, AgentEvent, GraphEvent.
    • Acceptance: UIs can reconstruct a running or completed agent turn from persisted TinyAgents events without relying only on transient AgentProgress.
  • Write a one-time OpenHuman session transcript migration into TinyAgents store/journal records. Done: src/openhuman/session_import/ (openhuman.session_import_run), design + as-built notes in docs/tinyagents-session-migration-design.md. Write-only Phase 1; read-side shadow/cutover remain.

    • OpenHuman files: src/openhuman/agent/harness/session/transcript.rs, src/openhuman/agent/harness/session/migration.rs, user workspace session_raw/*.jsonl, legacy Markdown session directories.
    • TinyAgents components: harness::store::{Store, AppendStore, FileStore, JsonlAppendStore}, harness::message::Message, harness event/status records with thread_id, run_id, root_run_id, and stream offsets.
    • Acceptance: old OpenHuman sessions are imported idempotently, preserving timestamps, transcript stems, parent/child session links, provider/model metadata, tool-call ids, native/XML/P-format history, and malformed-file warnings. Existing OpenHuman readers remain as compatibility projections until parity fixtures prove the migration.
  • Bridge TinyAgents events into DomainEvent as a compatibility projection.

    • OpenHuman files: src/core/event_bus/events.rs, src/openhuman/agent/bus.rs, src/openhuman/tinyagents/observability.rs.
    • Acceptance: existing subscribers continue to receive DomainEvent, but new code reads TinyAgents events/status first.
  • Persist graph run status and checkpoint metadata in OpenHuman run ledger.

    • OpenHuman files: src/openhuman/tinyagents/checkpoint.rs, src/openhuman/session_db/run_ledger/store.rs, src/openhuman/agent_orchestration/**.
    • TinyAgents components: Checkpoint, CheckpointMetadata, GraphRunStatus, GraphObservation.
    • Acceptance: command center, workflow runs, delegation, and team runs can list checkpoints, current node/task status, and replay offsets from one DB source.
  • Export graph topology for debugging and UI inspection.

    • OpenHuman files: built-in graph.rs files, agent_orchestration/*/graph.rs, model_council/graph.rs.
    • TinyAgents components: GraphTopology, to_json, to_mermaid, validation report.
    • Acceptance: every custom OpenHuman graph has a debug endpoint or test snapshot that exports topology and validates missing nodes/routes.
    • Done: tinyagents/topology.rsGraphTopologyReport (mermaid + JSON + validation errors/warnings), describe(), and all_graph_topologies(). Pattern: each graph exposes a build_*_graph (structure) reused by both the runner and a *_topology() that builds it with no-op stub closures and returns CompiledGraph::topology(). Exported graphs: agent_teams:member, delegation (extracted build_delegation_graph), workflow_runs:scheduler (build_scheduler_graph with injected select/run engine effects), subagent:pipeline, and spawn_parallel_agents. Generic map-reduce fan-outs such as council runs are still item-count-driven dispatch→N→collect patterns, not fixed named topologies, so they are intentionally not exported. Debug endpoint: agent.graph_topologies JSON-RPC controller (agent/schemas.rs) returning {name, ok, errors, warnings, mermaid, topology} per graph.

Phase 5 - Usage, Cost, And Budgets

  • [~] Replace bridge-only usage accounting with TinyAgents usage records.

    • OpenHuman files: src/openhuman/cost/*, src/openhuman/tinyagents/observability.rs, src/openhuman/inference/provider/traits.rs.
    • TinyAgents components: harness::usage::{Usage, UsageTotals}, harness::cost::CostTotals, AgentEvent::UsageRecorded.
    • Acceptance: input, output, cached input, reasoning, image/audio, embedding, tool/model call counts, and estimated/provider-reported source are recorded in normalized records.
    • Partial (real bug fixed): the bridge hardcoded charged_amount_usd: 0.0 and ProviderModel dropped cached tokens, so EVERY tinyagents turn recorded $0 cost. Now model.rs carries cached_input_tokens via crate Usage.cache_read_tokens, and the bridge estimates per-call cost from catalogued per-MTok rates (cost::catalog::estimate_cost_usd). Remaining: reasoning/image/audio/embedding token fields, model/tool call counts, and an explicit estimated-vs-provider-charged cost_source tag on TokenUsage (provider-charged preservation needs an out-of-band carry — crate Usage has no USD field).
  • [~] Move budget checks to pre-call TinyAgents middleware.

    • OpenHuman files: src/openhuman/cost/tracker.rs, src/openhuman/tinyagents/mod.rs.
    • TinyAgents components: RunPolicy, cost middleware, before_model, before_tool.
    • Acceptance: per-run, per-thread, daily, and monthly budgets can warn or fail before spend where enough data exists, then reconcile after provider usage is known.
    • Partial: CostBudgetMiddleware (before_model) fails the run before a model call when the global daily/monthly budget is already exceeded (CostTracker::check_budget), and logs on the warning threshold. Self-gating on config.enabled; previously daily/monthly enforcement was dormant on the tinyagents path. Remaining: per-run/per-thread budgets (need new CostConfig fields + thread-id threading into the runner) and projecting the next call's cost pre-spend (needs an input-token estimate).
  • [~] Add cost rollup across sub-agents and graphs.

    • OpenHuman files: src/openhuman/agent_orchestration/**, src/openhuman/cost/global.rs.
    • TinyAgents components: run ids, parent/root run lineage, SubAgentStarted, SubAgentCompleted, graph child runs.
    • Acceptance: parent run totals include child agent/model/tool usage without double counting dashboard totals.
    • Partial (audit + real gap fixed). Audit of the current mechanics: (1) parent-turn rollup — the turn_subagent_usage task-local collector wraps the turn future; run_typed_mode (the single sub-agent chokepoint) records every inline child, and graph fan-outs (spawn_parallel_graph, delegation, council) execute via join_all on the same task, so their children inherit the collector too; (2) dashboard totals — the global tracker is fed per model call by each run's own event bridge, and the parent's fold into LastTurnUsage/transcript never re-records to the tracker, so there is no double counting. Gap fixed: the tracker feed lived only in the bridge, so an unobserved (on_progress = None, fire-and-forget) turn's spend never reached the dashboard — the runner's cost fallback now records the aggregate via record_unobserved_turn_usage (mutually exclusive with the bridge → exactly-once). Remaining: crate run-id / parent-root lineage on cost records (needs TokenUsage schema fields), and rollup for detached background children beyond global-tracker capture (documented behavior today).

Phase 6 - Graph Runtime And Orchestration

  • Convert remaining ad hoc control loops into explicit TinyAgents graphs.

    • Candidate OpenHuman files: src/openhuman/agent_orchestration/*, src/openhuman/subconscious/*, src/openhuman/cron/*, src/openhuman/learning/*, src/openhuman/tools/ops.rs.
    • TinyAgents components: GraphBuilder, Command, conditional routing, reducers, Send, barriers, recursion policy.
    • Acceptance: every long-running multi-step orchestration has named nodes, route tests, recursion bounds, cancellation checks, and topology export.
  • Replace simple fanout helpers with graph Send where payload-specific fanout matters.

    • Current helper: src/openhuman/tinyagents/orchestration.rs.
    • TinyAgents components: Command::send, GraphInput, NodeContext::send_arg.
    • Acceptance: map-reduce style flows can schedule multiple invocations of the same node with distinct payloads instead of materializing one node per item.
  • Make spawn_parallel_agents a first-class TinyAgents graph tool.

    • Current OpenHuman files: src/openhuman/agent_orchestration/tools/spawn_parallel_agents.rs, src/openhuman/tinyagents/orchestration.rs, src/openhuman/agent_orchestration/worktree.rs, src/openhuman/agent_orchestration/workflow_runs/engine.rs.
    • Current behavior: validates at least two tasks, checks parent context, enforces max_parallel_tools, resolves AgentDefinitions, enforces the parent subagents.allowlist, optionally creates per-worker git worktrees, fans workers out through spawn_parallel_graph + map_reduce, collects results in input order, emits DomainEvent + AgentProgress, detects stale parent file reads and cross-worker changed-file overlaps, and returns a structured parallel_agents JSON payload.
    • TinyAgents components: GraphBuilder, Command, Send, NodeContext::send_arg, ChannelSet/reducers, GraphEventSink, SubAgentNode or OpenHuman run_subagent adapter, RecursionPolicy, RunPolicy, CancellationToken.
    • Required migration shape:
      • validate node: parse tasks, enforce min/max count, parent context, allowlist, toolkit requirements, and worktree preflight.
      • dispatch node: use Send to schedule one worker invocation per task instead of generating one static worker node per task.
      • worker node: run the OpenHuman sub-agent build pipeline with inherited model/tool/security policy, optional worktree_action_dir, child task id, and bounded turn/output budgets.
      • collect reducer: aggregate successes, failures, elapsed time, iterations, worktree status, changed files, and stale read markers in deterministic task order.
      • finalize node: emit compatibility DomainEvent/AgentProgress projections, overlap warnings, and the existing JSON result shape.
    • Acceptance: parallel agent runs are checkpointable, cancellable at graph boundaries, bounded by parent policy, observable through TinyAgents graph events/status, compatible with existing spawn_parallel_agents tool output, and able to run edit-capable workers in isolated worktrees without silently falling back to shared workspace.
  • Define parallel-agent ownership and scheduling policy explicitly.

    • OpenHuman files: src/openhuman/agent_registry/agents/orchestrator/prompt.md, src/openhuman/agent_orchestration/tools/spawn_parallel_agents.rs, src/openhuman/tools/traits.rs.
    • TinyAgents components: graph route metadata, RunPolicy, task metadata, tool safety metadata.
    • Required policy:
      • The parent must provide disjoint ownership boundaries for write-capable tasks, or the graph rejects/falls back to serial delegation.
      • Read-only workers may share the parent workspace; write-capable workers should request isolation = "worktree".
      • Parent and children inherit one root run id for cost/event rollups, but each child gets its own task id and optional worker-thread id.
      • A child cannot widen tools, model choice, sandbox mode, trusted roots, or budget beyond the parent-granted policy.
      • Cancellation, steering, and wait/collect must be delivered at graph or harness safe boundaries only.
    • Acceptance: the orchestrator can ask for parallelism without prompt-only conventions; policy violations become structured graph/tool errors.
  • Make per-agent graph.rs selectors real customization points.

    • OpenHuman files: src/openhuman/agent_registry/agents/*/graph.rs, src/openhuman/agent/harness/agent_graph.rs.
    • TinyAgents components: CompiledGraph, sub-agent nodes, graph testkit.
    • Acceptance: at least three agents get bespoke graphs where useful: orchestrator, researcher, and tool_maker are good first candidates.
  • [~] Keep durable orchestration stores OpenHuman-owned until the OpenHuman compatibility adapter is written.

    • Current OpenHuman files: running_subagents.rs, workflow_runs, agent_teams, command_center, subagent_sessions.
    • Updated TinyAgents status: current main has harness stores, JSONL append journals, lineage-aware harness/graph status, SubAgentSession, subgraph nodes, and sub-agent graph nodes. This means the blocker is no longer only "SDK lacks storage"; it is now the OpenHuman adapter/migration design and restart/resume parity.
    • Acceptance: migrate durable SQL/JSON state through a compatibility adapter, not by dropping records into in-memory task storage. OpenHuman controllers continue to read compatible projections while TinyAgents records become the canonical internal state.
  • Add graph interrupt/resume for human review points.

    • OpenHuman files: src/openhuman/approval/*, src/openhuman/agent_orchestration/workflow_runs/*, src/openhuman/tinyagents/delegation.rs.
    • TinyAgents components: Interrupt, ResumeTarget, Command::resume, checkpoints.
    • Acceptance: approval/review pauses are durable graph interrupts where the run can resume from the exact checkpoint after user action.

Phase 7 - Sub-Agents, Steering, And Recursion

  • Represent spawn_subagent, steer_subagent, wait_subagent, and follow-ups as TinyAgents steering commands plus OpenHuman projections.

    • OpenHuman files: src/openhuman/agent_orchestration/tools.rs, src/openhuman/agent_orchestration/running_subagents.rs, src/openhuman/agent/harness/subagent_runner/**.
    • TinyAgents components: SteeringCommand, SteeringHandle, SubAgentSession, SubAgentTool, recursion depth events.
    • Acceptance: mid-flight messages are delivered only at safe loop boundaries, accepted/rejected steering emits events, and tool/model allowlists can only narrow from parent policy.
  • Re-express the OpenHuman sub-agent pipeline as a TinyAgents subgraph.

    • OpenHuman files: src/openhuman/agent/harness/subagent_runner/**, src/openhuman/agent_orchestration/tools/spawn_subagent.rs, src/openhuman/agent_orchestration/subagent_sessions/**.
    • TinyAgents components: harness::subagent::{SubAgent, SubAgentSession, SubAgentTool}, graph::subagent_node::{SubAgentNode, SubAgentPolicy}, graph::subgraph::{shared_subgraph_node, adapter_subgraph_node}, CapabilityRegistry, graph status/observability.
    • Acceptance: definition resolution, tool filtering, prompt assembly, toolkit preflight, sandbox/action-root narrowing, handoff cache, checkpoint/awaiting-user handback, and worker-thread mirroring are explicit graph nodes or adapters. The final child run has TinyAgents lineage, status, usage/cost rollup, and transcript storage; OpenHuman keeps only product policy nodes and compatibility response formatting.
  • Reconcile OpenHuman spawn depth with TinyAgents recursion policy.

    • OpenHuman files: src/openhuman/agent/harness/spawn_depth_context.rs, src/openhuman/agent/harness/subagent_runner/**.
    • TinyAgents components: RecursionPolicy, RecursionStack, SubAgentDepth.
    • Acceptance: there is one authoritative recursion cap and one error shape, with compatibility conversion for existing UI/JSON-RPC responses.
  • [~] Keep OpenHuman's sub-agent build pipeline as product policy, but move the pipeline mechanics into TinyAgents graph/sub-agent primitives.

    • Product-owned pieces: agent definition resolution, prompt assembly, memory context, worker-thread mirroring, handoff cache, tool filtering, provider routing, sandbox scope.
    • TinyAgents components to adopt beneath it: SubAgentSession, SubAgentTool, SubAgentNode, SubAgentPolicy, ToolExecutionContext, event lineage, graph status, cancellation, usage/cost rollup.
    • Acceptance: use TinyAgents for execution, transcript/session tracking, graph structure, and lineage without flattening OpenHuman's agent registry semantics into generic SDK defaults.

Phase 8 - Registry And Capability Catalog

  • Build a CapabilityRegistry projection from OpenHuman registries.

    • OpenHuman files: agent_registry, tools, mcp_registry, inference, cost, approval, security, controller registry in src/core/all.rs.
    • TinyAgents components: CapabilityRegistry, ComponentId, ComponentKind, model catalog, graph/tool/agent/store/middleware entries.
    • Acceptance: OpenHuman models, tools, agents, graphs, stores, and middleware can be looked up through one policy-aware capability projection.
  • Add registry diagnostics for duplicate names and unsafe aliases.

    • OpenHuman files: src/openhuman/tools/generated.rs, mcp_registry, composio, agent_registry.
    • TinyAgents components: component names, ToolRegistry, diagnostics.
    • Acceptance: duplicate tool names, provider-specific aliases, MCP names, and generated tool names fail closed before model dispatch.
  • Use TinyAgents model catalog shape for local provider catalog snapshots.

    • OpenHuman files: docs/inference-provider-catalog.md, src/openhuman/inference/presets.rs, src/openhuman/inference/provider/factory.rs.
    • TinyAgents components: registry model catalog, local snapshots, price and context metadata.
    • Acceptance: model picker, router, budget estimator, and capability filter read one normalized catalog projection.

Phase 9 - Memory, Retrieval, Embeddings, Context, And Cache

  • Adapt OpenHuman memory/retrieval to TinyAgents retriever interfaces.

    • OpenHuman files: memory, memory_search, memory_tree, agent_memory/memory_loader.rs, context/*.
    • TinyAgents components: EmbeddingModel, Retriever, VectorStore, ScoredDoc, context events.
    • Acceptance: the agent harness can load retrieval context through a TinyAgents retriever facade while OpenHuman stores remain authoritative.
  • Move context compaction provenance into TinyAgents events.

    • OpenHuman files: context/README.md, tinyagents/summarize.rs, tinyagents/payload_summarizer.rs.
    • TinyAgents components: SummaryRecord, Compressed events, PromptCacheLayout, cache layout events.
    • Acceptance: every summary records source ids, before/after token estimates, policy version, and whether stable prompt prefix was preserved.
  • Normalize embedding usage/cost records.

    • OpenHuman files: embeddings, memory_sync, cost.
    • TinyAgents components: embedding usage fields, model catalog pricing.
    • Acceptance: embedding calls contribute usage and cost with provider/model, dimensions, vector count, and source.

Phase 10 - Dead Code And TinyAgents Re-Expression Audit

Do not delete these blindly. Treat this as an audit list for code that is dead, vestigial, or generic runtime behavior now expressible through TinyAgents harness/graph primitives. Delete only after call-site search, compatibility assessment, and migration coverage are complete.

  • Audit src/openhuman/agent/harness/engine/*.

    • Current role: surviving seams from the retired in-house turn loops: CheckpointStrategy, ProgressReporter, and TurnProgress.
    • TinyAgents expression: RunPolicy, AgentEvent, GraphEvent, EventSink, HarnessStatusStore, cap/stop middleware.
    • Candidate outcome: move max-iteration and progress projection into TinyAgents middleware/events, then delete engine/* if no product-specific compatibility seam remains.
  • Audit src/openhuman/agent/harness/agent_graph.rs and built-in agent_registry/agents/*/graph.rs default selectors.

    • Current role: per-agent graph hook, but most built-ins return AgentGraph::Default.
    • TinyAgents expression: a registry of CompiledGraph/graph factories keyed by agent id, with default graph supplied by the runtime and custom graphs registered only where they differ.
    • Candidate outcome: replace dozens of boilerplate default graph.rs files with registry defaults, keeping files only for agents with custom graphs.
  • Audit stale architecture references to removed in-house graph/loop code.

    • Current files: gitbooks/developing/architecture/agent-harness.md, src/openhuman/context/README.md.
    • Candidate stale names: src/openhuman/agent_graph/, GraphBlueprint, run_turn_engine, run_tool_call_loop, harness/tool_loop.rs, old context summarizer files.
    • TinyAgents expression: link to live src/openhuman/tinyagents/*, GraphBuilder, AgentHarness, ContextCompressionMiddleware, and graph-export/status surfaces.
    • Candidate outcome: move historical details into a short "pre-migration history" appendix or remove them from active architecture docs.
    • Partial: flagged the agent-harness.md agent_graph/GraphBlueprint section as HISTORICAL-removed (strong inline callout pointing at the live tinyagents surfaces); fixed context/README.md "Used by" line that still referenced the deleted reduce_before_call/ProviderSummarizer/ SegmentRecapSummarizer/unified_compaction_enabled. Sweep completed: every code doc-comment that described run_turn_engine/run_tool_call_loop/ tool_loop.rs as current behavior (≈20 sites across 15 files: tools, security, tokenjuice, triage, orchestration steering, task-local contexts, cron, host_runtime, event-bus example, test-file headers) now points at the live tinyagents surfaces; intentionally-historical "legacy X was removed / parity with" notes were kept. Domain READMEs (tools, tokenjuice, approval) fixed too — the tokenjuice one now records that compact_tool_output lost its only production caller with the retired loop (re-wiring it as an after_tool middleware; the stale default-Full wrapper is now deleted.
  • Audit src/openhuman/context/{pipeline,guard,microcompact}.rs.

    • Current role: context stats/session-memory bookkeeping plus older compaction concepts; live history reduction moved to ContextCompressionMiddleware and MessageTrimMiddleware.
    • TinyAgents expression: context-window middleware, Compressed events, PromptCacheLayout, cache layout events, SummaryRecord, usage/context pressure status.
    • Candidate outcome: keep only stats/session-memory state that remains OpenHuman-specific; move compression policy/provenance into TinyAgents middleware and delete unused reduction paths.
  • Audit src/openhuman/tinyagents/payload_summarizer.rs.

    • Current role: oversized tool-result compression via a summarizer sub-agent with a local circuit breaker.
    • TinyAgents expression: ToolMiddleware::after_tool, ContextCompressionMiddleware, SummaryRecord, tool artifact/result compaction events.
    • Candidate outcome: convert to middleware over TinyAgents tool results so the summarizer is no longer a separate OpenHuman-only hook.
  • Audit src/openhuman/agent_orchestration/running_subagents.rs.

    • Current role: bespoke live-task registry layered on TinyAgents InMemoryTaskStore, plus watch channels, abort handles, tombstones, task/session lookup, wait, steer, and cancel operations.
    • TinyAgents expression: TaskStore, SteeringCommand, SteeringHandle, CancellationToken, run tree/status store, durable OpenHuman ledger projections.
    • Candidate outcome: keep OpenHuman durable session/worker-thread records, but collapse transient lifecycle mechanics into TinyAgents task/status primitives once they can represent wait/steer/hard-abort needs.
  • Audit src/openhuman/agent_orchestration/tools/spawn_parallel_agents.rs after the graph-tool migration.

    • Current role: validation, preflight, fanout, worktree setup, event projection, overlap detection, result formatting.
    • TinyAgents expression: graph nodes (validate, dispatch, worker, collect, finalize) with Send fanout and reducers.
    • Candidate outcome: keep a thin tool wrapper that invokes the graph and formats the existing JSON payload; move orchestration mechanics into the graph module.
  • Audit hand-rolled join_all fanouts that are workflow orchestration, not simple IO batching.

    • Candidate files from current search: src/openhuman/mcp_registry/registry.rs, src/openhuman/learning/reflection.rs, src/openhuman/inference/local/service/ollama_admin/diagnostics.rs.
    • TinyAgents expression: only migrate fanouts that need agent/run lineage, checkpointing, policy, cancellation, or graph observability. Leave simple independent IO probes as ordinary Rust concurrency.
    • Candidate outcome: document why each fanout stays as join_all or move it to graph::parallel::map_reduce / graph Send.
  • Audit tool registry comments and docs that still describe retired direct-loop behavior.

    • Current files: src/openhuman/tools/traits.rs, src/openhuman/tools/README.md, src/openhuman/agent/harness/session/turn/tools.rs.
    • TinyAgents expression: ToolRegistry, ToolMiddleware, graph tool nodes, tool safety metadata, ToolExecutionContext.
    • Candidate outcome: update comments to describe the TinyAgents execution path and delete references to the retired serial harness::tool_loop dispatcher once no code path uses it.
    • Done: tools/traits.rs concurrency note and tools/README.md "Used by" now describe the tinyagents execution path (SharedToolAdapter / ToolPolicyMiddleware); session/turn/tools.rs had no stale references.

Phase 11 - Testing And Conformance

  • Create a focused parity test matrix for the current TinyAgents route.

    • OpenHuman files: src/openhuman/tinyagents/tests.rs, tests/agent_harness_e2e.rs, tests/agent_tool_loop_raw_coverage_e2e.rs.
    • TinyAgents components: AgentHarness, AgentEvent, ToolRegistry, ModelRegistry, MessageTrimMiddleware, ContextCompressionMiddleware.
    • Acceptance: chat turn, channel turn, sub-agent turn, unknown tool recovery, approval denial, streaming text, streaming tool args, reasoning deltas, early-exit pause, model-call cap, and cost footer all have tests on the TinyAgents path.
  • Add a TinyAgents adapter inventory test.

    • OpenHuman files: src/openhuman/tinyagents/mod.rs, src/openhuman/tinyagents/tests.rs.
    • Acceptance: one test asserts that the shared runner registers model, tools, middleware, event bridge, context compression, stop hooks, and unknown-tool policy in the intended order.
    • Done: harness assembly extracted from run_turn_via_tinyagents_shared into a testable assemble_turn_harness returning AssembledTurnHarness (harness + cursor/error-slot/halt-summary/outcome-sink/steering/early-exit seams). adapter_inventory_registers_model_tools_and_middleware asserts model registry, callable tools + unknown-tool policy, 9 lifecycle + 2 around-tool middlewares, steering handle, early-exit hook; adapter_inventory_gates_context_middleware_on_window proves the compression/trim gating. SDK gap: MiddlewareStack exposes lengths but not names, so exact ordering is documented at the registration sites and guarded by counts.
  • Port behavior clusters to TinyAgents testkit.

    • OpenHuman files: src/openhuman/agent/harness/*_tests.rs, tests/agent_*.
    • TinyAgents components: harness testkit, graph assert_graph, GraphEventRecorder, mock models/tools.
    • Acceptance: legacy assertions over loop wording are replaced by assertions over TinyAgents events, checkpoints, graph metadata, and final transcript.
  • Add cross-module e2e tests for graph/sub-agent/model/tool composition.

    • Candidate tests: workflow run with child sub-agents, delegation with review loop, council fanout, MCP tool call, Composio approval denial, memory retrieval plus summarization.
    • Acceptance: tests cover default features, openai-feature compatible code paths where available, and all-features where dependency constraints allow.
  • Add fuzz-style graph composition tests.

    • TinyAgents components: graph testkit, reducers, command routing, subgraph nodes, sub-agent fake nodes.
    • Acceptance: generated small graphs cover direct edges, conditional routes, Send fanout, joins, interrupts, recursion caps, and checkpoint resume.

Suggested Execution Order

  1. Confirm version/feature compatibility and SDK-extension gaps.
  2. Move tool safety/approval/security into TinyAgents middleware.
  3. Normalize usage/cost records and parent/child rollups.
  4. Persist TinyAgents event/status journals to the OpenHuman run ledger.
  5. Build the OpenHuman CapabilityRegistry projection.
  6. Convert one high-value built-in agent to a bespoke TinyAgents graph.
  7. Adapt memory/retrieval/context surfaces where the TinyAgents interfaces are a good fit.
  8. Audit and remove or re-express dead/vestigial runtime code through TinyAgents harness/graph primitives.
  9. Finish with parity, adapter-inventory, conformance, e2e, and fuzz-style tests.

Non-Goals For The First Migration Wave

  • Do not replace OpenHuman's durable run ledgers with TinyAgents SQLite storage until the one-time transcript/session migration, TinyAgents store/status adapter, and restart/resume parity tests are complete.
  • Do not expose TinyAgents' OpenAI provider directly to product code while OpenHuman provider config, credentials, OAuth, and billing classification are still the product source of truth.
  • Do not remove DomainEvent until all existing subscribers have a TinyAgents event/status replacement.