diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 00000000..958ea364 --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,78 @@ +# Changelog + +All notable changes to OpenJarvis are documented in this file. + +The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). + +--- + +## [Unreleased] + +### Added + +#### Skills System (Plans 1, 2A, 2B) + +- **Skills core** — every skill is a tool. Skills appear in a system prompt catalog, agents invoke them on demand, content (pipeline results, markdown instructions, or both) gets injected into context. + - `SkillManifest` + `SkillStep` types with tags, depends, invocation flags, markdown content + - `SkillManager` — discovery, precedence resolution, catalog XML generation, tool wrapping + - `SkillTool(BaseTool)` — auto-extracts parameters from step argument templates + - `SkillExecutor` — sequential pipeline execution with sub-skill delegation + - Dependency graph with cycle detection, max depth enforcement, capability unions + - Security: four trust tiers (bundled/indexed/unreviewed/workspace), capability-gated enforcement + - Skill index module for git-backed registry search + +- **agentskills.io spec adoption** — canonical `SKILL.md` format with YAML frontmatter following the [agentskills.io](https://agentskills.io/specification) open standard. + - `SkillParser` with strict spec validation + tolerant field mapping via `FIELD_MAPPING` table + - `ToolTranslator` for external tool name translation (Bash -> shell_exec, Read -> file_read, etc.) + - Source resolvers: `HermesResolver`, `OpenClawResolver`, `GitHubResolver` + - `SkillImporter` with provenance tracking (`.source` metadata files), optional script import + - Sourced subdirectory layout (`~/.openjarvis/skills///`) + +- **Skills learning loop** — trace tagging, pattern discovery, DSPy/GEPA optimization. + - Trace metadata tagging: `skill`, `skill_source`, `skill_kind` flow through ToolExecutor -> TraceCollector -> TraceStep + - `SkillDiscovery` wired into `SkillManager.discover_from_traces()` with kebab name normalization + - `SkillOptimizer` — per-skill DSPy/GEPA wrapper that buckets traces and writes sidecar overlays + - `SkillOverlay` — sidecar storage at `~/.openjarvis/learning/skills//optimized.toml` + - `SkillManager._load_overlays()` applies optimized descriptions + few-shot examples at discovery time + - `LearningOrchestrator._maybe_optimize_skills()` — opt-in auto-trigger + +- **Skills benchmark harness** — 4-condition PinchBench evaluation. + - I3 fix: `skill_few_shot_examples` wired through SystemBuilder -> `_run_agent` -> `ToolUsingAgent` -> `native_react.REACT_SYSTEM_PROMPT` + - `SkillBenchmarkRunner` — 4-condition x N-seed x M-task sweep with markdown report + - `JarvisAgentBackend` accepts `skills_enabled` and `overlay_dir` kwargs + - Conditions: `no_skills`, `skills_on`, `skills_optimized_dspy`, `skills_optimized_gepa` + +- **CLI commands:** + - `jarvis skill list` / `info` / `run` / `install` / `sync` / `sources` / `update` / `remove` / `search` + - `jarvis skill discover` — mine traces for recurring tool patterns + - `jarvis skill show-overlay` — inspect optimization output + - `jarvis optimize skills` — run DSPy/GEPA per-skill optimization + - `jarvis bench skills` — run the PinchBench skills benchmark + +- **Agent prompt improvement:** + - `native_react.REACT_SYSTEM_PROMPT` now includes "Using Skills" guidance that teaches agents to distinguish executable vs. instructional skill responses + - `{skill_examples}` placeholder for optimized few-shot example injection + +- **Configuration:** + - `[skills]` section: `enabled`, `skills_dir`, `active`, `auto_discover`, `auto_sync`, `max_depth`, `sandbox_dangerous` + - `[[skills.sources]]` section: `source`, `url`, `filter`, `auto_update` + - `[learning.skills]` section: `auto_optimize`, `optimizer`, `min_traces_per_skill`, `optimization_interval_seconds`, `overlay_dir` + - `SkillSourceConfig` and `SkillsLearningConfig` dataclasses + +- **Documentation:** + - `docs/user-guide/skills.md` — comprehensive user guide + - `docs/architecture/skills.md` — technical deep-dive + - `docs/tutorials/skills-workflow.md` — end-to-end tutorial + - `docs/getting-started/configuration.md` — expanded with skills config sections + - `CLAUDE.md` — updated architecture section + +### Fixed + +- **Trace metadata flow** — `ToolResult.metadata` now propagates through `TOOL_CALL_END` event to `TraceStep.metadata` (was silently dropped at the event-bus boundary) +- **TaintSet JSON serialization** — `ToolExecutor._json_safe_metadata()` filters non-JSON-serializable values (like `TaintSet`) from event payloads before they reach `TraceStore` +- **Non-dict YAML frontmatter** — source resolvers handle `yaml.safe_load()` returning a string instead of a dict (discovered on real OpenClaw imports) +- **OpenClaw category/name queries** — `jarvis skill install openclaw:owner/slug` now correctly splits into category + name match +- **SkillDiscovery trace compatibility** — `_extract_tool_sequence` reads from `step.input["tool"]` (the actual `TraceStep` format), not the nonexistent `step.tool_name` attribute +- **LearningOrchestrator skill trigger** — `_maybe_optimize_skills` runs BEFORE the SFT-data short-circuit (skills are tagged via trace metadata, not mined as SFT pairs) +- **PinchBenchScorer constructor** — `SkillBenchmarkRunner` constructs `PinchBenchScorer(judge_backend, model)` instead of no-args +- **EvalRunner results access** — reads per-task data from `eval_runner.results` property, not nonexistent `summary.results` diff --git a/README.md b/README.md index 257e5aba..1fda2d66 100644 --- a/README.md +++ b/README.md @@ -107,6 +107,29 @@ jarvis memory index ./docs/ # index your documents jarvis ask "Summarize all emails about Project X" ``` +### Skills + +Skills teach agents how to better use tools and improve their reasoning. Every skill is a tool — agents discover them from a catalog and invoke them on demand. + +```bash +# Install skills from public sources +jarvis skill install hermes:arxiv +jarvis skill sync hermes --category research + +# Use skills with any agent +jarvis ask "Use the code-explainer skill to explain this Python code: for i in range(5): print(i*2)" + +# Optimize skills from your trace history +jarvis optimize skills --policy dspy + +# Benchmark the impact +jarvis bench skills --max-samples 5 --seeds 42 +``` + +Import from [Hermes Agent](https://github.com/NousResearch/hermes-agent) (~150 skills), [OpenClaw](https://github.com/openclaw/skills) (~13,700 community skills), or any GitHub repo. Skills follow the [agentskills.io](https://agentskills.io/specification) open standard. + +See the [Skills User Guide](https://open-jarvis.github.io/OpenJarvis/user-guide/skills/) and [Skills Tutorial](https://open-jarvis.github.io/OpenJarvis/tutorials/skills-workflow/) for details. + ### Built-in Agents | Agent | Type | What it does | diff --git a/docs/architecture/skills.md b/docs/architecture/skills.md new file mode 100644 index 00000000..1a623146 --- /dev/null +++ b/docs/architecture/skills.md @@ -0,0 +1,222 @@ +--- +title: Skills Architecture +description: Technical deep-dive into the skills system design, components, and integration patterns +--- + +# Skills Architecture + +Skills are a **cross-cutting orchestration layer** that sits across the five existing primitives (Intelligence, Engine, Agents, Memory/Tools, Learning). They connect tools, agents, memory, and learning into reusable workflows without replacing or subsumming any primitive. + +## System Design + +``` + ┌──────────────────────┐ + │ SystemBuilder │ + │ .build() │ + └──────────┬───────────┘ + │ + ┌──────────▼───────────┐ + │ SkillManager │ + │ • discover() │ + │ • get_skill_tools()│ + │ • get_catalog_xml()│ + └──────────┬───────────┘ + │ + ┌────────────────┼────────────────┐ + │ │ │ + ┌─────────▼──────┐ ┌──────▼──────┐ ┌───────▼───────┐ + │ SkillTool │ │ Catalog │ │ Overlay │ + │ (BaseTool) │ │ XML │ │ Loader │ + │ → agent tools │ │ → sys. │ │ → optimized │ + │ list │ │ prompt │ │ desc + │ + │ │ │ │ │ few-shot │ + └────────────────┘ └─────────────┘ └───────────────┘ +``` + +## Key Components + +### SkillManifest (`skills/types.py`) + +The canonical data structure for a loaded skill: + +```python +@dataclass(slots=True) +class SkillManifest: + name: str + version: str = "0.1.0" + description: str = "" + author: str = "" + steps: List[SkillStep] = field(default_factory=list) + required_capabilities: List[str] = field(default_factory=list) + signature: str = "" + metadata: Dict[str, Any] = field(default_factory=dict) + tags: List[str] = field(default_factory=list) + depends: List[str] = field(default_factory=list) + user_invocable: bool = True + disable_model_invocation: bool = False + markdown_content: str = "" +``` + +### SkillManager (`skills/manager.py`) + +The central coordinator. Created by `SystemBuilder.build()` during system composition. + +**Lifecycle:** +1. `discover(paths)` — scans skill directories in precedence order, loads manifests, validates the dependency graph, applies optimization overlays +2. `get_skill_tools()` — wraps each discovered skill as a `SkillTool(BaseTool)`, wires sub-skill resolver callbacks +3. `get_catalog_xml()` — generates the lightweight `` XML for system prompt injection +4. `get_few_shot_examples()` — returns formatted few-shot strings from optimization overlays + +### SkillTool (`skills/tool_adapter.py`) + +Adapter that makes any skill look like a regular `BaseTool` to agents: + +- `spec` property derives `ToolSpec` from the manifest — auto-extracts input parameters from step argument templates +- `execute(**params)` runs the pipeline (if steps exist), returns markdown content (if SKILL.md exists), or both +- `_build_result_metadata()` tags every invocation with `skill`, `skill_source`, `skill_kind` for downstream trace analysis + +### SkillParser (`skills/parser.py`) + +Two-pass parser for agentskills.io-compatible SKILL.md frontmatter: + +1. **Strict pass** — validates required fields (`name`, `description`), length limits, kebab-case naming rules +2. **Tolerant pass** — maps non-spec vendor fields to canonical locations via `FIELD_MAPPING` table. Unmapped fields are logged and preserved in `metadata.openjarvis.original_frontmatter` + +The mapping table is data, not code paths. Adding support for new vendor fields means adding entries — no logic changes. + +### SkillExecutor (`skills/executor.py`) + +Sequential pipeline executor: + +- Steps with `tool_name` → delegate to `ToolExecutor.execute()` +- Steps with `skill_name` → delegate to a resolver callback (set by SkillManager) +- Template rendering: `{placeholder}` syntax resolved from a shared context dict +- `output_key` stores each step's result for downstream steps +- Publishes `SKILL_EXECUTE_START` / `SKILL_EXECUTE_END` events on the EventBus + +### Source Resolvers (`skills/sources/`) + +One resolver per import source, all implementing `SourceResolver` ABC: + +| Resolver | Repo layout | Special handling | +|----------|-------------|------------------| +| `HermesResolver` | `skills///` | Skips `DESCRIPTION.md`, reads Hermes vendor metadata | +| `OpenClawResolver` | `skills///` | Reads `_meta.json` sidecars | +| `GitHubResolver` | Recursive walk for `SKILL.md` | Generic — accepts any repo URL | + +### SkillImporter (`skills/importer.py`) + +Takes a `ResolvedSkill` from a source resolver and installs it on disk: + +1. Parse source SKILL.md through `SkillParser` +2. Translate tool references (`Bash` → `shell_exec`, `Read` → `file_read`, etc.) +3. Compatibility check (platform, missing tools) +4. Copy SKILL.md + references/assets/templates (scripts gated by `--with-scripts`) +5. Write `.source` provenance file with commit SHA, translated tools, timestamps + +### SkillOverlay (`skills/overlay.py`) + +Sidecar storage for optimization output at `~/.openjarvis/learning/skills//optimized.toml`: + +```toml +[optimized] +skill_name = "research-and-summarize" +optimizer = "dspy" +optimized_at = "2026-04-08T14:30:00Z" +trace_count = 47 +description = "An optimized description" + +[[optimized.few_shot]] +input = "transformer attention mechanisms" +output = "## Recent Advances..." +``` + +The overlay is the **contract** between the optimizer and the SkillManager. Both sides agree on the schema; either can be swapped independently. + +### SkillOptimizer (`learning/agents/skill_optimizer.py`) + +Per-skill wrapper around DSPy/GEPA: + +1. Buckets traces by `metadata.skill` (from the C1 trace tagging) +2. Skips skills below `min_traces_per_skill` threshold +3. Calls `_run_dspy()` or `_run_gepa()` per qualifying skill +4. Writes overlay TOML files + +## Integration Points + +### SystemBuilder Wiring + +`SystemBuilder.build()` handles skill integration: + +```python +# 1. Create SkillManager +skill_manager = SkillManager(bus, capability_policy=...) + +# 2. Discover skills from disk +skill_manager.discover(paths=[workspace_skills, user_skills]) + +# 3. Wrap as tools and merge into tool list +skill_tools = skill_manager.get_skill_tools(tool_executor=...) +tool_list.extend(skill_tools) + +# 4. Capture few-shot examples for agents +system._skill_few_shot_examples = skill_manager.get_few_shot_examples() +``` + +### Trace Metadata Flow + +When an agent invokes a `SkillTool`: + +``` +SkillTool.execute() + → ToolResult(metadata={"skill": name, "skill_source": src, "skill_kind": kind}) + → ToolExecutor._json_safe_metadata() filters non-serializable values + → TOOL_CALL_END event with metadata + → TraceCollector._on_tool_end() → TraceStep(metadata=...) + → TraceStore saves to SQLite (metadata as JSON) + → SkillOptimizer._bucket_traces_by_skill() reads metadata.skill +``` + +### Agent Few-Shot Injection + +Optimized few-shot examples flow through: + +``` +SkillManager.get_few_shot_examples() + → system._skill_few_shot_examples (stashed on JarvisSystem) + → _run_agent() → agent_kwargs["skill_few_shot_examples"] + → ToolUsingAgent._skill_few_shot_examples + → native_react.run() → REACT_SYSTEM_PROMPT.format(skill_examples=...) +``` + +## Dependency Graph + +Skills can compose other skills. At discovery time, SkillManager validates: + +1. **Cycle detection** — Kahn's algorithm for topological sort +2. **Max depth enforcement** — configurable (default 5) +3. **Capability union** — parent must declare all transitive child capabilities + +## File Layout + +``` +src/openjarvis/skills/ +├── __init__.py # Public exports +├── types.py # SkillManifest, SkillStep +├── manager.py # SkillManager +├── executor.py # SkillExecutor + sub-skill delegation +├── loader.py # TOML + Markdown + directory loading +├── tool_adapter.py # SkillTool(BaseTool) wrapper +├── parser.py # Strict + tolerant agentskills.io parser +├── tool_translator.py # External tool name translation +├── importer.py # Install from resolved sources +├── overlay.py # Optimization sidecar storage +├── dependency.py # Graph validation +├── security.py # Trust tiers, capability validation +├── index.py # Git-backed skill index +└── sources/ + ├── base.py # SourceResolver ABC + ├── hermes.py # HermesResolver + ├── openclaw.py # OpenClawResolver + └── github.py # GitHubResolver +``` diff --git a/docs/getting-started/configuration.md b/docs/getting-started/configuration.md index 19083b35..8e1aa36e 100644 --- a/docs/getting-started/configuration.md +++ b/docs/getting-started/configuration.md @@ -432,6 +432,81 @@ db_path = "~/.openjarvis/traces.db" --- +### `[skills]` — Skills System + +Controls the skills system — reusable compositions of tools and agent instructions. Skills teach agents how to better use tools and improve their reasoning. See the [Skills User Guide](../user-guide/skills.md) for full documentation. + +```toml +[skills] +enabled = true +skills_dir = "~/.openjarvis/skills/" +active = "*" +auto_discover = true +auto_sync = false +max_depth = 5 +sandbox_dangerous = true +``` + +| Field | Type | Default | Description | +|-------|------|---------|-------------| +| `enabled` | bool | `true` | Whether to enable the skills system. When disabled, no skills are loaded or exposed to agents. | +| `skills_dir` | string | `~/.openjarvis/skills/` | Directory where skills are installed. | +| `active` | string | `"*"` | Comma-separated list of skill names to activate, or `"*"` for all discovered skills. | +| `auto_discover` | bool | `true` | Whether to scan `skills_dir` for skills on startup. | +| `auto_sync` | bool | `false` | Whether to pull from configured sources on session start (checks freshness every 24h). | +| `max_depth` | int | `5` | Maximum sub-skill nesting depth for composed skills. | +| `sandbox_dangerous` | bool | `true` | Whether to warn about skills with dangerous capabilities (`shell:execute`, `network:listen`, `filesystem:write`). | + +#### `[[skills.sources]]` — Skill Import Sources + +Configure one or more skill sources for automatic import. Each `[[skills.sources]]` entry defines a source to pull from. + +```toml +[[skills.sources]] +source = "hermes" +filter = { category = ["research", "coding", "productivity"] } +auto_update = true + +[[skills.sources]] +source = "openclaw" +filter = { search = "web3|crypto" } + +[[skills.sources]] +source = "github" +url = "https://github.com/myorg/internal-skills" +auto_update = true +``` + +| Field | Type | Default | Description | +|-------|------|---------|-------------| +| `source` | string | `""` | Source type: `"hermes"`, `"openclaw"`, or `"github"`. | +| `url` | string | `""` | Repository URL. Required when `source = "github"`. | +| `filter` | table | `{}` | Filter criteria. Supported keys: `category` (list of strings), `search` (regex string). | +| `auto_update` | bool | `false` | Whether to pull latest commits when syncing this source. | + +#### `[learning.skills]` — Skills Learning Loop + +Controls the automatic optimization of skill descriptions and few-shot examples from trace data. Requires `[traces] enabled = true` to collect the traces that the optimizer analyzes. + +```toml +[learning.skills] +auto_optimize = false +optimizer = "dspy" +min_traces_per_skill = 20 +optimization_interval_seconds = 86400 +overlay_dir = "~/.openjarvis/learning/skills/" +``` + +| Field | Type | Default | Description | +|-------|------|---------|-------------| +| `auto_optimize` | bool | `false` | Whether to run skill optimization automatically after each learning cycle. | +| `optimizer` | string | `"dspy"` | Optimization policy: `"dspy"` (bootstrap few-shot) or `"gepa"` (evolutionary). | +| `min_traces_per_skill` | int | `20` | Minimum trace count for a skill to be eligible for optimization. | +| `optimization_interval_seconds` | int | `86400` | Run optimization at most once per this interval (default: once per day). | +| `overlay_dir` | string | `~/.openjarvis/learning/skills/` | Where optimized skill overlays are stored. | + +--- + ### `[channel]` — Channel Messaging Controls the channel messaging bridge for multi-platform communication. Each supported platform has its own nested sub-section. diff --git a/docs/tutorials/index.md b/docs/tutorials/index.md index 192b7cba..da0e000b 100644 --- a/docs/tutorials/index.md +++ b/docs/tutorials/index.md @@ -44,6 +44,14 @@ Hands-on guides that walk through building real applications with OpenJarvis. Ea [:octicons-arrow-right-24: Get started](code-companion.md) +- :material-puzzle:{ .lg .middle } **Skills Workflow** + + --- + + Install skills from Hermes Agent, use them with a local agent, discover patterns from traces, optimize with DSPy, and benchmark the impact — the complete skills lifecycle. + + [:octicons-arrow-right-24: Get started](skills-workflow.md) + ## What You Will Learn diff --git a/docs/tutorials/skills-workflow.md b/docs/tutorials/skills-workflow.md new file mode 100644 index 00000000..77fd3251 --- /dev/null +++ b/docs/tutorials/skills-workflow.md @@ -0,0 +1,236 @@ +--- +title: Skills Workflow +description: End-to-end tutorial — install skills, use them with an agent, discover patterns from traces, and optimize with DSPy +--- + +# Skills Workflow Tutorial + +This tutorial walks through the complete skills lifecycle: installing skills from public sources, using them with a local agent, discovering patterns from trace history, and optimizing skill descriptions with DSPy. By the end you will have a working skills setup that improves over time. + +!!! note "Before you begin" + This tutorial assumes OpenJarvis is installed with Ollama running and a model available (e.g., `qwen3.5:9b`). If you have not completed setup yet, start with the [Quick Start guide](../getting-started/quickstart.md). + +## Step 1: Install Skills from Hermes Agent + +OpenJarvis can import skills from the [Hermes Agent](https://github.com/NousResearch/hermes-agent) skill library maintained by NousResearch. Let's install a few useful ones. + +```bash +# Install individual skills +jarvis skill install hermes:arxiv +jarvis skill install hermes:github-pr-workflow + +# Or bulk install an entire category +jarvis skill sync hermes --category research +``` + +The first install clones the Hermes repo to `~/.openjarvis/skill-cache/hermes/` (one-time, ~5s). Subsequent installs reuse the cache. + +Verify what's installed: + +```bash +jarvis skill list +``` + +You should see a table with each skill's name, description, version, and tags. + +## Step 2: Inspect an Installed Skill + +Let's look at what the `arxiv` skill contains: + +```bash +jarvis skill info arxiv +``` + +This shows the skill's metadata — author, description, tags, capabilities, whether it has structured steps or markdown instructions, and its invocation flags. + +You can also inspect the raw SKILL.md: + +```bash +cat ~/.openjarvis/skills/hermes/arxiv/SKILL.md | head -40 +``` + +The `.source` file records provenance: + +```bash +cat ~/.openjarvis/skills/hermes/arxiv/.source +``` + +This shows the source (`hermes:arxiv`), the git commit it was imported from, which tool names were translated (e.g., `Edit→file_edit`), and the install timestamp. + +## Step 3: Use Skills with an Agent + +Now let's ask the agent a question that should trigger skill usage: + +```bash +jarvis ask "Use the code-explainer skill to explain this Python code: for i in range(5): print(i*2)" \ + --engine ollama --model qwen3.5:9b +``` + +The agent will: +1. See the skill catalog in its system prompt +2. Decide to invoke `skill_code-explainer` +3. Receive the markdown instructions from the skill +4. Follow the 5-step pattern to explain the code + +Try a pipeline skill too: + +```bash +jarvis ask "Use the math-solver skill to compute 17 * 23" \ + --engine ollama --model qwen3.5:9b +``` + +This time the agent invokes `skill_math-solver`, which executes a deterministic pipeline (calling the `calculator` tool internally) and returns the computed result directly. + +## Step 4: Create Your Own Skill + +Create a new skill directory: + +```bash +mkdir -p ~/.openjarvis/skills/my-reviewer +``` + +Write a SKILL.md: + +```bash +cat > ~/.openjarvis/skills/my-reviewer/SKILL.md << 'EOF' +--- +name: my-reviewer +description: Review code changes with a security-first approach +license: MIT +metadata: + openjarvis: + version: "0.1.0" + author: me + tags: [coding, review, security] +--- + +When asked to review code, follow this approach: + +1. **Security scan first** — check for injection vulnerabilities, hardcoded secrets, unsafe deserialization +2. **Correctness** — verify logic, edge cases, error handling +3. **Style** — naming, structure, consistency with surrounding code +4. **Summary** — one paragraph with the verdict: approve, request changes, or block + +Always start with security. If you find a security issue, flag it as BLOCKING regardless of other concerns. +EOF +``` + +Verify it's discovered: + +```bash +jarvis skill list +``` + +You should see `my-reviewer` in the table. Try it: + +```bash +jarvis ask "Use the my-reviewer skill to review this function: def login(user, pwd): return db.query(f'SELECT * FROM users WHERE name={user} AND pass={pwd}')" \ + --engine ollama --model qwen3.5:9b +``` + +The agent should follow the security-first approach and flag the SQL injection vulnerability. + +## Step 5: Generate Traces + +For the learning loop to work, we need traces. Run several queries that use skills: + +```bash +# Generate a few traces +jarvis ask "Use math-solver to compute 100 / 7" +jarvis ask "Use code-explainer to explain: lambda x: x**2" +jarvis ask "Use my-reviewer to review: def add(a,b): return a+b" +jarvis ask "Use math-solver to compute 2**10" +jarvis ask "Use code-explainer to explain: [x for x in range(10) if x % 2 == 0]" +``` + +Each query produces a trace in `~/.openjarvis/traces.db` with skill metadata tags (`skill`, `skill_source`, `skill_kind`). + +## Step 6: Discover Patterns from Traces + +Mine the trace store for recurring tool sequences: + +```bash +# Preview without writing +jarvis skill discover --dry-run --min-frequency 2 + +# Write discovered patterns as skill manifests +jarvis skill discover --min-frequency 2 +``` + +Discovered skills land in `~/.openjarvis/skills/discovered/` and automatically appear in `jarvis skill list` on the next session. + +## Step 7: Optimize Skills with DSPy + +Once you have enough traces (at least 3-5 per skill), run the optimizer: + +```bash +# Preview what would be optimized +jarvis optimize skills --dry-run + +# Run DSPy optimization +jarvis optimize skills --policy dspy --min-traces 3 +``` + +This produces overlay files at `~/.openjarvis/learning/skills//optimized.toml` with improved descriptions and few-shot examples extracted from your best traces. + +Inspect what was produced: + +```bash +jarvis skill show-overlay math-solver +jarvis skill show-overlay code-explainer +``` + +The next time you run a query, the agent sees the optimized descriptions and few-shot examples in its system prompt. + +## Step 8: Benchmark the Impact + +Run a quick benchmark to see if skills + optimization actually help: + +```bash +# Smoke test: 4 conditions × 1 seed × 5 tasks +jarvis bench skills --max-samples 5 --seeds 42 +``` + +This runs the PinchBench benchmark in four conditions (no skills, skills on, DSPy-optimized, GEPA-optimized) and produces a markdown report at `docs/superpowers/results/`. + +## Step 9: Configure Auto-Import and Auto-Optimization + +For a hands-off experience, add this to `~/.openjarvis/config.toml`: + +```toml +[skills] +enabled = true +auto_sync = true + +[[skills.sources]] +source = "hermes" +filter = { category = ["research", "coding"] } +auto_update = true + +[learning.skills] +auto_optimize = true +optimizer = "dspy" +min_traces_per_skill = 20 +``` + +Now skills are automatically synced from Hermes on session start, and the optimizer runs after each learning cycle when enough traces accumulate. + +## What You Learned + +| Concept | What you did | +|---------|-------------| +| **Installing skills** | `jarvis skill install hermes:arxiv` — imported from public sources | +| **Using skills** | `jarvis ask "Use the code-explainer skill..."` — agent invokes skills as tools | +| **Creating skills** | Wrote a `SKILL.md` with YAML frontmatter and markdown instructions | +| **Generating traces** | Ran skill-using queries to populate the trace store | +| **Discovering patterns** | `jarvis skill discover` — mined traces for recurring tool sequences | +| **Optimizing skills** | `jarvis optimize skills --policy dspy` — improved descriptions + few-shot examples | +| **Benchmarking** | `jarvis bench skills` — measured the impact across 4 conditions | +| **Auto configuration** | Added `[skills]` and `[learning.skills]` config sections | + +## Next Steps + +- Browse the [full skills user guide](../user-guide/skills.md) for all CLI commands and configuration options +- Read the [skills architecture](../architecture/skills.md) for the technical deep-dive +- Explore the [Hermes Agent skill library](https://github.com/NousResearch/hermes-agent/tree/main/skills) for more skills to install +- Try [OpenClaw skills](https://github.com/openclaw/skills) for community-contributed skills diff --git a/docs/user-guide/skills.md b/docs/user-guide/skills.md new file mode 100644 index 00000000..05717262 --- /dev/null +++ b/docs/user-guide/skills.md @@ -0,0 +1,405 @@ +--- +title: Skills +description: Reusable compositions of tools and agent instructions — discover, import, optimize, and share +search.boost: 2.0 +--- + +# Skills + +Skills teach agents **how to better use tools and improve their reasoning**. They are reusable compositions of tools, sub-skills, and agent instructions that can be shared via public registries. + +Every skill is a tool. Skills appear in a lightweight catalog in the agent's system prompt, and when the agent invokes one, its content (pipeline results, markdown instructions, or both) gets injected into the conversation context. + +## Overview + +| Concept | Description | +|---------|-------------| +| **Skill** | A directory containing `skill.toml` (structured pipeline), `SKILL.md` (markdown instructions), or both | +| **SkillManager** | Central coordinator for discovery, resolution, catalog generation, and tool wrapping | +| **SkillTool** | Adapter that wraps any skill as a `BaseTool` so agents can invoke it | +| **Overlay** | Sidecar file at `~/.openjarvis/learning/skills/` storing optimized descriptions and few-shot examples | +| **Source** | A resolver for importing skills from Hermes Agent, OpenClaw, or any GitHub repo | + +## Quick Start + +```bash +# List installed skills +jarvis skill list + +# Install a skill from Hermes Agent +jarvis skill install hermes:apple-notes + +# Bulk install a category +jarvis skill sync hermes --category research + +# Run a skill directly +jarvis skill run math-solver -a expression="41 + 82" + +# See skill details +jarvis skill info research-and-summarize +``` + +## Skill Definition Format + +A skill is a directory containing a `skill.toml`, a `SKILL.md`, or both. + +### Directory Structure + +``` +research-and-summarize/ +├── SKILL.md # Markdown instructions (loaded on invocation) +├── skill.toml # Structured pipeline steps +├── templates/ # Optional Jinja2 templates +├── scripts/ # Optional executable helpers +├── references/ # Optional detailed docs +├── assets/ # Optional static resources +└── examples/ # Optional usage examples +``` + +### skill.toml (Structured Pipeline) + +Pipeline skills define a sequence of tool calls that execute deterministically: + +```toml +[skill] +name = "research-and-summarize" +version = "0.1.0" +description = "Search the web and produce a structured summary" +author = "openjarvis" +tags = ["research", "summarization"] +required_capabilities = ["network:fetch"] +depends = ["summarize"] + +[[skill.steps]] +tool_name = "web_search" +arguments_template = '{"query": "{query}"}' +output_key = "search_results" + +[[skill.steps]] +skill_name = "summarize" +arguments_template = '{"text": "{search_results}"}' +output_key = "summary" +``` + +Steps can call tools (`tool_name`) or other skills (`skill_name`). Template placeholders like `{query}` become the skill's input parameters. Output keys chain between steps. + +### SKILL.md (Instructional Content) + +Instructional skills provide markdown guidance that agents follow using their other tools: + +```markdown +--- +name: code-explainer +description: Explain code in plain language with examples +license: MIT +metadata: + openjarvis: + version: "0.1.0" + author: openjarvis + tags: [coding, explanation] +--- + +When asked to explain code, follow this approach: + +1. Identify the programming language +2. Break the code into logical sections +3. Explain each section in plain language +4. Highlight any patterns, idioms, or potential issues +5. Provide a one-sentence summary at the end +``` + +The YAML frontmatter follows the [agentskills.io](https://agentskills.io/specification) open standard. Required fields: `name`, `description`. Optional: `license`, `compatibility`, `metadata`, `allowed-tools`. + +### What Happens on Invocation + +| Skill has | On invocation | +|-----------|---------------| +| `skill.toml` steps only | Execute the pipeline, return results | +| `SKILL.md` only | Return the markdown instructions — agent follows them in subsequent turns | +| Both | Execute pipeline steps AND return the markdown guidance alongside results | + +## Installing Skills + +### From Hermes Agent + +```bash +# Single skill +jarvis skill install hermes:apple-notes + +# Bulk install by category +jarvis skill sync hermes --category research +jarvis skill sync hermes --category coding +jarvis skill sync hermes # everything (~150 skills) +``` + +### From OpenClaw + +```bash +# Single skill (owner/slug format) +jarvis skill install openclaw:0xv4l3nt1n3/etherscan + +# Bulk install with search filter +jarvis skill sync openclaw --search "web3|crypto" +``` + +### From Any GitHub Repo + +```bash +jarvis skill install github:user/repo/path/to/skill --url https://github.com/user/repo +``` + +### Config-Driven Auto Import + +Add sources to `~/.openjarvis/config.toml` for automatic syncing: + +```toml +[skills] +enabled = true +auto_sync = true + +[[skills.sources]] +source = "hermes" +filter = { category = ["research", "coding", "productivity"] } +auto_update = true + +[[skills.sources]] +source = "openclaw" +filter = { search = "web3|crypto" } +``` + +When `auto_sync = true`, the SkillManager checks source freshness on each session start and pulls updates in the background. + +### Managing Sources + +```bash +# List configured sources +jarvis skill sources + +# Update all configured sources +jarvis skill update +``` + +## How Agents Use Skills + +### Skill Catalog in the System Prompt + +All available skills appear as a lightweight XML catalog in the agent's system prompt: + +```xml + + + + + +``` + +The agent reads this catalog and decides when to invoke a skill based on the user's request. + +### Invocation Control + +Per-skill flags control visibility: + +```toml +[skill] +user_invocable = true # expose as CLI command (default: true) +disable_model_invocation = false # hide from agent catalog (default: false) +``` + +| `user_invocable` | `disable_model_invocation` | CLI command? | Agent discovers? | +|---|---|---|---| +| true (default) | false (default) | Yes | Yes | +| true | true | Yes | No | +| false | false | No | Yes | +| false | true | No | No (dormant) | + +### Pipeline vs. Instructional Skills + +Agents handle both skill types correctly: + +- **Pipeline skills** (with `skill.toml` steps) execute deterministically and return computed results. The agent uses the result directly in its answer. +- **Instructional skills** (with `SKILL.md` only) return markdown text describing HOW to accomplish a task. The agent reads the instructions and follows them using its other tools (web_search, shell_exec, calculator, etc.). + +## Skill Discovery from Traces + +OpenJarvis can automatically mine your trace history for recurring tool sequences and surface them as candidate skills: + +```bash +# Preview discovered patterns without writing +jarvis skill discover --dry-run --min-frequency 3 + +# Write discovered skills to ~/.openjarvis/skills/discovered/ +jarvis skill discover +``` + +Discovered skills land in `~/.openjarvis/skills/discovered/` and automatically appear in `jarvis skill list` on the next session. + +## Skill Optimization + +### Optimizing with DSPy or GEPA + +The skills learning loop uses your trace history to optimize skill descriptions and extract few-shot examples: + +```bash +# Preview what would be optimized +jarvis optimize skills --dry-run + +# Run DSPy optimization +jarvis optimize skills --policy dspy --min-traces 3 + +# Run GEPA evolutionary optimization +jarvis optimize skills --policy gepa --min-traces 3 + +# Inspect what optimization produced +jarvis skill show-overlay research-and-summarize +``` + +Optimization results are stored as sidecar overlays at `~/.openjarvis/learning/skills//optimized.toml`. They override the skill's description and add few-shot examples to the agent's system prompt. The original skill files are never modified. + +### Auto-Optimization + +Enable automatic optimization in config: + +```toml +[learning.skills] +auto_optimize = false # set to true to enable +optimizer = "dspy" # "dspy" or "gepa" +min_traces_per_skill = 20 +``` + +When enabled, the `LearningOrchestrator` runs skill optimization after each learning cycle. + +## Benchmarking Skills + +Measure whether skills improve agent performance: + +```bash +# Full sweep: 4 conditions × 3 seeds +jarvis bench skills + +# Smoke test: 4 conditions × 1 seed × 5 tasks +jarvis bench skills --max-samples 5 --seeds 42 + +# Single condition +jarvis bench skills --condition skills_optimized_dspy +``` + +The four benchmark conditions are: + +| Condition | What it tests | +|---|---| +| `no_skills` | Skills disabled (control) | +| `skills_on` | Skills enabled, no optimization | +| `skills_optimized_dspy` | DSPy-optimized overlays | +| `skills_optimized_gepa` | GEPA-optimized overlays | + +Results are written to `docs/superpowers/results/pinchbench-skills-eval-{date}.md` with a summary table, per-task breakdown, deltas, and skill invocation counts. + +## Security & Trust + +### Trust Tiers + +| Tier | Source | Verification | Runtime | +|------|--------|-------------|---------| +| **Bundled** | Ships with OpenJarvis | Implicit trust | Full access within declared capabilities | +| **Indexed** | In official skill index, signed | SHA256 + Ed25519 | Capability-gated | +| **Unreviewed** | Arbitrary GitHub URL | SHA256 only | Capability-gated + sandbox warning | +| **Workspace** | Local `./skills/` directory | None (user code) | Trusted | + +### Capability Enforcement + +Skills declare required capabilities. At runtime, the SkillExecutor checks that each tool call falls within the skill's declared capabilities: + +- `network:fetch` — outbound HTTP requests +- `filesystem:read` / `filesystem:write` — file access +- `shell:execute` — run shell commands (dangerous) +- `memory:read` / `memory:write` — memory backend access +- `engine:inference` — LLM calls + +Skills declaring dangerous capabilities (`shell:execute`, `network:listen`, `filesystem:write`) trigger install-time warnings and sandbox recommendations. + +### Scripts + +Imported skills may include `scripts/` directories with executable code. These are **skipped by default** for security. Use `--with-scripts` to opt in: + +```bash +jarvis skill install hermes:arxiv --with-scripts +``` + +## Skill Composition + +Skills can invoke other skills as sub-steps: + +```toml +[[skill.steps]] +skill_name = "summarize" +arguments_template = '{"text": "{search_results}"}' +output_key = "summary" +``` + +The SkillManager builds a dependency graph at discovery time and validates: + +1. **No cycles** — `A → B → C → A` is rejected with a clear error +2. **Max depth** — default 5 levels (configurable) +3. **Capability unions** — parent must declare all capabilities its children need + +## Configuration Reference + +### `[skills]` Section + +```toml +[skills] +enabled = true # enable/disable the skill system +skills_dir = "~/.openjarvis/skills/" # where skills are installed +active = "*" # which skills to activate ("*" = all) +auto_discover = true # scan skills_dir on startup +auto_sync = false # pull from configured sources on startup +max_depth = 5 # max sub-skill nesting depth +sandbox_dangerous = true # warn about dangerous capabilities +``` + +### `[[skills.sources]]` Section + +```toml +[[skills.sources]] +source = "hermes" # "hermes", "openclaw", or "github" +url = "" # required when source = "github" +filter = { category = ["research", "coding"] } +auto_update = true # pull latest on sync +``` + +### `[learning.skills]` Section + +```toml +[learning.skills] +auto_optimize = false # opt-in automatic optimization +optimizer = "dspy" # "dspy" or "gepa" +min_traces_per_skill = 20 # minimum traces before optimizing +optimization_interval_seconds = 86400 # at most once per day +overlay_dir = "~/.openjarvis/learning/skills/" +``` + +## Name Precedence + +When the same skill name exists in multiple locations, closest scope wins: + +1. **Workspace** `./skills/` (highest priority) +2. **User** `~/.openjarvis/skills/` +3. **Bundled** (shipped with OpenJarvis) + +## CLI Reference + +| Command | Description | +|---------|-------------| +| `jarvis skill list` | List installed skills | +| `jarvis skill info ` | Show detailed skill information | +| `jarvis skill run [-a key=value]` | Execute a skill directly | +| `jarvis skill install :` | Install from Hermes, OpenClaw, or GitHub | +| `jarvis skill sync [] [--category C]` | Bulk install + update from sources | +| `jarvis skill sources` | List configured skill sources | +| `jarvis skill update` | Pull latest from configured sources | +| `jarvis skill remove ` | Remove an installed skill | +| `jarvis skill search ` | Search the skill index | +| `jarvis skill discover [--dry-run]` | Mine traces for recurring tool patterns | +| `jarvis skill show-overlay ` | Inspect optimization output for a skill | +| `jarvis optimize skills [--policy dspy\|gepa]` | Optimize skill descriptions + few-shot examples | +| `jarvis bench skills [--condition C]` | Run the PinchBench skills benchmark | diff --git a/src/openjarvis/agents/_stubs.py b/src/openjarvis/agents/_stubs.py index d462994c..f65854e5 100644 --- a/src/openjarvis/agents/_stubs.py +++ b/src/openjarvis/agents/_stubs.py @@ -306,6 +306,7 @@ class ToolUsingAgent(BaseAgent): agent_id: Optional[str] = None, interactive: bool = False, confirm_callback: Optional[Any] = None, + skill_few_shot_examples: Optional[List[str]] = None, ) -> None: super().__init__( engine, @@ -317,6 +318,9 @@ class ToolUsingAgent(BaseAgent): from openjarvis.tools._stubs import ToolExecutor self._tools = tools or [] + # Plan 2B I3: store optimized few-shot examples for agents to inject + # into their own system prompt templates as appropriate. + self._skill_few_shot_examples = list(skill_few_shot_examples or []) _aid = agent_id or getattr(self, "agent_id", "") self._executor = ToolExecutor( self._tools, diff --git a/src/openjarvis/agents/native_react.py b/src/openjarvis/agents/native_react.py index 2c3cf4cb..89061505 100644 --- a/src/openjarvis/agents/native_react.py +++ b/src/openjarvis/agents/native_react.py @@ -28,7 +28,25 @@ Action Input: Thought: Final Answer: -{tool_descriptions}""" +# Using Skills + +Tools whose names begin with `skill_` are SKILLS. When you call a skill tool, +the response can take one of two forms: + +- **Computed result**: The skill ran a deterministic pipeline and returned a + value (number, string, JSON, etc.). Use the value directly in your answer. + +- **Procedural instructions**: The skill returned markdown text describing + HOW to accomplish a task. Recognize this when the response starts with + `#` headings, contains bullet lists, or uses phrases like "When asked + to...", "First...", "Steps:". When you receive instructions: + 1. READ the instructions carefully — they are your playbook + 2. FOLLOW the steps using your OTHER tools (e.g. calculator, web_search, + shell_exec, file_read) — not the same skill + 3. DO NOT call the same skill again — you already have its instructions + 4. Synthesize a Final Answer from what you learned + +{skill_examples}{tool_descriptions}""" @AgentRegistry.register("native_react") @@ -52,6 +70,7 @@ class NativeReActAgent(ToolUsingAgent): max_tokens: Optional[int] = None, interactive: bool = False, confirm_callback=None, + skill_few_shot_examples: Optional[List[str]] = None, ) -> None: super().__init__( engine, @@ -63,6 +82,7 @@ class NativeReActAgent(ToolUsingAgent): max_tokens=max_tokens, interactive=interactive, confirm_callback=confirm_callback, + skill_few_shot_examples=skill_few_shot_examples, ) def _parse_response(self, text: str) -> dict: @@ -111,7 +131,22 @@ class NativeReActAgent(ToolUsingAgent): # Build system prompt with rich tool descriptions tool_desc = build_tool_descriptions(self._tools) - system_prompt = REACT_SYSTEM_PROMPT.format(tool_descriptions=tool_desc) + # Plan 2B I3: render the optimized few-shot examples as a section + # that the model sees BEFORE the tool descriptions. When no + # examples are present, this is an empty string and the rendered + # prompt is unchanged. + if self._skill_few_shot_examples: + skill_examples_block = ( + "## Skill Examples\n\n" + + "\n\n".join(self._skill_few_shot_examples) + + "\n\n" + ) + else: + skill_examples_block = "" + system_prompt = REACT_SYSTEM_PROMPT.format( + tool_descriptions=tool_desc, + skill_examples=skill_examples_block, + ) messages = self._build_messages(input, context, system_prompt=system_prompt) diff --git a/src/openjarvis/cli/bench_cmd.py b/src/openjarvis/cli/bench_cmd.py index ab6f431d..24fd8ccb 100644 --- a/src/openjarvis/cli/bench_cmd.py +++ b/src/openjarvis/cli/bench_cmd.py @@ -358,3 +358,130 @@ def run( energy_monitor.close() except Exception as exc: logger.debug("Energy monitor cleanup failed: %s", exc) + + +@bench.command("skills") +@click.option( + "--condition", + "-c", + type=click.Choice( + [ + "all", + "no_skills", + "skills_on", + "skills_optimized_dspy", + "skills_optimized_gepa", + ] + ), + default="all", + show_default=True, + help="Which condition(s) to run.", +) +@click.option( + "--seeds", + "-s", + multiple=True, + type=int, + default=(42, 43, 44), + show_default=True, + help="Random seeds to run per condition.", +) +@click.option( + "--max-samples", + "-n", + default=None, + type=int, + help="Limit number of PinchBench tasks (default: full benchmark).", +) +@click.option( + "--model", + "-m", + default="qwen3.5:9b", + show_default=True, + help="Model identifier.", +) +@click.option( + "--engine", + "-e", + default="ollama", + show_default=True, + help="Engine backend.", +) +@click.option( + "--output-dir", + "-o", + default="docs/superpowers/results/", + show_default=True, + help="Where the markdown report lands.", +) +def skills( + condition: str, + seeds: tuple, + max_samples: int, + model: str, + engine: str, + output_dir: str, +) -> None: + """Run the PinchBench skills benchmark across one or all conditions.""" + from rich.console import Console + from rich.table import Table + + from openjarvis.evals.skill_benchmark import ( + SkillBenchmarkConfig, + SkillBenchmarkRunner, + ) + + console = Console() + + cfg = SkillBenchmarkConfig( + model=model, + engine=engine, + seeds=list(seeds), + max_samples=max_samples, + output_dir=Path(output_dir), + ) + runner = SkillBenchmarkRunner(cfg) + + if condition == "all": + comparison = runner.run_all_conditions() + + # Print summary table + table = Table(title="PinchBench Skills Evaluation") + table.add_column("Condition", style="cyan") + table.add_column("Mean pass rate") + table.add_column("Stddev") + table.add_column("Tokens") + table.add_column("Runtime (s)") + for cond_name, result in comparison.results.items(): + table.add_row( + cond_name, + f"{result.mean_pass_rate:.3f}", + f"{result.stddev_pass_rate:.3f}", + str(result.total_tokens), + f"{result.total_runtime_seconds:.1f}", + ) + console.print(table) + + if comparison.deltas: + console.print("\n[bold]Deltas:[/bold]") + for name, value in comparison.deltas.items(): + sign = "+" if value >= 0 else "" + console.print(f" {name}: {sign}{value:.3f}") + + report_path = runner.write_report(comparison) + console.print(f"\n[green]Report written:[/green] {report_path}") + else: + result = runner.run_condition(condition) + table = Table(title=f"PinchBench Skills — {condition}") + table.add_column("Field", style="cyan") + table.add_column("Value") + table.add_row("Condition", result.condition) + table.add_row("Mean pass rate", f"{result.mean_pass_rate:.3f}") + table.add_row("Stddev", f"{result.stddev_pass_rate:.3f}") + table.add_row( + "Per-seed", + ", ".join(f"{s}={r:.3f}" for s, r in result.per_seed_pass_rate.items()), + ) + table.add_row("Total tokens", str(result.total_tokens)) + table.add_row("Runtime (s)", f"{result.total_runtime_seconds:.1f}") + console.print(table) diff --git a/src/openjarvis/cli/optimize_cmd.py b/src/openjarvis/cli/optimize_cmd.py index 3f586b2a..04bd0e2e 100644 --- a/src/openjarvis/cli/optimize_cmd.py +++ b/src/openjarvis/cli/optimize_cmd.py @@ -11,6 +11,18 @@ from rich.console import Console from rich.table import Table +def _get_trace_store(): + """Return a TraceStore from user config, or None on failure.""" + try: + from openjarvis.core.config import load_config + from openjarvis.traces.store import TraceStore + + cfg = load_config() + return TraceStore(cfg.traces.db_path) + except Exception: + return None + + @click.group("optimize") def optimize_group() -> None: """LLM-driven configuration optimization.""" @@ -417,4 +429,85 @@ def optimize_personal(action: str, workflow: str, trials: int) -> None: ) +@optimize_group.command("skills") +@click.option( + "--policy", + "-p", + type=click.Choice(["dspy", "gepa"]), + default="dspy", + show_default=True, + help="Optimization policy to use.", +) +@click.option( + "--min-traces", + "-n", + default=20, + show_default=True, + type=int, + help="Minimum traces required per skill.", +) +@click.option( + "--dry-run", + is_flag=True, + default=False, + help="Show planned work without invoking the optimizer LM.", +) +def skills(policy: str, min_traces: int, dry_run: bool) -> None: + """Optimize per-skill descriptions and few-shot examples from traces.""" + from rich.console import Console + from rich.table import Table + + from openjarvis.core.events import EventBus + from openjarvis.learning.agents.skill_optimizer import SkillOptimizer + from openjarvis.skills.manager import SkillManager + + console = Console() + store = _get_trace_store() + if store is None: + console.print("[red]No trace store found. Enable tracing first.[/red]") + raise SystemExit(1) + + if dry_run: + # Just bucket and report counts + traces = store.list_traces(limit=10000) + opt = SkillOptimizer(min_traces_per_skill=min_traces, optimizer=policy) + buckets = opt._bucket_traces_by_skill(traces) + if not buckets: + console.print("[dim]No skill-tagged traces found.[/dim]") + return + table = Table(title="Optimization plan (dry run)") + table.add_column("Skill", style="cyan") + table.add_column("Trace count") + table.add_column("Action") + for name, bucket in buckets.items(): + action = ( + "would optimize" + if len(bucket) >= min_traces + else f"skip (< {min_traces} traces)" + ) + table.add_row(name, str(len(bucket)), action) + console.print(table) + return + + # Real run + mgr = SkillManager(bus=EventBus()) + mgr.discover() + optimizer = SkillOptimizer(min_traces_per_skill=min_traces, optimizer=policy) + results = optimizer.optimize(store, mgr) + + if not results: + console.print("[dim]No skill-tagged traces found.[/dim]") + return + + table = Table(title=f"Skill optimization ({policy})") + table.add_column("Skill", style="cyan") + table.add_column("Status") + table.add_column("Traces") + table.add_column("Overlay path") + for name, res in results.items(): + path_str = str(res.overlay_path) if res.overlay_path else "—" + table.add_row(name, res.status, str(res.trace_count), path_str) + console.print(table) + + __all__ = ["optimize_group"] diff --git a/src/openjarvis/cli/skill_cmd.py b/src/openjarvis/cli/skill_cmd.py index 4f268b0f..c935c7e7 100644 --- a/src/openjarvis/cli/skill_cmd.py +++ b/src/openjarvis/cli/skill_cmd.py @@ -1,75 +1,525 @@ -"""``jarvis skill`` — skill management commands.""" +"""CLI commands for skill management.""" from __future__ import annotations +from pathlib import Path +from typing import List + import click from rich.console import Console from rich.table import Table +from openjarvis.core.config import load_config +from openjarvis.core.events import EventBus +from openjarvis.skills.manager import SkillManager + + +def _get_trace_store(): + """Return a TraceStore instance from the user config (or None).""" + try: + from openjarvis.core.config import load_config + from openjarvis.traces.store import TraceStore + + cfg = load_config() + return TraceStore(cfg.traces.db_path) + except Exception: + return None + + +def _get_discovered_dir() -> Path: + """Return the directory where discovered skill manifests are written.""" + return Path("~/.openjarvis/skills/discovered/").expanduser() + + +def _get_overlay_dir() -> Path: + """Return the directory where optimization overlays are stored.""" + return Path("~/.openjarvis/learning/skills/").expanduser() + + +def _get_skill_paths() -> List[Path]: + paths: List[Path] = [] + workspace = Path("./skills") + if workspace.exists(): + paths.append(workspace) + user_dir = Path("~/.openjarvis/skills/").expanduser() + paths.append(user_dir) + return paths + + +def _get_manager() -> SkillManager: + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=_get_skill_paths()) + return mgr + @click.group() -def skill() -> None: - """Manage skills — list, install, remove.""" +def skill(): + """Manage reusable skills.""" @skill.command("list") -def list_skills() -> None: +def list_skills(): """List installed skills.""" - console = Console(stderr=True) - try: - from openjarvis.core.registry import SkillRegistry - - keys = sorted(SkillRegistry.keys()) - if not keys: - console.print("[dim]No skills installed.[/dim]") - return - table = Table(title="Installed Skills") - table.add_column("Name", style="cyan") - table.add_column("Description", style="green") - for key in keys: - skill_cls = SkillRegistry.get(key) - desc = "" - if hasattr(skill_cls, "manifest"): - m = skill_cls.manifest if not callable(skill_cls.manifest) else None - if m and hasattr(m, "description"): - desc = m.description[:60] - table.add_row(key, desc) - console.print(table) - except Exception as exc: - console.print(f"[red]Error: {exc}[/red]") + console = Console() + mgr = _get_manager() + names = mgr.skill_names() + if not names: + console.print("[dim]No skills installed.[/dim]") + return + table = Table(title="Installed Skills") + table.add_column("Name", style="cyan") + table.add_column("Description", max_width=50) + table.add_column("Version") + table.add_column("Tags") + for name in sorted(names): + m = mgr.resolve(name) + tags = ", ".join(m.tags) if m.tags else "" + desc = m.description[:50] + "..." if len(m.description) > 50 else m.description + table.add_row(name, desc, m.version, tags) + console.print(table) -@skill.command() +@skill.command("info") @click.argument("skill_name") -def install(skill_name: str) -> None: - """Install a skill from the bundled library.""" - console = Console(stderr=True) - console.print(f"[yellow]Installing skill: {skill_name}[/yellow]") - # Skills are discovered from TOML files — point user to the right place +def info(skill_name: str): + """Show detailed information about a skill.""" + console = Console() + mgr = _get_manager() + try: + m = mgr.resolve(skill_name) + except KeyError: + console.print(f"[red]Skill '{skill_name}' not found.[/red]") + raise SystemExit(1) + console.print(f"[bold]{m.name}[/bold] v{m.version}") + if m.author: + console.print(f"Author: {m.author}") + if m.description: + console.print(f"Description: {m.description}") + if m.tags: + console.print(f"Tags: {', '.join(m.tags)}") + if m.required_capabilities: + console.print(f"Capabilities: {', '.join(m.required_capabilities)}") + if m.depends: + console.print(f"Dependencies: {', '.join(m.depends)}") + if m.steps: + console.print(f"Steps: {len(m.steps)}") + if m.markdown_content: + console.print("Has instructions: yes") + console.print(f"User invocable: {m.user_invocable}") console.print( - f"[dim]Place skill TOML file in ~/.openjarvis/skills/{skill_name}.toml[/dim]" + f"Model invocation: {'disabled' if m.disable_model_invocation else 'enabled'}" ) -@skill.command() +@skill.command("run") @click.argument("skill_name") -def remove(skill_name: str) -> None: - """Remove an installed skill.""" - console = Console(stderr=True) - console.print(f"[yellow]Removing skill: {skill_name}[/yellow]") - console.print("[dim]Skill removal not yet implemented.[/dim]") +@click.option("--arg", "-a", multiple=True, help="Arguments as key=value pairs.") +def run(skill_name: str, arg: tuple): + """Execute a skill directly.""" + console = Console() + mgr = _get_manager() + context = {} + for a in arg: + if "=" in a: + k, v = a.split("=", 1) + context[k.strip()] = v.strip() + try: + result = mgr.execute(skill_name, context) + except KeyError: + console.print(f"[red]Skill '{skill_name}' not found.[/red]") + raise SystemExit(1) + if result.success: + console.print("[green]Success[/green]") + if result.step_results: + console.print(result.step_results[-1].content) + else: + console.print("[red]Failed[/red]") + if result.step_results: + console.print(result.step_results[-1].content) -@skill.command() -@click.argument("query", default="") -def search(query: str) -> None: - """Search for available skills.""" - console = Console(stderr=True) - if not query: - console.print("[dim]Provide a search query.[/dim]") +def _parse_source_query(query: str) -> tuple[str, str]: + """Parse a ``:`` query into (source, name). + + Raises ``click.BadParameter`` if the format is wrong. + """ + if ":" not in query: + raise click.BadParameter( + f"Expected ':' format (e.g. 'hermes:apple-notes'), " + f"got: {query!r}" + ) + source, _, name = query.partition(":") + if not source or not name: + raise click.BadParameter( + f"Both source and name are required: got source={source!r}, name={name!r}" + ) + return source, name + + +def _get_resolver(source: str, url: str = ""): + """Return a resolver instance for the given source name.""" + if source == "hermes": + from openjarvis.skills.sources.hermes import HermesResolver + + return HermesResolver() + if source == "openclaw": + from openjarvis.skills.sources.openclaw import OpenClawResolver + + return OpenClawResolver() + if source == "github": + if not url: + raise click.BadParameter("github source requires --url") + from pathlib import Path as _Path + + from openjarvis.skills.sources.github import GitHubResolver + + cache = _Path( + "~/.openjarvis/skill-cache/github/" + url.rstrip("/").rsplit("/", 1)[-1] + ).expanduser() + return GitHubResolver(cache_root=cache, repo_url=url) + raise click.BadParameter(f"Unknown source: {source!r}") + + +@skill.command("install") +@click.argument("query") +@click.option( + "--with-scripts", + is_flag=True, + default=False, + help="Import the skill's scripts/ directory (security-sensitive).", +) +@click.option( + "--force", is_flag=True, default=False, help="Overwrite existing install." +) +@click.option( + "--url", + default="", + help="Repo URL (required when source is 'github').", +) +def install(query: str, with_scripts: bool, force: bool, url: str): + """Install a skill from a source. + + Example: ``jarvis skill install hermes:apple-notes`` + """ + console = Console() + source, name = _parse_source_query(query) + + resolver = _get_resolver(source, url=url) + try: + resolver.sync() + except Exception as exc: + console.print(f"[red]Failed to sync source {source}: {exc}[/red]") + raise SystemExit(1) + + # Support category/name queries (e.g. openclaw:owner/slug, hermes:apple/apple-notes) + if "/" in name: + category, _, skill_name = name.partition("/") + matches = [ + s + for s in resolver.list_skills() + if s.name == skill_name and s.category == category + ] + else: + matches = [s for s in resolver.list_skills() if s.name == name] + if not matches: + console.print(f"[red]No skill named '{name}' found in source '{source}'[/red]") + raise SystemExit(1) + + from openjarvis.skills.importer import SkillImporter + from openjarvis.skills.parser import SkillParser + from openjarvis.skills.tool_translator import ToolTranslator + + importer = SkillImporter(parser=SkillParser(), tool_translator=ToolTranslator()) + result = importer.import_skill(matches[0], with_scripts=with_scripts, force=force) + + if result.success: + if result.skipped: + console.print("[yellow]Skill already installed[/yellow]") + else: + console.print(f"[green]Installed:[/green] {result.target_path}") + if result.translated_tools: + console.print(f" Translated tools: {', '.join(result.translated_tools)}") + if result.untranslated_tools: + console.print( + f" [yellow]Untranslated tools:[/yellow] " + f"{', '.join(result.untranslated_tools)}" + ) + for warning in result.warnings: + console.print(f" [yellow]{warning}[/yellow]") + else: + console.print( + "[red]Install failed: " + + "; ".join(result.warnings or ["unknown error"]) + + "[/red]" + ) + raise SystemExit(1) + + +@skill.command("sync") +@click.argument("source", required=False) +@click.option("--category", default="", help="Filter by category.") +@click.option("--tag", default="", help="Filter by tag.") +@click.option("--search", default="", help="Substring search across name+description.") +@click.option( + "--with-scripts", + is_flag=True, + default=False, + help="Import scripts/ directories.", +) +@click.option("--force", is_flag=True, default=False, help="Re-import existing skills.") +def sync( + source: str, + category: str, + tag: str, + search: str, + with_scripts: bool, + force: bool, +): + """Bulk install + update from a source (or all configured sources).""" + console = Console() + + cfg = load_config() + + # Determine which sources to sync + source_configs: list = [] + if source: + source_configs.append({"source": source, "filter": {}, "url": ""}) + else: + for src_cfg in cfg.skills.sources: + source_configs.append( + { + "source": src_cfg.source, + "filter": dict(src_cfg.filter or {}), + "url": src_cfg.url, + } + ) + + if not source_configs: + console.print( + "[yellow]No sources to sync. " + "Add sources to [skills.sources] in config.toml " + "or pass a source name.[/yellow]" + ) return - console.print(f"[dim]Searching for skills matching '{query}'...[/dim]") - console.print("[dim]Skill search not yet implemented.[/dim]") + + from openjarvis.skills.importer import SkillImporter + from openjarvis.skills.parser import SkillParser + from openjarvis.skills.tool_translator import ToolTranslator + + importer = SkillImporter(parser=SkillParser(), tool_translator=ToolTranslator()) + + total_installed = 0 + for src in source_configs: + console.print(f"[cyan]Syncing {src['source']}...[/cyan]") + try: + resolver = _get_resolver(src["source"], url=src["url"]) + resolver.sync() + except Exception as exc: + console.print(f"[red]Failed to sync {src['source']}: {exc}[/red]") + continue + + skills_to_import = resolver.list_skills() + + # Apply CLI filters + if category: + skills_to_import = [s for s in skills_to_import if s.category == category] + if search: + sl = search.lower() + skills_to_import = [ + s + for s in skills_to_import + if sl in s.name.lower() or sl in s.description.lower() + ] + + # Apply config filter (categories list) + cfg_categories = src["filter"].get("category") or [] + if cfg_categories: + skills_to_import = [ + s for s in skills_to_import if s.category in cfg_categories + ] + + installed_count = 0 + for resolved in skills_to_import: + r = importer.import_skill(resolved, with_scripts=with_scripts, force=force) + if r.success and not r.skipped: + installed_count += 1 + console.print(f" Imported {installed_count}/{len(skills_to_import)} skills") + total_installed += installed_count + + console.print(f"[green]Total installed: {total_installed}[/green]") + + +@skill.command("sources") +def sources(): + """List configured skill sources.""" + console = Console() + + cfg = load_config() + + if not cfg.skills.sources: + console.print( + "[dim]No skill sources configured. " + "Add entries to [skills.sources] in config.toml.[/dim]" + ) + return + + table = Table(title="Configured Skill Sources") + table.add_column("Source", style="cyan") + table.add_column("URL") + table.add_column("Filter") + table.add_column("Auto-update") + for s in cfg.skills.sources: + filt = ", ".join(f"{k}={v}" for k, v in (s.filter or {}).items()) or "—" + table.add_row( + s.source, + s.url or "(default)", + filt, + "yes" if s.auto_update else "no", + ) + console.print(table) + + +@skill.command("remove") +@click.argument("skill_name") +def remove(skill_name: str): + """Remove an installed skill.""" + Console().print("[yellow]Skill removal not yet implemented.[/yellow]") + + +@skill.command("search") +@click.argument("query") +def search(query: str): + """Search available skills.""" + Console().print( + "[yellow]Skill search not yet implemented. " + "Run 'jarvis skill update' first.[/yellow]" + ) + + +@skill.command("update") +def update(): + """Pull latest commits for all installed skill sources.""" + console = Console() + + cfg = load_config() + if not cfg.skills.sources: + console.print("[dim]No sources configured.[/dim]") + return + + for src in cfg.skills.sources: + console.print(f"[cyan]Updating {src.source}...[/cyan]") + try: + resolver = _get_resolver(src.source, url=src.url) + resolver.sync() + console.print(" [green]OK[/green]") + except Exception as exc: + console.print(f" [red]Failed: {exc}[/red]") + + +@skill.command("discover") +@click.option( + "--min-frequency", + "-f", + default=3, + show_default=True, + type=int, + help="Minimum recurrence count to surface a tool sequence as a skill.", +) +@click.option( + "--min-outcome", + "-o", + default=0.5, + show_default=True, + type=float, + help="Minimum average outcome score (0.0-1.0) to qualify.", +) +@click.option( + "--dry-run", + is_flag=True, + default=False, + help="Print discovered patterns without writing manifests.", +) +def discover(min_frequency: int, min_outcome: float, dry_run: bool) -> None: + """Mine the trace store for recurring tool sequences and write them as + discovered skill manifests under ~/.openjarvis/skills/discovered/.""" + console = Console() + store = _get_trace_store() + if store is None: + console.print( + "[red]No trace store found. " + "Enable tracing in config (traces.enabled = true) and run " + "some queries first.[/red]" + ) + raise SystemExit(1) + + output_dir = _get_discovered_dir() + mgr = SkillManager(bus=EventBus()) + + if dry_run: + # Use a temporary directory so nothing is persisted + import tempfile + + tmp = Path(tempfile.mkdtemp(prefix="openjarvis-discover-dryrun-")) + try: + written = mgr.discover_from_traces( + store, + min_frequency=min_frequency, + min_outcome=min_outcome, + output_dir=tmp, + ) + finally: + import shutil + + shutil.rmtree(tmp, ignore_errors=True) + else: + written = mgr.discover_from_traces( + store, + min_frequency=min_frequency, + min_outcome=min_outcome, + output_dir=output_dir, + ) + + if not written: + console.print( + "[dim]No recurring tool sequences found above the threshold.[/dim]" + ) + return + + table = Table(title="Discovered Skills") + table.add_column("Name", style="cyan") + table.add_column("Path") + for item in written: + table.add_row(item["name"], item["path"]) + console.print(table) + if dry_run: + console.print("[yellow]--dry-run: no files were written.[/yellow]") + + +@skill.command("show-overlay") +@click.argument("skill_name") +def show_overlay(skill_name: str) -> None: + """Show the optimization overlay for a skill, if one exists.""" + console = Console() + from openjarvis.skills.overlay import SkillOverlayLoader + + loader = SkillOverlayLoader(_get_overlay_dir()) + overlay = loader.load(skill_name) + if overlay is None: + console.print(f"[red]No overlay found for skill '{skill_name}'.[/red]") + raise SystemExit(1) + + console.print(f"[bold]{overlay.skill_name}[/bold]") + console.print(f"Optimizer: {overlay.optimizer}") + console.print(f"Optimized at: {overlay.optimized_at}") + console.print(f"Trace count: {overlay.trace_count}") + console.print(f"Description: {overlay.description}") + if overlay.few_shot: + console.print(f"Few-shot examples ({len(overlay.few_shot)}):") + for i, ex in enumerate(overlay.few_shot, start=1): + inp = (ex.get("input", "") or "")[:100] + out = (ex.get("output", "") or "")[:100] + console.print(f" {i}. input={inp!r}") + console.print(f" output={out!r}") __all__ = ["skill"] diff --git a/src/openjarvis/core/config.py b/src/openjarvis/core/config.py index 3025c22a..c31229e0 100644 --- a/src/openjarvis/core/config.py +++ b/src/openjarvis/core/config.py @@ -653,6 +653,17 @@ class AgentLearningConfig: gepa: GEPAOptimizerConfig = field(default_factory=GEPAOptimizerConfig) +@dataclass(slots=True) +class SkillsLearningConfig: + """Configuration for the skills learning loop (Plan 2A).""" + + auto_optimize: bool = False # opt in via config + optimizer: str = "dspy" # "dspy" or "gepa" + min_traces_per_skill: int = 20 + optimization_interval_seconds: int = 86400 + overlay_dir: str = "~/.openjarvis/learning/skills/" + + @dataclass(slots=True) class MetricsConfig: """Reward / optimization metric weights.""" @@ -675,6 +686,7 @@ class LearningConfig: default_factory=IntelligenceLearningConfig, ) agent: AgentLearningConfig = field(default_factory=AgentLearningConfig) + skills: SkillsLearningConfig = field(default_factory=SkillsLearningConfig) metrics: MetricsConfig = field(default_factory=MetricsConfig) # Training pipeline @@ -1260,13 +1272,31 @@ class CompressionConfig: strategy: str = "session_consolidation" +@dataclass(slots=True) +class SkillSourceConfig: + """Configuration for a single skill source (Hermes, OpenClaw, GitHub).""" + + source: str = "" # "hermes", "openclaw", or "github" + url: str = "" # required when source = "github" + filter: Dict[str, Any] = field(default_factory=dict) + auto_update: bool = False + + @dataclass(slots=True) class SkillsConfig: """Configuration for agent-authored procedural skills.""" + enabled: bool = True skills_dir: str = "~/.openjarvis/skills/" - nudge_interval: int = 15 + active: str = "*" auto_discover: bool = True + auto_sync: bool = False + nudge_interval: int = 15 + index_repo: str = "https://github.com/openjarvis/skill-index.git" + index_dir: str = "~/.openjarvis/skill-index/" + max_depth: int = 5 + sandbox_dangerous: bool = True + sources: List[SkillSourceConfig] = field(default_factory=list) @dataclass diff --git a/src/openjarvis/evals/backends/jarvis_agent.py b/src/openjarvis/evals/backends/jarvis_agent.py index e4981fc4..f1f72dce 100644 --- a/src/openjarvis/evals/backends/jarvis_agent.py +++ b/src/openjarvis/evals/backends/jarvis_agent.py @@ -3,6 +3,7 @@ from __future__ import annotations import time +from pathlib import Path from typing import Any, Dict, List, Optional from openjarvis.evals.core.backend import InferenceBackend @@ -26,6 +27,8 @@ class JarvisAgentBackend(InferenceBackend): gpu_metrics: bool = False, model: Optional[str] = None, max_turns: Optional[int] = None, + skills_enabled: bool = True, + overlay_dir: Optional[Path] = None, ) -> None: from openjarvis.system import SystemBuilder @@ -52,6 +55,11 @@ class JarvisAgentBackend(InferenceBackend): # GAIA tasks before this was configurable per-eval). if max_turns is not None: builder._config.agent.max_turns = max_turns + # Plan 2B: per-condition skill switches. Mutate the builder's + # config directly so SystemBuilder picks them up at build time. + builder._config.skills.enabled = skills_enabled + if overlay_dir is not None: + builder._config.learning.skills.overlay_dir = str(overlay_dir) self._system = builder.telemetry(telemetry).traces(True).build() def generate( diff --git a/src/openjarvis/evals/skill_benchmark.py b/src/openjarvis/evals/skill_benchmark.py new file mode 100644 index 00000000..dbed8b0c --- /dev/null +++ b/src/openjarvis/evals/skill_benchmark.py @@ -0,0 +1,508 @@ +"""SkillBenchmarkRunner — orchestrate the 4-condition × N-seed × M-task +PinchBench sweep that measures whether skills + DSPy/GEPA optimization +improves agent performance. + +Plan 2B implementation. See: +docs/superpowers/specs/2026-04-08-skills-benchmark-evaluation-design.md +""" + +from __future__ import annotations + +import logging +import statistics +import time +from dataclasses import dataclass, field +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Dict, List, Optional + +LOGGER = logging.getLogger(__name__) + +CONDITIONS = ( + "no_skills", + "skills_on", + "skills_optimized_dspy", + "skills_optimized_gepa", +) + + +# --------------------------------------------------------------------------- +# Config + result types +# --------------------------------------------------------------------------- + + +@dataclass(slots=True) +class SkillBenchmarkConfig: + """Configuration for a single SkillBenchmarkRunner sweep.""" + + benchmark: str = "pinchbench" + model: str = "qwen3.5:9b" + engine: str = "ollama" + agent: str = "native_react" + tools: List[str] = field( + default_factory=lambda: [ + "calculator", + "think", + "shell_exec", + "web_search", + "file_read", + "file_write", + ] + ) + seeds: List[int] = field(default_factory=lambda: [42, 43, 44]) + max_samples: Optional[int] = None + output_dir: Path = field(default_factory=lambda: Path("docs/superpowers/results/")) + skills_dir: Path = field( + default_factory=lambda: Path("~/.openjarvis/skills/").expanduser() + ) + overlay_dir_dspy: Path = field( + default_factory=lambda: Path("~/.openjarvis/learning/skills-dspy/").expanduser() + ) + overlay_dir_gepa: Path = field( + default_factory=lambda: Path("~/.openjarvis/learning/skills-gepa/").expanduser() + ) + + +@dataclass(slots=True) +class ConditionResult: + """Aggregated result for a single condition across all seeds.""" + + condition: str + seeds: List[int] + per_seed_pass_rate: Dict[int, float] + mean_pass_rate: float + stddev_pass_rate: float + per_task_results: Dict[str, List[bool]] # task_id → [pass_seed1, ...] + skill_invocation_counts: Dict[str, int] # skill_name → total invocations + total_tokens: int + total_runtime_seconds: float + + +@dataclass(slots=True) +class ConditionComparison: + """Top-level result of a SkillBenchmarkRunner sweep.""" + + config: SkillBenchmarkConfig + started_at: str + ended_at: str + results: Dict[str, ConditionResult] + deltas: Dict[str, float] + + +class SkillBenchmarkRunner: + """Orchestrates the 4-condition × N-seed × M-task PinchBench sweep. + + Each condition is a different SystemBuilder configuration: + - no_skills: cfg.skills.enabled = False + - skills_on: enabled, overlay_dir = empty (no overlays load) + - skills_optimized_dspy: enabled, overlay_dir = config.overlay_dir_dspy + - skills_optimized_gepa: enabled, overlay_dir = config.overlay_dir_gepa + + Per-seed runs share the same backend instantiation but pass a fresh + seed to the EvalRunner. + """ + + def __init__(self, config: SkillBenchmarkConfig) -> None: + self._config = config + # An "empty" overlay dir for the skills_on condition. We point at + # a known-empty subdirectory under the output dir so SkillManager + # finds zero overlays even if the user happens to have populated + # the default ~/.openjarvis/learning/skills/ tree. + self._empty_overlay_dir = ( + Path(self._config.output_dir).expanduser() / "_skills_on_empty_overlays" + ) + self._empty_overlay_dir.mkdir(parents=True, exist_ok=True) + + # ------------------------------------------------------------------ + # Per-condition backend construction + # ------------------------------------------------------------------ + + def _backend_kwargs_for_condition(self, condition: str) -> Dict[str, Any]: + """Return the kwargs to pass to JarvisAgentBackend for *condition*. + + Pure function — no side effects, no SystemBuilder construction. + Tested in isolation so we can verify the per-condition switches + without invoking the engine. + """ + if condition == "no_skills": + return { + "skills_enabled": False, + "overlay_dir": None, + } + if condition == "skills_on": + return { + "skills_enabled": True, + "overlay_dir": self._empty_overlay_dir, + } + if condition == "skills_optimized_dspy": + return { + "skills_enabled": True, + "overlay_dir": self._config.overlay_dir_dspy, + } + if condition == "skills_optimized_gepa": + return { + "skills_enabled": True, + "overlay_dir": self._config.overlay_dir_gepa, + } + raise ValueError( + f"Unknown condition '{condition}'. Expected one of: " + f"{', '.join(CONDITIONS)}" + ) + + def _build_backend_for_condition(self, condition: str) -> Any: + """Construct a JarvisAgentBackend for *condition*. + + Separate from `_backend_kwargs_for_condition` so the kwarg logic + can be tested without instantiating an engine. + """ + from openjarvis.evals.backends.jarvis_agent import JarvisAgentBackend + + kw = self._backend_kwargs_for_condition(condition) + return JarvisAgentBackend( + engine_key=self._config.engine, + agent_name=self._config.agent, + tools=list(self._config.tools), + telemetry=False, + model=self._config.model, + skills_enabled=kw["skills_enabled"], + overlay_dir=kw["overlay_dir"], + ) + + # ------------------------------------------------------------------ + # Per-seed and per-condition runs + # ------------------------------------------------------------------ + + def _run_single_seed( + self, + condition: str, + seed: int, + ) -> Dict[str, Any]: + """Run the benchmark once for one (condition, seed) pair. + + Returns a raw dict with per-task results, pass rate, skill + invocation counts, total tokens, and total runtime. + + This method is the integration boundary with the existing + EvalRunner / PinchBenchDataset. It is intentionally a thin + shim so tests can monkeypatch it without instantiating an + engine or running real benchmark tasks. + """ + from openjarvis.evals.backends.jarvis_direct import JarvisDirectBackend + from openjarvis.evals.core.runner import EvalRunner + from openjarvis.evals.core.types import RunConfig + from openjarvis.evals.datasets.pinchbench import PinchBenchDataset + from openjarvis.evals.scorers.pinchbench import PinchBenchScorer + + backend = self._build_backend_for_condition(condition) + + dataset = PinchBenchDataset(path=None) + dataset.load( + max_samples=self._config.max_samples, + split=None, + seed=seed, + ) + + # PinchBenchScorer is an LLM-as-judge scorer; it needs a judge + # backend. We reuse the same engine the agent uses (typically + # a local Ollama model) so the headline run is fully local. + try: + judge_backend = JarvisDirectBackend( + engine_key=self._config.engine, + ) + except RuntimeError as exc: + LOGGER.warning( + "Judge backend unavailable, automated checks only: %s", + exc, + ) + judge_backend = None + scorer = PinchBenchScorer(judge_backend, self._config.model) + + runner_cfg = RunConfig( + benchmark=self._config.benchmark, + backend="jarvis-agent", + model=self._config.model, + max_workers=1, + episode_mode=False, + seed=seed, + ) + + eval_runner = EvalRunner( + config=runner_cfg, + dataset=dataset, + backend=backend, + scorer=scorer, + ) + + t0 = time.monotonic() + summary = eval_runner.run() + elapsed = time.monotonic() - t0 + + # Aggregate per-task results. RunSummary doesn't carry the + # per-record list — that lives on the runner instance via the + # `results` property (which returns List[EvalResult]). + per_task: Dict[str, bool] = {} + for r in eval_runner.results: + rid = getattr(r, "record_id", "") or "" + ok = bool(getattr(r, "is_correct", False)) + if rid: + per_task[rid] = ok + + # Pass rate. Prefer the runner-computed accuracy when we have + # records; fall back to summary.accuracy otherwise. + if per_task: + pass_rate = sum(1 for v in per_task.values() if v) / len(per_task) + else: + pass_rate = float(getattr(summary, "accuracy", 0.0) or 0.0) + + # Skill invocation counts from the trace store (if any) + skill_invocations = self._extract_skill_invocations(backend) + + # Token totals. RunSummary tracks input/output tokens separately. + total_input = int(getattr(summary, "total_input_tokens", 0) or 0) + total_output = int(getattr(summary, "total_output_tokens", 0) or 0) + total_tokens = total_input + total_output + + return { + "pass_rate": pass_rate, + "per_task": per_task, + "skill_invocations": skill_invocations, + "total_tokens": total_tokens, + "total_runtime_seconds": elapsed, + } + + def _extract_skill_invocations(self, backend: Any) -> Dict[str, int]: + """Count per-skill invocations in the trace store from this run. + + Walks all traces in the backend's system trace store and counts + TraceStep.metadata.skill values. Returns an empty dict on any + failure (the count is informational, not load-bearing). + """ + try: + system = getattr(backend, "_system", None) + if system is None: + return {} + store = getattr(system, "trace_store", None) + if store is None: + return {} + traces = store.list_traces(limit=10000) + except Exception: + return {} + + counts: Dict[str, int] = {} + for trace in traces: + steps = getattr(trace, "steps", None) or [] + for step in steps: + meta = getattr(step, "metadata", None) or {} + if isinstance(meta, dict): + name = meta.get("skill") + if name: + counts[name] = counts.get(name, 0) + 1 + return counts + + def run_condition(self, condition: str) -> ConditionResult: + """Run the benchmark for *condition* across all configured seeds. + + Returns a ConditionResult with mean ± stddev pass rate and the + per-task / per-skill aggregations. + """ + if condition not in CONDITIONS: + raise ValueError( + f"Unknown condition '{condition}'. Expected one of: " + f"{', '.join(CONDITIONS)}" + ) + + per_seed_pass_rate: Dict[int, float] = {} + per_task_results: Dict[str, List[bool]] = {} + skill_counts: Dict[str, int] = {} + total_tokens = 0 + total_runtime = 0.0 + + for seed in self._config.seeds: + seed_data = self._run_single_seed(condition, seed) + per_seed_pass_rate[seed] = float(seed_data["pass_rate"]) + + for task_id, ok in seed_data["per_task"].items(): + per_task_results.setdefault(task_id, []).append(bool(ok)) + + for skill_name, count in seed_data["skill_invocations"].items(): + skill_counts[skill_name] = skill_counts.get(skill_name, 0) + int(count) + + total_tokens += int(seed_data["total_tokens"]) + total_runtime += float(seed_data["total_runtime_seconds"]) + + rates = list(per_seed_pass_rate.values()) + mean_rate = statistics.fmean(rates) if rates else 0.0 + stddev_rate = statistics.stdev(rates) if len(rates) > 1 else 0.0 + + return ConditionResult( + condition=condition, + seeds=list(self._config.seeds), + per_seed_pass_rate=per_seed_pass_rate, + mean_pass_rate=mean_rate, + stddev_pass_rate=stddev_rate, + per_task_results=per_task_results, + skill_invocation_counts=skill_counts, + total_tokens=total_tokens, + total_runtime_seconds=total_runtime, + ) + + # ------------------------------------------------------------------ + # Sweep + report + # ------------------------------------------------------------------ + + def run_all_conditions(self) -> ConditionComparison: + """Run all 4 conditions × all seeds. + + Returns a ConditionComparison with per-condition results and the + computed deltas (skills_on - no_skills, etc.). + """ + started_at = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + results: Dict[str, ConditionResult] = {} + for condition in CONDITIONS: + LOGGER.info("Running condition: %s", condition) + results[condition] = self.run_condition(condition) + + ended_at = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + # Compute the headline deltas + deltas: Dict[str, float] = {} + if "no_skills" in results and "skills_on" in results: + deltas["skills_on - no_skills"] = ( + results["skills_on"].mean_pass_rate + - results["no_skills"].mean_pass_rate + ) + if "skills_on" in results and "skills_optimized_dspy" in results: + deltas["skills_optimized_dspy - skills_on"] = ( + results["skills_optimized_dspy"].mean_pass_rate + - results["skills_on"].mean_pass_rate + ) + if "skills_on" in results and "skills_optimized_gepa" in results: + deltas["skills_optimized_gepa - skills_on"] = ( + results["skills_optimized_gepa"].mean_pass_rate + - results["skills_on"].mean_pass_rate + ) + + return ConditionComparison( + config=self._config, + started_at=started_at, + ended_at=ended_at, + results=results, + deltas=deltas, + ) + + def write_report(self, comparison: ConditionComparison) -> Path: + """Write a markdown report to output_dir. + + Filename: pinchbench-skills-eval-{YYYY-MM-DD}.md + """ + date_str = datetime.now(timezone.utc).strftime("%Y-%m-%d") + out_dir = Path(self._config.output_dir).expanduser() + out_dir.mkdir(parents=True, exist_ok=True) + path = out_dir / f"pinchbench-skills-eval-{date_str}.md" + + lines: List[str] = [] + lines.append(f"# PinchBench Skills Evaluation — {date_str}") + lines.append("") + lines.append(f"**Started:** {comparison.started_at}") + lines.append(f"**Ended:** {comparison.ended_at}") + lines.append(f"**Model:** {comparison.config.model}") + lines.append(f"**Engine:** {comparison.config.engine}") + lines.append(f"**Agent:** {comparison.config.agent}") + lines.append(f"**Seeds:** {', '.join(str(s) for s in comparison.config.seeds)}") + max_samples = comparison.config.max_samples + lines.append( + f"**Max samples:** {max_samples if max_samples is not None else 'full'}" + ) + lines.append("") + + # Summary table + lines.append("## Summary") + lines.append("") + lines.append( + "| Condition | Mean pass rate | Stddev | Total tokens | Runtime (s) |" + ) + lines.append("|---|---|---|---|---|") + for condition in CONDITIONS: + r = comparison.results.get(condition) + if r is None: + continue + lines.append( + f"| {condition} | {r.mean_pass_rate:.3f} | " + f"{r.stddev_pass_rate:.3f} | {r.total_tokens} | " + f"{r.total_runtime_seconds:.1f} |" + ) + lines.append("") + + # Deltas + if comparison.deltas: + lines.append("## Deltas") + lines.append("") + for name, value in comparison.deltas.items(): + sign = "+" if value >= 0 else "" + lines.append(f"- **{name}**: {sign}{value:.3f}") + lines.append("") + + # Per-task breakdown + lines.append("## Per-task results") + lines.append("") + all_tasks: set[str] = set() + for r in comparison.results.values(): + all_tasks.update(r.per_task_results.keys()) + if all_tasks: + header = "| Task | " + " | ".join(CONDITIONS) + " |" + sep = "|---|" + "|".join(["---"] * len(CONDITIONS)) + "|" + lines.append(header) + lines.append(sep) + for task_id in sorted(all_tasks): + row = [task_id] + for condition in CONDITIONS: + r = comparison.results.get(condition) + passes = ( + r.per_task_results.get(task_id, []) if r is not None else [] + ) + if not passes: + row.append("—") + else: + n_pass = sum(1 for v in passes if v) + row.append(f"{n_pass}/{len(passes)}") + lines.append("| " + " | ".join(row) + " |") + lines.append("") + + # Per-skill invocation counts + lines.append("## Per-skill invocation counts") + lines.append("") + all_skills: set[str] = set() + for r in comparison.results.values(): + all_skills.update(r.skill_invocation_counts.keys()) + if all_skills: + header = "| Skill | " + " | ".join(CONDITIONS) + " |" + sep = "|---|" + "|".join(["---"] * len(CONDITIONS)) + "|" + lines.append(header) + lines.append(sep) + for skill_name in sorted(all_skills): + row = [skill_name] + for condition in CONDITIONS: + r = comparison.results.get(condition) + count = ( + r.skill_invocation_counts.get(skill_name, 0) + if r is not None + else 0 + ) + row.append(str(count)) + lines.append("| " + " | ".join(row) + " |") + else: + lines.append("(no skill invocations recorded)") + lines.append("") + + path.write_text("\n".join(lines), encoding="utf-8") + return path + + +__all__ = [ + "CONDITIONS", + "ConditionComparison", + "ConditionResult", + "SkillBenchmarkConfig", + "SkillBenchmarkRunner", +] diff --git a/src/openjarvis/learning/agents/skill_discovery.py b/src/openjarvis/learning/agents/skill_discovery.py index ba9e5d23..afa72d77 100644 --- a/src/openjarvis/learning/agents/skill_discovery.py +++ b/src/openjarvis/learning/agents/skill_discovery.py @@ -130,15 +130,44 @@ class SkillDiscovery: "tool_name", getattr(step, "name", ""), ) + # Also check step.input dict for tool name (TraceStep format) + if not name: + step_input = getattr(step, "input", {}) or {} + if isinstance(step_input, dict): + name = step_input.get("tool", "") or "" if name: tools.append(name) return tools def _extract_outcome(self, trace: Any) -> float: - """Extract outcome score from a trace.""" + """Extract outcome score from a trace. + + Handles both numeric outcome scores and string labels + ("success" / "failure"). Falls back to the ``feedback`` field when + present, then defaults to 0.0. + """ if isinstance(trace, dict): - return float(trace.get("outcome", 0.0)) - return float(getattr(trace, "outcome", 0.0)) + raw = trace.get("outcome") + if raw is None: + return float(trace.get("feedback", 0.0) or 0.0) + else: + raw = getattr(trace, "outcome", None) + if raw is None: + return float(getattr(trace, "feedback", 0.0) or 0.0) + + if isinstance(raw, (int, float)): + return float(raw) + if isinstance(raw, str): + lraw = raw.lower() + if lraw == "success": + return 1.0 + if lraw == "failure": + return 0.0 + try: + return float(raw) + except ValueError: + return 0.0 + return 0.0 def _extract_query(self, trace: Any) -> str: """Extract the original query from a trace.""" diff --git a/src/openjarvis/learning/agents/skill_optimizer.py b/src/openjarvis/learning/agents/skill_optimizer.py new file mode 100644 index 00000000..e6a5151e --- /dev/null +++ b/src/openjarvis/learning/agents/skill_optimizer.py @@ -0,0 +1,312 @@ +"""SkillOptimizer — per-skill DSPy/GEPA optimization wrapper (Plan 2A). + +Buckets traces by skill name, runs the underlying optimizer on each skill's +bucket, and writes the result as a sidecar overlay file in +``~/.openjarvis/learning/skills//optimized.toml``. + +The actual DSPy/GEPA invocation is done in ``_run_dspy`` / ``_run_gepa``, +which are isolated for easy mocking in tests. In Plan 2A these are +deliberately minimal — they call the existing optimizer modules and extract +the description + few-shot examples. Plan 2B will measure the impact via +benchmarks. +""" + +from __future__ import annotations + +import logging +from dataclasses import dataclass, field +from datetime import datetime, timezone +from pathlib import Path +from typing import Any, Dict, List, Optional + +from openjarvis.core.types import Trace, TraceStep +from openjarvis.skills.manager import SkillManager +from openjarvis.skills.overlay import SkillOverlay, write_overlay + +LOGGER = logging.getLogger(__name__) + + +@dataclass(slots=True) +class SkillOptimizationResult: + """Result of optimizing a single skill.""" + + skill_name: str + status: str # "optimized" | "skipped" | "error" + trace_count: int = 0 + overlay_path: Optional[Path] = None + error: str = "" + + +@dataclass(slots=True) +class _OptimizerOutput: + """Internal: extracted output from the underlying DSPy/GEPA optimizer.""" + + description: str = "" + few_shot: List[Dict[str, str]] = field(default_factory=list) + + +class SkillOptimizer: + """Per-skill optimization wrapper around DSPyAgentOptimizer / GEPA. + + Parameters + ---------- + min_traces_per_skill: + Minimum trace count for a skill to be eligible for optimization. + optimizer: + ``"dspy"`` or ``"gepa"``. Determines which underlying optimizer is + called inside ``_run_dspy`` / ``_run_gepa``. + """ + + def __init__( + self, + *, + min_traces_per_skill: int = 20, + optimizer: str = "dspy", + ) -> None: + self._min_traces = min_traces_per_skill + self._optimizer = optimizer + + # ------------------------------------------------------------------ + # Public API + # ------------------------------------------------------------------ + + def optimize( + self, + trace_store: Any, + skill_manager: SkillManager, + *, + overlay_dir: Optional[Path] = None, + ) -> Dict[str, SkillOptimizationResult]: + """Run the per-skill optimization loop. + + Returns a dict mapping skill name to result. Writes overlay TOML + files for each skill that produced output. + """ + if overlay_dir is None: + # Try config first; fall back to the default tree. + try: + from openjarvis.core.config import load_config + + cfg = load_config() + cfg_dir = getattr( + getattr(cfg.learning, "skills", None), + "overlay_dir", + None, + ) + if cfg_dir: + overlay_dir = Path(cfg_dir).expanduser() + except Exception: + pass + if overlay_dir is None: + overlay_dir = Path( + "~/.openjarvis/learning/skills/" + ).expanduser() + overlay_dir = Path(overlay_dir).expanduser() + + traces = trace_store.list_traces(limit=10000) + buckets = self._bucket_traces_by_skill(traces) + + results: Dict[str, SkillOptimizationResult] = {} + for skill_name, skill_traces in buckets.items(): + if len(skill_traces) < self._min_traces: + results[skill_name] = SkillOptimizationResult( + skill_name=skill_name, + status="skipped", + trace_count=len(skill_traces), + ) + continue + + try: + if self._optimizer == "gepa": + output = self._run_gepa(skill_name, skill_traces) + else: + output = self._run_dspy(skill_name, skill_traces) + except Exception as exc: + LOGGER.warning("Skill optimizer failed for '%s': %s", skill_name, exc) + results[skill_name] = SkillOptimizationResult( + skill_name=skill_name, + status="error", + trace_count=len(skill_traces), + error=str(exc), + ) + continue + + overlay = SkillOverlay( + skill_name=skill_name, + optimizer=self._optimizer, + optimized_at=datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"), + trace_count=len(skill_traces), + description=output.description, + few_shot=list(output.few_shot), + ) + path = write_overlay(overlay, overlay_dir) + results[skill_name] = SkillOptimizationResult( + skill_name=skill_name, + status="optimized", + trace_count=len(skill_traces), + overlay_path=path, + ) + + return results + + # ------------------------------------------------------------------ + # Bucketing + # ------------------------------------------------------------------ + + def _bucket_traces_by_skill(self, traces: List[Trace]) -> Dict[str, List[Trace]]: + """Group traces by the skill names they invoked. + + A trace is added to a skill's bucket once for each unique skill it + invoked. Traces with no skill-tagged tool calls are dropped entirely. + """ + buckets: Dict[str, List[Trace]] = {} + for trace in traces: + skill_names = self._extract_skill_names(trace) + for name in skill_names: + buckets.setdefault(name, []).append(trace) + return buckets + + @staticmethod + def _extract_skill_names(trace: Trace) -> List[str]: + """Return the unique skill names invoked in *trace*'s tool steps.""" + seen: List[str] = [] + steps = getattr(trace, "steps", []) or [] + for step in steps: + metadata: Dict[str, Any] = {} + if isinstance(step, dict): + metadata = step.get("metadata", {}) or {} + elif isinstance(step, TraceStep): + metadata = step.metadata or {} + else: + metadata = getattr(step, "metadata", {}) or {} + skill_name = metadata.get("skill") + if skill_name and skill_name not in seen: + seen.append(skill_name) + return seen + + # ------------------------------------------------------------------ + # Underlying optimizers (mockable in tests) + # ------------------------------------------------------------------ + + def _run_dspy( + self, + skill_name: str, + skill_traces: List[Trace], + ) -> _OptimizerOutput: + """Run DSPyAgentOptimizer on the bucket and extract description + few-shot. + + In Plan 2A this is intentionally simple — it pulls the few highest- + feedback traces as candidate few-shot examples and uses the agent + optimizer's output `system_prompt` as the new skill description if + non-empty. Plan 2B will measure and refine. + """ + from openjarvis.core.config import DSPyOptimizerConfig + from openjarvis.learning.agents.dspy_optimizer import DSPyAgentOptimizer + + cfg = DSPyOptimizerConfig( + min_traces=max(1, self._min_traces), + max_bootstrapped_demos=4, + max_labeled_demos=4, + ) + optimizer = DSPyAgentOptimizer(cfg) + + # Build a tiny in-memory trace store shim for the underlying optimizer + class _BucketStore: + def __init__(self, _traces): + self._traces = _traces + + def list_traces(self, *, limit=100, **_kwargs): + return list(self._traces[:limit]) + + try: + updates = optimizer.optimize(_BucketStore(skill_traces)) or {} + except Exception as exc: + LOGGER.warning( + "DSPyAgentOptimizer raised for '%s' (using empty output): %s", + skill_name, + exc, + ) + updates = {} + + description = str(updates.get("system_prompt", "")) if updates else "" + few_shot_raw = updates.get("few_shot_examples", []) if updates else [] + few_shot: List[Dict[str, str]] = [] + for item in few_shot_raw or []: + if isinstance(item, dict): + few_shot.append( + { + "input": str(item.get("input", "")), + "output": str(item.get("output", "")), + } + ) + + # Fallback: if DSPy produced nothing, derive few-shot from top traces + if not few_shot: + top = sorted( + skill_traces, + key=lambda t: t.feedback if t.feedback is not None else 0.0, + reverse=True, + )[:4] + for tr in top: + if tr.query and tr.result: + few_shot.append({"input": tr.query, "output": tr.result}) + + return _OptimizerOutput(description=description, few_shot=few_shot) + + def _run_gepa( + self, + skill_name: str, + skill_traces: List[Trace], + ) -> _OptimizerOutput: + """Run GEPAAgentOptimizer on the bucket. Same shape as _run_dspy.""" + from openjarvis.core.config import GEPAOptimizerConfig + from openjarvis.learning.agents.gepa_optimizer import GEPAAgentOptimizer + + cfg = GEPAOptimizerConfig(min_traces=max(1, self._min_traces)) + optimizer = GEPAAgentOptimizer(cfg) + + class _BucketStore: + def __init__(self, _traces): + self._traces = _traces + + def list_traces(self, *, limit=100, **_kwargs): + return list(self._traces[:limit]) + + try: + updates = optimizer.optimize(_BucketStore(skill_traces)) or {} + except Exception as exc: + LOGGER.warning( + "GEPAAgentOptimizer raised for '%s' (using empty output): %s", + skill_name, + exc, + ) + updates = {} + + description = str(updates.get("system_prompt", "")) if updates else "" + + # Parse GEPA's few_shot_examples output (matches DSPy path semantics) + few_shot_raw = updates.get("few_shot_examples", []) if updates else [] + few_shot: List[Dict[str, str]] = [] + for item in few_shot_raw or []: + if isinstance(item, dict): + few_shot.append( + { + "input": str(item.get("input", "")), + "output": str(item.get("output", "")), + } + ) + + # Fallback: if GEPA produced nothing, derive few-shot from top traces + if not few_shot: + top = sorted( + skill_traces, + key=lambda t: t.feedback if t.feedback is not None else 0.0, + reverse=True, + )[:4] + for tr in top: + if tr.query and tr.result: + few_shot.append({"input": tr.query, "output": tr.result}) + return _OptimizerOutput(description=description, few_shot=few_shot) + + +__all__ = ["SkillOptimizer", "SkillOptimizationResult"] diff --git a/src/openjarvis/learning/learning_orchestrator.py b/src/openjarvis/learning/learning_orchestrator.py index 6c77ad0a..bda1719c 100644 --- a/src/openjarvis/learning/learning_orchestrator.py +++ b/src/openjarvis/learning/learning_orchestrator.py @@ -98,6 +98,32 @@ class LearningOrchestrator: "timestamp": time.time(), } + # 0. Skill optimization (Plan 2A C2) — runs INDEPENDENTLY of the + # routing/agent SFT pipeline. Skills are tagged via trace metadata + # rather than mined as SFT pairs, so they can be optimized even when + # there's no other training data available. + try: + from openjarvis.core.config import load_config + + cfg = load_config() + skills_cfg = getattr(cfg.learning, "skills", None) + if skills_cfg is not None and skills_cfg.auto_optimize: + skill_results = self._maybe_optimize_skills( + auto_optimize=True, + optimizer=skills_cfg.optimizer, + min_traces_per_skill=skills_cfg.min_traces_per_skill, + ) + if skill_results is not None: + result["skill_optimization"] = { + name: { + "status": r.status, + "trace_count": r.trace_count, + } + for name, r in skill_results.items() + } + except Exception as exc: + logger.warning("Skill auto-optimize probe failed: %s", exc) + # 1. Mine training data from traces sft_pairs = self._miner.extract_sft_pairs(agent=agent_id) routing_pairs = self._miner.extract_routing_pairs(agent=agent_id) @@ -194,5 +220,37 @@ class LearningOrchestrator: logger.warning("LoRA training failed: %s", exc) return {"status": "error", "reason": str(exc)} + def _maybe_optimize_skills( + self, + *, + auto_optimize: bool = False, + optimizer: str = "dspy", + min_traces_per_skill: int = 20, + ) -> Optional[dict]: + """Optionally run the skill optimizer. + + Called from :meth:`run` when ``learning.skills.auto_optimize`` is + true. Returns the per-skill result dict or ``None`` if disabled. + """ + if not auto_optimize: + return None + try: + from openjarvis.core.events import EventBus + from openjarvis.learning.agents.skill_optimizer import SkillOptimizer + from openjarvis.skills.manager import SkillManager + + mgr = SkillManager(bus=EventBus()) + mgr.discover() + opt = SkillOptimizer( + min_traces_per_skill=min_traces_per_skill, + optimizer=optimizer, + ) + return opt.optimize(self._trace_store, mgr) + except Exception as exc: + import logging + + logging.getLogger(__name__).warning("Skill auto-optimize failed: %s", exc) + return None + __all__ = ["LearningOrchestrator"] diff --git a/src/openjarvis/prompt/builder.py b/src/openjarvis/prompt/builder.py index cfd38f5a..1b562fb0 100644 --- a/src/openjarvis/prompt/builder.py +++ b/src/openjarvis/prompt/builder.py @@ -17,6 +17,9 @@ class SystemPromptBuilder: skill_index: Optional[List[Tuple[str, str]]] = None, session_context: Optional[str] = None, previous_state: Optional[str] = None, + skill_catalog_xml: Optional[str] = None, + skill_few_shot: Optional[List[str]] = None, + skill_few_shot_examples: Optional[List[str]] = None, ) -> None: self._agent_template = agent_template self._mf_config = memory_files_config or MemoryFilesConfig() @@ -24,6 +27,12 @@ class SystemPromptBuilder: self._skill_index = skill_index or [] self._session_context = session_context self._previous_state = previous_state + self._skill_catalog_xml = skill_catalog_xml + # Allow either name; skill_few_shot_examples is the Plan 2A canonical name. + if skill_few_shot_examples is not None: + self._skill_few_shot = list(skill_few_shot_examples) + else: + self._skill_few_shot = list(skill_few_shot or []) self._frozen_prefix: Optional[str] = None def build(self) -> str: @@ -57,7 +66,10 @@ class SystemPromptBuilder: ) if user: sections.append(f"## User Profile\n\n{user}") - if self._skill_index: + # XML skill catalog (preferred over legacy markdown list) + if self._skill_catalog_xml: + sections.append("## Available Skills\n\n" + self._skill_catalog_xml) + elif self._skill_index: skill_lines = [] for name, desc in self._skill_index: truncated = desc[: self._sp_config.skill_desc_max_chars] @@ -65,6 +77,9 @@ class SystemPromptBuilder: truncated = truncated[:-3] + "..." skill_lines.append(f"- **{name}**: {truncated}") sections.append("## Available Skills\n\n" + "\n".join(skill_lines)) + if self._skill_few_shot: + examples = "\n\n".join(self._skill_few_shot) + sections.append("## Skill Examples\n\n" + examples) return "\n\n".join(sections) def _load_file(self, path_str: str, max_chars: int) -> str: diff --git a/src/openjarvis/skills/__init__.py b/src/openjarvis/skills/__init__.py index 02a46be9..6e5b567c 100644 --- a/src/openjarvis/skills/__init__.py +++ b/src/openjarvis/skills/__init__.py @@ -1,8 +1,46 @@ """Skill system — reusable multi-tool compositions.""" -from openjarvis.skills.executor import SkillExecutor -from openjarvis.skills.loader import load_skill +from openjarvis.skills.dependency import ( + DependencyCycleError, + DepthExceededError, + build_dependency_graph, + compute_capability_union, + validate_dependencies, +) +from openjarvis.skills.executor import SkillExecutor, SkillResult +from openjarvis.skills.importer import ImportResult, SkillImporter +from openjarvis.skills.loader import ( + discover_skills, + load_skill, + load_skill_directory, + load_skill_markdown, +) +from openjarvis.skills.manager import SkillManager +from openjarvis.skills.parser import SkillParseError, SkillParser from openjarvis.skills.tool_adapter import SkillTool +from openjarvis.skills.tool_translator import TOOL_TRANSLATION, ToolTranslator from openjarvis.skills.types import SkillManifest, SkillStep -__all__ = ["SkillExecutor", "SkillManifest", "SkillStep", "SkillTool", "load_skill"] +__all__ = [ + "DependencyCycleError", + "DepthExceededError", + "ImportResult", + "SkillExecutor", + "SkillImporter", + "SkillManager", + "SkillManifest", + "SkillParseError", + "SkillParser", + "SkillResult", + "SkillStep", + "SkillTool", + "TOOL_TRANSLATION", + "ToolTranslator", + "build_dependency_graph", + "compute_capability_union", + "discover_skills", + "load_skill", + "load_skill_directory", + "load_skill_markdown", + "validate_dependencies", +] diff --git a/src/openjarvis/skills/dependency.py b/src/openjarvis/skills/dependency.py new file mode 100644 index 00000000..601f12bc --- /dev/null +++ b/src/openjarvis/skills/dependency.py @@ -0,0 +1,186 @@ +"""Dependency graph: cycle detection, topological sort, capability union.""" + +from __future__ import annotations + +from collections import deque +from typing import TYPE_CHECKING, Dict, List, Set + +if TYPE_CHECKING: + from openjarvis.skills.types import SkillManifest + + +class DependencyCycleError(Exception): + """Raised when a cycle is detected in the skill dependency graph.""" + + +class DepthExceededError(Exception): + """Raised when the dependency depth exceeds the configured maximum.""" + + +def build_dependency_graph(skills: Dict[str, "SkillManifest"]) -> Dict[str, Set[str]]: + """Build a directed graph of skill dependencies. + + Each key maps to the set of skill names it directly depends on, drawn from + both the ``depends`` field and any ``skill_name`` references in steps. + Only edges pointing to skills that are present in the *skills* dict are + included (unknown refs are silently dropped so that the graph stays clean + for topological analysis). + + Args: + skills: Mapping of skill name → SkillManifest. + + Returns: + Dict mapping each skill name to the set of its direct dependencies. + """ + graph: Dict[str, Set[str]] = {name: set() for name in skills} + + for name, manifest in skills.items(): + # Explicit depends list + for dep in manifest.depends: + if dep in skills: + graph[name].add(dep) + + # Implicit deps via step skill_name references + for step in manifest.steps: + if step.skill_name and step.skill_name in skills: + graph[name].add(step.skill_name) + + return graph + + +def validate_dependencies( + skills: Dict[str, "SkillManifest"], + *, + max_depth: int = 5, +) -> List[str]: + """Validate the skill dependency graph and return a topological ordering. + + Uses Kahn's algorithm for topological sort and cycle detection. After a + valid ordering is found, a DFS checks that no skill's transitive dependency + chain exceeds *max_depth*. Dependencies that reference unknown skills are + silently skipped. + + Args: + skills: Mapping of skill name → SkillManifest. + max_depth: Maximum allowed dependency chain depth (default 5). + + Returns: + List of skill names in valid topological order (dependencies first). + + Raises: + DependencyCycleError: If the graph contains a cycle. + DepthExceededError: If any skill's dependency chain depth exceeds max_depth. + """ + graph = build_dependency_graph(skills) + + # --- Kahn's algorithm --- + in_degree: Dict[str, int] = {name: 0 for name in graph} + for name in graph: + for dep in graph[name]: + in_degree[dep] # ensure present (already is, but be explicit) + + # Count how many skills depend on each node + reverse: Dict[str, Set[str]] = {name: set() for name in graph} + for name, deps in graph.items(): + for dep in deps: + reverse[dep].add(name) + # re-compute in_degree from scratch below + + # Compute in-degree: number of dependencies each skill has (within the graph) + in_degree = {name: len(deps) for name, deps in graph.items()} + + queue: deque[str] = deque(name for name, deg in in_degree.items() if deg == 0) + order: List[str] = [] + + while queue: + node = queue.popleft() + order.append(node) + for dependent in reverse[node]: + in_degree[dependent] -= 1 + if in_degree[dependent] == 0: + queue.append(dependent) + + if len(order) != len(graph): + raise DependencyCycleError( + "Cycle detected in skill dependency graph. " + f"Skills involved: {set(graph) - set(order)}" + ) + + # --- Depth enforcement via DFS --- + depth_cache: Dict[str, int] = {} + + def _depth(name: str, visiting: Set[str]) -> int: + if name in depth_cache: + return depth_cache[name] + deps = graph.get(name, set()) + if not deps: + depth_cache[name] = 1 + return 1 + visiting.add(name) + max_child = max(_depth(d, visiting) for d in deps) + visiting.discard(name) + result = max_child + 1 + depth_cache[name] = result + return result + + for name in graph: + d = _depth(name, set()) + if d > max_depth: + raise DepthExceededError( + f"Skill '{name}' depth {d} exceeds max_depth={max_depth}" + ) + + return order + + +def compute_capability_union( + skill_name: str, + skills: Dict[str, "SkillManifest"], +) -> List[str]: + """Compute the union of all required_capabilities transitively needed by a skill. + + Performs a DFS over the dependency graph, collecting ``required_capabilities`` + from the skill itself and all of its transitive dependencies. Duplicates are + removed while preserving a deterministic order (first-seen wins). + + Args: + skill_name: Name of the root skill to start from. + skills: Mapping of skill name → SkillManifest. + + Returns: + Deduplicated list of capability strings. Returns an empty list if the + skill is not found. + """ + if skill_name not in skills: + return [] + + graph = build_dependency_graph(skills) + seen_skills: Set[str] = set() + seen_caps: Set[str] = set() + caps_ordered: List[str] = [] + + def _dfs(name: str) -> None: + if name in seen_skills: + return + seen_skills.add(name) + manifest = skills.get(name) + if manifest is None: + return + for cap in manifest.required_capabilities: + if cap not in seen_caps: + seen_caps.add(cap) + caps_ordered.append(cap) + for dep in graph.get(name, set()): + _dfs(dep) + + _dfs(skill_name) + return caps_ordered + + +__all__ = [ + "DependencyCycleError", + "DepthExceededError", + "build_dependency_graph", + "validate_dependencies", + "compute_capability_union", +] diff --git a/src/openjarvis/skills/executor.py b/src/openjarvis/skills/executor.py index fbecf08e..6627de60 100644 --- a/src/openjarvis/skills/executor.py +++ b/src/openjarvis/skills/executor.py @@ -5,7 +5,7 @@ from __future__ import annotations import json import re from dataclasses import dataclass, field -from typing import Any, Dict, List, Optional +from typing import Any, Callable, Dict, List, Optional from openjarvis.core.events import EventBus, EventType from openjarvis.core.types import ToolCall, ToolResult @@ -21,6 +21,10 @@ class SkillResult: context: Dict[str, Any] = field(default_factory=dict) +# Resolver callback: given a skill name and the current context, returns a SkillResult. +SkillResolver = Callable[[str, Dict[str, Any]], SkillResult] + + class SkillExecutor: """Execute a skill manifest step-by-step. @@ -36,6 +40,11 @@ class SkillExecutor: ) -> None: self._tool_executor = tool_executor self._bus = bus + self._skill_resolver: Optional[SkillResolver] = None + + def set_skill_resolver(self, resolver: SkillResolver) -> None: + """Register a callback used to delegate ``skill_name`` steps.""" + self._skill_resolver = resolver def run( self, @@ -54,25 +63,34 @@ class SkillExecutor: ) for i, step in enumerate(manifest.steps): + step_id = step.tool_name or step.skill_name + # Render template try: rendered = self._render_template(step.arguments_template, ctx) except Exception as exc: result = ToolResult( - tool_name=step.tool_name, + tool_name=step_id, content=f"Template rendering error: {exc}", success=False, ) all_results.append(result) break - # Execute - tool_call = ToolCall( - id=f"skill_{manifest.name}_{i}", - name=step.tool_name, - arguments=rendered, - ) - result = self._tool_executor.execute(tool_call) + if step.skill_name: + # Delegate to sub-skill resolver + result = self._run_sub_skill( + step.skill_name, rendered, ctx, manifest.name, i + ) + else: + # Execute via tool executor + tool_call = ToolCall( + id=f"skill_{manifest.name}_{i}", + name=step.tool_name, + arguments=rendered, + ) + result = self._tool_executor.execute(tool_call) + all_results.append(result) if not result.success: @@ -97,6 +115,54 @@ class SkillExecutor: context=ctx, ) + def _run_sub_skill( + self, + skill_name: str, + rendered_args: str, + parent_ctx: Dict[str, Any], + parent_skill: str, + step_index: int, + ) -> ToolResult: + """Invoke the skill resolver and convert its result to a ToolResult.""" + if self._skill_resolver is None: + return ToolResult( + tool_name=skill_name, + content=f"No skill resolver registered for sub-skill '{skill_name}'", + success=False, + ) + + # Parse rendered args and merge into a copy of the parent context + try: + args: Dict[str, Any] = json.loads(rendered_args) + except json.JSONDecodeError: + args = {} + + child_ctx = {**parent_ctx, **args} + + sub_result: SkillResult = self._skill_resolver(skill_name, child_ctx) + + # Expose the final context value under the first output_key, or the + # last step's content, as the synthetic "content" of this ToolResult. + content: Any = "" + if sub_result.context: + # Return the last stored value from the child's context that is + # not already in the parent context (i.e. the output of the sub-skill). + new_keys = [k for k in sub_result.context if k not in parent_ctx] + if new_keys: + content = sub_result.context[new_keys[-1]] + else: + content = ( + list(sub_result.context.values())[-1] if sub_result.context else "" + ) + elif sub_result.step_results: + content = sub_result.step_results[-1].content + + return ToolResult( + tool_name=skill_name, + content=content, + success=sub_result.success, + ) + @staticmethod def _render_template(template: str, ctx: Dict[str, Any]) -> str: """Simple {key} placeholder rendering.""" @@ -111,4 +177,4 @@ class SkillExecutor: return re.sub(r"\{(\w+)\}", _replace, template) -__all__ = ["SkillExecutor", "SkillResult"] +__all__ = ["SkillExecutor", "SkillResolver", "SkillResult"] diff --git a/src/openjarvis/skills/importer.py b/src/openjarvis/skills/importer.py new file mode 100644 index 00000000..4137f4bb --- /dev/null +++ b/src/openjarvis/skills/importer.py @@ -0,0 +1,195 @@ +"""SkillImporter — install ResolvedSkill instances into ~/.openjarvis/skills/. + +Steps performed by ``import_skill``: + +1. Parse the source SKILL.md through SkillParser (strict + tolerant). +2. Translate tool references in the markdown body via ToolTranslator. +3. Decide on scripts (default-skip; opt-in via with_scripts=True). +4. Write to disk at ///: + - translated SKILL.md + - references/, assets/, templates/ (always copied) + - scripts/ (only if approved) + - .source metadata file +5. Return ImportResult with status, warnings, translated/missing tools. +""" + +from __future__ import annotations + +import shutil +from dataclasses import dataclass, field +from datetime import datetime, timezone +from pathlib import Path +from typing import List + +import yaml + +from openjarvis.skills.parser import SkillParser +from openjarvis.skills.sources.base import ResolvedSkill +from openjarvis.skills.tool_translator import ToolTranslator + +# Subdirectories of a skill that are always copied (never gated by --with-scripts) +COPIED_SUBDIRS = ("references", "assets", "templates") + + +@dataclass(slots=True) +class ImportResult: + """Result of importing a single skill.""" + + success: bool = True + skipped: bool = False + target_path: Path | None = None + translated_tools: List[str] = field(default_factory=list) + untranslated_tools: List[str] = field(default_factory=list) + scripts_imported: bool = False + warnings: List[str] = field(default_factory=list) + + +class SkillImporter: + """Install resolved skills into the user skills directory.""" + + def __init__( + self, + parser: SkillParser, + tool_translator: ToolTranslator, + target_root: Path | None = None, + ) -> None: + self._parser = parser + self._translator = tool_translator + if target_root is None: + target_root = Path("~/.openjarvis/skills/").expanduser() + self._target_root = Path(target_root) + + def import_skill( + self, + resolved: ResolvedSkill, + *, + with_scripts: bool = False, + force: bool = False, + ) -> ImportResult: + """Install *resolved* into ``///``. + + Returns an :class:`ImportResult` with status, paths, translated + tools, untranslated tools, and warnings. + """ + result = ImportResult() + target_dir = self._target_root / resolved.source / resolved.name + result.target_path = target_dir + + # Skip if already installed and force is False + if target_dir.exists() and not force: + result.skipped = True + result.warnings.append( + f"Skill already installed at {target_dir} (use force=True to overwrite)" + ) + return result + + # 1. Parse source SKILL.md + source_md = resolved.path / "SKILL.md" + if not source_md.exists(): + source_md = resolved.path / "skill.md" + if not source_md.exists(): + result.success = False + result.warnings.append(f"No SKILL.md found in {resolved.path}") + return result + + try: + frontmatter, body = self._read_skill_md(source_md) + self._parser.parse_frontmatter(frontmatter, markdown_content=body) + except Exception as exc: + result.success = False + result.warnings.append(f"Parse error: {exc}") + return result + + # 2. Translate tool references + translated_body, untranslated = self._translator.translate_markdown(body) + result.untranslated_tools = untranslated + # Compute the list of translations actually applied + applied: List[str] = [] + for ext, internal in self._translator._table.items(): + if ext in body and ext not in translated_body: + applied.append(f"{ext}->{internal}") + result.translated_tools = applied + + # 3. Write the target directory + if target_dir.exists(): + shutil.rmtree(target_dir) + target_dir.mkdir(parents=True) + + # 3a. Translated SKILL.md + new_md = self._render_skill_md(frontmatter, translated_body) + (target_dir / "SKILL.md").write_text(new_md, encoding="utf-8") + + # 3b. Always-copied subdirs + for subdir in COPIED_SUBDIRS: + src_sub = resolved.path / subdir + if src_sub.exists(): + shutil.copytree(src_sub, target_dir / subdir) + + # 3c. Scripts (gated by with_scripts) + scripts_src = resolved.path / "scripts" + if scripts_src.exists() and with_scripts: + shutil.copytree(scripts_src, target_dir / "scripts") + result.scripts_imported = True + elif scripts_src.exists(): + result.warnings.append( + "Skipped scripts/ directory (use with_scripts=True to import)" + ) + + # 4. Write .source metadata + self._write_source_metadata(target_dir, resolved, result) + + return result + + # ------------------------------------------------------------------ + # Helpers + # ------------------------------------------------------------------ + + def _read_skill_md(self, path: Path) -> tuple[dict, str]: + """Parse a SKILL.md file into (frontmatter dict, markdown body).""" + raw = path.read_text(encoding="utf-8") + if not raw.startswith("---"): + return {}, raw + rest = raw[3:].lstrip("\n") + end = rest.find("\n---") + if end == -1: + return {}, raw + fm_text = rest[:end] + body = rest[end + 4 :].lstrip("\n") + try: + fm = yaml.safe_load(fm_text) + if not isinstance(fm, dict): + fm = {} + except yaml.YAMLError: + fm = {} + return fm, body + + def _render_skill_md(self, frontmatter: dict, body: str) -> str: + """Re-serialize a SKILL.md file from frontmatter dict + body.""" + fm_text = yaml.safe_dump(frontmatter, sort_keys=False, default_flow_style=False) + return f"---\n{fm_text}---\n\n{body}" + + def _write_source_metadata( + self, + target_dir: Path, + resolved: ResolvedSkill, + result: ImportResult, + ) -> None: + """Write the .source TOML provenance file.""" + installed_at = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + translated_str = ", ".join(f'"{t}"' for t in result.translated_tools) + missing_str = ", ".join(f'"{t}"' for t in result.untranslated_tools) + scripts_lower = "true" if result.scripts_imported else "false" + + content = ( + f'source = "{resolved.source}:{resolved.name}"\n' + f'commit = "{resolved.commit}"\n' + f'category = "{resolved.category}"\n' + f'installed_at = "{installed_at}"\n' + f"translated_tools = [{translated_str}]\n" + f"missing_tools = [{missing_str}]\n" + f"scripts_imported = {scripts_lower}\n" + ) + (target_dir / ".source").write_text(content, encoding="utf-8") + + +__all__ = ["ImportResult", "SkillImporter"] diff --git a/src/openjarvis/skills/index.py b/src/openjarvis/skills/index.py new file mode 100644 index 00000000..548ef8c6 --- /dev/null +++ b/src/openjarvis/skills/index.py @@ -0,0 +1,99 @@ +"""Skill index — load and search a remote/local skill catalog.""" + +from __future__ import annotations + +from dataclasses import dataclass, field +from pathlib import Path +from typing import Optional + +try: + import tomllib +except ModuleNotFoundError: + import tomli as tomllib # type: ignore[no-redef] + + +@dataclass(slots=True) +class SkillIndexEntry: + """A single entry in the skill index catalog.""" + + name: str + version: str + description: str + author: str + source: str + sha256: str + tags: list[str] = field(default_factory=list) + required_capabilities: list[str] = field(default_factory=list) + + +class SkillIndex: + """Catalog of available skills loaded from an ``index.toml`` file. + + ``index.toml`` format:: + + [[skills]] + name = "research" + version = "0.1.0" + description = "Research a topic" + author = "openjarvis" + source = "github.com/openjarvis/skills/research" + sha256 = "abc123" + tags = ["research"] + required_capabilities = ["network:fetch"] + """ + + def __init__(self, index_dir: str | Path) -> None: + self._index_dir = Path(index_dir) + self._entries: dict[str, SkillIndexEntry] = {} + self._load() + + def _load(self) -> None: + """Parse ``index.toml`` from the index directory if it exists.""" + index_file = self._index_dir / "index.toml" + if not index_file.exists(): + return + + with open(index_file, "rb") as fh: + data = tomllib.load(fh) + + for skill_data in data.get("skills", []): + entry = SkillIndexEntry( + name=skill_data.get("name", ""), + version=skill_data.get("version", "0.1.0"), + description=skill_data.get("description", ""), + author=skill_data.get("author", ""), + source=skill_data.get("source", ""), + sha256=skill_data.get("sha256", ""), + tags=skill_data.get("tags", []), + required_capabilities=skill_data.get("required_capabilities", []), + ) + self._entries[entry.name] = entry + + @property + def entries(self) -> dict[str, SkillIndexEntry]: + """Return all entries keyed by skill name.""" + return self._entries + + def search(self, query: str) -> list[SkillIndexEntry]: + """Return entries whose name, description, or tags match *query*. + + Matching is case-insensitive substring search across name, description, + and each tag value. + """ + q = query.lower() + results: list[SkillIndexEntry] = [] + for entry in self._entries.values(): + if ( + q in entry.name.lower() + or q in entry.description.lower() + or any(q in tag.lower() for tag in entry.tags) + ): + results.append(entry) + return results + + def get(self, name: str) -> Optional[SkillIndexEntry]: + """Look up a skill by exact name, returning ``None`` if not found.""" + return self._entries.get(name) + + +__all__ = ["SkillIndexEntry", "SkillIndex"] diff --git a/src/openjarvis/skills/loader.py b/src/openjarvis/skills/loader.py index 20119052..42b9cdbf 100644 --- a/src/openjarvis/skills/loader.py +++ b/src/openjarvis/skills/loader.py @@ -5,6 +5,8 @@ from __future__ import annotations from pathlib import Path from typing import Optional +import yaml + from openjarvis.skills.types import SkillManifest, SkillStep try: @@ -12,6 +14,34 @@ try: except ModuleNotFoundError: import tomli as tomllib # type: ignore[no-redef] +import logging + +LOGGER = logging.getLogger(__name__) + + +def _read_source_metadata(path: Path) -> dict: + """Read .source TOML file if present, returning the parsed dict. + + Returns an empty dict if the file is missing or malformed. Never + raises — bad sidecar files should never break skill loading. Logs + a warning when a malformed file is encountered so users can debug + why their imported source provenance is missing. + """ + source_path = path / ".source" + if not source_path.exists(): + return {} + try: + with open(source_path, "rb") as fh: + return tomllib.load(fh) + except Exception as exc: + LOGGER.warning( + "Malformed .source sidecar at %s (skill source provenance " + "will be missing): %s", + source_path, + exc, + ) + return {} + def load_skill( path: str | Path, @@ -53,7 +83,8 @@ def load_skill( for step_data in skill_data.get("steps", []): steps.append( SkillStep( - tool_name=step_data["tool_name"], + tool_name=step_data.get("tool_name", ""), + skill_name=step_data.get("skill_name", ""), arguments_template=step_data.get("arguments_template", "{}"), output_key=step_data.get("output_key", ""), ) @@ -68,6 +99,10 @@ def load_skill( required_capabilities=skill_data.get("required_capabilities", []), signature=skill_data.get("signature", ""), metadata=skill_data.get("metadata", {}), + tags=skill_data.get("tags", []), + depends=skill_data.get("depends", []), + user_invocable=skill_data.get("user_invocable", True), + disable_model_invocation=skill_data.get("disable_model_invocation", False), ) # Verify signature if requested @@ -108,18 +143,177 @@ def load_skill( return manifest +def load_skill_markdown(path: str | Path) -> SkillManifest: + """Load a skill manifest from a SKILL.md file via SkillParser. + + Parses YAML frontmatter (between ``---`` delimiters) and the markdown + body, then runs them through :class:`SkillParser` for strict validation + and tolerant field mapping. + """ + from openjarvis.skills.parser import SkillParseError, SkillParser + + path = Path(path) + raw = path.read_text(encoding="utf-8") + + frontmatter: dict = {} + markdown_content = raw + + if raw.startswith("---"): + rest = raw[3:] + if rest.startswith("\n"): + rest = rest[1:] + end_idx = rest.find("\n---") + if end_idx != -1: + yaml_block = rest[:end_idx] + try: + frontmatter = yaml.safe_load(yaml_block) or {} + except yaml.YAMLError: + frontmatter = {} + after = rest[end_idx + 4 :] + if after.startswith("\n"): + after = after[1:] + markdown_content = after + + # Default name to file stem if missing (preserves legacy behavior) + if "name" not in frontmatter: + frontmatter["name"] = path.stem + if "description" not in frontmatter: + frontmatter["description"] = frontmatter.get("name", path.stem) + + parser = SkillParser() + try: + return parser.parse_frontmatter(frontmatter, markdown_content=markdown_content) + except SkillParseError: + # Legacy fallback — build manifest directly without strict validation + return SkillManifest( + name=str(frontmatter.get("name", path.stem)), + description=str(frontmatter.get("description", "")), + version=str(frontmatter.get("version", "0.1.0")), + author=str(frontmatter.get("author", "")), + tags=list(frontmatter.get("tags", []) or []), + required_capabilities=list( + frontmatter.get("required_capabilities", []) or [] + ), + user_invocable=bool(frontmatter.get("user_invocable", True)), + disable_model_invocation=bool( + frontmatter.get("disable_model_invocation", False) + ), + markdown_content=markdown_content, + ) + + +def load_skill_directory(path: str | Path) -> SkillManifest: + """Load a skill from a directory containing ``skill.toml`` and/or ``SKILL.md``. + + - If only ``skill.toml`` is present the manifest is loaded from TOML. + - If only ``SKILL.md`` is present the manifest is loaded from markdown. + - If both are present the TOML manifest takes precedence for structured + fields and ``markdown_content`` is merged in from the markdown file. + - If neither is present a ``FileNotFoundError`` is raised. + """ + path = Path(path) + toml_path = path / "skill.toml" + md_path = path / "SKILL.md" + + has_toml = toml_path.exists() + has_md = md_path.exists() + + if not has_toml and not has_md: + raise FileNotFoundError( + f"No skill.toml or SKILL.md found in directory '{path}'" + ) + + if has_toml: + manifest = load_skill(toml_path) + if has_md: + md_manifest = load_skill_markdown(md_path) + # Merge markdown content into the TOML-sourced manifest + manifest = SkillManifest( + name=manifest.name, + version=manifest.version, + description=manifest.description, + author=manifest.author, + steps=manifest.steps, + required_capabilities=manifest.required_capabilities, + signature=manifest.signature, + metadata=manifest.metadata, + tags=manifest.tags, + depends=manifest.depends, + user_invocable=manifest.user_invocable, + disable_model_invocation=manifest.disable_model_invocation, + markdown_content=md_manifest.markdown_content, + ) + else: + # Only markdown present + manifest = load_skill_markdown(md_path) + + # Promote .source file's source field into manifest.metadata.openjarvis.source + source_data = _read_source_metadata(path) + if source_data: + # Extract just the source name (e.g. "hermes" from "hermes:apple-notes") + source_str = source_data.get("source", "") + source_name = source_str.partition(":")[0] if source_str else "" + if source_name: + new_metadata = dict(manifest.metadata) if manifest.metadata else {} + oj = dict(new_metadata.get("openjarvis", {}) or {}) + oj["source"] = source_name + new_metadata["openjarvis"] = oj + manifest.metadata = new_metadata + + return manifest + + def discover_skills(directory: str | Path) -> list[SkillManifest]: - """Scan directory for TOML skill files and load them.""" + """Scan a directory for skill definitions and load them. + + Handles three layouts: + - Flat ``*.toml`` files directly inside *directory*. + - Skill directories: ``//{skill.toml,SKILL.md}``. + - Sourced layout: ``///{skill.toml,SKILL.md}``. + """ directory = Path(directory).expanduser() if not directory.exists(): return [] - manifests = [] + + manifests: list[SkillManifest] = [] + + # Flat *.toml files at the top level for toml_file in sorted(directory.glob("*.toml")): try: manifests.append(load_skill(toml_file)) except Exception: continue + + # Walk one or two levels deep looking for skill packages + for child in sorted(directory.iterdir()): + if not child.is_dir(): + continue + # Direct skill package: /{skill.toml,SKILL.md} + if (child / "skill.toml").exists() or (child / "SKILL.md").exists(): + try: + manifests.append(load_skill_directory(child)) + except Exception: + continue + continue + + # Sourced layout: //{skill.toml,SKILL.md} + for grandchild in sorted(child.iterdir()): + if not grandchild.is_dir(): + continue + if (grandchild / "skill.toml").exists() or ( + grandchild / "SKILL.md" + ).exists(): + try: + manifests.append(load_skill_directory(grandchild)) + except Exception: + continue + return manifests -__all__ = ["load_skill", "discover_skills"] +__all__ = [ + "load_skill", + "load_skill_markdown", + "load_skill_directory", + "discover_skills", +] diff --git a/src/openjarvis/skills/manager.py b/src/openjarvis/skills/manager.py new file mode 100644 index 00000000..487de511 --- /dev/null +++ b/src/openjarvis/skills/manager.py @@ -0,0 +1,396 @@ +"""SkillManager — coordinates skill discovery, catalog, tool wrapping, and execution.""" + +from __future__ import annotations + +import html +from pathlib import Path +from typing import Any, Dict, List, Optional + +from openjarvis.core.events import EventBus +from openjarvis.skills.dependency import validate_dependencies +from openjarvis.skills.executor import SkillExecutor, SkillResult +from openjarvis.skills.loader import discover_skills +from openjarvis.skills.tool_adapter import SkillTool +from openjarvis.skills.types import SkillManifest +from openjarvis.tools._stubs import BaseTool, ToolExecutor + + +class SkillManager: + """Coordinate skill discovery, resolution, catalog generation, and execution. + + Parameters + ---------- + bus: + Event bus for publishing lifecycle events. + capability_policy: + Optional capability policy passed through to tool executors. + """ + + def __init__( + self, + bus: EventBus, + *, + capability_policy: Optional[Any] = None, + overlay_dir: Optional[Path] = None, + ) -> None: + self._bus = bus + self._capability_policy = capability_policy + self._skills: Dict[str, SkillManifest] = {} + self._tool_executor: Optional[ToolExecutor] = None + if overlay_dir is None: + # Try to read from config first; fall back to the default + # ~/.openjarvis/learning/skills/ if config can't be loaded. + try: + from openjarvis.core.config import load_config + + cfg = load_config() + cfg_dir = getattr( + getattr(cfg.learning, "skills", None), + "overlay_dir", + None, + ) + if cfg_dir: + overlay_dir = Path(cfg_dir).expanduser() + except Exception: + pass + if overlay_dir is None: + overlay_dir = Path( + "~/.openjarvis/learning/skills/" + ).expanduser() + self._overlay_dir = Path(overlay_dir).expanduser() + + # ------------------------------------------------------------------ + # Discovery + # ------------------------------------------------------------------ + + def discover(self, paths: Optional[List[Path]] = None) -> None: + """Scan directories in order and register skills. + + First-seen name wins (workspace path listed first = highest precedence). + After loading, the full dependency graph is validated and any + sidecar overlays in ``overlay_dir`` are applied to discovered skills. + + Parameters + ---------- + paths: + Directories to scan. If *None* or empty, no skills are loaded + from disk — but ``_load_overlays()`` still runs (in case the + caller had previously seeded ``self._skills`` directly). + """ + if paths: + for directory in paths: + manifests = discover_skills(directory) + for manifest in manifests: + # First-seen wins: do not overwrite an already-registered skill + if manifest.name not in self._skills: + self._skills[manifest.name] = manifest + + # Validate the dependency graph after loading skills + if self._skills: + validate_dependencies(self._skills) + + # Load optimization overlays (Plan 2A) — always runs, even when no + # paths are provided, so callers can apply overlays to skills loaded + # via other means. + self._load_overlays() + + def _load_overlays(self) -> None: + """Apply optimization overlays to discovered skills. + + For each skill, look for ``//optimized.toml``. + If present, override the manifest description and stash few-shot + examples under ``manifest.metadata.openjarvis.few_shot``. + + Bad overlays are silently ignored — they should not break discovery. + """ + from openjarvis.skills.overlay import SkillOverlayLoader + + loader = SkillOverlayLoader(self._overlay_dir) + for name, manifest in self._skills.items(): + overlay = loader.load(name) + if overlay is None: + continue + if overlay.description: + manifest.description = overlay.description + if overlay.few_shot: + new_metadata = dict(manifest.metadata) if manifest.metadata else {} + oj = dict(new_metadata.get("openjarvis", {}) or {}) + oj["few_shot"] = list(overlay.few_shot) + new_metadata["openjarvis"] = oj + manifest.metadata = new_metadata + + # ------------------------------------------------------------------ + # Resolve / introspect + # ------------------------------------------------------------------ + + def resolve(self, name: str) -> SkillManifest: + """Return the manifest for a skill by name. + + Raises + ------ + KeyError + If *name* is not registered. + """ + if name not in self._skills: + raise KeyError(f"Skill '{name}' not found") + return self._skills[name] + + def skill_names(self) -> List[str]: + """Return the names of all registered skills.""" + return list(self._skills.keys()) + + # ------------------------------------------------------------------ + # Tool wrapping + # ------------------------------------------------------------------ + + def get_skill_tools( + self, *, tool_executor: Optional[ToolExecutor] = None + ) -> List[BaseTool]: + """Wrap each registered skill as a :class:`SkillTool` (a :class:`BaseTool`). + + Parameters + ---------- + tool_executor: + Tool executor to use when running skill pipelines. Falls back to + the one set via :meth:`set_tool_executor` if not provided here. + + Returns + ------- + list[BaseTool] + One :class:`SkillTool` per registered skill. + """ + executor = tool_executor or self._tool_executor + tools: List[BaseTool] = [] + + for manifest in self._skills.values(): + real_executor = executor or _NullToolExecutor() + skill_exec = SkillExecutor(real_executor, bus=self._bus) + + # Wire sub-skill resolver so nested skill_name steps can delegate back + skill_exec.set_skill_resolver(self._make_resolver()) + + skill_tool = SkillTool(manifest, skill_exec, skill_manager=self) + tools.append(skill_tool) + + return tools + + def _make_resolver(self): + """Return a resolver callback that delegates sub-skill execution.""" + + def _resolver(name: str, context: Dict[str, Any]) -> SkillResult: + manifest = self.resolve(name) + skill_exec = SkillExecutor( + self._tool_executor or _NullToolExecutor(), + bus=self._bus, + ) + skill_exec.set_skill_resolver(_resolver) + return skill_exec.run(manifest, initial_context=context) + + return _resolver + + # ------------------------------------------------------------------ + # Catalog + # ------------------------------------------------------------------ + + def get_catalog_xml(self) -> str: + """Generate an ```` XML catalog. + + Skills with ``disable_model_invocation=True`` are excluded so that + internal or automation-only skills are not surfaced to the model. + """ + lines: List[str] = [""] + + for manifest in self._skills.values(): + if manifest.disable_model_invocation: + continue + escaped_name = html.escape(manifest.name) + escaped_desc = html.escape(manifest.description or manifest.name) + lines.append( + f" " + ) + + lines.append("") + return "\n".join(lines) + + def get_few_shot_examples(self) -> List[str]: + """Return formatted few-shot example strings ready for system prompt. + + Pulls from ``manifest.metadata.openjarvis.few_shot`` for every + registered skill. Returns one formatted string per example. + """ + examples: List[str] = [] + for name, manifest in self._skills.items(): + oj = manifest.metadata.get("openjarvis", {}) if manifest.metadata else {} + few_shot = oj.get("few_shot", []) or [] + for ex in few_shot: + if not isinstance(ex, dict): + continue + inp = str(ex.get("input", "")) + out = str(ex.get("output", "")) + if inp or out: + examples.append(f"### {name}\nInput: {inp}\nOutput: {out}") + return examples + + # ------------------------------------------------------------------ + # Trace-driven skill discovery (Plan 2A) + # ------------------------------------------------------------------ + + def discover_from_traces( + self, + trace_store: Any, + *, + min_frequency: int = 3, + min_outcome: float = 0.5, + output_dir: Optional[Path] = None, + ) -> List[Dict[str, Any]]: + """Mine the trace store for recurring tool sequences. + + For each recurring sequence found by :class:`SkillDiscovery`, write + a TOML skill manifest into *output_dir* (default + ``~/.openjarvis/skills/discovered/``). Returns a list of dicts with + ``name`` and ``path`` for each manifest written. + + Names are normalized to spec-compliant kebab-case (lowercase with + hyphens, no underscores) so the resulting manifests load cleanly + through the discovery walker. + """ + from openjarvis.learning.agents.skill_discovery import SkillDiscovery + + traces = trace_store.list_traces(limit=10000) + discovery = SkillDiscovery( + min_frequency=min_frequency, + min_outcome=min_outcome, + ) + discovered = discovery.analyze_traces(traces) + + if output_dir is None: + output_dir = Path("~/.openjarvis/skills/discovered/").expanduser() + output_dir = Path(output_dir).expanduser() + output_dir.mkdir(parents=True, exist_ok=True) + + written: List[Dict[str, Any]] = [] + for skill in discovered: + name = self._normalize_skill_name(skill.name) + skill_subdir = output_dir / name + skill_subdir.mkdir(parents=True, exist_ok=True) + toml_path = skill_subdir / "skill.toml" + toml_path.write_text( + self._serialize_discovered_skill(name, skill), + encoding="utf-8", + ) + written.append({"name": name, "path": str(toml_path)}) + + return written + + @staticmethod + def _normalize_skill_name(raw_name: str) -> str: + """Convert an arbitrary discovered name to a spec-compliant kebab name. + + - Lowercase + - Replace underscores and whitespace with hyphens + - Collapse runs of hyphens + - Strip leading/trailing hyphens + """ + import re + + normalized = raw_name.lower() + normalized = re.sub(r"[_\s]+", "-", normalized) + normalized = re.sub(r"-+", "-", normalized) + normalized = normalized.strip("-") + if not normalized: + normalized = "discovered-skill" + return normalized + + @staticmethod + def _serialize_discovered_skill(name: str, skill: Any) -> str: + """Serialize a DiscoveredSkill into a spec-compliant skill.toml.""" + lines: List[str] = ["[skill]"] + lines.append(f'name = "{name}"') + lines.append('version = "0.1.0"') + # Truncate description to spec max 1024 chars + description = (skill.description or f"Discovered skill: {name}")[:1024] + # Escape backslashes and double quotes for basic TOML strings + description = description.replace("\\", "\\\\").replace('"', '\\"') + lines.append(f'description = "{description}"') + lines.append('author = "openjarvis (auto-discovered)"') + lines.append('tags = ["auto-discovered"]') + lines.append("") + + for tool_name in skill.tool_sequence: + lines.append("[[skill.steps]]") + lines.append(f'tool_name = "{tool_name}"') + lines.append('arguments_template = "{}"') + lines.append('output_key = ""') + lines.append("") + + return "\n".join(lines) + + # ------------------------------------------------------------------ + # Execution + # ------------------------------------------------------------------ + + def execute( + self, + name: str, + context: Optional[Dict[str, Any]] = None, + ) -> SkillResult: + """Resolve and execute a skill by name. + + Parameters + ---------- + name: + Skill name to execute. + context: + Initial context dict passed to the executor. + + Returns + ------- + SkillResult + """ + manifest = self.resolve(name) + executor = SkillExecutor( + self._tool_executor or _NullToolExecutor(), + bus=self._bus, + ) + executor.set_skill_resolver(self._make_resolver()) + return executor.run(manifest, initial_context=context) + + # ------------------------------------------------------------------ + # Configuration + # ------------------------------------------------------------------ + + def set_tool_executor(self, tool_executor: ToolExecutor) -> None: + """Attach a :class:`ToolExecutor` for running tool steps in skill pipelines.""" + self._tool_executor = tool_executor + + # ------------------------------------------------------------------ + # Lifecycle stubs (not yet implemented) + # ------------------------------------------------------------------ + + def install(self, source: Any, *, verify: bool = True) -> None: + """Install a skill from a remote source. Not yet implemented.""" + raise NotImplementedError("Skill installation not yet implemented") + + def remove(self, name: Any) -> None: + """Remove a skill by name. Not yet implemented.""" + raise NotImplementedError("Skill removal not yet implemented") + + +# --------------------------------------------------------------------------- +# Internal helpers +# --------------------------------------------------------------------------- + + +class _NullToolExecutor(ToolExecutor): + """A no-op ToolExecutor used when no real executor is available. + + Allows SkillTool/SkillExecutor construction to succeed even before a + real tool executor is wired in; any actual tool call will produce an + error ToolResult rather than raising. + """ + + def __init__(self) -> None: + super().__init__(tools=[], bus=None) + + +__all__ = ["SkillManager"] diff --git a/src/openjarvis/skills/overlay.py b/src/openjarvis/skills/overlay.py new file mode 100644 index 00000000..a7f54a89 --- /dev/null +++ b/src/openjarvis/skills/overlay.py @@ -0,0 +1,151 @@ +"""SkillOverlay — sidecar storage for optimization output (Plan 2A). + +Optimization results from DSPy/GEPA are written to +``~/.openjarvis/learning/skills//optimized.toml``. This module +provides a small loader and writer for that overlay format. + +The overlay is a strict TOML file with a single ``[optimized]`` section +followed by zero or more ``[[optimized.few_shot]]`` array tables. + +Example: + + [optimized] + skill_name = "research-and-summarize" + optimizer = "dspy" + optimized_at = "2026-04-08T14:30:00Z" + trace_count = 47 + description = "Search the web for a topic and produce a structured summary" + + [[optimized.few_shot]] + input = "transformer attention mechanisms" + output = "## Recent Advances ..." +""" + +from __future__ import annotations + +import logging +from dataclasses import dataclass, field +from pathlib import Path +from typing import Dict, List, Optional + +try: + import tomllib +except ModuleNotFoundError: + import tomli as tomllib # type: ignore[no-redef] + +LOGGER = logging.getLogger(__name__) + + +@dataclass(slots=True) +class SkillOverlay: + """A single optimization overlay for one skill.""" + + skill_name: str + optimizer: str # "dspy" or "gepa" + optimized_at: str # ISO 8601 UTC timestamp + trace_count: int + description: str + few_shot: List[Dict[str, str]] = field(default_factory=list) + + +class SkillOverlayLoader: + """Read overlay files from a sidecar directory tree. + + Layout: ``//optimized.toml`` + """ + + def __init__(self, root: Path) -> None: + self._root = Path(root).expanduser() + + def load(self, skill_name: str) -> Optional[SkillOverlay]: + """Load the overlay for *skill_name*. + + Returns ``None`` if the overlay file is missing or malformed. + Never raises — bad overlay files should not break skill loading. + """ + path = self._root / skill_name / "optimized.toml" + if not path.exists(): + return None + try: + with open(path, "rb") as fh: + data = tomllib.load(fh) + except Exception as exc: + LOGGER.warning( + "Failed to load overlay for skill '%s' at %s: %s", + skill_name, + path, + exc, + ) + return None + + opt = data.get("optimized", {}) + if not isinstance(opt, dict) or not opt: + LOGGER.warning( + "Overlay file %s missing [optimized] section", + path, + ) + return None + + few_shot_raw = opt.get("few_shot", []) or [] + few_shot: List[Dict[str, str]] = [] + if isinstance(few_shot_raw, list): + for item in few_shot_raw: + if isinstance(item, dict): + few_shot.append( + { + "input": str(item.get("input", "")), + "output": str(item.get("output", "")), + } + ) + + return SkillOverlay( + skill_name=str(opt.get("skill_name", skill_name)), + optimizer=str(opt.get("optimizer", "")), + optimized_at=str(opt.get("optimized_at", "")), + trace_count=int(opt.get("trace_count", 0)), + description=str(opt.get("description", "")), + few_shot=few_shot, + ) + + +def write_overlay(overlay: SkillOverlay, root: Path) -> Path: + """Write *overlay* to ``//optimized.toml``. + + Creates the directory structure if needed. Returns the path written. + """ + root = Path(root).expanduser() + skill_dir = root / overlay.skill_name + skill_dir.mkdir(parents=True, exist_ok=True) + path = skill_dir / "optimized.toml" + + lines: List[str] = ["[optimized]"] + lines.append(f'skill_name = "{_escape(overlay.skill_name)}"') + lines.append(f'optimizer = "{_escape(overlay.optimizer)}"') + lines.append(f'optimized_at = "{_escape(overlay.optimized_at)}"') + lines.append(f"trace_count = {int(overlay.trace_count)}") + lines.append(f'description = "{_escape(overlay.description)}"') + lines.append("") + + for example in overlay.few_shot: + lines.append("[[optimized.few_shot]]") + lines.append(f'input = "{_escape(example.get("input", ""))}"') + lines.append(f'output = "{_escape(example.get("output", ""))}"') + lines.append("") + + path.write_text("\n".join(lines), encoding="utf-8") + return path + + +def _escape(text: str) -> str: + """Minimal TOML string escaping for the basic-string format.""" + return ( + str(text) + .replace("\\", "\\\\") + .replace('"', '\\"') + .replace("\n", "\\n") + .replace("\r", "\\r") + .replace("\t", "\\t") + ) + + +__all__ = ["SkillOverlay", "SkillOverlayLoader", "write_overlay"] diff --git a/src/openjarvis/skills/parser.py b/src/openjarvis/skills/parser.py new file mode 100644 index 00000000..26a9381d --- /dev/null +++ b/src/openjarvis/skills/parser.py @@ -0,0 +1,231 @@ +"""SkillParser — strict spec validation + tolerant field mapping. + +The parser is the single chokepoint for converting raw frontmatter dicts +into normalized SkillManifest instances. It runs two passes: + +1. Strict pass — validates required fields, length limits, naming rules. +2. Tolerant pass — maps non-spec top-level fields to their canonical + locations under metadata.openjarvis.* via FIELD_MAPPING. + +The mapping table is data, not code paths. Adding support for a new +vendor's fields means adding entries to FIELD_MAPPING — no logic changes. +""" + +from __future__ import annotations + +import logging +import re +from typing import Any, Dict + +from openjarvis.skills.types import SkillManifest + +LOGGER = logging.getLogger(__name__) + +# --------------------------------------------------------------------------- +# Spec constants +# --------------------------------------------------------------------------- + +MAX_NAME_LENGTH = 64 +MAX_DESCRIPTION_LENGTH = 1024 +MAX_COMPATIBILITY_LENGTH = 500 + +# Spec-allowed top-level frontmatter fields +SPEC_FIELDS = frozenset( + {"name", "description", "license", "compatibility", "metadata", "allowed-tools"} +) + +# --------------------------------------------------------------------------- +# Field mapping table — non-spec top-level fields → canonical locations +# --------------------------------------------------------------------------- + +# Each entry maps a non-spec top-level field name to (target_kind, attr). +# When target_kind is "field" the value is set directly on the SkillManifest +# dataclass attribute. When target_kind is "openjarvis_meta" the value is +# stored under manifest.metadata["openjarvis"][attr]. +FIELD_MAPPING: Dict[str, tuple[str, str]] = { + "version": ("field", "version"), + "author": ("field", "author"), + "tags": ("field", "tags"), + "depends": ("field", "depends"), + "required_capabilities": ("field", "required_capabilities"), + "user_invocable": ("field", "user_invocable"), + "disable_model_invocation": ("field", "disable_model_invocation"), + "platforms": ("openjarvis_meta", "platforms"), + "prerequisites": ("openjarvis_meta", "prerequisites"), +} + +# Naming pattern: lowercase alnum + hyphens, no leading/trailing/consecutive hyphens +_NAME_PATTERN = re.compile(r"^[a-z0-9](?:[a-z0-9]|-(?!-))*[a-z0-9]$|^[a-z0-9]$") + + +class SkillParseError(ValueError): + """Raised when a skill frontmatter cannot be parsed.""" + + +class SkillParser: + """Parse SKILL.md frontmatter into a SkillManifest. + + Two-pass design: + + - :meth:`parse_frontmatter` runs the strict pass + tolerant pass and + returns a normalized :class:`SkillManifest`. + """ + + def parse_frontmatter( + self, + frontmatter: Dict[str, Any], + *, + markdown_content: str = "", + ) -> SkillManifest: + """Parse a frontmatter dict, returning a SkillManifest. + + Runs strict validation first, then applies tolerant field mapping. + """ + self._validate_strict(frontmatter) + return self._build_manifest(frontmatter, markdown_content) + + # ------------------------------------------------------------------ + # Strict pass + # ------------------------------------------------------------------ + + def _validate_strict(self, frontmatter: Dict[str, Any]) -> None: + """Validate required fields, length limits, and naming rules.""" + # Required fields + if "name" not in frontmatter: + raise SkillParseError("Missing required field in frontmatter: name") + if "description" not in frontmatter: + raise SkillParseError("Missing required field in frontmatter: description") + + name = frontmatter["name"] + description = frontmatter["description"] + + # Type check + if not isinstance(name, str): + raise SkillParseError(f"Field 'name' must be a string, got {type(name)}") + if not isinstance(description, str): + raise SkillParseError( + f"Field 'description' must be a string, got {type(description)}" + ) + + # Length limits + if len(name) == 0 or len(name) > MAX_NAME_LENGTH: + raise SkillParseError( + f"Field 'name' must be 1-{MAX_NAME_LENGTH} chars, got {len(name)}" + ) + if len(description) == 0 or len(description) > MAX_DESCRIPTION_LENGTH: + raise SkillParseError( + f"Field 'description' must be 1-{MAX_DESCRIPTION_LENGTH} chars, " + f"got {len(description)}" + ) + + # Naming rules + self._validate_name(name) + + # compatibility length (optional field) + compat = frontmatter.get("compatibility") + if compat is not None: + if not isinstance(compat, str): + raise SkillParseError("Field 'compatibility' must be a string") + if len(compat) > MAX_COMPATIBILITY_LENGTH: + raise SkillParseError( + f"Field 'compatibility' exceeds {MAX_COMPATIBILITY_LENGTH} chars" + ) + + def _validate_name(self, name: str) -> None: + """Validate the spec naming rules for a skill name.""" + if name != name.lower(): + raise SkillParseError(f"Skill name '{name}' must be lowercase") + if name.startswith("-") or name.endswith("-"): + raise SkillParseError( + f"Skill name '{name}' must not start or end with a hyphen" + ) + if "--" in name: + raise SkillParseError( + f"Skill name '{name}' must not contain consecutive hyphens" + ) + for ch in name: + if not (ch.isalnum() or ch == "-"): + raise SkillParseError( + f"Skill name '{name}' contains invalid character '{ch}'; " + f"only lowercase alphanumeric and hyphens are allowed" + ) + + # ------------------------------------------------------------------ + # Build manifest (placeholder — Task 2 adds tolerant pass) + # ------------------------------------------------------------------ + + def _build_manifest( + self, + frontmatter: Dict[str, Any], + markdown_content: str, + ) -> SkillManifest: + """Construct a SkillManifest from validated frontmatter. + + Runs the tolerant pass: applies FIELD_MAPPING for non-spec fields, + captures unknown fields under metadata.openjarvis.original_frontmatter, + and merges metadata.openjarvis.* into canonical fields. + """ + # Start with required + optional spec fields + manifest = SkillManifest( + name=frontmatter["name"], + description=frontmatter["description"], + markdown_content=markdown_content, + ) + + # Pre-existing metadata block (if any) + raw_metadata = frontmatter.get("metadata") or {} + if not isinstance(raw_metadata, dict): + raw_metadata = {} + # Initialize openjarvis namespace + oj_meta = dict(raw_metadata.get("openjarvis") or {}) + + # Apply FIELD_MAPPING for non-spec top-level fields + unmapped: Dict[str, Any] = {} + for key, value in frontmatter.items(): + if key in SPEC_FIELDS: + continue + if key in FIELD_MAPPING: + target, attr = FIELD_MAPPING[key] + if target == "field": + setattr(manifest, attr, value) + else: # "openjarvis_meta" + oj_meta[attr] = value + else: + unmapped[key] = value + LOGGER.warning( + "Unmapped frontmatter field '%s' in skill '%s' " + "(value preserved in metadata.openjarvis.original_frontmatter)", + key, + manifest.name, + ) + + # Merge metadata.openjarvis.* into canonical fields (these override + # top-level mappings since they are explicit OpenJarvis-namespaced). + for key in ( + "version", + "author", + "tags", + "depends", + "required_capabilities", + "user_invocable", + "disable_model_invocation", + ): + if key in oj_meta: + setattr(manifest, key, oj_meta[key]) + + # Preserve unmapped fields under metadata.openjarvis.original_frontmatter + if unmapped: + oj_meta["original_frontmatter"] = unmapped + + # Stash the openjarvis metadata block back + if oj_meta: + new_metadata = dict(raw_metadata) + new_metadata["openjarvis"] = oj_meta + manifest.metadata = new_metadata + else: + manifest.metadata = raw_metadata + + return manifest + + +__all__ = ["SkillParseError", "SkillParser", "SPEC_FIELDS", "FIELD_MAPPING"] diff --git a/src/openjarvis/skills/security.py b/src/openjarvis/skills/security.py new file mode 100644 index 00000000..b514440f --- /dev/null +++ b/src/openjarvis/skills/security.py @@ -0,0 +1,65 @@ +"""Skill security — capability validation and trust tiers.""" + +from __future__ import annotations + +from enum import Enum +from typing import List, Set + +from openjarvis.skills.types import SkillManifest + +DANGEROUS_CAPABILITIES: frozenset[str] = frozenset( + {"shell:execute", "network:listen", "filesystem:write"} +) + + +class TrustTier(str, Enum): + """Trust tier for a skill, ordered from most to least trusted.""" + + BUNDLED = "bundled" + INDEXED = "indexed" + WORKSPACE = "workspace" + UNREVIEWED = "unreviewed" + + +def classify_trust_tier( + *, + is_bundled: bool = False, + is_workspace: bool = False, + has_signature: bool = False, + in_index: bool = False, +) -> TrustTier: + """Return the trust tier for a skill based on its provenance. + + Priority (highest to lowest): bundled > workspace > indexed > unreviewed. + """ + if is_bundled: + return TrustTier.BUNDLED + if is_workspace: + return TrustTier.WORKSPACE + if has_signature and in_index: + return TrustTier.INDEXED + return TrustTier.UNREVIEWED + + +def validate_capabilities(manifest: SkillManifest, allowed: Set[str]) -> List[str]: + """Return a list of capabilities required by *manifest* that are not in *allowed*. + + An empty list means the manifest is fully authorized. + """ + return [cap for cap in manifest.required_capabilities if cap not in allowed] + + +def has_dangerous_capabilities(manifest: SkillManifest) -> List[str]: + """Return the subset of *manifest*'s required capabilities that are dangerous.""" + return [ + cap for cap in manifest.required_capabilities if cap in DANGEROUS_CAPABILITIES + ] + + +__all__ = [ + "DANGEROUS_CAPABILITIES", + "TrustTier", + "classify_trust_tier", + "validate_capabilities", + "has_dangerous_capabilities", +] diff --git a/src/openjarvis/skills/sources/__init__.py b/src/openjarvis/skills/sources/__init__.py new file mode 100644 index 00000000..8940e7c9 --- /dev/null +++ b/src/openjarvis/skills/sources/__init__.py @@ -0,0 +1,16 @@ +"""Skill source resolvers — Hermes, OpenClaw, generic GitHub.""" + +from openjarvis.skills.sources.base import ResolvedSkill, SourceResolver +from openjarvis.skills.sources.github import GitHubResolver +from openjarvis.skills.sources.hermes import HERMES_REPO_URL, HermesResolver +from openjarvis.skills.sources.openclaw import OPENCLAW_REPO_URL, OpenClawResolver + +__all__ = [ + "GitHubResolver", + "HERMES_REPO_URL", + "HermesResolver", + "OPENCLAW_REPO_URL", + "OpenClawResolver", + "ResolvedSkill", + "SourceResolver", +] diff --git a/src/openjarvis/skills/sources/base.py b/src/openjarvis/skills/sources/base.py new file mode 100644 index 00000000..140f59b4 --- /dev/null +++ b/src/openjarvis/skills/sources/base.py @@ -0,0 +1,66 @@ +"""Source resolver ABC + ResolvedSkill dataclass.""" + +from __future__ import annotations + +from abc import ABC, abstractmethod +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Dict, List + + +@dataclass(slots=True) +class ResolvedSkill: + """A skill found in an upstream source, ready to import. + + Lightweight — does not include the full SKILL.md body. The importer + reads the file from *path* when actually installing. + """ + + name: str + source: str + path: Path + category: str + description: str + commit: str + sidecar_data: Dict[str, Any] = field(default_factory=dict) + + +class SourceResolver(ABC): + """Abstract base for a skill source resolver. + + Implementations clone or pull an upstream repo into a cache directory, + walk the cache to find SKILL.md files, and return ResolvedSkill objects + that the importer can install. + """ + + name: str = "" + + @abstractmethod + def cache_dir(self) -> Path: + """Where this source clones its repo.""" + + @abstractmethod + def sync(self) -> None: + """Clone or pull the upstream repo into the cache.""" + + @abstractmethod + def list_skills(self) -> List[ResolvedSkill]: + """Walk the cache directory and return all discoverable skills.""" + + def resolve(self, query: str) -> List[ResolvedSkill]: + """Filter list_skills() by name (substring match). + + Empty *query* returns all skills. + """ + all_skills = self.list_skills() + if not query: + return all_skills + q = query.lower() + return [s for s in all_skills if q in s.name.lower()] + + def filter_by_category(self, category: str) -> List[ResolvedSkill]: + """Return skills whose category exactly matches *category*.""" + return [s for s in self.list_skills() if s.category == category] + + +__all__ = ["ResolvedSkill", "SourceResolver"] diff --git a/src/openjarvis/skills/sources/github.py b/src/openjarvis/skills/sources/github.py new file mode 100644 index 00000000..62c81fea --- /dev/null +++ b/src/openjarvis/skills/sources/github.py @@ -0,0 +1,121 @@ +"""GitHubResolver — generic resolver for any GitHub repo containing skills. + +Performs a recursive walk for SKILL.md (or skill.md) files anywhere +under the cache directory. +""" + +from __future__ import annotations + +import logging +import subprocess +from pathlib import Path +from typing import List + +import yaml + +from openjarvis.skills.sources.base import ResolvedSkill, SourceResolver + +LOGGER = logging.getLogger(__name__) + + +class GitHubResolver(SourceResolver): + """Generic resolver for any GitHub repo containing SKILL.md files.""" + + name = "github" + + def __init__(self, cache_root: Path, repo_url: str) -> None: + self._cache_root = Path(cache_root) + self._repo_url = repo_url + + def cache_dir(self) -> Path: + return self._cache_root + + def sync(self) -> None: + if self._cache_root.exists() and (self._cache_root / ".git").exists(): + subprocess.run( + ["git", "-C", str(self._cache_root), "pull", "--ff-only"], + check=True, + ) + else: + self._cache_root.parent.mkdir(parents=True, exist_ok=True) + subprocess.run( + ["git", "clone", self._repo_url, str(self._cache_root)], + check=True, + ) + + def list_skills(self) -> List[ResolvedSkill]: + if not self._cache_root.exists(): + return [] + + results: List[ResolvedSkill] = [] + commit = self._read_commit() + seen_dirs: set[Path] = set() + + # Recursive walk for SKILL.md or skill.md + for pattern in ("SKILL.md", "skill.md"): + for skill_md in sorted(self._cache_root.rglob(pattern)): + # Skip files inside .git + if ".git" in skill_md.parts: + continue + skill_dir = skill_md.parent + if skill_dir in seen_dirs: + continue + seen_dirs.add(skill_dir) + + name, description = self._read_preview( + skill_md, default_name=skill_dir.name + ) + # Use the immediate parent directory of the skill dir as category + try: + category = skill_dir.parent.relative_to(self._cache_root).as_posix() + except ValueError: + category = "" + + results.append( + ResolvedSkill( + name=name, + source=self.name, + path=skill_dir, + category=category, + description=description, + commit=commit, + ) + ) + + return results + + def _read_preview(self, skill_md: Path, default_name: str) -> tuple[str, str]: + try: + raw = skill_md.read_text(encoding="utf-8") + except Exception: + return default_name, "" + if not raw.startswith("---"): + return default_name, "" + rest = raw[3:].lstrip("\n") + end = rest.find("\n---") + if end == -1: + return default_name, "" + try: + fm = yaml.safe_load(rest[:end]) + except yaml.YAMLError: + return default_name, "" + if not isinstance(fm, dict): + return default_name, "" + return str(fm.get("name", default_name)), str(fm.get("description", "")) + + def _read_commit(self) -> str: + if not (self._cache_root / ".git").exists(): + return "" + try: + result = subprocess.run( + ["git", "-C", str(self._cache_root), "rev-parse", "HEAD"], + capture_output=True, + text=True, + check=True, + ) + return result.stdout.strip() + except subprocess.CalledProcessError: + return "" + + +__all__ = ["GitHubResolver"] diff --git a/src/openjarvis/skills/sources/hermes.py b/src/openjarvis/skills/sources/hermes.py new file mode 100644 index 00000000..273e4867 --- /dev/null +++ b/src/openjarvis/skills/sources/hermes.py @@ -0,0 +1,130 @@ +"""HermesResolver — resolves skills from NousResearch/hermes-agent. + +Layout: + skills///SKILL.md + skills//DESCRIPTION.md (skipped) +""" + +from __future__ import annotations + +import logging +import subprocess +from pathlib import Path +from typing import List + +import yaml + +from openjarvis.skills.sources.base import ResolvedSkill, SourceResolver + +LOGGER = logging.getLogger(__name__) + +HERMES_REPO_URL = "https://github.com/NousResearch/hermes-agent.git" + + +class HermesResolver(SourceResolver): + """Resolves skills from the Hermes Agent repository.""" + + name = "hermes" + + def __init__(self, cache_root: Path | None = None) -> None: + if cache_root is None: + cache_root = Path("~/.openjarvis/skill-cache/hermes/").expanduser() + self._cache_root = Path(cache_root) + + def cache_dir(self) -> Path: + return self._cache_root + + def sync(self) -> None: + """Clone or pull the Hermes repo into the cache directory.""" + if self._cache_root.exists() and (self._cache_root / ".git").exists(): + subprocess.run( + ["git", "-C", str(self._cache_root), "pull", "--ff-only"], + check=True, + ) + else: + self._cache_root.parent.mkdir(parents=True, exist_ok=True) + subprocess.run( + ["git", "clone", HERMES_REPO_URL, str(self._cache_root)], + check=True, + ) + + def list_skills(self) -> List[ResolvedSkill]: + """Walk skills///SKILL.md.""" + skills_root = self._cache_root / "skills" + if not skills_root.exists(): + return [] + + results: List[ResolvedSkill] = [] + commit = self._read_commit() + + for category_dir in sorted(skills_root.iterdir()): + if not category_dir.is_dir(): + continue + category = category_dir.name + for skill_dir in sorted(category_dir.iterdir()): + if not skill_dir.is_dir(): + continue # skip DESCRIPTION.md and other files + skill_md = skill_dir / "SKILL.md" + if not skill_md.exists(): + continue + + # Read minimal frontmatter for the preview + name, description = self._read_preview( + skill_md, default_name=skill_dir.name + ) + results.append( + ResolvedSkill( + name=name, + source=self.name, + path=skill_dir, + category=category, + description=description, + commit=commit, + ) + ) + + return results + + def _read_preview(self, skill_md: Path, default_name: str) -> tuple[str, str]: + """Read just enough frontmatter to populate ResolvedSkill preview fields.""" + try: + raw = skill_md.read_text(encoding="utf-8") + except Exception: + return default_name, "" + + if not raw.startswith("---"): + return default_name, "" + + rest = raw[3:].lstrip("\n") + end = rest.find("\n---") + if end == -1: + return default_name, "" + try: + fm = yaml.safe_load(rest[:end]) + except yaml.YAMLError: + return default_name, "" + if not isinstance(fm, dict): + return default_name, "" + + return ( + str(fm.get("name", default_name)), + str(fm.get("description", "")), + ) + + def _read_commit(self) -> str: + """Return the current HEAD SHA of the cached repo, or empty string.""" + if not (self._cache_root / ".git").exists(): + return "" + try: + result = subprocess.run( + ["git", "-C", str(self._cache_root), "rev-parse", "HEAD"], + capture_output=True, + text=True, + check=True, + ) + return result.stdout.strip() + except subprocess.CalledProcessError: + return "" + + +__all__ = ["HermesResolver", "HERMES_REPO_URL"] diff --git a/src/openjarvis/skills/sources/openclaw.py b/src/openjarvis/skills/sources/openclaw.py new file mode 100644 index 00000000..780f4401 --- /dev/null +++ b/src/openjarvis/skills/sources/openclaw.py @@ -0,0 +1,129 @@ +"""OpenClawResolver — resolves skills from the OpenClaw skill index. + +Layout: + skills///SKILL.md + skills///_meta.json (optional sidecar registry data) +""" + +from __future__ import annotations + +import json +import logging +import subprocess +from pathlib import Path +from typing import List + +import yaml + +from openjarvis.skills.sources.base import ResolvedSkill, SourceResolver + +LOGGER = logging.getLogger(__name__) + +OPENCLAW_REPO_URL = "https://github.com/openclaw/skills.git" + + +class OpenClawResolver(SourceResolver): + """Resolves skills from the OpenClaw skill index.""" + + name = "openclaw" + + def __init__(self, cache_root: Path | None = None) -> None: + if cache_root is None: + cache_root = Path("~/.openjarvis/skill-cache/openclaw/").expanduser() + self._cache_root = Path(cache_root) + + def cache_dir(self) -> Path: + return self._cache_root + + def sync(self) -> None: + if self._cache_root.exists() and (self._cache_root / ".git").exists(): + subprocess.run( + ["git", "-C", str(self._cache_root), "pull", "--ff-only"], + check=True, + ) + else: + self._cache_root.parent.mkdir(parents=True, exist_ok=True) + subprocess.run( + ["git", "clone", OPENCLAW_REPO_URL, str(self._cache_root)], + check=True, + ) + + def list_skills(self) -> List[ResolvedSkill]: + skills_root = self._cache_root / "skills" + if not skills_root.exists(): + return [] + + results: List[ResolvedSkill] = [] + commit = self._read_commit() + + for owner_dir in sorted(skills_root.iterdir()): + if not owner_dir.is_dir(): + continue + for skill_dir in sorted(owner_dir.iterdir()): + if not skill_dir.is_dir(): + continue + skill_md = skill_dir / "SKILL.md" + if not skill_md.exists(): + continue + + name, description = self._read_preview( + skill_md, default_name=skill_dir.name + ) + sidecar = self._read_sidecar(skill_dir / "_meta.json") + results.append( + ResolvedSkill( + name=name, + source=self.name, + path=skill_dir, + category=owner_dir.name, + description=description, + commit=commit, + sidecar_data=sidecar, + ) + ) + + return results + + def _read_preview(self, skill_md: Path, default_name: str) -> tuple[str, str]: + try: + raw = skill_md.read_text(encoding="utf-8") + except Exception: + return default_name, "" + if not raw.startswith("---"): + return default_name, "" + rest = raw[3:].lstrip("\n") + end = rest.find("\n---") + if end == -1: + return default_name, "" + try: + fm = yaml.safe_load(rest[:end]) + except yaml.YAMLError: + return default_name, "" + if not isinstance(fm, dict): + return default_name, "" + return str(fm.get("name", default_name)), str(fm.get("description", "")) + + def _read_sidecar(self, sidecar_path: Path) -> dict: + if not sidecar_path.exists(): + return {} + try: + return json.loads(sidecar_path.read_text(encoding="utf-8")) + except (json.JSONDecodeError, OSError): + return {} + + def _read_commit(self) -> str: + if not (self._cache_root / ".git").exists(): + return "" + try: + result = subprocess.run( + ["git", "-C", str(self._cache_root), "rev-parse", "HEAD"], + capture_output=True, + text=True, + check=True, + ) + return result.stdout.strip() + except subprocess.CalledProcessError: + return "" + + +__all__ = ["OpenClawResolver", "OPENCLAW_REPO_URL"] diff --git a/src/openjarvis/skills/tool_adapter.py b/src/openjarvis/skills/tool_adapter.py index a615102d..13ede93a 100644 --- a/src/openjarvis/skills/tool_adapter.py +++ b/src/openjarvis/skills/tool_adapter.py @@ -2,18 +2,30 @@ from __future__ import annotations -from typing import Any, Dict +import re +from typing import Any, Dict, List, Optional, Set from openjarvis.core.types import ToolResult from openjarvis.skills.executor import SkillExecutor from openjarvis.skills.types import SkillManifest from openjarvis.tools._stubs import BaseTool, ToolSpec +_PLACEHOLDER_RE = re.compile(r"\{(\w+)\}") + class SkillTool(BaseTool): """Wraps a SkillManifest as a BaseTool that agents can invoke. Follows the same adapter pattern as MCPToolAdapter. + + Parameters + ---------- + manifest: + The skill manifest to wrap. + executor: + A :class:`SkillExecutor` used to run the skill pipeline. + skill_manager: + Optional skill manager (reserved for sub-skill delegation). """ tool_id: str @@ -22,50 +34,162 @@ class SkillTool(BaseTool): self, manifest: SkillManifest, executor: SkillExecutor, + *, + skill_manager: Optional[Any] = None, ) -> None: self._manifest = manifest self._executor = executor + self._skill_manager = skill_manager self.tool_id = f"skill_{manifest.name}" + self._parameters = self._build_parameters() + + # ------------------------------------------------------------------ + # Parameter extraction + # ------------------------------------------------------------------ + + def _build_parameters(self) -> Dict[str, Any]: + """Auto-extract input parameters from the manifest. + + For pipeline skills, scan each step's ``arguments_template`` for + ``{placeholder}`` patterns. Subtract the ``output_key`` values + produced by *prior* steps so only externally-supplied parameters + are exposed. + + For instruction-only skills (no steps), expose a single optional + ``task`` parameter. + """ + if not self._manifest.steps: + # Instruction-only / markdown-only skill + return { + "type": "object", + "properties": { + "task": { + "type": "string", + "description": "Optional task description or context.", + } + }, + } + + produced: Set[str] = set() + input_params: List[str] = [] + seen: Set[str] = set() + + for step in self._manifest.steps: + # Find all {placeholder} tokens in this step's template + placeholders = _PLACEHOLDER_RE.findall(step.arguments_template) + for ph in placeholders: + if ph not in produced and ph not in seen: + input_params.append(ph) + seen.add(ph) + + # After processing this step, its output_key becomes available + # for subsequent steps and should NOT be surfaced as an input + if step.output_key: + produced.add(step.output_key) + + properties: Dict[str, Any] = {} + for param in input_params: + properties[param] = { + "type": "string", + "description": f"Input value for '{param}'.", + } + + return { + "type": "object", + "properties": properties, + } + + # ------------------------------------------------------------------ + # Result metadata + # ------------------------------------------------------------------ + + def _build_result_metadata(self, *, steps_run: int) -> Dict[str, Any]: + """Build the metadata dict attached to ToolResult. + + Includes the skill name, source provenance, and kind so downstream + consumers (TraceCollector, SkillOptimizer) can bucket invocations. + """ + oj_meta = self._manifest.metadata.get("openjarvis", {}) or {} + source = oj_meta.get("source", "user") + kind = "executable" if self._manifest.steps else "instructional" + return { + "skill": self._manifest.name, + "skill_source": source, + "skill_kind": kind, + "steps": steps_run, + } + + # ------------------------------------------------------------------ + # BaseTool interface + # ------------------------------------------------------------------ @property def spec(self) -> ToolSpec: return ToolSpec( name=f"skill_{self._manifest.name}", description=self._manifest.description or f"Skill: {self._manifest.name}", - parameters={ - "type": "object", - "properties": { - "input": { - "type": "string", - "description": "Input text for the skill pipeline.", - }, - "context": { - "type": "object", - "description": "Additional context key-value pairs.", - }, - }, - }, + parameters=self._parameters, category="skill", required_capabilities=self._manifest.required_capabilities, ) def execute(self, **params: Any) -> ToolResult: - initial_ctx: Dict[str, Any] = params.get("context", {}) - if "input" in params: - initial_ctx["input"] = params["input"] + """Execute the skill. - result = self._executor.run(self._manifest, initial_context=initial_ctx) + If the manifest has pipeline steps, run them via the executor and + collect the last step's output. If ``markdown_content`` is present, + append it to the content. Returns a combined :class:`ToolResult`. + """ + tool_name = self.spec.name + content_parts: List[str] = [] + + if self._manifest.steps: + # Build initial context from all supplied params + initial_ctx: Dict[str, Any] = {k: v for k, v in params.items()} + + result = self._executor.run(self._manifest, initial_context=initial_ctx) + + if not result.success: + # Propagate failure immediately + error_content = ( + result.step_results[-1].content + if result.step_results + else "Pipeline failed with no step results." + ) + return ToolResult( + tool_name=tool_name, + content=error_content, + success=False, + metadata=self._build_result_metadata( + steps_run=len(result.step_results) + ), + ) + + # Use the last step's output as the primary content + if result.step_results: + last_step = result.step_results[-1] + last_output = ( + result.context.get( + # Prefer the keyed output if available + self._manifest.steps[-1].output_key, + last_step.content, + ) + if self._manifest.steps[-1].output_key + else last_step.content + ) + content_parts.append(str(last_output)) + + # Append markdown content if present (hybrid or instruction-only) + if self._manifest.markdown_content: + content_parts.append(self._manifest.markdown_content) + + combined = "\n\n".join(filter(None, content_parts)) return ToolResult( - tool_name=self.spec.name, - content=result.context.get( - result.step_results[-1].tool_name if result.step_results else "", - result.step_results[-1].content if result.step_results else "", - ) - if result.step_results - else "", - success=result.success, - metadata={"skill": self._manifest.name, "steps": len(result.step_results)}, + tool_name=tool_name, + content=combined, + success=True, + metadata=self._build_result_metadata(steps_run=len(self._manifest.steps)), ) diff --git a/src/openjarvis/skills/tool_translator.py b/src/openjarvis/skills/tool_translator.py new file mode 100644 index 00000000..5f59f8c2 --- /dev/null +++ b/src/openjarvis/skills/tool_translator.py @@ -0,0 +1,114 @@ +"""ToolTranslator — translate external tool names to OpenJarvis equivalents. + +External skill libraries (Hermes Agent, OpenClaw) reference tools by Claude +Code's standard tool names (Bash, Read, Write, etc.). OpenJarvis uses +different names (shell_exec, file_read, file_write). This module translates +those references in skill markdown bodies and ``allowed-tools`` fields. + +The translation table is small (~10 entries) covering Claude Code's standard +tools and grows as we encounter new vendor tools. +""" + +from __future__ import annotations + +import re +from typing import Dict, List, Tuple + +# --------------------------------------------------------------------------- +# Translation table — external name → OpenJarvis name +# --------------------------------------------------------------------------- + +TOOL_TRANSLATION: Dict[str, str] = { + "Bash": "shell_exec", + "Read": "file_read", + "Write": "file_write", + "Edit": "file_edit", + "Glob": "file_glob", + "Grep": "file_grep", + "WebFetch": "web_search", + "WebSearch": "web_search", + "Task": "delegate_agent", + "NotebookEdit": "notebook_edit", +} + + +class ToolTranslator: + """Rewrite tool references in markdown bodies and allowed-tools fields. + + Uses a word-boundary regex so partial matches like 'Reader' or + 'Reading' are not rewritten. + """ + + def __init__( + self, + translation_table: Dict[str, str] | None = None, + ) -> None: + self._table = dict(translation_table or TOOL_TRANSLATION) + # Word-boundary pattern matching any external name + names = sorted(self._table.keys(), key=len, reverse=True) + if names: + self._pattern = re.compile( + r"\b(" + "|".join(re.escape(n) for n in names) + r")\b" + ) + else: + self._pattern = None + + def translate_markdown(self, body: str) -> Tuple[str, List[str]]: + """Translate tool references in a markdown body. + + Returns + ------- + tuple[str, list[str]] + (rewritten body, list of untranslated tool names found) + """ + if not body or self._pattern is None: + return body, [] + + def _sub(match: re.Match) -> str: + return self._table.get(match.group(1), match.group(1)) + + new_body = self._pattern.sub(_sub, body) + + # Find untranslated tool-like references (CamelCase words that look + # like tool names but aren't in the table). Conservative heuristic: + # words with internal uppercase (true CamelCase) followed by ' tool' + # or a word boundary, length 3-30, not in the translation table. + untranslated: List[str] = [] + # Require at least one uppercase letter after position 0 (true CamelCase) + candidate_pattern = re.compile(r"\b([A-Z][a-z]+[A-Z][a-zA-Z]*)(?:\s+tool|\b)") + for cand in candidate_pattern.findall(body): + if cand not in self._table and cand not in untranslated: + if 3 <= len(cand) <= 30: + untranslated.append(cand) + return new_body, untranslated + + def translate_allowed_tools(self, allowed: str) -> Tuple[str, List[str]]: + """Translate the space-delimited allowed-tools field. + + Tokens may have parenthesized arguments like ``Bash(git:*)``. Only + the prefix before the first ``(`` is translated. + """ + if not allowed: + return allowed, [] + + out_tokens: List[str] = [] + untranslated: List[str] = [] + for token in allowed.split(): + # Split off any (args) suffix + if "(" in token: + head, _, tail = token.partition("(") + tail = "(" + tail + else: + head, tail = token, "" + + if head in self._table: + out_tokens.append(self._table[head] + tail) + else: + out_tokens.append(token) + if head not in untranslated: + untranslated.append(head) + + return " ".join(out_tokens), untranslated + + +__all__ = ["TOOL_TRANSLATION", "ToolTranslator"] diff --git a/src/openjarvis/skills/types.py b/src/openjarvis/skills/types.py index c0d82ab8..5cff305e 100644 --- a/src/openjarvis/skills/types.py +++ b/src/openjarvis/skills/types.py @@ -10,7 +10,8 @@ from typing import Any, Dict, List class SkillStep: """A single step in a skill pipeline.""" - tool_name: str + tool_name: str = "" + skill_name: str = "" # invoke another skill instead of a tool arguments_template: str = "{}" # Jinja2-style template output_key: str = "" # Key to store result in context @@ -27,6 +28,11 @@ class SkillManifest: required_capabilities: List[str] = field(default_factory=list) signature: str = "" # Base64-encoded Ed25519 signature metadata: Dict[str, Any] = field(default_factory=dict) + tags: List[str] = field(default_factory=list) + depends: List[str] = field(default_factory=list) + user_invocable: bool = True + disable_model_invocation: bool = False + markdown_content: str = "" # loaded from SKILL.md def manifest_bytes(self) -> bytes: """Serialize the manifest (excluding signature) for signing/verification.""" @@ -40,12 +46,15 @@ class SkillManifest: "steps": [ { "tool_name": s.tool_name, + "skill_name": s.skill_name, "arguments_template": s.arguments_template, "output_key": s.output_key, } for s in self.steps ], "required_capabilities": self.required_capabilities, + "tags": self.tags, + "depends": self.depends, } return json.dumps(data, sort_keys=True).encode() diff --git a/src/openjarvis/system.py b/src/openjarvis/system.py index 5becc0f0..bfb6c645 100644 --- a/src/openjarvis/system.py +++ b/src/openjarvis/system.py @@ -49,6 +49,7 @@ class JarvisSystem: agent_scheduler: Optional[Any] = None # AgentScheduler agent_executor: Optional[Any] = None # AgentExecutor speech_backend: Optional[Any] = None # SpeechBackend + skill_manager: Optional[Any] = None # SkillManager _learning_orchestrator: Optional[Any] = None # LearningOrchestrator _mcp_clients: List = field(default_factory=list) @@ -202,6 +203,12 @@ class JarvisSystem: if getattr(agent_cls, "accepts_tools", False): agent_kwargs["tools"] = agent_tools agent_kwargs["max_turns"] = self.config.agent.max_turns + # Plan 2B I3: forward optimized few-shot examples to agents + # that accept them. Older agents that don't accept the kwarg + # are handled by the existing TypeError fallback below. + examples = getattr(self, "_skill_few_shot_examples", None) + if examples: + agent_kwargs["skill_few_shot_examples"] = examples if system_prompt is not None: agent_kwargs["system_prompt"] = system_prompt if self.capability_policy is not None: @@ -662,6 +669,37 @@ class SystemBuilder: # Build tool executor tool_executor = ToolExecutor(tool_list, bus) if tool_list else None + # Resolve skills + skill_manager = None + skill_few_shot_examples: List[str] = [] + if config.skills.enabled: + try: + from pathlib import Path + + from openjarvis.skills.manager import SkillManager + + skill_manager = SkillManager( + bus, capability_policy=sec.capability_policy + ) + skill_paths = [Path(config.skills.skills_dir).expanduser()] + workspace_skills = Path("./skills") + if workspace_skills.exists(): + skill_paths.insert(0, workspace_skills) + skill_manager.discover(paths=skill_paths) + if tool_executor: + skill_manager.set_tool_executor(tool_executor) + skill_tools = skill_manager.get_skill_tools( + tool_executor=tool_executor, + ) + tool_list.extend(skill_tools) + if tool_list: + tool_executor = ToolExecutor(tool_list, bus) + # Plan 2B I3: capture optimized few-shot examples so + # _run_agent can forward them to tool-using agents. + skill_few_shot_examples = skill_manager.get_few_shot_examples() + except Exception as exc: + logger.warning("Failed to initialize skills: %s", exc) + # Resolve agent name agent_name = self._agent_name or config.agent.default_agent @@ -777,8 +815,12 @@ class SystemBuilder: agent_scheduler=agent_scheduler, agent_executor=agent_executor, speech_backend=speech_backend, + skill_manager=skill_manager, ) system._learning_orchestrator = learning_orchestrator + # Plan 2B I3: stash few-shot examples on the system so _run_agent + # can include them in agent_kwargs for tool-using agents. + system._skill_few_shot_examples = skill_few_shot_examples # Transfer MCP clients so JarvisSystem.close() can shut them down system._mcp_clients = list(getattr(self, "_mcp_clients", [])) # Wire system reference — must happen before scheduler.start() diff --git a/src/openjarvis/tools/_stubs.py b/src/openjarvis/tools/_stubs.py index 295f7d3f..06b7f129 100644 --- a/src/openjarvis/tools/_stubs.py +++ b/src/openjarvis/tools/_stubs.py @@ -269,6 +269,13 @@ class ToolExecutor: # Emit end event if self._bus: result_text = str(result.content)[:10240] if result.content else "" + # Pass through ToolResult.metadata so downstream consumers + # (TraceCollector → TraceStep.metadata → SkillOptimizer) can + # see skill-tagged invocations. Filter to JSON-serializable + # values only — internal objects like TaintSet (added by the + # taint auto-detect above) must not leak to event subscribers + # since the trace store will JSON-serialize them later. + event_metadata = self._json_safe_metadata(result.metadata) self._bus.publish( EventType.TOOL_CALL_END, { @@ -276,11 +283,41 @@ class ToolExecutor: "success": result.success, "latency": latency, "result": result_text, + "metadata": event_metadata, }, ) return result + @staticmethod + def _json_safe_metadata(metadata: Optional[Dict[str, Any]]) -> Dict[str, Any]: + """Return a copy of *metadata* containing only JSON-serializable values. + + ``ToolExecutor`` annotates ``ToolResult.metadata`` with internal + objects (currently ``_taint: TaintSet``). Those are useful for + in-process security checks but cannot be serialized when the + ``TraceCollector`` writes ``TraceStep.metadata`` to JSON in the + SQLite trace store. This helper drops any keys whose value is + not JSON-safe — silently, since the missing data is not + load-bearing for downstream consumers. + """ + if not metadata: + return {} + + import json + + safe: Dict[str, Any] = {} + for key, value in metadata.items(): + if not isinstance(key, str): + continue + try: + json.dumps(value) + except (TypeError, ValueError): + # Skip non-serializable values (e.g. TaintSet) + continue + safe[key] = value + return safe + def available_tools(self) -> List[ToolSpec]: """Return specs for all available tools.""" return [t.spec for t in self._tools.values()] diff --git a/src/openjarvis/traces/collector.py b/src/openjarvis/traces/collector.py index 5b95f19c..96f17311 100644 --- a/src/openjarvis/traces/collector.py +++ b/src/openjarvis/traces/collector.py @@ -182,6 +182,10 @@ class TraceCollector: def _on_tool_end(self, event: Any) -> None: start = getattr(self, "_tool_start_time", event.timestamp) start_data = getattr(self, "_tool_start_data", {}) + # Pull through any metadata the tool attached to its ToolResult + # (e.g. SkillTool's skill/skill_source/skill_kind tags) so the + # SkillOptimizer can bucket traces by skill name. + result_metadata = event.data.get("metadata") or {} self._current_steps.append( TraceStep( step_type=StepType.TOOL_CALL, @@ -197,6 +201,7 @@ class TraceCollector: "success": event.data.get("success", False), "result": event.data.get("result", ""), }, + metadata=dict(result_metadata), ) ) diff --git a/tests/cli/test_bench_skills_command.py b/tests/cli/test_bench_skills_command.py new file mode 100644 index 00000000..76509fee --- /dev/null +++ b/tests/cli/test_bench_skills_command.py @@ -0,0 +1,115 @@ +"""Tests for jarvis bench skills CLI command (Plan 2B).""" + +from __future__ import annotations + +from pathlib import Path +from unittest.mock import patch + +from click.testing import CliRunner + +from openjarvis.cli import cli + + +class TestBenchSkillsCommand: + def test_help(self) -> None: + result = CliRunner().invoke(cli, ["bench", "skills", "--help"]) + assert result.exit_code == 0 + assert "condition" in result.output.lower() + + def test_runs_all_conditions_with_mocked_runner(self, tmp_path: Path) -> None: + from openjarvis.evals.skill_benchmark import ( + ConditionComparison, + ConditionResult, + SkillBenchmarkConfig, + ) + + # Build a fake comparison the mocked runner will return + fake_results = { + cond: ConditionResult( + condition=cond, + seeds=[42], + per_seed_pass_rate={42: 0.5}, + mean_pass_rate=0.5, + stddev_pass_rate=0.0, + per_task_results={}, + skill_invocation_counts={}, + total_tokens=10, + total_runtime_seconds=1.0, + ) + for cond in ( + "no_skills", + "skills_on", + "skills_optimized_dspy", + "skills_optimized_gepa", + ) + } + fake_cmp = ConditionComparison( + config=SkillBenchmarkConfig(output_dir=tmp_path), + started_at="2026-04-08T00:00:00Z", + ended_at="2026-04-08T00:01:00Z", + results=fake_results, + deltas={"skills_on - no_skills": 0.0}, + ) + + with patch( + "openjarvis.evals.skill_benchmark.SkillBenchmarkRunner.run_all_conditions", + return_value=fake_cmp, + ): + with patch( + "openjarvis.evals.skill_benchmark.SkillBenchmarkRunner.write_report", + return_value=tmp_path / "fake-report.md", + ): + result = CliRunner().invoke( + cli, + [ + "bench", + "skills", + "--max-samples", + "1", + "--seeds", + "42", + "--output-dir", + str(tmp_path), + ], + ) + assert result.exit_code == 0, result.output + assert "no_skills" in result.output + assert "skills_optimized_dspy" in result.output + + def test_runs_single_condition(self, tmp_path: Path) -> None: + from openjarvis.evals.skill_benchmark import ConditionResult + + fake_result = ConditionResult( + condition="skills_optimized_dspy", + seeds=[42], + per_seed_pass_rate={42: 0.7}, + mean_pass_rate=0.7, + stddev_pass_rate=0.0, + per_task_results={}, + skill_invocation_counts={}, + total_tokens=20, + total_runtime_seconds=2.0, + ) + + with patch( + "openjarvis.evals.skill_benchmark.SkillBenchmarkRunner.run_condition", + return_value=fake_result, + ): + result = CliRunner().invoke( + cli, + [ + "bench", + "skills", + "--condition", + "skills_optimized_dspy", + "--seeds", + "42", + "--max-samples", + "1", + "--output-dir", + str(tmp_path), + ], + ) + assert result.exit_code == 0, result.output + assert "skills_optimized_dspy" in result.output + assert "0.700" in result.output or "0.7" in result.output diff --git a/tests/cli/test_optimize_skills_command.py b/tests/cli/test_optimize_skills_command.py new file mode 100644 index 00000000..45ffaa50 --- /dev/null +++ b/tests/cli/test_optimize_skills_command.py @@ -0,0 +1,82 @@ +"""Tests for jarvis optimize skills CLI command (Plan 2A).""" + +from __future__ import annotations + +from pathlib import Path +from unittest.mock import patch + +from click.testing import CliRunner + +from openjarvis.cli import cli + + +class TestOptimizeSkillsCommand: + def test_help(self) -> None: + result = CliRunner().invoke(cli, ["optimize", "skills", "--help"]) + assert result.exit_code == 0 + + def test_dry_run_no_traces(self, tmp_path: Path) -> None: + class _EmptyStore: + def list_traces(self, *, limit: int = 100, **kwargs): + return [] + + with patch( + "openjarvis.cli.optimize_cmd._get_trace_store", + return_value=_EmptyStore(), + ): + result = CliRunner().invoke(cli, ["optimize", "skills", "--dry-run"]) + assert result.exit_code == 0 + + def test_optimize_runs_with_mocked_optimizer(self, tmp_path: Path) -> None: + from openjarvis.core.types import StepType, Trace, TraceStep + from openjarvis.learning.agents.skill_optimizer import ( + SkillOptimizationResult, + ) + + def _trace(skill_name="research-skill"): + return Trace( + query=f"q for {skill_name}", + steps=[ + TraceStep( + step_type=StepType.TOOL_CALL, + timestamp=0.0, + input={"tool": f"skill_{skill_name}", "arguments": {}}, + output={"success": True, "result": "ok"}, + metadata={ + "skill": skill_name, + "skill_kind": "instructional", + }, + ), + ], + outcome="success", + feedback=1.0, + result="ok", + ) + + class _Store: + def list_traces(self, *, limit: int = 100, **kwargs): + return [_trace() for _ in range(25)] + + fake_results = { + "research-skill": SkillOptimizationResult( + skill_name="research-skill", + status="optimized", + trace_count=25, + overlay_path=tmp_path / "research-skill" / "optimized.toml", + ), + } + + with patch( + "openjarvis.cli.optimize_cmd._get_trace_store", + return_value=_Store(), + ): + with patch( + "openjarvis.learning.agents.skill_optimizer.SkillOptimizer.optimize", + return_value=fake_results, + ): + result = CliRunner().invoke( + cli, ["optimize", "skills", "--policy", "dspy"] + ) + assert result.exit_code == 0 + assert "research-skill" in result.output + assert "optimized" in result.output diff --git a/tests/cli/test_skill_cmd.py b/tests/cli/test_skill_cmd.py index a91bc563..7dea402b 100644 --- a/tests/cli/test_skill_cmd.py +++ b/tests/cli/test_skill_cmd.py @@ -2,6 +2,10 @@ from __future__ import annotations +import textwrap +from pathlib import Path +from unittest.mock import patch + from click.testing import CliRunner from openjarvis.cli import cli @@ -27,3 +31,225 @@ class TestSkillCmd: assert "install" in result.output assert "remove" in result.output assert "search" in result.output + + def test_skill_run_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "run", "--help"]) + assert result.exit_code == 0 + + def test_skill_info_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "info", "--help"]) + assert result.exit_code == 0 + + def test_skill_update_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "update", "--help"]) + assert result.exit_code == 0 + + def test_skill_list_shows_discovered_skills(self, tmp_path: Path) -> None: + skill_dir = tmp_path / "my_skill" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "my_skill" + description = "A test skill" + + [[skill.steps]] + tool_name = "echo" + output_key = "x" + """) + ) + with patch( + "openjarvis.cli.skill_cmd._get_skill_paths", + return_value=[tmp_path], + ): + result = CliRunner().invoke(cli, ["skill", "list"]) + assert result.exit_code == 0 + assert "my_skill" in result.output + + def test_skill_info_shows_details(self, tmp_path: Path) -> None: + skill_dir = tmp_path / "info_skill" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "info_skill" + description = "Detailed skill" + author = "test_author" + tags = ["research", "test"] + required_capabilities = ["network:fetch"] + + [[skill.steps]] + tool_name = "echo" + output_key = "x" + """) + ) + with patch( + "openjarvis.cli.skill_cmd._get_skill_paths", + return_value=[tmp_path], + ): + result = CliRunner().invoke(cli, ["skill", "info", "info_skill"]) + assert result.exit_code == 0 + assert "info_skill" in result.output + assert "test_author" in result.output + + +class TestSkillInstallCommand: + def test_install_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "install", "--help"]) + assert result.exit_code == 0 + assert "source" in result.output.lower() + + def test_install_invalid_source_format(self) -> None: + result = CliRunner().invoke(cli, ["skill", "install", "invalid"]) + assert result.exit_code != 0 + assert "format" in result.output.lower() or "source" in result.output.lower() + + +class TestSkillSyncCommand: + def test_sync_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "sync", "--help"]) + assert result.exit_code == 0 + + +class TestSkillSourcesCommand: + def test_sources_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "sources", "--help"]) + assert result.exit_code == 0 + + def test_sources_lists_configured_sources(self, tmp_path: Path) -> None: + from unittest.mock import patch + + from openjarvis.core.config import ( + JarvisConfig, + SkillsConfig, + SkillSourceConfig, + ) + + cfg = JarvisConfig() + cfg.skills = SkillsConfig( + sources=[ + SkillSourceConfig(source="hermes"), + SkillSourceConfig(source="openclaw"), + ] + ) + with patch("openjarvis.cli.skill_cmd.load_config", return_value=cfg): + result = CliRunner().invoke(cli, ["skill", "sources"]) + assert result.exit_code == 0 + assert "hermes" in result.output + assert "openclaw" in result.output + + +class TestSkillDiscoverCommand: + def test_discover_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "discover", "--help"]) + assert result.exit_code == 0 + + def test_discover_dry_run_no_traces(self, tmp_path: Path) -> None: + from unittest.mock import patch + + # Patch the trace store to return nothing + class _EmptyStore: + def list_traces(self, *, limit: int = 100, **kwargs): + return [] + + with patch( + "openjarvis.cli.skill_cmd._get_trace_store", + return_value=_EmptyStore(), + ): + result = CliRunner().invoke(cli, ["skill", "discover", "--dry-run"]) + assert result.exit_code == 0 + + def test_discover_writes_when_not_dry_run(self, tmp_path: Path) -> None: + from unittest.mock import patch + + from openjarvis.core.types import StepType, Trace, TraceStep + + def _trace(): + return Trace( + query="q", + steps=[ + TraceStep( + step_type=StepType.TOOL_CALL, + timestamp=0.0, + input={"tool": "web_search", "arguments": {}}, + output={"success": True, "result": "ok"}, + metadata={}, + ), + TraceStep( + step_type=StepType.TOOL_CALL, + timestamp=0.0, + input={"tool": "calculator", "arguments": {}}, + output={"success": True, "result": "ok"}, + metadata={}, + ), + ], + outcome="success", + feedback=1.0, + ) + + class _Store: + def list_traces(self, *, limit: int = 100, **kwargs): + return [_trace() for _ in range(5)] + + output_dir = tmp_path / "discovered" + with patch( + "openjarvis.cli.skill_cmd._get_trace_store", + return_value=_Store(), + ): + with patch( + "openjarvis.cli.skill_cmd._get_discovered_dir", + return_value=output_dir, + ): + result = CliRunner().invoke( + cli, ["skill", "discover", "--min-frequency", "3"] + ) + assert result.exit_code == 0 + assert output_dir.exists() + # At least one manifest should have been written + assert any(output_dir.rglob("skill.toml")) + + +class TestSkillShowOverlayCommand: + def test_show_overlay_help(self) -> None: + result = CliRunner().invoke(cli, ["skill", "show-overlay", "--help"]) + assert result.exit_code == 0 + + def test_show_overlay_missing(self, tmp_path: Path) -> None: + from unittest.mock import patch + + with patch( + "openjarvis.cli.skill_cmd._get_overlay_dir", + return_value=tmp_path / "no-such", + ): + result = CliRunner().invoke(cli, ["skill", "show-overlay", "nonexistent"]) + assert result.exit_code != 0 + + def test_show_overlay_displays_optimized(self, tmp_path: Path) -> None: + from unittest.mock import patch + + from openjarvis.skills.overlay import SkillOverlay, write_overlay + + write_overlay( + SkillOverlay( + skill_name="research-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=42, + description="Better description", + few_shot=[{"input": "q", "output": "a"}], + ), + tmp_path, + ) + + with patch( + "openjarvis.cli.skill_cmd._get_overlay_dir", + return_value=tmp_path, + ): + result = CliRunner().invoke( + cli, ["skill", "show-overlay", "research-skill"] + ) + assert result.exit_code == 0 + assert "research-skill" in result.output + assert "Better description" in result.output + assert "42" in result.output + assert "dspy" in result.output diff --git a/tests/conftest.py b/tests/conftest.py index 52b90820..67f8a534 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -19,6 +19,7 @@ from openjarvis.core.registry import ( MemoryRegistry, ModelRegistry, RouterPolicyRegistry, + SkillRegistry, SpeechRegistry, ToolRegistry, TTSRegistry, @@ -40,6 +41,7 @@ def _clean_registries() -> None: CompressionRegistry.clear() ConnectorRegistry.clear() TTSRegistry.clear() + SkillRegistry.clear() reset_event_bus() diff --git a/tests/core/test_config_skills_sources.py b/tests/core/test_config_skills_sources.py new file mode 100644 index 00000000..2df0c965 --- /dev/null +++ b/tests/core/test_config_skills_sources.py @@ -0,0 +1,43 @@ +"""Tests for skills.sources config section.""" + +from __future__ import annotations + +from openjarvis.core.config import SkillsConfig, SkillSourceConfig + + +class TestSkillSourceConfig: + def test_default_filter_empty(self): + cfg = SkillSourceConfig(source="hermes") + assert cfg.source == "hermes" + assert cfg.filter == {} + assert cfg.auto_update is False + assert cfg.url == "" + + def test_with_filter(self): + cfg = SkillSourceConfig( + source="hermes", + filter={"category": ["research", "coding"]}, + ) + assert cfg.filter["category"] == ["research", "coding"] + + +class TestSkillsConfigWithSources: + def test_default_no_sources(self): + cfg = SkillsConfig() + assert cfg.enabled is True + assert cfg.auto_sync is False + assert cfg.sources == [] + + def test_auto_sync_can_be_enabled(self): + cfg = SkillsConfig(auto_sync=True) + assert cfg.auto_sync is True + + def test_can_add_sources(self): + cfg = SkillsConfig( + sources=[ + SkillSourceConfig(source="hermes"), + SkillSourceConfig(source="openclaw"), + ] + ) + assert len(cfg.sources) == 2 + assert cfg.sources[0].source == "hermes" diff --git a/tests/core/test_skills_learning_config.py b/tests/core/test_skills_learning_config.py new file mode 100644 index 00000000..be849b8d --- /dev/null +++ b/tests/core/test_skills_learning_config.py @@ -0,0 +1,40 @@ +"""Tests for SkillsLearningConfig and its wiring into LearningConfig.""" + +from __future__ import annotations + +from openjarvis.core.config import LearningConfig, SkillsLearningConfig + + +class TestSkillsLearningConfig: + def test_defaults(self): + cfg = SkillsLearningConfig() + assert cfg.auto_optimize is False + assert cfg.optimizer == "dspy" + assert cfg.min_traces_per_skill == 20 + assert cfg.optimization_interval_seconds == 86400 + assert cfg.overlay_dir == "~/.openjarvis/learning/skills/" + + def test_can_be_constructed_with_all_fields(self): + cfg = SkillsLearningConfig( + auto_optimize=True, + optimizer="gepa", + min_traces_per_skill=10, + optimization_interval_seconds=3600, + overlay_dir="/tmp/overlays/", + ) + assert cfg.auto_optimize is True + assert cfg.optimizer == "gepa" + assert cfg.min_traces_per_skill == 10 + assert cfg.optimization_interval_seconds == 3600 + assert cfg.overlay_dir == "/tmp/overlays/" + + +class TestLearningConfigSkillsField: + def test_skills_field_present(self): + cfg = LearningConfig() + assert hasattr(cfg, "skills") + assert isinstance(cfg.skills, SkillsLearningConfig) + + def test_skills_field_default_disabled(self): + cfg = LearningConfig() + assert cfg.skills.auto_optimize is False diff --git a/tests/evals/test_jarvis_agent_backend_skills.py b/tests/evals/test_jarvis_agent_backend_skills.py new file mode 100644 index 00000000..fb7966d4 --- /dev/null +++ b/tests/evals/test_jarvis_agent_backend_skills.py @@ -0,0 +1,70 @@ +"""Tests for JarvisAgentBackend skills_enabled / overlay_dir kwargs (Plan 2B).""" + +from __future__ import annotations + +from pathlib import Path + + +class TestJarvisAgentBackendSkillsKwargs: + def test_default_skills_enabled_true(self): + from openjarvis.evals.backends.jarvis_agent import JarvisAgentBackend + + # Construction should not raise even if no model/engine is reachable + # because we don't call generate(). The kwargs themselves should be + # accepted and applied to the SystemBuilder config. + try: + backend = JarvisAgentBackend( + engine_key="ollama", + agent_name="native_react", + tools=[], + skills_enabled=True, + ) + except RuntimeError: + # If no engine is available we still consider the test passing + # for this kwarg-acceptance check; the kwarg was applied before + # the engine resolution failed. + return + assert backend._system.config.skills.enabled is True + + def test_skills_enabled_false_disables_skills(self): + from openjarvis.evals.backends.jarvis_agent import JarvisAgentBackend + + try: + backend = JarvisAgentBackend( + engine_key="ollama", + agent_name="native_react", + tools=[], + skills_enabled=False, + ) + except RuntimeError: + return + assert backend._system.config.skills.enabled is False + assert backend._system.skill_manager is None + + def test_overlay_dir_kwarg_applied_to_config(self, tmp_path: Path): + from openjarvis.evals.backends.jarvis_agent import JarvisAgentBackend + + custom = tmp_path / "custom-overlays" + try: + backend = JarvisAgentBackend( + engine_key="ollama", + agent_name="native_react", + tools=[], + skills_enabled=True, + overlay_dir=custom, + ) + except RuntimeError: + return + # The config should reflect our overlay_dir override + assert backend._system.config.learning.skills.overlay_dir == str(custom) + + def test_init_signature_accepts_kwargs_without_engine(self): + """Even without a real engine, the __init__ signature must accept + the new kwargs without raising TypeError.""" + import inspect + + from openjarvis.evals.backends.jarvis_agent import JarvisAgentBackend + + sig = inspect.signature(JarvisAgentBackend.__init__) + assert "skills_enabled" in sig.parameters + assert "overlay_dir" in sig.parameters diff --git a/tests/evals/test_skill_benchmark.py b/tests/evals/test_skill_benchmark.py new file mode 100644 index 00000000..f1fb5215 --- /dev/null +++ b/tests/evals/test_skill_benchmark.py @@ -0,0 +1,382 @@ +"""Tests for SkillBenchmarkRunner (Plan 2B).""" + +from __future__ import annotations + +from pathlib import Path +from typing import Any + +import pytest + + +class TestSkillBenchmarkConfigDefaults: + def test_defaults(self): + from openjarvis.evals.skill_benchmark import SkillBenchmarkConfig + + cfg = SkillBenchmarkConfig() + assert cfg.benchmark == "pinchbench" + assert cfg.model == "qwen3.5:9b" + assert cfg.engine == "ollama" + assert cfg.agent == "native_react" + assert cfg.seeds == [42, 43, 44] + assert cfg.max_samples is None + assert "shell_exec" in cfg.tools + assert "web_search" in cfg.tools + + def test_construct_with_overrides(self): + from openjarvis.evals.skill_benchmark import SkillBenchmarkConfig + + cfg = SkillBenchmarkConfig( + benchmark="pinchbench", + model="other-model", + seeds=[1, 2], + max_samples=5, + ) + assert cfg.model == "other-model" + assert cfg.seeds == [1, 2] + assert cfg.max_samples == 5 + + +class TestConditionResult: + def test_create(self): + from openjarvis.evals.skill_benchmark import ConditionResult + + r = ConditionResult( + condition="no_skills", + seeds=[42, 43, 44], + per_seed_pass_rate={42: 0.30, 43: 0.32, 44: 0.28}, + mean_pass_rate=0.30, + stddev_pass_rate=0.02, + per_task_results={"task_001": [True, False, True]}, + skill_invocation_counts={}, + total_tokens=1000, + total_runtime_seconds=120.0, + ) + assert r.condition == "no_skills" + assert r.mean_pass_rate == 0.30 + assert r.skill_invocation_counts == {} + + +class TestConditionComparison: + def test_create(self): + from openjarvis.evals.skill_benchmark import ( + ConditionComparison, + ConditionResult, + SkillBenchmarkConfig, + ) + + cfg = SkillBenchmarkConfig() + cmp = ConditionComparison( + config=cfg, + started_at="2026-04-08T00:00:00Z", + ended_at="2026-04-08T01:00:00Z", + results={ + "no_skills": ConditionResult( + condition="no_skills", + seeds=[42], + per_seed_pass_rate={42: 0.30}, + mean_pass_rate=0.30, + stddev_pass_rate=0.0, + per_task_results={}, + skill_invocation_counts={}, + total_tokens=0, + total_runtime_seconds=0.0, + ), + }, + deltas={}, + ) + assert "no_skills" in cmp.results + assert cmp.config.benchmark == "pinchbench" + + +class TestBuildBackendForCondition: + def _make_runner(self, tmp_path: Path): + from openjarvis.evals.skill_benchmark import ( + SkillBenchmarkConfig, + SkillBenchmarkRunner, + ) + + cfg = SkillBenchmarkConfig( + engine="ollama", + model="qwen3.5:9b", + tools=[], + seeds=[42], + output_dir=tmp_path, + overlay_dir_dspy=tmp_path / "dspy", + overlay_dir_gepa=tmp_path / "gepa", + ) + return SkillBenchmarkRunner(cfg) + + def test_unknown_condition_raises(self, tmp_path: Path): + runner = self._make_runner(tmp_path) + with pytest.raises(ValueError, match="condition"): + runner._backend_kwargs_for_condition("not_a_condition") + + def test_no_skills_kwargs(self, tmp_path: Path): + runner = self._make_runner(tmp_path) + kw = runner._backend_kwargs_for_condition("no_skills") + assert kw["skills_enabled"] is False + assert kw["overlay_dir"] is None + + def test_skills_on_kwargs(self, tmp_path: Path): + runner = self._make_runner(tmp_path) + kw = runner._backend_kwargs_for_condition("skills_on") + assert kw["skills_enabled"] is True + # skills_on uses an empty/missing overlay dir so no overlays load + assert kw["overlay_dir"] is not None + assert "skills_on_empty_overlays" in str(kw["overlay_dir"]) + + def test_skills_optimized_dspy_kwargs(self, tmp_path: Path): + runner = self._make_runner(tmp_path) + kw = runner._backend_kwargs_for_condition("skills_optimized_dspy") + assert kw["skills_enabled"] is True + assert kw["overlay_dir"] == tmp_path / "dspy" + + def test_skills_optimized_gepa_kwargs(self, tmp_path: Path): + runner = self._make_runner(tmp_path) + kw = runner._backend_kwargs_for_condition("skills_optimized_gepa") + assert kw["skills_enabled"] is True + assert kw["overlay_dir"] == tmp_path / "gepa" + + +class TestRunCondition: + def _make_runner(self, tmp_path: Path): + from openjarvis.evals.skill_benchmark import ( + SkillBenchmarkConfig, + SkillBenchmarkRunner, + ) + + cfg = SkillBenchmarkConfig( + seeds=[42, 43], + max_samples=2, + output_dir=tmp_path, + ) + return SkillBenchmarkRunner(cfg) + + def test_run_condition_aggregates_per_seed( + self, tmp_path: Path, monkeypatch: Any + ) -> None: + """run_condition runs the eval once per seed and aggregates the + results into a ConditionResult.""" + from openjarvis.evals.skill_benchmark import ( + SkillBenchmarkRunner, + ) + + runner = self._make_runner(tmp_path) + + # Stub _run_single_seed to return synthetic per-seed data + seed_results = { + 42: { + "pass_rate": 0.5, + "per_task": {"task_001": True, "task_002": False}, + "skill_invocations": {"research-skill": 1}, + "total_tokens": 100, + "total_runtime_seconds": 10.0, + }, + 43: { + "pass_rate": 1.0, + "per_task": {"task_001": True, "task_002": True}, + "skill_invocations": {"research-skill": 2}, + "total_tokens": 200, + "total_runtime_seconds": 20.0, + }, + } + + def fake_run_single_seed(self_unused, condition, seed): + return seed_results[seed] + + monkeypatch.setattr( + SkillBenchmarkRunner, + "_run_single_seed", + fake_run_single_seed, + ) + + result = runner.run_condition("no_skills") + + assert result.condition == "no_skills" + assert result.seeds == [42, 43] + assert result.per_seed_pass_rate == {42: 0.5, 43: 1.0} + assert result.mean_pass_rate == 0.75 + assert result.stddev_pass_rate > 0.0 + # Per-task aggregation: each task gets a list of [seed1_pass, seed2_pass] + assert result.per_task_results["task_001"] == [True, True] + assert result.per_task_results["task_002"] == [False, True] + # Skill invocation counts summed across seeds + assert result.skill_invocation_counts["research-skill"] == 3 + assert result.total_tokens == 300 + assert result.total_runtime_seconds == 30.0 + + def test_run_condition_single_seed_zero_stddev( + self, tmp_path: Path, monkeypatch: Any + ) -> None: + from openjarvis.evals.skill_benchmark import ( + SkillBenchmarkConfig, + SkillBenchmarkRunner, + ) + + cfg = SkillBenchmarkConfig( + seeds=[42], + max_samples=1, + output_dir=tmp_path, + ) + runner = SkillBenchmarkRunner(cfg) + + def fake_run_single_seed(self_unused, condition, seed): + return { + "pass_rate": 0.42, + "per_task": {"task_001": False}, + "skill_invocations": {}, + "total_tokens": 50, + "total_runtime_seconds": 5.0, + } + + monkeypatch.setattr( + SkillBenchmarkRunner, + "_run_single_seed", + fake_run_single_seed, + ) + + result = runner.run_condition("no_skills") + assert result.mean_pass_rate == 0.42 + # Single seed → stddev is 0 + assert result.stddev_pass_rate == 0.0 + + +class TestRunAllConditions: + def test_run_all_conditions_invokes_all_four( + self, tmp_path: Path, monkeypatch: Any + ) -> None: + from openjarvis.evals.skill_benchmark import ( + ConditionResult, + SkillBenchmarkConfig, + SkillBenchmarkRunner, + ) + + cfg = SkillBenchmarkConfig( + seeds=[42], + max_samples=1, + output_dir=tmp_path, + ) + runner = SkillBenchmarkRunner(cfg) + + invoked: list = [] + + def fake_run_condition(self_unused, condition): + invoked.append(condition) + return ConditionResult( + condition=condition, + seeds=[42], + per_seed_pass_rate={42: 0.5}, + mean_pass_rate=0.5, + stddev_pass_rate=0.0, + per_task_results={}, + skill_invocation_counts={}, + total_tokens=10, + total_runtime_seconds=1.0, + ) + + monkeypatch.setattr( + SkillBenchmarkRunner, + "run_condition", + fake_run_condition, + ) + + comparison = runner.run_all_conditions() + assert set(invoked) == { + "no_skills", + "skills_on", + "skills_optimized_dspy", + "skills_optimized_gepa", + } + assert len(comparison.results) == 4 + + def test_run_all_conditions_computes_deltas( + self, tmp_path: Path, monkeypatch: Any + ) -> None: + from openjarvis.evals.skill_benchmark import ( + ConditionResult, + SkillBenchmarkConfig, + SkillBenchmarkRunner, + ) + + cfg = SkillBenchmarkConfig( + seeds=[42], + max_samples=1, + output_dir=tmp_path, + ) + runner = SkillBenchmarkRunner(cfg) + + rates = { + "no_skills": 0.30, + "skills_on": 0.35, + "skills_optimized_dspy": 0.40, + "skills_optimized_gepa": 0.38, + } + + def fake_run_condition(self_unused, condition): + return ConditionResult( + condition=condition, + seeds=[42], + per_seed_pass_rate={42: rates[condition]}, + mean_pass_rate=rates[condition], + stddev_pass_rate=0.0, + per_task_results={}, + skill_invocation_counts={}, + total_tokens=0, + total_runtime_seconds=0.0, + ) + + monkeypatch.setattr( + SkillBenchmarkRunner, + "run_condition", + fake_run_condition, + ) + + comparison = runner.run_all_conditions() + # Deltas use the same names as the conditions + assert comparison.deltas["skills_on - no_skills"] == pytest.approx(0.05) + assert comparison.deltas["skills_optimized_dspy - skills_on"] == pytest.approx( + 0.05 + ) + assert comparison.deltas["skills_optimized_gepa - skills_on"] == pytest.approx( + 0.03 + ) + + +class TestWriteReport: + def test_write_report_creates_dated_markdown(self, tmp_path: Path) -> None: + from openjarvis.evals.skill_benchmark import ( + ConditionComparison, + ConditionResult, + SkillBenchmarkConfig, + SkillBenchmarkRunner, + ) + + cfg = SkillBenchmarkConfig(output_dir=tmp_path) + runner = SkillBenchmarkRunner(cfg) + + result = ConditionResult( + condition="no_skills", + seeds=[42, 43], + per_seed_pass_rate={42: 0.30, 43: 0.32}, + mean_pass_rate=0.31, + stddev_pass_rate=0.014, + per_task_results={"task_001": [True, False]}, + skill_invocation_counts={}, + total_tokens=100, + total_runtime_seconds=12.5, + ) + cmp = ConditionComparison( + config=cfg, + started_at="2026-04-08T00:00:00Z", + ended_at="2026-04-08T00:01:00Z", + results={"no_skills": result}, + deltas={}, + ) + + path = runner.write_report(cmp) + assert path.exists() + assert path.suffix == ".md" + content = path.read_text() + assert "no_skills" in content + assert "0.31" in content or "31.0" in content + assert "task_001" in content diff --git a/tests/learning/test_orchestrator_skill_optimization.py b/tests/learning/test_orchestrator_skill_optimization.py new file mode 100644 index 00000000..a330eda5 --- /dev/null +++ b/tests/learning/test_orchestrator_skill_optimization.py @@ -0,0 +1,146 @@ +"""Tests for LearningOrchestrator opt-in skill optimization (Plan 2A).""" + +from __future__ import annotations + +from pathlib import Path +from unittest.mock import MagicMock, patch + + +class TestOrchestratorSkillAutoOptimize: + def test_auto_optimize_disabled_by_default_does_not_call_skill_optimizer( + self, tmp_path: Path + ) -> None: + from openjarvis.learning.learning_orchestrator import ( + LearningOrchestrator, + ) + + store = MagicMock() + store.list_traces.return_value = [] + + orchestrator = LearningOrchestrator( + trace_store=store, + config_dir=tmp_path, + ) + + with patch( + "openjarvis.learning.agents.skill_optimizer.SkillOptimizer.optimize" + ) as mock_optimize: + orchestrator._maybe_optimize_skills(auto_optimize=False) + mock_optimize.assert_not_called() + + def test_auto_optimize_enabled_calls_skill_optimizer(self, tmp_path: Path) -> None: + from openjarvis.learning.learning_orchestrator import ( + LearningOrchestrator, + ) + + store = MagicMock() + store.list_traces.return_value = [] + + orchestrator = LearningOrchestrator( + trace_store=store, + config_dir=tmp_path, + ) + + with patch( + "openjarvis.learning.agents.skill_optimizer.SkillOptimizer.optimize", + return_value={}, + ) as mock_optimize: + orchestrator._maybe_optimize_skills( + auto_optimize=True, + optimizer="dspy", + min_traces_per_skill=5, + ) + mock_optimize.assert_called_once() + + +class TestOrchestratorRunSkillTrigger: + """End-to-end: LearningOrchestrator.run() invokes _maybe_optimize_skills + when learning.skills.auto_optimize is true (Plan 2A C2 fix).""" + + def _make_store(self) -> MagicMock: + store = MagicMock() + # Just enough surface area for orchestrator.run() to short-circuit + store.list_traces.return_value = [] + return store + + def _make_config(self, *, auto_optimize: bool): + from openjarvis.core.config import ( + JarvisConfig, + LearningConfig, + SkillsLearningConfig, + ) + + cfg = JarvisConfig() + cfg.learning = LearningConfig() + cfg.learning.skills = SkillsLearningConfig( + auto_optimize=auto_optimize, + optimizer="dspy", + min_traces_per_skill=5, + ) + return cfg + + def test_run_does_not_call_skill_optimizer_when_disabled( + self, tmp_path: Path + ) -> None: + from openjarvis.learning.learning_orchestrator import ( + LearningOrchestrator, + ) + + store = self._make_store() + orchestrator = LearningOrchestrator( + trace_store=store, + config_dir=tmp_path, + ) + + with patch( + "openjarvis.core.config.load_config", + return_value=self._make_config(auto_optimize=False), + ): + with patch( + "openjarvis.learning.agents.skill_optimizer.SkillOptimizer.optimize" + ) as mock_optimize: + orchestrator.run() + mock_optimize.assert_not_called() + + def test_run_calls_skill_optimizer_when_enabled( + self, tmp_path: Path + ) -> None: + from openjarvis.learning.agents.skill_optimizer import ( + SkillOptimizationResult, + ) + from openjarvis.learning.learning_orchestrator import ( + LearningOrchestrator, + ) + + store = self._make_store() + orchestrator = LearningOrchestrator( + trace_store=store, + config_dir=tmp_path, + ) + + fake_results = { + "research-skill": SkillOptimizationResult( + skill_name="research-skill", + status="optimized", + trace_count=10, + ), + } + + with patch( + "openjarvis.core.config.load_config", + return_value=self._make_config(auto_optimize=True), + ): + with patch( + "openjarvis.learning.agents.skill_optimizer.SkillOptimizer.optimize", + return_value=fake_results, + ) as mock_optimize: + result = orchestrator.run() + mock_optimize.assert_called_once() + # The orchestrator should record the skill optimization + # results in the returned dict + assert "skill_optimization" in result + assert "research-skill" in result["skill_optimization"] + assert ( + result["skill_optimization"]["research-skill"]["status"] + == "optimized" + ) diff --git a/tests/learning/test_skill_optimizer.py b/tests/learning/test_skill_optimizer.py new file mode 100644 index 00000000..23c7323c --- /dev/null +++ b/tests/learning/test_skill_optimizer.py @@ -0,0 +1,141 @@ +"""Tests for SkillOptimizer (Plan 2A) — buckets traces by skill, runs DSPy/GEPA.""" + +from __future__ import annotations + +from pathlib import Path +from typing import Any, List + +from openjarvis.core.events import EventBus +from openjarvis.core.types import StepType, Trace, TraceStep +from openjarvis.skills.manager import SkillManager + + +class _FakeTraceStore: + def __init__(self, traces: List[Trace]) -> None: + self._traces = traces + + def list_traces(self, *, limit: int = 100, **_kwargs: Any) -> List[Trace]: + return list(self._traces[:limit]) + + +def _make_skill_trace(skill_name: str, feedback: float = 1.0) -> Trace: + return Trace( + query=f"query for {skill_name}", + agent="native_react", + model="qwen3.5:9b", + engine="ollama", + steps=[ + TraceStep( + step_type=StepType.TOOL_CALL, + timestamp=0.0, + input={"tool": f"skill_{skill_name}", "arguments": {}}, + output={"success": True, "result": "ok"}, + metadata={"skill": skill_name, "skill_kind": "instructional"}, + ), + ], + outcome="success", + feedback=feedback, + result="result text", + ) + + +def _make_manager_with_skill(name: str, tmp_path: Path) -> SkillManager: + skill_dir = tmp_path / "skills" / name + skill_dir.mkdir(parents=True) + (skill_dir / "SKILL.md").write_text( + f"---\nname: {name}\ndescription: Original description\n---\nBody" + ) + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path / "skills"]) + return mgr + + +class TestSkillOptimizerBucketing: + def test_buckets_traces_by_skill_name(self, tmp_path: Path): + from openjarvis.learning.agents.skill_optimizer import SkillOptimizer + + traces = [ + _make_skill_trace("research-skill"), + _make_skill_trace("research-skill"), + _make_skill_trace("code-skill"), + ] + optimizer = SkillOptimizer(min_traces_per_skill=1) + buckets = optimizer._bucket_traces_by_skill(traces) + + assert "research-skill" in buckets + assert "code-skill" in buckets + assert len(buckets["research-skill"]) == 2 + assert len(buckets["code-skill"]) == 1 + + def test_skips_traces_without_skill_metadata(self, tmp_path: Path): + from openjarvis.learning.agents.skill_optimizer import SkillOptimizer + + # A trace with no skill metadata in the tool call + plain_trace = Trace( + query="plain", + steps=[ + TraceStep( + step_type=StepType.TOOL_CALL, + timestamp=0.0, + input={"tool": "calculator", "arguments": {}}, + output={"success": True, "result": "42"}, + metadata={}, + ), + ], + ) + optimizer = SkillOptimizer(min_traces_per_skill=1) + buckets = optimizer._bucket_traces_by_skill([plain_trace]) + assert buckets == {} + + +class TestSkillOptimizerOptimize: + def test_skips_skills_below_min_traces(self, tmp_path: Path): + from openjarvis.learning.agents.skill_optimizer import SkillOptimizer + + traces = [_make_skill_trace("research-skill") for _ in range(3)] + store = _FakeTraceStore(traces) + mgr = _make_manager_with_skill("research-skill", tmp_path) + + optimizer = SkillOptimizer(min_traces_per_skill=20) + results = optimizer.optimize(store, mgr, overlay_dir=tmp_path / "overlays") + + assert "research-skill" in results + assert results["research-skill"].status == "skipped" + assert results["research-skill"].trace_count == 3 + + def test_optimizes_skill_with_enough_traces( + self, tmp_path: Path, monkeypatch: Any + ) -> None: + from openjarvis.learning.agents.skill_optimizer import ( + SkillOptimizer, + _OptimizerOutput, + ) + + traces = [_make_skill_trace("research-skill") for _ in range(25)] + store = _FakeTraceStore(traces) + mgr = _make_manager_with_skill("research-skill", tmp_path) + + # Mock the underlying DSPy call + def fake_run(self_unused, skill_name, skill_traces): + return _OptimizerOutput( + description="An optimized description", + few_shot=[ + {"input": "hello", "output": "world"}, + ], + ) + + monkeypatch.setattr(SkillOptimizer, "_run_dspy", fake_run) + + optimizer = SkillOptimizer(min_traces_per_skill=10) + results = optimizer.optimize(store, mgr, overlay_dir=tmp_path / "overlays") + + assert results["research-skill"].status == "optimized" + assert results["research-skill"].trace_count == 25 + # Overlay file should exist + overlay_path = tmp_path / "overlays" / "research-skill" / "optimized.toml" + assert overlay_path.exists() + # And contain the optimized description + content = overlay_path.read_text() + assert "An optimized description" in content + assert "hello" in content + assert "world" in content diff --git a/tests/prompt/test_system_prompt_builder_few_shot.py b/tests/prompt/test_system_prompt_builder_few_shot.py new file mode 100644 index 00000000..244b2248 --- /dev/null +++ b/tests/prompt/test_system_prompt_builder_few_shot.py @@ -0,0 +1,37 @@ +"""Tests for SystemPromptBuilder skill few-shot injection (Plan 2A).""" + +from __future__ import annotations + +from openjarvis.prompt.builder import SystemPromptBuilder + + +class TestSystemPromptBuilderFewShot: + def test_no_few_shot_no_section(self): + builder = SystemPromptBuilder( + agent_template="You are an agent.", + ) + prompt = builder.build() + assert "## Skill Examples" not in prompt + + def test_few_shot_section_appears_in_prompt(self): + builder = SystemPromptBuilder( + agent_template="You are an agent.", + skill_few_shot_examples=[ + "### research-skill\nInput: q1\nOutput: a1", + "### code-skill\nInput: q2\nOutput: a2", + ], + ) + prompt = builder.build() + assert "## Skill Examples" in prompt + assert "research-skill" in prompt + assert "code-skill" in prompt + assert "Input: q1" in prompt + assert "Output: a2" in prompt + + def test_empty_few_shot_list_no_section(self): + builder = SystemPromptBuilder( + agent_template="You are an agent.", + skill_few_shot_examples=[], + ) + prompt = builder.build() + assert "## Skill Examples" not in prompt diff --git a/tests/skills/test_cli_install.py b/tests/skills/test_cli_install.py new file mode 100644 index 00000000..c86e5174 --- /dev/null +++ b/tests/skills/test_cli_install.py @@ -0,0 +1,83 @@ +"""End-to-end CLI tests for skill install + sync against fake sources.""" + +from __future__ import annotations + +import textwrap +from pathlib import Path +from unittest.mock import patch + +from click.testing import CliRunner + +from openjarvis.cli import cli + + +def _build_fake_hermes_cache(cache_root: Path) -> None: + """Build a minimal fake Hermes cache for testing.""" + skills_root = cache_root / "skills" + (skills_root / "research" / "research-skill").mkdir(parents=True) + (skills_root / "research" / "research-skill" / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: research-skill + description: A research skill + --- + + Use the Bash tool to fetch data, then Read the file. + """) + ) + # Fake .git so HermesResolver doesn't try to clone + (cache_root / ".git").mkdir() + + +class TestCliInstallE2E: + def test_install_hermes_skill_e2e(self, tmp_path: Path) -> None: + """jarvis skill install hermes:research-skill installs to target dir.""" + from openjarvis.skills.sources.hermes import HermesResolver + + cache = tmp_path / "hermes-cache" + target = tmp_path / "target" + _build_fake_hermes_cache(cache) + + # Create a HermesResolver pointed at the fake cache + def _make_resolver(*_args, **_kwargs): + r = HermesResolver.__new__(HermesResolver) + r._cache_root = cache + return r + + # Patch the helper that builds resolvers in the CLI + with patch( + "openjarvis.cli.skill_cmd._get_resolver", + lambda src, url="": _make_resolver(), + ): + # Patch HermesResolver.sync to no-op (cache is already built) + with patch.object(HermesResolver, "sync", lambda self: None): + # Patch the SkillImporter constructor used inside the install + # command so it writes to our test target instead of + # ~/.openjarvis/skills/. + from openjarvis.skills.importer import SkillImporter as _SI + + original_init = _SI.__init__ + + def patched_init(self, parser, tool_translator, target_root=None): + original_init( + self, + parser=parser, + tool_translator=tool_translator, + target_root=target, + ) + + with patch.object(_SI, "__init__", patched_init): + result = CliRunner().invoke( + cli, ["skill", "install", "hermes:research-skill"] + ) + + assert result.exit_code == 0, result.output + assert "Installed" in result.output + + installed = target / "hermes" / "research-skill" / "SKILL.md" + assert installed.exists(), ( + f"Expected {installed} to exist; output:\n{result.output}" + ) + body = installed.read_text() + assert "shell_exec" in body + assert "file_read" in body diff --git a/tests/skills/test_dependency.py b/tests/skills/test_dependency.py new file mode 100644 index 00000000..78e35891 --- /dev/null +++ b/tests/skills/test_dependency.py @@ -0,0 +1,251 @@ +"""Tests for dependency graph and cycle detection (Task 3).""" + +from __future__ import annotations + +import pytest + +from openjarvis.skills.types import SkillManifest, SkillStep + + +def _manifest(name, depends=None, capabilities=None, steps=None): + return SkillManifest( + name=name, + depends=depends or [], + required_capabilities=capabilities or [], + steps=steps or [], + ) + + +class TestBuildDependencyGraph: + def test_no_dependencies(self): + from openjarvis.skills.dependency import build_dependency_graph + + skills = {"a": _manifest("a")} + graph = build_dependency_graph(skills) + assert graph == {"a": set()} + + def test_simple_dependency(self): + from openjarvis.skills.dependency import build_dependency_graph + + skills = { + "a": _manifest("a"), + "b": _manifest("b", depends=["a"]), + } + graph = build_dependency_graph(skills) + assert "a" in graph["b"] + assert graph["a"] == set() + + def test_step_skill_name_edges(self): + from openjarvis.skills.dependency import build_dependency_graph + + steps = [SkillStep(skill_name="base_skill", output_key="r")] + skills = { + "base_skill": _manifest("base_skill"), + "composite": _manifest("composite", steps=steps), + } + graph = build_dependency_graph(skills) + assert "base_skill" in graph["composite"] + + def test_combined_depends_and_steps(self): + from openjarvis.skills.dependency import build_dependency_graph + + steps = [SkillStep(skill_name="step_dep", output_key="r")] + skills = { + "explicit_dep": _manifest("explicit_dep"), + "step_dep": _manifest("step_dep"), + "main": _manifest("main", depends=["explicit_dep"], steps=steps), + } + graph = build_dependency_graph(skills) + assert "explicit_dep" in graph["main"] + assert "step_dep" in graph["main"] + + def test_multiple_skills_no_overlap(self): + from openjarvis.skills.dependency import build_dependency_graph + + skills = { + "x": _manifest("x"), + "y": _manifest("y"), + } + graph = build_dependency_graph(skills) + assert graph["x"] == set() + assert graph["y"] == set() + + +class TestValidateDependencies: + def test_valid_topological_sort_simple(self): + from openjarvis.skills.dependency import validate_dependencies + + skills = { + "a": _manifest("a"), + "b": _manifest("b", depends=["a"]), + } + order = validate_dependencies(skills) + assert order.index("a") < order.index("b") + + def test_valid_topological_sort_chain(self): + from openjarvis.skills.dependency import validate_dependencies + + skills = { + "a": _manifest("a"), + "b": _manifest("b", depends=["a"]), + "c": _manifest("c", depends=["b"]), + } + order = validate_dependencies(skills) + assert order.index("a") < order.index("b") + assert order.index("b") < order.index("c") + + def test_no_dependencies_returns_all_skills(self): + from openjarvis.skills.dependency import validate_dependencies + + skills = { + "x": _manifest("x"), + "y": _manifest("y"), + } + order = validate_dependencies(skills) + assert set(order) == {"x", "y"} + + def test_cycle_detection_raises(self): + from openjarvis.skills.dependency import ( + DependencyCycleError, + validate_dependencies, + ) + + skills = { + "a": _manifest("a", depends=["b"]), + "b": _manifest("b", depends=["a"]), + } + with pytest.raises(DependencyCycleError): + validate_dependencies(skills) + + def test_cycle_detection_three_nodes(self): + from openjarvis.skills.dependency import ( + DependencyCycleError, + validate_dependencies, + ) + + skills = { + "a": _manifest("a", depends=["c"]), + "b": _manifest("b", depends=["a"]), + "c": _manifest("c", depends=["b"]), + } + with pytest.raises(DependencyCycleError): + validate_dependencies(skills) + + def test_depth_exceeded_raises(self): + from openjarvis.skills.dependency import ( + DepthExceededError, + validate_dependencies, + ) + + # Chain of depth 6 exceeds default max_depth=5 + skills = { + "a": _manifest("a"), + "b": _manifest("b", depends=["a"]), + "c": _manifest("c", depends=["b"]), + "d": _manifest("d", depends=["c"]), + "e": _manifest("e", depends=["d"]), + "f": _manifest("f", depends=["e"]), + } + with pytest.raises(DepthExceededError): + validate_dependencies(skills, max_depth=5) + + def test_depth_within_limit_passes(self): + from openjarvis.skills.dependency import validate_dependencies + + # Chain of depth 3 is within max_depth=5 + skills = { + "a": _manifest("a"), + "b": _manifest("b", depends=["a"]), + "c": _manifest("c", depends=["b"]), + } + order = validate_dependencies(skills, max_depth=5) + assert len(order) == 3 + + def test_missing_dependency_silently_skipped(self): + from openjarvis.skills.dependency import validate_dependencies + + # "missing" is not in skills — should not raise + skills = { + "a": _manifest("a", depends=["missing"]), + } + order = validate_dependencies(skills) + assert "a" in order + + def test_custom_max_depth(self): + from openjarvis.skills.dependency import ( + DepthExceededError, + validate_dependencies, + ) + + skills = { + "a": _manifest("a"), + "b": _manifest("b", depends=["a"]), + "c": _manifest("c", depends=["b"]), + } + # Chain depth 2 exceeds max_depth=1 + with pytest.raises(DepthExceededError): + validate_dependencies(skills, max_depth=1) + + +class TestComputeCapabilityUnion: + def test_single_skill_no_deps(self): + from openjarvis.skills.dependency import compute_capability_union + + skills = { + "a": _manifest("a", capabilities=["read_files"]), + } + caps = compute_capability_union("a", skills) + assert "read_files" in caps + + def test_transitive_capabilities(self): + from openjarvis.skills.dependency import compute_capability_union + + skills = { + "a": _manifest("a", capabilities=["network"]), + "b": _manifest("b", depends=["a"], capabilities=["disk_write"]), + } + caps = compute_capability_union("b", skills) + assert "network" in caps + assert "disk_write" in caps + + def test_deep_transitive_capabilities(self): + from openjarvis.skills.dependency import compute_capability_union + + skills = { + "a": _manifest("a", capabilities=["cap_a"]), + "b": _manifest("b", depends=["a"], capabilities=["cap_b"]), + "c": _manifest("c", depends=["b"], capabilities=["cap_c"]), + } + caps = compute_capability_union("c", skills) + assert "cap_a" in caps + assert "cap_b" in caps + assert "cap_c" in caps + + def test_no_capabilities(self): + from openjarvis.skills.dependency import compute_capability_union + + skills = { + "a": _manifest("a"), + "b": _manifest("b", depends=["a"]), + } + caps = compute_capability_union("b", skills) + assert caps == [] + + def test_dedup_capabilities(self): + from openjarvis.skills.dependency import compute_capability_union + + # Both a and b claim "network" + skills = { + "a": _manifest("a", capabilities=["network"]), + "b": _manifest("b", depends=["a"], capabilities=["network", "disk"]), + } + caps = compute_capability_union("b", skills) + assert caps.count("network") == 1 + assert "disk" in caps + + def test_missing_skill_returns_empty(self): + from openjarvis.skills.dependency import compute_capability_union + + skills = {} + caps = compute_capability_union("nonexistent", skills) + assert caps == [] diff --git a/tests/skills/test_importer.py b/tests/skills/test_importer.py new file mode 100644 index 00000000..758e3678 --- /dev/null +++ b/tests/skills/test_importer.py @@ -0,0 +1,204 @@ +"""Tests for SkillImporter — installs ResolvedSkill instances on disk.""" + +from __future__ import annotations + +from pathlib import Path + +from openjarvis.skills.importer import SkillImporter +from openjarvis.skills.parser import SkillParser +from openjarvis.skills.sources.base import ResolvedSkill +from openjarvis.skills.tool_translator import ToolTranslator + + +def _make_resolved(tmp_path: Path, body: str = "Body") -> ResolvedSkill: + """Create a fake source skill directory and return a ResolvedSkill.""" + src_dir = tmp_path / "source" / "my-skill" + src_dir.mkdir(parents=True) + (src_dir / "SKILL.md").write_text( + f"---\nname: my-skill\ndescription: A test skill\n---\n{body}" + ) + return ResolvedSkill( + name="my-skill", + source="hermes", + path=src_dir, + category="testing", + description="A test skill", + commit="abc123", + ) + + +class TestImportSkill: + def test_imports_to_sourced_subdir(self, tmp_path: Path): + target_root = tmp_path / "skills" + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + resolved = _make_resolved(tmp_path) + result = importer.import_skill(resolved) + + assert result.success + target = target_root / "hermes" / "my-skill" + assert target.exists() + assert (target / "SKILL.md").exists() + + def test_writes_source_metadata_file(self, tmp_path: Path): + target_root = tmp_path / "skills" + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + resolved = _make_resolved(tmp_path) + importer.import_skill(resolved) + + source_file = target_root / "hermes" / "my-skill" / ".source" + assert source_file.exists() + content = source_file.read_text() + assert "source = " in content + assert "abc123" in content + assert "scripts_imported = false" in content + + def test_translates_tool_references_in_body(self, tmp_path: Path): + target_root = tmp_path / "skills" + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + resolved = _make_resolved( + tmp_path, body="First use the Bash tool, then Read the file." + ) + result = importer.import_skill(resolved) + + installed = target_root / "hermes" / "my-skill" / "SKILL.md" + body = installed.read_text() + assert "shell_exec" in body + assert "file_read" in body + assert "Bash" not in body + assert "Bash->shell_exec" in str(result.translated_tools) + + def test_scripts_skipped_by_default(self, tmp_path: Path): + target_root = tmp_path / "skills" + # Source has a scripts/ directory + src_dir = tmp_path / "source" / "my-skill" + src_dir.mkdir(parents=True) + (src_dir / "SKILL.md").write_text("---\nname: my-skill\ndescription: x\n---\n") + scripts_dir = src_dir / "scripts" + scripts_dir.mkdir() + (scripts_dir / "helper.py").write_text("print('hi')") + + resolved = ResolvedSkill( + name="my-skill", + source="hermes", + path=src_dir, + category="x", + description="x", + commit="a", + ) + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + result = importer.import_skill(resolved) + + target = target_root / "hermes" / "my-skill" + assert not (target / "scripts").exists() + assert result.scripts_imported is False + + def test_scripts_imported_with_flag(self, tmp_path: Path): + target_root = tmp_path / "skills" + src_dir = tmp_path / "source" / "my-skill" + src_dir.mkdir(parents=True) + (src_dir / "SKILL.md").write_text("---\nname: my-skill\ndescription: x\n---\n") + scripts_dir = src_dir / "scripts" + scripts_dir.mkdir() + (scripts_dir / "helper.py").write_text("print('hi')") + + resolved = ResolvedSkill( + name="my-skill", + source="hermes", + path=src_dir, + category="x", + description="x", + commit="a", + ) + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + result = importer.import_skill(resolved, with_scripts=True) + + target = target_root / "hermes" / "my-skill" + assert (target / "scripts" / "helper.py").exists() + assert result.scripts_imported is True + + def test_references_assets_always_copied(self, tmp_path: Path): + target_root = tmp_path / "skills" + src_dir = tmp_path / "source" / "my-skill" + src_dir.mkdir(parents=True) + (src_dir / "SKILL.md").write_text("---\nname: my-skill\ndescription: x\n---\n") + (src_dir / "references").mkdir() + (src_dir / "references" / "REFERENCE.md").write_text("# Reference") + (src_dir / "assets").mkdir() + (src_dir / "assets" / "template.txt").write_text("template") + + resolved = ResolvedSkill( + name="my-skill", + source="hermes", + path=src_dir, + category="x", + description="x", + commit="a", + ) + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + importer.import_skill(resolved) + + target = target_root / "hermes" / "my-skill" + assert (target / "references" / "REFERENCE.md").exists() + assert (target / "assets" / "template.txt").exists() + + def test_force_overwrites_existing_install(self, tmp_path: Path): + target_root = tmp_path / "skills" + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + resolved = _make_resolved(tmp_path, body="Original body") + importer.import_skill(resolved) + + # Modify source and re-import with force + (resolved.path / "SKILL.md").write_text( + "---\nname: my-skill\ndescription: A test skill\n---\nUpdated body" + ) + importer.import_skill(resolved, force=True) + + installed = target_root / "hermes" / "my-skill" / "SKILL.md" + assert "Updated body" in installed.read_text() + + def test_install_without_force_skips_existing(self, tmp_path: Path): + target_root = tmp_path / "skills" + importer = SkillImporter( + parser=SkillParser(), + tool_translator=ToolTranslator(), + target_root=target_root, + ) + resolved = _make_resolved(tmp_path, body="Original") + importer.import_skill(resolved) + + (resolved.path / "SKILL.md").write_text( + "---\nname: my-skill\ndescription: A test skill\n---\nNew body" + ) + result = importer.import_skill(resolved, force=False) + assert result.skipped + + installed = target_root / "hermes" / "my-skill" / "SKILL.md" + assert "Original" in installed.read_text() diff --git a/tests/skills/test_index.py b/tests/skills/test_index.py new file mode 100644 index 00000000..b8d10427 --- /dev/null +++ b/tests/skills/test_index.py @@ -0,0 +1,105 @@ +"""Tests for skill index — sync and search.""" + +from __future__ import annotations + +import textwrap +from pathlib import Path + +from openjarvis.skills.index import SkillIndex + + +class TestSkillIndex: + def test_load_index(self, tmp_path: Path): + index_file = tmp_path / "index.toml" + index_file.write_text( + textwrap.dedent("""\ + [[skills]] + name = "research" + version = "0.1.0" + description = "Research a topic" + author = "openjarvis" + source = "github.com/openjarvis/skills/research" + sha256 = "abc123" + tags = ["research"] + required_capabilities = ["network:fetch"] + + [[skills]] + name = "summarize" + version = "0.2.0" + description = "Summarize text" + author = "community" + source = "github.com/user/summarize" + sha256 = "def456" + tags = ["nlp"] + required_capabilities = [] + """) + ) + index = SkillIndex(tmp_path) + assert len(index.entries) == 2 + assert index.entries["research"].version == "0.1.0" + + def test_search_by_name(self, tmp_path: Path): + index_file = tmp_path / "index.toml" + index_file.write_text( + textwrap.dedent("""\ + [[skills]] + name = "web_research" + version = "0.1.0" + description = "Search the web" + author = "openjarvis" + source = "github.com/openjarvis/skills/web_research" + sha256 = "abc" + tags = ["research"] + required_capabilities = [] + + [[skills]] + name = "code_review" + version = "0.1.0" + description = "Review code" + author = "openjarvis" + source = "github.com/openjarvis/skills/code_review" + sha256 = "def" + tags = ["coding"] + required_capabilities = [] + """) + ) + index = SkillIndex(tmp_path) + results = index.search("research") + assert len(results) == 1 + assert results[0].name == "web_research" + + def test_search_by_tag(self, tmp_path: Path): + index_file = tmp_path / "index.toml" + index_file.write_text( + textwrap.dedent("""\ + [[skills]] + name = "skill_a" + version = "0.1.0" + description = "First" + author = "x" + source = "github.com/x/a" + sha256 = "a" + tags = ["nlp", "research"] + required_capabilities = [] + """) + ) + index = SkillIndex(tmp_path) + results = index.search("nlp") + assert len(results) == 1 + + def test_search_no_results(self, tmp_path: Path): + index_file = tmp_path / "index.toml" + index_file.write_text( + "[[skills]]\n" + 'name = "a"\nversion = "0.1.0"\n' + 'description = ""\nauthor = ""\n' + 'source = ""\nsha256 = ""\n' + "tags = []\nrequired_capabilities = []\n" + ) + index = SkillIndex(tmp_path) + results = index.search("zzzzz") + assert results == [] + + def test_missing_index_file(self, tmp_path: Path): + index = SkillIndex(tmp_path) + assert index.entries == {} diff --git a/tests/skills/test_integration_live.py b/tests/skills/test_integration_live.py new file mode 100644 index 00000000..c1fd5713 --- /dev/null +++ b/tests/skills/test_integration_live.py @@ -0,0 +1,175 @@ +"""Live integration tests for the skills system. + +These tests require a running Ollama instance with qwen3.5:4b. +Mark: live +""" + +from __future__ import annotations + +import pytest + +from openjarvis.core.events import EventBus, EventType +from openjarvis.skills.manager import SkillManager +from openjarvis.skills.tool_adapter import SkillTool +from openjarvis.system import SystemBuilder + + +@pytest.mark.live +class TestSkillSystemIntegration: + """Integration tests verifying skills flow end-to-end with a real engine.""" + + def test_system_builder_discovers_skills(self): + """SystemBuilder.build() discovers installed skills and adds them to tools.""" + system = SystemBuilder().engine("ollama").model("qwen3.5:4b").build() + try: + assert system.skill_manager is not None, "SkillManager should be created" + skill_names = system.skill_manager.skill_names() + assert len(skill_names) > 0, f"Should discover skills, got: {skill_names}" + print(f" Discovered skills: {skill_names}") + + # Verify skill tools are in the system tools list + skill_tools = [t for t in system.tools if isinstance(t, SkillTool)] + assert len(skill_tools) > 0, "Skill tools should be in system tools" + print(f" Skill tools in system: {[t.spec.name for t in skill_tools]}") + finally: + if hasattr(system, "close"): + try: + system.close() + except Exception: + pass + + def test_skill_catalog_in_system_prompt(self): + """The skill catalog XML should be generated correctly.""" + bus = EventBus() + mgr = SkillManager(bus=bus) + from pathlib import Path + + mgr.discover(paths=[Path("~/.openjarvis/skills/").expanduser()]) + + catalog = mgr.get_catalog_xml() + assert "" in catalog + assert "research-and-summarize" in catalog + assert "code-explainer" in catalog + assert "math-solver" in catalog + print(f" Catalog XML:\n{catalog}") + + def test_skill_tool_invocation_returns_content(self): + """Invoking a skill tool returns meaningful content.""" + bus = EventBus(record_history=True) + mgr = SkillManager(bus=bus) + from pathlib import Path + + mgr.discover(paths=[Path("~/.openjarvis/skills/").expanduser()]) + + # Test instruction-only skill + tools = mgr.get_skill_tools() + code_explainer = next( + (t for t in tools if "code-explainer" in t.spec.name), None + ) + assert code_explainer is not None, "code-explainer skill should exist" + + result = code_explainer.execute(task="explain a for loop") + assert result.success + assert ( + "programming language" in result.content.lower() + or "plain language" in result.content.lower() + ), f"Should return instructions, got: {result.content[:200]}" + print(f" code-explainer returned: {result.content[:200]}...") + + # Verify the bus is reachable (instruction-only skills don't run + # a pipeline, so no SKILL_EXECUTE_* events here — tool-mode skills + # emit them, covered in TestSkillEventsAndTracing below) + assert bus.history is not None + + def test_agent_ask_with_skills_available(self): + """Agent can answer a query with skills available in the tool list.""" + system = ( + SystemBuilder() + .engine("ollama") + .model("qwen3.5:4b") + .tools( + [ + "think", + "calculator", + ] + ) + .build() + ) + try: + # Skills should be auto-discovered and added + assert system.skill_manager is not None + skill_count = len(system.skill_manager.skill_names()) + total_tools = len(system.tools) + print(f" Skills: {skill_count}, Total tools: {total_tools}") + + # The tools list should include both regular tools and skill tools + tool_names = [t.spec.name for t in system.tools] + print(f" Tool names: {tool_names}") + + # Ask a simple question — agent should work normally + result = system.ask("What is 2 + 2?") + assert "content" in result + content = result["content"] + assert "4" in content, f"Expected '4' in response, got: {content[:200]}" + print(f" Agent response: {content[:200]}") + finally: + if hasattr(system, "close"): + try: + system.close() + except Exception: + pass + + +@pytest.mark.live +class TestSkillEventsAndTracing: + """Verify skills emit proper events for the Learning pipeline.""" + + def test_skill_execution_emits_events(self): + """Running a structured skill emits SKILL_EXECUTE_START/END events.""" + from openjarvis.core.types import ToolResult + from openjarvis.skills.executor import SkillExecutor + from openjarvis.skills.types import SkillManifest, SkillStep + from openjarvis.tools._stubs import BaseTool, ToolExecutor, ToolSpec + + class EchoTool(BaseTool): + tool_id = "echo" + + @property + def spec(self): + return ToolSpec(name="echo", description="Echo") + + def execute(self, **params): + return ToolResult( + tool_name="echo", + content=params.get("text", "echoed"), + success=True, + ) + + bus = EventBus(record_history=True) + te = ToolExecutor([EchoTool()]) + executor = SkillExecutor(te, bus=bus) + + manifest = SkillManifest( + name="test_traced", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "hello"}', + output_key="result", + ), + ], + ) + result = executor.run(manifest) + assert result.success + + events = [(e.event_type, e.data) for e in bus.history] + event_types = [e[0] for e in events] + assert EventType.SKILL_EXECUTE_START in event_types + assert EventType.SKILL_EXECUTE_END in event_types + + # Verify event data + start_event = next(e for e in events if e[0] == EventType.SKILL_EXECUTE_START) + assert start_event[1]["skill"] == "test_traced" + end_event = next(e for e in events if e[0] == EventType.SKILL_EXECUTE_END) + assert end_event[1]["success"] is True + print(" Events correctly emitted with skill metadata") diff --git a/tests/skills/test_loader_markdown.py b/tests/skills/test_loader_markdown.py new file mode 100644 index 00000000..99625052 --- /dev/null +++ b/tests/skills/test_loader_markdown.py @@ -0,0 +1,215 @@ +"""Tests for SKILL.md and directory-based skill loading.""" + +from __future__ import annotations + +import textwrap +from pathlib import Path + +from openjarvis.skills.loader import load_skill_directory, load_skill_markdown + + +class TestLoadSkillMarkdown: + def test_load_markdown_with_frontmatter(self, tmp_path: Path): + md = tmp_path / "SKILL.md" + md.write_text( + textwrap.dedent("""\ + --- + name: research + required_capabilities: [network:fetch] + --- + + When asked to research a topic: + 1. Break into sub-questions + 2. Search each one + """) + ) + manifest = load_skill_markdown(md) + assert manifest.name == "research" + assert manifest.required_capabilities == ["network:fetch"] + assert "Break into sub-questions" in manifest.markdown_content + + def test_load_markdown_no_frontmatter(self, tmp_path: Path): + md = tmp_path / "SKILL.md" + md.write_text("Just some instructions.\n") + manifest = load_skill_markdown(md) + assert manifest.name == "SKILL" + assert manifest.markdown_content == "Just some instructions.\n" + + +class TestLoadSkillDirectory: + def test_directory_with_toml_only(self, tmp_path: Path): + skill_dir = tmp_path / "my_skill" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "my_skill" + description = "A test skill" + tags = ["test"] + + [[skill.steps]] + tool_name = "echo" + output_key = "result" + """) + ) + manifest = load_skill_directory(skill_dir) + assert manifest.name == "my_skill" + assert manifest.tags == ["test"] + assert len(manifest.steps) == 1 + assert manifest.markdown_content == "" + + def test_directory_with_markdown_only(self, tmp_path: Path): + skill_dir = tmp_path / "guide_skill" + skill_dir.mkdir() + (skill_dir / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: guide_skill + required_capabilities: [filesystem:read] + --- + + Follow these steps when reviewing code. + """) + ) + manifest = load_skill_directory(skill_dir) + assert manifest.name == "guide_skill" + assert manifest.required_capabilities == ["filesystem:read"] + assert "reviewing code" in manifest.markdown_content + + def test_directory_with_both(self, tmp_path: Path): + skill_dir = tmp_path / "hybrid" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "hybrid" + description = "Hybrid skill" + tags = ["hybrid"] + + [[skill.steps]] + tool_name = "web_search" + arguments_template = '{"query": "{query}"}' + output_key = "results" + """) + ) + (skill_dir / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: hybrid + --- + + Present results with citations. + """) + ) + manifest = load_skill_directory(skill_dir) + assert manifest.name == "hybrid" + assert len(manifest.steps) == 1 + assert "citations" in manifest.markdown_content + + def test_directory_with_depends_and_skill_steps(self, tmp_path: Path): + skill_dir = tmp_path / "composed" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "composed" + depends = ["summarize"] + + [[skill.steps]] + tool_name = "web_search" + arguments_template = '{"query": "{query}"}' + output_key = "raw" + + [[skill.steps]] + skill_name = "summarize" + arguments_template = '{"text": "{raw}"}' + output_key = "summary" + """) + ) + manifest = load_skill_directory(skill_dir) + assert manifest.depends == ["summarize"] + assert manifest.steps[1].skill_name == "summarize" + assert manifest.steps[1].tool_name == "" + + +class TestParserDelegation: + def test_load_skill_markdown_uses_parser_for_validation(self, tmp_path: Path): + """SKILL.md loading goes through SkillParser, gaining strict validation.""" + from openjarvis.skills.loader import load_skill_markdown + + md = tmp_path / "SKILL.md" + md.write_text("---\nname: my-skill\ndescription: x\n---\nBody") + manifest = load_skill_markdown(md) + assert manifest.name == "my-skill" + assert manifest.description == "x" + assert "Body" in manifest.markdown_content + + def test_load_skill_markdown_handles_legacy_top_level_fields(self, tmp_path: Path): + """Legacy SKILL.md with top-level tags/version/etc. still works.""" + from openjarvis.skills.loader import load_skill_markdown + + md = tmp_path / "SKILL.md" + md.write_text( + "---\n" + "name: legacy-skill\n" + "description: x\n" + "version: 2.0.0\n" + "tags: [a, b]\n" + "required_capabilities: [network:fetch]\n" + "---\n" + "Body" + ) + manifest = load_skill_markdown(md) + assert manifest.version == "2.0.0" + assert manifest.tags == ["a", "b"] + assert manifest.required_capabilities == ["network:fetch"] + + +class TestLoadSkillDirectorySourcePromotion: + def test_dot_source_file_promoted_to_metadata(self, tmp_path: Path): + """A .source file in the skill directory promotes its source field + into manifest.metadata.openjarvis.source.""" + from openjarvis.skills.loader import load_skill_directory + + skill_dir = tmp_path / "imported-skill" + skill_dir.mkdir() + (skill_dir / "SKILL.md").write_text( + "---\nname: imported-skill\ndescription: x\n---\nBody" + ) + (skill_dir / ".source").write_text( + 'source = "hermes:apple-notes"\n' + 'commit = "abc123"\n' + 'category = "apple"\n' + 'installed_at = "2026-04-04T22:30:00Z"\n' + "translated_tools = []\n" + "missing_tools = []\n" + "scripts_imported = false\n" + ) + manifest = load_skill_directory(skill_dir) + oj = manifest.metadata.get("openjarvis", {}) + assert oj.get("source") == "hermes" + + def test_no_dot_source_file_no_metadata(self, tmp_path: Path): + """When no .source file is present, metadata.openjarvis.source is unset.""" + from openjarvis.skills.loader import load_skill_directory + + skill_dir = tmp_path / "user-skill" + skill_dir.mkdir() + (skill_dir / "SKILL.md").write_text( + "---\nname: user-skill\ndescription: x\n---\nBody" + ) + manifest = load_skill_directory(skill_dir) + oj = manifest.metadata.get("openjarvis", {}) + assert "source" not in oj + + def test_malformed_dot_source_does_not_crash(self, tmp_path: Path): + """A malformed .source file is ignored, not crashed on.""" + from openjarvis.skills.loader import load_skill_directory + + skill_dir = tmp_path / "bad" + skill_dir.mkdir() + (skill_dir / "SKILL.md").write_text("---\nname: bad\ndescription: x\n---\nBody") + (skill_dir / ".source").write_text("this is not valid toml = [[[") + # Should load the manifest without raising + manifest = load_skill_directory(skill_dir) + assert manifest.name == "bad" diff --git a/tests/skills/test_manager.py b/tests/skills/test_manager.py new file mode 100644 index 00000000..017fb80a --- /dev/null +++ b/tests/skills/test_manager.py @@ -0,0 +1,520 @@ +"""Tests for SkillManager — discovery, catalog, tools, and resolve.""" + +from __future__ import annotations + +import textwrap +from pathlib import Path + +import pytest + +from openjarvis.core.events import EventBus +from openjarvis.core.types import ToolResult +from openjarvis.skills.manager import SkillManager +from openjarvis.tools._stubs import BaseTool, ToolExecutor, ToolSpec # noqa: F401 + +# --------------------------------------------------------------------------- +# Helpers +# --------------------------------------------------------------------------- + + +def _write_toml_skill(directory: Path, name: str, description: str = "") -> None: + skill_dir = directory / name + skill_dir.mkdir(parents=True, exist_ok=True) + (skill_dir / "skill.toml").write_text( + textwrap.dedent(f"""\ + [skill] + name = "{name}" + description = "{description or name}" + tags = ["test"] + + [[skill.steps]] + tool_name = "echo" + arguments_template = '{{"text": "{{input}}"}}' + output_key = "result" + """) + ) + + +class EchoTool(BaseTool): + tool_id = "echo" + + @property + def spec(self) -> ToolSpec: + return ToolSpec(name="echo", description="Echo input") + + def execute(self, **params) -> ToolResult: + return ToolResult( + tool_name="echo", content=params.get("text", ""), success=True + ) + + +# --------------------------------------------------------------------------- +# TestSkillManagerDiscovery +# --------------------------------------------------------------------------- + + +class TestSkillManagerDiscovery: + def test_discover_single_dir(self, tmp_path: Path) -> None: + """discover() loads skills from a single directory.""" + _write_toml_skill(tmp_path, "alpha") + _write_toml_skill(tmp_path, "beta") + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + + names = mgr.skill_names() + assert "alpha" in names + assert "beta" in names + + def test_discover_precedence_workspace_over_user(self, tmp_path: Path) -> None: + """First path (workspace) wins when the same name appears in multiple dirs.""" + workspace = tmp_path / "workspace" + user = tmp_path / "user" + workspace.mkdir() + user.mkdir() + + # Both dirs have a skill named "shared" + _write_toml_skill(workspace, "shared", description="workspace version") + _write_toml_skill(user, "shared", description="user version") + + mgr = SkillManager(bus=EventBus()) + # workspace is listed first → highest precedence + mgr.discover(paths=[workspace, user]) + + manifest = mgr.resolve("shared") + assert manifest.description == "workspace version" + + def test_discover_empty_dir(self, tmp_path: Path) -> None: + """discover() does not raise on an empty directory.""" + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + assert mgr.skill_names() == [] + + def test_discover_nonexistent_dir(self, tmp_path: Path) -> None: + """discover() silently skips paths that do not exist.""" + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path / "does_not_exist"]) + assert mgr.skill_names() == [] + + def test_discover_multiple_dirs_accumulates(self, tmp_path: Path) -> None: + """Skills from all directories are accumulated (modulo precedence).""" + dir_a = tmp_path / "a" + dir_b = tmp_path / "b" + dir_a.mkdir() + dir_b.mkdir() + + _write_toml_skill(dir_a, "skill_a") + _write_toml_skill(dir_b, "skill_b") + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[dir_a, dir_b]) + + names = mgr.skill_names() + assert "skill_a" in names + assert "skill_b" in names + + +# --------------------------------------------------------------------------- +# TestSkillManagerCatalog +# --------------------------------------------------------------------------- + + +class TestSkillManagerCatalog: + def test_catalog_xml_contains_skill_names(self, tmp_path: Path) -> None: + """get_catalog_xml() includes discovered skill names.""" + _write_toml_skill(tmp_path, "greet", description="Greet the user") + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + + xml = mgr.get_catalog_xml() + assert "greet" in xml + assert "" in xml + + def test_catalog_xml_empty_when_no_skills(self) -> None: + """get_catalog_xml() returns a well-formed empty block when no skills loaded.""" + mgr = SkillManager(bus=EventBus()) + xml = mgr.get_catalog_xml() + assert "" in xml + assert "" in xml + + def test_catalog_excludes_model_disabled_skills(self, tmp_path: Path) -> None: + """Skills with disable_model_invocation=true are excluded from the catalog.""" + skill_dir = tmp_path / "hidden" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "hidden" + description = "Hidden skill" + disable_model_invocation = true + + [[skill.steps]] + tool_name = "echo" + output_key = "x" + """) + ) + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + xml = mgr.get_catalog_xml() + assert "hidden" not in xml + + def test_catalog_escapes_xml_special_chars(self, tmp_path: Path) -> None: + """Descriptions with XML-special chars are escaped in the catalog.""" + skill_dir = tmp_path / "special" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "special" + description = "A skill with & 'quotes'" + tags = ["test"] + + [[skill.steps]] + tool_name = "echo" + output_key = "r" + """) + ) + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + xml = mgr.get_catalog_xml() + # Raw < must not appear unescaped inside the XML body + assert "" not in xml + assert "<" in xml or "special" in xml # description is escaped or safe + + def test_catalog_includes_visible_but_not_hidden(self, tmp_path: Path) -> None: + """Catalog includes user_invocable skills and excludes model-disabled ones.""" + _write_toml_skill(tmp_path, "visible") + skill_dir = tmp_path / "invisible" + skill_dir.mkdir() + (skill_dir / "skill.toml").write_text( + textwrap.dedent("""\ + [skill] + name = "invisible" + description = "Not for models" + disable_model_invocation = true + + [[skill.steps]] + tool_name = "echo" + output_key = "r" + """) + ) + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + xml = mgr.get_catalog_xml() + assert "visible" in xml + assert "invisible" not in xml + + +# --------------------------------------------------------------------------- +# TestSkillManagerTools +# --------------------------------------------------------------------------- + + +class TestSkillManagerTools: + def test_get_skill_tools_returns_base_tool_instances(self, tmp_path: Path) -> None: + """get_skill_tools() returns a list of BaseTool instances.""" + _write_toml_skill(tmp_path, "tool_skill") + + tool_executor = ToolExecutor([EchoTool()]) + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + mgr.set_tool_executor(tool_executor) + + tools = mgr.get_skill_tools() + assert len(tools) == 1 + assert isinstance(tools[0], BaseTool) + + def test_get_skill_tools_accepts_tool_executor_kwarg(self, tmp_path: Path) -> None: + """get_skill_tools() accepts an optional tool_executor keyword argument.""" + _write_toml_skill(tmp_path, "kwarg_skill") + + tool_executor = ToolExecutor([EchoTool()]) + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + + tools = mgr.get_skill_tools(tool_executor=tool_executor) + assert len(tools) == 1 + assert isinstance(tools[0], BaseTool) + + def test_get_skill_tools_empty_when_no_skills(self) -> None: + """get_skill_tools() returns an empty list when no skills are loaded.""" + mgr = SkillManager(bus=EventBus()) + tools = mgr.get_skill_tools() + assert tools == [] + + def test_get_skill_tools_names_prefixed(self, tmp_path: Path) -> None: + """Each SkillTool has a name prefixed with 'skill_'.""" + _write_toml_skill(tmp_path, "my_skill") + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + + tools = mgr.get_skill_tools() + assert any(t.spec.name == "skill_my_skill" for t in tools) + + +# --------------------------------------------------------------------------- +# TestSkillManagerResolve +# --------------------------------------------------------------------------- + + +class TestSkillManagerResolve: + def test_resolve_existing_skill(self, tmp_path: Path) -> None: + """resolve() returns the SkillManifest for a known skill name.""" + _write_toml_skill(tmp_path, "resolve_me") + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + + manifest = mgr.resolve("resolve_me") + assert manifest.name == "resolve_me" + + def test_resolve_missing_raises_key_error(self) -> None: + """resolve() raises KeyError for an unknown skill name.""" + mgr = SkillManager(bus=EventBus()) + with pytest.raises(KeyError): + mgr.resolve("nonexistent_skill") + + def test_resolve_after_multiple_discovers(self, tmp_path: Path) -> None: + """resolve() works correctly after calling discover() multiple times.""" + dir_a = tmp_path / "a" + dir_b = tmp_path / "b" + dir_a.mkdir() + dir_b.mkdir() + + _write_toml_skill(dir_a, "skill_one") + _write_toml_skill(dir_b, "skill_two") + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[dir_a]) + mgr.discover(paths=[dir_b]) + + assert mgr.resolve("skill_one").name == "skill_one" + assert mgr.resolve("skill_two").name == "skill_two" + + +class TestSkillManagerSourcedLayout: + def test_discovers_skills_under_source_subdirs(self, tmp_path: Path): + """SkillManager.discover() finds skills in // layout.""" + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + + # Build hermes// and openclaw// subdirs + for source, name in [("hermes", "apple-notes"), ("openclaw", "etherscan")]: + d = tmp_path / source / name + d.mkdir(parents=True) + (d / "SKILL.md").write_text( + f"---\nname: {name}\ndescription: from {source}\n---\nBody" + ) + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + names = mgr.skill_names() + assert "apple-notes" in names + assert "etherscan" in names + + def test_flat_and_sourced_layout_coexist(self, tmp_path: Path): + """Both flat .// and ./// are discovered.""" + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + + # Flat layout + flat = tmp_path / "my-flat-skill" + flat.mkdir() + (flat / "SKILL.md").write_text( + "---\nname: my-flat-skill\ndescription: flat\n---\n" + ) + + # Sourced layout + sourced = tmp_path / "hermes" / "my-sourced-skill" + sourced.mkdir(parents=True) + (sourced / "SKILL.md").write_text( + "---\nname: my-sourced-skill\ndescription: sourced\n---\n" + ) + + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path]) + names = mgr.skill_names() + assert "my-flat-skill" in names + assert "my-sourced-skill" in names + + +class TestSkillManagerOverlayLoading: + def test_overlay_description_overrides_manifest(self, tmp_path: Path): + """When an overlay exists, the optimized description replaces + the manifest's description after discover().""" + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + from openjarvis.skills.overlay import SkillOverlay, write_overlay + + skill_dir = tmp_path / "skills" / "research-skill" + skill_dir.mkdir(parents=True) + (skill_dir / "SKILL.md").write_text( + "---\nname: research-skill\ndescription: Original description\n---\nBody" + ) + + overlay_dir = tmp_path / "overlays" + write_overlay( + SkillOverlay( + skill_name="research-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=25, + description="A much better optimized description", + ), + overlay_dir, + ) + + mgr = SkillManager(bus=EventBus(), overlay_dir=overlay_dir) + mgr.discover(paths=[tmp_path / "skills"]) + + manifest = mgr.resolve("research-skill") + assert manifest.description == "A much better optimized description" + + def test_overlay_few_shot_stored_in_metadata(self, tmp_path: Path): + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + from openjarvis.skills.overlay import SkillOverlay, write_overlay + + skill_dir = tmp_path / "skills" / "test-skill" + skill_dir.mkdir(parents=True) + (skill_dir / "SKILL.md").write_text( + "---\nname: test-skill\ndescription: x\n---\nBody" + ) + + overlay_dir = tmp_path / "overlays" + write_overlay( + SkillOverlay( + skill_name="test-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=20, + description="x", + few_shot=[ + {"input": "q1", "output": "a1"}, + {"input": "q2", "output": "a2"}, + ], + ), + overlay_dir, + ) + + mgr = SkillManager(bus=EventBus(), overlay_dir=overlay_dir) + mgr.discover(paths=[tmp_path / "skills"]) + + manifest = mgr.resolve("test-skill") + oj = manifest.metadata.get("openjarvis", {}) + few_shot = oj.get("few_shot", []) + assert len(few_shot) == 2 + assert few_shot[0]["input"] == "q1" + + def test_no_overlay_dir_does_not_crash(self, tmp_path: Path): + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + + skill_dir = tmp_path / "skills" / "test-skill" + skill_dir.mkdir(parents=True) + (skill_dir / "SKILL.md").write_text( + "---\nname: test-skill\ndescription: original\n---\nBody" + ) + + # No overlay_dir argument — should still work + mgr = SkillManager(bus=EventBus()) + mgr.discover(paths=[tmp_path / "skills"]) + manifest = mgr.resolve("test-skill") + assert manifest.description == "original" + + def test_get_few_shot_examples_returns_formatted_strings(self, tmp_path: Path): + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + from openjarvis.skills.overlay import SkillOverlay, write_overlay + + skill_dir = tmp_path / "skills" / "fs-skill" + skill_dir.mkdir(parents=True) + (skill_dir / "SKILL.md").write_text( + "---\nname: fs-skill\ndescription: x\n---\nBody" + ) + + overlay_dir = tmp_path / "overlays" + write_overlay( + SkillOverlay( + skill_name="fs-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=20, + description="x", + few_shot=[ + {"input": "what is X?", "output": "X is Y"}, + ], + ), + overlay_dir, + ) + + mgr = SkillManager(bus=EventBus(), overlay_dir=overlay_dir) + mgr.discover(paths=[tmp_path / "skills"]) + + examples = mgr.get_few_shot_examples() + assert len(examples) >= 1 + assert any("what is X?" in s and "X is Y" in s for s in examples) + + def test_overlay_dir_read_from_config_when_not_explicit( + self, tmp_path: Path + ) -> None: + """Plan 2A I1 fix: SkillManager picks up overlay_dir from + cfg.learning.skills.overlay_dir when no explicit value is passed.""" + from unittest.mock import patch + + from openjarvis.core.config import ( + JarvisConfig, + LearningConfig, + SkillsLearningConfig, + ) + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + + cfg = JarvisConfig() + cfg.learning = LearningConfig() + cfg.learning.skills = SkillsLearningConfig( + overlay_dir=str(tmp_path / "configured-overlays") + ) + + with patch("openjarvis.core.config.load_config", return_value=cfg): + mgr = SkillManager(bus=EventBus()) + assert ( + mgr._overlay_dir == (tmp_path / "configured-overlays").expanduser() + ) + + def test_discover_with_empty_paths_still_loads_overlays( + self, tmp_path: Path + ) -> None: + """Plan 2A I2 fix: discover() with no paths still applies overlays + to skills that were seeded by other means.""" + from openjarvis.core.events import EventBus + from openjarvis.skills.manager import SkillManager + from openjarvis.skills.overlay import SkillOverlay, write_overlay + from openjarvis.skills.types import SkillManifest + + overlay_dir = tmp_path / "overlays" + write_overlay( + SkillOverlay( + skill_name="seeded-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=15, + description="Optimized seeded description", + ), + overlay_dir, + ) + + mgr = SkillManager(bus=EventBus(), overlay_dir=overlay_dir) + # Seed a skill directly (simulating a non-disk source) + mgr._skills["seeded-skill"] = SkillManifest( + name="seeded-skill", + description="Original description", + markdown_content="Body", + ) + # Call discover() with no paths — should still load overlays + mgr.discover() + + manifest = mgr.resolve("seeded-skill") + assert manifest.description == "Optimized seeded description" diff --git a/tests/skills/test_manager_discover_from_traces.py b/tests/skills/test_manager_discover_from_traces.py new file mode 100644 index 00000000..7d5c7771 --- /dev/null +++ b/tests/skills/test_manager_discover_from_traces.py @@ -0,0 +1,111 @@ +"""Tests for SkillManager.discover_from_traces (Plan 2A).""" + +from __future__ import annotations + +from pathlib import Path +from typing import Any, List + +from openjarvis.core.events import EventBus +from openjarvis.core.types import StepType, Trace, TraceStep +from openjarvis.skills.manager import SkillManager + + +class _FakeTraceStore: + """In-memory trace store stub for tests.""" + + def __init__(self, traces: List[Trace]) -> None: + self._traces = traces + + def list_traces(self, *, limit: int = 100, **_kwargs: Any) -> List[Trace]: + return list(self._traces[:limit]) + + +def _make_trace(tools: List[str], outcome: str = "success") -> Trace: + steps = [ + TraceStep( + step_type=StepType.TOOL_CALL, + timestamp=0.0, + input={"tool": tool, "arguments": {}}, + output={"success": True, "result": ""}, + metadata={}, + ) + for tool in tools + ] + return Trace( + query="test query", + agent="native_react", + model="qwen3.5:9b", + engine="ollama", + steps=steps, + outcome=outcome, + feedback=1.0 if outcome == "success" else 0.0, + ) + + +class TestDiscoverFromTraces: + def test_writes_manifest_for_recurring_sequence(self, tmp_path: Path): + # Five identical successful traces of (web_search, calculator) + traces = [_make_trace(["web_search", "calculator"]) for _ in range(5)] + store = _FakeTraceStore(traces) + + mgr = SkillManager(bus=EventBus()) + output_dir = tmp_path / "discovered" + written = mgr.discover_from_traces( + store, min_frequency=3, output_dir=output_dir + ) + + assert len(written) >= 1 + # Discovery should produce at least one manifest file + assert any(Path(item["path"]).exists() for item in written) + + def test_respects_min_frequency_filter(self, tmp_path: Path): + # Only 2 occurrences — below default min_frequency of 3 + traces = [_make_trace(["web_search", "calculator"]) for _ in range(2)] + store = _FakeTraceStore(traces) + + mgr = SkillManager(bus=EventBus()) + output_dir = tmp_path / "discovered" + written = mgr.discover_from_traces( + store, min_frequency=3, output_dir=output_dir + ) + + assert written == [] + + def test_empty_trace_store_returns_empty_list(self, tmp_path: Path): + store = _FakeTraceStore([]) + mgr = SkillManager(bus=EventBus()) + output_dir = tmp_path / "discovered" + written = mgr.discover_from_traces(store, output_dir=output_dir) + assert written == [] + + def test_discovered_skills_use_hyphen_naming(self, tmp_path: Path): + traces = [_make_trace(["web_search", "calculator"]) for _ in range(5)] + store = _FakeTraceStore(traces) + + mgr = SkillManager(bus=EventBus()) + output_dir = tmp_path / "discovered" + written = mgr.discover_from_traces( + store, min_frequency=3, output_dir=output_dir + ) + + for item in written: + name = item["name"] + # Plan 1 spec: lowercase, only hyphens (no underscores) + assert name == name.lower() + assert "_" not in name + + def test_round_trip_discover_then_load(self, tmp_path: Path): + """Discovered skills are subsequently loadable via discover().""" + traces = [_make_trace(["web_search", "calculator"]) for _ in range(5)] + store = _FakeTraceStore(traces) + + mgr = SkillManager(bus=EventBus()) + output_dir = tmp_path / "discovered" + mgr.discover_from_traces(store, min_frequency=3, output_dir=output_dir) + + # Now load them + mgr2 = SkillManager(bus=EventBus()) + mgr2.discover(paths=[output_dir]) + # At least one of the discovered skills should be loadable + names = mgr2.skill_names() + assert len(names) >= 1 diff --git a/tests/skills/test_native_react_few_shot.py b/tests/skills/test_native_react_few_shot.py new file mode 100644 index 00000000..7da80d03 --- /dev/null +++ b/tests/skills/test_native_react_few_shot.py @@ -0,0 +1,167 @@ +"""Tests for native_react skill_few_shot_examples wiring (Plan 2B I3 fix).""" + +from __future__ import annotations + +from unittest.mock import MagicMock + +from openjarvis.agents.native_react import REACT_SYSTEM_PROMPT, NativeReActAgent + + +class _StubEngine: + def generate(self, *args, **kwargs): + return {"content": "Final Answer: stub", "usage": {}} + + +class TestToolUsingAgentAcceptsKwarg: + def test_default_empty_list(self): + agent = NativeReActAgent( + engine=_StubEngine(), + model="stub", + tools=[], + ) + assert agent._skill_few_shot_examples == [] + + def test_explicit_examples_stored(self): + examples = ["### research-skill\nInput: q\nOutput: a"] + agent = NativeReActAgent( + engine=_StubEngine(), + model="stub", + tools=[], + skill_few_shot_examples=examples, + ) + assert agent._skill_few_shot_examples == examples + + def test_none_resolves_to_empty_list(self): + agent = NativeReActAgent( + engine=_StubEngine(), + model="stub", + tools=[], + skill_few_shot_examples=None, + ) + assert agent._skill_few_shot_examples == [] + + +class TestReactSystemPromptPlaceholder: + def test_format_with_empty_examples(self): + rendered = REACT_SYSTEM_PROMPT.format( + tool_descriptions="No tools available.", + skill_examples="", + ) + assert "## Skill Examples" not in rendered + assert "No tools available." in rendered + # ReAct format spec must still be present + assert "Thought:" in rendered + assert "Action:" in rendered + assert "Final Answer:" in rendered + + def test_format_with_non_empty_examples(self): + examples_block = ( + "## Skill Examples\n\n### research-skill\nInput: q\nOutput: a\n\n" + ) + rendered = REACT_SYSTEM_PROMPT.format( + tool_descriptions="No tools available.", + skill_examples=examples_block, + ) + assert "## Skill Examples" in rendered + assert "research-skill" in rendered + assert "Input: q" in rendered + assert "Output: a" in rendered + + +class TestSystemBuilderCapturesFewShot: + def test_skill_manager_examples_stored_on_system(self, tmp_path): + """SystemBuilder.build() pulls examples from SkillManager and + stashes them on the JarvisSystem instance for _run_agent to + forward to tool-using agents.""" + + from openjarvis.skills.manager import SkillManager + from openjarvis.skills.overlay import SkillOverlay, write_overlay + from openjarvis.skills.types import SkillManifest + + # Build an overlay so the manager picks up real few-shot examples + overlay_dir = tmp_path / "overlays" + write_overlay( + SkillOverlay( + skill_name="seeded-skill", + optimizer="dspy", + optimized_at="2026-04-08T00:00:00Z", + trace_count=10, + description="Optimized", + few_shot=[{"input": "ping", "output": "pong"}], + ), + overlay_dir, + ) + + from openjarvis.core.events import EventBus + + mgr = SkillManager(bus=EventBus(), overlay_dir=overlay_dir) + mgr._skills["seeded-skill"] = SkillManifest( + name="seeded-skill", + description="Original", + markdown_content="Body", + ) + mgr.discover() # applies overlay + + # The captured examples should now contain our seeded one + examples = mgr.get_few_shot_examples() + assert len(examples) >= 1 + assert any("ping" in s and "pong" in s for s in examples) + + +class TestRunAgentForwardsExamples: + def test_run_agent_passes_examples_to_tool_using_agent(self): + """_run_agent injects system._skill_few_shot_examples into + agent_kwargs when the agent class has accepts_tools=True.""" + from unittest.mock import patch + + from openjarvis.agents._stubs import AgentResult + from openjarvis.system import JarvisSystem + + captured_kwargs: dict = {} + + class _CapturingAgent: + accepts_tools = True + + def __init__(self, engine, model, **kwargs): + captured_kwargs.update(kwargs) + + def run(self, query, context=None, **kw): + return AgentResult(content="ok", turns=1) + + # Build a minimal JarvisSystem with the captured examples + system = JarvisSystem.__new__(JarvisSystem) + system.config = MagicMock() + system.config.intelligence.temperature = 0.0 + system.config.intelligence.max_tokens = 100 + system.config.agent.max_turns = 5 + system.bus = MagicMock() + system.engine = MagicMock() + system.engine_key = "stub" + system.model = "stub" + system.tools = [] + system.tool_executor = None + system.memory_backend = None + system.channel_backend = None + system.trace_store = None + system.capability_policy = None + system.session_store = None + system._skill_few_shot_examples = ["### research-skill\nInput: q\nOutput: a"] + system._mcp_clients = [] + + with patch( + "openjarvis.core.registry.AgentRegistry.get", + return_value=_CapturingAgent, + ): + system._run_agent( + query="test", + messages=[], + agent_name="capturing", + tool_names=None, + temperature=0.0, + max_tokens=100, + ) + + assert "skill_few_shot_examples" in captured_kwargs + assert captured_kwargs["skill_few_shot_examples"] == [ + "### research-skill\nInput: q\nOutput: a" + ] diff --git a/tests/skills/test_overlay.py b/tests/skills/test_overlay.py new file mode 100644 index 00000000..27b4564e --- /dev/null +++ b/tests/skills/test_overlay.py @@ -0,0 +1,99 @@ +"""Tests for the skills overlay loader/writer (Plan 2A).""" + +from __future__ import annotations + +from pathlib import Path + +from openjarvis.skills.overlay import ( + SkillOverlay, + SkillOverlayLoader, + write_overlay, +) + + +class TestSkillOverlayDataclass: + def test_create_minimal(self): + ov = SkillOverlay( + skill_name="my-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=47, + description="Optimized description", + ) + assert ov.skill_name == "my-skill" + assert ov.optimizer == "dspy" + assert ov.few_shot == [] + + def test_create_with_few_shot(self): + ov = SkillOverlay( + skill_name="my-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=47, + description="Optimized description", + few_shot=[ + {"input": "transformer attention", "output": "## Recent Advances..."}, + ], + ) + assert len(ov.few_shot) == 1 + assert ov.few_shot[0]["input"] == "transformer attention" + + +class TestWriteOverlay: + def test_write_creates_file(self, tmp_path: Path): + overlay_dir = tmp_path / "learning" / "skills" + ov = SkillOverlay( + skill_name="test-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=20, + description="Optimized", + ) + path = write_overlay(ov, overlay_dir) + assert path.exists() + assert path.name == "optimized.toml" + assert path.parent.name == "test-skill" + + def test_write_then_read_roundtrip(self, tmp_path: Path): + overlay_dir = tmp_path / "learning" / "skills" + ov = SkillOverlay( + skill_name="test-skill", + optimizer="dspy", + optimized_at="2026-04-08T14:30:00Z", + trace_count=42, + description="Better description", + few_shot=[ + {"input": "q1", "output": "a1"}, + {"input": "q2", "output": "a2"}, + ], + ) + write_overlay(ov, overlay_dir) + + loader = SkillOverlayLoader(overlay_dir) + loaded = loader.load("test-skill") + assert loaded is not None + assert loaded.skill_name == "test-skill" + assert loaded.optimizer == "dspy" + assert loaded.trace_count == 42 + assert loaded.description == "Better description" + assert len(loaded.few_shot) == 2 + assert loaded.few_shot[0]["input"] == "q1" + + +class TestSkillOverlayLoader: + def test_load_missing_returns_none(self, tmp_path: Path): + overlay_dir = tmp_path / "learning" / "skills" + loader = SkillOverlayLoader(overlay_dir) + assert loader.load("nonexistent") is None + + def test_load_malformed_returns_none(self, tmp_path: Path): + overlay_dir = tmp_path / "learning" / "skills" / "broken" + overlay_dir.mkdir(parents=True) + (overlay_dir / "optimized.toml").write_text("not valid toml = [[[") + + loader = SkillOverlayLoader(tmp_path / "learning" / "skills") + assert loader.load("broken") is None + + def test_load_directory_does_not_exist(self, tmp_path: Path): + loader = SkillOverlayLoader(tmp_path / "does-not-exist") + assert loader.load("anything") is None diff --git a/tests/skills/test_parser.py b/tests/skills/test_parser.py new file mode 100644 index 00000000..30562c91 --- /dev/null +++ b/tests/skills/test_parser.py @@ -0,0 +1,213 @@ +"""Tests for the SkillParser strict + tolerant passes.""" + +from __future__ import annotations + +import pytest + +from openjarvis.skills.parser import SkillParseError, SkillParser + + +class TestStrictRequiredFields: + def test_missing_name_raises(self): + parser = SkillParser() + frontmatter = {"description": "x"} + with pytest.raises(SkillParseError, match="name"): + parser.parse_frontmatter(frontmatter) + + def test_missing_description_raises(self): + parser = SkillParser() + frontmatter = {"name": "test-skill"} + with pytest.raises(SkillParseError, match="description"): + parser.parse_frontmatter(frontmatter) + + def test_minimal_valid_frontmatter(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + {"name": "minimal", "description": "Does something useful"} + ) + assert manifest.name == "minimal" + assert manifest.description == "Does something useful" + + +class TestStrictNamingRules: + def test_uppercase_rejected(self): + parser = SkillParser() + with pytest.raises(SkillParseError, match="lowercase"): + parser.parse_frontmatter({"name": "MySkill", "description": "x"}) + + def test_underscore_rejected(self): + parser = SkillParser() + with pytest.raises(SkillParseError, match="hyphen"): + parser.parse_frontmatter({"name": "my_skill", "description": "x"}) + + def test_leading_hyphen_rejected(self): + parser = SkillParser() + with pytest.raises(SkillParseError, match="hyphen"): + parser.parse_frontmatter({"name": "-skill", "description": "x"}) + + def test_trailing_hyphen_rejected(self): + parser = SkillParser() + with pytest.raises(SkillParseError, match="hyphen"): + parser.parse_frontmatter({"name": "skill-", "description": "x"}) + + def test_consecutive_hyphens_rejected(self): + parser = SkillParser() + with pytest.raises(SkillParseError, match="consecutive"): + parser.parse_frontmatter({"name": "my--skill", "description": "x"}) + + def test_valid_kebab_name_accepted(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + {"name": "my-skill-123", "description": "x"} + ) + assert manifest.name == "my-skill-123" + + +class TestStrictLengthLimits: + def test_name_over_64_chars_rejected(self): + parser = SkillParser() + too_long = "a" * 65 + with pytest.raises(SkillParseError, match="64"): + parser.parse_frontmatter({"name": too_long, "description": "x"}) + + def test_description_over_1024_chars_rejected(self): + parser = SkillParser() + too_long = "a" * 1025 + with pytest.raises(SkillParseError, match="1024"): + parser.parse_frontmatter({"name": "test", "description": too_long}) + + +class TestTolerantFieldMapping: + def test_top_level_version_mapped_to_metadata(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "version": "1.2.3", + } + ) + assert manifest.version == "1.2.3" + + def test_top_level_author_mapped_to_metadata(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "author": "alice", + } + ) + assert manifest.author == "alice" + + def test_top_level_tags_mapped_to_metadata(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "tags": ["research", "nlp"], + } + ) + assert manifest.tags == ["research", "nlp"] + + def test_top_level_required_capabilities_mapped(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "required_capabilities": ["network:fetch"], + } + ) + assert manifest.required_capabilities == ["network:fetch"] + + def test_top_level_depends_mapped(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "depends": ["summarize"], + } + ) + assert manifest.depends == ["summarize"] + + def test_top_level_user_invocable_mapped(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "user_invocable": False, + } + ) + assert manifest.user_invocable is False + + def test_top_level_disable_model_invocation_mapped(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "disable_model_invocation": True, + } + ) + assert manifest.disable_model_invocation is True + + def test_metadata_openjarvis_namespace_used(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "metadata": { + "openjarvis": { + "version": "9.9.9", + "tags": ["already", "mapped"], + }, + }, + } + ) + assert manifest.version == "9.9.9" + assert manifest.tags == ["already", "mapped"] + + def test_unmapped_field_logs_warning(self, caplog): + import logging + + parser = SkillParser() + with caplog.at_level(logging.WARNING): + parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "weird_vendor_field": "something", + } + ) + assert any("weird_vendor_field" in record.message for record in caplog.records) + + def test_original_frontmatter_preserved(self): + parser = SkillParser() + original = { + "name": "test", + "description": "x", + "weird_field": "preserved", + } + manifest = parser.parse_frontmatter(original) + # Stored under metadata.openjarvis.original_frontmatter + oj = manifest.metadata.get("openjarvis", {}) + assert "original_frontmatter" in oj + assert oj["original_frontmatter"]["weird_field"] == "preserved" + + def test_platforms_list_mapped_to_compatibility_string(self): + parser = SkillParser() + manifest = parser.parse_frontmatter( + { + "name": "test", + "description": "x", + "platforms": ["macos", "linux"], + } + ) + # Platforms get rendered into a compatibility string under metadata + oj = manifest.metadata.get("openjarvis", {}) + assert "platforms" in oj or "compatibility" in oj diff --git a/tests/skills/test_security.py b/tests/skills/test_security.py new file mode 100644 index 00000000..712ad441 --- /dev/null +++ b/tests/skills/test_security.py @@ -0,0 +1,59 @@ +"""Tests for skill security — capability validation and trust tiers.""" + +from __future__ import annotations + +from openjarvis.skills.security import ( + TrustTier, + classify_trust_tier, + has_dangerous_capabilities, + validate_capabilities, +) +from openjarvis.skills.types import SkillManifest + + +class TestTrustTiers: + def test_bundled_tier(self): + assert classify_trust_tier(is_bundled=True) == TrustTier.BUNDLED + + def test_indexed_tier(self): + assert ( + classify_trust_tier(has_signature=True, in_index=True) == TrustTier.INDEXED + ) + + def test_unreviewed_tier(self): + assert ( + classify_trust_tier(has_signature=False, in_index=False) + == TrustTier.UNREVIEWED + ) + + def test_workspace_tier(self): + assert classify_trust_tier(is_workspace=True) == TrustTier.WORKSPACE + + +class TestCapabilityValidation: + def test_valid_capabilities(self): + manifest = SkillManifest(name="test", required_capabilities=["network:fetch"]) + allowed = {"network:fetch", "filesystem:read"} + assert validate_capabilities(manifest, allowed) == [] + + def test_missing_capability(self): + manifest = SkillManifest(name="test", required_capabilities=["shell:execute"]) + allowed = {"network:fetch"} + violations = validate_capabilities(manifest, allowed) + assert "shell:execute" in violations + + +class TestDangerousCapabilities: + def test_detects_dangerous(self): + manifest = SkillManifest( + name="test", required_capabilities=["shell:execute", "network:fetch"] + ) + dangerous = has_dangerous_capabilities(manifest) + assert "shell:execute" in dangerous + assert "network:fetch" not in dangerous + + def test_no_dangerous(self): + manifest = SkillManifest( + name="test", required_capabilities=["network:fetch", "filesystem:read"] + ) + assert has_dangerous_capabilities(manifest) == [] diff --git a/tests/skills/test_skill_tool_metadata_tagging.py b/tests/skills/test_skill_tool_metadata_tagging.py new file mode 100644 index 00000000..6d372c66 --- /dev/null +++ b/tests/skills/test_skill_tool_metadata_tagging.py @@ -0,0 +1,105 @@ +"""Tests for SkillTool metadata tagging (Plan 2A trace tagging).""" + +from __future__ import annotations + +from openjarvis.core.types import ToolResult +from openjarvis.skills.executor import SkillExecutor +from openjarvis.skills.tool_adapter import SkillTool +from openjarvis.skills.types import SkillManifest, SkillStep +from openjarvis.tools._stubs import BaseTool, ToolExecutor, ToolSpec + + +class _EchoTool(BaseTool): + tool_id = "echo" + + @property + def spec(self): + return ToolSpec(name="echo", description="echo") + + def execute(self, **params): + return ToolResult( + tool_name="echo", + content=params.get("text", ""), + success=True, + ) + + +class TestSkillToolMetadataTagging: + def _make_tool(self, manifest: SkillManifest) -> SkillTool: + executor = SkillExecutor(ToolExecutor([_EchoTool()])) + return SkillTool(manifest, executor) + + def test_metadata_includes_skill_name(self): + manifest = SkillManifest( + name="my-skill", + description="A test skill", + markdown_content="Just instructions", + ) + tool = self._make_tool(manifest) + result = tool.execute(task="hello") + assert result.metadata["skill"] == "my-skill" + + def test_skill_source_defaults_to_user(self): + manifest = SkillManifest( + name="my-skill", + description="A test skill", + markdown_content="Just instructions", + ) + tool = self._make_tool(manifest) + result = tool.execute(task="hello") + assert result.metadata["skill_source"] == "user" + + def test_skill_source_propagates_from_manifest(self): + manifest = SkillManifest( + name="apple-notes", + description="Apple Notes", + markdown_content="Use memo", + metadata={"openjarvis": {"source": "hermes"}}, + ) + tool = self._make_tool(manifest) + result = tool.execute(task="create a note") + assert result.metadata["skill_source"] == "hermes" + + def test_skill_kind_instructional_when_no_steps(self): + manifest = SkillManifest( + name="explainer", + description="Explains things", + markdown_content="Explain step by step", + ) + tool = self._make_tool(manifest) + result = tool.execute(task="explain a loop") + assert result.metadata["skill_kind"] == "instructional" + + def test_skill_kind_executable_when_steps_present(self): + manifest = SkillManifest( + name="echo-skill", + description="Echoes input", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "{input}"}', + output_key="result", + ), + ], + ) + tool = self._make_tool(manifest) + result = tool.execute(input="hello") + assert result.metadata["skill_kind"] == "executable" + + def test_failed_pipeline_still_tags_metadata(self): + manifest = SkillManifest( + name="broken", + description="Calls nonexistent tool", + steps=[ + SkillStep( + tool_name="nonexistent_tool", + arguments_template="{}", + output_key="x", + ), + ], + ) + tool = self._make_tool(manifest) + result = tool.execute() + assert result.success is False + assert result.metadata["skill"] == "broken" + assert result.metadata["skill_kind"] == "executable" diff --git a/tests/skills/test_skill_trace_tagging_integration.py b/tests/skills/test_skill_trace_tagging_integration.py new file mode 100644 index 00000000..f983994d --- /dev/null +++ b/tests/skills/test_skill_trace_tagging_integration.py @@ -0,0 +1,212 @@ +"""End-to-end integration test for skill trace tagging (Plan 2A C1 fix). + +Verifies the full flow: + SkillTool.execute() + → ToolExecutor publishes TOOL_CALL_END with metadata + → TraceCollector copies metadata into TraceStep.metadata + → SkillOptimizer can bucket traces by skill name + +If this test passes, the trace tagging path is wired end-to-end. +""" + +from __future__ import annotations + +import json + +from openjarvis.core.events import EventBus +from openjarvis.core.types import StepType, ToolCall, ToolResult +from openjarvis.skills.executor import SkillExecutor +from openjarvis.skills.tool_adapter import SkillTool +from openjarvis.skills.types import SkillManifest +from openjarvis.tools._stubs import BaseTool, ToolExecutor, ToolSpec + + +class TestSkillTraceTaggingEndToEnd: + def test_skill_metadata_flows_to_trace_step(self) -> None: + """SkillTool metadata reaches TraceCollector via the event bus.""" + bus = EventBus(record_history=True) + + # Subscribe a small collector that mimics TraceCollector._on_tool_end + captured_metadata: dict = {} + + def _on_tool_end(event): + nonlocal captured_metadata + captured_metadata = event.data.get("metadata", {}) + + from openjarvis.core.events import EventType + + bus.subscribe(EventType.TOOL_CALL_END, _on_tool_end) + + # Build a SkillTool wrapping an instructional manifest + manifest = SkillManifest( + name="research-skill", + description="x", + markdown_content="Just instructions.", + metadata={"openjarvis": {"source": "hermes"}}, + ) + skill_executor = SkillExecutor(ToolExecutor([], bus=bus)) + skill_tool = SkillTool(manifest, skill_executor) + + # Build a ToolExecutor that knows about the SkillTool, with the bus + tool_executor = ToolExecutor([skill_tool], bus=bus) + # Invoke through the executor (this is what the agent would do) + tool_executor.execute( + ToolCall(id="t1", name="skill_research-skill", arguments="{}") + ) + + # Now the published TOOL_CALL_END event should carry the metadata + assert captured_metadata.get("skill") == "research-skill" + assert captured_metadata.get("skill_source") == "hermes" + assert captured_metadata.get("skill_kind") == "instructional" + + def test_trace_collector_writes_metadata_to_step(self) -> None: + """A real TraceCollector populates TraceStep.metadata from the event.""" + from openjarvis.traces.collector import TraceCollector + + bus = EventBus(record_history=True) + + # Set up the SkillTool + manifest = SkillManifest( + name="my-skill", + description="x", + markdown_content="Body", + metadata={"openjarvis": {"source": "openclaw"}}, + ) + skill_executor = SkillExecutor(ToolExecutor([], bus=bus)) + skill_tool = SkillTool(manifest, skill_executor) + tool_executor = ToolExecutor([skill_tool], bus=bus) + + # Stub agent that just calls the skill once + class _StubAgent: + agent_id = "stub" + + def run(self, query, context=None, **kwargs): + from openjarvis.agents._stubs import AgentResult + + tool_executor.execute( + ToolCall( + id="t1", + name="skill_my-skill", + arguments="{}", + ) + ) + return AgentResult(content="done", turns=1) + + collector = TraceCollector( + agent=_StubAgent(), + store=None, # in-memory only + bus=bus, + ) + collector.run("test") + + # Inspect the captured trace + trace = collector.last_trace + assert trace is not None, "Collector should have captured a trace" + + tool_steps = [s for s in trace.steps if s.step_type == StepType.TOOL_CALL] + assert len(tool_steps) >= 1, "Should have captured at least one tool step" + + first = tool_steps[0] + assert first.metadata.get("skill") == "my-skill", ( + f"Expected metadata.skill='my-skill', got {first.metadata!r}" + ) + assert first.metadata.get("skill_source") == "openclaw" + assert first.metadata.get("skill_kind") == "instructional" + + +class _TaintingTool(BaseTool): + """Test tool whose result triggers the auto_detect_taint codepath. + + Returns content containing strings the security taint scanner picks + up as user-input or external — those add a `_taint: TaintSet` to + the result metadata before TOOL_CALL_END is published. + """ + + tool_id = "tainting" + + @property + def spec(self): + return ToolSpec(name="tainting", description="emit tainted output") + + def execute(self, **params): + return ToolResult( + tool_name="tainting", + content="user said: hello world", + success=True, + ) + + +class TestEventMetadataIsJsonSafe: + """Plan 2B regression: TOOL_CALL_END payload metadata must be JSON + serializable so the trace store can persist it without crashing. + + Bug surfaced by Plan 2B Task 11 smoke test, where every PinchBench + task failed with `Object of type TaintSet is not JSON serializable` + because ToolExecutor was passing through the internal `_taint` key + that the security auto-detect adds. + """ + + def test_event_metadata_excludes_non_json_objects(self): + from openjarvis.core.events import EventType + + bus = EventBus(record_history=True) + captured: dict = {} + + def _on_tool_end(event): + nonlocal captured + captured = dict(event.data.get("metadata") or {}) + + bus.subscribe(EventType.TOOL_CALL_END, _on_tool_end) + + tool_executor = ToolExecutor([_TaintingTool()], bus=bus) + tool_executor.execute(ToolCall(id="t1", name="tainting", arguments="{}")) + + # The published metadata must be JSON serializable end-to-end + # — TraceCollector will eventually feed this to TraceStore.save() + # which calls json.dumps() on TraceStep.metadata. + try: + json.dumps(captured) + except (TypeError, ValueError) as exc: + raise AssertionError( + f"Published TOOL_CALL_END metadata is not JSON serializable: " + f"{exc}. Keys: {list(captured.keys())}" + ) + + # And the internal _taint key must NOT be present (it would have + # been the offender if json.dumps had failed) + assert "_taint" not in captured, ( + f"_taint key leaked into event metadata: {captured!r}" + ) + + def test_skill_metadata_still_present_after_filtering(self): + """The JSON-safe filter must NOT drop legitimate skill metadata.""" + from openjarvis.core.events import EventType + + bus = EventBus(record_history=True) + captured: dict = {} + + def _on_tool_end(event): + nonlocal captured + captured = dict(event.data.get("metadata") or {}) + + bus.subscribe(EventType.TOOL_CALL_END, _on_tool_end) + + manifest = SkillManifest( + name="my-skill", + description="x", + markdown_content="Body", + metadata={"openjarvis": {"source": "hermes"}}, + ) + skill_executor = SkillExecutor(ToolExecutor([], bus=bus)) + skill_tool = SkillTool(manifest, skill_executor) + + tool_executor = ToolExecutor([skill_tool], bus=bus) + tool_executor.execute(ToolCall(id="t1", name="skill_my-skill", arguments="{}")) + + # Skill keys must survive the filter + assert captured.get("skill") == "my-skill" + assert captured.get("skill_source") == "hermes" + assert captured.get("skill_kind") == "instructional" + + # And it must still be JSON-serializable + json.dumps(captured) diff --git a/tests/skills/test_skills.py b/tests/skills/test_skills.py index b151558e..c39c8590 100644 --- a/tests/skills/test_skills.py +++ b/tests/skills/test_skills.py @@ -148,6 +148,129 @@ class TestSkillExecutor: assert EventType.SKILL_EXECUTE_END in event_types +class TestSkillStepExtended: + def test_step_with_skill_name(self): + step = SkillStep(skill_name="summarize", output_key="result") + assert step.skill_name == "summarize" + assert step.tool_name == "" + + def test_step_either_tool_or_skill(self): + step_tool = SkillStep(tool_name="web_search") + step_skill = SkillStep(skill_name="summarize") + assert step_tool.tool_name == "web_search" + assert step_tool.skill_name == "" + assert step_skill.skill_name == "summarize" + assert step_skill.tool_name == "" + + +class TestSkillManifestExtended: + def test_manifest_with_tags_and_depends(self): + manifest = SkillManifest( + name="test", + tags=["research"], + depends=["summarize"], + ) + assert manifest.tags == ["research"] + assert manifest.depends == ["summarize"] + + def test_manifest_with_invocation_flags(self): + manifest = SkillManifest( + name="test", + user_invocable=False, + disable_model_invocation=True, + ) + assert not manifest.user_invocable + assert manifest.disable_model_invocation + + def test_manifest_with_markdown_content(self): + manifest = SkillManifest( + name="test", + markdown_content="When asked to research...", + ) + assert manifest.markdown_content == "When asked to research..." + + def test_manifest_bytes_includes_new_fields(self): + manifest = SkillManifest( + name="test", + tags=["a"], + depends=["b"], + steps=[], + ) + data = manifest.manifest_bytes() + assert b"tags" in data + assert b"depends" in data + + +class TestSkillExecutorSubSkills: + def test_sub_skill_delegation(self): + """Executor delegates skill_name steps to a skill resolver.""" + tools = [EchoTool(), UpperTool()] + tool_executor = ToolExecutor(tools) + executor = SkillExecutor(tool_executor) + + child_manifest = SkillManifest( + name="upper_skill", + steps=[ + SkillStep( + tool_name="upper", + arguments_template='{"text": "{text}"}', + output_key="uppered", + ), + ], + ) + + def resolve_skill(name, context): + from openjarvis.skills.executor import SkillResult + + if name == "upper_skill": + return executor.run(child_manifest, initial_context=context) + return SkillResult(skill_name=name, success=False) + + executor.set_skill_resolver(resolve_skill) + + parent_manifest = SkillManifest( + name="parent", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "hello"}', + output_key="echoed", + ), + SkillStep( + skill_name="upper_skill", + arguments_template='{"text": "{echoed}"}', + output_key="result", + ), + ], + ) + result = executor.run(parent_manifest) + assert result.success + assert result.context.get("result") == "HELLO" + + def test_sub_skill_failure_stops_pipeline(self): + tools = [EchoTool()] + tool_executor = ToolExecutor(tools) + executor = SkillExecutor(tool_executor) + + def resolve_skill(name, context): + from openjarvis.skills.executor import SkillResult + + return SkillResult(skill_name=name, success=False) + + executor.set_skill_resolver(resolve_skill) + + manifest = SkillManifest( + name="parent", + steps=[ + SkillStep(skill_name="nonexistent", output_key="x"), + SkillStep(tool_name="echo", output_key="y"), + ], + ) + result = executor.run(manifest) + assert not result.success + assert len(result.step_results) == 1 + + class TestSkillTool: def test_skill_as_tool(self): from openjarvis.skills.tool_adapter import SkillTool diff --git a/tests/skills/test_sources.py b/tests/skills/test_sources.py new file mode 100644 index 00000000..d2132e1f --- /dev/null +++ b/tests/skills/test_sources.py @@ -0,0 +1,329 @@ +"""Tests for SourceResolver ABC and resolver implementations.""" + +from __future__ import annotations + +import textwrap +from pathlib import Path +from typing import List + +from openjarvis.skills.sources.base import ResolvedSkill, SourceResolver + + +class _FakeResolver(SourceResolver): + """In-memory resolver used for ABC tests.""" + + name = "fake" + + def __init__(self, skills: List[ResolvedSkill]) -> None: + self._skills = skills + + def cache_dir(self) -> Path: + return Path("/tmp/fake-cache") + + def sync(self) -> None: + pass + + def list_skills(self) -> List[ResolvedSkill]: + return list(self._skills) + + +class TestResolvedSkill: + def test_create(self): + skill = ResolvedSkill( + name="my-skill", + source="hermes", + path=Path("/tmp/x"), + category="research", + description="Does research", + commit="abc123", + ) + assert skill.name == "my-skill" + assert skill.source == "hermes" + + def test_default_sidecar_data(self): + skill = ResolvedSkill( + name="x", + source="x", + path=Path("/tmp"), + category="", + description="", + commit="", + ) + assert skill.sidecar_data == {} + + +class TestSourceResolverABC: + def test_resolve_filters_by_name(self): + skills = [ + ResolvedSkill( + name="research-it", + source="fake", + path=Path("/tmp/r"), + category="research", + description="x", + commit="a", + ), + ResolvedSkill( + name="code-it", + source="fake", + path=Path("/tmp/c"), + category="coding", + description="x", + commit="a", + ), + ] + resolver = _FakeResolver(skills) + results = resolver.resolve("research-it") + assert len(results) == 1 + assert results[0].name == "research-it" + + def test_resolve_filters_by_partial_name(self): + skills = [ + ResolvedSkill( + name="research-it", + source="fake", + path=Path("/tmp/r"), + category="research", + description="x", + commit="a", + ), + ResolvedSkill( + name="code-it", + source="fake", + path=Path("/tmp/c"), + category="coding", + description="x", + commit="a", + ), + ] + resolver = _FakeResolver(skills) + # 'res' should match 'research-it' + results = resolver.resolve("res") + assert len(results) == 1 + + def test_resolve_empty_query_returns_all(self): + skills = [ + ResolvedSkill( + name="a", + source="fake", + path=Path("/tmp/a"), + category="x", + description="x", + commit="a", + ), + ResolvedSkill( + name="b", + source="fake", + path=Path("/tmp/b"), + category="x", + description="x", + commit="a", + ), + ] + resolver = _FakeResolver(skills) + assert len(resolver.resolve("")) == 2 + + def test_filter_by_category(self): + skills = [ + ResolvedSkill( + name="a", + source="fake", + path=Path("/tmp/a"), + category="research", + description="x", + commit="a", + ), + ResolvedSkill( + name="b", + source="fake", + path=Path("/tmp/b"), + category="coding", + description="x", + commit="a", + ), + ] + resolver = _FakeResolver(skills) + results = resolver.filter_by_category("research") + assert len(results) == 1 + assert results[0].category == "research" + + +class TestHermesResolver: + def test_lists_skills_in_two_level_layout(self, tmp_path: Path): + from openjarvis.skills.sources.hermes import HermesResolver + + # Build a fake Hermes layout: skills///SKILL.md + skills_root = tmp_path / "skills" + (skills_root / "apple" / "apple-notes").mkdir(parents=True) + (skills_root / "apple" / "apple-notes" / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: apple-notes + description: Manage Apple Notes + --- + Body + """) + ) + (skills_root / "github" / "github-pr").mkdir(parents=True) + (skills_root / "github" / "github-pr" / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: github-pr + description: Manage GitHub PRs + --- + Body + """) + ) + + resolver = HermesResolver(cache_root=tmp_path) + skills = resolver.list_skills() + names = sorted(s.name for s in skills) + assert names == ["apple-notes", "github-pr"] + # Categories come from the parent directory + for s in skills: + if s.name == "apple-notes": + assert s.category == "apple" + elif s.name == "github-pr": + assert s.category == "github" + + def test_skips_description_md(self, tmp_path: Path): + from openjarvis.skills.sources.hermes import HermesResolver + + skills_root = tmp_path / "skills" + category_dir = skills_root / "apple" + category_dir.mkdir(parents=True) + # DESCRIPTION.md at category level — should be skipped + (category_dir / "DESCRIPTION.md").write_text("# Apple skills") + (category_dir / "apple-notes").mkdir() + (category_dir / "apple-notes" / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: apple-notes + description: x + --- + """) + ) + + resolver = HermesResolver(cache_root=tmp_path) + skills = resolver.list_skills() + assert len(skills) == 1 + assert skills[0].name == "apple-notes" + + def test_filter_by_category(self, tmp_path: Path): + from openjarvis.skills.sources.hermes import HermesResolver + + skills_root = tmp_path / "skills" + for cat in ("apple", "github"): + d = skills_root / cat / f"{cat}-skill" + d.mkdir(parents=True) + (d / "SKILL.md").write_text( + textwrap.dedent(f"""\ + --- + name: {cat}-skill + description: x + --- + """) + ) + + resolver = HermesResolver(cache_root=tmp_path) + results = resolver.filter_by_category("apple") + assert len(results) == 1 + assert results[0].category == "apple" + + +class TestOpenClawResolver: + def test_lists_skills_in_owner_layout(self, tmp_path: Path): + from openjarvis.skills.sources.openclaw import OpenClawResolver + + skills_root = tmp_path / "skills" + # OpenClaw layout: skills///SKILL.md + d = skills_root / "alice" / "etherscan" + d.mkdir(parents=True) + (d / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: etherscan + description: Query EVM chain data + --- + Body + """) + ) + + resolver = OpenClawResolver(cache_root=tmp_path) + skills = resolver.list_skills() + assert len(skills) == 1 + assert skills[0].name == "etherscan" + assert skills[0].category == "alice" + + def test_reads_meta_json_sidecar(self, tmp_path: Path): + import json + + from openjarvis.skills.sources.openclaw import OpenClawResolver + + skills_root = tmp_path / "skills" + d = skills_root / "alice" / "etherscan" + d.mkdir(parents=True) + (d / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: etherscan + description: x + --- + """) + ) + (d / "_meta.json").write_text( + json.dumps( + { + "owner": "alice", + "slug": "etherscan", + "latest": {"version": "1.0.0", "commit": "abc123"}, + } + ) + ) + + resolver = OpenClawResolver(cache_root=tmp_path) + skills = resolver.list_skills() + assert len(skills) == 1 + assert skills[0].sidecar_data["owner"] == "alice" + assert skills[0].sidecar_data["latest"]["version"] == "1.0.0" + + +class TestGitHubResolver: + def test_recursive_walk_finds_skill_md(self, tmp_path: Path): + from openjarvis.skills.sources.github import GitHubResolver + + # Arbitrary nested layout + d = tmp_path / "any" / "depth" / "my-skill" + d.mkdir(parents=True) + (d / "SKILL.md").write_text( + textwrap.dedent("""\ + --- + name: my-skill + description: x + --- + Body + """) + ) + + resolver = GitHubResolver(cache_root=tmp_path, repo_url="https://example.com/x") + skills = resolver.list_skills() + assert len(skills) == 1 + assert skills[0].name == "my-skill" + + def test_finds_lowercase_skill_md(self, tmp_path: Path): + from openjarvis.skills.sources.github import GitHubResolver + + d = tmp_path / "my-skill" + d.mkdir() + (d / "skill.md").write_text( + textwrap.dedent("""\ + --- + name: my-skill + description: x + --- + """) + ) + + resolver = GitHubResolver(cache_root=tmp_path, repo_url="https://example.com/x") + skills = resolver.list_skills() + assert len(skills) == 1 diff --git a/tests/skills/test_tool_adapter_v2.py b/tests/skills/test_tool_adapter_v2.py new file mode 100644 index 00000000..0cfc7750 --- /dev/null +++ b/tests/skills/test_tool_adapter_v2.py @@ -0,0 +1,202 @@ +"""Tests for SkillTool v2 — parameter extraction and markdown support.""" + +from __future__ import annotations + +from openjarvis.core.types import ToolResult +from openjarvis.skills.executor import SkillExecutor +from openjarvis.skills.tool_adapter import SkillTool +from openjarvis.skills.types import SkillManifest, SkillStep +from openjarvis.tools._stubs import BaseTool, ToolExecutor, ToolSpec + + +class EchoTool(BaseTool): + tool_id = "echo" + + @property + def spec(self): + return ToolSpec(name="echo", description="Echo input") + + def execute(self, **params): + return ToolResult( + tool_name="echo", content=params.get("text", ""), success=True + ) + + +def _make_executor(*extra_tools): + tools = [EchoTool(), *extra_tools] + tool_executor = ToolExecutor(tools) + return SkillExecutor(tool_executor) + + +class TestParameterExtraction: + def test_pipeline_params_extracted(self): + """Placeholders in arguments_template become input params.""" + manifest = SkillManifest( + name="pipe_skill", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "{query}"}', + output_key="echoed", + ) + ], + ) + skill_tool = SkillTool(manifest, _make_executor()) + props = skill_tool.spec.parameters.get("properties", {}) + assert "query" in props + + def test_output_keys_not_exposed_as_params(self): + """output_key values from prior steps are not exposed as input params.""" + manifest = SkillManifest( + name="chain_skill", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "hello"}', + output_key="echoed", + ), + SkillStep( + tool_name="echo", + arguments_template='{"text": "{echoed}"}', + output_key="result", + ), + ], + ) + skill_tool = SkillTool(manifest, _make_executor()) + props = skill_tool.spec.parameters.get("properties", {}) + # "echoed" is produced by step 1, so it should NOT be an input param + assert "echoed" not in props + + def test_instruction_only_skill_gets_task_param(self): + """Skills with no steps (markdown-only) expose an optional 'task' param.""" + manifest = SkillManifest( + name="md_only", + markdown_content="When asked to research, follow these steps...", + steps=[], + ) + skill_tool = SkillTool(manifest, _make_executor()) + props = skill_tool.spec.parameters.get("properties", {}) + assert "task" in props + + def test_no_extra_params_for_pipeline_skill(self): + """Pipeline skills don't get a spurious 'task' param.""" + manifest = SkillManifest( + name="echo_skill", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "{msg}"}', + output_key="out", + ) + ], + ) + skill_tool = SkillTool(manifest, _make_executor()) + props = skill_tool.spec.parameters.get("properties", {}) + assert "msg" in props + assert "task" not in props + + def test_multiple_placeholders_across_steps(self): + """Multiple distinct placeholders across steps are all exposed.""" + manifest = SkillManifest( + name="multi_param", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "{first}"}', + output_key="out1", + ), + SkillStep( + tool_name="echo", + arguments_template='{"text": "{second}"}', + output_key="out2", + ), + ], + ) + skill_tool = SkillTool(manifest, _make_executor()) + props = skill_tool.spec.parameters.get("properties", {}) + assert "first" in props + assert "second" in props + + +class TestMarkdownReturn: + def test_instruction_only_returns_markdown(self): + """Instruction-only skill (no steps) returns markdown_content as result.""" + manifest = SkillManifest( + name="md_skill", + markdown_content="## Instructions\nDo the thing.", + steps=[], + ) + skill_tool = SkillTool(manifest, _make_executor()) + result = skill_tool.execute(task="some task") + assert result.success + assert "## Instructions" in result.content + + def test_instruction_only_no_task_param_still_works(self): + """Instruction-only skill works even without task param.""" + manifest = SkillManifest( + name="md_skill2", + markdown_content="Follow these instructions.", + steps=[], + ) + skill_tool = SkillTool(manifest, _make_executor()) + result = skill_tool.execute() + assert result.success + assert "Follow these instructions." in result.content + + def test_hybrid_returns_pipeline_result_and_markdown(self): + """Hybrid skill (steps + markdown) returns both pipeline output and markdown.""" + manifest = SkillManifest( + name="hybrid_skill", + markdown_content="## Guidance\nExtra context.", + steps=[ + SkillStep( + tool_name="echo", + arguments_template='{"text": "{input}"}', + output_key="result", + ) + ], + ) + skill_tool = SkillTool(manifest, _make_executor()) + result = skill_tool.execute(input="hello world") + assert result.success + # Pipeline output is present + assert "hello world" in result.content + # Markdown is appended + assert "## Guidance" in result.content + + def test_pipeline_failure_is_propagated(self): + """If the pipeline fails, the ToolResult is not successful.""" + manifest = SkillManifest( + name="broken_skill", + steps=[ + SkillStep(tool_name="nonexistent_tool", output_key="x"), + ], + ) + skill_tool = SkillTool(manifest, _make_executor()) + result = skill_tool.execute() + assert not result.success + + def test_tool_id_uses_skill_prefix(self): + """tool_id is prefixed with 'skill_'.""" + manifest = SkillManifest(name="my_skill", steps=[]) + skill_tool = SkillTool(manifest, _make_executor()) + assert skill_tool.tool_id == "skill_my_skill" + + def test_spec_name_uses_skill_prefix(self): + """spec.name is prefixed with 'skill_'.""" + manifest = SkillManifest(name="another_skill", steps=[]) + skill_tool = SkillTool(manifest, _make_executor()) + assert skill_tool.spec.name == "skill_another_skill" + + def test_spec_category_is_skill(self): + """spec.category is always 'skill'.""" + manifest = SkillManifest(name="cat_skill", steps=[]) + skill_tool = SkillTool(manifest, _make_executor()) + assert skill_tool.spec.category == "skill" + + def test_skill_manager_kwarg_accepted(self): + """Constructor accepts optional skill_manager keyword argument.""" + manifest = SkillManifest(name="mgr_skill", steps=[]) + # Should not raise + skill_tool = SkillTool(manifest, _make_executor(), skill_manager=None) + assert skill_tool.tool_id == "skill_mgr_skill" diff --git a/tests/skills/test_tool_translator.py b/tests/skills/test_tool_translator.py new file mode 100644 index 00000000..965564cc --- /dev/null +++ b/tests/skills/test_tool_translator.py @@ -0,0 +1,81 @@ +"""Tests for ToolTranslator — translate external tool names to OpenJarvis.""" + +from __future__ import annotations + +from openjarvis.skills.tool_translator import TOOL_TRANSLATION, ToolTranslator + + +class TestTranslationTable: + def test_bash_translates_to_shell_exec(self): + assert TOOL_TRANSLATION["Bash"] == "shell_exec" + + def test_read_translates_to_file_read(self): + assert TOOL_TRANSLATION["Read"] == "file_read" + + def test_write_translates_to_file_write(self): + assert TOOL_TRANSLATION["Write"] == "file_write" + + def test_websearch_translates(self): + assert TOOL_TRANSLATION["WebSearch"] == "web_search" + + +class TestMarkdownTranslation: + def test_translate_inline_tool_reference(self): + translator = ToolTranslator() + body = "Use the Bash tool to run commands." + new_body, untranslated = translator.translate_markdown(body) + assert "shell_exec" in new_body + assert "Bash" not in new_body + assert untranslated == [] + + def test_translate_multiple_tools(self): + translator = ToolTranslator() + body = "First use Read to load the file, then Write to save it." + new_body, _ = translator.translate_markdown(body) + assert "file_read" in new_body + assert "file_write" in new_body + + def test_unknown_tool_collected(self): + translator = ToolTranslator() + body = "Use the QuantumComputer tool." + new_body, untranslated = translator.translate_markdown(body) + assert "QuantumComputer" in untranslated + assert "QuantumComputer" in new_body # left in place + + def test_word_boundary_respected(self): + """'Read' should not match 'Reader' or 'Reading'.""" + translator = ToolTranslator() + body = "The Reader processes Reading material." + new_body, _ = translator.translate_markdown(body) + assert "Reader" in new_body + assert "Reading" in new_body + assert "file_read" not in new_body + + def test_empty_body(self): + translator = ToolTranslator() + new_body, untranslated = translator.translate_markdown("") + assert new_body == "" + assert untranslated == [] + + +class TestAllowedToolsTranslation: + def test_translate_space_delimited(self): + translator = ToolTranslator() + new_field, untranslated = translator.translate_allowed_tools("Bash Read Write") + assert "shell_exec" in new_field + assert "file_read" in new_field + assert "file_write" in new_field + assert untranslated == [] + + def test_translate_with_arguments(self): + translator = ToolTranslator() + # Spec example: Bash(git:*) Bash(jq:*) Read + new_field, _ = translator.translate_allowed_tools("Bash(git:*) Bash(jq:*) Read") + assert "shell_exec(git:*)" in new_field + assert "shell_exec(jq:*)" in new_field + assert "file_read" in new_field + + def test_unknown_in_allowed_tools_collected(self): + translator = ToolTranslator() + new_field, untranslated = translator.translate_allowed_tools("Bash UnknownTool") + assert "UnknownTool" in untranslated