OpenJarvis can run vision-capable local models (gemma3, qwen2.5-vl), but the CLI had no way to send them a picture -- the Ollama engine only serialized text. This adds end-to-end image input. What's new - `jarvis ask -i/--image <file>` attaches one or more images to the query. - `jarvis ask -S/--screen` captures the primary monitor (dependency-free on Windows via .NET; mss/Pillow fallback elsewhere). - Vision auto-routes to direct-to-engine mode; with an explicit --agent it warns rather than silently dropping the image. - Privacy guard: warns before sending an image to a non-local engine, keeping OpenJarvis local-first by default. - Context-window default raised 8k -> 16k (JARVIS_NUM_CTX) so an image plus a conversation fit. Implementation - Message.images carries base64 data; messages_to_dicts() forwards it to Ollama's /api/chat "images" field. Text-only messages are unchanged. - GuardrailsEngine preserves images when it rewrites a flagged message. Tests (tests/test_vision.py, 6/6 pass, ruff-clean) - payload forwarding, text path untouched, num_ctx override, guardrail image preservation. Verified on AMD RX 9070 XT (Ollama/Vulkan, 100% GPU) with gemma3:4b: solid-color image, file image, and live screen capture all described. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Jon Saad-Falcon <jonsaadfalcon@gmail.com>
19 KiB
Changelog
All notable changes to OpenJarvis are documented in this file.
The format is based on Keep a Changelog.
[Unreleased]
Added
Vision input for jarvis ask — attach images to a query with
-i/--image (repeatable) or capture the current screen with
-S/--screen, for vision-capable models such as gemma3:4b. Images flow
through Message.images into Ollama's /api/chat images field; text-only
requests are unaffected. A privacy guard warns before any image is sent to a
non-local engine, and the security guardrail now preserves images when it
sanitizes a flagged prompt. Screen capture uses the built-in Windows .NET
stack with mss/Pillow fallbacks on other platforms. Adds the
JARVIS_NUM_CTX environment variable to tune the Ollama context window
(default 16384).
[1.0.2] - 2026-05-24
A patch release that fixes a packaging bug which broke the v1.0.1
wheel on PyPI, silences a noisy startup warning, restores a working
install path while openjarvis.ai is down, improves desktop
first-boot diagnostics on Windows, and ships the RAM-detection fix
for Windows that missed the v1.0.1 cutoff.
Fixed
openjarvis/traces/ missing from the v1.0.1 PyPI wheel (#372).
The .gitignore carried an unanchored traces/ pattern, which
hatchling honored at wheel-build time and matched the runtime module
src/openjarvis/traces/ — silently dropping the whole package. Every
fresh pip install openjarvis==1.0.1 then failed at import with
ModuleNotFoundError: No module named 'openjarvis.traces' on the
first jarvis ask, learning, or server call. Anchored the pattern to
/traces/. Verified: a clean uv build now produces a wheel
containing all four traces/ files.
pynvml deprecation FutureWarning on every command (#389).
Switched the dependency from the legacy pynvml package to NVIDIA's
official nvidia-ml-py (same pynvml module name, no warning shim),
and added defensive warnings.filterwarnings at every import pynvml
site to suppress the warning even when pynvml is pulled in
transitively.
Windows RAM detection returning 0.0 GB (#373). The Windows
branch of _total_ram_gb() (via GlobalMemoryStatusEx) landed after
the v1.0.1 cutoff, so v1.0.1 users still saw 0.0 GB from jarvis init. Now shipping in the wheel. A new windows-latest CI job runs
the real GlobalMemoryStatusEx path on every PR as a regression
guard.
Desktop first-boot hung on "did not become healthy in time"
(#331). The Tauri boot path ran uv sync with stderr discarded and
the exit code ignored, so a failed dependency install surfaced only
as a generic 600-second health-check timeout. Now captures stderr,
checks the exit status, and surfaces the actual uv sync error
(with the diagnostic tail) before the long wait. The error-formatting
logic is covered by unit tests.
Changed
Install URL moved to GitHub Pages (#337, #352). The documented
openjarvis.ai/install.sh URL was failing with sslv3 alert handshake failure (the domain is community-operated and had a broken
TLS config). The canonical installer is now served from the
project-controlled GitHub Pages site at
https://open-jarvis.github.io/OpenJarvis/install.sh, generated from
the same scripts/install/install.sh at docs-build time. The README
also documents the WSL2 path for Windows and the uv prerequisite
for the desktop binary, and the installer bails early with a clear
message when run under Git Bash / MSYS2 / Cygwin.
[1.0.1] - 2026-05-17
A patch release that closes the auto-update gap so the analytics module added in #351 actually reaches users on the desktop, adds runtime opt-out for that analytics, fixes the misleading upgrade hint the CLI was printing, and lands the ACE optimizer alongside DSPy and GEPA.
Added
ACE agent optimizer (learning/agents/ace_optimizer.py). Adds
ACE as a third agent-learning
policy alongside DSPy and GEPA. Where DSPy bootstraps few-shot
examples and GEPA evolves prompt populations, ACE evolves a textual
playbook of strategies the agent reads at inference time, updated
by a Generator / Reflector / Curator triad. Pick via
[learning.agent] policy = "ace". Setup is manual (ACE isn't on
PyPI and isn't a properly-packaged Python project as of v1.0.1) —
see docs/learning/ace.md for the install path and trace-adapter
behavior.
jarvis self-update subcommand. Detects how OpenJarvis was
installed (pip, uv tool, editable git checkout) by inspecting
openjarvis.__file__, then runs the right upgrade command. Supports
--check (print the command without running) and -y (skip the
confirmation prompt). The post-command "new version available" hint
now points users at this command instead of guessing at the right
flow.
Desktop auto-update endpoint wired to the rolling
desktop-latest GitHub release. The Tauri updater plugin was
configured on the build side (createUpdaterArtifacts: true,
includeUpdaterJson: true, signing key in TAURI_SIGNING_PRIVATE_KEY)
but inert on the runtime side (active: false, endpoints: []). The
installed desktop app would never check. Both are now fixed; the app
polls releases/download/desktop-latest/latest.json every 30 minutes
and signature-verifies downloads against the minisign pubkey baked
into the app. Full flow, key-rotation runbook, and dev escape hatch
(OPENJARVIS_NO_UPDATER=1) documented in docs/desktop-auto-update.md.
Analytics env-var opt-out (DO_NOT_TRACK, OPENJARVIS_NO_ANALYTICS).
Tanvir's analytics module (#351) only respected the
[analytics] enabled config-file setting. Both env vars are now
honored in is_analytics_enabled() and in the install.sh beacon
script. Any truthy value (1, true, yes, on) disables for
that process; env opt-out takes precedence over the config file.
Documented under a new "Opting out" section in docs/telemetry.md.
Changed
Version-check trigger widened. The "new version available" hint
in _version_check.py used to fire only on {ask, chat, serve} and
hardcoded the wrong upgrade command (git pull && uv sync — only
correct for editable installs). Now fires on every interactive
command (doctor, init, quickstart, model, agents, skill,
memory, bench, telemetry, config, eval, optimize, plus
the original three) and uses install-detection to print the right
upgrade command. Honors JARVIS_NO_UPDATE_CHECK=1 and CI=true to
stay silent in automation.
Desktop app version bumped 0.1.0 → 1.0.1 across
tauri.conf.json, frontend/package.json, and
frontend/src-tauri/Cargo.toml so the Python and desktop release
streams are aligned and the auto-updater has a real version to
compare against.
Migration from 1.0.0
- Importing
is_analytics_enabled? Same signature; behavior now short-circuits on env opt-out before checking the config. Callers that want the raw "is the config flag set" semantic should readcfg.enableddirectly. - Editable-git users running
jarvis self-updateget the detectedgit pull && uv synccommand pointed at their actual checkout, not~/OpenJarvis. If you'd come to rely on the hardcoded path, update your muscle memory.
[1.0.0] - 2026-05-15
The five-primitive architecture (Intelligence, Engine, Agents, Tools & Memory, Learning) is now stable, with efficiency and on-device learning as first-class capabilities alongside accuracy. Companion blog post: From Minions to OpenJarvis: A Retrospective on Two Years in Local AI.
Highlights
Five composable primitives. Intelligence, Engine, Agents, Tools & Memory,
and Learning each sit behind a single typed interface — any slot is
substitutable without touching the rest. The composition layer is
JarvisSystem in src/openjarvis/system.py, driven by a TOML config.
Built-in agents across three execution modes. Eight agents spanning a single-turn chat baseline, a deep-research agent with inline citations, a CodeAct-style coder, and a continuous monitor with memory compression for long-horizon workflows. Execution modes cover on-demand, scheduled, and continuous.
Starter presets. Eight preset configs installable via
jarvis init --preset <name> bundle an agent with a hardware-appropriate
engine, connectors, and tools. Variants cover Apple Silicon, Linux GPU
servers, and CPU-only laptops, plus a quickstart for LLM-guided spec search.
Inference engines. Four first-class local engines (Ollama, vLLM, SGLang,
llama.cpp) and five cloud providers (OpenAI, Anthropic, Google Gemini,
OpenRouter, MiniMax) sit behind a single Engine interface. Discovery
in engine/_discovery.py picks a sensible default per host.
Added — hybrid local-cloud capabilities
Per-query routing via a query-complexity analyzer
(src/openjarvis/learning/routing/complexity.py). Produces a 0.0–1.0
complexity score with code/math/reasoning signals and a suggested token
budget, populating RoutingContext so easy queries stay local and only
queries that need frontier capability escalate.
LLM-guided spec search (src/openjarvis/learning/spec_search/).
SpecSearchOrchestrator wires diagnose → plan → execute → gate into a
single learning session: a frontier model reads traces, proposes
coordinated edits across all five primitives, and a held-out benchmark
gate (gate/benchmark_gate.py, gate/regression.py, gate/cold_start.py)
accepts only non-regressing edits. Ships with the spec-search-quickstart
preset and a runnable tutorial at examples/openjarvis/spec_search_quickstart.py.
Six hybrid coordination paradigms in src/openjarvis/agents/hybrid/.
Each paradigm pairs a local student with a frontier cloud teacher under
a different orchestration shape, as LocalCloudAgent subclasses:
minions— reactive single-local + single-cloud loopconductor— static DAG planneradvisors— executor ↔ advisor looparchon— generate → rank → fuseskillorchestra— per-query router across local skillstoolorchestra— RL'd local model with a tool pool
A runner CLI (python -m openjarvis.agents.hybrid.runner --cell <name>)
and a 35-cell experiment registry (one TOML per method × benchmark ×
model triple) let researchers run, score, and compare these on equal
footing. Includes a Modal-backed SWE-bench-Verified harness scorer
(evals/scorers/swebench_harness.py).
Added — efficiency as a first-class constraint
Hardware-agnostic energy telemetry at 50ms resolution across NVIDIA
(telemetry/energy_nvidia.py), AMD (telemetry/energy_amd.py), Apple
Silicon (telemetry/energy_apple.py), and Intel RAPL
(telemetry/energy_rapl.py). Energy, dollar cost, FLOPs, and latency
are treated as evaluation targets alongside accuracy.
Instrumentation for FLOPs, batch, steady-state, ITL, phase energy, and
vLLM-specific metrics. Joined per-query by the aggregator
(telemetry/aggregator.py) so traces carry accuracy + efficiency together.
Added — local learning loop
Closed-loop optimization across the stack — model weights via SFT
(learning/intelligence/sft_trainer.py) and GRPO
(learning/intelligence/grpo_trainer.py plus an orchestrator-specific
variant under learning/intelligence/orchestrator/), prompts via DSPy
(learning/agents/dspy_optimizer.py), agent logic via GEPA
(learning/agents/gepa_optimizer.py), and engine + stack configuration
via LLM-guided spec search. LearningOrchestrator coordinates triggers
and applies optimizer overlays at discovery time so improvements compound
across primitives.
Added — cross-framework evaluation
External agentic-framework evaluation via subprocess. The
evals/backends/external/ subpackage wraps Hermes Agent and OpenClaw as
one-shot subprocess backends behind the existing InferenceBackend ABC.
The evals/comparison/ toolkit provides path + commit-pin enforcement
(third_party.py), config templating (make_configs.py), and LaTeX
table generation (table_gen.py).
Ships with a new optional extra framework-comparison (depends on
polars), a live_external pytest marker for integration tests
requiring real foreign-framework installations, and a ToolOrchestra
evaluation dataset (evals/datasets/toolorchestra.py) alongside the
existing 30+ benchmark suite.
Added — Skills System (Plans 1, 2A, 2B)
-
Skills core — every skill is a tool. Skills appear in a system prompt catalog, agents invoke them on demand, content (pipeline results, markdown instructions, or both) gets injected into context.
SkillManifest+SkillSteptypes with tags, depends, invocation flags, markdown contentSkillManager— discovery, precedence resolution, catalog XML generation, tool wrappingSkillTool(BaseTool)— auto-extracts parameters from step argument templatesSkillExecutor— sequential pipeline execution with sub-skill delegation- Dependency graph with cycle detection, max depth enforcement, capability unions
- Security: four trust tiers (bundled/indexed/unreviewed/workspace), capability-gated enforcement
- Skill index module for git-backed registry search
-
agentskills.io spec adoption — canonical
SKILL.mdformat with YAML frontmatter following the agentskills.io open standard.SkillParserwith strict spec validation + tolerant field mapping viaFIELD_MAPPINGtableToolTranslatorfor external tool name translation (Bash -> shell_exec, Read -> file_read, etc.)- Source resolvers:
HermesResolver,OpenClawResolver,GitHubResolver SkillImporterwith provenance tracking (.sourcemetadata files), optional script import- Sourced subdirectory layout (
~/.openjarvis/skills/<source>/<name>/)
-
Skills learning loop — trace tagging, pattern discovery, DSPy/GEPA optimization.
- Trace metadata tagging:
skill,skill_source,skill_kindflow through ToolExecutor -> TraceCollector -> TraceStep SkillDiscoverywired intoSkillManager.discover_from_traces()with kebab name normalizationSkillOptimizer— per-skill DSPy/GEPA wrapper that buckets traces and writes sidecar overlaysSkillOverlay— sidecar storage at~/.openjarvis/learning/skills/<name>/optimized.tomlSkillManager._load_overlays()applies optimized descriptions + few-shot examples at discovery timeLearningOrchestrator._maybe_optimize_skills()— opt-in auto-trigger
- Trace metadata tagging:
-
Skills benchmark harness — 4-condition PinchBench evaluation.
- I3 fix:
skill_few_shot_exampleswired through SystemBuilder ->_run_agent->ToolUsingAgent->native_react.REACT_SYSTEM_PROMPT SkillBenchmarkRunner— 4-condition x N-seed x M-task sweep with markdown reportJarvisAgentBackendacceptsskills_enabledandoverlay_dirkwargs- Conditions:
no_skills,skills_on,skills_optimized_dspy,skills_optimized_gepa
- I3 fix:
-
CLI commands:
jarvis skill list/info/run/install/sync/sources/update/remove/searchjarvis skill discover— mine traces for recurring tool patternsjarvis skill show-overlay— inspect optimization outputjarvis optimize skills— run DSPy/GEPA per-skill optimizationjarvis bench skills— run the PinchBench skills benchmark
-
Agent prompt improvement:
native_react.REACT_SYSTEM_PROMPTnow includes "Using Skills" guidance that teaches agents to distinguish executable vs. instructional skill responses{skill_examples}placeholder for optimized few-shot example injection
-
Configuration:
[skills]section:enabled,skills_dir,active,auto_discover,auto_sync,max_depth,sandbox_dangerous[[skills.sources]]section:source,url,filter,auto_update[learning.skills]section:auto_optimize,optimizer,min_traces_per_skill,optimization_interval_seconds,overlay_dirSkillSourceConfigandSkillsLearningConfigdataclasses
-
Documentation:
docs/user-guide/skills.md— comprehensive user guidedocs/architecture/skills.md— technical deep-divedocs/tutorials/skills-workflow.md— end-to-end tutorialdocs/getting-started/configuration.md— expanded with skills config sectionsCLAUDE.md— updated architecture section
Examples & Tutorials
examples/openjarvis/spec_search_quickstart.py— runnable end-to-end LLM-guided spec search session.docs/user-guide/llm-guided-spec-search.md— paper-aligned user guide.docs/architecture/learning.md— Learning primitive deep-dive covering routing, spec search, optimizers, and the orchestrator.docs/tutorials/— code-companion, deep-research, messaging-hub, scheduled-ops, and skills-workflow walkthroughs.src/openjarvis/agents/hybrid/registry/*.toml— 35-cell registry of paradigm × benchmark × model experiments.
Migration from 0.x
learning/distillation/is nowlearning/spec_search/. The subsystem was renamed to match the LLM-guided spec search semantics documented in the companion paper. Update any imports (from openjarvis.learning.distillation.*→from openjarvis.learning.spec_search.*). Thejarvis distillationCLI command is removed; usespec_search-prefixed config keys instead._third_party.tomlno longer ships default paths. SetHERMES_AGENT_PATHandOPENCLAW_PATHenv vars to point at your local checkouts before running the framework-comparison harness; missing or empty paths now raiseThirdPartyNotFoundErrorwith an actionable hint.- Engine
generate_fullreturn shape extended.JarvisAgentBackend.generate_fullandJarvisDirectBackend.generate_fullnow return the spec §6.2 extended fields (energy_joules,peak_power_w,tool_calls,turn_count,framework,framework_commit,error). Existing callers that didn't read these fields are unaffected; new callers can rely on cross-framework parity.
Fixed
- Trace metadata flow —
ToolResult.metadatanow propagates throughTOOL_CALL_ENDevent toTraceStep.metadata(was silently dropped at the event-bus boundary). - TaintSet JSON serialization —
ToolExecutor._json_safe_metadata()filters non-JSON-serializable values (likeTaintSet) from event payloads before they reachTraceStore. - Non-dict YAML frontmatter — source resolvers handle
yaml.safe_load()returning a string instead of a dict (discovered on real OpenClaw imports). - OpenClaw category/name queries —
jarvis skill install openclaw:owner/slugnow correctly splits into category + name match. - SkillDiscovery trace compatibility —
_extract_tool_sequencereads fromstep.input["tool"](the actualTraceStepformat), not the nonexistentstep.tool_nameattribute. - LearningOrchestrator skill trigger —
_maybe_optimize_skillsruns BEFORE the SFT-data short-circuit (skills are tagged via trace metadata, not mined as SFT pairs). - PinchBenchScorer constructor —
SkillBenchmarkRunnerconstructsPinchBenchScorer(judge_backend, model)instead of no-args. - EvalRunner results access — reads per-task data from
eval_runner.resultsproperty, not nonexistentsummary.results.