cloud.py:
- Skip server_tool_use blocks from tool_calls (they're already in
content_blocks; agents would otherwise try to execute Anthropic
server-side tools like web_search as if they were local).
- Drop the synthetic tool_use_id fallback (f"{btype}_{len(...)}").
Anthropic always returns ids for tool_use blocks; the fallback
only masked missing ids with fakes that fail Anthropic's
tool_use_id matching on the next assistant turn. Now we log a
warning and skip the block.
- Streaming/non-streaming parity: emit content_blocks and
tool_results at end-of-stream in _stream_full_anthropic so
streaming callers see the same rich blocks as generate().
StreamChunk gets two new optional fields.
swebench_harness.py:
- Log a warning when set_cpu_quota is missing instead of silently
no-op'ing — previously a swebench API change would have produced
a silent 0-score sweep with no diagnostic.
- Delete the prior report JSON before invoking the subprocess so a
stale file from a crashed run can't be re-read as fresh.
archon.py:
- Token tally is now thread-local (threading.local). The runner
reuses one ArchonAgent across a ThreadPoolExecutor; the previous
module-global dict let concurrent tasks reset and read each
other's counters, corrupting per-task cost_usd / tokens_*.
toolorchestra.py:
- Wrap the orchestration loop in try/finally so shared_workdir is
always cleaned up. Previously rmtree only ran on the success
path; any exception in the turn loop, worker call, or diff
extraction leaked the cloned repo (hundreds of MB at n=500).
Matches the pattern conductor.py and mini_swe_agent.py use.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Why OpenJarvis?
Personal AI agents are exploding in popularity, but nearly all of them still route intelligence through cloud APIs. Your "personal" AI continues to depend on someone else's server. At the same time, our Intelligence Per Watt research showed that local language models already handle 88.7% of single-turn chat and reasoning queries, with intelligence efficiency improving 5.3× from 2023 to 2025. The models and hardware are increasingly ready. What has been missing is the software stack to make local-first personal AI practical.
OpenJarvis is that stack. It is an opinionated framework for local-first personal AI, built around three core ideas: shared primitives for building on-device agents; evaluations that treat energy, FLOPs, latency, and dollar cost as first-class constraints alongside accuracy; and a learning loop that improves models using local trace data. The goal is simple: make it possible to build personal AI agents that run locally by default, calling the cloud only when truly necessary. OpenJarvis aims to be both a research platform and a production foundation for local AI, in the spirit of PyTorch.
Installation
curl -fsSL https://openjarvis.ai/install.sh | bash
That's it. The installer handles everything: uv, the Python venv, Ollama, and pulling a small starter model. About 3 minutes on a typical broadband connection. Then:
jarvis
The Rust extension and bigger models continue downloading in the background while you chat. Run jarvis doctor to see status.
Platforms: macOS (Intel + Apple Silicon), Linux, WSL2 on Windows.
Manual install / contributors: see docs/getting-started/install.md.
Quick Start
curl -fsSL https://openjarvis.ai/install.sh | bash
jarvis
jarvis init --preset <name> switches to a starter config. Available presets: morning-digest-mac, morning-digest-linux, morning-digest-minimal, deep-research, code-assistant, scheduled-monitor, chat-simple.
Starter Configs
Install any preset with one command:
uv run jarvis init --preset morning-digest-mac # or any preset below
Prefix every
jarvis ...invocation withuv run, or activate the venv first (source .venv/bin/activate) so plainjarvis ...works for the rest of your shell session.
| Preset | Use Case | What it does |
|---|---|---|
morning-digest-mac |
Daily Briefing (Mac) | Spoken briefing from email, calendar, health, news with Jarvis voice |
morning-digest-linux |
Daily Briefing (Linux) | Same, with vLLM support for GPU servers |
morning-digest-minimal |
Daily Briefing (minimal) | Just Gmail + Calendar, runs on any machine |
deep-research |
Research Assistant | Multi-hop research across indexed docs with citations |
code-assistant |
Code Companion | Agent with code execution, file I/O, and shell access |
scheduled-monitor |
Persistent Monitor | Stateful agent that runs on a schedule with memory |
chat-simple |
Simple Chat | Lightweight conversation, no tools needed |
# Example: Morning Digest on Mac
uv run jarvis init --preset morning-digest-mac
uv run jarvis connect gdrive # one OAuth flow covers Gmail, Calendar, Tasks
uv run jarvis digest --fresh # generate and play your first briefing
# Example: Deep Research
uv run jarvis init --preset deep-research
uv run jarvis memory index ./docs/ # requires the Rust extension — see Setup above
uv run jarvis ask "Summarize all emails about Project X"
Skills
Skills teach agents how to better use tools and improve their reasoning. Every skill is a tool — agents discover them from a catalog and invoke them on demand.
# Install skills from public sources
jarvis skill install hermes:arxiv
jarvis skill sync hermes --category research
# Use skills with any agent
jarvis ask "Use the code-explainer skill to explain this Python code: for i in range(5): print(i*2)"
# Optimize skills from your trace history
jarvis optimize skills --policy dspy
# Benchmark the impact
jarvis bench skills --max-samples 5 --seeds 42
Import from Hermes Agent (~150 skills), OpenClaw (~13,700 community skills), or any GitHub repo. Skills follow the agentskills.io open standard.
See the Skills User Guide and Skills Tutorial for details.
Built-in Agents
| Agent | Type | What it does |
|---|---|---|
morning_digest |
Scheduled | Daily briefing from email, calendar, health, news — with TTS audio |
deep_research |
On-demand | Multi-hop research with citations across web and local docs |
monitor_operative |
Continuous | Long-horizon monitoring with memory, compression, and retrieval |
orchestrator |
On-demand | Multi-turn reasoning with automatic tool selection |
native_react |
On-demand | ReAct (Thought-Action-Observation) loop agent |
operative |
Continuous | Persistent autonomous agent with state management |
native_openhands |
On-demand | CodeAct — generates and executes Python code |
simple |
On-demand | Single-turn chat, no tools |
See the User Guide and Tutorials for detailed setup instructions.
Full documentation — including Docker deployment, cloud engines, development setup, and tutorials — at open-jarvis.github.io/OpenJarvis.
Contributing
We welcome contributions! See the Contributing Guide for incentives, contribution types, and the PR process.
Quick start for contributors:
git clone https://github.com/open-jarvis/OpenJarvis.git
cd OpenJarvis
uv sync --extra dev
uv run pre-commit install
uv run pytest tests/ -v
Browse the Roadmap for areas where help is needed. Comment "take" on any issue to get auto-assigned.
About
OpenJarvis is part of Intelligence Per Watt, a research initiative studying the efficiency of on-device AI systems. The project is developed at Hazy Research and the Scaling Intelligence Lab at Stanford SAIL.
Sponsors
Laude Institute • Stanford Marlowe • Google Cloud Platform • Lambda Labs • Ollama • IBM Research • Stanford HAI
Citation
@misc{saadfalcon2026openjarvis,
title={OpenJarvis: Personal AI, On Personal Devices},
author={Jon Saad-Falcon and Avanika Narayan and Herumb Shandilya and Hakki Orhun Akengin and Robby Manihani and Gabriel Bo and John Hennessy and Christopher R\'{e} and Azalia Mirhoseini},
year={2026},
howpublished={\url{https://scalingintelligence.stanford.edu/blogs/openjarvis/}},
}
