1. Agent responses in Interact tab now render as markdown (headers,
bold, lists, etc.) instead of raw text with ** and ##
2. Instruction section is its own box above Configuration in Overview
3. Edit button is inline next to "Instruction" heading
4. User messages stay as plain text (no markdown needed)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1. Template cards highlight with accent border on hover
2. Text aligned top-left, emoji inline with title
3. Schedule is now an editable dropdown (not hardcoded display)
4. Overview tab shows current instruction with Edit button
5. Instruction is editable and saved to agent config
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Resolve tool names from config["tools"] into actual tool instances
using ToolRegistry, inject runtime deps, and pass them to the agent
constructor instead of hardcoded tools=[].
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace the recompute approach from #146 with per-token outlier
detection. Entries whose energy, FLOPs, or dollar savings exceed
generous per-token thresholds (~1000x legitimate values) are hidden
from the leaderboard entirely. All displayed values come directly
from the database — no rewriting of submitted data.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Sub-project A spec covering: 2-step wizard, smart template defaults,
rich system prompt templates, tool wiring fix, recommended model
endpoint, and Advanced settings collapse.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Energy and FLOPs are now derived from total_tokens on the leaderboard
display using Claude Opus 4.6 constants, rather than trusting submitted
DB values. This prevents gaming (e.g. TotallyNoire submitting 14B Wh
from 15M tokens). Dollar savings are clamped at the theoretical max
($25/1M tokens) on both the leaderboard and frontend submission side.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Comprehensive step-by-step guide covering Homebrew, uv, Rust, llama.cpp,
model download, Python 3.12 pin (PyO3 compat), and common pitfalls.
Cherry-picked from PR #131 by @gridworks — cleaned up to include only
the docs content (removed duplicate files, binary artifacts, and
unrelated lockfile changes from the original PR).
Co-Authored-By: gridworks <5502067+gridworks@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Companion to #143 which fixed the frontend to use Claude Opus 4.6 as
the sole baseline. This migration recomputes existing Supabase rows
using the exact closed-form: new = T/3.8M + 10*old/19, derived from
the original triple-provider formula.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Remove cent sign from cost display, use $X.XXXX format
- Compact stat cards (horizontal icon+value layout)
- Tighter config grid spacing with bolder labels
- Reduce padding and gaps throughout overview tab
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Dollar savings were previously summed across all three cloud providers
(GPT-5.3 + Claude Opus 4.6 + Gemini 3.1 Pro), inflating the reported
number by ~3x. Now uses only Claude Opus 4.6 pricing as the baseline,
with an asterisk footnote on the leaderboard explaining this.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds `codex/` prefixed model support using the OpenAI Responses API —
the same protocol used by zeroclaw and other Codex-compatible tools.
Live-tested against gpt-5-mini-2025-08-07 with:
- Generate (non-streaming): confirmed working
- System prompt → instructions mapping: confirmed working
- SSE streaming: confirmed working (9 chunks)
- End-to-end via `jarvis ask`: confirmed working
Implementation:
- Default endpoint: api.openai.com/v1/responses (standard API key)
- Override via OPENAI_CODEX_BASE_URL for ChatGPT OAuth tokens
(e.g. chatgpt.com/backend-api/codex)
- Auth via OPENAI_CODEX_API_KEY env var
- Responses API format: input array, instructions field, output_text extraction
- Handles reasoning+message output blocks correctly
- SSE streaming parses response.output_text.delta events
Models: codex/gpt-4o, codex/gpt-4o-mini, codex/o3-mini,
codex/gpt-5-mini, codex/gpt-5-mini-2025-08-07
Usage:
export OPENAI_CODEX_API_KEY="your-api-key-or-oauth-token"
jarvis ask "Hello" --model codex/gpt-5-mini-2025-08-07
Closes#134
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Overview tab:
- Show Intelligence (model name) with click-to-change dropdown
- Split "Total Tokens" into "Input Tokens" and "Output Tokens"
- Model can be switched for existing agents via dropdown
Backend:
- Add input_tokens/output_tokens columns to managed_agents
- Track prompt_tokens and completion_tokens separately in executor
- Disable Ollama thinking by default (think:false) to prevent
empty responses from token exhaustion
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Root cause of empty responses: Qwen3.5's extended thinking mode
consumes all tokens (4096) on hidden <think> tags, leaving zero
visible content. The /no_think text tag was unreliable.
Fix: pass think=false in the Ollama API payload, which properly
disables thinking at the API level. Drops token usage from ~4096
to ~5-200 per response and eliminates empty content.
Also:
- Fix token tracking to read total_tokens from metadata (was looking
for tokens_used which is never set)
- Remove the /no_think system prompt hack (superseded by API param)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Incorporates the useful new features from PR #135 (by @gridworks) on top
of the existing PrivacyScanner implementation:
- Add DNS configuration check (macOS, via scutil --dns)
- Add --json flag to `jarvis scan` for machine-readable output
- Add --no-scan flag to `jarvis init` to skip the post-init audit
- Expand remote-access process list (ngrok, tailscaled, cloudflared, ZeroTier)
- Upgrade `jarvis scan` output from plain text to Rich table
- Add GET /v1/security/scan API endpoint
- Add tests for all new features (30 tests, all passing)
Closes#133
Co-Authored-By: gridworks <5502067+gridworks@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add current_activity field to managed agents that the executor updates
at each phase of a tick (loading model, delivering messages, generating
response, retrying, finalizing). The frontend polls this every 2s and
displays the live status instead of static "Agent is thinking...".
Backend:
- Add current_activity column to managed_agents (migration)
- Add _set_activity helper to AgentExecutor
- Update activity at: start_tick, model load, message delivery,
generation, retry, finalize
- Clear activity on end_tick
Frontend:
- InteractTab polls both messages and agent status in parallel
- Shows current_activity text with pulsing indicator
- Falls back to "Agent is thinking..." if activity is empty
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two fixes for the last 4 failing PinchBench tasks:
1. Tool arguments (tasks 05, 07, 13): Arguments were lost in the
pipeline — ToolExecutor stored them but system.py stripped them
from the tool_results dict, so the LLM judge saw write_file({})
instead of the actual content. Now pipe arguments through:
_stubs.py → system.py → scorer transcript.
2. Multi-session tasks (task 22): Parse `sessions` field from task
frontmatter, use first session's prompt as record.problem, and
execute remaining sessions sequentially within the workspace
context in _process_one().
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
get_engine() probes all engines with health checks, which can load
different models and interfere with in-flight Ollama requests, causing
intermittent empty responses. Create a plain OllamaEngine directly
from config instead.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Append /no_think to system prompt to prevent Qwen3.5 from consuming
all tokens on extended thinking and producing empty visible output
- Retry once if agent returns empty content
- Add debug logging to _make_lightweight_system for engine diagnostics
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Gemini 3.1+ reasoning models require a thought_signature field in
function_call parts when replaying conversation history. Without it,
the API returns 400 INVALID_ARGUMENT on every multi-turn tool call.
Changes:
- CloudEngine: capture thought_signature from Gemini responses and
store in _thought_sigs dict keyed by tool_call id
- CloudEngine: replay thought_signature when building function_call
parts for Gemini conversation history
- native_openhands: thread thought_signature through via side dict
(ToolCall uses slots, can't add dynamic attributes)
- Add PinchBench eval configs for Claude Opus 4.6, Gemini 3.1 Pro,
Nemotron-3-Super, Qwen 122B, and Qwen 35B
Impact: Gemini 3.1 Pro PinchBench score 4% → 78%
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The server's engine is wrapped in MultiEngine → InstrumentedEngine →
GuardrailsEngine. When reused from a background thread for agent ticks,
this chain returns empty content. Create a fresh OllamaEngine for each
tick instead, which reliably returns model output.
Also fixes Run Now endpoint to use the same lightweight system approach.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Update STARTUP_MODEL from 2b to 4b for better quality on first launch.
Also update preferred_model() to prefer STARTUP_MODEL when it fits,
rather than always picking the third-largest model. This gives a
consistent default across machines while still falling back to
RAM-appropriate sizing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Root cause: _run_tick and _immediate_tick called SystemBuilder().build()
which picks the first model from Ollama (qwen3.5:35b) instead of the
model the server was started with (e.g. qwen3.5:9b). A 0.6B query was
running on a 35B model, causing 5+ minute stalls.
Fix: reuse the server's engine/model from app.state via a lightweight
system facade instead of rebuilding the full JarvisSystem.
Also:
- Add detailed logging to AgentExecutor (model, pending messages,
timing, content length, errors with tracebacks)
- Add logging to immediate tick lifecycle
- Remove Queue button from Interact tab (single Send button)
- Enter key now sends immediately instead of queueing
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Backend:
- Immediate-mode messages now trigger a background tick so the agent
actually processes and responds (previously they were just stored)
Frontend (Interact tab):
- Reverse message order so newest appear at bottom near the input box
- Filter out agent responses with empty content (blank bubbles)
- Add "Agent is thinking..." indicator with pulsing dot while processing
- Show timestamps instead of raw mode/status labels
- Poll for new messages every 3s so responses appear automatically
- Only auto-scroll to bottom on initial tab load, not on every poll
update (prevents hijacking the user's scroll position)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
_apply_toml_section only normalized TOML arrays to comma-separated strings
for real dataclass fields, but backward-compat property setters like
reward_weights also expect string input. When a user's config.toml had
an array value for a property-backed attribute, the raw list was passed
to the setter which called .split(",") on it, causing:
'list' object has no attribute 'split'
This also hardens serve.py against the same issue when reading
config.agent.tools.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Merge main into fix/ssrf-check, keeping both the auto-recover
logic for error-state agents and the async streaming support
for the send_message endpoint.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>