Files
OpenJarvis/tests
Tarun SureshandClaude Opus 4.6 bf24ffc527 feat: joint multi-benchmark optimization with LLM-guided search
Add Claude Opus-guided optimization loop that jointly optimizes agent
configs across multiple benchmarks (TerminalBench-native, GAIA, HLE)
with weighted accuracy aggregation and Pareto frontier computation.

Key additions:
- MultiBenchTrialRunner with native terminal-bench v2 Docker execution
- LLM optimizer: fixed params injection, structured trial feedback
- OptimizationStore (SQLite) with per-benchmark score persistence
- System prompt passthrough from optimizer → eval → agent
- Custom terminal-bench agent (OpenJarvisTerminalBenchAgent) avoiding
  LiteLLM serialization issues with agent_import_path
- TOML configs for Qwen3-235B joint agentic optimization
- CLI: `jarvis optimize` command with dry-run support

Early results on Qwen3-235B-A22B-Instruct-2507-FP8 (4xA100):
  Best config: native_openhands, temp=0.0, 20 turns → 5.6% weighted acc
  (TB2=5%, GAIA=6%, HLE=6%)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-06 18:44:24 +00:00
..
2026-03-04 19:35:33 -08:00