Files
OpenJarvis/tests
Andrew Park 08fe3a6c98 evals: switch the default LLM judge to claude-haiku-4-5
Replaces gpt-5-mini-2025-08-07 as the default judge everywhere: the 147
eval configs plus the hardcoded defaults in the evals CLI, recipe composer,
comparison config generator, taubench user-simulator, trial runner and
OptimizeConfig.

Haiku is roughly 3x faster on the judging path and we are out of OpenAI
quota, so gpt-5-mini was failing closed on long eval runs.
2026-07-11 11:59:32 -07:00
..
2026-05-27 09:36:26 -07:00
2026-04-17 15:22:17 -07:00
2026-03-31 19:25:29 +05:30
2026-03-12 17:29:39 +00:00