Andrew Park
08fe3a6c98
evals: switch the default LLM judge to claude-haiku-4-5
...
Replaces gpt-5-mini-2025-08-07 as the default judge everywhere: the 147
eval configs plus the hardcoded defaults in the evals CLI, recipe composer,
comparison config generator, taubench user-simulator, trial runner and
OptimizeConfig.
Haiku is roughly 3x faster on the judging path and we are out of OpenAI
quota, so gpt-5-mini was failing closed on long eval runs.
2026-07-11 11:59:32 -07:00
..
2026-05-25 23:51:29 +00:00
2026-07-11 11:55:59 -07:00
2026-05-27 09:36:26 -07:00
2026-03-20 18:35:18 -07:00
2026-06-03 19:54:03 -07:00
2026-06-01 10:26:11 -07:00
2026-05-25 11:07:35 -07:00
2026-06-02 20:14:42 -07:00
2026-03-17 00:13:49 +00:00
2026-05-25 23:56:55 +00:00
2026-06-01 09:23:22 -07:00
2026-06-01 13:17:38 -07:00
2026-05-15 21:31:13 -07:00
2026-04-17 15:22:17 -07:00
2026-05-20 03:36:55 +00:00
2026-06-03 16:59:29 -07:00
2026-03-20 18:35:18 -07:00
2026-04-13 15:03:03 -07:00
2026-07-11 11:59:32 -07:00
2026-06-03 18:27:00 -07:00
2026-03-20 18:35:18 -07:00
2026-05-14 20:11:24 -07:00
2026-04-17 11:02:56 -07:00
2026-05-14 20:11:24 -07:00
2026-06-03 16:59:47 -07:00
2026-03-20 18:35:18 -07:00
2026-03-20 18:35:18 -07:00
2026-04-03 11:05:41 -07:00
2026-04-11 13:19:19 -07:00
2026-06-01 11:57:44 -07:00
2026-06-01 11:31:19 -07:00
2026-03-31 19:25:29 +05:30
2026-04-13 15:03:39 -07:00
2026-04-13 15:03:03 -07:00
2026-05-20 20:22:44 +00:00
2026-03-20 18:35:18 -07:00
2026-07-11 11:55:59 -07:00
2026-06-02 19:31:05 -07:00
2026-04-01 21:06:20 -07:00
2026-03-20 18:35:18 -07:00
2026-03-12 17:29:39 +00:00
2026-05-05 19:11:15 -07:00
2026-04-03 11:05:41 -07:00
2026-04-17 11:02:56 -07:00