Repository navigation
Add argrouter submission - #220
Harunercul wants to merge 3 commits into
Conversation
|
/evaluate |
|
Fixed the validation failure: four generations with no answer (reasoning hit the provider's max output length) are now marked /evaluate |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Evaluation completed by RouterArena automated workflow |
|
As suggested by the evaluation bot, the 4 entries without a valid generation were regenerated once with the same model and the same |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Evaluation completed by RouterArena automated workflow |
Summary
Leaderboard submission for argrouter (
argrouter). One OpenRouter call per query, no cascades, no ensembling.Method
Noneplaceholders with generic patterns. It reads no RouterArena config, template or dataset file.argmax P(correct) − λ · expected cost.google/gemini-3.8-flash-reasoning-lowis gemini-3.8-flash with OpenRouterreasoning.effort = low. The small addition to_call_openroutermaps a generic-reasoning-<effort>suffix to that parameter; all other calls are unchanged.Calibration data
livecodebench_util.Local result (RouterArena scripts, full split)
Four generations originally returned no answer (the model's reasoning reached the provider's maximum output length). As the evaluation bot suggested, they were regenerated once with the same model and the same
ModelInferencecall: three now have answers; one (MMLUPro_math_6607,xiaomi/mimo-v2.6-flash) again returned no answer and stayssuccess: false(scored 0). The figures in the table above are from before the regeneration and include the billed cost of the failed attempts.Prices
New
model_cost.jsonentries use OpenRouter list prices as of 2026-10-07 (USD per 1M input / output tokens):google/gemini-3.8-flash-reasoning-lowz-ai/glm-5.3-flashopenai/gpt-6-lunaxiaomi/mimo-v2.6-flashgoogle/gemma-4-31b-ituses its existing entry.Leaderboard display request (if accepted)