Skip to content

Add argrouter submission - #220

Open
Harunercul wants to merge 3 commits into
RouteWorks:mainfrom
Harunercul:argrouter-submission
Open

Harunercul wants to merge 3 commits into
RouteWorks:mainfrom
Harunercul:argrouter-submission

Conversation

@Harunercul

@Harunercul Harunercul commented Oct 7, 2026 •

Copy link
Copy Markdown

Summary

Leaderboard submission for argrouter (argrouter). One OpenRouter call per query, no cascades, no ensembling.

Method

  • Routing signal: the question content only. The router drops the first and last paragraph of the prompt (instruction and answer-format text) and strips line-leading field labels, option letters and None placeholders with generic patterns. It reads no RouterArena config, template or dataset file.
  • For each candidate model it predicts P(correct) and the expected cost from external calibration data, where cost is the billed token usage of that model on similar calibration items (reasoning tokens included), priced at the table prices. It picks argmax P(correct) − λ · expected cost.
  • Reasoning effort is part of the choice: google/gemini-3.8-flash-reasoning-low is gemini-3.8-flash with OpenRouter reasoning.effort = low. The small addition to _call_openrouter maps a generic -reasoning-<effort> suffix to that parameter; all other calls are unchanged.
  • The pool (5 models) and λ were selected by cross-validation on the calibration set and fixed before RouterArena was routed.

Calibration data

  • Held-out items from the same 23 public source benchmarks. Every RouterArena item (full, sub_10 and robustness) is excluded by Global Index and by normalised question text; overlap checked as zero.
  • Graded with RouterArena's own scorers; LiveCodeBench answers executed in a no-network container with livecodebench_util.
  • Items were sampled per source with fixed quotas; sources are weighted equally when fitting and when selecting the pool and λ.

Local result (RouterArena scripts, full split)

Metric Value
Arena score 0.7629
Accuracy 79.40%
Cost per 1K queries $0.558
Robustness 0.676

Four generations originally returned no answer (the model's reasoning reached the provider's maximum output length). As the evaluation bot suggested, they were regenerated once with the same model and the same ModelInference call: three now have answers; one (MMLUPro_math_6607, xiaomi/mimo-v2.6-flash) again returned no answer and stays success: false (scored 0). The figures in the table above are from before the regeneration and include the billed cost of the failed attempts.

Prices

New model_cost.json entries use OpenRouter list prices as of 2026-10-07 (USD per 1M input / output tokens):

Model Price
google/gemini-3.8-flash-reasoning-low 0.75 / 3.75 (same as gemini-3.8-flash)
z-ai/glm-5.3-flash 0.15 / 0.50
openai/gpt-6-luna 0.10 / 0.50
xiaomi/mimo-v2.6-flash 0.14 / 0.28

google/gemma-4-31b-it uses its existing entry.

Leaderboard display request (if accepted)

  • Router name: argrouter
  • Affiliation: @Harunercul
  • Links: [Code] (inference library; the training pipeline is not public)
  • Type: closed-source

@Harunercul

Copy link
Copy Markdown
Author

/evaluate

@Harunercul

Copy link
Copy Markdown
Author

Fixed the validation failure: four generations with no answer (reasoning hit the provider's max output length) are now marked success: false with an error note; check_config_prediction_files.py argrouter full --check-generated-result passes locally. Their billed cost is described in the PR description.

/evaluate

@Harunercul

Copy link
Copy Markdown
Author

/evaluate

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

Router Evaluation Results

Router: argrouter
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7629
Accuracy 79.34%
Total Cost $4.544507
Avg Cost per Query $0.000541
Avg Cost per 1K Queries $0.5410
Number of Queries 8400
Abnormal Entries 4
Robustness Score 0.6762

⚠️ 4 of 8400 queries (0.0%) had no valid generation (inference failed / empty answer) and were scored as incorrect (0). These queries still count toward the denominator, so accuracy and cost reflect the full query set. Please regenerate predictions for these queries and resubmit for a complete evaluation.


Evaluation completed by RouterArena automated workflow

@Harunercul

Copy link
Copy Markdown
Author

As suggested by the evaluation bot, the 4 entries without a valid generation were regenerated once with the same model and the same ModelInference.infer call. 3 now have answers; MMLUPro_math_6607 (xiaomi/mimo-v2.6-flash) again ran out of output length with no answer and remains success: false. check_config_prediction_files.py argrouter full --check-generated-result passes locally.

@Harunercul

Copy link
Copy Markdown
Author

/evaluate

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

Router Evaluation Results

Router: argrouter
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7630
Accuracy 79.36%
Total Cost $4.584784
Avg Cost per Query $0.000546
Avg Cost per 1K Queries $0.5458
Number of Queries 8400
Abnormal Entries 1
Robustness Score 0.6762

⚠️ 1 of 8400 queries (0.0%) had no valid generation (inference failed / empty answer) and were scored as incorrect (0). These queries still count toward the denominator, so accuracy and cost reflect the full query set. Please regenerate predictions for these queries and resubmit for a complete evaluation.


Evaluation completed by RouterArena automated workflow

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant