Skip to content

Commit cff9659

Browse files
yl231Louie Luclaude
authored
Add cruq-router submission + leaderboard (#197) (#200)
* Add cruq-router submission (#197) Self-consistency cascade (qwen3-235b K=4 probes at T=0.7 -> escalate to deepseek-v4-flash on low agreement). Acc 71.35%, $0.18/1K, arena 0.7077. Audited: 8400 unique rows, real tokens (all probes charged), standard market pricing. Resolved from cruq-ai/RouterArena:cruq-router (fork not pushable). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Leaderboard: add cruq-router (#197) at #16 Arena 70.77 / acc 71.35% / $0.18 / robustness 81.67. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Louie Lu <yl231@datalab2.cs.rice.edu> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
1 parent 06e6890 commit cff9659

9 files changed

Lines changed: 149 additions & 18 deletions

File tree

‎README.md‎

Lines changed: 17 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -52,22 +52,23 @@ For more details, please see our [website](https://routeworks.github.io/leaderbo
5252
| 13 | [Hybrid Router]() | 👤&nbsp;[@mikemao27](https://github.com/mikemao27) | 72.08 | 71.38 | $0.04 | 89.87 | 94.19 | 92.81 | — | 96.67 |
5353
| 14 | [R2-Router](https://arxiv.org/abs/2602.02823/) | 🎓&nbsp;UCF | 71.60 | 71.23 | $0.06 | 24.51 | 48.70 | 99.85 | — | 45.71 |
5454
| 15 | [LLM Router](https://github.com/ypollak2/llm-router)&nbsp;[[PyPI]](https://pypi.org/project/llm-routing/) | 👤&nbsp;[@ypollak2](https://github.com/ypollak2) | 71.26 | 72.05 | $0.20 | 18.01 | 20.46 | 89.13 | — | 30.00 |
55-
| 16 | [chuzom-solo-v32]() | 👤&nbsp;[@ypollak2](https://github.com/ypollak2) | 70.61 | 70.59 | $0.10 | — | — | — | — | 100.00 |
56-
| 17 | [Azure-Model-Router](https://ai.azure.com/catalog/models/model-router)&nbsp;[[Web]](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/model-router) | 💼&nbsp;Microsoft | 70.42 | 72.94 | $0.73 | — | — | — | — | 71.43 |
57-
| 18 | [Auto Router]() | 👤&nbsp;[@cxf2015](https://github.com/cxf2015) | 70.05 | 70.17 | $0.12 | 37.58 | 40.02 | 86.04 | — | 49.52 |
58-
| 19 | [Lynkr]() | 👤&nbsp;[@vishalveerareddy123](https://github.com/vishalveerareddy123) | 67.65 | 68.41 | $0.29 | 10.97 | 16.08 | 84.48 | — | 92.38 |
59-
| 20 | [MIRT‑BERT](https://arxiv.org/pdf/2506.01048)&nbsp;[[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓&nbsp;USTC | 66.89 | 66.88 | $0.15 | 3.44 | 19.62 | 78.18 | 27.03 | 61.19 |
60-
| 21 | [NIRT‑BERT](https://arxiv.org/pdf/2506.01048)&nbsp;[[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓&nbsp;USTC | 66.12 | 66.34 | $0.21 | 3.83 | 14.04 | 77.88 | 10.42 | 49.29 |
61-
| 22 | [AsiaInfo-Router]() | 👤&nbsp;[@Uncle-LL](https://github.com/Uncle-LL) | 65.87 | 75.20 | $8.54 | — | — | — | — | 69.52 |
62-
| 23 | [GPT‑5](https://openai.com/index/introducing-gpt-5/) | 💼&nbsp;OpenAI | 64.32 | 73.96 | $10.02 | — | — | — | — | — |
63-
| 24 | [CARROT](https://arxiv.org/abs/2502.03261)&nbsp;[[Code]](https://github.com/somerstep/CARROT)&nbsp;[[HF]](https://huggingface.co/CARROT-LLM-Routing) | 🎓&nbsp;UMich | 63.87 | 67.21 | $2.06 | 2.68 | 6.77 | 78.63 | 1.50 | 89.05 |
64-
| 25 | [Chayan](https://huggingface.co/adaptive-classifier/chayan)&nbsp;[[HF]](https://huggingface.co/adaptive-classifier/chayan) | 🎓&nbsp;Adaptive&nbsp;Classifier | 63.83 | 64.89 | $0.56 | 43.03 | 43.75 | 88.74 | — | — |
65-
| 26 | [RouterBench‑MLP](https://arxiv.org/pdf/2403.12031)&nbsp;[[Code]](https://github.com/withmartian/routerbench)&nbsp;[[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓&nbsp;Martian | 57.56 | 61.62 | $4.83 | 13.39 | 24.45 | 83.32 | 90.91 | 80.00 |
66-
| 27 | [NotDiamond](https://www.notdiamond.ai/) | 💼&nbsp;NotDiamond | 57.29 | 60.83 | $4.10 | 1.55 | 2.14 | 76.81 | — | 55.91 |
67-
| 28 | [GraphRouter](https://arxiv.org/abs/2410.03834)&nbsp;[[Code]](https://github.com/ulab-uiuc/GraphRouter) | 🎓&nbsp;UIUC | 57.22 | 57.00 | $0.34 | 4.73 | 38.33 | 74.25 | 2.70 | 94.29 |
68-
| 29 | [RouterBench‑KNN](https://arxiv.org/pdf/2403.12031)&nbsp;[[Code]](https://github.com/withmartian/routerbench)&nbsp;[[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓&nbsp;Martian | 55.48 | 58.69 | $4.27 | 13.09 | 25.49 | 78.77 | 1.33 | 83.33 |
69-
| 30 | [RouteLLM](https://arxiv.org/abs/2406.18665)&nbsp;[[Code]](https://github.com/lm-sys/RouteLLM)&nbsp;[[HF]](https://huggingface.co/routellm) | 🎓&nbsp;Berkeley | 48.07 | 47.04 | $0.27 | 99.72 | 99.63 | 68.76 | 0.40 | 100.00 |
70-
| 31 | [RouterDC](https://arxiv.org/abs/2409.19886)&nbsp;[[Code]](https://github.com/shuhao02/RouterDC) | 🎓&nbsp;SUSTech | 33.75 | 32.01 | $0.07 | 39.84 | 73.00 | 49.05 | 10.75 | 85.24 |
55+
| 16 | [cruq-router]() | 👤&nbsp;[@nabaruns](https://github.com/nabaruns) | 70.77 | 71.35 | $0.18 | — | — | — | — | 81.67 |
56+
| 17 | [chuzom-solo-v32]() | 👤&nbsp;[@ypollak2](https://github.com/ypollak2) | 70.61 | 70.59 | $0.10 | — | — | — | — | 100.00 |
57+
| 18 | [Azure-Model-Router](https://ai.azure.com/catalog/models/model-router)&nbsp;[[Web]](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/model-router) | 💼&nbsp;Microsoft | 70.42 | 72.94 | $0.73 | — | — | — | — | 71.43 |
58+
| 19 | [Auto Router]() | 👤&nbsp;[@cxf2015](https://github.com/cxf2015) | 70.05 | 70.17 | $0.12 | 37.58 | 40.02 | 86.04 | — | 49.52 |
59+
| 20 | [Lynkr]() | 👤&nbsp;[@vishalveerareddy123](https://github.com/vishalveerareddy123) | 67.65 | 68.41 | $0.29 | 10.97 | 16.08 | 84.48 | — | 92.38 |
60+
| 21 | [MIRT‑BERT](https://arxiv.org/pdf/2506.01048)&nbsp;[[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓&nbsp;USTC | 66.89 | 66.88 | $0.15 | 3.44 | 19.62 | 78.18 | 27.03 | 61.19 |
61+
| 22 | [NIRT‑BERT](https://arxiv.org/pdf/2506.01048)&nbsp;[[Code]](https://github.com/Mercidaiha/IRT-Router) | 🎓&nbsp;USTC | 66.12 | 66.34 | $0.21 | 3.83 | 14.04 | 77.88 | 10.42 | 49.29 |
62+
| 23 | [AsiaInfo-Router]() | 👤&nbsp;[@Uncle-LL](https://github.com/Uncle-LL) | 65.87 | 75.20 | $8.54 | — | — | — | — | 69.52 |
63+
| 24 | [GPT‑5](https://openai.com/index/introducing-gpt-5/) | 💼&nbsp;OpenAI | 64.32 | 73.96 | $10.02 | — | — | — | — | — |
64+
| 25 | [CARROT](https://arxiv.org/abs/2502.03261)&nbsp;[[Code]](https://github.com/somerstep/CARROT)&nbsp;[[HF]](https://huggingface.co/CARROT-LLM-Routing) | 🎓&nbsp;UMich | 63.87 | 67.21 | $2.06 | 2.68 | 6.77 | 78.63 | 1.50 | 89.05 |
65+
| 26 | [Chayan](https://huggingface.co/adaptive-classifier/chayan)&nbsp;[[HF]](https://huggingface.co/adaptive-classifier/chayan) | 🎓&nbsp;Adaptive&nbsp;Classifier | 63.83 | 64.89 | $0.56 | 43.03 | 43.75 | 88.74 | — | — |
66+
| 27 | [RouterBench‑MLP](https://arxiv.org/pdf/2403.12031)&nbsp;[[Code]](https://github.com/withmartian/routerbench)&nbsp;[[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓&nbsp;Martian | 57.56 | 61.62 | $4.83 | 13.39 | 24.45 | 83.32 | 90.91 | 80.00 |
67+
| 28 | [NotDiamond](https://www.notdiamond.ai/) | 💼&nbsp;NotDiamond | 57.29 | 60.83 | $4.10 | 1.55 | 2.14 | 76.81 | — | 55.91 |
68+
| 29 | [GraphRouter](https://arxiv.org/abs/2410.03834)&nbsp;[[Code]](https://github.com/ulab-uiuc/GraphRouter) | 🎓&nbsp;UIUC | 57.22 | 57.00 | $0.34 | 4.73 | 38.33 | 74.25 | 2.70 | 94.29 |
69+
| 30 | [RouterBench‑KNN](https://arxiv.org/pdf/2403.12031)&nbsp;[[Code]](https://github.com/withmartian/routerbench)&nbsp;[[HF]](https://huggingface.co/datasets/withmartian/routerbench) | 🎓&nbsp;Martian | 55.48 | 58.69 | $4.27 | 13.09 | 25.49 | 78.77 | 1.33 | 83.33 |
70+
| 31 | [RouteLLM](https://arxiv.org/abs/2406.18665)&nbsp;[[Code]](https://github.com/lm-sys/RouteLLM)&nbsp;[[HF]](https://huggingface.co/routellm) | 🎓&nbsp;Berkeley | 48.07 | 47.04 | $0.27 | 99.72 | 99.63 | 68.76 | 0.40 | 100.00 |
71+
| 32 | [RouterDC](https://arxiv.org/abs/2409.19886)&nbsp;[[Code]](https://github.com/shuhao02/RouterDC) | 🎓&nbsp;SUSTech | 33.75 | 32.01 | $0.07 | 39.84 | 73.00 | 49.05 | 10.75 | 85.24 |
7172

7273
🎓 Open-source  💼 Closed-source 
7374

‎leaderboard_manifest.yaml‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -202,6 +202,15 @@ routers:
202202
github_url: "https://github.com/KT-A-Autonomous-Tech-Team/ModelRouter"
203203
type: "open-source"
204204

205+
- readme_name: "cruq-router"
206+
website_name: "cruq-router"
207+
prediction: "cruq-router"
208+
category_key: "cruq-router"
209+
flip_key: "cruq-router"
210+
website:
211+
affiliation: "@nabaruns"
212+
github_url: "https://github.com/nabaruns"
213+
205214
# --- Externally-evaluated baselines (headline from README; derived data on
206215
# the website is preserved as-is) ---
207216
- readme_name: "MIRT-BERT"

‎model_cost/model_cost.json‎

Lines changed: 22 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -332,8 +332,8 @@
332332
"output_token_price_per_million": 1.5
333333
},
334334
"MiniMax-M3": {
335-
"input_token_price_per_million": 0.60,
336-
"output_token_price_per_million": 2.40
335+
"input_token_price_per_million": 0.6,
336+
"output_token_price_per_million": 2.4
337337
},
338338
"agnes-2.0-flash": {
339339
"input_token_price_per_million": 0.03,
@@ -363,6 +363,26 @@
363363
"input_token_price_per_million": 0.14,
364364
"output_token_price_per_million": 0.28
365365
},
366+
"openai/gpt-5-mini": {
367+
"input_token_price_per_million": 0.25,
368+
"output_token_price_per_million": 2.0
369+
},
370+
"anthropic/claude-sonnet-4.5": {
371+
"input_token_price_per_million": 3.0,
372+
"output_token_price_per_million": 15.0
373+
},
374+
"openai/gpt-4o-mini": {
375+
"input_token_price_per_million": 0.15,
376+
"output_token_price_per_million": 0.6
377+
},
378+
"google/gemini-2.5-flash-lite": {
379+
"input_token_price_per_million": 0.1,
380+
"output_token_price_per_million": 0.4
381+
},
382+
"google/gemini-2.5-pro": {
383+
"input_token_price_per_million": 1.25,
384+
"output_token_price_per_million": 10.0
385+
},
366386
"google/gemma-4-31b-it": {
367387
"input_token_price_per_million": 0.08,
368388
"output_token_price_per_million": 0.35
Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
{"pipeline_params":{"router_name":"cruq-router","router_cls_name":"CruqSCRouter","models":["qwen/qwen3-235b-a22b-2507","deepseek/deepseek-v4-flash"],"k":4,"tau":0.6,"probe_model":"qwen/qwen3-235b-a22b-2507","escalate_model":"deepseek/deepseek-v4-flash"}}

‎router_inference/predictions/cruq-router-robustness.json‎

Lines changed: 1 addition & 0 deletions
Large diffs are not rendered by default.

‎router_inference/predictions/cruq-router.json‎

Lines changed: 1 addition & 0 deletions
Large diffs are not rendered by default.

‎router_inference/router/__init__.py‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@
1010
from router_inference.router.chuzom_solo_v32 import ChuzomSoloV32Router
1111
from router_inference.router.llm_router import LLMRouter
1212
from router_inference.router.lynkr_router import LynkrRouter
13+
from router_inference.router.cruq_sc_router import CruqSCRouter
1314

1415
__all__ = [
1516
"BaseRouter",
@@ -19,4 +20,5 @@
1920
"LLMRouter",
2021
"ChuzomSoloV32Router",
2122
"LynkrRouter",
23+
"CruqSCRouter",
2224
]
Lines changed: 93 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,93 @@
1+
# SPDX-FileCopyrightText: Copyright contributors to the cruq.ai project
2+
# SPDX-License-Identifier: Apache-2.0
3+
4+
"""
5+
cruq self-consistency cascade router.
6+
7+
The router sees only the prompt string. It probes the cheapest capable model
8+
(qwen3-235b-a22b) K times at temperature>0 and measures self-consistency: the
9+
fraction of the K samples that agree on the final \\boxed{} answer. High
10+
agreement is a strong, prompt-independent confidence signal, so:
11+
12+
- consistency >= tau -> keep qwen's majority-vote answer (cost = K cheap probes)
13+
- consistency < tau -> escalate to deepseek-v4-flash (a stronger model)
14+
15+
This is an *inference-time* signal, not a prompt classifier. Across the whole
16+
research programme, no prompt-only selector (lexical difficulty, per-model
17+
P(correct) heads, calibrated thresholds, or a domain classifier) beat the best
18+
single model, because per-query difficulty is close to unpredictable from prompt
19+
text. Model-agreement at inference is what breaks that ceiling.
20+
21+
Compliance: nothing here is fit on RouterArena data. K and tau are priors set in
22+
the config; the probe model and escalation model are fixed choices. No labels,
23+
metadata, or ground truth are read at decision time.
24+
25+
Reproducibility: the K qwen probes per query are cached to
26+
phase2/data/qwen_sc_full.jsonl by phase2/sample_qwen_sc_full.py. This class reads
27+
that cache to make the same keep/escalate decision the submitted prediction file
28+
encodes. The final prediction file (with the majority-vote answer and honest
29+
K-probe token accounting) is assembled by phase2/build_sc_submission.py.
30+
"""
31+
32+
import json
33+
import os
34+
import re
35+
import collections
36+
from typing import Dict, List
37+
38+
from router_inference.router.base_router import BaseRouter
39+
40+
_BOXED = re.compile(r"\\boxed\{+([^{}]*)\}+")
41+
42+
43+
def _norm(s: str) -> str:
44+
return re.sub(r"[^a-z0-9]", "", str(s).lower())
45+
46+
47+
class CruqSCRouter(BaseRouter):
48+
"""Self-consistency cascade: probe the cheap model K times, escalate on disagreement."""
49+
50+
def __init__(self, router_name: str):
51+
super().__init__(router_name)
52+
params = self.config["pipeline_params"]
53+
self.k = int(params.get("k", 4))
54+
self.tau = float(params.get("tau", 0.6))
55+
self.probe_model = params.get("probe_model", "qwen/qwen3-235b-a22b-2507")
56+
self.escalate_model = params.get("escalate_model", "deepseek/deepseek-v4-flash")
57+
58+
here = os.path.dirname(os.path.abspath(__file__))
59+
root = os.path.dirname(os.path.dirname(here))
60+
# prompt -> global index, so a raw query string can find its cached probes
61+
self._prompt_to_gi: Dict[str, str] = {}
62+
for path in ("dataset/router_data.json", "dataset/router_data_10.json"):
63+
p = os.path.join(root, path)
64+
if os.path.exists(p):
65+
for e in json.load(open(p, encoding="utf-8")):
66+
self._prompt_to_gi[e["prompt_formatted"]] = e["global index"]
67+
68+
# global index -> list of normalized boxed answers from the K probes
69+
self._samples: Dict[str, List[str]] = collections.defaultdict(list)
70+
cache = os.path.join(root, "phase2", "data", "qwen_sc_full.jsonl")
71+
if os.path.exists(cache):
72+
for line in open(cache, encoding="utf-8"):
73+
try:
74+
r = json.loads(line)
75+
except Exception:
76+
continue
77+
if r["s"] < self.k:
78+
self._samples[r["gi"]].append(_norm(r["boxed"]))
79+
80+
def _get_prediction(self, query: str) -> str:
81+
gi = self._prompt_to_gi.get(query)
82+
if gi is None:
83+
# Unknown query (no cached probes): fall back to the cheap probe model.
84+
return self.probe_model
85+
samples = self._samples.get(gi, [])
86+
boxes = [b for b in samples if b]
87+
if not boxes:
88+
# Free-form dataset (no \boxed answer): self-consistency can't apply, so keep
89+
# the cheap probe model rather than pay to escalate on a signal we don't have.
90+
return self.probe_model
91+
top = collections.Counter(boxes).most_common(1)[0][1]
92+
consistency = top / len(samples)
93+
return self.probe_model if consistency >= self.tau else self.escalate_model

‎universal_model_names.py‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -41,6 +41,9 @@
4141
"gemini-2.5-flash",
4242
"gemini-2.5-pro",
4343
"google/gemini-3.1-flash-lite",
44+
"google/gemini-2.5-flash-lite",
45+
"google/gemini-2.5-pro",
46+
"anthropic/claude-sonnet-4.5",
4447
"gemini-3-flash-preview",
4548
# Mistral models
4649
"mistral-medium",

0 commit comments

Comments
 (0)