Skip to content

Paix2-router (#164): integrity review, suspended from the leaderboard pending the author's response #211

Description

@yl231

Summary

Paix2-router (#164, @xufan866) is temporarily suspended from the leaderboard while we review it under the README's evaluation-only rule:

Submissions that train, fit, or tune any router component on RouterArena data (including the label files) will be rejected, and any accepted submission found in violation will be withdrawn.

The submission's own prediction file shows routing choices on the 809 public queries that match which candidate answers were graded correct, with no exceptions. We have not concluded how this happened. We are asking @xufan866 to explain it and to provide a reproducible pipeline. This issue is the place for that response; the deadline is 2026-10-11.

Suspension is not a final decision. The prediction files stay in the repository; only the leaderboard row is removed for now.

Evidence

Every number below comes from the submission's own files and the issues linked here, and can be re-run with the script at the end of this issue.

1. No missed correct answers on the 809 public queries

For each of the 809 queries in dataset/router_data_10.json, the submission includes all four candidate models' answers and grades (3,236 candidate records). On the 630 queries where at least one candidate is graded correct, the routed model is correct on all 630. This is why the bot reported Opt.Acc = 1.0000. The submission's own Jul 14 version, built on the same stored answers, picked a wrong model on 5 of these 630.

A router that sees only the question cannot know in advance which candidate will answer correctly, so some misses are expected.

2. The Jul 27 revision re-picked among already-graded answers (#190)

As @loswald first reported in #190, compare commit 460cfcbf83 (Jul 14, paix-router.json) with commit 4620694fb6 (Jul 27–28, Paix2-router.json, identical to the file on main):

  • all 3,236 candidate records are unchanged;
  • the routed model changes on 294 of the 809 public queries;
  • each of the 294 new routed answers is byte-identical to the Jul 14 stored answer for that model, so no new answers were generated for these queries.

How the 294 changes fall, using the grades already present in the Jul 14 file:

Change Count
Correct → cheapest correct candidate 163
Correct → another correct candidate 33
Wrong → correct (all 5 of the Jul 14 misses) 5
Correct → wrong 0
No candidate correct (accuracy unaffected) 93

Two examples:

  • Ethics_deontology_2: only GLM-4-9B's stored answer is correct; agnes, MiniMax-M3 and DeepSeek-R1-8B are wrong. The Jul 14 version picked MiniMax-M3; the Jul 27 version switched to GLM-4-9B.
  • ArcMMLU_358: agnes, MiniMax-M3 and DeepSeek-R1-8B are correct and GLM-4-9B is wrong. The Jul 14 version picked MiniMax-M3; the Jul 27 version switched to agnes, the cheapest correct one.

A change to routing logic that reads only the question would be expected to turn some correct picks into wrong ones. None of the 201 changes on queries with a correct candidate did.

3. The MiniMax-M3 results do not reproduce (#203)

@nilam-divyam re-ran MiniMax-M3 on the 760 queries Paix2 routed to it, with byte-identical prompts, and got 62.8–69.5% across settings, against 84.1% in the submission (#203). 718 of those 760 queries are outside the public 809, so the gap is not limited to the public queries.

A gap like this is what we would expect if a model were chosen only where its answer was already known to be correct: a fresh answer to the same question is correct less often.

Correction to our earlier reply on #190

In August we replied on #190 that the public and remaining queries behaved alike, so the routing generalized. That assumed the remaining queries were not available to submitters. As @loswald pointed out, all 8,400 queries and their labels are public, so that argument does not hold. Thank you to @loswald and @nilam-divyam for careful, well-documented reports.

What we are asking from @xufan866

By 2026-10-11, please reply in this issue with:

  1. An explanation of how the routing decisions were produced, and in particular how the Jul 27 version chose the 294 changed picks on the public queries.
  2. A statement of whether any router component was trained, fit, tuned or selected using RouterArena labels, or using graded results on RouterArena queries.
  3. A reproducible pipeline: the routing component (code, a container, or an endpoint we can call) that chooses a model from the prompt alone, plus each model's inference settings (API or provider, temperature, reasoning mode, max tokens). The answers do not need to match word for word. When we re-run it, the routing decisions and each model's accuracy on its routed queries should be comparable to the submitted results.
  4. Cost: if the router calls more than one model per query, the usage of every call.

What happens next

  • If the pipeline reproduces the routing from prompts alone and the results are comparable, the entry is restored.
  • If it does not, or if there is no response by 2026-10-11, the entry is withdrawn under the evaluation-only rule. A new submission built without RouterArena labels is welcome.

The maintainers make the final decision after reviewing the response.

Re-running this analysis

From a RouterArena checkout:

git fetch origin pull/164/head
python paix2_check.py
paix2_check.py
"""Re-run the Paix2-router (#164) integrity checks from public repository files.

Usage (from a RouterArena checkout):
    git fetch origin pull/164/head
    python paix2_check.py
"""
import json
import subprocess
from collections import Counter

P = "router_inference/predictions/"
JUL14 = "460cfcbf83:" + P + "paix-router.json"   # first version, Jul 14
FINAL = "main:" + P + "Paix2-router.json"         # identical to commit 4620694fb6


def load(spec):
    out = subprocess.run(["git", "show", spec], capture_output=True, text=True, check=True)
    return json.loads(out.stdout)


def tables(rows):
    """Routed model per query, and every graded candidate on queries that carry alternates."""
    routed, cand = {}, {}
    for r in rows:
        q, m, acc = r["global index"], r["prediction"], r.get("accuracy")
        if not r.get("for_optimality"):
            routed[q] = m
        if isinstance(acc, (int, float)):
            ans = (r.get("generated_result") or {}).get("generated_answer")
            cand.setdefault(q, {})[m] = (acc, r.get("cost"), ans)
    return routed, {q: c for q, c in cand.items() if len(c) > 1 and q in routed}


def misses(spec):
    routed, cand = tables(load(spec))
    solvable = [q for q, c in cand.items() if any(v[0] >= 1 for v in c.values())]
    missed = [q for q in solvable if cand[q][routed[q]][0] < 1]
    return len(solvable), len(missed)


print("1) Routed model wrong although a candidate was correct (public queries)")
for name, spec in [("Paix2, final (on main)", FINAL), ("Paix2, Jul 14", JUL14)]:
    solvable, missed = misses(spec)
    print(f"   {name:24} {missed:3} of {solvable} ({100 * missed / solvable:.1f}%)")

print("\n2) Jul 14 -> final: changed picks on the public queries")
r14, c14 = tables(load(JUL14))
rF, cF = tables(load(FINAL))
same = sum(cF[q].get(m, (None,))[0] == c14[q][m][0] for q in c14 for m in c14[q])
print(f"   candidate grades unchanged: {same} of {sum(len(c) for c in c14.values())}")
changed = [q for q in c14 if r14[q] != rF[q]]
reused = sum(cF[q][rF[q]][2] == c14[q][rF[q]][2] for q in changed)
print(f"   changed picks: {len(changed)}; new answer identical to the stored Jul 14 answer: {reused}")


def cheapest_correct(q):
    ok = [(v[1], m) for m, v in c14[q].items() if v[0] >= 1]
    return min(ok)[1] if ok else None


kinds = Counter()
for q in changed:
    old_ok, new_ok, best = c14[q][r14[q]][0] >= 1, c14[q][rF[q]][0] >= 1, cheapest_correct(q)
    if best is None:
        kinds["no candidate correct"] += 1
    elif old_ok and not new_ok:
        kinds["correct -> wrong"] += 1
    elif not old_ok and new_ok:
        kinds["wrong -> correct"] += 1
    elif rF[q] == best:
        kinds["correct -> cheapest correct"] += 1
    else:
        kinds["correct -> another correct"] += 1
for k in ["correct -> cheapest correct", "correct -> another correct", "wrong -> correct",
          "correct -> wrong", "no candidate correct"]:
    print(f"   {k:30} {kinds[k]}")

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions