Skip to content

fix(paperbench): handle uneven seed counts in run parsing - #173

Open
sylvesterkaczmarek wants to merge 2 commits into
openai:mainfrom
sylvesterkaczmarek:fix/paperbench-uneven-seed-runs
Open

fix(paperbench): handle uneven seed counts in run parsing#173
sylvesterkaczmarek wants to merge 2 commits into
openai:mainfrom
sylvesterkaczmarek:fix/paperbench-uneven-seed-runs

Conversation

@sylvesterkaczmarek

Copy link
Copy Markdown

Summary

Make PaperBench run reconstruction handle papers with uneven numbers of recorded seeds without raising IndexError.

parse_run_data() builds seed-indexed EvaluationRun objects up to the largest per-paper seed count. For a paper with fewer runs, the loop currently skips only when:

seed == len(data)

If another paper has two or more additional seeds, a later iteration reaches seed > len(data) and still evaluates data[seed], crashing the metrics parser.

Fix

Treat every out-of-range seed as unavailable:

if seed >= len(data):
    continue

This preserves the intended behavior described by the surrounding comment: incomplete evaluation runs are retained with only the paper evaluations that exist for that seed.

Regression coverage

Added a JSONL-level regression with one paper containing three recorded runs and another containing one. The test verifies that parsing produces three evaluation runs, with the shorter paper present only in the first run, instead of crashing on the third seed.

Risk

Low. Equal-length run sets are unchanged. The only behavior change is for uneven per-paper seed counts that would otherwise index past the end of a paper's data.

Issue number

N/A. Found while auditing PaperBench metrics reconstruction.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant