Parent: #39
Context
The open Martian code-review benchmark evaluates major reviewers on 50 PRs and documents how to add another tool. An external corpus would provide stronger third-party comparability than Juror's own benchmark alone.
Reference: https://github.com/withmartian/code-review-benchmark/blob/main/offline/README.md
Acceptance criteria
- Reproduce the benchmark locally with a version-pinned Juror configuration.
- Record raw Juror outputs, model versions, preset, cost, latency, and run failures.
- Document any adaptation required to fit the benchmark protocol.
- Publish reproducible results without changing the benchmark's evaluation rules.
- Submit Juror upstream if the maintainers accept new reviewers.
- Link the external result from Juror's benchmarking documentation, including limitations and date.
Parent: #39
Context
The open Martian code-review benchmark evaluates major reviewers on 50 PRs and documents how to add another tool. An external corpus would provide stronger third-party comparability than Juror's own benchmark alone.
Reference: https://github.com/withmartian/code-review-benchmark/blob/main/offline/README.md
Acceptance criteria