ZKGadgetEval variance run: no build-vs-LSP gap; range-check-8 is the one hard cell - #13
Merged
Conversation
…he one hard cell 102 sessions (17 gadgets x 2 modes x 3 samples, budget 5, claude-sonnet-5, 256 min): build 44/45, lsp 45/45, negatives refused 12/12, zero alarms. The first run's 'build beats LSP' headline did not replicate — LSP went 45/45 (including range-check-8, its only prior failure) and the aggregate wall-time gap reversed; mean iterations are identical at 1.07. Verdict recorded in RESULTS.md: no measurable quality difference between feedback modes on this suite; the run-1 claim is kept with a superseded note as a worked example of why single-sample comparisons are not citable. Real signal: range-check-8 is the sole failure in BOTH runs (different modes each time) and the only cell with pass@1 < 1 (0.67, pass@2 = 1.0, Wilson 95% [0.21, 0.94]) — the suite's one frontier obligation short of lookups.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
3-sample variance run (102 sessions, 17 gadgets × 2 modes × 3 samples, budget 5, claude-sonnet-5, 256 min):
Negatives refused 12/12, zero soundness alarms.
Two findings, both recorded in
benchmark/RESULTS.md:Combined across both runs: 118/120 positive sessions proved, 16/16 false specs refused. Raw JSON committed as
benchmark/variance-run-2026-07-17.json; README pointer updated.