Benchmarks - #38
Merged
Merged
Benchmarks#38
Conversation
adrn
force-pushed
the
benchmarks
branch
3 times, most recently
from
September 14, 2026 16:13
ad2903b to
6cfcb97
Compare
The jaxtyping/beartype hooks in the root conftest are session-wide, so any pytest run under this rootdir gets an instrumented harv.models. That is what the test suite wants, but it distorts the benchmark harness's first-call number: beartype decorates Python-level functions, and log_prob is called inside jax.vmap under eqx.filter_jit, so the checks run once at trace time and land entirely in compile time. HARV_NO_TYPECHECK=1 skips them. Unset everywhere else; behavior is unchanged for every existing caller. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nothing in the repo measured the sampler's throughput, so two things the spec asserts were never checked: that batch_size=100_000 is right on CPU and batch_size=n_prior_samples on GPU, and what first-call compile actually costs. benchmarks/ holds a star (one-axis-at-a-time) grid over the six model/parameterization pairs, n_obs 8..256, prior libraries 1e4..1e7, batch_size, and in-memory vs streamed HDF5 caches. 66 measurements after deduplicating the baseline points the curves share. Every cell uses top_k, which gives a static output shape so the vmap over the conditional-linear solve is traced once. Under plain rejection the accepted count varies with the data, and the timings would partly measure recompilation rather than the axis under test. Timed calls are wrapped in block_until_ready: JAX dispatch is async, so without it the GPU numbers would measure enqueue speed. float64 is pinned before harv is imported, since float32 changes the sampler's arithmetic and would not describe how anyone runs harv. These never run on CI, three ways over: benchmarks/ is outside testpaths, so CI's bare pytest does not see it; pytest-benchmark lives in a new `bench` group that CI's --group test does not install; and --bench is required, so a stray `pytest benchmarks/` cannot start a multi-hour run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pytest-benchmark prints a table but persists nothing, so a finished run would otherwise leave nothing behind. report.py merges every JSON in benchmarks/results/ by device and renders the docs page: per-curve tables, a fitted log-log slope for each numeric axis, a CPU-vs-GPU speedup column when both are present, a first-call compile table, and log-log figures. Curve membership comes from grid.curve_definitions rather than from the data, because curves share baseline cells and those are measured only once. The run records its grid mode so the mapping is exact instead of inferred. Results from a --bench-smoke run get a prominent warning banner: they exist to prove the pipeline works, and must not be mistaken for a description of harv's performance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
running-benchmarks.md is the developer page: the exact invocations, what each flag and environment variable is for, the grid, the cost, and why none of it runs on CI. benchmarks.md is the generated results page, committed because ReadTheDocs has no GPU and must never try to rebuild it. It currently holds smoke-test numbers behind a warning banner. The toctree needs the file to exist, and a real workstation run replaces it. Spec updates: HARV_NO_TYPECHECK in the beartype section; the batch_size guidance now points at the measurement that checks it, and notes that batch_size is the only device knob harv has; and the TTFX section, since first-call compile time is now measured even though it is not yet optimized. Also fixes report.py emitting a doubled trailing newline, which the end-of-file-fixer hook would reject on every regeneration. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
benchmarks/results/ held only a gitignored smoke file, so git did not track the directory and it did not exist on a fresh checkout. --benchmark-json then blew up with FileNotFoundError. The .gitkeep fixes that, but the real problem is worse: pytest-benchmark writes the JSON from pytest_sessionfinish, after every benchmark has already run. A missing directory or a typo'd path therefore does not fail fast, it discards the whole session -- hours, on a real run. So pytest_configure now creates the parent directory and write-probes it before collection, turning a late data-loss bug into an immediate UsageError. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Applied uniformly so cells stay comparable, and required for EcoswEsinwRV: its default prior puts ~21% of draws outside the unit disk (e >= 1), where the period-dependent K prior is NaN, and a single NaN leaves the recorded evidence metadata NaN for the whole cell. The prior itself is documented in docs/sharp-bits.md and tracked in the spec's "Known bugs" section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Following the documented steps hit "results mix grid modes ['smoke', 'star']". The TL;DR told you to write the smoke run into benchmarks/results/, which is the directory report.py globs, and the warning to clear stale files sat below the code block. smoke.json is also gitignored, so it lingered invisibly. The smoke step additionally overwrote the committed docs/benchmarks.md. Fixed at both ends. report.py now skips smoke-mode files by default (--include-smoke opts back in) and, when modes really are mixed, names each file and its mode instead of just asserting they are. The docs write the smoke run to benchmarks/smoke/ -- outside the glob, gitignored -- and render it to its own file, with the rm of stale results now inside the code block where it cannot be skipped past. Also caught a related silent-data-loss case. Nothing about a filename makes a run a GPU run: report.py labels devices from jax.devices()[0], so when CUDA-enabled jaxlib is missing and jax falls back to CPU, gpu.json and cpu.json collide on one label and one overwrites the other. That is now a named error, and the docs open with a device check so the fallback is caught before the hours are spent, not after. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The page told you to check for a CudaDevice and mentioned jax[cuda12] in prose, but never said how to install it -- and the install step it did give, `uv sync --group bench`, can never produce one. jax ships CPU-only and the bench group adds no CUDA jaxlib, so following the instructions exactly guaranteed a CPU-only environment and a "gpu" run that was really a second CPU run. bench-cuda is jax[cuda12] behind a sys_platform == 'linux' marker: the jax-cuda12-plugin wheels are Linux-only, and without the marker `uv lock` cannot resolve on macOS or Windows. With it the group is a no-op off Linux, so the TL;DR can carry one install command that is correct on every machine. Kept out of `bench` so a CPU-only run does not pull several GB of NVIDIA wheels, and out of CI, which installs only the `test` group. Also documents why to install via the group rather than `uv pip install`: uv sync removes anything absent from the lockfile, so the bare pip route quietly loses GPU support on the next sync. Plus a short checklist for the case where it still reports CpuDevice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two gaps the page left. It never said that one install serves both runs -- the
CPU/GPU split is JAX_PLATFORMS, not a second environment -- so installing CUDA
looked like it might foreclose the CPU half. And bench-cuda pinned cuda12 while
jax 0.8 publishes cuda12, cuda12-local, cuda13 and cuda13-local; cuda13 is the
current line, with cuda12 documented as the fallback for older drivers.
Also adds --bench-expect={cpu,gpu}, checked against jax.default_backend() before
collection. A filename says nothing about which device ran -- report.py labels
devices from jax.devices()[0] -- so a GPU run that silently fell back to CPU used
to be detectable only afterwards, from a device-label collision, after the hours
were already spent. Now it aborts in seconds. Both real runs in the TL;DR pass it.
default_backend() rather than Device.platform: that has spelled CUDA both "gpu"
and "cuda" across versions.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The page had tables and no conclusions. It now opens with a computed findings section: nine readings, each figure derived from the loaded results so that regenerating on other hardware updates the numbers instead of contradicting the prose. Only the interpretation is fixed text. What the current data says. The GPU advantage grows with library size, 3.1-8.3x at M=1e4 to 29.6-52.1x at M=1e7, because a small library cannot fill the device; below ~1e5 a call is fixed cost, 9 ms against the CPU's 41 ms. Throughput at M=1e7 spans 5.8 to 26.0 M samples/s depending on parameterization. Evidence statistics agree between devices to 1.4e-10, so the speedup is not a precision trade. Two findings contradict what we assumed. batch_size spans only 1.04-1.13x on GPU against 1.55-1.84x on CPU, so the spec's "set batch_size = n_prior_samples on GPU" was backwards about which device the knob is for. And streaming from HDF5 costs the GPU 1.35-1.76x against the CPU's 1.05-1.22x -- same I/O, but on GPU there is no slow compute left to hide it behind. Two claims are hedged in the text because the data does not support the obvious reading. The n_obs=256 cliff is measured but its remedy is not: the batch_size sweep only ran at n_obs=64, so the working-set explanation is labelled a hypothesis. And every CPU cell is a single process on an idle node, so those are per-core ceilings, not per-node predictions -- under contention the batch_size optimum is known to invert. Comparative headings and prose are suppressed when only one device's results are loaded, so a CPU-only regeneration cannot assert a comparison it has no evidence for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
harv is described as a backbone for population science and had no guidance on population-scale runs. This distils two independent bodies of evidence into one page: the CPU/GPU benchmark grid, and a production campaign over a simulated Gaia DR4 catalog whose notes lived in an untracked scratch file. The page is explicit about which source each number comes from and never blends them -- different hardware, different data, and a reader sizing an allocation needs to know which is which. "Making it fast" gives the cost model, budgeting tables in device-hours, and four hardware paths as tabs: one workstation, one GPU, many CPU nodes, many GPU nodes. The advice genuinely differs between them. batch_size is nearly irrelevant on GPU and inverts under MPI contention. HDF5 streaming is cheap on CPU and expensive on GPU. And harv shards nothing itself, so scaling out is one process per shard of sources -- worth saying plainly, since the alternative assumption is expensive. "Making it right" carries the failure modes that only appear at scale, each with its evidence table: weighted output that must be renormalized and stored as log-weights, ESS as a resolution diagnostic that runs backwards, uncertainty calibration beating a fitted jitter on every axis, the amplitude prior silently setting the detection threshold, and completeness binned on detectable rather than recorded SNR. Several of these cost that campaign a week each, and none announces itself. Also corrects the spec's batch_size guidance, which the benchmark contradicts. HARV_AT_SCALE.md is removed; it was untracked, so this page is now the only record of its content -- which is why the evidence tables came across rather than just the conclusions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
series_of labels a row by every coordinate that differs from the baseline, but the library-size curve *derives* one of them: batch_size is min(baseline, M), because a batch cannot exceed the library. Its M=1e4 cell therefore carries batch_size=1e4, was read as a different series, and split every row in two. The visible damage was worse than a cosmetic split. The M=1e4 point vanished from the plot entirely -- a one-point series is skipped -- so the figure started at 1e5, and every slope was fitted on three points instead of four. That mattered: CPU slopes read 0.95-0.99, and over the full range they are 0.89-0.94. Now a coordinate the curve derives from its own axis no longer distinguishes a series. The plot spans 1e4 to 1e7, where the GPU's fixed-cost regime is visible as a flat segment below 1e5 -- which is the thing the narrative claims, now actually shown. Found by looking at the rendered figure. The findings section read the data directly and so was unaffected, which is why the tables and the prose disagreed without either being obviously wrong. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TODO: