Skip to content

Benchmarks - #38

Merged
adrn merged 20 commits into
mainfrom
benchmarks
Sep 14, 2026
Merged

adrn merged 20 commits into
mainfrom
benchmarks

Conversation

@adrn

@adrn adrn commented Aug 31, 2026 •

Copy link
Copy Markdown
Owner

TODO:

@adrn
adrn force-pushed the benchmarks branch 3 times, most recently from ad2903b to 6cfcb97 Compare September 14, 2026 16:13
adrn and others added 20 commits September 14, 2026 15:07
The jaxtyping/beartype hooks in the root conftest are session-wide, so any
pytest run under this rootdir gets an instrumented harv.models. That is what
the test suite wants, but it distorts the benchmark harness's first-call
number: beartype decorates Python-level functions, and log_prob is called
inside jax.vmap under eqx.filter_jit, so the checks run once at trace time and
land entirely in compile time.

HARV_NO_TYPECHECK=1 skips them. Unset everywhere else; behavior is unchanged
for every existing caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nothing in the repo measured the sampler's throughput, so two things the spec
asserts were never checked: that batch_size=100_000 is right on CPU and
batch_size=n_prior_samples on GPU, and what first-call compile actually costs.

benchmarks/ holds a star (one-axis-at-a-time) grid over the six
model/parameterization pairs, n_obs 8..256, prior libraries 1e4..1e7,
batch_size, and in-memory vs streamed HDF5 caches. 66 measurements after
deduplicating the baseline points the curves share.

Every cell uses top_k, which gives a static output shape so the vmap over the
conditional-linear solve is traced once. Under plain rejection the accepted
count varies with the data, and the timings would partly measure recompilation
rather than the axis under test. Timed calls are wrapped in block_until_ready:
JAX dispatch is async, so without it the GPU numbers would measure enqueue
speed. float64 is pinned before harv is imported, since float32 changes the
sampler's arithmetic and would not describe how anyone runs harv.

These never run on CI, three ways over: benchmarks/ is outside testpaths, so
CI's bare pytest does not see it; pytest-benchmark lives in a new `bench`
group that CI's --group test does not install; and --bench is required, so a
stray `pytest benchmarks/` cannot start a multi-hour run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pytest-benchmark prints a table but persists nothing, so a finished run would
otherwise leave nothing behind. report.py merges every JSON in
benchmarks/results/ by device and renders the docs page: per-curve tables, a
fitted log-log slope for each numeric axis, a CPU-vs-GPU speedup column when
both are present, a first-call compile table, and log-log figures.

Curve membership comes from grid.curve_definitions rather than from the data,
because curves share baseline cells and those are measured only once. The run
records its grid mode so the mapping is exact instead of inferred.

Results from a --bench-smoke run get a prominent warning banner: they exist to
prove the pipeline works, and must not be mistaken for a description of harv's
performance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
running-benchmarks.md is the developer page: the exact invocations, what each
flag and environment variable is for, the grid, the cost, and why none of it
runs on CI. benchmarks.md is the generated results page, committed because
ReadTheDocs has no GPU and must never try to rebuild it.

It currently holds smoke-test numbers behind a warning banner. The toctree
needs the file to exist, and a real workstation run replaces it.

Spec updates: HARV_NO_TYPECHECK in the beartype section; the batch_size
guidance now points at the measurement that checks it, and notes that
batch_size is the only device knob harv has; and the TTFX section, since
first-call compile time is now measured even though it is not yet optimized.

Also fixes report.py emitting a doubled trailing newline, which the
end-of-file-fixer hook would reject on every regeneration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
benchmarks/results/ held only a gitignored smoke file, so git did not track the
directory and it did not exist on a fresh checkout. --benchmark-json then blew
up with FileNotFoundError.

The .gitkeep fixes that, but the real problem is worse: pytest-benchmark writes
the JSON from pytest_sessionfinish, after every benchmark has already run. A
missing directory or a typo'd path therefore does not fail fast, it discards the
whole session -- hours, on a real run. So pytest_configure now creates the parent
directory and write-probes it before collection, turning a late data-loss bug
into an immediate UsageError.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Applied uniformly so cells stay comparable, and required for EcoswEsinwRV:
its default prior puts ~21% of draws outside the unit disk (e >= 1), where the
period-dependent K prior is NaN, and a single NaN leaves the recorded evidence
metadata NaN for the whole cell.

The prior itself is documented in docs/sharp-bits.md and tracked in the spec's
"Known bugs" section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Following the documented steps hit "results mix grid modes ['smoke', 'star']".
The TL;DR told you to write the smoke run into benchmarks/results/, which is the
directory report.py globs, and the warning to clear stale files sat below the
code block. smoke.json is also gitignored, so it lingered invisibly. The smoke
step additionally overwrote the committed docs/benchmarks.md.

Fixed at both ends. report.py now skips smoke-mode files by default
(--include-smoke opts back in) and, when modes really are mixed, names each file
and its mode instead of just asserting they are. The docs write the smoke run to
benchmarks/smoke/ -- outside the glob, gitignored -- and render it to its own
file, with the rm of stale results now inside the code block where it cannot be
skipped past.

Also caught a related silent-data-loss case. Nothing about a filename makes a run
a GPU run: report.py labels devices from jax.devices()[0], so when CUDA-enabled
jaxlib is missing and jax falls back to CPU, gpu.json and cpu.json collide on one
label and one overwrites the other. That is now a named error, and the docs open
with a device check so the fallback is caught before the hours are spent, not
after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The page told you to check for a CudaDevice and mentioned jax[cuda12] in prose,
but never said how to install it -- and the install step it did give,
`uv sync --group bench`, can never produce one. jax ships CPU-only and the bench
group adds no CUDA jaxlib, so following the instructions exactly guaranteed a
CPU-only environment and a "gpu" run that was really a second CPU run.

bench-cuda is jax[cuda12] behind a sys_platform == 'linux' marker: the
jax-cuda12-plugin wheels are Linux-only, and without the marker `uv lock` cannot
resolve on macOS or Windows. With it the group is a no-op off Linux, so the
TL;DR can carry one install command that is correct on every machine. Kept out of
`bench` so a CPU-only run does not pull several GB of NVIDIA wheels, and out of
CI, which installs only the `test` group.

Also documents why to install via the group rather than `uv pip install`: uv sync
removes anything absent from the lockfile, so the bare pip route quietly loses
GPU support on the next sync. Plus a short checklist for the case where it still
reports CpuDevice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two gaps the page left. It never said that one install serves both runs -- the
CPU/GPU split is JAX_PLATFORMS, not a second environment -- so installing CUDA
looked like it might foreclose the CPU half. And bench-cuda pinned cuda12 while
jax 0.8 publishes cuda12, cuda12-local, cuda13 and cuda13-local; cuda13 is the
current line, with cuda12 documented as the fallback for older drivers.

Also adds --bench-expect={cpu,gpu}, checked against jax.default_backend() before
collection. A filename says nothing about which device ran -- report.py labels
devices from jax.devices()[0] -- so a GPU run that silently fell back to CPU used
to be detectable only afterwards, from a device-label collision, after the hours
were already spent. Now it aborts in seconds. Both real runs in the TL;DR pass it.

default_backend() rather than Device.platform: that has spelled CUDA both "gpu"
and "cuda" across versions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The page had tables and no conclusions. It now opens with a computed findings
section: nine readings, each figure derived from the loaded results so that
regenerating on other hardware updates the numbers instead of contradicting the
prose. Only the interpretation is fixed text.

What the current data says. The GPU advantage grows with library size, 3.1-8.3x
at M=1e4 to 29.6-52.1x at M=1e7, because a small library cannot fill the device;
below ~1e5 a call is fixed cost, 9 ms against the CPU's 41 ms. Throughput at
M=1e7 spans 5.8 to 26.0 M samples/s depending on parameterization. Evidence
statistics agree between devices to 1.4e-10, so the speedup is not a precision
trade.

Two findings contradict what we assumed. batch_size spans only 1.04-1.13x on GPU
against 1.55-1.84x on CPU, so the spec's "set batch_size = n_prior_samples on
GPU" was backwards about which device the knob is for. And streaming from HDF5
costs the GPU 1.35-1.76x against the CPU's 1.05-1.22x -- same I/O, but on GPU
there is no slow compute left to hide it behind.

Two claims are hedged in the text because the data does not support the obvious
reading. The n_obs=256 cliff is measured but its remedy is not: the batch_size
sweep only ran at n_obs=64, so the working-set explanation is labelled a
hypothesis. And every CPU cell is a single process on an idle node, so those are
per-core ceilings, not per-node predictions -- under contention the batch_size
optimum is known to invert.

Comparative headings and prose are suppressed when only one device's results are
loaded, so a CPU-only regeneration cannot assert a comparison it has no evidence
for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
harv is described as a backbone for population science and had no guidance on
population-scale runs. This distils two independent bodies of evidence into one
page: the CPU/GPU benchmark grid, and a production campaign over a simulated Gaia
DR4 catalog whose notes lived in an untracked scratch file.

The page is explicit about which source each number comes from and never blends
them -- different hardware, different data, and a reader sizing an allocation
needs to know which is which.

"Making it fast" gives the cost model, budgeting tables in device-hours, and four
hardware paths as tabs: one workstation, one GPU, many CPU nodes, many GPU nodes.
The advice genuinely differs between them. batch_size is nearly irrelevant on GPU
and inverts under MPI contention. HDF5 streaming is cheap on CPU and expensive on
GPU. And harv shards nothing itself, so scaling out is one process per shard of
sources -- worth saying plainly, since the alternative assumption is expensive.

"Making it right" carries the failure modes that only appear at scale, each with
its evidence table: weighted output that must be renormalized and stored as
log-weights, ESS as a resolution diagnostic that runs backwards, uncertainty
calibration beating a fitted jitter on every axis, the amplitude prior silently
setting the detection threshold, and completeness binned on detectable rather
than recorded SNR. Several of these cost that campaign a week each, and none
announces itself.

Also corrects the spec's batch_size guidance, which the benchmark contradicts.
HARV_AT_SCALE.md is removed; it was untracked, so this page is now the only
record of its content -- which is why the evidence tables came across rather than
just the conclusions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
series_of labels a row by every coordinate that differs from the baseline, but
the library-size curve *derives* one of them: batch_size is min(baseline, M),
because a batch cannot exceed the library. Its M=1e4 cell therefore carries
batch_size=1e4, was read as a different series, and split every row in two.

The visible damage was worse than a cosmetic split. The M=1e4 point vanished
from the plot entirely -- a one-point series is skipped -- so the figure started
at 1e5, and every slope was fitted on three points instead of four. That
mattered: CPU slopes read 0.95-0.99, and over the full range they are 0.89-0.94.

Now a coordinate the curve derives from its own axis no longer distinguishes a
series. The plot spans 1e4 to 1e7, where the GPU's fixed-cost regime is visible
as a flat segment below 1e5 -- which is the thing the narrative claims, now
actually shown.

Found by looking at the rendered figure. The findings section read the data
directly and so was unaffected, which is why the tables and the prose disagreed
without either being obviously wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@adrn
adrn enabled auto-merge September 14, 2026 19:29
@adrn
adrn merged commit b37be98 into main Sep 14, 2026
13 checks passed
@adrn
adrn deleted the benchmarks branch September 14, 2026 19:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant