Skip to content

PJRT Executable Cache - #88

Closed
csvance wants to merge 2 commits into
mainfrom
feature/pjrt-capi-executable-cache
Closed

csvance wants to merge 2 commits into
mainfrom
feature/pjrt-capi-executable-cache

Conversation

@csvance

@csvance csvance commented Sep 12, 2026 •

Copy link
Copy Markdown
Collaborator

Implements #14 and removes the need for #60

  • Add support for a full PJRT executable cache keyed on Reactant jll version + device
  • This comes at the cost of adding our own C PJRT backend, but it's needed due to time constraints with an upcoming deployment which cannot wait on upstreaming the additional C++ bindings to Reactant.
  • Reset high watermark during memory measurement to fix problem of estimation being inaccurate on cold start that compiled models before measuring memory usage

The plan is to get this all working with our own bindings first, then let that shape an upstream PR to add the bindings we need to Reactant so we don't need to maintain our own backend / worry so much about Reactant jll compatibility.

csvance and others added 2 commits September 12, 2026 16:10
…C API

Restarting a worker recompiled every model, which is the bulk of production downtime after a
restart. XLA can serialize a compiled program and load it back in tens of milliseconds, but
Reactant binds neither call. The CUDA build of Reactant_jll is itself a PJRT GPU plugin and
exports the PJRT_Api function table (GetPjrtApi), and Reactant ships generated Julia bindings
for that table without including them. Driving the table directly from Julia therefore needs no
library change, so this adds a second backend behind the existing AbstractBackend protocol
rather than waiting on an upstream rebuild.

PJRTCAPIBackend (runtime/pjrt_capi.jl, runtime/pjrt_capi_backend.jl) owns client, buffers,
compile and execute through the table. It compiles the same program with the same options as
the Reactant backend and produces byte-identical outputs (verified on four bundles), and it is
selected automatically on CUDA workers by runtime.engine=auto once the library's table size and
API version match the bindings. That gate matters: Reactant 0.2.270 shipped bindings that
lagged its own JLL, and a silent layout mismatch would corrupt memory rather than fail. CPU
keeps the Reactant backend because the JLL exports no CPU table.

The cache lives under each bundle's .cache/ directory. Programs are partitioned by the
Reactant_jll build, CUDA runtime, device kind and compute capability, because XLA validates
only the compute capability on load: an older build was measured loading a newer build's
program without complaint. Entries are named by the MLIR source's content hash, and
.cache/mlir_hashes.json records every module's hash so a changed module has its programs
dropped while a weights-only update keeps them. The watcher fingerprints only bundle files, so
cache writes never trigger a reload.

The memory probe now resets the allocator high-water mark before each model. Autotuning
scratch during compile had been setting the monotone peak far above any model's real run
scratch (698 MB against 30 MB on resnet50), which inflated the weight budget until an operator
restarted with a warm autotune cache. The C API exposes the reset directly; the Reactant
backend reaches the same method through an exported symbol, so production on the current JLL
benefits too. A model loaded by the watcher is probed and the budget re-resolved in place.

Compat is relaxed to the current majors (gRPCClient "1", gRPCServer "0") and the Reactant
floor raised to 0.2.285 / Reactant_jll 0.0.407, the first pair with matching bindings and
thunk-serializing XLA.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0164kicU5aD1XRpkQcqZPSog
@csvance

csvance commented Sep 12, 2026

Copy link
Copy Markdown
Collaborator Author

Closing this in favor of trying to push changes through upstream (and also as part of our internal bazel build until then)

@csvance csvance closed this Sep 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant