Conversation
…C API Restarting a worker recompiled every model, which is the bulk of production downtime after a restart. XLA can serialize a compiled program and load it back in tens of milliseconds, but Reactant binds neither call. The CUDA build of Reactant_jll is itself a PJRT GPU plugin and exports the PJRT_Api function table (GetPjrtApi), and Reactant ships generated Julia bindings for that table without including them. Driving the table directly from Julia therefore needs no library change, so this adds a second backend behind the existing AbstractBackend protocol rather than waiting on an upstream rebuild. PJRTCAPIBackend (runtime/pjrt_capi.jl, runtime/pjrt_capi_backend.jl) owns client, buffers, compile and execute through the table. It compiles the same program with the same options as the Reactant backend and produces byte-identical outputs (verified on four bundles), and it is selected automatically on CUDA workers by runtime.engine=auto once the library's table size and API version match the bindings. That gate matters: Reactant 0.2.270 shipped bindings that lagged its own JLL, and a silent layout mismatch would corrupt memory rather than fail. CPU keeps the Reactant backend because the JLL exports no CPU table. The cache lives under each bundle's .cache/ directory. Programs are partitioned by the Reactant_jll build, CUDA runtime, device kind and compute capability, because XLA validates only the compute capability on load: an older build was measured loading a newer build's program without complaint. Entries are named by the MLIR source's content hash, and .cache/mlir_hashes.json records every module's hash so a changed module has its programs dropped while a weights-only update keeps them. The watcher fingerprints only bundle files, so cache writes never trigger a reload. The memory probe now resets the allocator high-water mark before each model. Autotuning scratch during compile had been setting the monotone peak far above any model's real run scratch (698 MB against 30 MB on resnet50), which inflated the weight budget until an operator restarted with a warm autotune cache. The C API exposes the reset directly; the Reactant backend reaches the same method through an exported symbol, so production on the current JLL benefits too. A model loaded by the watcher is probed and the budget re-resolved in place. Compat is relaxed to the current majors (gRPCClient "1", gRPCServer "0") and the Reactant floor raised to 0.2.285 / Reactant_jll 0.0.407, the first pair with matching bindings and thunk-serializing XLA. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0164kicU5aD1XRpkQcqZPSog
Collaborator
Author
|
Closing this in favor of trying to push changes through upstream (and also as part of our internal bazel build until then) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #14 and removes the need for #60
The plan is to get this all working with our own bindings first, then let that shape an upstream PR to add the bindings we need to Reactant so we don't need to maintain our own backend / worry so much about Reactant jll compatibility.