Skip to content

Bazel Build + Vendor Executable Cache w/ new JLL - #90

Merged
csvance merged 5 commits into
mainfrom
executable-cache-vendored
Sep 28, 2026
Merged

csvance merged 5 commits into
mainfrom
executable-cache-vendored

Conversation

@csvance

@csvance csvance commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

This will close #14, vendoring the API into ReactantServer.jl until it can be merged upstream. Once it's merged upstream, we will switch over to the upstream API and set the compat bounds appropriately.

Also provides Bazel build infrastructure to support building container images for formal releases and for integration into polyglot repo build systems.

csvance and others added 5 commits September 28, 2026 15:36
…actant's bindings

Reactant is gaining executable serialization and allocator-statistics control
(EnzymeAD/Reactant.jl branch feature/executable-serialization-memory-stats:
XLA.serialize_executable, XLA.load_serialized_executable, XLA.clear_memory_stats!,
XLA.compiled_memory_stats). With those, the Reactant backend can do everything the PJRT C
API backend on feature/pjrt-capi-executable-cache was carried for, on CPU as well as CUDA,
without a second binding layer. This branch is the validation of that claim, built off main
so it stands on its own.

- runtime/executable_cache.jl: the per-bundle cache (layout, mlir_hashes.json, invalidation,
  atomic writes, counters), unchanged from the C API branch.
- ReactantBackend: compile_artifact looks the program up under <bundle>/.cache/, loads it
  with a compile-options override on a hit, compiles and serializes it on a miss, drops and
  recompiles an entry that fails to load. The cache key hashes the post-numerics portable
  artifact; because serializing lowers the module to VHLO in place, the module handed to XLA
  is deserialized from those same bytes. clear_memory_stats! and compiled_memory_stats call
  the new Reactant functions; the mangled-symbol hack is gone. All of it is feature-detected,
  so on an older Reactant the worker compiles everything and probes as before.
- Scheduler probe, post-compile peak reset, hot-load re-probe, watcher .cache/ rule,
  Prometheus counters, runtime.executable_cache config and env override, docs: ported from
  the C API branch minus runtime.engine and the C API backend.
- Tests: the filesystem and probe tests from the C API branch, plus a Reactant round trip on
  the CPU client: compile, store, load, run with identical output; corrupt entry dropped and
  recompiled; weights-only change still hits; cache off touches nothing. Skips with a warning
  on a Reactant without the bindings.

Verified with Reactant from the feature branch and a locally built libReactantExtra:
ReactantServerCore and ReactantServer suites pass (755 tests), the round trip included.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0187bPmjLReauAiN6YP2a3cd
The executable cache and the resettable memory probe need the C API from
EnzymeAD/Reactant.jl#3277, which first ships in Reactant_jll 0.0.410 (built
from a9df0ce, two commits after the merge). Its Julia half, #3305, is still
open, and a production release cannot wait for it. So this ports #3305 into
runtime/xla_serialization.jl as `VendoredXLA`: its own functions calling
Reactant_jll.libReactantExtra directly with @CCall, not methods on
Reactant.XLA. That avoids type piracy and cannot clash once #3305 ships, and
it needs only the jll symbols, not the MLIR.API wrappers Reactant 0.2.288
first generates.

With the symbols guaranteed, the feature detection goes:
supports_executable_cache is unconditionally true for the Reactant backend,
and the tests stop skipping.

The floors are the tested versions: Reactant 0.2.287 (the first release on
jll 0.0.410) in ReactantServer and ReactantServerExport, kept in sync, and
a new direct Reactant_jll >= 0.0.410. The jll bound is a floor, because a
plain "0.0.410" would admit only 0.0.410.

To retire the port: raise the Reactant floor to the release carrying #3305,
point `_XS` at Reactant's own XLA module, and delete the file.

Verified at Reactant 0.2.287 / jll 0.0.410: ReactantServerCore 384/384,
ReactantServer 767/767, Gateway 711/711, Node 106/106, Client all pass. On
an A6000: compile 4.3 s vs load 0.10 s, output bit identical, and
clear_memory_stats! resets the peak from 105 MB to the 786 KB in use.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RaRcwKGd7YfoRSjoZDZKSN
docker/Dockerfile resolves at build time, so the same commit can yield a
different image next week. A validated deployment needs the opposite, so
//deploy:image assembles the node image from committed locks with no resolve
anywhere in the build. It has nothing site-specific in it: no private
registry, remote cache, or custom JLL.

Two locks. deploy/Manifest.toml is the production lock for the root
Project.toml, while the root manifest stays gitignored so development and CI
keep resolving fresh. //deploy:manifest_current fails when a compat change
leaves it stale; //deploy:relock moves it. It pins Reactant 0.2.287 /
Reactant_jll 0.0.410, the versions tested. deploy/debs.lock.json pins the
Ubuntu packages the CUDA base lacks: curl for the healthcheck and tini as
PID 1, plus libcurl's libraries. //deploy:relock_debs resolves them with apt
inside the pinned base, so the lock is exactly the delta and never
overwrites a base package. The URLs are snapshot.ubuntu.com ones, because
archive URLs vanish once a version is superseded.

The image starts without precompiling. Julia records a cached package's
sources relative to their depot, so caches built in a temporary tree only
load from /opt if every source lives in a depot. The workspace's own
packages do not, so the image lists the workspace root as a depot as well.
//deploy:compiled_layer then precompiles the four projects the image starts
Julia in (supervisor, worker, gateway, worker healthcheck) against the
unpacked depot and app layers. It uses the official Julia x86_64 CPU target
set, so the caches load on any x86_64 host. Loading Reactant needs
libcuda.so.1, so the build takes the driver stub from the pinned CUDA base
itself; the stub never reaches the image. //deploy:image_check fails if any
entry project would recompile or reject a cache.

Paths match the Dockerfile image: tini runs /usr/local/bin/entrypoint.node.sh,
the healthchecks are in /usr/local/bin, the node file is at
/etc/reactantserver/node.yaml, and 8001/8002 are exposed. The healthcheck
itself cannot be baked in: the OCI image config has no healthcheck field, so
deploy/README.md gives the run flags.

Repository names are prefixed (reactantserver_julia, _cuda_base, _debs) and
the targets public, so a deployment repository can take a bazel_dep on this
module without collisions instead of maintaining its own image.

Verified on an A6000 with four model bundles: a cold start (empty
executable caches) is ready in 78 s against 352 s for the same image
without baked caches, all 6 programs compiled and stored; a warm start is
ready in 57 s with 6/6 cache hits; the healthcheck passes; and podman stop
exits 0 in 1 to 2 s through tini.

Also fixes the .gitignore append: the file lacked a trailing newline, so the
new block had fused onto the PLAN.md rule.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RaRcwKGd7YfoRSjoZDZKSN
Built through a bazel_dep from another module, the image put the whole
workspace one level too deep. Inside a dependency, $(locations) yields
external/<repo>/<path> (../<repo>/<path> with the sibling layout), and
app_layer and stage_workspace copied each file to that execroot path. So
the tree landed at /opt/reactantserver/external/reactant_server+/, where
no entrypoint looks, and the first missing file (docker/*.sh) failed the
build. stage_workspace has the same flaw, which would have handed Pkg a
project with no member packages beside it.

workspace_path strips the repository prefix; files of the root module pass
through unchanged. The app layer is byte-identical built from this
repository or from a consumer, and so are the debs, depot and julia layers.
Only the compiled layer differs, since Julia's package images are not
bit-reproducible; the consumer's image passes //deploy:image_check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RaRcwKGd7YfoRSjoZDZKSN
Reactant 0.2.289 traces every Number a `@trace while` body captures. The
ROI align loop in roi_align_fpn captures the helper k3, which captured the
ROI count K, so K became a TracedRNumber{Int64}, and reshape(v, 1, 1, K)
has no method for a traced dimension. All four detector export testsets
error with that MethodError (74 passed, 4 errored), on main as on this
branch: main last passed CI on 2026-09-25, on 0.2.288, and the export
package's compat ("0.2.264" before, "0.2.287" now) admits 0.2.289 either
way.

K now reaches k3 through Val(K), a type parameter with no fields for the
loop capture to trace, so the reshape dimension stays static and the
exported program is unchanged.

Verified with the export suite (REACTANTSERVER_SKIP_PYTORCH=true, as the
export-lux job runs it): 103/103 on Reactant 0.2.289 / jll 0.0.413, and
103/103 on the 0.2.287 / 0.0.410 floor.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RaRcwKGd7YfoRSjoZDZKSN
@csvance
csvance merged commit 32068b0 into main Sep 28, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Persistent Kernel Cache

1 participant