Conversation
maleadt
added this pull request to stack #29
September 25, 2026 14:59
maleadt
marked this pull request as ready for review
September 25, 2026 15:11
Generate each LLVM atomic instruction (load, store, atomicrmw, cmpxchg and fence) as an llvmcall of IR built for the pointer type, value type, ordering, syncscope, volatility, alignment and metadata, all passed as type parameters. This covers Ptr and LLVMPtr in any address space, on Julia's typed (before 1.12) and opaque llvmcall pointers, without depending on LLVM.jl. cmpxchg returns its result directly instead of through memory, and the metadata hook can attach `!mmra` tags. Nothing uses the generator yet; the next commits move the API onto it.
The LLVMPtr methods lived in a package extension on LLVM.jl, so they
didn't exist unless LLVM.jl happened to be loaded: calls then failed
with a MethodError, and inferred as Union{}. The extension also built
IR with an LLVM context per specialisation, and returned the success
flag of `cas!` through memory.
Implement them with the generator instead, and drop the extension and
the LLVM.jl dependency. Operations follow one table: `nand` and the
Bool, BFloat16 and pointer cases that have an atomicrmw instruction no
longer take the compare-and-swap loop. The named RMW functions only
return the old value, so they don't compute `op(old, x)` in Julia.
On LLVMPtr, `max` and `min` on floats used `atomicrmw fmax`/`fmin`, which implement IEEE maxNum/minNum and so ignore NaN operands, while `modify!` returned `old => max(old, x)` computed with Julia's `max`, which propagates them: the result didn't match memory. On Ptr they were already a compare-and-swap loop with Julia's semantics. Use `atomicrmw fmaximum`/`fminimum`, which match Julia's `max` and `min`, where LLVM has them (LLVM 21, Julia 1.14), and the compare-and-swap loop before. Callers that want the hardware's maxNum and minNum get them explicitly, with the `UnsafeAtomics.fmax`/`fmin` operations and `fmax!`/`fmin!`.
Ptr atomics were generated per value type, ordering and scope, and everything else went through fallbacks that bitcast to an unsigned integer and dispatched again. Implement them with the generator, like LLVMPtr, which removes those fallbacks and the method tables. In the system scope, keep using Julia's intrinsics where they can express the operation (not for weak or volatile operations, other alignments, or a failure ordering stronger than the success ordering), as before for the integer and floating-point types. Arguments other than orderings and scopes are now a MethodError.
maleadt
force-pushed
the
tb/generator
branch
from
September 25, 2026 15:12
c21de16 to
b7324a8
Compare
llvmcall
maleadt
force-pushed
the
tb/generator
branch
from
September 28, 2026 12:05
b7324a8 to
cd11c71
Compare
Orderings can now also be passed as a Symbol, with LLVM's or Julia's names, and the canonical scopes as `:singlethread`, `:subgroup`, `:workgroup`, `:device` or `:system`. Like Julia's atomic intrinsics, orderings and scopes have to be constants: they are turned into `Val`s for the generator, which folds for constants and is a dynamic call otherwise. Invalid orderings throw a ConcurrencyViolationError, also for the system-scope fence, which never passes them to Julia's intrinsic, and invalid scopes an ArgumentError. The methods are now only defined for Ptr and LLVMPtr.
maleadt
force-pushed
the
tb/generator
branch
from
September 28, 2026 13:17
cd11c71 to
b11578c
Compare
There was no way to emit a volatile atomic, a weak cmpxchg, or an alignment other than the size of the value. Add `volatile` and `align` keywords to all operations on pointers, and `weak` to `cas!`. The compare-and-swap loop passes them on, and uses a weak cmpxchg itself. Like orderings, the keywords have to be constants to fold into the instruction. They only stay constants through a function that forwards them if Julia inlines that function; back-ends that can't rely on that can call the `Val`-based primitives directly.
LLVM has atomicrmw operations that GPU atomics libraries need, like CUDA's `atomicInc`/`atomicDec`: `uinc_wrap`, `udec_wrap`, `usub_cond` and `usub_sat`. Add them as `inc_wrap!`, `dec_wrap!`, `sub_cond!` and `sub_sat!`, with `UnsafeAtomics.inc_wrap` etc. as the operations for `modify!`, which compute the same results in Julia. They are only used from LLVM 22 (Julia 1.14): earlier versions parse them, but the AArch64 back-end can't compile them. Before that, `modify!` uses the compare-and-swap loop.
Without a scope, atomics used the system scope, also on GPU memory, which is accessed through LLVMPtr. That is stronger and slower than needed (AMDGPU.jl overrides Atomix to use its agent scope), and not supported everywhere (e.g. CUDA on Pascal with Windows). Default to the device scope there instead, like CUDA C, HIP, SYCL and MSL do; on CPUs, every scope but `singlethread` is the system scope. Ptr keeps the system scope. Expose the choice as `UnsafeAtomics.default_scope(ptr)`, and make `failure_order`, which `cas!` uses for its default failure ordering, public too.
The README only had a title and there were no docstrings. Document what the package guarantees: one LLVM atomic instruction per call, which orderings are valid for which operation, what the syncscopes and the default scope mean, which operations are native and when the compare-and-swap loop is used, and the Val-based primitives that back-ends can rely on. List what changed from 0.3.
The tests ran on a single thread. Check that concurrent atomicrmw instructions and compare-and-swap loops from several threads don't lose updates, in a process with four threads. Also test the public functions on a type without fields.
maleadt
force-pushed
the
tb/generator
branch
from
September 28, 2026 13:25
b11578c to
f49b398
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: #28 → #25 → #27 → this PR.
Motivation
UnsafeAtomics is where Julia's GPU back-ends (Metal.jl, SPIRVIntrinsics, AMDGPU.jl, and eventually CUDA.jl) and Atomix/KernelAbstractions get their atomics from. The intended contract is simple: every call emits exactly one LLVM atomic instruction with the requested ordering and syncscope, and GPUCompiler lowers that instruction for each target. The implementation fell short of that in ways each back-end worked around separately:
llvmcallstrings forPtr, and an LLVM.jl-based extension forLLVMPtr. They had drifted apart in which types and operations they supported, how they fell back, and which bugs they had. Most of the fixes in Fix scoped atomics, CAS fallbacks, and GPU fences for 0.3.3 #25 applied to only one of them.Bool,BFloat16and pointers always went through a compare-and-swap loop.Approach
Both pointer types now share one implementation that generates the instruction as an
llvmcallstring. Usingllvmcallrather than LLVM.jl is what makes a single implementation possible: thePtrpath is used by CPU packages that shouldn't need to depend on LLVM.jl. It also decouples UnsafeAtomics from LLVM.jl's release cycle and IR builder API, which lags behind new LLVM atomic operations. The price is handling textual IR differences across Julia versions (typed vs. opaque pointers), which the generator takes care of and the tests cover.Orderings and the canonical scopes can also be passed as
Symbols. Like with Julia's own atomic intrinsics (and in MSL), orderings, scopes and the keywords have to be compile-time constants; anything else is a dynamic call. Operations without a native instruction on the current LLVM version fall back to a compare-and-swap loop, so every operation works for every supported type.CPU users should see no difference: system-scope
Ptroperations still useCore.Intrinsicswhere possible.Breaking changes
LLVMPtroperations default to thedevicescope instead ofsystem, as in CUDA, HIP, SYCL and Metal. On CPUs the two are equivalent.maxandminnow follow Julia's semantics (NaN propagates). The IEEE maxNum/minNum behavior is available asfmax!/fmin!; SPIRVIntrinsics should switch to those.LLVMPtrsupport is always available.Also new:
volatile,alignandweakkeywords;inc_wrap!,dec_wrap!,sub_cond!andsub_sat!;default_scopeandfailure_order; and support for 8-bit,Bool,Float16,BFloat16,Int128and pointer values. The README describes the full contract.Known limitations
Keyword arguments only stay constant through a wrapper that Julia inlines. Mark such wrappers
@inlineorBase.@constprop :aggressive, or call theVal-basedInternal.llvm_*primitives, as back-ends that forward arguments through their own functions should.Testing
Tested on Julia 1.10 to 1.13 on aarch64 macOS (an earlier version of this PR also on 1.14-DEV and x86_64 Linux). The tests now check the generated IR across instructions, types, address spaces, scopes and flags, not just the results. Metal.jl's atomics and KernelAbstractions tests pass against this branch (with Metal.jl#977). Atomix works unchanged, but its compat bound needs to include 0.4.