Skip to content

UnsafeAtomics 0.4: generate atomic instructions with llvmcall - #30

Open
maleadt wants to merge 11 commits into
tb/scopesfrom
tb/generator
Open

maleadt wants to merge 11 commits into
tb/scopesfrom
tb/generator

Conversation

@maleadt

@maleadt maleadt commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Stack: #28 → #25 → #27 → this PR.

Motivation

UnsafeAtomics is where Julia's GPU back-ends (Metal.jl, SPIRVIntrinsics, AMDGPU.jl, and eventually CUDA.jl) and Atomix/KernelAbstractions get their atomics from. The intended contract is simple: every call emits exactly one LLVM atomic instruction with the requested ordering and syncscope, and GPUCompiler lowers that instruction for each target. The implementation fell short of that in ways each back-end worked around separately:

  • There were two implementations: hand-written llvmcall strings for Ptr, and an LLVM.jl-based extension for LLVMPtr. They had drifted apart in which types and operations they supported, how they fell back, and which bugs they had. Most of the fixes in Fix scoped atomics, CAS fallbacks, and GPU fences for 0.3.3 #25 applied to only one of them.
  • There was no way to request volatile, weak or specifically aligned atomics, several LLVM operations weren't exposed, and types like Bool, BFloat16 and pointers always went through a compare-and-swap loop.

Approach

Both pointer types now share one implementation that generates the instruction as an llvmcall string. Using llvmcall rather than LLVM.jl is what makes a single implementation possible: the Ptr path is used by CPU packages that shouldn't need to depend on LLVM.jl. It also decouples UnsafeAtomics from LLVM.jl's release cycle and IR builder API, which lags behind new LLVM atomic operations. The price is handling textual IR differences across Julia versions (typed vs. opaque pointers), which the generator takes care of and the tests cover.

Orderings and the canonical scopes can also be passed as Symbols. Like with Julia's own atomic intrinsics (and in MSL), orderings, scopes and the keywords have to be compile-time constants; anything else is a dynamic call. Operations without a native instruction on the current LLVM version fall back to a compare-and-swap loop, so every operation works for every supported type.

CPU users should see no difference: system-scope Ptr operations still use Core.Intrinsics where possible.

Breaking changes

  • LLVMPtr operations default to the device scope instead of system, as in CUDA, HIP, SYCL and Metal. On CPUs the two are equivalent.
  • Floating-point max and min now follow Julia's semantics (NaN propagates). The IEEE maxNum/minNum behavior is available as fmax!/fmin!; SPIRVIntrinsics should switch to those.
  • The LLVM.jl extension is removed; LLVMPtr support is always available.

Also new: volatile, align and weak keywords; inc_wrap!, dec_wrap!, sub_cond! and sub_sat!; default_scope and failure_order; and support for 8-bit, Bool, Float16, BFloat16, Int128 and pointer values. The README describes the full contract.

Known limitations

Keyword arguments only stay constant through a wrapper that Julia inlines. Mark such wrappers @inline or Base.@constprop :aggressive, or call the Val-based Internal.llvm_* primitives, as back-ends that forward arguments through their own functions should.

Testing

Tested on Julia 1.10 to 1.13 on aarch64 macOS (an earlier version of this PR also on 1.14-DEV and x86_64 Linux). The tests now check the generated IR across instructions, types, address spaces, scopes and flags, not just the results. Metal.jl's atomics and KernelAbstractions tests pass against this branch (with Metal.jl#977). Atomix works unchanged, but its compat bound needs to include 0.4.

@maleadt
maleadt added this pull request to stack #29 September 25, 2026 14:59
@maleadt
maleadt marked this pull request as ready for review September 25, 2026 15:11
Generate each LLVM atomic instruction (load, store, atomicrmw, cmpxchg
and fence) as an llvmcall of IR built for the pointer type, value type,
ordering, syncscope, volatility, alignment and metadata, all passed as
type parameters. This covers Ptr and LLVMPtr in any address space, on
Julia's typed (before 1.12) and opaque llvmcall pointers, without
depending on LLVM.jl. cmpxchg returns its result directly instead of
through memory, and the metadata hook can attach `!mmra` tags.

Nothing uses the generator yet; the next commits move the API onto it.
The LLVMPtr methods lived in a package extension on LLVM.jl, so they
didn't exist unless LLVM.jl happened to be loaded: calls then failed
with a MethodError, and inferred as Union{}. The extension also built
IR with an LLVM context per specialisation, and returned the success
flag of `cas!` through memory.

Implement them with the generator instead, and drop the extension and
the LLVM.jl dependency. Operations follow one table: `nand` and the
Bool, BFloat16 and pointer cases that have an atomicrmw instruction no
longer take the compare-and-swap loop. The named RMW functions only
return the old value, so they don't compute `op(old, x)` in Julia.
On LLVMPtr, `max` and `min` on floats used `atomicrmw fmax`/`fmin`,
which implement IEEE maxNum/minNum and so ignore NaN operands, while
`modify!` returned `old => max(old, x)` computed with Julia's `max`,
which propagates them: the result didn't match memory. On Ptr they
were already a compare-and-swap loop with Julia's semantics.

Use `atomicrmw fmaximum`/`fminimum`, which match Julia's `max` and
`min`, where LLVM has them (LLVM 21, Julia 1.14), and the
compare-and-swap loop before. Callers that want the hardware's maxNum
and minNum get them explicitly, with the `UnsafeAtomics.fmax`/`fmin`
operations and `fmax!`/`fmin!`.
Ptr atomics were generated per value type, ordering and scope, and
everything else went through fallbacks that bitcast to an unsigned
integer and dispatched again. Implement them with the generator, like
LLVMPtr, which removes those fallbacks and the method tables.

In the system scope, keep using Julia's intrinsics where they can
express the operation (not for weak or volatile operations, other
alignments, or a failure ordering stronger than the success
ordering), as before for the integer and floating-point types.
Arguments other than orderings and scopes are now a MethodError.
Orderings can now also be passed as a Symbol, with LLVM's or Julia's
names, and the canonical scopes as `:singlethread`, `:subgroup`,
`:workgroup`, `:device` or `:system`. Like Julia's atomic intrinsics,
orderings and scopes have to be constants: they are turned into `Val`s
for the generator, which folds for constants and is a dynamic call
otherwise. Invalid orderings throw a ConcurrencyViolationError, also
for the system-scope fence, which never passes them to Julia's
intrinsic, and invalid scopes an ArgumentError.

The methods are now only defined for Ptr and LLVMPtr.
There was no way to emit a volatile atomic, a weak cmpxchg, or an
alignment other than the size of the value. Add `volatile` and `align`
keywords to all operations on pointers, and `weak` to `cas!`. The
compare-and-swap loop passes them on, and uses a weak cmpxchg itself.

Like orderings, the keywords have to be constants to fold into the
instruction. They only stay constants through a function that forwards
them if Julia inlines that function; back-ends that can't rely on that
can call the `Val`-based primitives directly.
LLVM has atomicrmw operations that GPU atomics libraries need, like
CUDA's `atomicInc`/`atomicDec`: `uinc_wrap`, `udec_wrap`, `usub_cond`
and `usub_sat`. Add them as `inc_wrap!`, `dec_wrap!`, `sub_cond!` and
`sub_sat!`, with `UnsafeAtomics.inc_wrap` etc. as the operations for
`modify!`, which compute the same results in Julia.

They are only used from LLVM 22 (Julia 1.14): earlier versions parse
them, but the AArch64 back-end can't compile them. Before that,
`modify!` uses the compare-and-swap loop.
Without a scope, atomics used the system scope, also on GPU memory,
which is accessed through LLVMPtr. That is stronger and slower than
needed (AMDGPU.jl overrides Atomix to use its agent scope), and not
supported everywhere (e.g. CUDA on Pascal with Windows). Default to
the device scope there instead, like CUDA C, HIP, SYCL and MSL do; on
CPUs, every scope but `singlethread` is the system scope. Ptr keeps
the system scope.

Expose the choice as `UnsafeAtomics.default_scope(ptr)`, and make
`failure_order`, which `cas!` uses for its default failure ordering,
public too.
The README only had a title and there were no docstrings. Document
what the package guarantees: one LLVM atomic instruction per call,
which orderings are valid for which operation, what the syncscopes and
the default scope mean, which operations are native and when the
compare-and-swap loop is used, and the Val-based primitives that
back-ends can rely on. List what changed from 0.3.
The tests ran on a single thread. Check that concurrent atomicrmw
instructions and compare-and-swap loops from several threads don't
lose updates, in a process with four threads. Also test the public
functions on a type without fields.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant