Cross-platform GPU memory (VRAM) detection for Rust — no vendor SDKs, nothing to install beyond your GPU driver.
| Vendor | Linux | Windows | macOS | Backend |
|---|---|---|---|---|
| NVIDIA | ✅ | ✅ | ✅† | NVML · system_profiler |
| AMD | ✅ | — | ✅† | DRM sysfs · KFD · system_profiler |
| Intel | ✅ | — | ✅† | DRM sysfs · system_profiler |
| Apple | — | — | ✅ | system_profiler · sysctl · vm_stat |
† Intel Macs only — discrete and integrated GPUs are read from system_profiler.
Best-effort: you get an empty list on unsupported platforms, never an error.
Note: Verified on NVIDIA, AMD and Apple Silicon hardware. The Intel path is implemented but not yet confirmed on a real device, as are the ROCm and oneAPI install probes — if something doesn't work, please open an issue. Help from the community confirming detection on Intel GPUs is very much appreciated.
[dependencies]
gpu-probe = "0.1.1"NVIDIA support pulls in nvml-wrapper. For AMD/Apple-only builds, drop it:
gpu-probe = { version = "0.1.1", default-features = false }for gpu in gpu_probe::detect() {
println!("{gpu}");
// NVIDIA GeForce RTX 3090 (NVIDIA, sm_86): 24.0 GiB total, 9.8 GiB free
// AMD cyan_skillfish (AMD, gfx1013): 14.5 GiB total, 14.0 GiB free
}detect() returns Vec<GpuInfo>:
pub struct GpuInfo {
pub name: String,
pub vendor: Vendor, // Nvidia | Amd | Intel | Apple | Unknown
pub total_bytes: u64,
pub free_bytes: Option<u64>,
pub used_bytes: Option<u64>,
pub arch_target: Option<ArchTarget>, // Gfx | Sm | Xe | Apple
}Check whether a model fits, or pick the emptiest GPU:
let need = 16 * 1024 * 1024 * 1024; // 16 GiB
let fits = gpu_probe::detect()
.iter()
.any(|g| g.free_bytes.unwrap_or(g.total_bytes) >= need);
let emptiest = gpu_probe::detect()
.into_iter()
.max_by_key(|g| g.free_bytes.unwrap_or(g.total_bytes));Or run the bundled example: cargo run --example detect.
Which prebuilt artifact a GPU can run is reported per device, as one field carrying whichever form the vendor uses. A GPU has at most one, enforced by the type rather than by convention:
use gpu_probe::ArchTarget;
for gpu in gpu_probe::detect() {
match gpu.arch_target {
// AMD: the ROCm/HIP `--offload-arch` value
Some(ArchTarget::Gfx(gfx)) => println!("build {gfx}"), // gfx1013
// NVIDIA: the CUDA compute capability
Some(ArchTarget::Sm(sm)) => println!("build sm_{}{}", sm.major, sm.minor),
// Intel: the architecture family
Some(ArchTarget::Xe(arch)) => println!("build {arch}"), // xe-hpg
// Apple: the Metal feature tier
Some(ArchTarget::Apple(family)) => println!("targets {family}"), // apple8
// `ArchTarget` is #[non_exhaustive], so a wildcard is required — new
// vendors land as new variants without breaking this match.
_ => {}
}
}The four are not equally precise, and the table says why:
| vendor | value | source | selects a build? |
|---|---|---|---|
| AMD | gfx1013 |
KFD sysfs | yes — --offload-arch |
| NVIDIA | sm_86 |
NVML | yes — -arch |
| Intel | xe-hpg |
PCI device id | family only; ocloc -device is finer |
| Apple | apple8 |
chip name | no — a capability tier, .metallib is not per-family |
gfx(), sm(), xe(), and apple() pull out one vendor's form when that's
all you need:
let amd_targets: Vec<_> = gpu_probe::detect()
.iter()
.filter_map(|gpu| gpu.arch_target.and_then(ArchTarget::gfx))
.collect();ArchTarget itself is not ordered — comparing an AMD target to an NVIDIA one is
meaningless — but GfxTarget and ComputeCapability both order major first,
so a minimum requirement compares directly:
use gpu_probe::{ComputeCapability, GfxTarget};
let ampere_or_newer = ComputeCapability::new(8, 0);
let rdna2_or_newer = GfxTarget::new(10, 3, 0);The AMD form comes from KFD sysfs, published by the amdgpu kernel driver —
no ROCm install is required, and it is reported on machines that have none.
The Intel and Apple forms are derived rather than queried, because neither
platform publishes an architecture a caller can read: Intel's comes from a PCI
device id table, Apple's from the chip name — the latter following Apple's
published Metal Feature Set Tables. Both report None for anything their table
does not recognise rather than guessing, and the Intel table has not been
verified against real hardware yet.
Toolchain properties describe the machine rather than any one GPU, so they're returned separately. They are not the same measurement:
| function | reports | source |
|---|---|---|
cuda_host() |
CUDA driver version | NVML, kernel-side |
rocm_host() |
ROCm userspace release | $ROCM_PATH/.info/version → /opt/rocm |
oneapi_host() |
oneAPI toolkit release | $ONEAPI_ROOT/compiler/latest → /opt/intel/oneapi |
vulkan_host() |
Vulkan driver-advertised API version | loader (libvulkan.so.1) + ICD manifests under /usr/local/share/vulkan/icd.d, /usr/share/vulkan/icd.d |
Only NVIDIA exposes a driver version. amdgpu declares no MODULE_VERSION and
KFD publishes only a topology counter, so the AMD and Intel probes can report the
userspace install and nothing more — a property of the drivers, not an omission.
Vulkan is a runtime rather than a toolkit, and — unlike the other three — has no
per-GPU architecture to match: SPIR-V is portable and the driver compiles it at
load time, so vulkan_host() carries no arch_target counterpart.
use gpu_probe::{ComputeCapability, RocmVersion};
if let Some(cuda) = gpu_probe::cuda_host() {
println!("{} / CUDA {}", cuda.compute_capability, cuda.driver_version);
// 8.6 / CUDA 13.3
if cuda.compute_capability >= ComputeCapability::new(8, 0) {
// pick an Ampere-or-newer build
}
}
if let Some(rocm) = gpu_probe::rocm_host() {
println!("ROCm {}", rocm.version); // ROCm 6.2.4
if rocm.version >= RocmVersion::new(6, 0, 0) {
// pick a ROCm 6 build
}
}
if let Some(oneapi) = gpu_probe::oneapi_host() {
println!("oneAPI {}", oneapi.version); // oneAPI 2024.2.1
}
if let Some(vulkan) = gpu_probe::vulkan_host() {
println!("Vulkan {}", vulkan.api_version); // Vulkan 1.3.280
}Some is the signal that the stack is installed and the host can run its
builds. None is weaker: it means the stack was not found where this crate
looks — NVML unavailable for cuda_host() (no driver, no device, the nvidia
feature disabled, or unusable values), or no userspace install at the standard
prefixes for the others. A distro shipping ROCm into /usr rather than
/opt/rocm, or a container carrying only the runtime libraries, will report
None despite working. Treat Some as proof and None as "probably not,
worth confirming".
What None does not tell you is that the GPU is unusable. The kernel and
userspace halves are independent: arch_target comes from the amdgpu driver
and is reported with no ROCm installed at all.
vulkan_host() inverts that asymmetry, and the difference matters when you are
choosing a build. Its Some is the weak answer: the loader and an ICD manifest
are enough to satisfy it, and a software rasterizer — Mesa's lavapipe, installed
by default on many distributions — is an ICD like any other. So a machine with
no GPU at all can report a Vulkan API version, and a consumer reading that as
"this host can compute on a GPU" will pick a Vulkan artifact and run it on the
CPU through LLVM, slower than the CPU build it passed over. Pair it with
detect() when the question is capability rather than "is a runtime present":
an empty GPU list alongside Some is precisely the software-rasterizer case.
Its version needs the same care. api_version is the highest any installed
driver advertises in its ICD manifest — a static declaration on disk, read
without linking anything — and that is neither of the two numbers it resembles:
- Not the loader's instance version.
vulkaninfoandvkEnumerateInstanceVersionreport the loader's, so the two routinely disagree: a Mesa driver advertising1.4.354behind a1.4.357loader is reported here as1.4.354. The driver's is the one that binds, since loaders track the current headers while drivers implement features on their own schedule. The loader only becomes the constraint when the two are sourced separately — a container whose base image carries a stale loader against drivers bind-mounted from the host. - Not any single GPU's. With two drivers installed,
vulkan_host()is the higher of what they advertise and may describe neither card. Per-device versions come fromvkGetPhysicalDeviceProperties, which requires linking the loader and creating an instance — the vendor-SDK-free trade this crate exists to make.
Gate on major/minor in either case. Vulkan's patch number is the spec
header revision, not a feature level, so 1.4.354 and 1.4.357 are equally
Vulkan 1.4 and a patch-sensitive check only rejects builds that would have run.
cuda_host().compute_capability is device 0's — the same value that GPU reports
as ArchTarget::Sm in its own arch_target.
Readiness is four separate questions, and the pieces above answer each one:
use gpu_probe::{ArchTarget, GfxTarget};
let need = 16 * 1024 * 1024 * 1024; // 16 GiB
let built_for = GfxTarget::new(10, 1, 3); // this artifact is gfx1013
// The runtime has to be installed to execute anything.
let runtime_ready = gpu_probe::rocm_host().is_some();
let device_ready = gpu_probe::detect().iter().any(|gpu| {
// The kernel driver has to expose the GPU for compute — on AMD, a
// target at all means KFD is live.
gpu.arch_target == Some(ArchTarget::Gfx(built_for))
// And the weights have to fit.
&& gpu.free_bytes.unwrap_or(gpu.total_bytes) >= need
});
if runtime_ready && device_ready {
// load the model
}For NVIDIA the runtime check is cuda_host().is_some(), which tests the
driver — the usual thing to verify, since frameworks like PyTorch ship their
own CUDA runtime. For Intel, oneapi_host() tests for the toolkit and does not
see a distro-packaged Level Zero runtime, so it is the least complete of the
three.
total_bytesis dedicated VRAM on discrete GPUs. On integrated/unified GPUs (Intel iGPUs, AMD APUs, Apple Silicon) it's the shared system-memory ceiling, andfree_bytes/used_bytesare oftenNone.- Apple Silicon is an exception: because CPU and GPU share one pool,
free_bytes/used_bytescome from system-wide paging statistics (vm_stat), counting active + wired + compressor pages. That is the same figure Activity Monitor reports as "Memory Used", so it reflects the whole machine rather than the GPU alone. Discrete GPUs on Intel Macs report their own VRAM and are left untouched. - AMD APUs are a special case: their
mem_info_vram_totalis only a BIOS carveout (as little as 512 MiB), sototal_bytesadds the GTT pool they really allocate from — sized by the kernel'sttm.pages_limit— andfree_bytes/used_bytescover both pools. The result matches Vulkan/RADV to the byte; ROCm reports ~512 MiB less, since KFD publishes only the GTT-backed bank. - AMD GPU names come from the KFD ASIC codename (
AMD cyan_skillfish), falling back to the DRM node (AMD GPU (card1)) when KFD reports none. oneapi_host()covers the toolkit, not the GPU runtime: a host using a distro-packaged Level Zero driver with no toolkit reportsNone, since reading that runtime's version needs linking rather than a file read. Not yet verified against a real install.- NVIDIA detection reads NVML from the installed driver at runtime — the CUDA toolkit is not required.
- NVML is initialized once per process and intentionally never shut down. Cycling
nvmlInit/nvmlShutdownleaks a file descriptor each time, sodetect()is safe to poll on a timer: descriptor use is flat, and each call still returns live memory values.