If you can't reach Prometheus from your machine — network policy, mTLS mesh, RBAC, air-gapped cluster — that's expected. Hand these queries to whoever has access, or export from Grafana. The analyzer only needs the CSVs.
Use a uniform step. GPU-hours are sum(step × value), so irregular samples
silently corrupt the dollar figure. 300s over 7 days is a good default.
Run each as a range query and export as CSV.
DCGM_FI_PROF_SM_ACTIVE
DCGM_FI_PROF_GR_ENGINE_ACTIVE
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE
DCGM_FI_DEV_FB_USED
DCGM_FI_DEV_FB_FREE
DCGM_FI_DEV_POWER_USAGE
Join on timestamp + UUID into these columns:
| Column | From |
|---|---|
timestamp |
sample time, RFC3339 or epoch seconds |
gpu_uuid |
UUID label |
node |
Hostname label |
gpu_model |
modelName label |
namespace |
exported_namespace or namespace |
pod |
exported_pod or pod |
container |
exported_container or container |
sm_active |
DCGM_FI_PROF_SM_ACTIVE (0–1) |
gr_engine_active |
DCGM_FI_PROF_GR_ENGINE_ACTIVE (0–1) |
tensor_active |
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE (0–1) |
fb_used_mib |
DCGM_FI_DEV_FB_USED |
fb_total_mib |
FB_USED + FB_FREE |
power_watts |
DCGM_FI_DEV_POWER_USAGE |
DCGM_FI_DEV_GPU_UTIL is not used and should not be substituted. It reports
that a kernel is resident, not that useful work is happening — a process pinning
the device with a trivial kernel reads 100%.
If the PROF_* metrics return nothing, DCGM profiling is disabled. The tool
still works; ghost-work and memory-parked detection are suppressed and it says
so in the report.
If there are no pod labels, dcgm-exporter isn't configured with Kubernetes device-plugin mapping. You'll get node-level attribution only.
kube_pod_container_resource_requests{resource="nvidia_com_gpu"}
kube_pod_container_resource_limits{resource="nvidia_com_gpu"}
kube_pod_container_status_restarts_total
| Column | From |
|---|---|
timestamp |
sample time |
namespace, pod, container |
labels |
node |
node label |
gpu_requested |
resource_requests value |
gpu_limit |
resource_limits value |
restarts |
restarts_total value |
Without this file there is no waste calculation and no dollar figure — only a utilization report.
- Every timestamp gap is identical. Uneven sampling is the one failure that produces a plausible-looking wrong number.
fb_total_mibis non-zero.sm_activeis in the range 0–1, not 0–100.
Then:
gpuwaste analyze --data ./export --step 300 --pricing pricing.yamlSharing an export with someone outside your organization? Run
gpuwaste anonymize first, then gpuwaste inspect to see exactly what's in it.