Skip to content

Latest commit

 

History

History
92 lines (70 loc) · 2.84 KB

File metadata and controls

92 lines (70 loc) · 2.84 KB

Exporting without gpuwaste extract

If you can't reach Prometheus from your machine — network policy, mTLS mesh, RBAC, air-gapped cluster — that's expected. Hand these queries to whoever has access, or export from Grafana. The analyzer only needs the CSVs.

Use a uniform step. GPU-hours are sum(step × value), so irregular samples silently corrupt the dollar figure. 300s over 7 days is a good default.


utilization.csv

Run each as a range query and export as CSV.

DCGM_FI_PROF_SM_ACTIVE
DCGM_FI_PROF_GR_ENGINE_ACTIVE
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE
DCGM_FI_DEV_FB_USED
DCGM_FI_DEV_FB_FREE
DCGM_FI_DEV_POWER_USAGE

Join on timestamp + UUID into these columns:

Column From
timestamp sample time, RFC3339 or epoch seconds
gpu_uuid UUID label
node Hostname label
gpu_model modelName label
namespace exported_namespace or namespace
pod exported_pod or pod
container exported_container or container
sm_active DCGM_FI_PROF_SM_ACTIVE (0–1)
gr_engine_active DCGM_FI_PROF_GR_ENGINE_ACTIVE (0–1)
tensor_active DCGM_FI_PROF_PIPE_TENSOR_ACTIVE (0–1)
fb_used_mib DCGM_FI_DEV_FB_USED
fb_total_mib FB_USED + FB_FREE
power_watts DCGM_FI_DEV_POWER_USAGE

DCGM_FI_DEV_GPU_UTIL is not used and should not be substituted. It reports that a kernel is resident, not that useful work is happening — a process pinning the device with a trivial kernel reads 100%.

If the PROF_* metrics return nothing, DCGM profiling is disabled. The tool still works; ghost-work and memory-parked detection are suppressed and it says so in the report.

If there are no pod labels, dcgm-exporter isn't configured with Kubernetes device-plugin mapping. You'll get node-level attribution only.


allocation.csv

kube_pod_container_resource_requests{resource="nvidia_com_gpu"}
kube_pod_container_resource_limits{resource="nvidia_com_gpu"}
kube_pod_container_status_restarts_total
Column From
timestamp sample time
namespace, pod, container labels
node node label
gpu_requested resource_requests value
gpu_limit resource_limits value
restarts restarts_total value

Without this file there is no waste calculation and no dollar figure — only a utilization report.


Sanity checks before you analyze

  • Every timestamp gap is identical. Uneven sampling is the one failure that produces a plausible-looking wrong number.
  • fb_total_mib is non-zero.
  • sm_active is in the range 0–1, not 0–100.

Then:

gpuwaste analyze --data ./export --step 300 --pricing pricing.yaml

Sharing an export with someone outside your organization? Run gpuwaste anonymize first, then gpuwaste inspect to see exactly what's in it.