gpuwaste synth writes CSVs directly, so it tests the analyzer but not the
extract path. This rig runs a fake dcgm-exporter behind a real Prometheus,
using the same label spellings a live cluster produces — so the query,
label-mapping and join code all get exercised for real.
Runs on any machine, Apple Silicon included. No NVIDIA hardware involved.
cd dev
docker compose up -d
# let it accumulate a few minutes of scrapes.
# time is compressed: 10 real seconds = 1 simulated hour, so a full
# day/night cycle passes every 4 minutes.
sleep 300
gpuwaste extract --prometheus http://localhost:9090 --days 1 --step 15 --out ./local
gpuwaste analyze --data ./local --step 15 --cluster local-testExpect the same seven findings synth produces. If extract returns them, the
Prometheus path works and the only thing left untested is your cluster's
specific label schema.
Check the raw metrics directly:
curl -s localhost:9400/metrics | head -30
open http://localhost:9090 # Prometheus UITear down with docker compose down.
Does: the range queries are correct, DCGM label variants resolve, the
utilization/allocation join works, fb_total is derived properly, and uniform
sampling is preserved end to end.
Doesn't: your cluster's actual label spellings, whether DCGM profiling is enabled there, or how kube-state-metrics is configured. Those only surface against the real thing — which is the point of running it internally.