GPU Cluster Monitoring (GCM): Large-Scale AI Research Cluster Monitoring
-
Updated
Oct 6, 2026 - Python
GPU Cluster Monitoring (GCM): Large-Scale AI Research Cluster Monitoring
GPU Observability with workload attribution. One OTLP agent per node ties hardware metrics (NVIDIA, AMD, Intel Gaudi) to the K8s pod or Slurm job burning the GPU.
Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.
Autoscale NVIDIA Triton on Kubernetes by GPU utilisation — DCGM, Prometheus adapter, HPA, GPU time-slicing. Tested end-to-end.
top-like TUI for per-pod GPU usage in Kubernetes (zero cluster footprint)
GPU 风扇控制台:按 GPU 温度闭环调速服务器机箱风扇(IPMI raw + FastAPI + Vue3 + SQLite,DCGM/Prometheus 数据源,带安全护栏/看门狗),附带按 CPU 温度控速的 Dell PowerEdge 旧版脚本 | Web-based IPMI server fan control driven by GPU temperature
Single-file interactive study tool for the NVIDIA NCP-AI Operations exam: quiz, flashcards, guided labs, and a stateful BCM/Slurm/Kubernetes/DCGM command sandbox.
Simulate NVIDIA GPUs for testing. 7 behavior profiles, scale to 1000+ GPUs, Docker-ready Prometheus exporter using DCGM
Moved to freelensapp/freelens-gpu-extension (npm @freelensapp/gpu-extension)
End-to-end observability for disaggregated LLM inference on EKS — DCGM metrics, KEDA autoscaling on GPU signals, per-namespace cost attribution, multi-agent OTel tracing, and Istio mTLS. Reference implementation for the OpenTelemetry AI Inference Platform blueprint.
GPU-native agent-swarm orchestration for the NVIDIA AI stack — NeMo, NIM, Triton, DCGM, NGC, NIXL, OpenShell. Spawn GPU-pinned agent teams across DGX/HGX nodes with NVLink-aware scheduling, task DAGs, adaptive scheduling, and full observability.
Find out how many GPU-hours you're paying for and not using
Automated acceptance toolkit for Linux deep learning GPU servers
One-command NVIDIA DGX Spark (GB10) cluster GPU monitoring with Grafana + Prometheus + DCGM + node_exporter + vLLM. Track GPU temperature, utilization, power, SM clock, memory, disk, network and LLM inference throughput in a pre-built dashboard. 一条命令搭建 DGX Spark 集群监控
Production LLM serving infrastructure using Triton Inference Server, vLLM, and Ray Serve with OpenAI-compatible endpoints. Includes Kubernetes autoscaling configs driven by DCGM GPU metrics and a BentoML packaging path for portable model deployment.
Freelens extension for GPUs
Prometheus exporter for hardware telemetry from DMTF Redfish-capable BMCs. Multi-target probe pattern. Demo stack included.
kubectl plugin that compares requested GPU resources against DCGM Exporter utilization metrics and generates rightsizing recommendations with projected monthly cost savings. Supports nvidia.com/gpu and amd.com/gpu — the gap VPA leaves open.
To associate your repository with the dcgm topic, visit your repo's landing page and select "manage topics."