CoreWeave provides managed Grafana dashboards for metrics and logs.
- Deploy the vLLM example with the Helm chart.
- Open the Grafana URL from the CoreWeave console.
- Navigate to Kubernetes / Workloads / Pods and filter by your namespace.
Check GPU utilization, memory, network throughput, and pod restarts while invoking the model.
- The
kube-state-metricsandnode-exporterdashboards show cluster health. - vLLM exports Prometheus metrics such as
vllm_engine_execution_time. - Ensure Prometheus scrapes the vLLM service on port
8000to populate Grafana. - Logs are collected via Loki; search by
app=vllm. - For a local demo use the
observability/example which spins up Prometheus and Grafana with Docker Compose.
Screenshots can be added to docs/img/ for presentations.