This document describes the design decisions, component interactions, and operational patterns for the model-serving-stack.
flowchart TD
Client(["Client / Application"])
GW["FastAPI Gateway<br/>OpenAI-compatible, :8080<br/>api/main.py"]
MW["StructuredLoggingMiddleware<br/>request IDs, latency, status"]
subgraph Backends ["Inference Backends"]
VLLM["vLLM<br/>PagedAttention<br/>vllm/"]
Triton["Triton Inference Server<br/>Python backend<br/>triton/"]
Ray["Ray Serve<br/>autoscaling<br/>ray_serve/"]
end
subgraph Packaging ["Model Packaging"]
BentoML["BentoML<br/>bentoml/"]
end
subgraph Autoscaling ["GPU Autoscaling"]
DCGM["DCGM Exporter<br/>DaemonSet"]
PromAdapter["Prometheus Adapter"]
HPA["Kubernetes HPA<br/>DCGM_FI_DEV_GPU_UTIL target 70"]
end
subgraph Observability ["Observability"]
Prom["Prometheus"]
Graf["Grafana Dashboard"]
end
Client --> GW
GW --> MW
MW -->|"vllm/client.py"| VLLM
Client -->|"triton/client.py"| Triton
Client -->|"serve run"| Ray
Ray -->|"vllm/client.py"| VLLM
BentoML -->|"vllm/client.py"| VLLM
DCGM -->|"DCGM_FI_DEV_GPU_UTIL"| Prom
Prom --> PromAdapter
PromAdapter --> HPA
HPA -->|"scale replicas"| VLLM
Prom -->|"scrapes :8000 / :8002"| VLLM
Prom --> Graf
The gateway has exactly one backend: api/main.py constructs a VLLMClient at
import time and every chat completion goes to vLLM. Triton and Ray Serve are
first-class parts of the stack but are addressed directly, not through the
gateway. See configuration.md.
Each backend serves a different operational profile:
- vLLM — best for dynamic batching at high concurrency. PagedAttention prevents KV cache fragmentation. It is the default and the only backend the gateway calls.
- Triton — best for latency-critical serving and multi-model hosting. In the
committed baseline it showed ~31% lower TTFT p50 than native vLLM at
concurrency 32 (142 ms -> 98 ms); see benchmarks.md for the
caveats on that measurement. The model repository here uses Triton's Python
backend wrapping a vLLM
AsyncLLMEngine, which is the flexible path; a TRT-LLM engine would be compiled per GPU and swapped in behind the sameconfig.pbtxtcontract. - Ray Serve — best when you need programmatic autoscaling policies, Python
request preprocessing, or multi-model routing (see
ray_serve/multi_model_router.py). The baseline recorded ~10% overhead versus native vLLM from actor scheduling.
Switching the gateway's backend currently requires a code change in
api/main.py; there is no environment-variable selector.
LLM inference is GPU-bound, not CPU/memory-bound. A standard HPA on CPU
utilization triggers far too late (GPU saturated, CPU still low) or far too
early (GPU idle, CPU busy with Python overhead). GPU memory is no better as a
signal because vLLM pre-allocates its KV cache to gpu_memory_utilization, so
usage barely moves with load. DCGM_FI_DEV_GPU_UTIL measures GPU execution
engine utilization directly, which tracks real serving pressure.
The 70% target provides a buffer for request spikes before saturation turns into queuing latency. Full rationale, dependencies, and tuning: autoscaling.md.
Centralizing request logging at the gateway (rather than inside each backend) means:
- Every request gets a correlation
request_id, returned asX-Request-ID. - Latency is measured from the client's perspective, not the backend's internal view.
- Requests that fail before reaching a backend are still logged.
/health and /metrics are excluded from this logging so probes do not flood
the log. The middleware also stamps X-Backend: vllm on responses. It is
implemented in api/main.py; logs are emitted as logging records on the
api.access logger with structured extra fields, not exported to Prometheus.
triton/model_repository/
└── llama3/
├── config.pbtxt # model name, backend, transaction policy, tensors
└── 1/
└── model.py # Python backend wrapping a vLLM AsyncLLMEngine
config.pbtxt specifies:
name: "llama3",backend: "python"max_batch_size: 0— batching is handled by the vLLM engine inside the model, not by Triton's dynamic batchermodel_transaction_policy { decoupled: true }— allows streaming/decoupled responses- Inputs:
text_input(TYPE_STRING), optionalmax_tokens(TYPE_INT32), optionaltemperature(TYPE_FP32) - Output:
text_output(TYPE_STRING) instance_group: oneKIND_GPUinstance on GPU 0
kubernetes/hpa-dcgm.yaml is the canonical HPA. It uses an external metric
sourced from the DCGM Prometheus Exporter:
metrics:
- type: External
external:
metric:
name: DCGM_FI_DEV_GPU_UTIL
selector:
matchLabels:
gpu: "0"
target:
type: AverageValue
averageValue: "70"It also sets scaling behavior: a 60-second scale-up stabilization window at
one pod per minute, and a 300-second scale-down window at one pod per two
minutes. This requires the
Prometheus Adapter to
bridge DCGM metrics from Prometheus into the Kubernetes external metrics API.
- Client sends
POST /v1/chat/completionsto the FastAPI gateway on port 8080. StructuredLoggingMiddlewareassigns arequest_idand records the start time.- The handler validates the body against the Pydantic schemas in
api/models.py. VLLMClient.chat()forwards the messages to the vLLM OpenAI-compatible server (VLLM_HOST/VLLM_PORT) using theopenaiSDK.- vLLM processes the request with PagedAttention KV cache management.
- The gateway returns a complete
chat.completionresponse. Streaming (SSE) is not implemented on this path; the request schema has nostreamfield. - The middleware logs method, path, status, latency, client IP and user agent,
and sets
X-Request-IDandX-Backendon the response.
Prometheus scrapes the backends directly — vllm-server:8000 and
triton-server:8002 per monitoring/prometheus/prometheus.yml. The gateway
itself does not currently expose a /metrics endpoint.
Sibling reference projects by the same author; they are independent repositories, not runtime dependencies of this one.
| Repo | Relationship |
|---|---|
llm-finetuning-lab |
LoRA/QLoRA fine-tuning pipeline producing weights a stack like this serves |
inference-optimization-bench |
Quantization and inference benchmarking that informs backend selection |
nvidia-nim-agent-toolkit |
Agent toolkit over NVIDIA NIM — the client side of a serving stack |