A RunPod serverless worker that runs Ollama with multiple models, selectable per request.
- Fork or clone this repo
- Connect your GitHub repo to a RunPod Serverless endpoint via Custom Deployment
- Set the following environment variables in your endpoint configuration:
| Variable | Description | Example |
|---|---|---|
| OLLAMA_MODELS | Comma-separated list of models to pull on startup | qwen3.5:9b,qwen3-embedding:0.6b |
- Attach a network volume at
/runpod-volumeto cache models across cold starts - Set container disk to at least 40 GB
- Create a GitHub release to trigger the build
{
"input": {
"model": "qwen3.5:9b",
"method": "chat",
"payload": {
"messages": [
{ "role": "user", "content": "Hello" }
]
}
}
}{
"input": {
"model": "qwen3.5:4b",
"method": "generate",
"payload": {
"prompt": "Why is the sky blue?"
}
}
}{
"input": {
"model": "qwen3-embedding:0.6b",
"method": "embed",
"payload": {
"input": "Hello world"
}
}
}- Any model available on ollama.com/library can be used
- Models not in
OLLAMA_MODELSwill be pulled on first request and cached to the network volume - Streaming is not supported in serverless mode