Skip to content

About

run multiple ollama models in runpod

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

runpod-ollama-multi

A RunPod serverless worker that runs Ollama with multiple models, selectable per request.

Deploy on RunPod

  1. Fork or clone this repo
  2. Connect your GitHub repo to a RunPod Serverless endpoint via Custom Deployment
  3. Set the following environment variables in your endpoint configuration:
Variable Description Example
OLLAMA_MODELS Comma-separated list of models to pull on startup qwen3.5:9b,qwen3-embedding:0.6b
  1. Attach a network volume at /runpod-volume to cache models across cold starts
  2. Set container disk to at least 40 GB
  3. Create a GitHub release to trigger the build

Request Format

Chat

{
  "input": {
    "model": "qwen3.5:9b",
    "method": "chat",
    "payload": {
      "messages": [
        { "role": "user", "content": "Hello" }
      ]
    }
  }
}

Generate

{
  "input": {
    "model": "qwen3.5:4b",
    "method": "generate",
    "payload": {
      "prompt": "Why is the sky blue?"
    }
  }
}

Embed

{
  "input": {
    "model": "qwen3-embedding:0.6b",
    "method": "embed",
    "payload": {
      "input": "Hello world"
    }
  }
}

Notes

  • Any model available on ollama.com/library can be used
  • Models not in OLLAMA_MODELS will be pulled on first request and cached to the network volume
  • Streaming is not supported in serverless mode

About

run multiple ollama models in runpod

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages