base64_module.cpp— C++ extension: fast file → base64 encodersetup.py— builds the C++ extension into a Python-importable modulemain.py— Python script that calls the NVIDIA chat completions API, using the C++ extension for any file/image encoding
pip install pybind11
python setup.py build_ext --inplaceThis produces a fastb64*.so file in this directory that main.py imports
directly.
Do not put the key in the source code. Export it as an environment variable instead:
export NVIDIA_API_KEY="nvapi-..."(If you pasted a key into this chat earlier, treat it as compromised — revoke/regenerate it in your NVIDIA NGC account.)
python main.pyTo send an image with the prompt:
call_api("What's in this image?", image_path="photo.png")The C++ extension handles reading + base64-encoding the file; Python handles the HTTP request and JSON response.
Files added for this: Dockerfile, .dockerignore, requirements.txt,
app.py, render.yaml.
app.py is a small Flask app (served by gunicorn in the container) with
two routes:
GET /— health check, used by Render and to confirm the service is livePOST /chat— body{"prompt": "..."}(optionally"image_path": "..."), calls the NVIDIA API and returns the JSON response
This fits Render's free Web Service tier — no payment info required. The tradeoff: free web services spin down after ~15 min of inactivity and take a few seconds to wake on the next request. Fine for testing/light use; upgrade to a paid instance later if you need it always-on.
Option A — Blueprint (render.yaml), one click:
- Push this folder to a GitHub/GitLab repo.
- In Render: New > Blueprint → pick the repo. Render reads
render.yamland creates thenvidia-qwen-webservice for you. - Go to the new service's Environment tab and set
NVIDIA_API_KEY(it's markedsync: falseso it isn't committed to your repo). - Deploy.
Option B — Manual:
- Push this folder to a repo.
- In Render: New > Web Service → pick the repo.
- Runtime: Docker (Render finds the
Dockerfileautomatically). - Instance type: Free.
- Add environment variable
NVIDIA_API_KEYunder the Environment tab. - Create the service.
Once deployed, test it:
curl https://<your-service>.onrender.com/
curl -X POST https://<your-service>.onrender.com/chat \
-H "Content-Type: application/json" \
-d '{"prompt": "Hello!"}'Note: main.py checks for NVIDIA_API_KEY at startup, so if it's missing
the app will fail to boot with a clear error in the logs rather than run
with a broken key — set the env var and redeploy to fix.