Thank you for your interest in contributing! This document outlines the development workflow and standards.
git clone https://github.com/TylrDn/model-serving-stack.git
cd model-serving-stack
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install -e ".[dev]"
cp .env.template .env # see docs/configuration.mdCI (.github/workflows/ci.yml) enforces two gates on every push and pull request to main:
ruff— lint rulesE,F,Iat line length 100 ([tool.ruff]inpyproject.toml)pytest— with--cov-fail-under=80overapi,vllm,triton,evals
mypy is configured in pyproject.toml but is not run by CI. Running it locally is welcome; a clean mypy run is not a merge requirement today.
Run exactly what CI runs before submitting a PR:
ruff check .
pytest tests/ -v -m "not gpu" \
--cov=api --cov=vllm --cov=triton --cov=evals \
--cov-report=term-missing \
--cov-fail-under=80- Branch from
main:git checkout -b feat/your-feature - Make changes with focused commits using Conventional Commits:
feat:,fix:,docs:,test:,refactor: - Open a PR against
main— CI runs automatically - All CI checks must be green before merge
To add a new serving backend (e.g., TensorRT-LLM standalone):
- Create a directory under the backend name:
tensorrt_llm/ - Implement a client with the same surface as
vllm/client.py(chat()andhealth_check()), which is whatapi/main.py,ray_serve/deployment.pyandbentoml/service.pyconsume - Add the backend key to
configs/models.yaml - Add unit tests in
tests/test_<backend>.py, mocking the client library the waytests/test_triton_client.pydoes, and mark anything needing real hardware with@pytest.mark.gpu - Document any new environment variables in
.env.templateanddocs/configuration.md
# All non-GPU tests
pytest tests/ -m "not gpu"
# Specific test file
pytest tests/test_api.py -v
# HTML coverage report
pytest --cov=. --cov-report=html
open htmlcov/index.html
# Load test against a running gateway on :8080
python -m evals.load_test --requests 500 --concurrency 32 --output-dir results/load_testsSee docs/testing.md for the coverage map and the gpu marker convention.
- Bump
versioninpyproject.tomlfollowing SemVer. - Move the
## [Unreleased]entries inCHANGELOG.mdunder a new## [X.Y.Z] - YYYY-MM-DDheading and leave a fresh empty## [Unreleased]section. - Commit with
chore(release): vX.Y.Zand confirm CI is green onmain. - Tag and push:
git tag -a vX.Y.Z -m "vX.Y.Z" git push origin main --follow-tags - If the release changes benchmark numbers, commit the new artifact under
results/using the schema in results/README.md and update the README table in the same commit.
Open an issue on this repository.