A practical, self-hostable, and AWS/Azure-portable curated list of 180 tools covering the complete AI/ML delivery path — from idea to production.
AI/ML development is no longer only about training a model. A usable product also needs data collection, labeling, retrieval, prompts, agents, evaluation, privacy, model serving, billing, observability, CI/CD, and a reliable user experience. The challenge is choosing enough software to move quickly without creating an unmaintainable platform.
This list follows a practical FOSS-first philosophy: prefer focused and composable tools, keep data and model artifacts portable, use standard APIs, avoid unnecessary platform complexity, and introduce GPUs or Kubernetes only when the workload justifies them.
The fastest path is not "use every AI tool." Start with one model, one dataset, one evaluation set, one API, and one observable deployment. Add components when a real bottleneck appears.
- Legend
- How to Use This List
- AI Application Platforms and Chat Interfaces
- Agent Frameworks and Orchestration
- Local Model Runtimes and Model Clients
- RAG, Vector Search, and Retrieval
- Dataset Creation, Labeling, and Data Quality
- Classical ML and Deep Learning Foundations
- Fine-Tuning, Alignment, and Distributed Training
- Experiment Tracking, MLOps, and Pipelines
- Model Serving and Inference Optimization
- AI Evaluation, Tracing, Safety, and Observability
- Voice, Speech, and Audio AI
- Computer Vision, OCR, Documents, and Multimodal AI
- Synthetic Data, Privacy, Governance, and AI Security
- Git, CI/CD, Kubernetes, and AI Infrastructure
- Productization, AI Operations, and Platform Integration
- MVP-to-Production Paths
- Git-Connected AI Development Loop
- AWS and Azure Migration
- AI/ML Production Checklist
- Contributing
- License
| Label | Meaning |
|---|---|
| 🟢 Now | Usually practical for an MVP or a small team. |
| 🟡 Evaluate | Useful when a concrete problem appears; validate complexity and fit first. |
| 🔵 Production | Appropriate when throughput, latency, reliability, or operational requirements justify it. |
| 🟣 Later | Powerful but usually requires clusters, GPUs, complex operations, or a larger team. |
| Review current license, model license, edition, trademark, and hosted-service terms before resale or white-labeling. | |
| Application/Platform | A deployable project or service. |
| Framework/Library | A component used inside your application or training pipeline. |
The catalog is deliberately split by responsibility. Choose one or two tools from a category, not the entire category. A small team may need only eight to twelve tools for its first product. Every tool has a direct link, a short purpose, a type, and a suggested adoption stage.
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 1 | Open WebUI | Self-hosted chat interface for local and remote models | Application | 🟢 Now |
| 2 | LibreChat | Multi-provider conversational AI interface | Application | 🟢 Now |
| 3 | AnythingLLM | Private document chat and knowledge assistants | Application | 🟢 Now |
| 4 | Dify | Visual LLM apps, workflows, and chatbots | Platform | 🟢 Now |
| 5 | Flowise | Visual LLM flows and agent prototypes | Platform | 🟢 Now |
| 6 | Langflow | Visual orchestration for LLM applications | Platform | 🟢 Now |
| 7 | Rasa | Controlled conversational assistants and intent workflows | Platform | 🟢 Now |
| 8 | Onyx | Enterprise search and knowledge assistant | Application | 🟡 Evaluate |
| 9 | Khoj | Self-hosted personal knowledge assistant | Application | 🟢 Now |
| 10 | Jan | Local-first desktop AI assistant | Application | 🟢 Now |
| 11 | PrivateGPT | Private document question answering | Application | 🟡 Evaluate |
| 12 | LobeChat | Open-source AI chat and agent workspace | Application | 🟢 Now |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 13 | LangChain | Chains, tool calls, retrieval, and agents | Framework | 🟢 Now |
| 14 | LlamaIndex | Data connectors, indexing, and RAG | Framework | 🟢 Now |
| 15 | Haystack | Search, RAG, pipelines, and LLM applications | Framework | 🟢 Now |
| 16 | Semantic Kernel | AI orchestration, plugins, memory, and tools | Framework | 🟡 Evaluate |
| 17 | AutoGen | Multi-agent conversations and task orchestration | Framework | 🟡 Evaluate |
| 18 | CrewAI | Role-based multi-agent workflows | Framework | 🟡 Evaluate |
| 19 | DSPy | Programmatic prompt and LM-pipeline optimization | Framework | 🟡 Evaluate |
| 20 | PydanticAI | Typed Python agents and tool calling | Framework | 🟢 Now |
| 21 | Letta | Stateful agents and long-term memory | Platform | 🟡 Evaluate |
| 22 | smolagents | Lightweight code and tool-using agents | Framework | 🟢 Now |
| 23 | OpenHands | Software-engineering agents for repositories and tools | Application | 🟡 Evaluate |
| 24 | Browser Use | Browser automation for AI agents | Framework | 🟡 Evaluate |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 25 | Ollama | Local model download and inference | Runtime | 🟢 Now |
| 26 | LocalAI | OpenAI-compatible local model API | Server | 🟢 Now |
| 27 | LM Studio | Desktop local model execution and API serving | Application | 🟢 Now |
| 28 | llama.cpp | Portable CPU/GPU inference runtime | Runtime | 🟢 Now |
| 29 | vLLM | High-throughput LLM serving | Server | 🔵 Production |
| 30 | Text Generation Inference | Production text-generation serving | Server | 🔵 Production |
| 31 | SGLang | Efficient LLM and multimodal serving | Server | 🔵 Production |
| 32 | MLC LLM | Deploy LLMs across local and edge hardware | Runtime | 🟡 Evaluate |
| 33 | MLX | Machine learning framework for Apple Silicon | Framework | 🟢 Now (Mac) |
| 34 | GPT4All | Private local chat and model execution | Application | 🟢 Now |
| 35 | KoboldCpp | Local GGUF model server and UI | Application | 🟡 Evaluate |
| 36 | Text Generation WebUI | Local model experimentation interface | Application | 🟡 Evaluate |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 37 | Qdrant | Vector search and semantic retrieval | Server | 🟢 Now |
| 38 | Weaviate | Vector search and AI retrieval | Server | 🟡 Evaluate |
| 39 | Milvus | Distributed vector database | Server | 🔵 Production |
| 40 | Chroma | Local-first embeddings and vector search | Database | 🟢 Now |
| 41 | LanceDB | Embedded vector and multimodal data search | Database | 🟡 Evaluate |
| 42 | pgvector | Vector similarity inside PostgreSQL | Database extension | 🟢 Now |
| 43 | Vespa | Large-scale search, ranking, and serving | Platform | 🔵 Production |
| 44 | OpenSearch | Search, analytics, and vector retrieval | Platform | 🟡 Evaluate |
| 45 | Elasticsearch | Search, analytics, and vector retrieval | Platform | 🟡 Evaluate |
| 46 | R2R | RAG ingestion, search, and retrieval APIs | Platform | 🟡 Evaluate |
| 47 | txtai | Semantic search and language-model workflows | Framework | 🟡 Evaluate |
| 48 | Vald | Cloud-native distributed vector search | Platform | 🟣 Later |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 49 | DVC | Git-like versioning for datasets and models | Platform | 🟢 Now |
| 50 | Label Studio | Text, image, audio, document, and LLM annotation | Application | 🟢 Now |
| 51 | CVAT | Computer-vision annotation and datasets | Application | 🟢 Now |
| 52 | FiftyOne | Inspect, curate, search, and evaluate vision data | Application | 🟡 Evaluate |
| 53 | Roboflow | Vision dataset preparation and deployment | Platform | 🟡 Evaluate |
| 54 | Hugging Face Datasets | Load, transform, and share datasets | Framework | 🟢 Now |
| 55 | lakeFS | Git-like branching for data lakes | Platform | 🟣 Later |
| 56 | Pachyderm | Data versioning and reproducible pipelines | Platform | 🟣 Later |
| 57 | Great Expectations | Dataset validation and quality expectations | Framework | 🟢 Now |
| 58 | Cleanlab | Find label errors and improve dataset quality | Framework | 🟡 Evaluate |
| 59 | DataHub | Metadata catalog, lineage, and data discovery | Platform | 🟣 Later |
| 60 | OpenMetadata | Metadata, governance, discovery, and lineage | Platform | 🟣 Later |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 61 | PyTorch | Deep-learning model training | Framework | 🟢 Now |
| 62 | TensorFlow | Model training and deployment | Framework | 🟡 Evaluate |
| 63 | JAX | High-performance numerical computing and ML | Framework | 🟡 Evaluate |
| 64 | scikit-learn | Classical ML and preprocessing | Framework | 🟢 Now |
| 65 | XGBoost | Gradient-boosted models for tabular data | Framework | 🟢 Now |
| 66 | LightGBM | Efficient gradient boosting | Framework | 🟢 Now |
| 67 | CatBoost | Gradient boosting with categorical features | Framework | 🟢 Now |
| 68 | Keras | High-level deep-learning API | Framework | 🟢 Now |
| 69 | fastai | Practical deep-learning training | Framework | 🟢 Now |
| 70 | statsmodels | Statistical models and inference | Framework | 🟢 Now |
| 71 | Prophet | Time-series forecasting | Framework | 🟡 Evaluate |
| 72 | RAPIDS | GPU-accelerated data science | Framework | 🟣 Later |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 73 | Transformers | Use and fine-tune language, vision, and audio models | Framework | 🟢 Now |
| 74 | TRL | Transformer reinforcement learning and alignment | Framework | 🟡 Evaluate |
| 75 | Unsloth | Memory-efficient LLM fine-tuning | Framework | 🟡 Evaluate |
| 76 | Axolotl | Configuration-driven LLM fine-tuning | Tool | 🟡 Evaluate |
| 77 | LLaMA-Factory | Fine-tune and align language models | Tool | 🟡 Evaluate |
| 78 | PEFT | Parameter-efficient fine-tuning | Framework | 🟢 Now |
| 79 | DeepSpeed | Distributed and memory-efficient training | Framework | 🟣 Later |
| 80 | Composer | Efficient neural-network training methods | Framework | 🟡 Evaluate |
| 81 | OpenRLHF | RLHF and preference-training workflows | Framework | 🟣 Later |
| 82 | LitGPT | Train and fine-tune language models | Framework | 🟡 Evaluate |
| 83 | torchtune | PyTorch-native LLM fine-tuning | Framework | 🟡 Evaluate |
| 84 | Colossal-AI | Distributed training and inference | Framework | 🟣 Later |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 85 | MLflow | Experiment tracking, model registry, and AI lifecycle | Platform | 🟢 Now |
| 86 | Kubeflow | Kubernetes ML pipelines and training | Platform | 🟣 Later |
| 87 | ClearML | Experiment management and ML orchestration | Platform | 🟡 Evaluate |
| 88 | Metaflow | Reproducible data-science workflows | Platform | 🟡 Evaluate |
| 89 | Dagster | Data and asset-oriented pipelines | Platform | 🟡 Evaluate |
| 90 | Apache Airflow | Scheduled data and ML pipelines | Platform | 🟢 Now (when needed) |
| 91 | Flyte | Reproducible scalable workflow orchestration | Platform | 🟣 Later |
| 92 | ZenML | MLOps framework connecting experiments and deployment | Framework | 🟡 Evaluate |
| 93 | Polyaxon | ML experimentation and orchestration | Platform | 🟣 Later |
| 94 | Feast | Feature store for training and online inference | Platform | 🟣 Later |
| 95 | Aim | Open-source experiment tracking | Platform | 🟢 Now |
| 96 | TensorBoard | Training metrics and model visualization | Application | 🟢 Now |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 97 | KServe | Kubernetes-native model serving | Platform | 🟣 Later |
| 98 | Seldon Core | Model deployment and inference graphs | Platform | 🟣 Later |
| 99 | BentoML | Package and deploy models as APIs | Platform | 🟢 Now |
| 100 | Ray Serve | Distributed model and Python service serving | Platform | 🟣 Later |
| 101 | NVIDIA Triton | Multi-framework GPU model serving | Server | 🔵 Production |
| 102 | MLServer | Standardized model inference server | Server | 🟡 Evaluate |
| 103 | TorchServe | PyTorch model serving | Server | 🟡 Evaluate |
| 104 | TorchX | Distributed ML job launching | Framework | 🟣 Later |
| 105 | OpenVINO | Inference optimization across hardware | Runtime | 🟡 Evaluate |
| 106 | ONNX Runtime | Cross-platform model inference | Runtime | 🟢 Now |
| 107 | Apache TVM | Compiler stack for ML deployment | Framework | 🟣 Later |
| 108 | TensorRT-LLM | Optimized NVIDIA LLM inference | Runtime | 🔵 Production |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 109 | Langfuse | LLM traces, prompts, evaluations, experiments, and cost | Platform | 🟢 Now |
| 110 | Phoenix | LLM tracing and evaluation | Platform | 🟢 Now |
| 111 | Promptfoo | Prompt/model tests and red teaming | Tool | 🟢 Now |
| 112 | Ragas | RAG evaluation metrics and test sets | Framework | 🟢 Now |
| 113 | DeepEval | LLM testing and evaluation | Framework | 🟡 Evaluate |
| 114 | Evidently | Data quality, drift, and ML monitoring | Platform | 🟡 Evaluate |
| 115 | TruLens | Evaluate and trace LLM applications | Framework | 🟡 Evaluate |
| 116 | Giskard | Test ML/LLM models for quality and risk | Platform | 🟡 Evaluate |
| 117 | OpenLLMetry | OpenTelemetry instrumentation for LLM apps | Framework | 🟢 Now |
| 118 | Agenta | Prompt experimentation, evaluation, and tracing | Platform | 🟡 Evaluate |
| 119 | Helicone | LLM observability, logging, and cost tracking | Platform | 🟡 Evaluate |
| 120 | NeMo Guardrails | Programmable conversational guardrails | Framework | 🟡 Evaluate |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 121 | Pipecat | Real-time voice and multimodal conversational agents | Framework | 🟢 Now |
| 122 | LiveKit Agents | Real-time voice and multimodal agents over WebRTC | Platform | 🟡 Evaluate |
| 123 | Vocode | Voice-based conversational applications | Framework | 🟡 Evaluate |
| 124 | Whisper | Speech-to-text transcription | Model/tool | 🟢 Now |
| 125 | faster-whisper | Efficient Whisper inference | Model/tool | 🟢 Now |
| 126 | Vosk | Offline speech recognition | Runtime | 🟡 Evaluate |
| 127 | Piper | Local text-to-speech | Runtime | 🟡 Evaluate |
| 128 | Silero | Speech and voice-activity models | Model/tool | 🟡 Evaluate |
| 129 | Coqui TTS | Text-to-speech and voice-model tooling | Framework | 🟡 Evaluate |
| 130 | OpenVoice | Voice cloning and controllable speech | Model/tool | 🟡 Evaluate |
| 131 | Kokoro | Lightweight open text-to-speech model | Model/tool | 🟡 Evaluate |
| 132 | SpeechBrain | Speech and audio research toolkit | Framework | 🟡 Evaluate |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 133 | PaddleOCR | OCR and document extraction | Toolkit | 🟢 Now |
| 134 | Tesseract | Open-source OCR engine | Engine | 🟢 Now |
| 135 | docTR | Deep-learning document text recognition | Framework | 🟡 Evaluate |
| 136 | Surya | OCR, layout analysis, and reading order | Toolkit | 🟡 Evaluate |
| 137 | Marker | Convert PDFs and documents to structured Markdown | Toolkit | 🟢 Now |
| 138 | Unstructured | Document parsing and ingestion pipelines | Platform | 🟢 Now |
| 139 | LayoutLM | Document understanding models | Framework/model | 🟡 Evaluate |
| 140 | Detectron2 | Computer-vision detection and segmentation | Framework | 🟡 Evaluate |
| 141 | Ultralytics YOLO | Real-time object detection and vision models | Framework | 🟢 Now |
| 142 | OpenCV | Computer vision and image processing | Framework | 🟢 Now |
| 143 | MMDetection | OpenMMLab detection toolbox | Framework | 🟡 Evaluate |
| 144 | GroundingDINO | Open-vocabulary object detection | Model/tool | 🟡 Evaluate |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 145 | SDV | Synthetic tabular, relational, and time-series data | Framework | 🟡 Evaluate |
| 146 | Synthea | Synthetic patient data generation | Application | 🟡 Evaluate |
| 147 | ydata-synthetic | Synthetic data generation methods | Framework | 🟡 Evaluate |
| 148 | Synthcity | Synthetic data and privacy research | Framework | 🟡 Evaluate |
| 149 | Twinify | Privacy-preserving synthetic data | Framework | 🟡 Evaluate |
| 150 | SmartNoise | Differential privacy tooling | Platform | 🟣 Later |
| 151 | OpenDP | Differential privacy framework | Framework | 🟣 Later |
| 152 | Microsoft Presidio | PII detection and anonymization | Platform | 🟢 Now |
| 153 | garak | LLM vulnerability scanning | Tool | 🟢 Now |
| 154 | LLM Guard | Input/output scanners for LLM security | Framework | 🟡 Evaluate |
| 155 | Guardrails AI | Validation and safety guardrails | Framework | 🟡 Evaluate |
| 156 | DeepTeam | LLM red teaming and safety testing | Tool | 🟡 Evaluate |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 157 | Docker | Reproducible packaging for models and services | Infrastructure | 🟢 Now |
| 158 | Podman | Daemonless OCI container workflows | Infrastructure | 🟡 Evaluate |
| 159 | GitLab | Git hosting, CI/CD, registry, and security pipelines | Platform | 🟢 Now |
| 160 | Forgejo | Lightweight self-hosted Git forge | Platform | 🟢 Now |
| 161 | Woodpecker CI | Open-source container-based CI/CD | Platform | 🟢 Now |
| 162 | Jenkins | Extensible build, test, and release automation | Platform | 🟡 Evaluate |
| 163 | Argo CD | GitOps continuous delivery for Kubernetes | Platform | 🟣 Later |
| 164 | K3s | Lightweight Kubernetes for cloud VMs and edge | Platform | 🟣 Later |
| 165 | OpenTofu | Infrastructure as code for AWS, Azure, and GCP | Infrastructure | 🟢 Now |
| 166 | MinIO | S3-compatible model and dataset object storage | Infrastructure | 🟡 Evaluate |
| 167 | JupyterHub | Multi-user notebook environments | Platform | 🟢 Now (teams) |
| 168 | NVIDIA GPU Operator | GPU drivers and workloads in Kubernetes | Platform | 🟣 Later |
| # | Tool | What it helps with | Type | Stage |
|---|---|---|---|---|
| 169 | Supabase | Application database, Auth, APIs, Storage, and Realtime | Platform | 🟢 Now |
| 170 | Appwrite | Backend services for Auth, Storage, Functions, and Realtime | Platform | 🟢 Now (alternative) |
| 171 | LiteLLM | Unified model gateway, routing, budgets, and fallbacks | Platform | 🟢 Now |
| 172 | n8n | Visual integration and AI workflow automation | Platform | 🟢 Now |
| 173 | Novu | Notification workflows for AI products and SaaS | Platform | 🟢 Now |
| 174 | PostHog | Product analytics, experiments, and feature flags | Platform | 🟢 Now |
| 175 | Chatwoot | Customer support and omnichannel conversations | Platform | 🟢 Now |
| 176 | Libredesk | Lightweight self-hosted support desk | Application | 🟢 Now |
| 177 | listmonk | Self-hosted newsletters and mailing lists | Application | 🟢 Now |
| 178 | Formbricks | Surveys, feedback, and product research | Platform | 🟢 Now |
| 179 | Appsmith | Internal tools and admin panels | Platform | 🟢 Now |
| 180 | OpenTelemetry | Vendor-neutral traces, metrics, and logs | Standard/tooling | 🟢 Now |
Start with Open WebUI, Dify, or a small LlamaIndex/Haystack service; use pgvector or Qdrant; use Ollama or a hosted model; record traces with Langfuse; and create a small Promptfoo or Ragas evaluation set. This is enough for a useful document assistant without Kubeflow or a large agent platform.
Use Libredesk or Chatwoot for the human support surface, Rasa/Dify/LangChain for the assistant, a retrieval system for product documentation, a model gateway for routing, and explicit escalation to a human. Store tenant, user, conversation, tool, and consent metadata. Do not let a support agent issue refunds or change account data without server-side authorization and confirmation.
Use LiveKit or Pipecat for realtime transport, Whisper or faster-whisper for speech-to-text, an LLM gateway for reasoning, Piper or a hosted TTS provider for speech, and Langfuse/OpenTelemetry for tracing. Add interruption handling, timeouts, call recording policy, and human handoff before calling it production-ready.
Use pandas or DuckDB for data preparation, scikit-learn/XGBoost/LightGBM/CatBoost for a baseline, DVC for data and model versioning, MLflow for experiments, Great Expectations for data checks, and BentoML or a small API for serving. This path is usually cheaper and easier to operate than fine-tuning a large language model.
Use Transformers, Datasets, PEFT, Unsloth/Axolotl or torchtune, DVC, MLflow, a held-out evaluation set, and a GPU runner. Confirm data rights and model-license compatibility before training or redistributing the result. Serve with vLLM, BentoML, SGLang, or another inference server only after measuring the workload.
Use Git, DVC, object storage, MLflow, a CI system, containerized training, OpenTelemetry, a model server, and a rollback procedure. Add Kubeflow, Flyte, KServe, Feast, K3s, GPU Operator, or a feature store only when reproducibility, multi-user scheduling, or scale demands them.
A practical team loop is:
- Store application code, prompts, evaluation cases, and pipeline definitions in Git.
- Store large datasets and model artifacts in DVC-backed object storage rather than ordinary Git history.
- Run formatting, unit tests, data checks, prompt tests, security scans, and a small evaluation suite in CI.
- Record the commit, dataset version, model version, configuration, hardware, metrics, latency, and cost for each meaningful run.
- Build a versioned container and deploy it to a development environment.
- Review traces, failures, user feedback, and evaluation regressions.
- Promote through a feature flag or approval gate, then keep rollback artifacts available.
This loop is often more valuable than starting with a large MLOps platform.
A self-hosted AI system can move to AWS or Azure when its containers, artifacts, data, secrets, and operational configuration are portable. The migration includes more than application images.
| Concern | Portable Boundary | AWS Examples | Azure Examples |
|---|---|---|---|
| Model API | ModelGateway or InferenceService |
ECS/EKS, EC2 GPU, Batch, SageMaker endpoint | Container Apps/AKS, GPU VM, Batch, Azure ML endpoint |
| Dataset/model artifacts | DVC plus ArtifactStore |
S3, EFS, FSx | Blob Storage, Azure Files |
| Experiment tracking | MLflow API and artifact store | S3/RDS/EC2 | Blob/PostgreSQL/VM |
| Vector search | Qdrant/pgvector/OpenSearch interface | EC2/EKS/OpenSearch | AKS/VM/Azure AI Search adapter |
| Training | Containerized entrypoint | EC2 GPU, Batch, EKS | GPU VM, Batch, AKS |
| Secrets | Environment-independent secret interface | Secrets Manager/SSM | Key Vault |
| CI/CD | Git and container pipeline | GitLab Runner, ECS/EKS | GitLab Runner, Container Apps/AKS |
| Observability | OpenTelemetry | CloudWatch or self-hosted backends | Azure Monitor or self-hosted backends |
| Infrastructure | OpenTofu modules | AWS provider | Azure provider |
Keep provider-specific SDKs at the infrastructure edge. Do not spread AWS or Azure types through product, model, or training modules. Model weights, prompts, evaluation data, vector indexes, and secrets all need an export and restore plan.
A model or agent is not production-ready merely because it returns good answers in a notebook. Establish:
- Data provenance and permission checks
- Tenant isolation
- Input validation and output validation
- Rate limits, timeouts, and cost budgets
- Prompt-injection defenses
- Human escalation paths
- Structured logs, traces, and metrics
- Backup procedures and rollback artifacts
Additional checks by domain:
- Agents — restrict tool permissions and validate every tool argument.
- RAG — retain source references and measure retrieval quality.
- Voice systems — test interruption, silence, accents, latency, recordings, and handoff.
- Classical ML — monitor drift and label quality.
- All AI systems — define what happens when the model is unavailable or wrong.
⚠️ Review the current license and commercial terms of every tool and every model before using it in a customer-facing SaaS, agency template, managed service, or white-label product.
Contributions welcome! Please read the contribution guidelines first. This list follows the awesome-lint format — one PR per addition, alphabetized within a category where practical, no self-promotion without disclosure.
To the extent possible under law, the contributors have waived all copyright and related or neighboring rights to this work. See LICENSE for details.
