A living encyclopedia of observability, logging & monitoring tools — independently surveyed, benchmark-ready, and continuously updated.
Authored by Team Ardur · CC BY 4.0
| Tool | Type | License | Tier | Notes |
|---|---|---|---|---|
| AgentOps | Agent Observability | Partially open | A | Agent session lifecycles; tool calls; replay; CrewAI/AutoGen native |
| Arize AX | Enterprise AI Monitoring | Proprietary | A | Production-scale telemetry; 50+ research-backed eval metrics |
| Arize Phoenix | AI Observability + Evaluation | Elastic-2.0 | A | 2.5M+ monthly downloads; RAG debugging; local-first OTel |
| Arthur AI | Model Performance + Governance | Proprietary | A | $60M raised; risk, bias, drift; compliance dashboards |
| Braintrust | Eval-first Observability | Proprietary | A | $124M raised; Notion/Stripe/Vercel; CI/CD quality gates |
| Datadog LLM Observability | APM Extension | Proprietary | A | Unified infra + LLM tracing; agentless; OTel GenAI support |
| DeepEval | LLM Evaluation Framework | Apache-2.0 | A | 50+ metrics; CI-native; pytest integration; RAG/agent/multi-turn |
| Fiddler | AI Governance + Monitoring | Proprietary | A | $100M raised; explainability, bias, compliance; ML+LLM unified |
| Galileo | AI Evaluation + Guardrails | Proprietary | A | $68M raised; Luna-2 evaluators; sub-200ms scoring; runtime guardrails |
| Helicone | LLM Gateway + Observability | Apache-2.0 (partial) | A | YC W23; proxy-based; one-line integration; 100+ models |
| Langfuse | LLM Engineering Platform | MIT (core) | A | 21K+ stars; self-hosted; acquired by ClickHouse Jan 2026 |
| LangSmith | Observability + Evaluation | Proprietary | A | 1B+ events/day; ~35% Fortune 500; LangChain native |
| LiteLLM | LLM Gateway + Proxy | Open source | A | 18K+ stars; 100+ LLM APIs in OpenAI format; YC W23 |
| OpenLIT | AI Engineering Platform | Apache-2.0 | A | LLM + GPU + VectorDB observability; 60+ integrations; self-hosted |
| Portkey | AI Gateway + Observability | Partially open | A | 1T tokens/day; 250+ models; 20-40ms latency; MCP Gateway |
| Promptfoo | Eval + Red Teaming | MIT | A | YAML-based assertions; 50+ vulnerability types; GitHub Actions; always free OSS |
| Pydantic Logfire | AI-native Observability | Proprietary | A | 10M free spans/mo; $2/M after; OTel-native; SQL-queryable |
| Ragas | RAG Evaluation | Open source | A | Faithfulness, context precision/recall; research-backed; pairs with Phoenix |
| TruLens | RAG Evaluation Framework | MIT | A | RAG Triad: groundedness, relevance, correctness; OTel traces |
| Weights & Biases Weave | LLM Tracing + Eval | Proprietary | A | ML + LLM unified; $50/user/mo; auto-instruments major frameworks |
| Comet Opik | ML + LLM Observability | Apache-2.0 (Opik OSS) | B | Comet heritage; PyTest CI; guardrails; PII detection |
| Dynatrace | AIOps + AI Observability | Proprietary | B | AI-powered root cause analysis; full-stack monitoring |
| Fastn | MCP Gateway + Observability | Proprietary | B | Managed MCP gateway; 1,000+ integrations; Adaptive Context Layer |
| Future AGI | Unified Eval + Observability | Apache-2.0 | B | traceAI auto-instrumentation; 100+ metrics; multimodal |
| Grafana Cloud | AI Observability | Open core + SaaS | B | OpenLIT integration; 5 prebuilt dashboards for GenAI |
| Guardrails AI | GenAI Reliability | Proprietary | B | $8M raised; runtime guardrails; structured output validation |
| Honeycomb | Observability (OTel-native) | Proprietary | B | Distributed tracing; high-cardinality; GenAI semantic conventions |
| Kong AI Gateway | API Gateway + AI | Open core + SaaS | B | Multi-LLM routing; semantic caching; rate limiting |
| Lunary | LLM Analytics + Observability | Open source / SaaS | B | Prompt management, PII masking, agent tracing; EU data center |
| Maxim AI | Full Lifecycle Platform | Proprietary | B | Simulation + eval + observability; Bifrost open-source LLM gateway (Go) |
| Netdata | Real-Time GPU Monitoring | Open source + SaaS | B | Per-second monitoring; NVIDIA/Intel GPU; DCGM collector; MCP server |
| New Relic | APM + LLM Observability | Proprietary | B | OpenLIT integration; pre-built dashboards; OTLP endpoint |
| NVIDIA DCGM | GPU Monitoring | Open source (NVIDIA) | B | GPU health, utilization, performance; Prometheus exporter |
| OpenRouter | LLM Routing + Proxy | Open source | B | Standardized API for 100+ open/commercial models |
| Splunk (Cisco) | Infrastructure + AI Observability | Proprietary | B | Cisco acquisition 2025; GPU-to-application monitoring; AI-ready PODs |
| TrueFoundry | Full-Stack AI Infra | Proprietary | B | AI Gateway, token-level cost tracking, FinOps; hybrid/on-prem |
| Vercel AI Gateway | Serverless AI Gateway | Proprietary | B | Unified API for AI providers; serverless edge deployment |
| AI Score | AI Governance (Europe) | Proprietary | C | $1M raised; enterprise visibility, compliance monitoring, risk scoring |
| Aigentsphere | Agent Governance Infra | Proprietary | C | $4M raised; monitoring, policy enforcement, compliance reporting |
| AIM Intelligence | AI Red Teaming + Guardrails | Proprietary | C | APAC; $7M Series A; automated red teaming, real-time safety |
| Alinia | AI Compliance (Europe) | Proprietary | C | $7.5M raised; guardrails API, auditing, policy enforcement |
| Amazon SageMaker | AWS-native ML/LLM Ops | Proprietary | C | Model monitoring; batch/real-time eval jobs; managed infrastructure |
| Aporia | AI Observability + Guardrails | Acquired | C | $30M raised; ML monitoring, guardrails, drift detection; acquired |
| Arize AI | ML + LLM Monitoring | Mixed | C | $131M total raised; unified ML and LLM observability; drift detection |
| Atla AI | AI Agent Evaluation | Proprietary | C | $6M raised; agent evaluation insights, scoring |
| Aveni | AI Assurance (Financial) | Proprietary | C | Europe ($16.2M); AI-agent conduct risk assessment; financial oversight |
| Azure Machine Learning | Azure-native AI Platform | Proprietary | C | Dataset-driven evaluations; governance; access controls |
| CalypsoAI | Enterprise AI Inference Security | Acquired | C | $41M raised; inference security, model validation; acquired |
| Capsule Security | Runtime Trust Layer | Proprietary | C | $7M raised; agent monitoring, control, manipulation prevention |
| Ciphero | AI Verification Layer | Proprietary | C | $2.5M raised; AI interaction capture, verification, governance |
| Complyance | AI-native GRC | Proprietary | C | $20M raised; governance, risk, compliance automation |
| Confident AI | Eval Platform (DeepEval) | Proprietary + OSS | C | 50+ metrics; cross-functional workflows; red teaming |
| Coxwave | AI Trust + Verification | Proprietary | C | APAC; $4.8M raised Jan 2026; agent verification, reliability, governance |
| Cranium | AI Governance + Security | Proprietary | C | $32M raised; AI security, compliance, risk management |
| Credo AI | AI Governance Platform | Proprietary | C | $41M raised; governance, risk, compliance workflows; EU AI Act |
| Darwin AI | Public Sector AI Governance | Proprietary | C | $15M raised; government AI adoption, transparency, compliance |
| Databricks MLflow | Unified ML + LLM Lifecycle | Open source (MLflow) | C | Experiment tracking; model registry; GenAI tracing; Spark scalability |
| Deepchecks | ML Validation + Monitoring | Acquired | C | $14M raised; ML model validation, data drift; acquired |
| DeepKeep | Enterprise AI Security | Proprietary | C | $10M raised; AI model security, runtime protection |
| Deeploy | Responsible AI Oversight | Proprietary | C | $9M raised; AI oversight, explainability, compliance |
| Distributional | AI Testing Platform | Proprietary | C | $30M raised; enterprise AI testing, statistical validation |
| Geordie AI | AI Governance + Observability | Proprietary | C | Europe ($30M); agent posture, observability, compliance, controls |
| Giskard | AI Quality Testing | Open source | C | Bias detection, robustness testing, explainability |
| Google ADK | Google Agent Development | Open source (SDK) | C | Structured agent eval hooks; tool/action tracing; GCP integration |
| Harmonic Security | AI Data Protection | Proprietary | C | $24M raised; data leakage prevention, shadow AI visibility |
| HiddenLayer | AI Model Security | Proprietary | C | $56M raised; AI model security platform, runtime protection |
| Iridius | Compliance-by-Design | Proprietary | C | $8.6M raised; regulated enterprise workflows, validation |
| JetStream Security | AI Governance + Control | Proprietary | C | $34M raised (Seed); enterprise AI visibility, risk controls, blueprints |
| Kolena | AI Model Testing | Proprietary | C | $21M raised; model testing, benchmarking, quality assurance |
| Lakera | GenAI Application Security | Acquired | C | $30M raised; real-time guardrails, PII protection; acquired by Check Point |
| Laminar | AI Agent Observability | Open source | C | YC 2026; trace workflows; replay agent runs; anomaly detection |
| Lasso Security | LLM Cybersecurity | Proprietary | C | $6M raised; LLM security platform, data protection |
| LatticeFlow | AI Model Quality | Proprietary | C | $15M raised; model quality, robustness, fairness testing |
| Middleware | Full-Stack Cloud Observability | Proprietary | C | YC W2023; AI-based anomaly detection; GPT-4 error resolution |
| Mindgard | AI Security Testing | Proprietary | C | $12M raised; adversarial testing, model security |
| Modulos | AI Governance Compliance | Proprietary | C | $11M raised; compliance-by-design for AI systems |
| Mona | AI Model Monitoring | Proprietary | C | $7M raised; custom metric monitoring for AI models |
| Nexos.ai | AI Orchestration + Governance | Proprietary | C | Europe ($35M); secure adoption, observability, cost control |
| Numalis | Formal AI Validation | Proprietary | C | $6M raised; mathematical validation, formal verification |
| NVIDIA NeMo Guardrails | Runtime Guardrails | Open source (NVIDIA) | C | Rule-and-rail framework; fact-checking; grounding |
| Portal26 | GenAI Adoption Governance | Proprietary | C | $15M raised; shadow AI visibility, responsible adoption |
| Prompt Security | GenAI Runtime Security | Acquired | C | $23M raised; runtime GenAI security, prompt injection; acquired |
| Protect AI | ML/ML Security | Acquired | C | $108M raised; AI/ML system security, model scanning; acquired by Palo Alto |
| RAGAS | RAG Evaluation Metrics | Open source | C | Faithfulness, context precision, context recall; research-backed |
| Robust Intelligence | AI Model Security | Acquired | C | $44M raised; AI model security firewall; acquired by Cisco |
| Seldon | MLOps + Governance | Proprietary | C | $33M raised; model governance, explainability, monitoring |
| Singulr AI | AI Governance Security | Proprietary | C | $10M raised; agent security, governance, observability |
| Traceloop / OpenLLMetry | OTel-based LLM Tracing | Open source | C | OpenTelemetry-native LLM application monitoring |
| TrojAI | AI Attack Protection | Proprietary | C | $9M raised; adversarial attack detection, model hardening |
| Trustible | AI Governance | Proprietary | C | $6M raised; enterprise AI governance, compliance |
| Tynapse | Runtime Agent Security | Proprietary | C | APAC; $3.17M seed; blocking hallucinations, jailbreaks, data leakage |
| Vectara HHEM-2.1 | Hallucination Detection | Open weights | C | Cross-encoder hallucination model; RAG groundedness |
| Vijil | AI Agent Resilience | Proprietary | C | $23M raised; reliability testing, security, governance |
| White Circle | AI Control Platform | Proprietary | C | Europe ($11M); monitor, protect, test, improve AI models |
| WitnessAI | AI Security + Governance | Proprietary | C | $58M raised (Series B); runtime security, guardrails, policy enforcement |
| Xenos Labs | LLM Cost/Latency Monitoring | Proprietary | C | Japan (Antler-backed); Y240M pre-seed; LLM cost/latency/quality |
| Zania | AI Compliance (GRC) | Proprietary | C | $18M raised; agentic AI for governance, risk, compliance |
- Atla AI (Tier C)
- Laminar (Tier C)
- Vijil (Tier C)
- Aveni (Tier C)
- TrojAI (Tier C)
- Alinia (Tier C)
- Zania (Tier C)
- White Circle (Tier C)
- Harmonic Security (Tier C)
- OpenLIT (Tier A)
- Galileo (Tier A)
- Portkey (Tier A)
- Trustible (Tier C)
- AI Score (Tier C)
- JetStream Security (Tier C)
- Fiddler (Tier A)
- Geordie AI (Tier C)
- Cranium (Tier C)
- Modulos (Tier C)
- Credo AI (Tier C)
- Singulr AI (Tier C)
- Mona (Tier C)
- LatticeFlow (Tier C)
- HiddenLayer (Tier C)
- Robust Intelligence (Tier C)
- Kolena (Tier C)
- Grafana Cloud (Tier B)
- Arize Phoenix (Tier A)
- Aporia (Tier C)
- Nexos.ai (Tier C)
- Giskard (Tier C)
- AIM Intelligence (Tier C)
- WitnessAI (Tier C)
- Mindgard (Tier C)
- Distributional (Tier C)
- Coxwave (Tier C)
- Ciphero (Tier C)
- Complyance (Tier C)
- Pydantic Logfire (Tier A)
- Dynatrace (Tier B)
- Kong AI Gateway (Tier B)
- New Relic (Tier B)
- Datadog LLM Observability (Tier A)
- Amazon SageMaker (Tier C)
- Aigentsphere (Tier C)
- AgentOps (Tier A)
- Azure Machine Learning (Tier C)
- Iridius (Tier C)
- CalypsoAI (Tier C)
- Arize AX (Tier A)
- DeepKeep (Tier C)
- Promptfoo (Tier A)
- Confident AI (Tier C)
- Braintrust (Tier A)
- Numalis (Tier C)
- Maxim AI (Tier B)
- TrueFoundry (Tier B)
- Middleware (Tier C)
- NVIDIA DCGM (Tier B)
- Portal26 (Tier C)
- Lakera (Tier C)
- Guardrails AI (Tier B)
- Prompt Security (Tier C)
- Google ADK (Tier C)
- Vectara HHEM-2.1 (Tier C)
- Splunk (Cisco) (Tier B)
- Lunary (Tier B)
- Xenos Labs (Tier C)
- Lasso Security (Tier C)
- Langfuse (Tier A)
- DeepEval (Tier A)
- Helicone (Tier A)
- LiteLLM (Tier A)
- OpenRouter (Tier B)
- Weights & Biases Weave (Tier A)
- Fastn (Tier B)
- Arize AI (Tier C)
- Comet Opik (Tier B)
- Deepchecks (Tier C)
- Protect AI (Tier C)
- Seldon (Tier C)
- Arthur AI (Tier A)
- Traceloop / OpenLLMetry (Tier C)
- Honeycomb (Tier B)
- LangSmith (Tier A)
- Darwin AI (Tier C)
- Ragas (Tier A)
- TruLens (Tier A)
- RAGAS (Tier C)
- Netdata (Tier B)
- Deeploy (Tier C)
- Tynapse (Tier C)
- NVIDIA NeMo Guardrails (Tier C)
- Capsule Security (Tier C)
- Vercel AI Gateway (Tier B)
- Future AGI (Tier B)
- Databricks MLflow (Tier C)
- Browse tools — Start with the Tier A tools above; they are the most production-ready.
- Read the methodology — See
methodology/for how tools are evaluated and benchmarked. - Check editions — Quarterly editions provide curated analysis and trend reports.
- Contribute — See
CONTRIBUTING.mdfor how to add or update tools.
Latest: 2026-06.md
Tools are evaluated across multiple dimensions including accuracy, latency, token economics, scale behavior, ops burden, developer experience, and data sovereignty. See methodology/ for details.
See CONTRIBUTING.md for guidelines on adding tools, updating analyses, and submitting benchmarks.
Content: CC BY 4.0 — Authored by Team Ardur Code/Benchmarks: As noted in individual files