Skip to content

Repository files navigation

AI Observability Almanac

A living encyclopedia of observability, logging & monitoring tools — independently surveyed, benchmark-ready, and continuously updated.

Authored by Team Ardur · CC BY 4.0


Tools (97 total)

Tool Type License Tier Notes
AgentOps Agent Observability Partially open A Agent session lifecycles; tool calls; replay; CrewAI/AutoGen native
Arize AX Enterprise AI Monitoring Proprietary A Production-scale telemetry; 50+ research-backed eval metrics
Arize Phoenix AI Observability + Evaluation Elastic-2.0 A 2.5M+ monthly downloads; RAG debugging; local-first OTel
Arthur AI Model Performance + Governance Proprietary A $60M raised; risk, bias, drift; compliance dashboards
Braintrust Eval-first Observability Proprietary A $124M raised; Notion/Stripe/Vercel; CI/CD quality gates
Datadog LLM Observability APM Extension Proprietary A Unified infra + LLM tracing; agentless; OTel GenAI support
DeepEval LLM Evaluation Framework Apache-2.0 A 50+ metrics; CI-native; pytest integration; RAG/agent/multi-turn
Fiddler AI Governance + Monitoring Proprietary A $100M raised; explainability, bias, compliance; ML+LLM unified
Galileo AI Evaluation + Guardrails Proprietary A $68M raised; Luna-2 evaluators; sub-200ms scoring; runtime guardrails
Helicone LLM Gateway + Observability Apache-2.0 (partial) A YC W23; proxy-based; one-line integration; 100+ models
Langfuse LLM Engineering Platform MIT (core) A 21K+ stars; self-hosted; acquired by ClickHouse Jan 2026
LangSmith Observability + Evaluation Proprietary A 1B+ events/day; ~35% Fortune 500; LangChain native
LiteLLM LLM Gateway + Proxy Open source A 18K+ stars; 100+ LLM APIs in OpenAI format; YC W23
OpenLIT AI Engineering Platform Apache-2.0 A LLM + GPU + VectorDB observability; 60+ integrations; self-hosted
Portkey AI Gateway + Observability Partially open A 1T tokens/day; 250+ models; 20-40ms latency; MCP Gateway
Promptfoo Eval + Red Teaming MIT A YAML-based assertions; 50+ vulnerability types; GitHub Actions; always free OSS
Pydantic Logfire AI-native Observability Proprietary A 10M free spans/mo; $2/M after; OTel-native; SQL-queryable
Ragas RAG Evaluation Open source A Faithfulness, context precision/recall; research-backed; pairs with Phoenix
TruLens RAG Evaluation Framework MIT A RAG Triad: groundedness, relevance, correctness; OTel traces
Weights & Biases Weave LLM Tracing + Eval Proprietary A ML + LLM unified; $50/user/mo; auto-instruments major frameworks
Comet Opik ML + LLM Observability Apache-2.0 (Opik OSS) B Comet heritage; PyTest CI; guardrails; PII detection
Dynatrace AIOps + AI Observability Proprietary B AI-powered root cause analysis; full-stack monitoring
Fastn MCP Gateway + Observability Proprietary B Managed MCP gateway; 1,000+ integrations; Adaptive Context Layer
Future AGI Unified Eval + Observability Apache-2.0 B traceAI auto-instrumentation; 100+ metrics; multimodal
Grafana Cloud AI Observability Open core + SaaS B OpenLIT integration; 5 prebuilt dashboards for GenAI
Guardrails AI GenAI Reliability Proprietary B $8M raised; runtime guardrails; structured output validation
Honeycomb Observability (OTel-native) Proprietary B Distributed tracing; high-cardinality; GenAI semantic conventions
Kong AI Gateway API Gateway + AI Open core + SaaS B Multi-LLM routing; semantic caching; rate limiting
Lunary LLM Analytics + Observability Open source / SaaS B Prompt management, PII masking, agent tracing; EU data center
Maxim AI Full Lifecycle Platform Proprietary B Simulation + eval + observability; Bifrost open-source LLM gateway (Go)
Netdata Real-Time GPU Monitoring Open source + SaaS B Per-second monitoring; NVIDIA/Intel GPU; DCGM collector; MCP server
New Relic APM + LLM Observability Proprietary B OpenLIT integration; pre-built dashboards; OTLP endpoint
NVIDIA DCGM GPU Monitoring Open source (NVIDIA) B GPU health, utilization, performance; Prometheus exporter
OpenRouter LLM Routing + Proxy Open source B Standardized API for 100+ open/commercial models
Splunk (Cisco) Infrastructure + AI Observability Proprietary B Cisco acquisition 2025; GPU-to-application monitoring; AI-ready PODs
TrueFoundry Full-Stack AI Infra Proprietary B AI Gateway, token-level cost tracking, FinOps; hybrid/on-prem
Vercel AI Gateway Serverless AI Gateway Proprietary B Unified API for AI providers; serverless edge deployment
AI Score AI Governance (Europe) Proprietary C $1M raised; enterprise visibility, compliance monitoring, risk scoring
Aigentsphere Agent Governance Infra Proprietary C $4M raised; monitoring, policy enforcement, compliance reporting
AIM Intelligence AI Red Teaming + Guardrails Proprietary C APAC; $7M Series A; automated red teaming, real-time safety
Alinia AI Compliance (Europe) Proprietary C $7.5M raised; guardrails API, auditing, policy enforcement
Amazon SageMaker AWS-native ML/LLM Ops Proprietary C Model monitoring; batch/real-time eval jobs; managed infrastructure
Aporia AI Observability + Guardrails Acquired C $30M raised; ML monitoring, guardrails, drift detection; acquired
Arize AI ML + LLM Monitoring Mixed C $131M total raised; unified ML and LLM observability; drift detection
Atla AI AI Agent Evaluation Proprietary C $6M raised; agent evaluation insights, scoring
Aveni AI Assurance (Financial) Proprietary C Europe ($16.2M); AI-agent conduct risk assessment; financial oversight
Azure Machine Learning Azure-native AI Platform Proprietary C Dataset-driven evaluations; governance; access controls
CalypsoAI Enterprise AI Inference Security Acquired C $41M raised; inference security, model validation; acquired
Capsule Security Runtime Trust Layer Proprietary C $7M raised; agent monitoring, control, manipulation prevention
Ciphero AI Verification Layer Proprietary C $2.5M raised; AI interaction capture, verification, governance
Complyance AI-native GRC Proprietary C $20M raised; governance, risk, compliance automation
Confident AI Eval Platform (DeepEval) Proprietary + OSS C 50+ metrics; cross-functional workflows; red teaming
Coxwave AI Trust + Verification Proprietary C APAC; $4.8M raised Jan 2026; agent verification, reliability, governance
Cranium AI Governance + Security Proprietary C $32M raised; AI security, compliance, risk management
Credo AI AI Governance Platform Proprietary C $41M raised; governance, risk, compliance workflows; EU AI Act
Darwin AI Public Sector AI Governance Proprietary C $15M raised; government AI adoption, transparency, compliance
Databricks MLflow Unified ML + LLM Lifecycle Open source (MLflow) C Experiment tracking; model registry; GenAI tracing; Spark scalability
Deepchecks ML Validation + Monitoring Acquired C $14M raised; ML model validation, data drift; acquired
DeepKeep Enterprise AI Security Proprietary C $10M raised; AI model security, runtime protection
Deeploy Responsible AI Oversight Proprietary C $9M raised; AI oversight, explainability, compliance
Distributional AI Testing Platform Proprietary C $30M raised; enterprise AI testing, statistical validation
Geordie AI AI Governance + Observability Proprietary C Europe ($30M); agent posture, observability, compliance, controls
Giskard AI Quality Testing Open source C Bias detection, robustness testing, explainability
Google ADK Google Agent Development Open source (SDK) C Structured agent eval hooks; tool/action tracing; GCP integration
Harmonic Security AI Data Protection Proprietary C $24M raised; data leakage prevention, shadow AI visibility
HiddenLayer AI Model Security Proprietary C $56M raised; AI model security platform, runtime protection
Iridius Compliance-by-Design Proprietary C $8.6M raised; regulated enterprise workflows, validation
JetStream Security AI Governance + Control Proprietary C $34M raised (Seed); enterprise AI visibility, risk controls, blueprints
Kolena AI Model Testing Proprietary C $21M raised; model testing, benchmarking, quality assurance
Lakera GenAI Application Security Acquired C $30M raised; real-time guardrails, PII protection; acquired by Check Point
Laminar AI Agent Observability Open source C YC 2026; trace workflows; replay agent runs; anomaly detection
Lasso Security LLM Cybersecurity Proprietary C $6M raised; LLM security platform, data protection
LatticeFlow AI Model Quality Proprietary C $15M raised; model quality, robustness, fairness testing
Middleware Full-Stack Cloud Observability Proprietary C YC W2023; AI-based anomaly detection; GPT-4 error resolution
Mindgard AI Security Testing Proprietary C $12M raised; adversarial testing, model security
Modulos AI Governance Compliance Proprietary C $11M raised; compliance-by-design for AI systems
Mona AI Model Monitoring Proprietary C $7M raised; custom metric monitoring for AI models
Nexos.ai AI Orchestration + Governance Proprietary C Europe ($35M); secure adoption, observability, cost control
Numalis Formal AI Validation Proprietary C $6M raised; mathematical validation, formal verification
NVIDIA NeMo Guardrails Runtime Guardrails Open source (NVIDIA) C Rule-and-rail framework; fact-checking; grounding
Portal26 GenAI Adoption Governance Proprietary C $15M raised; shadow AI visibility, responsible adoption
Prompt Security GenAI Runtime Security Acquired C $23M raised; runtime GenAI security, prompt injection; acquired
Protect AI ML/ML Security Acquired C $108M raised; AI/ML system security, model scanning; acquired by Palo Alto
RAGAS RAG Evaluation Metrics Open source C Faithfulness, context precision, context recall; research-backed
Robust Intelligence AI Model Security Acquired C $44M raised; AI model security firewall; acquired by Cisco
Seldon MLOps + Governance Proprietary C $33M raised; model governance, explainability, monitoring
Singulr AI AI Governance Security Proprietary C $10M raised; agent security, governance, observability
Traceloop / OpenLLMetry OTel-based LLM Tracing Open source C OpenTelemetry-native LLM application monitoring
TrojAI AI Attack Protection Proprietary C $9M raised; adversarial attack detection, model hardening
Trustible AI Governance Proprietary C $6M raised; enterprise AI governance, compliance
Tynapse Runtime Agent Security Proprietary C APAC; $3.17M seed; blocking hallucinations, jailbreaks, data leakage
Vectara HHEM-2.1 Hallucination Detection Open weights C Cross-encoder hallucination model; RAG groundedness
Vijil AI Agent Resilience Proprietary C $23M raised; reliability testing, security, governance
White Circle AI Control Platform Proprietary C Europe ($11M); monitor, protect, test, improve AI models
WitnessAI AI Security + Governance Proprietary C $58M raised (Series B); runtime security, guardrails, policy enforcement
Xenos Labs LLM Cost/Latency Monitoring Proprietary C Japan (Antler-backed); Y240M pre-seed; LLM cost/latency/quality
Zania AI Compliance (GRC) Proprietary C $18M raised; agentic AI for governance, risk, compliance

Tool Categories

AI Agent Evaluation (1 tools)

AI Agent Observability (1 tools)

AI Agent Resilience (1 tools)

AI Assurance (Financial) (1 tools)

AI Attack Protection (1 tools)

AI Compliance (Europe) (1 tools)

AI Compliance (GRC) (1 tools)

AI Control Platform (1 tools)

AI Data Protection (1 tools)

AI Engineering Platform (1 tools)

AI Evaluation + Guardrails (1 tools)

AI Gateway + Observability (1 tools)

AI Governance (1 tools)

AI Governance (Europe) (1 tools)

AI Governance + Control (1 tools)

AI Governance + Monitoring (1 tools)

AI Governance + Observability (1 tools)

AI Governance + Security (1 tools)

AI Governance Compliance (1 tools)

AI Governance Platform (1 tools)

AI Governance Security (1 tools)

AI Model Monitoring (1 tools)

AI Model Quality (1 tools)

AI Model Security (2 tools)

AI Model Testing (1 tools)

AI Observability (1 tools)

AI Observability + Evaluation (1 tools)

AI Observability + Guardrails (1 tools)

AI Orchestration + Governance (1 tools)

AI Quality Testing (1 tools)

AI Red Teaming + Guardrails (1 tools)

AI Security + Governance (1 tools)

AI Security Testing (1 tools)

AI Testing Platform (1 tools)

AI Trust + Verification (1 tools)

AI Verification Layer (1 tools)

AI-native GRC (1 tools)

AI-native Observability (1 tools)

AIOps + AI Observability (1 tools)

API Gateway + AI (1 tools)

APM + LLM Observability (1 tools)

APM Extension (1 tools)

AWS-native ML/LLM Ops (1 tools)

Agent Governance Infra (1 tools)

Agent Observability (1 tools)

Azure-native AI Platform (1 tools)

Compliance-by-Design (1 tools)

Enterprise AI Inference Security (1 tools)

Enterprise AI Monitoring (1 tools)

Enterprise AI Security (1 tools)

Eval + Red Teaming (1 tools)

Eval Platform (DeepEval) (1 tools)

Eval-first Observability (1 tools)

Formal AI Validation (1 tools)

Full Lifecycle Platform (1 tools)

Full-Stack AI Infra (1 tools)

Full-Stack Cloud Observability (1 tools)

GPU Monitoring (1 tools)

GenAI Adoption Governance (1 tools)

GenAI Application Security (1 tools)

GenAI Reliability (1 tools)

GenAI Runtime Security (1 tools)

Google Agent Development (1 tools)

Hallucination Detection (1 tools)

Infrastructure + AI Observability (1 tools)

LLM Analytics + Observability (1 tools)

LLM Cost/Latency Monitoring (1 tools)

LLM Cybersecurity (1 tools)

LLM Engineering Platform (1 tools)

LLM Evaluation Framework (1 tools)

LLM Gateway + Observability (1 tools)

LLM Gateway + Proxy (1 tools)

LLM Routing + Proxy (1 tools)

LLM Tracing + Eval (1 tools)

MCP Gateway + Observability (1 tools)

ML + LLM Monitoring (1 tools)

ML + LLM Observability (1 tools)

ML Validation + Monitoring (1 tools)

ML/ML Security (1 tools)

MLOps + Governance (1 tools)

Model Performance + Governance (1 tools)

OTel-based LLM Tracing (1 tools)

Observability (OTel-native) (1 tools)

Observability + Evaluation (1 tools)

Public Sector AI Governance (1 tools)

RAG Evaluation (1 tools)

RAG Evaluation Framework (1 tools)

RAG Evaluation Metrics (1 tools)

Real-Time GPU Monitoring (1 tools)

Responsible AI Oversight (1 tools)

Runtime Agent Security (1 tools)

Runtime Guardrails (1 tools)

Runtime Trust Layer (1 tools)

Serverless AI Gateway (1 tools)

Unified Eval + Observability (1 tools)

Unified ML + LLM Lifecycle (1 tools)


Quick Start

  1. Browse tools — Start with the Tier A tools above; they are the most production-ready.
  2. Read the methodology — See methodology/ for how tools are evaluated and benchmarked.
  3. Check editions — Quarterly editions provide curated analysis and trend reports.
  4. Contribute — See CONTRIBUTING.md for how to add or update tools.

Recent Edition

Latest: 2026-06.md


Methodology

Tools are evaluated across multiple dimensions including accuracy, latency, token economics, scale behavior, ops burden, developer experience, and data sovereignty. See methodology/ for details.


Contributing

See CONTRIBUTING.md for guidelines on adding tools, updating analyses, and submitting benchmarks.


License

Content: CC BY 4.0 — Authored by Team Ardur Code/Benchmarks: As noted in individual files

About

No description, website, or topics provided.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors