CUGAR Agent is a modular agent stack for enterprise use: a planner and worker pair wired for LangGraph, LlamaIndex-backed RAG over pluggable vector stores, MCP tool packs behind a sandboxed registry, and first-class observability through Langfuse/OpenInference/Traceloop. Guardrails — budget ceilings, risk-tiered human approval gates, and an import/exec allowlist — are part of the core loop rather than a wrapper around it.
Status: working system, suite not green. As of 9f8a0e2: 863 tests
collected, 765 passing, 74 failing, 22 erroring, 55.8% measured line coverage of
src/. The orchestrator, routing, failure-taxonomy and approval subsystems are
well covered (90.5% / 96.6% / 91.4% / 98.0%); src/cuga/tools/ has no test at
all. Full provenance in
docs/testing/TEST_COVERAGE_MAP.md. Treat
PRODUCTION_READINESS.md as the honest inventory: every
claim there carries an evidence pointer, and anything unevidenced sits in a
clearly-marked "Planned" section.
This repository is a fork of cuga-project/cuga-agent.
- Upstream provides the core CUGA agent, the FastAPI backend, the Vue
frontend, and the MCP tools environment. Every release tag reachable from this
history (
v0.2.6…v0.3.2) is an upstream tag; this fork has published none of its own. - This fork adds the layers an enterprise deployment needs and upstream does
not ship: the orchestrator protocol and failure taxonomy
(
src/cuga/orchestrator/), risk-tiered HITL approval gates, guardrail policy with budget enforcement (src/cuga/backend/guardrails/), the governance and secrets hardening insrc/cuga/security/, the observability layer (src/cuga/observability/), Kubernetes manifests underops/k8s/, and the documentation set indocs/. - Version numbers here are fork milestone labels, not releases. See CHANGELOG.md.
To track upstream: git remote add upstream https://github.com/cuga-project/cuga-agent.git.
- Composable agent graph — Planner → Tool/User executor → Memory + observability hooks, wired for LangGraph.
- RAG-ready — LlamaIndex loader/retriever with pluggable vector stores (Chroma, Qdrant, Weaviate, Milvus);
localneeds no service. - Guardrails in the loop — budget ceilings, risk-tiered HITL approval gates (
READ/WRITE/DELETE/FINANCIAL/EXTERNAL), import and exec allowlists. - Observability-first — Langfuse/OpenInference emitters, golden-signal metrics on
/metrics, structured audit logs. - Developer experience — Typer CLI, Makefile tasks, uv env management, Ruff + mypy, pytest, pre-commit.
- Deployment — Dockerfile, Kubernetes manifests in
ops/k8s/, GitHub Actions CI,.env.examplefor cloud or on-prem.
┌──────────────────────────┐
│ Controller │
│ (policy + correlation ID)│
└────────────┬─────────────┘
│
plan(goal, registry)
│
┌──────────────┐ ┌─────────▼─────────┐ ┌────────────────────┐
│ Registry/CFG │──sandbox▶│ Planner │──steps──▶│ Executor/Tools │
│ (Hydra/Dyn) │ │ (ReAct/Plan&Exec) │ │ (LCEL, MCP, HTTP) │
└──────────────┘ └─────────┬─────────┘ └─────────┬──────────┘
│ │
traces + memory writes Langfuse/OpenInference
│ │
┌───────▼────────┐ ┌─▼────────┐
│ Memory / RAG │◀────context───────│ Clients │
│ (LlamaIndex) │ │ (CLI/API)│
└────────────────┘ └──────────┘
For a role-by-role, mode-aware walkthrough of how the controller, planners, executors, and MCP tool packs fit together (plus configuration keys), see docs/agents/architecture.md. For an MCP + LangChain web stack overview that covers the FastAPI backend, Vue 3 frontend, streaming flows, and configuration surfaces, see docs/MCP_LANGCHAIN_OVERVIEW.md. A step-by-step stable local launch checklist (registry + sandbox + Langflow readiness) lives in docs/local_stable_launch.md.
| Document | What it covers |
|---|---|
| System Execution Narrative | Request → response flow across all three entry points (CLI, FastAPI, MCP) |
| Architecture · agent roles | How controller, planners, executors and tool packs fit together |
| Orchestrator interface | Lifecycle callbacks, failure taxonomy, retry semantics, routing authority |
| FastAPI role | Why FastAPI is transport only, not orchestration |
| Developer onboarding | Setup → first agent → custom tool → custom agent, with working examples |
| Enterprise workflows | End-to-end examples: onboarding, incident response, data pipelines |
| Observability guide | Structured logging, tracing, metrics, dashboards, troubleshooting |
| Test coverage map | Measured coverage per component, with provenance and gaps |
| Production readiness | Evidence-backed checklist; unevidenced claims quarantined |
| Security · guardrails | Secret handling, sandboxing, allowlists, change-management rules |
| Planning · history | Roadmap and task tracker; archived development logs |
# 1) Install (Python >=3.10)
uv sync --all-extras --dev
uv run playwright install --with-deps chromium
# 2) Configure environment
cp .env.example .env
# set OPENAI_API_KEY / LANGFUSE_SECRET / etc inside .env
# 3) Smallest complete loop: registry -> memory -> plan -> execute
# No credentials, no network, ~80 lines.
uv run python examples/minimal_enterprise_agent/run.py
# 4) Run demo agent locally
uv run cuga start demo
# 5) LangGraph orchestration example
uv run python examples/run_langgraph_demo.py --goal "triage a support ticket"Start with examples/minimal_enterprise_agent/ — it is the shortest path to understanding how the pieces connect.
- Dependencies:
uv(orpip), optional browsers for Playwright, optional vector DB service (Chroma/Weaviate/Qdrant/Milvus). - Development:
uv sync --all-extras --devinstalls dev + optional extras (memory,sandbox,groq, etc.). - Pre-commit:
uv run pre-commit installthenuv run pre-commit run --all-files.
.env.example lists every variable the LLM, tracing and storage layers read;
src/cuga/config/validators.py enforces the required subset per environment mode
(LOCAL / SERVICE / MCP / TEST) at startup.
Three configuration directories exist. They look redundant and are not — each is
read by a different code path in src/cuga/config/resolver.py, which is why they
have not been merged:
| Directory | Read by | Contents |
|---|---|---|
config/ |
resolver.py (config/registry.yaml), scripts/verify_guardrails.py |
MCP/tool registry defaults and server fragments. Hydra-composed, with .guardrails-inherit markers. |
configs/ |
resolver.py YAML sources (precedence tier YAML) |
Runtime profiles you are expected to edit: agent.demo.yaml, memory.yaml, rag.yaml, observability.yaml, deploy/. |
configurations/ |
resolver.py (_shared/*.yaml, precedence tier DEFAULT), src/cuga/security/governance_loader.py |
Shipped defaults and policy: tenant capabilities, budget ceilings, governance policies, tool fragments, TOML profiles. Treat as read-only. |
Precedence, lowest to highest, is defined by the Precedence enum in
resolver.py: DEFAULT (configurations/_shared/) → YAML (configs/*.yaml,
config/registry.yaml) → environment → explicit overrides.
Run uv run python scripts/verify_guardrails.py --base origin/main before
shipping any change that touches registry or guardrail configuration.
- Review AGENTS.md before altering planners, tools, or registry entries; it is the single source of truth for allowlists, sandbox expectations, budgets, and redaction.
- Guardrail and registry changes are enforced by CI:
scripts/verify_guardrails.py --base <branch>collects diffs and fails ifREADME.md,PRODUCTION_READINESS.md,CHANGELOG.md, ordocs/planning/TASKS.mdare not updated alongside guardrail changes or if## vNextlacks a guardrail note. - Keep production checklists (PRODUCTION_READINESS.md) and security docs in sync with guardrail adjustments so downstream users understand the default policies and where to override them.
- Developer checklist: ensure registry entries declare sandboxes +
/workdirpinning for exec scopes, budget/observability env keys (AGENT_*,OTEL_*, LangFuse/LangSmith, Traceloop) are wired,docs/mcp/tiers.mdis regenerated fromdocs/mcp/registry.yaml, and new/updated tests exercise planner ranking, import guardrails, and registry hot-swap determinism.
- Planner: ReAct or Plan-and-Execute; emits steps with policy-aware cost/latency hints.
- Tool Executor: LCEL/LangChain tools, MCP adapters, HTTP/OpenAPI runners with sandboxed registry resolution.
- RAG/Data Agent: LlamaIndex loader+retriever (docs in
rag/), vector memory connectors inmemory/. - Coordinator: CrewAI/AutoGen-like orchestrator for multi-agent hand-offs.
- Observer: Langfuse/OpenInference emitters with correlation IDs and redaction hooks.
See AGENTS.md for role details and USAGE.md for end-to-end flows.
- Drop documents into
rag/sources/or configure a remote store. - Choose a backend in
configs/memory.yaml(chroma|qdrant|weaviate|milvus|local). - Run
uv run python scripts/load_corpus.py --source rag/sources --backend chroma. - Query via
uv run python examples/rag_query.py --query "How do I add a new MCP tool?".
memory/exposesVectorMemory(in-memory fallback), summarization hooks, and profile-scoped stores.- State keys are namespaced by profile to preserve sandbox isolation.
- Persistence is opt-in; see
configs/memory.yamlandTESTING.mdfor guidance.
- Langfuse client is wired via
observability/langfuse.pywith sampling + PII redaction hooks. - OpenInference/Traceloop emitters are optional and can be toggled per profile.
- Structured audit logs live under
logs/when enabled; avoid committing artifacts. - Watsonx Granite calls validate credentials up front and append JSONL audit rows with timestamp, actor, parameters, and outcome for offline review.
- The FastAPI orchestrator exposes a Prometheus-compatible metrics endpoint at
/metrics(default port 8000). This endpoint exports golden-signal metrics such ascuga_requests_total,cuga_success_rate,cuga_latency_ms{percentile="p50|p95|p99"},cuga_tool_error_rate,cuga_budget_warnings_total, andcuga_budget_exceeded_total. - Configure OpenTelemetry (OTLP) or console exporters via environment variables. Common envs:
OTEL_EXPORTER_OTLP_ENDPOINT— OTLP HTTP/gRPC endpoint for traces/metrics (optional; when unset the console exporter is used).OTEL_SERVICE_NAME— service name to appear in traces (default:cuga-orchestrator).OTEL_TRACES_EXPORTER/OTEL_METRICS_EXPORTER— exporter type (otlp, logging, none).
Example: curl the metrics endpoint locally
# If running the orchestrator locally on port 8000
curl -sS http://localhost:8000/metrics | head -n 80
# Expected sample lines (Prometheus format):
# cuga_requests_total 42
# cuga_success_rate 0.95
# cuga_latency_ms{percentile="p50"} 150.0
# cuga_latency_ms{percentile="p95"} 450.0
# cuga_tool_error_rate 0.02
# cuga_budget_warnings_total 3
# cuga_budget_exceeded_total 0agents/outlines planner/worker/tool-user patterns and how to register them with CrewAI/AutoGen.examples/multi_agent_dispatch.pydemonstrates round-robin delegation with shared vector context.- Hand-offs carry correlation IDs and redacted summaries, not raw prompts.
uv run ruff check .
uv run mypy src
uv run pytest tests -q --cov=src \
--ignore=tests/unit/test_config_precedence.py \
--ignore=tests/unit/test_http_client.pyNote the two --ignore flags: test_config_precedence.py aborts collection on an
ImportError and tests/unit/test_http_client.py collides on basename with
tests/unit/security/test_http_client.py. Both are open defects.
pytest.ini sets testpaths = tests/unit tests/scenario, so a bare pytest runs
451 of the 863 tests in tests/. Pass tests explicitly for the full suite.
CI workflows: ci.yml (lint + mypy + tests on Python 3.10/3.11/3.12, plus safety
and demo-smoke jobs), tests.yml, lint.yml, guardrails.yml, secret-scan.yml,
stability-tests.yml, release.yml. ci.yml gates on --cov-fail-under=80,
which measured coverage of 55.8% does not currently meet.
Details and per-component numbers: docs/TESTING.md, docs/testing/TEST_COVERAGE_MAP.md.
CUGAR Agent enforces security-first design with deny-by-default policies per AGENTS.md:
- Allowlist-First Tool Selection: Only explicitly allowed tools from
cuga.modular.tools.*can execute - Deny-by-Default Network: Network egress restricted to domain allowlist; localhost/private networks blocked by default
- Sandbox Isolation: All tool execution in isolated sandboxes (py/node slim|full, orchestrator profiles) with read-only mounts
- Budget Enforcement: Cost ceilings (default: 100 units/task) with
warnorblockpolicies - Human-in-the-Loop Approval: High-risk operations (DELETE, FINANCIAL) require explicit approval before execution
Request → Budget Guard → Tool Allowlist → Parameter Validation → Network Policy → Sandbox Execution
↓ ↓ ↓ ↓ ↓
(ceiling=100) (cuga.modular (type/range/pattern) (domain allowlist) (read-only)
.tools.* only) (no localhost)
Approval Flow (HITL):
- Low-risk (READ): Auto-approved, logged
- Medium-risk (WRITE): Auto-approved with audit trail
- High-risk (DELETE, FINANCIAL): Requires human approval (5min timeout, reject on timeout)
Budget Policy:
AGENT_BUDGET_CEILING=100(default): Max cost units per taskAGENT_BUDGET_POLICY=warn|block: Warn and continue, or block executionAGENT_ESCALATION_MAX=2: Max approval escalations before admin approval required
See SECURITY.md for complete security controls and docs/security/GOVERNANCE.md for governance architecture.
CUGAR Agent enforces security-first design with deny-by-default policies per AGENTS.md § 4 Sandbox Expectations:
- Policy Gates: HITL approval points for WRITE/DELETE/FINANCIAL actions (Slack send, file delete, stock orders)
- Per-Tenant Capability Maps: 8 organizational roles (marketing/trading/engineering/support) with tool allowlists/denylists
- Runtime Health Checks: Tool discovery ping, schema drift detection, cache TTLs to prevent huge cold-start lists
- Layered Access Control: Tool registration → Tenant map → Tool-level restrictions → Rate limits
See docs/security/GOVERNANCE.md for complete governance architecture, configuration files, and integration patterns.
- No eval/exec: All
eval()andexec()calls eliminated from production code paths - AST-based expression evaluation: Use
safe_eval_expression()fromcuga.backend.tools_env.code_sandbox.safe_evalfor mathematical expressions- Allowlisted operators: Add/Sub/Mul/Div/FloorDiv/Mod/Pow
- Allowlisted functions: math.sin/cos/tan/sqrt/log/exp, abs/round/min/max/sum
- Denies: assignments, imports, attribute access, eval/exec/import
- SafeCodeExecutor: All code execution routed through
SafeCodeExecutororsafe_execute_code()fromcuga.backend.tools_env.code_sandbox.safe_exec- Import allowlist: Only
cuga.modular.tools.*permitted - Import denylist: os/sys/subprocess/socket/pickle/eval/exec/compile
- Restricted builtins: Safe operations (math/types/iteration) allowed; eval/exec/open/import denied
- Filesystem deny-default: No file operations unless explicitly allowed
- Timeout enforcement: 30s default, configurable
- Audit trail: All imports/executions logged with trace_id
- Import allowlist: Only
- SafeClient wrapper: All HTTP requests MUST use
SafeClientfromcuga.security.http_client- Enforced timeouts: 10.0s read, 5.0s connect, 10.0s write, 10.0s total
- Automatic retry: Exponential backoff (4 attempts max, 8s max wait)
- URL redaction: Query params and credentials stripped from logs
- Env-only secrets: Credentials MUST be loaded from environment variables
- CI enforces
.env.exampleparity validation (no missing keys) - Secret scanning: trufflehog + gitleaks on every push/PR
- Hardcoded API keys/tokens trigger CI failure
- CI enforces
- Import restrictions: Dynamic imports limited to
cuga.modular.tools.*namespace only - Profile isolation: Memory and tool access namespaced per profile; no cross-profile leakage
- Sandbox profiles: All registry entries declare sandbox profile (py/node slim|full, orchestrator)
- Read-only defaults: Mounts are read-only by default;
/workdirpinning for exec scopes
See AGENTS.md for complete guardrail specifications and docs/security/ for detailed security controls.
CUGAR Agent provides production-grade observability with structured events, golden signals, and multi-backend export:
- Structured Events:
plan_created,route_decision,tool_call_start/complete/error,budget_warning/exceeded,approval_requested/received/timeout - Golden Signals: Success rate (%), latency (P50/P95/P99), tool error rate (%), mean steps/task, approval wait time, budget utilization
- Trace Propagation:
trace_idflows through CLI → planner → worker → coordinator → tools with parent-child relationships - PII Redaction: Auto-redact sensitive keys (
secret,token,password,api_key,credential,auth) before emission
# Prometheus metrics endpoint (scrape target)
curl http://localhost:8000/metrics
# Expected metrics:
cuga_requests_total # Total requests handled
cuga_success_rate # % successful requests
cuga_latency_ms{percentile} # P50/P95/P99 latency
cuga_tool_error_rate # % failed tool calls
cuga_steps_per_task # Mean planning steps
cuga_budget_warnings_total # Budget warnings emitted
cuga_budget_exceeded_total # Budget hard blocks
cuga_approval_requests_total{status} # Approval flow tracking- OpenTelemetry (OTLP): Set
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4318for Jaeger/Zipkin/Tempo - LangFuse: Set
LANGFUSE_PUBLIC_KEY,LANGFUSE_SECRET_KEY,LANGFUSE_HOSTfor LLM tracing - LangSmith: Set
LANGCHAIN_API_KEY,LANGCHAIN_PROJECT,LANGCHAIN_ENDPOINT - Console (Default): Offline-first JSON logs to stdout (no network required)
Import pre-built dashboard from observability/grafana_dashboard.json:
- Request rate & success rate panels
- Latency percentile charts (P50/P95/P99)
- Tool error breakdown by tool/type
- Budget utilization gauge
- Approval queue depth
- Event timeline with filtering
Configuration:
# Enable OTEL export
export OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4318"
export OTEL_SERVICE_NAME="cuga-orchestrator"
# Start with observability
uv run cuga start demoSee docs/observability/OBSERVABILITY_GUIDE.md for detailed instrumentation guide and PRODUCTION_READINESS.md for metrics scraping setup.
- Which LLMs are supported?
- OpenAI (GPT-4o, GPT-4 Turbo)
- Azure OpenAI
- Anthropic (Claude 3.5 Sonnet, Opus, Haiku)
- IBM Watsonx / Granite 4.0 (granite-4-h-small, granite-4-h-micro, granite-4-h-tiny) — Default provider with deterministic temperature=0.0
- Groq (Mixtral)
- Google GenAI
- Any LangChain-compatible model via adapters
- Do I need a vector DB? Not for quickstarts; an in-memory store is bundled. For production use Chroma/Qdrant/Weaviate/Milvus.
- How do I add a new tool? Implement
ToolSpecintools/registry.pyor wrap an MCP server; seeUSAGE.md. - Is this production-ready? Core stack follows sandboxed, profile-scoped design with observability. Harden configs before internet-facing use.
- How do I configure Watsonx/Granite? Set environment variables:
WATSONX_API_KEY,WATSONX_PROJECT_ID, and optionallyWATSONX_URL. Seedocs/configuration/ENVIRONMENT_MODES.mdfor details.
Read CONTRIBUTING.md for the development workflow, commit conventions and release process, and AGENTS.md before touching planners, tools or registry entries.
Planned work lives in docs/planning/ROADMAP.md and docs/planning/TASKS.md; unevidenced production claims are tracked in the "Planned / not yet evidenced" section of PRODUCTION_READINESS.md.
Apache 2.0. See LICENSE.
