A deployment risk analyzer. Give it a change set — a git diff, or a PR's file list — and it scores how risky that change is to deploy: infrastructure changes, schema migrations, dependency bumps, API-surface changes, security-sensitive files, oversized diffs, and more. Each finding cites the evidence that triggered it and suggests what to check.
Pure Python 3.10+. Zero runtime dependencies. Rule-based and auditable, with an optional LLM layer that reads the PR description.
git diff / PR ─> deploy-risk ─> risk score + band + findings
│ (markdown / json)
rule engine exit 1 if too risky (CI gate)
│
optional LLM note (reads the PR prose)
Not every deploy is equally risky, but most teams treat them the same — same review, same process, whether it's a one-word copy fix or a schema migration that drops a table. The information needed to tell them apart is right there in the diff; nobody's looking at it systematically.
deploy-risk looks at it systematically. It encodes the patterns an
experienced reviewer reacts to — "wait, this touches Terraform," "this
is a migration with a DROP," "this adds a .env file" — as
deterministic rules that produce a risk score and a checklist.
The analysis is rule-based, not LLM-based, for the same reason as the rest of this stack: an auditable rule you can point at beats a model that produces a confident number from nowhere. There's an optional LLM layer, but it only reads the PR's prose for signals the file list can't show — it never sets the score.
This is the deployment-risk subsystem from ai-ops-design, and it shares the finding/evidence shape with perf-advisor, regression-radar, and incident-pilot.
git clone https://github.com/anshikapundeel/deploy-risk
cd deploy-risk
pip install -e .Or run without installing:
python3 -m deploy_risk analyze --json examples/risky_pr.jsonRequires Python 3.10+. No other dependencies.
# Analyze your current branch against main (shells out to git):
deploy-risk analyze --base origin/main
# Analyze a JSON change-set fixture:
deploy-risk analyze --json examples/risky_pr.json
# Pipe a raw git diff:
git diff --numstat origin/main | deploy-risk analyze --stdin
# Machine-readable output:
deploy-risk analyze --json examples/risky_pr.json --format json
# CI gate: exit 1 if the risk band is "high" or above
deploy-risk analyze --base origin/main --fail-at high
# List the rules
deploy-risk list-rulesSample output (the bundled risky fixture):
# deploy-risk report
- Change set: PR #482: migrate auth to Redis sessions
- Risk score: 51 → 🔴 high
- Findings: 6
## 1. Database schema migration
Severity: 🔴 CRITICAL Confidence: high Rule: db.migration
Summary. 1 database migration file(s) changed: migrations/0042_drop_sessions_table.sql.
Some migrations contain deletions, which may drop columns or tables...
| Rule | Flags | Severity |
|---|---|---|
infra.change |
Terraform, k8s, Helm, Dockerfiles | HIGH (Terraform) / MEDIUM |
db.migration |
Schema migration files | CRITICAL (with deletions) / HIGH |
deps.change |
Dependency manifests / lockfiles | MEDIUM (manifest) / LOW (lockfile) |
api.surface_change |
protobuf, OpenAPI, GraphQL, route files | HIGH (deletions) / MEDIUM |
security.sensitive_change |
auth, crypto, secrets, committed key files | CRITICAL (secret) / HIGH |
size.large_changeset |
Oversized diffs | MEDIUM / LOW |
size.deletion_heavy |
Mostly-deletions or whole-file removals | MEDIUM / LOW |
ci.config_change |
CI pipelines, release config | MEDIUM |
Each rule lives in its own ~80-line file under deploy_risk/rules/,
produces findings with cited evidence, and is independently tested.
Adding a rule is a small change — see docs/ADDING_RULES.md.
Each finding contributes points by severity (info 0, low 1, medium 4, high 9, critical 16 — super-linear, so one CRITICAL outweighs several LOWs). The total maps to a coarse band:
| Score | Band |
|---|---|
| 0 | 🟢 minimal |
| 1–3 | 🟢 low |
| 4–9 | 🟡 moderate |
| 10–18 | 🟠 elevated |
| 19+ | 🔴 high |
The bands are deliberately coarse. The goal is to sort deploys into "ship it / look first / be careful," not to imply false precision in a number.
The intended workflow: run it on the PR's diff and gate the merge.
- name: Deployment risk check
run: |
deploy-risk analyze --base origin/${{ github.base_ref }} \
--fail-at elevated \
--output risk-report.md
- name: Post report to PR
if: always()
run: gh pr comment "$PR" --body-file risk-report.md--fail-at elevated exits 1 when the band reaches elevated or high, so
the job fails and flags the PR for extra scrutiny. The report posts
either way. Teams pick their own threshold — block at high, or just
advise by not setting --fail-at at all.
deploy-risk ships with no API keys, no default provider, and no outbound calls. The rule engine produces the authoritative score. An LLM, if configured, adds one thing the rules can't: it reads the PR title and body in natural language and surfaces risk the prose reveals (e.g. "the description says this changes retry behavior under load — watch for a thundering herd post-deploy").
export DEPLOY_RISK_LLM_PROVIDER=ollama
export DEPLOY_RISK_LLM_BASE_URL=http://localhost:11434
export DEPLOY_RISK_LLM_MODEL=llama3.1:8b
deploy-risk analyze --json examples/risky_pr.json --llm-noteAlso supports openai-compatible. The note is clearly labeled and
cannot change the score — the rules decide, the LLM advises. Without
the env vars, the note is skipped and the deterministic report still
prints.
For the JSON path, a change set looks like:
{
"label": "PR #482",
"pr_title": "...",
"pr_body": "...",
"files": [
{"path": "infra/terraform/redis.tf", "status": "added",
"additions": 88, "deletions": 0}
]
}status is one of added, modified, deleted, renamed. The
git-diff and stdin paths populate this automatically from
git diff --numstat plus --name-status.
Is: a fast, auditable pre-deploy risk check you can run locally or in CI. It encodes the "wait, this is risky because…" reactions an experienced reviewer has, as rules anyone can read and extend.
Isn't:
- A replacement for code review. It flags categories of risk from file paths and churn; it does not read your logic. A subtle bug in a one-line change scores "minimal" — correctly, because the risk it measures is deployment-shaped, not correctness-shaped.
- A deploy tool. It scores; it doesn't ship or roll back.
- A secret scanner. It flags a newly added secret-shaped file, but it doesn't scan file contents for embedded credentials — use a dedicated secret scanner for that.
- An LLM judge. The LLM only reads the PR prose for extra signal; the score is the rules'.
deploy_risk/
model.py ChangeSet, FileChange, RiskFinding, RiskReport (+ scoring)
engine.py runs rules, enforces the evidence contract
rules/ one file per risk rule
loader.py git numstat / name-status parsing, JSON fixtures
report.py Markdown + JSON renderers
llm.py optional reviewer-note hook; no keys, no defaults
__main__.py CLI
tests/ 27 tests: each rule, the engine, end-to-end, the loader
examples/ risky + safe PR fixtures
docs/ DESIGN.md, ADDING_RULES.md
pip install pytest
pytest -v # 27 testsCI runs on Python 3.10, 3.11, 3.12.
- 8 risk rules with cited evidence
- Score + bands,
--fail-atCI gate - git numstat / stdin / JSON inputs
- Optional LLM reviewer note (no keys baked in)
- 27 tests
- Read file contents for higher-signal rules (feature-flag checks, etc.)
- GitHub Action wrapper that posts the report as a PR check
- Per-repo config for custom rules and thresholds
- Historical calibration — learn which file paths actually caused incidents
- Hand high-risk deploys to incident-pilot as a pre-emptive watch
MIT.