OWASP GenAI Data Security compliance scanning for AI & agentic codebases
Audits your GenAI app against all 21 controls of the OWASP GenAI Data Security Risks & Mitigations 2026 framework. A deterministic engine owns the pattern matching, so results are reproducible and secrets never leave your machine — the LLM only orchestrates and writes prose.
Quick start · How it works · Features · Usage · CI/CD · The report · Contributing
Most scanners either miss GenAI-specific risk (generic SAST) or hand the judgment to an LLM that produces a different answer every run. DSGAI Scanner splits the job: a single-file, dependency-free Python CLI (cli/dsgai_scan.py) runs 107 PCRE rules over your code via ripgrep and emits identical findings every time; the Claude Code skill orchestrates the run, classifies ambiguous cases, and renders the report. You get a reproducible compliance artifact — and, if you want it, a $0 LLM-free path that drops straight into CI as SARIF.
Table of contents
- Quick start
- How it works
- Features
- The 21 DSGAI controls
- Usage
- CI/CD & Code Scanning
- The report
- How redaction works
- Reference — checkpoint file · PDF export · scope annotation · remediation tiers
- Cost & runtime
- Contributing · Non-goals
- License & attribution
Deterministic CLI — $0, no LLM, no account. Reproducible findings + SARIF in seconds. Needs only Python 3.10+ and ripgrep with PCRE2.
git clone --depth 1 https://github.com/GenAI-Security-Project/GenAI-Data-Security-Initiative
python GenAI-Data-Security-Initiative/dsgai_scanner_tool/cli/dsgai_scan.py scan . \
--sarif DSGAI-scan.sarif --json-out DSGAI-scan.jsonClaude Code skill — full report with narrative + remediation.
# 1. Install the skill
curl -fsSL https://raw.githubusercontent.com/GenAI-Security-Project/GenAI-Data-Security-Initiative/main/dsgai_scanner_tool/dsgai_scanner_tool.md \
-o ~/.claude/commands/dsgai_scanner_tool.md
# 2. Scan your repo
cd ~/my-genai-app && claude
/dsgai_scanner_tool
# 3. Open the timestamped report
open dsgai-reports/DSGAI-report-*.html # macOS (xdg-open on Linux, start on Windows)Tip
Check your environment first with python cli/dsgai_scan.py doctor — it verifies Python, ripgrep + PCRE2, and that the ruleset loads.
The deterministic CLI does all pattern matching and CVE lookups; the skill (or the CI's report renderer) turns the checkpoint into a report. Nothing but package names + versions ever leaves your machine.
flowchart TD
Repo[(Your GenAI repo)] --> Detect{"GenAI / agentic<br/>code detected?"}
Detect -->|No| NA["Minimal report<br/>21 × NOT APPLICABLE"]
Detect -->|Yes| Engine
subgraph Engine["Deterministic CLI — cli/dsgai_scan.py · stdlib + ripgrep · no LLM"]
direction TB
Scan[Scan 21 DSGAI controls<br/>107 PCRE rules]
CVE[CVE enrichment<br/>6 ecosystems · cached]
Scan --> Out[DSGAI-scan.json<br/>+ SARIF 2.1.0]
CVE --> Out
end
CVE -. package + version only .-> Sources((OSV · NVD))
Out --> Skill[Claude Code skill<br/>classify · judge · prose]
Skill --> Report[Timestamped HTML report<br/>STRICT-redacted]
Out --> Renderer[cli/dsgai_report.py<br/>deterministic HTML]
Renderer --> Report
Out --> CodeScan[GitHub Code Scanning]
Report --> Auditor([Auditor / GRC])
Report --> Ticket([Jira / backlog])
CodeScan --> PR([PR annotations])
style Engine fill:#0f172a,color:#fff
style Skill fill:#7c3aed,color:#fff
style Report fill:#16a34a,color:#fff
style CodeScan fill:#0b7285,color:#fff
style NA fill:#6b7280,color:#fff
style Sources fill:#e0e7ff,color:#1e1b4b
Value-bearing rules never surface the secret. Credential/PII patterns run through ripgrep in erase-the-match mode (rg -o --replace ''), so the matched value is destroyed inside ripgrep before anything is emitted — it cannot reach the report, the checkpoint, or a tool call. See How redaction works.
| 🎯 Deterministic engine | 107 PCRE rules, run via rg --pcre2. Identical findings on identical input — a compliance report you can diff, not an LLM opinion. |
| 🔒 Redaction by construction | Value-bearing matches are erased inside ripgrep; the checkpoint schema forbids content fields (machine-checked). |
| 🧾 SARIF 2.1.0 → Code Scanning | Native GitHub Code Scanning annotations on changed lines, plus a portable artifact. |
| 🐛 CVE enrichment, no hallucination | The CLI queries OSV (+ NVD for CVSS) per pinned version across 6 ecosystems; the LLM never transcribes CVE data. Cached, offline-capable. |
| 🌐 Multi-language | Python, JS/TS, Java, Kotlin, Go, and credential coverage for C#, Rust, Ruby. |
| 🧰 Ships with the ecosystem | A gitleaks rule pack and a Semgrep export so incumbent toolchains carry the framework. |
| 🎚️ Team-scale controls | Inline # dsgai-ignore suppressions (with reasons), a baseline so CI gates only on new findings, and incremental --diff scans. |
| ✅ Tested & gated | A public vulnerable fixture app with a line-pinned answer sheet and a CI self-test that makes external rule PRs safe to merge. |
| 🗣️ Honest reporting | STRICT mode renders stable file IDs, never the secret; the report declares its own residual risk instead of implying it can be published as-is. |
Every control is rated PASS / WARN / FAIL / NOT VALIDATED / NOT APPLICABLE / VENDOR ATTESTATION REQUIRED — click to expand the full list.
| Risk | Control | Responsibility |
|---|---|---|
| DSGAI01 | Training Data Privacy | BOTH |
| DSGAI02 | Agentic Identity & Credential Management | BUILD |
| DSGAI03 | Shadow AI & Unauthorized Data Flows | BOTH |
| DSGAI04 | AI Supply Chain Security | BUILD |
| DSGAI05 | RAG Data Security | BUILD |
| DSGAI06 | MCP & Plugin Security | BUILD |
| DSGAI07 | Data Lifecycle Management | BUILD |
| DSGAI08 | Regulatory & Privacy Compliance | BOTH |
| DSGAI09 | Multimodal AI Data Security | BOTH |
| DSGAI10 | Synthetic Data Security | BUILD |
| DSGAI11 | Multi-Tenant Data Isolation | BUILD |
| DSGAI12 | Database Agent Security | BUILD |
| DSGAI13 | Vector Store Security | BUILD |
| DSGAI14 | AI Telemetry & Observability Security | BUILD |
| DSGAI15 | Context Window Data Security | BUILD |
| DSGAI16 | AI IDE Plugin & Extension Security | BUILD |
| DSGAI17 | AI System Resilience & Availability | BUILD |
| DSGAI18 | Model Output Data Security | BUILD |
| DSGAI19 | AI Data Labeling Security | BUILD |
| DSGAI20 | Inference API Security | BOTH |
| DSGAI21 | Knowledge Store Security | BUILD |
[BUILD] you implement (scanned mechanically) · [BUY] the vendor is responsible (emits a Vendor Attestation Required callout) · [BOTH] shared. See the Reference → Scope annotation section below.
cli/dsgai_scan.py is a single stdlib-only file that shells out to ripgrep. Subcommands:
| Command | What it does |
|---|---|
scan [path] |
Run the ruleset; emit DSGAI-scan.json, SARIF, and/or a table. |
detect [path] |
Exit 0 if the repo contains GenAI/agentic signals, 1 otherwise. |
doctor |
Check Python, ripgrep + PCRE2, and that the ruleset loads. |
baseline [path] |
Snapshot current findings to a baseline file. |
cve [path] |
CVE enrichment only, from dependency manifests. |
Key scan flags (all combinable):
| Flag | Effect |
|---|---|
--sarif FILE |
Write SARIF 2.1.0 for GitHub Code Scanning. |
--json-out FILE |
Write the DSGAI-scan.json checkpoint. |
--internal |
Render full paths (default STRICT renders file IDs). |
--scope PATH |
Restrict the scan to a sub-directory. |
--exclude PATH |
Exclude a path/glob (repeatable). |
--diff REF |
Incremental: only files changed vs REF. |
--baseline FILE |
Gate only on findings not in the baseline. |
--no-cve / --offline / --refresh-cve |
Control CVE enrichment & caching. |
--fail-on {fail,warn} |
Exit non-zero at this threshold (CI gating). |
Suppress a finding inline (a reason is required, and it stays visible in a Suppressed section):
db.execute(query) # dsgai-ignore: P12.1 reason="reviewed 2026-07 — ORM param binding"The skill prefers the bundled CLI and falls back to an in-context scan if it isn't present; the report header names which engine ran.
| Command | What it does |
|---|---|
/dsgai_scanner_tool |
Full scan. STRICT redaction, CVE enrichment, whole repo. |
/dsgai_scanner_tool --internal |
Full file paths (team-internal report). |
/dsgai_scanner_tool --no-cve |
Skip CVE lookups (air-gapped / offline). |
/dsgai_scanner_tool --scope app/agents/ |
Scan one sub-directory (monorepos). |
/dsgai_scanner_tool --diff main --baseline dsgai-baseline.json |
Incremental, gated on new findings. |
The tool-neutral variant dsgai_scanner_prompt.md is generated from the skill (so the two never drift) and works with any assistant that can run shell commands — Cursor, Copilot Chat, ChatGPT, Gemini. Paste it as instructions and give the model access to your files and a shell.
GitHub Action — integrations/dsgai-scan.yml
A hardened, two-job workflow. The design keeps secrets away from any job that reads untrusted code:
scan— runs the deterministic CLI only: no LLM, no API key, no network egress with a secret. Safe on fork PRs. Uploads SARIF to Code Scanning (same-repo) and the checkpoint as an artifact.narrate— renders the HTML report deterministically (cli/dsgai_report.py); no LLM, no secret either.
All third-party actions are pinned by commit SHA. Gating is report-only by default — set the repo variable DSGAI_FAIL_ON to fail or warn to break the build. No ANTHROPIC_API_KEY is required.
Important
CI always scans in STRICT mode (file IDs, never full paths or values). Don't pass --internal in a workflow whose artifacts or logs are visible beyond your team.
Block hardcoded credentials before they land. Recommended path is the gitleaks rule pack (entropy-aware, cross-platform); a zero-dependency ripgrep script is provided as a bash-3.2-safe fallback. See integrations/pre-commit-hook.md.
dist/dsgai.semgrep.yaml exports the STRUCTURAL rules as a Semgrep pack — run the DSGAI framework inside a toolchain you already have.
It's built to complement the tools you already run, not replace them — it even ships packs for them.
- Framework-native. The reference implementation of the OWASP GenAI Data Security 2026 framework, maintained inside the initiative that authors it. Findings map 1:1 to the 21 controls, so the report reads as compliance evidence, not generic lint output.
- Deterministic where it matters. Pattern matching, CVE lookup, and report structure are code (stdlib CLI + ripgrep) — not model output. Identical input yields byte-identical findings, and every rule is data pinned to a known-answer fixture corpus. The optional LLM layer only adds prose; it never decides what was found.
- Redaction by construction, not by policy. Value-bearing rules erase the match inside ripgrep, and the checkpoint schema formally forbids content fields — so the guarantee is machine-checked in CI, not left to discipline.
- GenAI-specific depth, not SAST breadth. Vector-store auth, RAG access control, MCP transport, prompt-logging PII, unsafe model deserialization, agentic credential handling — patterns general rulesets don't carry.
- Feeds your toolchain. SARIF into GitHub Code Scanning; exported Semgrep and gitleaks packs so the framework rides tools you already operate.
Rendered deterministically by cli/dsgai_report.py from a scan of the public fixture app — STRICT mode, file IDs + line numbers only, zero real-repo disclosure. Fully reproducible.
The self-contained HTML report (no CDN, no external fonts; prints cleanly to PDF) contains:
- Executive summary & an obfuscation-mode badge (🛡️ STRICT / 🔓 INTERNAL)
- Compliance dashboard — counts across all 21 controls; every status carries a symbol + text label, not colour alone
- AI component inventory — detected frameworks, vector stores, LLM providers, MCP servers
- MITRE ATLAS techniques relevant to the detected stack (from a versioned static map)
- Findings — one card per risk with locations, line numbers, and tiered remediation
- CVE advisory panel — advisories for your exact dependency versions, grouped by DSGAI risk
- Compliance checklist mappable to GDPR, EU AI Act, SOC 2, ISO 42001
Warning
Residual risk. STRICT mode is designed to minimize disclosure — file IDs + line numbers only, value-bearing matches never shown. It does not make the report public-safe: the existence and location of failing controls is itself information. Handle it like any security assessment.
The redaction guarantee is structural, not a matter of remembering to be careful:
- Location-only matching. Value-bearing rules (DSGAI02/13/14/15 — credentials, vector-store tokens, telemetry PII, system-prompt secrets) run
rg -o --replace ''. Ripgrep erases the matched text before it emits anything, so the output ispath:line:— the secret never leaves the ripgrep process. - Stable file IDs in STRICT mode. Findings render as
F07:12; theF## → pathmap is written to a gitignoredDSGAI-filemap.jsonthat is never embedded in the report. - A schema that forbids leakage.
schemas/dsgai-scan.schema.jsonrejects any finding carryingmatch_text,content,value, orraw_grep_output, and the CLI self-validates before writing. The guarantee is machine-checked by CI.
Checkpoint file — DSGAI-scan.json
The CLI writes a schema-validated checkpoint alongside the report: framework/ruleset versions, per-control statuses, findings (control, rule ID, rendered path, line, status), suppressed findings, and CVEs. It carries no match content by construction. In STRICT mode it holds file IDs, not full paths — but it still records which controls fail and where, so treat it as a security artifact (commit only if your threat model allows, or .gitignore it).
Export to PDF
Browser: open the report in Chrome/Edge → Ctrl/Cmd+P → Save as PDF (cards expand for print).
Headless:
google-chrome --headless=new --print-to-pdf=DSGAI-report.pdf \
--print-to-pdf-no-header "file://$(pwd)/dsgai-reports/DSGAI-report-<timestamp>.html"Scope annotation — BUILD / BUY / BOTH
Each control is tagged by responsibility. [BUILD] controls are scanned mechanically. [BUY] controls (the vendor's responsibility) emit a VENDOR ATTESTATION REQUIRED callout listing exactly what to request (e.g. SOC 2 report, data-retention policy, rate-limit documentation). [BOTH] controls scan the BUILD portion and emit an attestation callout for the BUY portion. Controls with BUY aspects (DSGAI01, 08, 09, 20) consolidate into one Vendor Attestations to Request card.
Remediation tiers
| Tier | Meaning | Example |
|---|---|---|
| 🔴 Tier 1 — fix today | FAIL items + exploitable CVEs | Hardcoded sk- key; unauthenticated vector store; unsafe torch.load() |
| 🟡 Tier 2 — architecture backlog | WARN + structural gaps | Add a circuit breaker; centralize PII redaction; route via an internal LLM gateway |
| 🔵 Tier 3 — maturity program | NOT VALIDATED (process evidence) | Quarterly red-team; complete a DPIA; commission an AppSec review |
- CLI-only mode: $0.
python cli/dsgai_scan.py scan .uses no LLM — just ripgrep. Reproducible findings + SARIF in seconds. This is what the Action runs on every PR (including forks, since it needs no secrets). - Skill mode: a full scan of the fixture app renders in ~1–3 minutes in Claude Code; token cost scales with findings and report prose, not lines of code — the deterministic engine does the matching.
- Everything runs locally. Only package names + pinned versions go to public CVE databases, and only if you don't pass
--no-cve.
Tip
Found a wrong result? That's a contribution. Run the scan on your repo and file a false-positive or false-negative issue — every accepted report becomes a permanent, credited test case in the fixture corpus. No code required.
Detection rules live as data in rules/dsgai-rules.yaml (schema-validated, compiled to JSON). To add or fix a rule, see rules/README.md and the full guide in CONTRIBUTING.md; the ROADMAP.md tracks what's planned. Every rule ships with positive and negative fixture cases — the CI self-test is the gate.
To stay maintainable and trustworthy, some things are deliberately out of scope:
- No general-purpose secret scanning — we ship a gitleaks pack and lean on battle-tested tooling instead.
- No general SAST — scope is the 21 DSGAI controls, not every code smell.
- No rules without fixture tests — a rule with no positive and negative case has no defined precision.
- No changes that weaken the redaction guarantee — that property is non-negotiable.
Based on the OWASP GenAI Data Security Risks and Mitigations 2026 (v1.0, March 2026), licensed under Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0). (A split that keeps framework text under CC BY-SA 4.0 while re-licensing the executable code under Apache-2.0 is under review with OWASP leadership.)
Framework & scanner by the OWASP GenAI Data Security Initiative, led by Emmanuel Guilherme Junior. v0.1 groundwork by Harish Ramachandran.