Catch risky content before it reaches your RAG index. Find:
- Indirect instruction indicators — e.g., “ignore previous instructions”, “override policy”, benign canaries like
[TESTMARK] - Risky HTML/Markdown —
<script>,<iframe>,javascript:URIs (common XSS vectors if rendered) - PII/secret‑like strings — email/phone formats, prefixes like
sk_live_/sk_test_
Outputs JSON/CSV and returns non‑zero if findings meet your threshold — perfect for CI gates on docs/KB changes.
- Install
- Quick Start
- Configuration (
rules.yaml) - Exit Codes & Thresholds
- Output Schema
- Examples Included
- Use in CI/CD (GitHub Actions)
- Security Mapping (OWASP / NIST / CISA / SSDF)
- Notes & Safety
- License
# From repo root (editable install for local development)
pip install -e .
# (Optional) With pipx
# pipx install .Requires Python 3.10+.
# Pretty JSON to stdout; fail on >= "med"
rag-scan examples --format json --fail-on med
# CSV to file (for spreadsheets)
rag-scan examples --format csv -o findings.csv --fail-on med
# Always print a one-line severity summary to stderr
rag-scan examples --summaryTune the ruleset without code changes: disable, override severities, or add/replace patterns.
# rules.yaml
disable:
- INJ003 # Ignore benign [TESTMARK] canary in this repo
severity_overrides:
HTML003: med # Treat 'javascript:' as medium
rules:
- code: USR001
desc: Custom indirect phrase
pattern: "\\bcomply with the note below\\b"
severity: med
ignore_case: true
- code: USR002
desc: ZWSP-obfuscated IGNORE (I\u200bG\u200bN\u200bO\u200bR\u200bE)
pattern: "I\\u200bG\\u200bN\\u200bO\\u200bR\\u200bE"
severity: medUse it:
rag-scan docs/ -c rules.yaml --fail-on med --summary0→ No findings at/above threshold (--fail-on low|med|high)1→ At least one finding meets threshold2→ Usage error (e.g., bad path)
# Fail only on high-severity issues
rag-scan examples --fail-on high
# Print summary regardless of pass/fail
rag-scan examples --summaryJSON/CSV columns: doc_id, code, severity, desc, evidence
doc_id— file path scannedcode— rule ID (e.g.,INJ001,HTML003,SEC001)severity—low | med | highevidence— short, single‑line snippet around the match
Example JSON (truncated):
[
{
"doc_id": "examples/unsafe.html",
"code": "HTML003",
"severity": "high",
"desc": "javascript: URI scheme present",
"evidence": "<a href=\"javascript:alert(1)\">bad link</a>"
}
]examples/poison.md→ INJ001/INJ003 (indirect instruction, benign marker)examples/unsafe.html→ HTML001/002/003 (script/iframe/javascript:)examples/pii.txt→ PII001/PII002/SEC001 (email, US phone, secret prefix)
Run quick check:
rag-scan examples --format json --fail-on med --summary > /tmp/findings.jsonAdd a step to your workflow to gate PRs that change docs/KB:
- name: RAG hygiene scan (docs)
run: |
rag-scan docs/ --format json --fail-on med --summary > rag_findings.json- The job fails when exit code is non‑zero (findings ≥ threshold).
- Upload
rag_findings.jsonas an artifact for review.
This tool helps enforce controls that reduce real LLM‑app risks:
| Category | What this tool detects or enables | Standard(s) it supports |
|---|---|---|
| Prompt / Indirect Injection | Flags instruction‑shaped text in content (e.g., “ignore previous…”), supports allowlisting/transform design via rules | OWASP LLM Top‑10 A01 (Prompt Injection); NIST AI 600‑1: Content filtering before model use |
| Insecure Output Handling | Identifies risky HTML/MD (<script>, <iframe>, javascript:) likely to cause XSS if rendered |
OWASP LLM Top‑10 A02; SSDF 800‑218: Taint/sanitize untrusted outputs |
| Training‑Data / RAG Poisoning | Scans KB for malicious patterns and provenance hints; pairs with allowlists/metadata filters in your RAG | OWASP LLM Top‑10 A03; NIST AI 600‑1: Data provenance & pre‑processing |
| Sensitive Info Disclosure | Heuristics for PII and secret‑like tokens to prevent accidental indexing/exposure | OWASP LLM Top‑10 A07; NIST AI 600‑1: Privacy controls |
| SDLC / CI Gate | Thresholded exit codes + summary → easy TEVV integration as part of normal releases | CISA TEVV (AI red‑team/testing in SDLC); SSDF 800‑218 (release gates) |
This scanner is one layer in a defense‑in‑depth approach. Pair it with retrieval allowlists, index‑time sanitization, tuned guardrails, least‑privilege tool schemas, and front‑end output sanitizers.
- Use on non‑production content.
- Patterns are indicators, not perfect classification—tune with
rules.yaml. - Keep benign canaries (e.g.,
[TESTMARK]) for regression tests; you can disable them per‑repo.
MIT — see LICENSE.