Dynamic honeypot and cryptographic canary token sentinel detecting prompt extraction and triggering session freezes
-
Updated
Sep 26, 2026 - Python
Dynamic honeypot and cryptographic canary token sentinel detecting prompt extraction and triggering session freezes
Egress data exfiltration firewall detecting obfuscated credentials, private keys, and API tokens in payloads
Egress data exfiltration firewall detecting obfuscated credentials, private keys, and API tokens in payloads
Indirect prompt injection sanitizer and delimiter boundary firewall neutralizing embedded command overrides
Role-based access control and tool privilege attenuator preventing unauthorized tool invocations
Adversarial prompt jailbreak detector analyzing framing evasion, hypothetical bypasses, and token smuggling
Role-based access control and tool privilege attenuator preventing unauthorized tool invocations
Dynamic honeypot and cryptographic canary token sentinel detecting prompt extraction and triggering session freezes
Indirect prompt injection sanitizer and delimiter boundary firewall neutralizing embedded command overrides
Adversarial prompt jailbreak detector analyzing framing evasion, hypothetical bypasses, and token smuggling
Cross-Modal, Temporal, and Agent-Delegation Defenses Against Jailbreaks in Text-to-Image and Text-to-Video Generation
The language-level attack defense skill every agent should keep. Covers 12 attack categories. Works with Claude, GPT, Gemini, Copilot, and any LLM.
DPO/SafeDPO/OPAD training + eval for teaching tool-using LLMs to refuse falsely-benign MCP exploits
Risk-adaptive prompt-injection defense layer for commercial APIs and local LLMs.
Pre-registered adversarial robustness study testing whether model-initiated session termination provides defensive coverage beyond refusal training against multi-turn attacks. Minimal Python harness with scorer and analysis pipeline. Pilot on Gemma 4 26b. Preliminary findings in FINDINGS.md.
Sub-millisecond semantic WAF and cognitive drift circuit breaker for frontier reasoning agents (Claude Opus 5.5, GPT-6 Astra, Gemini 3.8 Flash Cyber).
[ICLR 2025] Reinforced Blue Teaming for VLMs Against Jailbreak Attacks
Tactical AI security posture and prompt injection vulnerability scanner for AI system instructions.
Semantic-layer prompt injection defence that separates untrusted instructions from authority while preserving the legitimate task.
Self-evolving prompt-injection defense: breach -> synthesize antibody -> two-sided gate -> promote -> broadcast.
To associate your repository with the jailbreak-defense topic, visit your repo's landing page and select "manage topics."