Open Source Reliability Harness: Make your agents follow rules. One line of code to enforce, trace, and improve.
-
Updated
May 23, 2026 - Python
Open Source Reliability Harness: Make your agents follow rules. One line of code to enforce, trace, and improve.
Review and moderation, your way. Online safety dashboard, queues, routing and automatic enforcement rules, and integrations.
🛡️ Programmable Guardrails for LLM Applications in Java. A framework-agnostic toolkit for input/output validation, PII masking, and jailbreak detection. The Java alternative to NVIDIA NeMo Guardrails.
A JavaScript-based content safety system designed to detect and filter sensitive media in real-time, ensuring platform compliance and user protection.
An intelligent task management assistant built with .NET, Next.js, Microsoft Agent Framework, AG-UI protocol, and Azure OpenAI, demonstrating Clean Architecture and autonomous AI agent capabilities
Step-by-Step tutorial that teaches you how to use Azure Safety Content - the prebuilt AI service that helps ensure that content sent to user is filtered to safeguard them from risky or undesirable outcomes
Benchmark LLM jailbreak resilience across providers with standardized tests, adversarial mode, rich analytics, and a clean Web UI.
│ Real-time NSFW & harmful content detection as a service
Arabic Content Moderator — scan text for toxicity, hate speech, spam. Dialect-aware. Fully offline.
Transform uncertainty into absolute confidence.
轻卫:基于 Agent 的中文内容安全检测系统
Free AI safety stack + frontier adversarial red teaming. Policy engine, content scanner, behavioral monitor, MCP gateway. 350+ vulnerabilities found across NVIDIA, Microsoft, Meta, Google. MIT licensed.
A DCSA's presentation journal: cinematic, doc-grounded customer briefings on Azure AI Services — one self-contained HTML deck per service.
抖音视频审核检测|同行举报分析工具|抖音视频风控|抖音风控||优化视频|举报同行|视频监测|视频检测
🛡️ Secure your LLM applications with PromptShields, a framework designed for real-time protection against prompt injection and data leaks.
AI application firewall for LLM-powered apps — multi-layered detection (heuristic, ML classifier, semantic, LLM-judge) against prompt injection, jailbreaks, and data leakage - inferwall.com
Open skill system for humane AI — 9 reusable specs + MCP runtime
Real-time text security pipeline for educational chat platforms — detects PII, self-harm signals, prompt injection & cyberattacks in ~30ms per message. Zero GPU, ~430MB RAM, no LLM calls.
Production-Grade LLM Alignment Engine (TruthProbe + ADT)
Technical presentations with hands-on demos
Add a description, image, and links to the content-safety topic page so that developers can more easily learn about it.
To associate your repository with the content-safety topic, visit your repo's landing page and select "manage topics."