Automate the obvious and investigate the ambiguous. High-performance safety rules engine for real-time event processing at scale.
-
Updated
Sep 5, 2026 - Python
Automate the obvious and investigate the ambiguous. High-performance safety rules engine for real-time event processing at scale.
Directory of open source tools for online trust and safety
Making open safety AI models accessible and beneficial to the safety community
Review and moderation, your way. Online safety dashboard, queues, routing and automatic enforcement rules, and integrations.
Software and Datasets for Mitigating Online Gender Based Violence in India
ActivityPub Trust and Safety Taskforce
Reproducible OpenEnv-compatible reinforcement-learning environments for social-media integrity moderation.
A self-hostable alternative to Ozone designed to work with Coop
A trust and safety agent that interacts with Osprey for investigation, real-time analysis, and prevention implementations
🛡️ Let your Rails users report and block each other (Trust & Safety)
Site with documentation and policies for the ROOST open source community. File non-technical or ROOST-wide issues here!
Self-hosted Trust & Safety policy engine with A/B testing, replay, and full audit trails
The safe version of OpenClaw. Same agent, with safety and accountability built in. Powered by AEP.
A guide to engineering AI errors out of agentic workflows
Proof-backed AI agent for checking suspicious job posts, recruiter messages, and apply links, now live with case study and pilot intake.
A browser extension to detect deepfakes and any other visual content you choose to filter out, right on the webpage
CaseLinker is a project designed to aggregate, structure, analyze, and visualize statistical and contextual information from cases involving crimes against children and child sexual exploitation & abuse (CSEA)
Supply-side ad ops diagnostic agent built on AAMP standards. MCP server with bid rate analysis, IVT detection, ads.txt compliance, and demand diagnostics.
Artifact for "Love, Lies, and Language Models" (USENIX Security '26) — evaluating whether commercial content moderation detects romance-baiting fraud. It largely doesn't.
To associate your repository with the trust-and-safety topic, visit your repo's landing page and select "manage topics."