[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"
-
Updated
Oct 3, 2026 - Python
[JMLR 2026] "UQLM: A Python Package for Uncertainty Quantification in Large Language Models"
Span-level grounding verification for RAG, code, and tool-grounded AI outputs.
up-to-date curated list of state-of-the-art Large vision language models hallucinations research work, papers & resources
[ACL 2024] User-friendly evaluation framework: Eval Suite & Benchmarks: UHGEval, HaluEval, HalluQA, etc.
HaluMem is the first operation level hallucination evaluation benchmark tailored to agent memory systems.
🔎Official code for our paper: "VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation".
[ACL 2026] Curated papers on video LLM hallucination, with benchmarks, mitigation methods, and an interactive browser. Updated monthly.
Unofficial implementation of Microsoft’s Claimify Paper: extracts specific, verifiable, decontextualized claims from LLM Q&A to be used for Hallucination, Groundedness, Relevancy and Truthfulness detection
A benchmark for evaluating hallucinations in large visual language models
Repository for the paper "HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models" https://arxiv.org/abs/2605.19341
Code release for THRONE, a CVPR 2024 paper on measuring object hallucinations in LVLM generated text.
A robust hybrid pipeline for detecting hallucinated citations in academic papers and research documents. The system combines exact bibliographic lookup, fuzzy matching, and optional LLM verification to classify citations as valid, partially valid, or hallucinated.
When AI makes $10M decisions, hallucinations aren't bugs—they're business risks. We built the verification infrastructure that makes AI agents accountable without slowing them down.
MCP server for Eruka — anti-hallucination context memory for AI agents
TrustScoreEval: Trust Scores for AI/LLM Responses — Detect hallucinations, flags misinformation & Validate outputs. Build trustworthy AI.
A blind benchmark for legal citation verification — 4-label classification over IL + federal primary law
Open-source, self-hosted Fastest AI security gateway for LLM and agent apps: guardrails, agentic PDP, semantic caching, MCP gateway.
A comprehensive study on reducing hallucinations in Large Language Models through strategic prompt engineering techniques. (COV + COT + Hybrid)
HALLUCINATED BY CURSOR WITh CODEX PLUGIN:::BEWARE:::::BaseX Coding Language - Revolutionary Base 5.10 Quantum Teleportation & Infinite Storage System by Joshua Hendricks Cole
To associate your repository with the hallucination-evaluation topic, visit your repo's landing page and select "manage topics."