Inference-time scaling for LLMs-as-a-judge.
-
Updated
Nov 5, 2025 - Jupyter Notebook
Inference-time scaling for LLMs-as-a-judge.
Self-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model training. arXiv 2605.08703.
An graph-eval framework for LLM's
OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.
An end-to-end AI agent project that transcribes audio files, embeds user queries, and searches in Qdrant and web browser via the Brave API. A Streamlit interface powered by OpenAI GPT models delivers actionable health insights from both the archive and the latest research.
Open-source benchmark for evaluating LLMs on domain-specific geological reasoning, AI for geology, and source-grounded geological knowledge.
StructAI offers a robust toolkit for LLM interaction—such as structured outputs, context management, and parallel execution.
ProductionOS v1.0 — Claude Code plugin with 76 agents, 39 commands, and 12 hooks. Deploys specialized agents that review, score, and improve your entire codebase. Smart routing, recursive convergence, self-evaluation.
Alexey Vorobey's experimental expert circle for Claude Code — 16 distilled mental models (Eric Seufert, Andrew Chen, Elena Verna, ...) that auto-refresh on cadence from LinkedIn, RSS, YouTube, podcasts.
Extensible benchmarking suite for evaluating AI coding agents on web search tasks. Compare native search vs MCP servers (You.com, expanding) across multiple agents (Claude Code, Gemini, Droid, Codex, expanding) with automated Docker workflows and statistical analysis.
Autonomous judge agents fix your app's debt, build what's missing, and verify every change through a browser or gate suite before it lands, then find the next thing and keep going until you say stop. Each fixer owns a file-disjoint slice of the tree, so no two can ever collide, and nothing unverified reaches git. A Claude Code plugin.
An LLM judge routes every Pi agent turn to the right model — fast execution for routine work, smart reasoning for hard problems — with multi-model failover and automatic orchestration.
A real-time, zero-trust governance layer that sits between AI agents and the outside world.
EvalBot — local-first chatbot security & quality evaluation (FastAPI + Next.js). Evaluate chatbot answers against your own docs & guidelines with ML/NLP + AI-judge scoring. Apache-2.0.
QVS-RAG
Eval-driven Customer Support FTE using OpenAI Agents SDK. Multi-agent routing, guardrails, and systematic quality evaluation.
AI数字人外呼多轮对话评测系统。将任务指令编译为状态机, 用覆盖率反馈驱动用户模拟,同等预算下P0风险发现率翻倍。
First-place official system for JOKER 2026 (CLEF) Task 1 English: a three-stage humor retrieval pipeline pairing hybrid sparse-dense retrieval and cross-encoder reranking with a rationale-distilled LLM judge ensemble. 0.6347 MAP. Team VANGUARD.
Benchmark evaluating LLM responses across 9 safety/quality dimensions — rule-based checks + validated LLM-judge ensemble, free-tier APIs
To associate your repository with the llm-judge topic, visit your repo's landing page and select "manage topics."