Skip to content
#

llm-judge

Here are 112 public repositories matching this topic...

oh-my-knowledge

OMK — Evidence-backed evaluation and observability for prompts, RAG, skills, agents, and workflows. Native Codex, Claude Code, and DeepSeek Harness support.

  • Updated Sep 21, 2026
  • TypeScript

ProductionOS v1.0 — Claude Code plugin with 76 agents, 39 commands, and 12 hooks. Deploys specialized agents that review, score, and improve your entire codebase. Smart routing, recursive convergence, self-evaluation.

  • Updated Apr 16, 2026
  • TypeScript

Extensible benchmarking suite for evaluating AI coding agents on web search tasks. Compare native search vs MCP servers (You.com, expanding) across multiple agents (Claude Code, Gemini, Droid, Codex, expanding) with automated Docker workflows and statistical analysis.

  • Updated Feb 27, 2026
  • TypeScript

Autonomous judge agents fix your app's debt, build what's missing, and verify every change through a browser or gate suite before it lands, then find the next thing and keep going until you say stop. Each fixer owns a file-disjoint slice of the tree, so no two can ever collide, and nothing unverified reaches git. A Claude Code plugin.

  • Updated Jul 29, 2026
  • JavaScript

First-place official system for JOKER 2026 (CLEF) Task 1 English: a three-stage humor retrieval pipeline pairing hybrid sparse-dense retrieval and cross-encoder reranking with a rationale-distilled LLM judge ensemble. 0.6347 MAP. Team VANGUARD.

  • Updated Jun 20, 2026
  • Python

Add this topic to your repo

To associate your repository with the llm-judge topic, visit your repo's landing page and select "manage topics."

Learn more