Skip to content
#

verbosity-bias

Here are 3 public repositories matching this topic...

Language: All
Filter by language

Code, per-item results and figures for a study dissociating prompt quality from response compliance in automated prompt-engineering assessment. Four-agent MATLAB evaluator on a locally hosted Qwen 2.5-7B judge, over 498 prompts from IFEval, LMSYS-Chat-1M and WildChat.

  • Updated Sep 19, 2026
  • MATLAB

SSIT: a label-free, gold-free test for whether LLM-judge position x verbosity bias corrections actually compose. No human labels, no model of the judge. Code, 7-judge/6-family pilot data, and pre-registered protocol for the NeurIPS 2026 JUDGe workshop paper.

  • Updated Sep 16, 2026
  • Python

Add this topic to your repo

To associate your repository with the verbosity-bias topic, visit your repo's landing page and select "manage topics."

Learn more