Skip to content

Latest commit

ย 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

2026 LLM Evaluation ๋…ผ๋ฌธ ์Šคํ„ฐ๋””

Index ์ฃผ์ œ ๋…ผ๋ฌธ/์ž๋ฃŒ ๋ฐœํ‘œ์ž ๋ฐœํ‘œ์ž๋ฃŒ & ์˜์ƒ
1 LLM ํ‰๊ฐ€ ๊ฐœ๊ด„ A Survey on Evaluation of Large Language Models [๋งํฌ]
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations [๋งํฌ]
๊น€๊ธฐ๋ฒ” ๐Ÿ“„ ๐ŸŽฅ
2 Long-Context Needle in a Haystack
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
InfiniteBench: Extending Long Context Evaluation Beyond 100K Tokens
RULER: Whatโ€™s the Real Context Size of Your Long-Context Language Models?
A Controllable Examination for Long-Context Language Models
์กฐ๋™ํ—Œ ๐Ÿ“„ ๐ŸŽฅ
2-1 Long-Context(Satellite) LongBench pro [๋งํฌ] ๊น€๊ธฐ๋ฒ” ๐Ÿ“„ ๐ŸŽฅ
3 ์ง€ํ‘œ ๋ถ•๊ดด, Goodhart's law The Leaderboard Illusion [๋งํฌ]
Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models [๋งํฌ]
ํ•œ์™„๊ทœ ๐Ÿ“„ ๐ŸŽฅ
4 Software Engineering Evaluating Large Language Models Trained on Code (HumanEval)[๋งํฌ]
SWE-bench[๋งํฌ]
SWE-bench Verified[๋งํฌ]
Multi-SWE-bench[๋งํฌ]
SWE-Bench Illusion[๋งํฌ]
SWE-rebench[๋งํฌ]
SWE-bench Pro[๋งํฌ]
๋ฐ•์ง„์šฐ ๐Ÿ“„ ๐ŸŽฅ
4-1 Tuning Coding Agents Improving Deep Agents with harness engineering[๋งํฌ] ๊น€๊ธฐ๋ฒ” ๐Ÿ“„ ๐ŸŽฅ
5 Agents - End to End GAIA: A Benchmark for General AI Assistants [๋งํฌ]
WebArena: A Realistic Web Environment for Building Autonomous Agents [๋งํฌ]
An Illusion of Progress? Assessing the Current State of Web Agents(Online-Mind2Web) [๋งํฌ]
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering[๋งํฌ]
๊น€๋™ํ˜„
6 Knowledge GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks [๋งํฌ]
GPQA: A Graduate-Level Google-Proof Q&A Benchmark [๋งํฌ]
๊น€์ •ํ›ˆ
7 Textual & Content Safety HarmBench: A Standardized Evaluation Framework for Automated Red Teaming[๋งํฌ]
A StrongREJECT for Empty Jailbreaks[๋งํฌ]
TrustLLM: Trustworthiness in LLMs [๋งํฌ]
ํ™์†Œํ˜„ ๐ŸŽฅ
_ Pi-mono, harness engineering Pi Monorepo: Tools for building AI agents.[๋งํฌ] Sigrid Jin ๐ŸŽฅ
8 Agentic & Behavioral Safety Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents [๋งํฌ]
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents[๋งํฌ]
sudo rm -rf agentic_security [๋งํฌ]
์ด๋™๊ฑด ๐Ÿ“„๐ŸŽฅ
9 LLM-as-a-Judge Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models [๋งํฌ]
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena [๋งํฌ]
JudgeBench: A Benchmark for Evaluating Judges [๋งํฌ]
์กฐ์„ฑ๊ตญ ๐Ÿ“„๐ŸŽฅ
10 Agents - Tool Use AGENTBENCH: Evaluating LLMs as Agents[๋งํฌ]
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models[๋งํฌ]
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers[๋งํฌ]
๊น€๊ฐ•๋ฏผ ๐Ÿ“„
- Twenty Questions Benchmark Twenty Questions Benchmark[๋งํฌ] ๊น€๊ธฐ๋ฒ” ๐ŸŽฅ
11 Thinking Process & Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning[๋งํฌ]
Evaluating Mathematical Reasoning Beyond Accuracy[๋งํฌ]
๋ฐ•์ง„ํ˜• ๐ŸŽฅ
12 Multimodal Reasoning MMMU Benchmark [๋งํฌ]
Humanity's Last Exam [๋งํฌ]
์ตœ๋™ํ˜ ๐Ÿ“„๐ŸŽฅ

About

2026 llm evaluation study

Resources

Stars

26 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors