| Index | ์ฃผ์ | ๋ ผ๋ฌธ/์๋ฃ | ๋ฐํ์ | ๋ฐํ์๋ฃ & ์์ |
|---|---|---|---|---|
| 1 | LLM ํ๊ฐ ๊ฐ๊ด | A Survey on Evaluation of Large Language Models [๋งํฌ] A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations [๋งํฌ] |
๊น๊ธฐ๋ฒ | ๐ ๐ฅ |
| 2 | Long-Context | Needle in a Haystack LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks InfiniteBench: Extending Long Context Evaluation Beyond 100K Tokens RULER: Whatโs the Real Context Size of Your Long-Context Language Models? A Controllable Examination for Long-Context Language Models |
์กฐ๋ํ | ๐ ๐ฅ |
| 2-1 | Long-Context(Satellite) | LongBench pro [๋งํฌ] | ๊น๊ธฐ๋ฒ | ๐ ๐ฅ |
| 3 | ์งํ ๋ถ๊ดด, Goodhart's law | The Leaderboard Illusion [๋งํฌ] Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models [๋งํฌ] |
ํ์๊ท | ๐ ๐ฅ |
| 4 | Software Engineering | Evaluating Large Language Models Trained on Code (HumanEval)[๋งํฌ] SWE-bench[๋งํฌ] SWE-bench Verified[๋งํฌ] Multi-SWE-bench[๋งํฌ] SWE-Bench Illusion[๋งํฌ] SWE-rebench[๋งํฌ] SWE-bench Pro[๋งํฌ] |
๋ฐ์ง์ฐ | ๐ ๐ฅ |
| 4-1 | Tuning Coding Agents | Improving Deep Agents with harness engineering[๋งํฌ] | ๊น๊ธฐ๋ฒ | ๐ ๐ฅ |
| 5 | Agents - End to End | GAIA: A Benchmark for General AI Assistants [๋งํฌ] WebArena: A Realistic Web Environment for Building Autonomous Agents [๋งํฌ] An Illusion of Progress? Assessing the Current State of Web Agents(Online-Mind2Web) [๋งํฌ] MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering[๋งํฌ] |
๊น๋ํ | |
| 6 | Knowledge | GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks [๋งํฌ] GPQA: A Graduate-Level Google-Proof Q&A Benchmark [๋งํฌ] |
๊น์ ํ | |
| 7 | Textual & Content Safety | HarmBench: A Standardized Evaluation Framework for Automated Red Teaming[๋งํฌ] A StrongREJECT for Empty Jailbreaks[๋งํฌ] TrustLLM: Trustworthiness in LLMs [๋งํฌ] |
ํ์ํ | ๐ฅ |
| _ | Pi-mono, harness engineering | Pi Monorepo: Tools for building AI agents.[๋งํฌ] | Sigrid Jin | ๐ฅ |
| 8 | Agentic & Behavioral Safety | Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents [๋งํฌ] AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents[๋งํฌ] sudo rm -rf agentic_security [๋งํฌ] |
์ด๋๊ฑด | ๐๐ฅ |
| 9 | LLM-as-a-Judge | Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models [๋งํฌ] Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena [๋งํฌ] JudgeBench: A Benchmark for Evaluating Judges [๋งํฌ] |
์กฐ์ฑ๊ตญ | ๐๐ฅ |
| 10 | Agents - Tool Use | AGENTBENCH: Evaluating LLMs as Agents[๋งํฌ] StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models[๋งํฌ] MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers[๋งํฌ] |
๊น๊ฐ๋ฏผ | ๐ |
| - | Twenty Questions Benchmark | Twenty Questions Benchmark[๋งํฌ] | ๊น๊ธฐ๋ฒ | ๐ฅ |
| 11 | Thinking Process & Reasoning | Measuring Faithfulness in Chain-of-Thought Reasoning[๋งํฌ] Evaluating Mathematical Reasoning Beyond Accuracy[๋งํฌ] |
๋ฐ์งํ | ๐ฅ |
| 12 | Multimodal Reasoning | MMMU Benchmark [๋งํฌ] Humanity's Last Exam [๋งํฌ] |
์ต๋ํ | ๐๐ฅ |
Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|
ย | ย | |||
ย | ย | |||
ย | ย | |||