Skip to content
#

truthfulqa

Here are 19 public repositories matching this topic...

Enterprise-grade LLM evaluation framework | Multi-model benchmarking, honest dashboards, system profiling | Academic metrics: MMLU, TruthfulQA, HellaSwag | Zero fake data | PyPI: llm-benchmark-toolkit | Blog: https://dev.to/nahuelgiudizi/building-an-honest-llm-evaluation-framework-from-fake-metrics-to-real-benchmarks-2b90

  • Updated Jul 22, 2026
  • Python

Does instruction tuning make language models more sycophantic? A paired causal study across Qwen, Llama, and Gemma on TruthfulQA, showing the effect is family-dependent in both magnitude and direction. 7,200 evaluations, 12 ATE estimates with paired t-tests and bootstrap CIs.

  • Updated May 29, 2026
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the truthfulqa topic, visit your repo's landing page and select "manage topics."

Learn more