Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
-
Updated
Oct 1, 2026 - Python
Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
GitHub Release asset quality gate and npm CLI: catch missing platforms, wrong architectures, empty installers, version drift, and missing checksums.
Worked as a project-based AI Agent Specialist with Handshake AI through the Dynamo project, working on structured software engineering tasks involving AI coding agents, GitHub repositories, task execution, testing, validation, and evaluation.
Python CLI for paper-reproduction workflows with PDF extraction, artifact manifests, opt-in provider readiness, human-gated experiment scaffolds, benchmark suites, run comparison, reports, and agent handoff.
Reusable scientific workflow skills with explicit artifact contracts, validators, and local repair.
Lightweight, agent-friendly inspection and contract validation for bioinformatics artifacts.
Evaluate an operation’s artifacts against explicit completion rules and retain a structured result/resume record when verification is incomplete.
Mechanistic Workbench (mwb): Local-first mechanistic interpretability workbench for IPython research, agent-readable state, artifact validation, evidence graphs, claim-safe MechanismCards, provenance, run ledgers, SAE/TransformerLens workflows, and reproducible MI experiments
Download Trimble RealWorks for Windows 10/11 (64-bit) and install it as Administrator for a complete setup experience.
To associate your repository with the artifact-validation topic, visit your repo's landing page and select "manage topics."