Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
dashboard mlops github-actions huggingface azure-openai llm-evaluation benchmark-automation llm-benchmark gdpval real-world-tasks professional-tasks artifact-validation
-
Updated
Jul 29, 2026 - Python