You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
A task-level audit of the grading rubrics in OpenAI's GDPval. 10,453 criteria opened: score scales vary tenfold across tasks, and one untestable line carries up to 21.7% of a task's score.