GDPval shifts model evaluation away from short, exam-like questions and toward the artifacts professionals actually produce. The benchmark asks whether an AI system can create useful work products, not merely recover the right fact, and uses experienced practitioners to judge the result.
01
Why real work changes the evaluation problem
Traditional benchmarks are useful for isolating capabilities, but they rarely resemble the work people complete inside an organization. GDPval starts with economically significant industries and knowledge-work occupations, then evaluates models on deliverables such as legal briefs, engineering concepts, care plans, presentations, spreadsheets, and customer-support responses.
This makes quality multidimensional. A correct answer can still be a poor deliverable if it is incomplete, badly structured, visually confusing, or impractical. Evaluating finished artifacts therefore requires judgment about accuracy, usability, reasoning, and presentation together.
02
How the benchmark was constructed
The first release spans 44 occupations in nine industries that contribute substantially to United States GDP. Occupations were filtered toward knowledge work and selected using labor and wage data. Experienced professionals then designed tasks based on the work products common in their fields.
Each occupation contributes 30 fully reviewed tasks to the complete set. A 220-task gold subset (five tasks per occupation) was released for wider use. Task authors averaged more than 14 years of professional experience, and tasks went through repeated reviews for clarity, realism, and feasibility.
- Tasks include the context and reference files a professional would expect.
- Outputs can be documents, slides, diagrams, spreadsheets, or multimedia, not text alone.
- The benchmark emphasizes realistic work products rather than synthetic exam questions.
03
Expert comparison instead of a single answer key
GDPval relies on professionals from the relevant occupations to compare model-generated work with human-produced reference work. Reviewers are blinded to the source and judge whether an AI deliverable is better than, comparable to, or worse than the human artifact. Detailed occupational rubrics add consistency across comparisons.
OpenAI also developed an automated grader intended to approximate expert preference. It can make repeated experimentation faster, but the publication treats it as an experimental aid rather than a replacement for domain experts. That distinction matters whenever quality depends on tacit professional standards.
04
What the early results suggest
In blind comparisons on the open gold set, leading frontier models were approaching expert-quality work on a meaningful share of tasks. Different systems showed different strengths: some produced stronger visual presentation, while others were more accurate on domain-specific details. The results reinforce the value of inspecting capability by task and quality dimension rather than relying on one aggregate score.
The study also reported large differences in raw inference time and cost compared with expert production. Those figures do not include the human review, iteration, integration, or accountability required in a workplace, so they are best read as evidence of potential leverage, not as a complete model of deployment economics.
05
The limits of one-shot professional tasks
GDPval is an early benchmark. Its tasks are largely one-shot and clearly specified, while real professional work often unfolds through discovery, feedback, revision, and coordination. A strong score does not show that a model can manage an ambiguous project over days or recognize when the requested deliverable is the wrong one.
The next evaluation frontier is therefore interactive: persistent context, changing requirements, tool use, revision after critique, and consequences that emerge across a workflow. GDPval provides a strong foundation, but it also makes the missing dimensions visible.
Alpheva perspective
What we take from the work
- Evaluate the artifact people use, not only the model response that produced it.
- Recruit domain experts to define tasks and calibrate quality judgments.
- Separate raw generation efficiency from the total cost of reliable deployment.
- Extend one-shot tests with iterative, stateful workflows before drawing broad conclusions about job-level capability.
