
Measuring model performance on economically valuable work
What GDPval reveals about evaluating frontier models on realistic professional deliverables across major industries.
Research
Research we follow
We follow research grounded in expert judgment, economically meaningful tasks, and reliable performance across complete professional workflows.

What GDPval reveals about evaluating frontier models on realistic professional deliverables across major industries.

How teams can design agent evaluations that measure outcomes, survive non-determinism, and improve alongside the product.

How preference data, reward modeling, and reinforcement learning helped smaller models produce summaries people preferred.
External benchmarks
We study leading independent benchmarks to understand how complex capabilities are measured, where current approaches fall short, and what real-world performance demands. These insights shape the benchmarks and evaluation environments we develop at Alpheva.
A useful reference for cross-application workflows, stateful tasks, and execution-based evaluation in real computer environments.
02A benchmark for evaluating whether AI systems can resolve real-world GitHub issues by modifying codebases and passing repository tests.
03A benchmark of hard, realistic terminal tasks with purpose-built environments, human-written solutions, and comprehensive test-based verification.
Work with Alpheva
We collaborate on evaluation and training questions grounded in real work.
Build with us