Open leaderboards

Benchmarks built for work that has consequences.

Compare how leading AI models perform across realistic software engineering, financial reasoning, and legal analysis tasks.

01

Alpheva SWE: Software Engineering Benchmark

Software engineering agents in real repositories.

Alpheva SWE evaluates autonomous coding agents across repository navigation, implementation, testing, debugging, and review. It measures whether an agent can recover from errors and complete a verified workflow in a versioned environment.

View benchmark details →
ModelMean
GPT-6 Astra Max (OpenAI)
82.6%
Claude Fable 5.1 Max (Anthropic)
82.1%
Claude Opus 5 Max (Anthropic)
80.4%
Claude Opus 5 High (Anthropic)
79.8%
Kimi K3 Max (Moonshot AI)
78.7%
Qwen 3.8 Max (Alibaba)
78.1%
Claude Fable 5 High (Anthropic)
77.6%
View all 15 models →
02

Alpheva Finance: Financial Reasoning Benchmark

Financial reasoning under incomplete information.

Alpheva Finance evaluates how models interpret primary filings, construct defensible assumptions, revise analysis as evidence changes, and communicate uncertainty while following professional accounting standards.

View benchmark details →
ModelMean
GPT-5.5 (OpenAI)
79.4%
Claude Opus 4.7 (Anthropic)
78.9%
DeepSeek V4 Pro (DeepSeek)
77.5%
Gemini 3.1 Pro (Google)
76.8%
Kimi K2.6 (Moonshot AI)
75.1%
Gemini 3.8 Flash (Google)
73.9%
Grok 4.3 (xAI)
73.6%
View all 16 models →
03

Alpheva Legal: Legal Reasoning Benchmark

Legal reasoning across evolving matters.

Alpheva Legal evaluates how models identify controlling issues, apply authorities to the record, revise conclusions as facts and arguments change, and ground legal work in text and visual evidence across multi-stage matters.

View benchmark details →
ModelMean
Claude Opus 5 Max (Anthropic)
62.1%
Claude Fable 5.1 Max (Anthropic)
60.8%
Claude Fable 5 (Anthropic)
55.7%
Kimi K3 Max (Moonshot AI)
54.1%
Grok 4.6 High (xAI)
53.8%
GPT-5.6 Sol Max (OpenAI)
51.1%
Muse Spark 1.3 xHigh (Meta)
50.3%
View all 16 models →

Work with Alpheva

Evaluate the capability your product depends on.

Work with Alpheva to design an open benchmark or a private enterprise evaluation.

Design a benchmark