← All benchmarks
Alpheva Finance: Financial Reasoning Benchmark
Financial reasoning under incomplete information.
Can a model make defensible financial decisions when information is incomplete?GPT-5.5(OpenAI)
79.4%Claude Opus 4.7(Anthropic)
78.9%DeepSeek V4 Pro(DeepSeek)
77.5%Gemini 3.1 Pro(Google)
76.8%Kimi K2.6(Moonshot AI)
75.1%Gemini 3.8 Flash(Google)
73.9%Grok 4.3(xAI)
73.6%GLM-5.1(Zhipu AI)
72.4%Claude Fable 5(Anthropic)
71.5%MiniMax M2.7(MiniMax)
69.6%Claude Opus 5(Anthropic)
68.2%GPT-5.6 Sol(OpenAI)
67.4%Grok 4.6(xAI)
66.8%GPT-5.6 Luna(OpenAI)
65.3%Muse Spark 1.2(Meta)
64.8%Claude Sonnet 5(Anthropic)
63.9%Mean
Overview
Alpheva Finance evaluates how models interpret primary filings, construct defensible assumptions, revise analysis as evidence changes, and communicate uncertainty while following professional accounting standards.
Evaluation focus
The benchmark is organized around the capabilities that determine whether a model can complete the full workflow reliably.
- 01Primary-source financial analysis
- 02Logical and defensible assumptions
- 03Accounting-standard adherence
- 04Exact calculation accuracy
Protocol
- Protocol
- Alpheva v0.1
- Materials
- Filings + workpapers
- Scope
- Evidence-aware judgment
Model results
| Model | Provider | Mean |
|---|---|---|
| GPT-5.5 | OpenAI | 79.4% |
| Claude Opus 4.7 | Anthropic | 78.9% |
| DeepSeek V4 Pro | DeepSeek | 77.5% |
| Gemini 3.1 Pro | 76.8% | |
| Kimi K2.6 | Moonshot AI | 75.1% |
| Gemini 3.8 Flash | 73.9% | |
| Grok 4.3 | xAI | 73.6% |
| GLM-5.1 | Zhipu AI | 72.4% |
| Claude Fable 5 | Anthropic | 71.5% |
| MiniMax M2.7 | MiniMax | 69.6% |
| Claude Opus 5 | Anthropic | 68.2% |
| GPT-5.6 Sol | OpenAI | 67.4% |
| Grok 4.6 | xAI | 66.8% |
| GPT-5.6 Luna | OpenAI | 65.3% |
| Muse Spark 1.2 | Meta | 64.8% |
| Claude Sonnet 5 | Anthropic | 63.9% |
Results and methodology
Scores are reported against the protocol and versioned environment defined above. Compare results within this benchmark and metric only.