← All benchmarks

Alpheva Finance: Financial Reasoning Benchmark

Financial reasoning under incomplete information.

Can a model make defensible financial decisions when information is incomplete?
GPT-5.5(OpenAI)
79.4%
Claude Opus 4.7(Anthropic)
78.9%
DeepSeek V4 Pro(DeepSeek)
77.5%
Gemini 3.1 Pro(Google)
76.8%
Kimi K2.6(Moonshot AI)
75.1%
Gemini 3.8 Flash(Google)
73.9%
Grok 4.3(xAI)
73.6%
GLM-5.1(Zhipu AI)
72.4%
Claude Fable 5(Anthropic)
71.5%
MiniMax M2.7(MiniMax)
69.6%
Claude Opus 5(Anthropic)
68.2%
GPT-5.6 Sol(OpenAI)
67.4%
Grok 4.6(xAI)
66.8%
GPT-5.6 Luna(OpenAI)
65.3%
Muse Spark 1.2(Meta)
64.8%
Claude Sonnet 5(Anthropic)
63.9%
Mean

Overview

Alpheva Finance evaluates how models interpret primary filings, construct defensible assumptions, revise analysis as evidence changes, and communicate uncertainty while following professional accounting standards.

Evaluation focus

The benchmark is organized around the capabilities that determine whether a model can complete the full workflow reliably.

  1. 01Primary-source financial analysis
  2. 02Logical and defensible assumptions
  3. 03Accounting-standard adherence
  4. 04Exact calculation accuracy

Protocol

Protocol
Alpheva v0.1
Materials
Filings + workpapers
Scope
Evidence-aware judgment

Model results

ModelProviderMean
GPT-5.5OpenAI79.4%
Claude Opus 4.7Anthropic78.9%
DeepSeek V4 ProDeepSeek77.5%
Gemini 3.1 ProGoogle76.8%
Kimi K2.6Moonshot AI75.1%
Gemini 3.8 FlashGoogle73.9%
Grok 4.3xAI73.6%
GLM-5.1Zhipu AI72.4%
Claude Fable 5Anthropic71.5%
MiniMax M2.7MiniMax69.6%
Claude Opus 5Anthropic68.2%
GPT-5.6 SolOpenAI67.4%
Grok 4.6xAI66.8%
GPT-5.6 LunaOpenAI65.3%
Muse Spark 1.2Meta64.8%
Claude Sonnet 5Anthropic63.9%

Results and methodology

Scores are reported against the protocol and versioned environment defined above. Compare results within this benchmark and metric only.