← All benchmarks
Alpheva Legal: Legal Reasoning Benchmark
Legal reasoning across evolving matters.
Can a model revise a grounded legal work product as the record evolves?Claude Opus 5 Max(Anthropic)
62.1%Claude Fable 5.1 Max(Anthropic)
60.8%Claude Fable 5(Anthropic)
55.7%Kimi K3 Max(Moonshot AI)
54.1%Grok 4.6 High(xAI)
53.8%GPT-5.6 Sol Max(OpenAI)
51.1%Muse Spark 1.3 xHigh(Meta)
50.3%Gemini 3.8 Flash High(Google)
47.9%Grok 4.5 High(xAI)
43.2%Muse Spark 1.1 xHigh(Meta)
42.3%Claude Sonnet 5(Anthropic)
40.9%Claude Opus 4.8(Anthropic)
38.2%Claude Opus 4.7(Anthropic)
35.8%GPT-5.5(OpenAI)
35.1%Gemini 3.1 Pro(Google)
21.9%Inkling Medium(Thinking Machines)
21.9%Mean
Overview
Alpheva Legal evaluates how models identify controlling issues, apply authorities to the record, revise conclusions as facts and arguments change, and ground legal work in text and visual evidence across multi-stage matters.
Evaluation focus
The benchmark is organized around the capabilities that determine whether a model can complete the full workflow reliably.
- 01Issue, rule, application, and conclusion
- 02Revision across multi-stage matters
- 03Grounding in legal authorities and records
- 04Reasoning over visual exhibits
Protocol
- Protocol
- Alpheva v0.1
- Materials
- Authorities + case files
- Scope
- Multi-stage matters
Model results
| Model | Provider | Mean |
|---|---|---|
| Claude Opus 5 Max | Anthropic | 62.1% |
| Claude Fable 5.1 Max | Anthropic | 60.8% |
| Claude Fable 5 | Anthropic | 55.7% |
| Kimi K3 Max | Moonshot AI | 54.1% |
| Grok 4.6 High | xAI | 53.8% |
| GPT-5.6 Sol Max | OpenAI | 51.1% |
| Muse Spark 1.3 xHigh | Meta | 50.3% |
| Gemini 3.8 Flash High | 47.9% | |
| Grok 4.5 High | xAI | 43.2% |
| Muse Spark 1.1 xHigh | Meta | 42.3% |
| Claude Sonnet 5 | Anthropic | 40.9% |
| Claude Opus 4.8 | Anthropic | 38.2% |
| Claude Opus 4.7 | Anthropic | 35.8% |
| GPT-5.5 | OpenAI | 35.1% |
| Gemini 3.1 Pro | 21.9% | |
| Inkling Medium | Thinking Machines | 21.9% |
Results and methodology
Scores are reported against the protocol and versioned environment defined above. Compare results within this benchmark and metric only.