← All benchmarks

Alpheva Legal: Legal Reasoning Benchmark

Legal reasoning across evolving matters.

Can a model revise a grounded legal work product as the record evolves?
Claude Opus 5 Max(Anthropic)
62.1%
Claude Fable 5.1 Max(Anthropic)
60.8%
Claude Fable 5(Anthropic)
55.7%
Kimi K3 Max(Moonshot AI)
54.1%
Grok 4.6 High(xAI)
53.8%
GPT-5.6 Sol Max(OpenAI)
51.1%
Muse Spark 1.3 xHigh(Meta)
50.3%
Gemini 3.8 Flash High(Google)
47.9%
Grok 4.5 High(xAI)
43.2%
Muse Spark 1.1 xHigh(Meta)
42.3%
Claude Sonnet 5(Anthropic)
40.9%
Claude Opus 4.8(Anthropic)
38.2%
Claude Opus 4.7(Anthropic)
35.8%
GPT-5.5(OpenAI)
35.1%
Gemini 3.1 Pro(Google)
21.9%
Inkling Medium(Thinking Machines)
21.9%
Mean

Overview

Alpheva Legal evaluates how models identify controlling issues, apply authorities to the record, revise conclusions as facts and arguments change, and ground legal work in text and visual evidence across multi-stage matters.

Evaluation focus

The benchmark is organized around the capabilities that determine whether a model can complete the full workflow reliably.

  1. 01Issue, rule, application, and conclusion
  2. 02Revision across multi-stage matters
  3. 03Grounding in legal authorities and records
  4. 04Reasoning over visual exhibits

Protocol

Protocol
Alpheva v0.1
Materials
Authorities + case files
Scope
Multi-stage matters

Model results

ModelProviderMean
Claude Opus 5 MaxAnthropic62.1%
Claude Fable 5.1 MaxAnthropic60.8%
Claude Fable 5Anthropic55.7%
Kimi K3 MaxMoonshot AI54.1%
Grok 4.6 HighxAI53.8%
GPT-5.6 Sol MaxOpenAI51.1%
Muse Spark 1.3 xHighMeta50.3%
Gemini 3.8 Flash HighGoogle47.9%
Grok 4.5 HighxAI43.2%
Muse Spark 1.1 xHighMeta42.3%
Claude Sonnet 5Anthropic40.9%
Claude Opus 4.8Anthropic38.2%
Claude Opus 4.7Anthropic35.8%
GPT-5.5OpenAI35.1%
Gemini 3.1 ProGoogle21.9%
Inkling MediumThinking Machines21.9%

Results and methodology

Scores are reported against the protocol and versioned environment defined above. Compare results within this benchmark and metric only.