← All benchmarks
Alpheva SWE: Software Engineering Benchmark
Software engineering agents in real repositories.
Can an agent sustain a complete engineering workflow, not just generate a patch?GPT-6 Astra Max(OpenAI)
82.6%Claude Fable 5.1 Max(Anthropic)
82.1%Claude Opus 5 Max(Anthropic)
80.4%Claude Opus 5 High(Anthropic)
79.8%Kimi K3 Max(Moonshot AI)
78.7%Qwen 3.8 Max(Alibaba)
78.1%Claude Fable 5 High(Anthropic)
77.6%Muse Spark 1.3 Max(Meta)
76.9%GPT-5.6 Sol xHigh(OpenAI)
76.5%Qwen 3.8 Flash Next(Alibaba)
75.8%Hy4 Preview(Tencent)
75.2%Grok 4.6 High(xAI)
74.7%GLM-5.3 Max(Zhipu AI)
74.2%DeepSeek V4.1 Flash Max(DeepSeek)
73.8%GLM-5.3 Flash(Zhipu AI)
72.9%Mean
Overview
Alpheva SWE evaluates autonomous coding agents across repository navigation, implementation, testing, debugging, and review. It measures whether an agent can recover from errors and complete a verified workflow in a versioned environment.
Evaluation focus
The benchmark is organized around the capabilities that determine whether a model can complete the full workflow reliably.
- 01Repository navigation
- 02Implementation and repair
- 03Testing and regression safety
- 04Tool use and recovery
Protocol
- Protocol
- Alpheva v0.1
- Environment
- Versioned repositories
- Scope
- End-to-end workflows
Model results
| Model | Provider | Mean |
|---|---|---|
| GPT-6 Astra Max | OpenAI | 82.6% |
| Claude Fable 5.1 Max | Anthropic | 82.1% |
| Claude Opus 5 Max | Anthropic | 80.4% |
| Claude Opus 5 High | Anthropic | 79.8% |
| Kimi K3 Max | Moonshot AI | 78.7% |
| Qwen 3.8 Max | Alibaba | 78.1% |
| Claude Fable 5 High | Anthropic | 77.6% |
| Muse Spark 1.3 Max | Meta | 76.9% |
| GPT-5.6 Sol xHigh | OpenAI | 76.5% |
| Qwen 3.8 Flash Next | Alibaba | 75.8% |
| Hy4 Preview | Tencent | 75.2% |
| Grok 4.6 High | xAI | 74.7% |
| GLM-5.3 Max | Zhipu AI | 74.2% |
| DeepSeek V4.1 Flash Max | DeepSeek | 73.8% |
| GLM-5.3 Flash | Zhipu AI | 72.9% |
Results and methodology
Scores are reported against the protocol and versioned environment defined above. Compare results within this benchmark and metric only.