← All benchmarks

Alpheva SWE: Software Engineering Benchmark

Software engineering agents in real repositories.

Can an agent sustain a complete engineering workflow, not just generate a patch?
GPT-6 Astra Max(OpenAI)
82.6%
Claude Fable 5.1 Max(Anthropic)
82.1%
Claude Opus 5 Max(Anthropic)
80.4%
Claude Opus 5 High(Anthropic)
79.8%
Kimi K3 Max(Moonshot AI)
78.7%
Qwen 3.8 Max(Alibaba)
78.1%
Claude Fable 5 High(Anthropic)
77.6%
Muse Spark 1.3 Max(Meta)
76.9%
GPT-5.6 Sol xHigh(OpenAI)
76.5%
Qwen 3.8 Flash Next(Alibaba)
75.8%
Hy4 Preview(Tencent)
75.2%
Grok 4.6 High(xAI)
74.7%
GLM-5.3 Max(Zhipu AI)
74.2%
DeepSeek V4.1 Flash Max(DeepSeek)
73.8%
GLM-5.3 Flash(Zhipu AI)
72.9%
Mean

Overview

Alpheva SWE evaluates autonomous coding agents across repository navigation, implementation, testing, debugging, and review. It measures whether an agent can recover from errors and complete a verified workflow in a versioned environment.

Evaluation focus

The benchmark is organized around the capabilities that determine whether a model can complete the full workflow reliably.

  1. 01Repository navigation
  2. 02Implementation and repair
  3. 03Testing and regression safety
  4. 04Tool use and recovery

Protocol

Protocol
Alpheva v0.1
Environment
Versioned repositories
Scope
End-to-end workflows

Model results

ModelProviderMean
GPT-6 Astra MaxOpenAI82.6%
Claude Fable 5.1 MaxAnthropic82.1%
Claude Opus 5 MaxAnthropic80.4%
Claude Opus 5 HighAnthropic79.8%
Kimi K3 MaxMoonshot AI78.7%
Qwen 3.8 MaxAlibaba78.1%
Claude Fable 5 HighAnthropic77.6%
Muse Spark 1.3 MaxMeta76.9%
GPT-5.6 Sol xHighOpenAI76.5%
Qwen 3.8 Flash NextAlibaba75.8%
Hy4 PreviewTencent75.2%
Grok 4.6 HighxAI74.7%
GLM-5.3 MaxZhipu AI74.2%
DeepSeek V4.1 Flash MaxDeepSeek73.8%
GLM-5.3 FlashZhipu AI72.9%

Results and methodology

Scores are reported against the protocol and versioned environment defined above. Compare results within this benchmark and metric only.