Image
AI-assisted software engineering has seen the emergence of several benchmarks to measure the capabilities of LLMs. Android developers face specific challenges that aren't covered by existing benchmarks, so we created one that focuses on a north star of high quality Android development.
Model Score (%) info Average percentage of 100 test cases successfully resolved across 5 runs for each model
arrow_range Cl range (%) info Expected performance range, reflecting the results' statistical reliability (p-value < 0.05)
Avg latency (h) info Average time taken to solve 100 tasks across 5 runs
Avg cost ($) info Average cost per full benchmark run
Image Claude Opus 5 91.8
87.4 — 95.7 13.2 $175.4
Image Claude Fable 5 90.9
86.2 — 95.4 8.7 $144.3
Image GPT 5.6 Sol 90.8
85.6 — 95.6 14.2 $150.0
Image Kimi K3 90.2
85.0 — 94.6 29.6 $112.9
Image GPT 5.6 Luna 87.6
81.8 — 92.3 10.6 $7.2
Image Qwen3.8 Max 87.0
82.2 — 91.6 41.2 $196.7
Image GPT 5.6 Terra 86.8
81.7 — 91.2 7.6 $48.0
Image GPT 5.5 80.2
73.0 — 86.6 11.4 $138.3
Image Claude Sonnet 5 76.2
68.8 — 82.5 12.3 $99.9
Image Gemini 3.6 Flash 75.6
67.5 — 82.1 10.2 $135.5
Image Gemini 3.1 Pro Preview 74.3
66.8 — 81.5 10.5 $86.6
Image GPT 5.4 74.1
66.6 — 81.2 8.4 $83.4
Image Claude Opus 4.8 72.4
65.1 — 79.2 6.7 $88.0
Image GLM 5.2 72.2
65.3 — 79.1 38.9 $117.0
Image Gemini 3.5 Flash 71.1
63.8 — 78.5 28.3 $165.6
Image Kimi K2.7 Code 70.4
63.4 — 77.1 31.8 $48.1
Image Claude Opus 4.7 68.7
60.9 — 76.4 7.0 $96.5
Image Kimi K2.6 67.6
59.8 — 74.8 57.2 $49.4
Image Claude Sonnet 4.6 67.0
58.5 — 75.1 16.9 $127.6
Image Minimax M3 63.6
56.1 — 70.7 26.0 $41.7
Image GLM 5.1 63.2
55.3 — 70.2 17.6 $53.5
Image Gemini 3 Flash Preview 62.5
54.7 — 69.7 13.1 $30.1
Image Mimo V2.5 Pro 60.8
53.2 — 68.6 13.6 $9.2
Image Deepseek V4 Pro 59.5
51.4 — 67.9 9.0 $3.7
Image Qwen3.7 Plus 57.7
50.2 — 65.6 18.5 $18.6
Image Deepseek V4 Flash 54.7
46.9 — 62.8 8.9 $1.5
Image Qwen3.7 Max 54.2
46.9 — 62.1 14.2 $58.3
Image Gemini 3.5 Flash-Lite 50.6
42.3 — 58.4 5.4 $34.1
Image Qwen3.6 27b 45.1
37.6 — 52.8 25.8 $97.3
Image MiniMax M2.7 41.6
34.7 — 49.4 18.2 $14.9
Image Gemma 4 31b IT 37.1
29.9 — 44.5 36.3 $10.4
Image Qwen3.6 35b A3b 37.0
29.1 — 44.3 16.3 $17.8
Latest results as of August 11th.
View archived leaderboards and check back periodically for updates.
Track the latest AI model benchmarks, newly introduced agent architectures, and continuous performance evaluations on the platform. Stay updated with our routine methodology updates and release logs.
  • New models • Aug 11th
    Image Claude Opus 5
  • New models • Aug 11th
    Image GPT 5.6 Sol, GPT 5.6 Luna, GPT 5.6 Terra
  • New models • Aug 11th
    Image Kimi K3
  • New models • Aug 11th
    Image Qwen3.8 Max
  • New models • Aug 11th
    Image Gemini 3.6 Flash, Gemini 3.5 Flash Lite
  • Archived models • Aug 11th
    Image Gemma 4 26B A4B IT
Image

Learn more about Android Bench

Learn more about how we created a set of common Android developer tasks.
Many of the tasks are based on how we define high quality Android development, which is detailed in our developer documentation.
See the full dataset on Harbor dashboard.