48 AI models. The same 20 jobs. One strict ruler.
Which AI model is cheap and good? Every model on ZOLT took the same 20 jobs — 5 writing, 5 questions, 5 code, 5 logic — at temperature 0. Every answer was scored by rule, no human judging, and the code answers were actually run and tested. Same jobs, same ruler, no favorites.
MiMo V2.6 Flash · GLM 5.3 Flash · GLM 5.3 FlashX
MiMo V2.6 Flash and GLM 5.3 Flash (5/5 in Writing, Questions, Code, and Logic, at about 0.001 credits per call)
GPT-6 Luna — 18/20 with perfect Code and Logic, about 0.001 credits per call
Claude Sonnet 5 — 0.356 credits per call, scored 17/20
| # | Model | Score | Writing | Questions | Code | Logic | Median time | Credits / call |
|---|---|---|---|---|---|---|---|---|
| 1 | MiMo v2.6 Flash | 20/20 | 5/5 | 5/5 | 5/5 | 5/5 | 4.8s | 0.001 |
| 2 | GLM 5.3 Flash | 20/20 | 5/5 | 5/5 | 5/5 | 5/5 | 4.3s | 0.001 |
| 3 | GLM 5.3 FlashX | 20/20 | 5/5 | 5/5 | 5/5 | 5/5 | 4.8s | 0.005 |
| 4 | HY4 Preview | 19/20 | 5/5 | 4/5 | 5/5 | 5/5 | 4.5s | 0.006 |
| 5 | Kimi K2.8 | 19/20 | 5/5 | 4/5 | 5/5 | 5/5 | 12.9s | 0.012 |
| 6 | GPT-6 Luna | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 3.5s | 0.001 |
| 7 | DeepSeek V4.1 Flash | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.4s | 0.003 |
| 8 | Qwen 3.8 Omni Flash | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.5s | 0.003 |
| 9 | Grok 4.6 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.3s | 0.004 |
| 10 | Gemini 3.7 Flash | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 3.9s | 0.005 |
| 11 | GLM 5.1 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 3.6s | 0.006 |
| 12 | GPT-6 Sol | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.0s | 0.007 |
| 13 | MiniMax M3 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.2s | 0.008 |
| 14 | DeepSeek V4 Pro | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.1s | 0.009 |
| 15 | Claude Haiku 4.5 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.0s | 0.010 |
| 16 | Qwen 3.7 Max | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 6.1s | 0.019 |
| 17 | GPT-5.5 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.0s | 0.019 |
| 18 | Claude Opus 5 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 3.9s | 0.019 |
| 19 | Claude Opus 4.8 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 3.7s | 0.020 |
| 20 | Qwen 3.8 Max | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.5s | 0.020 |
| 21 | GLM 5.2 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.8s | 0.020 |
| 22 | MiniMax M2.7 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 10.3s | 0.022 |
| 23 | Gemini 3.8 Flash | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 6.4s | 0.024 |
| 24 | Claude Opus 4.7 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.5s | 0.024 |
| 25 | Qwen 3.8 Max Preview | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 6.0s | 0.026 |
| 26 | GLM 5.3 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.8s | 0.031 |
| 27 | GPT-5.6 Terra | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.4s | 0.033 |
| 28 | Grok 4.7 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.5s | 0.034 |
| 29 | Kimi K3 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 7.9s | 0.041 |
| 30 | Grok 4.5 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.3s | 0.047 |
| 31 | GPT-6.1 Sol | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.0s | 0.067 |
| 32 | Gemini 3.1 Pro Preview | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.8s | 0.068 |
| 33 | Claude Sonnet 4.6 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 4.7s | 0.090 |
| 34 | Claude Sonnet 5.5 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 6.0s | 0.106 |
| 35 | GPT-5.6 Sol | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 13.5s | 0.111 |
| 36 | GPT-6 Astra | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.7s | 0.148 |
| 37 | Claude Fable 5 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 6.3s | 0.170 |
| 38 | Claude Opus 5.5 | 18/20 | 4/5 | 4/5 | 5/5 | 5/5 | 5.0s | 0.212 |
| 39 | Kimi K2.6 | 17/20 | 3/5 | 4/5 | 5/5 | 5/5 | 4.8s | 0.012 |
| 40 | GPT-5.6 | 17/20 | 3/5 | 4/5 | 5/5 | 5/5 | 46.8s | 0.066 |
| 41 | Gemini 3 Pro Preview | 17/20 | 3/5 | 4/5 | 5/5 | 5/5 | 7.1s | 0.082 |
| 42 | Claude Sonnet 5 | 17/20 | 4/5 | 4/5 | 5/5 | 4/5 | 6.3s | 0.356 |
| 43 | Kimi K2.7 Code | 16/20 | 3/5 | 4/5 | 4/5 | 5/5 | 4.7s | 0.036 |
| 44 | Grok 4.3 | 15/20 | 4/5 | 4/5 | 4/5 | 3/5 | 3.7s | 0.002 |
| 45 | HY3 | 15/20 | 2/5 | 4/5 | 4/5 | 5/5 | 5.6s | 0.003 |
| 46 | MiMo v2.5 | 14/20 | 3/5 | 4/5 | 4/5 | 3/5 | 7.2s | 0.003 |
| 47 | GPT-5.4 Mini | 10/20 | 3/5 | 4/5 | 1/5 | 2/5 | 5.7s | 0.006 |
| 48 | GPT-5.4 | 10/20 | 3/5 | 3/5 | 1/5 | 3/5 | 4.3s | 0.019 |
Writing jobs were checked for required points, word limits, and banned words. Questions each have one right answer, checked exactly (including the unit where the job asked for one). Code answers were executed in a locked-down sandbox and tested against the expected behavior. Logic answers were matched exactly. One retry pass was run for 6 calls that errored on the first pass (merchant-side failures, including all 5 of one model's writing calls); the retried answers are the ones scored above. Costs are computed from each call's real token usage at ZOLT's live credit rates and cross-checked against the gateway balance: 41.91 credits burned in total, measured. Full raw results (every answer, every score) are available on request.