ZOLT Model Test

48 AI models. The same 20 jobs. One strict ruler.

Run on the live ZOLT gateway · October 11, 2026 · 960 calls · Total measured cost: 41.91 credits


The question

Which AI model is cheap and good? Every model on ZOLT took the same 20 jobs — 5 writing, 5 questions, 5 code, 5 logic — at temperature 0. Every answer was scored by rule, no human judging, and the code answers were actually run and tested. Same jobs, same ruler, no favorites.

Winners

Perfect score — 20/20

MiMo V2.6 Flash · GLM 5.3 Flash · GLM 5.3 FlashX

Cheapest good model — all 4 categories

MiMo V2.6 Flash and GLM 5.3 Flash (5/5 in Writing, Questions, Code, and Logic, at about 0.001 credits per call)

Cheapest strong all-rounder

GPT-6 Luna — 18/20 with perfect Code and Logic, about 0.001 credits per call

Most expensive per call

Claude Sonnet 5 — 0.356 credits per call, scored 17/20

What the test showed

Full results

#ModelScoreWritingQuestionsCodeLogicMedian timeCredits / call
1MiMo v2.6 Flash20/205/55/55/55/54.8s0.001
2GLM 5.3 Flash20/205/55/55/55/54.3s0.001
3GLM 5.3 FlashX20/205/55/55/55/54.8s0.005
4HY4 Preview19/205/54/55/55/54.5s0.006
5Kimi K2.819/205/54/55/55/512.9s0.012
6GPT-6 Luna18/204/54/55/55/53.5s0.001
7DeepSeek V4.1 Flash18/204/54/55/55/54.4s0.003
8Qwen 3.8 Omni Flash18/204/54/55/55/54.5s0.003
9Grok 4.618/204/54/55/55/54.3s0.004
10Gemini 3.7 Flash18/204/54/55/55/53.9s0.005
11GLM 5.118/204/54/55/55/53.6s0.006
12GPT-6 Sol18/204/54/55/55/55.0s0.007
13MiniMax M318/204/54/55/55/54.2s0.008
14DeepSeek V4 Pro18/204/54/55/55/54.1s0.009
15Claude Haiku 4.518/204/54/55/55/55.0s0.010
16Qwen 3.7 Max18/204/54/55/55/56.1s0.019
17GPT-5.518/204/54/55/55/54.0s0.019
18Claude Opus 518/204/54/55/55/53.9s0.019
19Claude Opus 4.818/204/54/55/55/53.7s0.020
20Qwen 3.8 Max18/204/54/55/55/54.5s0.020
21GLM 5.218/204/54/55/55/54.8s0.020
22MiniMax M2.718/204/54/55/55/510.3s0.022
23Gemini 3.8 Flash18/204/54/55/55/56.4s0.024
24Claude Opus 4.718/204/54/55/55/54.5s0.024
25Qwen 3.8 Max Preview18/204/54/55/55/56.0s0.026
26GLM 5.318/204/54/55/55/55.8s0.031
27GPT-5.6 Terra18/204/54/55/55/55.4s0.033
28Grok 4.718/204/54/55/55/54.5s0.034
29Kimi K318/204/54/55/55/57.9s0.041
30Grok 4.518/204/54/55/55/55.3s0.047
31GPT-6.1 Sol18/204/54/55/55/55.0s0.067
32Gemini 3.1 Pro Preview18/204/54/55/55/55.8s0.068
33Claude Sonnet 4.618/204/54/55/55/54.7s0.090
34Claude Sonnet 5.518/204/54/55/55/56.0s0.106
35GPT-5.6 Sol18/204/54/55/55/513.5s0.111
36GPT-6 Astra18/204/54/55/55/55.7s0.148
37Claude Fable 518/204/54/55/55/56.3s0.170
38Claude Opus 5.518/204/54/55/55/55.0s0.212
39Kimi K2.617/203/54/55/55/54.8s0.012
40GPT-5.617/203/54/55/55/546.8s0.066
41Gemini 3 Pro Preview17/203/54/55/55/57.1s0.082
42Claude Sonnet 517/204/54/55/54/56.3s0.356
43Kimi K2.7 Code16/203/54/54/55/54.7s0.036
44Grok 4.315/204/54/54/53/53.7s0.002
45HY315/202/54/54/55/55.6s0.003
46MiMo v2.514/203/54/54/53/57.2s0.003
47GPT-5.4 Mini10/203/54/51/52/55.7s0.006
48GPT-5.410/203/53/51/53/54.3s0.019

How it was scored

Writing jobs were checked for required points, word limits, and banned words. Questions each have one right answer, checked exactly (including the unit where the job asked for one). Code answers were executed in a locked-down sandbox and tested against the expected behavior. Logic answers were matched exactly. One retry pass was run for 6 calls that errored on the first pass (merchant-side failures, including all 5 of one model's writing calls); the retried answers are the ones scored above. Costs are computed from each call's real token usage at ZOLT's live credit rates and cross-checked against the gateway balance: 41.91 credits burned in total, measured. Full raw results (every answer, every score) are available on request.