v1.1: Judge calibration improvements
Updated judges and rubrics to fix accidental rewarding of misleading responses. This makes the benchmark significantly more difficult overall, with Anthropic models especially impacted.
Last updated: September 22, 2026
A frontier benchmark for complex data work and analytical reasoning
DataBench score
| Model | Average cost per task | DataBench score |
|---|---|---|
| GPT-5.6 Luna XHigh | $0.04 | 35.3% |
| GPT-5.6 Luna High | $0.03 | 31.3% |
| GPT-5.6 Luna Medium | $0.02 | 24% |
| GPT-5.6 Luna Low | $0.01 | 19% |
| GPT-5.6 Terra XHigh | $0.30 | 39.3% |
| GPT-5.6 Terra High | $0.20 | 32.7% |
| GPT-5.6 Terra Medium | $0.15 | 29.3% |
| GPT-5.6 Terra Low | $0.14 | 27% |
| GPT-5.6 Sol XHigh | $0.98 | 40.3% |
| GPT-5.6 Sol High | $0.75 | 38% |
| GPT-5.6 Sol Medium | $0.61 | 39.7% |
| GPT-5.6 Sol Low | $0.37 | 29.7% |
| GPT-6 Astra Max | $2.11 | 55% |
| GPT-6 Astra XHigh | $1.75 | 52.3% |
| GPT-6 Astra High | $1.62 | 52% |
| GPT-6 Astra Medium | $1.24 | 52.7% |
| GPT-6 Astra Low | $0.85 | 55% |
| Sonnet 5 Max | $1.07 | 12.3% |
| Sonnet 5 XHigh | $0.62 | 14.3% |
| Sonnet 5 High | $0.49 | 13.7% |
| Sonnet 5 Medium | $0.36 | 13.3% |
| Sonnet 5 Low | $0.26 | 11.7% |
| Opus 4.8 Max | $1.78 | 15% |
| Opus 4.8 XHigh | $1.20 | 18.7% |
| Opus 4.8 High | $0.86 | 14% |
| Opus 4.8 Medium | $0.76 | 17% |
| Opus 4.8 Low | $0.60 | 12% |
| Opus 5 Max | $2.11 | 19.7% |
| Opus 5 XHigh | $1.89 | 17.3% |
| Opus 5 High | $1.61 | 18% |
| Opus 5 Medium | $1.10 | 17% |
| Opus 5 Low | $0.74 | 15.7% |
| Fable 5 Max | $3.49 | 18% |
| Fable 5 High | $1.88 | 16% |
| Fable 5 Medium | $1.50 | 14.7% |
| Fable 5 Low | $1.14 | 15% |
| Fable 5.1 Max | $3.36 | 27% |
| Fable 5.1 XHigh | $2.50 | 21.7% |
| Fable 5.1 High | $1.57 | 20.3% |
| Fable 5.1 Medium | $1.18 | 21% |
| Fable 5.1 Low | $0.92 | 16.3% |
| GLM 5.3 Max | $0.64 | 16.3% |
| GLM 5.3 High | $0.36 | 14.7% |
| GLM 5.3 Low | $0.20 | 8.7% |
| GLM 5.2 Max | $0.32 | 9% |
| GLM 5.2 High | $0.25 | 13.3% |
| DeepSeek V4.1 Flash High | $0.14 | 14% |
| DeepSeek V4.1 Flash Low | $0.10 | 13.7% |
| Kimi K2.7 Medium | $0.28 | 12.7% |
| Opus 5.5 Max | $3.57 | 70.5% |
| Opus 5.5 XHigh | $2.16 | 67% |
| Opus 5.5 High | $1.25 | 65% |
| Opus 5.5 Medium | $0.97 | 62.5% |
| Opus 5.5 Low | $0.61 | 54% |
| GPT-6 Luna Max | $0.06 | 52.3% |
| GPT-6 Luna XHigh | $0.04 | 46% |
| GPT-6 Luna High | $0.03 | 46% |
| GPT-6 Luna Medium | $0.02 | 42.7% |
| GPT-6 Luna Low | $0.01 | 30.3% |
| GPT-6 Sol Max | $0.98 | 57% |
| GPT-6 Sol XHigh | $0.58 | 61.3% |
| GPT-6 Sol High | $0.39 | 54.7% |
| GPT-6 Sol Medium | $0.27 | 51.3% |
| GPT-6 Sol Low | $0.17 | 50.3% |
Average cost per task
Each is evaluated against a rubric measuring analytical reasoning, tool use, evidence quality, and the final recommendation.
| 1Opus 5.5 · Max | 70.5% | $3.57 | 115.1K | 677.2s |
|---|---|---|---|---|
| 2Opus 5.5 · XHigh | 67.0% | $2.16 | 62.8K | 387.6s |
| 3Opus 5.5 · High | 65.0% | $1.25 | 30.5K | 220.1s |
| 4Opus 5.5 · Medium | 62.5% | $0.97 | 22.5K | 169.6s |
| 5GPT-6 Sol · XHigh | 61.3% | $0.58 | 14.6K | 326.7s |
| 6GPT-6 Sol · Max | 57.0% | $0.98 | 29.9K | 617.5s |
| 7GPT-6 Astra · Low | 55.0% | $0.85 | 2.7K | 88.5s |
| 8GPT-6 Astra · Max | 55.0% | $2.11 | 14.8K | 313.9s |
| 9GPT-6 Sol · High | 54.7% | $0.39 | 8.5K | 238s |
| 10Opus 5.5 · Low | 54.0% | $0.61 | 13K | 110.8s |
| 11GPT-6 Astra · Medium | 52.7% | $1.24 | 4.9K | 128.3s |
| 12GPT-6 Astra · XHigh | 52.3% | $1.75 | 9.3K | 224.3s |
| 13GPT-6 Luna · Max | 52.3% | $0.06 | 37.9K | 353.2s |
| 14GPT-6 Astra · High | 52.0% | $1.62 | 7.7K | 186.5s |
| 15GPT-6 Sol · Medium | 51.3% | $0.27 | 4.8K | 146.2s |
| 16GPT-6 Sol · Low | 50.3% | $0.17 | 2.8K | 97.6s |
| 17GPT-6 Luna · High | 46.0% | $0.03 | 15.5K | 175.4s |
| 18GPT-6 Luna · XHigh | 46.0% | $0.04 | 19.3K | 188.7s |
| 19GPT-6 Luna · Medium | 42.7% | $0.02 | 9.1K | 116.6s |
| 20GPT-5.6 Sol · XHigh | 40.3% | $0.98 | 11.4K | 210.2s |
| 21GPT-5.6 Sol · Medium | 39.7% | $0.61 | 5.6K | 134.3s |
| 22GPT-5.6 Terra · XHigh | 39.3% | $0.30 | 9.5K | 123.9s |
| 23GPT-5.6 Sol · High | 38.0% | $0.75 | 8K | 173.3s |
| 24GPT-5.6 Luna · XHigh | 35.3% | $0.04 | 11K | 145.3s |
| 25GPT-5.6 Terra · High | 32.7% | $0.20 | 5.4K | 95.2s |
| 26GPT-5.6 Luna · High | 31.3% | $0.03 | 7.9K | 118.1s |
| 27GPT-6 Luna · Low | 30.3% | $0.01 | 2.6K | 63.5s |
| 28GPT-5.6 Sol · Low | 29.7% | $0.37 | 2.8K | 76.2s |
| 29GPT-5.6 Terra · Medium | 29.3% | $0.15 | 3.2K | 69.4s |
| 30GPT-5.6 Terra · Low | 27.0% | $0.14 | 2.8K | 64.5s |
| 31Fable 5.1 · Max | 27.0% | $3.36 | 31.3K | 419s |
| 32GPT-5.6 Luna · Medium | 24.0% | $0.02 | 3.8K | 78.8s |
| 33Fable 5.1 · XHigh | 21.7% | $2.50 | 21.1K | 296.7s |
| 34Fable 5.1 · Medium | 21.0% | $1.18 | 8.1K | 126.3s |
| 35Fable 5.1 · High | 20.3% | $1.57 | 11.4K | 192.6s |
| 36Opus 5 · Max | 19.7% | $2.11 | 29.1K | 392.7s |
| 37GPT-5.6 Luna · Low | 19.0% | $0.01 | 2.3K | 61.3s |
| 38Opus 4.8 · XHigh | 18.7% | $1.20 | 20K | 251s |
| 39Opus 5 · High | 18.0% | $1.61 | 18.9K | 246.5s |
| 40Fable 5 · Max | 18.0% | $3.49 | 27.5K | 380s |
| 41Opus 5 · XHigh | 17.3% | $1.89 | 24.2K | 337.5s |
| 42Opus 4.8 · Medium | 17.0% | $0.76 | 10.1K | 133.4s |
| 43Opus 5 · Medium | 17.0% | $1.10 | 11.4K | 149.5s |
| 44Fable 5.1 · Low | 16.3% | $0.92 | 6.1K | 101.7s |
| 45GLM 5.3 · Max | 16.3% | $0.64 | 38.8K | 518.5s |
| 46Fable 5 · High | 16.0% | $1.88 | 10.2K | 156.7s |
| 47Opus 5 · Low | 15.7% | $0.74 | 6.9K | 98.2s |
| 48Opus 4.8 · Max | 15.0% | $1.78 | 34.7K | 433.9s |
| 49Fable 5 · Low | 15.0% | $1.14 | 4.7K | 79.8s |
| 50Fable 5 · Medium | 14.7% | $1.50 | 7.2K | 125s |
| 51GLM 5.3 · High | 14.7% | $0.36 | 14.9K | 215.1s |
| 52Sonnet 5 · XHigh | 14.3% | $0.62 | 20.2K | 213.6s |
| 53Opus 4.8 · High | 14.0% | $0.86 | 12.2K | 186.7s |
| 54DeepSeek V4.1 Flash · High | 14.0% | $0.14 | 31.7K | 168.4s |
| 55Sonnet 5 · High | 13.7% | $0.49 | 14.3K | 170.4s |
| 56DeepSeek V4.1 Flash · Low | 13.7% | $0.10 | 23.3K | 117.7s |
| 57Sonnet 5 · Medium | 13.3% | $0.36 | 9.1K | 114s |
| 58GLM 5.2 · High | 13.3% | $0.25 | 10.5K | 136.7s |
| 59Kimi K2.7 · Medium | 12.7% | $0.28 | 11.8K | 156.4s |
| 60Sonnet 5 · Max | 12.3% | $1.07 | 43.1K | 450.9s |
| 61Opus 4.8 · Low | 12.0% | $0.60 | 6.9K | 96.2s |
| 62Sonnet 5 · Low | 11.7% | $0.26 | 5.9K | 75.4s |
| 63GLM 5.2 · Max | 9.0% | $0.32 | 19.1K | 184s |
| 64GLM 5.3 · Low | 8.7% | $0.20 | 5.3K | 94.7s |
Updated judges and rubrics to fix accidental rewarding of misleading responses. This makes the benchmark significantly more difficult overall, with Anthropic models especially impacted.