Published benchmark scores
Published scores. Tinted cells mark the best score on each test, including ties.
| Benchmark | Pareto 26.9 | Fable 5.1 | GPT 6 Astra | DeepSeek 4.1 Flash |
|---|---|---|---|---|
| DeepSWE | 74 | 67 | 74 | 74 |
| Terminal-Bench 4.0 | 51 | 56 | 58 | 31 |
| MMMU-Pro | 78 | 81 | 87 | 77 |
| HLE (no tools) | 49 | 55 | 54 | 39 |
| ArXivMath | 88 | 72 | 91 | 28 |
Measured task costs and a composite score have not been published for this release. Scores do not establish cost per completed task.
Details and pricing
Pareto accepts text and image inputs. Use pareto as the model identifier. The published comparison set is Fable 5.1, GPT 6 Astra, and DeepSeek 4.1 Flash.
Per million tokens: $2.50 input, $0.25 cached input, and $7.50 output. Cached input is billed at 0.1 times the input rate. Actual spend depends on token usage and cache hits.