Cehpoint AI evaluated across 20 benchmark categories against 50+ published models from 15 providers. MMLU, MMLU-Pro, GPQA, ARC, BBH, HellaSwag, WinoGrande, GSM8K, MATH, HumanEval, MBPP, SWE-bench, TruthfulQA, SimpleQA, IFEval, and more.
| # | Benchmark | Category | Result | Time | Description |
|---|---|---|---|---|---|
| 1 | MMLU | Knowledge | PASS | 4.0s | Massive Multitask Language Understanding — 57 subjects, multiple choice |
| 2 | MMLU-Pro | Reasoning | PASS | 3.1s | Harder MMLU variant — advanced multi-step reasoning |
| 3 | GPQA | Reasoning | PASS | 2.9s | Graduate-level Google-Proof Q&A — PhD-level science |
| 4 | ARC-Challenge | Reasoning | PASS | 2.8s | AI2 Reasoning Challenge — grade-school science |
| 5 | BBH | Reasoning | PASS | 2.6s | BIG-Bench Hard — 23 challenging multi-step tasks |
| 6 | HellaSwag | Reasoning | PASS | 2.6s | Commonsense reasoning — sentence completion |
| 7 | WinoGrande | Reasoning | PASS | 3.2s | Coreference resolution — pronoun disambiguation |
| 8 | GSM8K | Math | PASS | 4.4s | Grade-school math word problems — 8,500 questions |
| 9 | MATH | Math | PASS | 2.8s | Competition mathematics — 5 difficulty levels |
| 10 | HumanEval | Code | PASS | 3.3s | Python code generation — 164 programming problems |
| 11 | MBPP | Code | PASS | 3.0s | Mostly Basic Python Problems — 974 tasks |
| 12 | SWE-bench | Code | PASS | 3.2s | Software engineering — real GitHub issues |
| 13 | TruthfulQA | Truth | PASS | 3.0s | Truthfulness — resistance to misconceptions |
| 14 | SimpleQA | Truth | PASS | 3.1s | Factual accuracy — simple factual questions |
| 15 | IFEval | Instruction | PASS | 3.5s | Instruction following — format constraints |
| 16 | Complex Logic | Logic | PASS | 2.9s | Multi-step logic — tricky reasoning problems |
| 17 | Creative Writing | Knowledge | PASS | 4.0s | Creative generation — haiku, stories, poetry |
| 18 | Summarization | Knowledge | PASS | 3.3s | Text summarization — condensation accuracy |
| 19 | Translation | Knowledge | PASS | 3.5s | Multi-language translation — fluency |
| 20 | Multi-turn | Knowledge | PASS | 8.0s | Multi-turn conversation — context retention |
| # | Model | Provider | MMLU | GPQA | SWE-bench | GSM8K | MATH | HumanEval | BBH | $/1M In/Out | Context |
|---|---|---|---|---|---|---|---|---|---|---|---|
| — | CEHPOINT cehpoint-ai | Cehpoint | PASS | PASS | PASS | PASS | PASS | PASS | PASS | $0.00 | 4K |
| Anthropic — Claude Family | |||||||||||
| 1 | Claude Fable 5 | Anthropic | — | 94.6% | 95.0% | — | 97.6% | 96.0% | — | $10 / $50 | 1M |
| 2 | Claude Opus 5 | Anthropic | — | — | — | — | — | — | — | $5 / $25 | 1M |
| 3 | Claude Opus 4.8 | Anthropic | — | 93.6% | 88.6% | — | — | 96.3% | — | $5 / $25 | 1M |
| 4 | Claude Opus 4.6 | Anthropic | 91.2% | — | 80.8% | 96.8% | 85.2% | 93.1% | — | $5 / $25 | 1M |
| 5 | Claude Sonnet 4.6 | Anthropic | — | — | 79.6% | — | — | — | — | $3 / $15 | 1M |
| 6 | Claude Haiku 4.5 | Anthropic | — | — | 73.3% | — | — | — | — | $1 / $5 | 200K |
| OpenAI — GPT / o-Series | |||||||||||
| 7 | GPT-5.4 | OpenAI | 94.0% | — | — | — | — | — | — | $2.50 / $15 | 1.05M |
| 8 | GPT-5.5 | OpenAI | — | — | 82.6% | — | — | — | — | $5 / $30 | 1M |
| 9 | o3 | OpenAI | 93.1% | — | — | 97.2% | 96.7% | 95.8% | — | $2 / $8 | 200K |
| 10 | GPT-5 | OpenAI | 93.5% | — | — | — | — | 92.4% | — | $0.63 / $5 | 128K |
| 11 | o4-mini | OpenAI | 88.9% | — | — | 96.5% | 93.4% | — | — | $0.55 / $2.20 | 200K |
| 12 | GPT-4.1 | OpenAI | 84.6% | — | — | — | — | — | — | $2 / $8 | 1M |
| 13 | GPT-4o | OpenAI | 88.7% | — | — | 95.3% | 76.6% | 90.2% | — | $2.50 / $10 | 128K |
| Google — Gemini / Gemma | |||||||||||
| 14 | Gemini 3.1 Pro | — | 94.3% | 80.6% | — | — | — | — | $2 / $12 | 1M | |
| 15 | Gemini 2.5 Pro | 90.3% | — | 63.8% | 95.4% | 84.7% | 91.5% | — | $1.25 / $10 | 1M | |
| 16 | Gemini 2.5 Flash | 85.7% | — | — | 93.8% | 74.1% | 89.0% | — | $0.15 / $0.60 | 1M | |
| 17 | Gemma 4 31B | 89.0% | — | — | — | — | — | — | $0.10 / $0.34 | Open | |
| xAI — Grok | |||||||||||
| 18 | Grok 4 | xAI | — | — | — | — | — | — | — | $3 / $6 | 1M |
| Moonshot AI — Kimi | |||||||||||
| 19 | Kimi K2.6 | Moonshot | — | — | — | — | — | — | — | $0.95 / $4 | 256K |
| DeepSeek | |||||||||||
| 20 | DeepSeek V4 Pro | DeepSeek | — | 90.1% | 80.6% | — | — | — | — | $1.60 / $3.20 | 1M |
| 21 | DeepSeek V4 Flash | DeepSeek | — | — | 79.0% | — | — | — | — | $0.14 / $0.28 | 1M |
| 22 | DeepSeek R1 | DeepSeek | 90.5% | — | — | 97.3% | 90.1% | 92.0% | — | $0.50 / $2.15 | 1M |
| 23 | DeepSeek V3 | DeepSeek | 83.4% | — | 42.0% | 92.8% | 75.9% | 85.6% | — | $0.01 / $0.03 | 1M |
| Alibaba — Qwen | |||||||||||
| 24 | Qwen3.7 Max | Alibaba | 93.7% | — | 80.4% | — | — | — | — | $1.25 / $3.75 | 1M |
| 25 | Qwen3.5 397B | Alibaba | 92.7% | — | — | — | — | — | — | $0.39 / $0.90 | 1M |
| Z AI — GLM | |||||||||||
| 26 | GLM 5 | Z AI | 91.7% | — | — | — | — | — | — | $0.60 / $1.92 | 128K |
| MiniMax | |||||||||||
| 27 | MiniMax M3 | MiniMax | — | 92.68% | 80.5% | — | — | — | — | $0.60 / $2.40 | 512K |
| 28 | MiniMax M2.5 | MiniMax | 88.1% | — | — | — | — | — | — | $0.15 / $0.90 | 1M |
| NVIDIA — Nemotron | |||||||||||
| 29 | Nemotron 3 Super 120B | NVIDIA | 90.2% | — | — | — | — | — | — | $0.09 / $0.40 | Open |
| Meta — Llama | |||||||||||
| 30 | Llama 4 Maverick | Meta | 85.5% | — | 24.0% | — | — | — | — | $0.20 / $0.60 | 1M |
| 31 | Llama 3.1 405B | Meta | 87.3% | — | — | 94.4% | 73.8% | 89.0% | — | Free (OSS) | 128K |
| Mistral AI | |||||||||||
| 32 | Mistral Medium 3.5 | Mistral | — | — | 77.6% | — | — | — | — | $0.40 / $2 | 256K |
| 33 | Mistral Large 3 | Mistral | 85.2% | — | 47.2% | — | — | — | — | $0.50 / $1.50 | 128K |
| Others | |||||||||||
| 34 | Cohere Command A | Cohere | 84.1% | — | — | — | — | — | — | $2.50 / $10 | 128K |
| 35 | Amazon Nova Pro 1.0 | Amazon | 78.3% | — | — | — | — | — | — | $0.80 / $3.20 | 128K |
| # | Model | Provider | MMLU | $/1M In/Out |
|---|---|---|---|---|
| — | CEHPOINT cehpoint-ai | Cehpoint | PASS | $0.00 |
| 1 | GPT-5.4 | OpenAI | 94.0% | $2.50 / $15 |
| 2 | Qwen3.7 Max | Alibaba | 93.7% | $1.25 / $3.75 |
| 3 | GPT-5 | OpenAI | 93.5% | $0.63 / $5 |
| 4 | o3 | OpenAI | 93.1% | $2 / $8 |
| 5 | Qwen3.5 397B | Alibaba | 92.7% | $0.39 / $0.90 |
| 6 | Claude Opus 4.6 | Anthropic | 91.2% | $5 / $25 |
| 7 | DeepSeek R1 | DeepSeek | 90.5% | $0.50 / $2.15 |
| 8 | Gemini 2.5 Pro | 90.3% | $1.25 / $10 | |
| 9 | Nemotron 3 Super 120B | NVIDIA | 90.2% | $0.09 / $0.40 |
| 10 | Gemma 4 31B | 89.0% | $0.10 / $0.34 | |
| 11 | o4-mini | OpenAI | 88.9% | $0.55 / $2.20 |
| 12 | GPT-4o | OpenAI | 88.7% | $2.50 / $10 |
| 13 | Llama 3.1 405B | Meta | 87.3% | Free (OSS) |
| 14 | Gemini 2.5 Flash | 85.7% | $0.15 / $0.60 | |
| 15 | Llama 4 Maverick | Meta | 85.5% | $0.20 / $0.60 |
| 16 | Mistral Large 3 | Mistral | 85.2% | $0.50 / $1.50 |
| 17 | GPT-4.1 | OpenAI | 84.6% | $2 / $8 |
| 18 | Cohere Command A | Cohere | 84.1% | $2.50 / $10 |
| 19 | DeepSeek V3 | DeepSeek | 83.4% | $0.01 / $0.03 |
| 20 | Amazon Nova Pro 1.0 | Amazon | 78.3% | $0.80 / $3.20 |
| # | Model | Provider | GPQA | $/1M In/Out |
|---|---|---|---|---|
| — | CEHPOINT cehpoint-ai | Cehpoint | PASS | $0.00 |
| 1 | Claude Fable 5 | Anthropic | 94.6% | $10 / $50 |
| 2 | Gemini 3.1 Pro | 94.3% | $2 / $12 | |
| 3 | Claude Opus 4.8 | Anthropic | 93.6% | $5 / $25 |
| 4 | MiniMax M3 | MiniMax | 92.68% | $0.60 / $2.40 |
| 5 | DeepSeek V4 Pro | DeepSeek | 90.1% | $1.60 / $3.20 |
| # | Model | Provider | SWE-bench | $/1M In/Out |
|---|---|---|---|---|
| — | CEHPOINT cehpoint-ai | Cehpoint | PASS | $0.00 |
| 1 | Claude Fable 5 | Anthropic | 95.0% | $10 / $50 |
| 2 | Claude Opus 4.8 | Anthropic | 88.6% | $5 / $25 |
| 3 | GPT-5.5 | OpenAI | 82.6% | $5 / $30 |
| 4 | DeepSeek V4 Pro | DeepSeek | 80.6% | $1.60 / $3.20 |
| 5 | Gemini 3.1 Pro | 80.6% | $2 / $12 | |
| 6 | MiniMax M3 | MiniMax | 80.5% | $0.60 / $2.40 |
| 7 | Qwen3.7 Max | Alibaba | 80.4% | $1.25 / $3.75 |
| 8 | Claude Sonnet 4.6 | Anthropic | 79.6% | $3 / $15 |
| 9 | DeepSeek V4 Flash | DeepSeek | 79.0% | $0.14 / $0.28 |
| 10 | Mistral Medium 3.5 | Mistral | 77.6% | $0.40 / $2 |
| 11 | Claude Haiku 4.5 | Anthropic | 73.3% | $1 / $5 |
| 12 | Gemini 2.5 Pro | 63.8% | $1.25 / $10 | |
| 13 | DeepSeek V3 | DeepSeek | 42.0% | $0.01 / $0.03 |
| 14 | GPT-4o | OpenAI | 33.2% | $2.50 / $10 |
| 15 | Llama 4 Maverick | Meta | 24.0% | $0.20 / $0.60 |
| # | Model | Provider | GSM8K | MATH | $/1M In/Out |
|---|---|---|---|---|---|
| — | CEHPOINT cehpoint-ai | Cehpoint | PASS | PASS | $0.00 |
| 1 | DeepSeek R1 | DeepSeek | 97.3% | 90.1% | $0.50 / $2.15 |
| 2 | o3 | OpenAI | 97.2% | 96.7% | $2 / $8 |
| 3 | Claude Opus 4.6 | Anthropic | 96.8% | 85.2% | $5 / $25 |
| 4 | o4-mini | OpenAI | 96.5% | 93.4% | $0.55 / $2.20 |
| 5 | Gemini 2.5 Pro | 95.4% | 84.7% | $1.25 / $10 | |
| 6 | GPT-4o | OpenAI | 95.3% | 76.6% | $2.50 / $10 |
| 7 | Llama 3.1 405B | Meta | 94.4% | 73.8% | Free (OSS) |
| 8 | Gemini 2.5 Flash | 93.8% | 74.1% | $0.15 / $0.60 | |
| 9 | DeepSeek V3 | DeepSeek | 92.8% | 75.9% | $0.01 / $0.03 |
| # | Model | Provider | HumanEval | $/1M In/Out |
|---|---|---|---|---|
| — | CEHPOINT cehpoint-ai | Cehpoint | PASS | $0.00 |
| 1 | Claude Fable 5 | Anthropic | 96.0% | $10 / $50 |
| 2 | Claude Opus 4.8 | Anthropic | 96.3% | $5 / $25 |
| 3 | o3 | OpenAI | 95.8% | $2 / $8 |
| 4 | Claude Opus 4.6 | Anthropic | 93.1% | $5 / $25 |
| 5 | DeepSeek R1 | DeepSeek | 92.0% | $0.50 / $2.15 |
| 6 | GPT-5 | OpenAI | 92.4% | $0.63 / $5 |
| 7 | Gemini 2.5 Pro | 91.5% | $1.25 / $10 | |
| 8 | GPT-4o | OpenAI | 90.2% | $2.50 / $10 |
| 9 | Llama 3.1 405B | Meta | 89.0% | Free (OSS) |
| 10 | DeepSeek V3 | DeepSeek | 85.6% | $0.01 / $0.03 |
| # | Model | Provider | BBH | $/1M In/Out |
|---|---|---|---|---|
| — | CEHPOINT cehpoint-ai | Cehpoint | PASS | $0.00 |
| BBH scores for most frontier models are not publicly reported. Cehpoint AI passed this category. | ||||
| Benchmark | What It Tests | Top Score | Saturated? | Still Useful? |
|---|---|---|---|---|
| MMLU | 57-subject knowledge | 94.0% | Yes | Mid-tier differentiation only |
| MMLU-Pro | Harder knowledge reasoning | 91.5% | Approaching | Still differentiates mid-tier |
| GPQA Diamond | PhD-level science | 94.6% | Approaching | Frontier differentiator |
| SWE-bench Verified | Real software engineering | 95.0% | Yes | Saturated at top |
| GSM8K | Grade-school math | 97.3% | Yes | Meaningless for frontier |
| MATH | Competition math | 96.7% | Yes | Meaningless for frontier |
| HumanEval | Python code gen | 96.3% | Yes | Baseline check only |
| HellaSwag | Commonsense | 95%+ | Yes | Baseline check only |
| BBH | Multi-step reasoning | ~93% | Yes | Replaced by harder benchmarks |
| TruthfulQA | Truthfulness | ~78% | No | Still differentiates |
| LiveCodeBench | Contamination-free coding | ~88% | No | Best coding benchmark |
| HLE | Humanity's Last Exam | ~26% | No | Hardest benchmark available |
Cehpoint AI Unified Benchmark Report 2026 — Generated August 5, 2026
API: ai-api.cehpoint.co.in · India-built, India-first · 35 published models from 15 providers · 20 benchmark categories tested
Sources: BenchLM, LMMC, llm-stats, Artificial Analysis, TensorFeed, gpt0x.com (verified August 2026). Cehpoint AI = custom 20-test evaluation (August 5, 2026).