LLM Leaderboard
Comprehensive benchmark scores for top Large Language Models. Compare performance across Coding, Reasoning, and Creative tasks.
| Model | Context | Platform Price | Official Price | GPQA | AIME 2025 | SWE-Bench | ARC-AGI v2 | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| No models found matching your criteria | |||||||||||
Metric Definitions
LLM
- GPQA
- Graduate-level science questions requiring expert knowledge.
- AIME 2025
- Recent math competition problems.
- SWE-Bench
- Real GitHub issues requiring code changes.
- ARC-AGI v2
- Abstract reasoning problems.