LLM Leaderboard

Comprehensive benchmark scores for top Large Language Models. Compare performance across Coding, Reasoning, and Creative tasks.

ModelContextPlatform PriceOfficial Price
GPQA
AIME 2025
SWE-Bench
ARC-AGI v2
No models found matching your criteria

Metric Definitions

LLM

GPQA
Graduate-level science questions requiring expert knowledge.
AIME 2025
Recent math competition problems.
SWE-Bench
Real GitHub issues requiring code changes.
ARC-AGI v2
Abstract reasoning problems.