Model Evaluation Leaderboard
Compare frontier models across benchmarks, metrics and evaluation conditions. Click a model for its full performance detail.
Metric Summary (Avg)
Leaderboard
Model Comparison
Side-by-side benchmark scores for the models you selected. Click a benchmark to explore its underlying data.
Compared Models
Evaluation Conditions
All scores are averaged over 3 independent runs on dataset v1.3 under Default conditions: temperature 0, identical prompts, system messages and stop sequences for every model.
BLEU is computed on detokenized output; win rate comes from pairwise head-to-head judgments on a shared prompt set.