Tag: benchmarks
-

Kimi K3 Just Beat Claude at Code
Moonshot AI’s Kimi K3 hit number one on Arena.ai’s Frontend Code Arena this week with a 76 percent pairwise win rate against Claude Fable 5. That is a startling jump from K2.6, which ranked 18th in the same benchmark just months ago. This is not marginal improvement. This is a model that learned something specific…
-

Claude Just Hit a New Ceiling on What’s Possible
Claude Fable 5 is now the top-ranked model on the major benchmarks. A hundred out of a hundred on composite quality. That’s not second place. That’s the leaderboard. More important than the ranking is what the ranking measures. It’s not a trick benchmark designed to favor one architecture. It’s the aggregation of real, hard problems:…
-

Google Just Dropped the Model That Beats Everyone on Science
Google released Gemini 2.5 Pro with Deep Think on June 22, and the benchmarks are not close anymore. They are not even in the same building. Here are the numbers: 89.8 percent on MMLU-Pro. 82.4 percent on GPQA Diamond. GPQA is the science one: the test where you need actual knowledge, not pattern-matching. Gemini is…
-

A 3B Model Just Outran the Megamodels at Reasoning
Weibo AI shipped VibeThinker-3B last month and it’s blowing the reasoning-to-parameters ratio to pieces. A three-billion-parameter model that scores 94.3 on AIME26, an 80.2 on LiveCodeBench v6, and passes recent unseen LeetCode contests at 96.1%. Here’s the punch: these are the benchmarks that GPT-5.5, Opus 4.8, o1, and the full-size reasoning models fight over. VibeThinker-3B…
