Kimi K3 found GPT-5.6 Sol’s weak spot
|
After many failed attempts, Kimi K3 finally stayed online long enough to finish DocBench Arena, my document and presentation creation benchmark. It completed every task. More importantly, 400+ Redditors have now cast 3,000+ blind votes on the 236 generated files, so we finally have enough data to compare K3 with GPT-5.6 Sol without relying on vibes. Overall, it isn’t even close. Sol remains #1 with an Arena score of 1306, while K3 enters at #6 with 1192. Most of that gap comes from PowerPoint: – GPT-5.6 Sol: 1373 (#1) – Kimi K3: 1181 (#7) But Word documents tell a very different story: – Kimi K3: 1214 (#5) – GPT-5.6 Sol: 1211 (#6) That is effectively a tie, but it means K3 has already caught Sol on one half of the benchmark. It also costs less, averaging $0.44 per finished document versus $0.56 for Sol. At first glance, that makes K3 look like the better value. Then you look at how long it takes: Sol finishes a task in roughly 2 minutes while K3 needs 5x longer, also using more tokens and more agent steps along the way. The Pareto charts capture this pretty well. We can see Sol appears on four of the six quality-efficiency frontiers. K3 appears on none. In other words, Sol repeatedly occupies a position where no model is both better and more efficient. K3 always has another model somewhere above and to its left. So Sol still has the crown. It is #1 overall, destroys K3 at presentations and finishes more than 5x faster. But K3 found the opening: Word documents. It already matches Sol there while costing around 22% less. PS: K3 also has only 186 comparisons versus Sol’s 423, so this could still move in either direction. These results only exist because Redditors judged the files blindly, so thanks again to everyone who voted. If you want to help settle K3’s position, cast your vote here https://docbench.sprintos.co submitted by /u/ell-hol1 |