Try solving visual puzzles used to sort humans from bots and see how they’ve gotten harder as AI has gotten smarter.
MirrorCode benchmark's August 2026 leaderboard reveals Claude Fable 5 leads all frontier models at 64%, while GPT-5.5's ...