
Steven Bower
@bowerblu · Jul 13, 2026
Everyone still stuck on software engineering benchmarks.
Prediction: within a year, all models will be so good at software we will stop caring about measuring it, and start focusing on other surprising capabilities.
0xMarioNawfal@RoundtableSpace· Jul 12, 2026Gemini 3.5 Pro benchmark leak just dropped and the numbers are turning heads.
> Reportedly outperforming Claude Fable 5 and GPT-5.6 in internal evals
> Significant zero-shot performance improvements over 3.1 Pro
> Currently in private validation and testing
- Public rollout

Elon Musk
@elonmusk
True
01:43 PM · July 21, 2026 · 33.2K views
61
18
201