
Avi Chawla
@_avichawla · Mar 16, 2026
Big release from Kimi!
They just released a new way to handle residual connections in Transformers.
In a standard Transformer, every sub-layer (attention or MLP) computes an output and adds it back to the input via a residual connection.
If you consider this across 40+ layers,

Kimi.ai@Kimi_Moonshot· Mar 16, 2026Introducing 𝑨𝒕𝒕𝒆𝒏𝒕𝒊𝒐𝒏 𝑹𝒆𝒔𝒊𝒅𝒖𝒂𝒍𝒔: Rethinking depth-wise aggregation.
Residual connections have long relied on fixed, uniform accumulation. Inspired by the duality of time and depth, we introduce Attention Residuals, replacing standard depth-wise recurrence with

Elon Musk
@elonmusk
Impressive work from Kimi
12:58 PM · March 16, 2026 · 490.3K views
293
171
3.6K