Loop the Loopies!

Loop the Loopies!

July 20, 2026 路 20 min 路 Episode 2063

About this episode

The episode discusses the Loopie series of Mixture-of-Experts models and their performance in comparison to traditional Transformers.

馃 Upvotes: 56 | cs.CL, cs.AI Authors: Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai Title: Loop the Loopies! Arxiv: http://arxiv.org/abs/2607.16051v2 Abstract: We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. With a novel post-training method, Loopie develops strong reasoning abilities and achieves frontier-level reasoning performance.

More episodes of Daily Paper Cast

Explore listener stats, chart rankings, contacts and more on the Daily Paper Cast podcast page.