What comes after attention? This startup says it already knows.

What comes after attention? This startup says it already knows.

July 7, 2026 · 20 min · Episode 1632

About this episode

The episode discusses Subquadratic's innovative Sparse Attention architecture and its implications for long-context performance in AI models.

Subquadratic is beginning to back up its ambitious claims with benchmarks and third-party validation for its SubQ 1.1 Small model, which uses its proprietary Sparse Attention (SSA) architecture to dramatically improve long-context performance. Rather than comparing every token to every other token, SSA selectively processes relationships, enabling near-linear scaling while maintaining high accuracy across context windows of up to 12 million tokens. The company reports near-perfect retrieval performance, competitive coding and reasoning benchmarks, and compute savings of up to 1,000x at maximum context lengths.

Topics covered

Keywords

Mentioned in this episode

Organizations: Subquadratic

Products: SubQ 1.1 Small, Sparse Attention (SSA)

More episodes of The New Stack Podcast

Explore listener stats, chart rankings, contacts and more on the The New Stack Podcast podcast page.