Solving the Memory Wall: A Deep Dive into AI Inference with Sandra Rivera

Solving the Memory Wall: A Deep Dive into AI Inference with Sandra Rivera

April 10, 2026 · 17 min · Episode 497

About this episode

The episode features a discussion with Sandra Rivera about AI inference and how VSORA's technology addresses the memory wall.

This week, I'm excited to welcome Sandra Rivera from VSORA! We dive into a discussion on why AI inference is essential for deployment at scale, specifically focusing on how VSORA’s patented software architecture addresses the "memory wall" by collapsing memory layers. We explore their recent tape-out, which promises approximately 3X the performance at half the power of leading GPUs. We also chat about deployment use cases, the need for low latency and high determinism, future plans for OEM modules and MLPerf benchmarking, and even get a brief look into Sandra’s family llama farm.

People in this episode

Host: Amelia

Guest: Sandra Rivera

Topics covered

Mentioned in this episode

Organizations: VSORA, EEJournal.com

More episodes of Amelia's Weekly Fish Fry

Explore listener stats, chart rankings, contacts and more on the Amelia's Weekly Fish Fry podcast page.