
Ethan He discusses the future of video agent models and their reliance on LLMs in this episode of Latent Space.
We’re announcing AIEWF speakers this week! Take the AI Engineering Survey ! Today’s guest Ethan first joined us for the LS Paper Club as the lead on NVIDIA Cosmos World Model , but then joined xAI and built Grok Imagine in 3 months: He comes back on Latent Space with some nuclear hot takes: that Video Models primarily get their intelligence from LLMs , not from training on video data, and that the next frontier for truly interactive, realtime, long-horizon world models is to work on LLMs (perhaps Interaction Models as well…) Put it this way: In the near term, the next Sora won’t be a better video model, but a video agent . Generative Media may more closely follow the evolution of AI coding which went from focusing on one-shot output performance and cost, to multiturn reasoning and planning models for agents and systems that can plan, edit, test, debug, and submit PRs. At a certain point, coding models got so good that the only significant next step to improve performance was handling the orchestration of these models. Now as the performance of video models increases significantly across realism, consistency, & prompt adherence while becoming more cost efficient, the next…
Hosts: swyx, Vibhu
Guest: Ethan He
Organizations: NVIDIA, xAI
Products: Grok Imagine
Explore listener stats, chart rankings, contacts and more on the Latent Space: The AI Engineer Podcast podcast page.