
James Governor talks with Paul Brookes about optimizing AI inference and the challenges in AI performance engineering.
In this MonkCast conversation, RedMonk's James Governor talks with Paul Brookes, a senior AI engineer at TurinTech, about making AI inference faster and cheaper. TurinTech predates the generative AI boom, having spent years optimizing complex code with genetic algorithms, and now points those tools at the models themselves. Brookes walks through techniques like kernel fusion and model compilation that squeeze more tokens per second out of specific hardware, drawing on the company's work with Intel on OpenVINO and vLLM. The two get into running capable open models such as Qwen on local machines, the spiraling cost of AI , and why judging engineers by tokens burned misses the point. Brookes also describes Artemis and its discovery harness, which lets agents learn from past results, and traces his own route from quantum physics into low-level performance engineering. This RedMonk conversation is sponsored by TurinTech. Show notes: https://redmonk.com/videos/paul-brookes/ Chapters 00:00 Introduction to AI and TurinTech 01:42 Optimization Challenges in AI 04:30 Working with Semiconductor Companies 08:46 The Cost of AI and Local Model Deployment 11:29 Local Model Performance and…
Host: James Governor
Guest: Paul Brookes
TurinTech
Organizations: Intel
Products: OpenVINO, vLLM, Qwen, Artemis
Places: quantum physics
Explore listener stats, chart rankings, contacts and more on the The MonkCast podcast page.