
Hugo Bowne-Anderson and Bryan Bischof discuss the shortcomings of AI agents in data science evaluations and propose more realistic problem-solving approaches.
I often see what I would consider to be b******t evals , especially in data, like write this dumb SQL . Almost every one of these dumb SQL questions that I’ve seen for benchmarks are just so either obviously easy or overwhelmingly adversarial. They just, they don’t feel valuable as a data scientist , it’s something that you probably would never ask a real data scientist to do. So I went out my way to create real ones. Let me read one to you. Bryan Bischof , Head of AI at Theory Ventures , joins Hugo to talk about what happened when 150 people spent six hours using AI agents to answer real data science questions across SQL tables , log files , and 750,000 PDFs . They Discuss: * Failure Funnels , pinpoint where agent reasoning breaks down using causal-chain binary evaluations instead of vague 1-5 scales; * Median Score: 23 out of 65 , what happened when world-class engineers turned agents loose on real data work, and why general-purpose coding agents with human prodding beat fancy frameworks; * Zero-Cost Submissions Kill Trust , without a penalty for wrong answers, agents hill-climb to correct submissions through brute force instead of building confidence; * Data Science is…
Host: Hugo Bowne-Anderson
Guest: Bryan Bischof
Organizations: Theory Ventures
Explore listener stats, chart rankings, contacts and more on the Vanishing Gradients podcast page.