
The episode discusses how to evaluate generative AI through various metrics and challenges.
We unpack how to evaluate AI that writes and creates, not just predicts. Why perplexity captures surprise, why a low perplexity score isn’t a guarantee of correctness, and how precision, recall, and the harmonic F1 balance model performance. We compare BLEU and ROUGE, explore Retrieval-Augmented Generation to stay faithful to private data, and discuss out-of-domain challenges, agentic AI, and the guardrails shaping the future. Note: This podcast was AI-generated, and sometimes AI can m...
Organizations: AI, BLEU
Explore listener stats, chart rankings, contacts and more on the Intellectually Curious podcast page.