The Model Found a Way Out - with Florian Brand (Prime Intellect)
Florian Brand of Prime Intellect discusses how to evaluate agents when their tools and execution environments affect the result. The episode examines benchmark gaming, preventing shortcuts around scoring rules and the statistical limits of expensive evaluation runs. It also asks whether subjective impressions of reliability can be translated into useful measurements of model behavior. Hosted by Ravid Shwartz Ziv and Allen Roush. Watch the full conversation on the publisher’s YouTube channel.
Kaggle Grandmasters, Agent Skills, and Why Everyone Is Overfitting with Jean-Francois Puget (Nvidia)
NVIDIA researcher Jean-Francois Puget discusses evaluating agent skills and distinguishing real improvements from benchmark overfitting. Drawing on Kaggle competitions, he emphasizes validation that separates development feedback from final evaluation. The episode also considers coding agents, multi-agent software development, model-release claims and his team’s approach to the ARC-AGI competition. Hosted by Ravid Shwartz Ziv and Allen Roush. Watch the full conversation on the publisher’s YouTube channel.