Pierre-Carl Langlais on Building Models from Data You Can Account For
Pierre-Carl Langlais
Hosted by Ravid Shwartz Ziv, Allen Roush
Published Jul 23, 2026
1 h 6 min
Description
Pleias co-founder Pierre-Carl Langlais describes language models built from documented open and public-domain sources alongside synthetic data. The conversation examines the SYNTH dataset, gaps in common web crawls, preservation of source material and ethical choices beyond copyright status. It also covers compact deployed models, benchmark incentives, reasoning-trace access and competing approaches to national AI development. Hosted by Ravid Shwartz Ziv and Allen Roush. Watch the full conversation on the publisher’s YouTube channel.
Keywords
More from this series
- Tiny Recursive Models Beat the Giants - Alexia Jolicoeur-Martineau (Microsoft)Sep 15, 2026 · 52 min
- Why Deep Learning Finally Works on Tables | Frank Hutter (Prior Labs)Aug 24, 2026 · 1 h 18 min
- Surya Ganguli: The Physics of IntelligenceAug 17, 2026 · 1 h 25 min
- Nathan Lambert: Inside Post-Training and the Open Model FightAug 8, 2026 · 1 h 15 min
- Sara Hooker on the End of Static AISep 16, 2026 · 1 h 36 min