The Information Bottleneck
Pleias co-founder Pierre-Carl Langlais describes language models built from documented open and public-domain sources alongside synthetic data. The conversation examines the SYNTH dataset, gaps in common web crawls, preservation of source material and ethical choices beyond copyright status. It also covers compact deployed models, benchmark incentives, reasoning-trace access and competing approaches to national AI development. Hosted by Ravid Shwartz Ziv and Allen Roush. Watch the full conversation on the publisher’s YouTube channel.