Pierre-Carl Langlais on Building Models from Data You Can Account For

Pierre-Carl Langlais

Hosted by Ravid Shwartz Ziv, Allen Roush

Published Jul 23, 2026
1 h 6 min

Description

Pleias co-founder Pierre-Carl Langlais describes language models built from documented open and public-domain sources alongside synthetic data. The conversation examines the SYNTH dataset, gaps in common web crawls, preservation of source material and ethical choices beyond copyright status. It also covers compact deployed models, benchmark incentives, reasoning-trace access and competing approaches to national AI development. Hosted by Ravid Shwartz Ziv and Allen Roush. Watch the full conversation on the publisher’s YouTube channel.

Keywords

training data provenancepublic-domain corporasynthetic data

More from this series

We use essential cookies to run the site. Optional analytics and public-page session replay help us improve World Wide. Learn more.