Infrastructure for AI at Scale - With Benny Chen (Fireworks AI)
Benny Chen
Hosted by Ravid Shwartz Ziv, Allen Roush
Published Jun 24, 2026
1 h 6 min
Description
Fireworks AI co-founder Benny Chen explains the infrastructure required to serve machine-learning models at scale, from GPU procurement and kernels to runtimes and request routing. He discusses power constraints, hardware utilization, mixture-of-experts models, quantization, reinforcement-learning efficiency and speculative decoding. The conversation also covers distributed training across data centres, sparse autoencoders, the constraints on open-source interpretability research, and the relationship between open models and compute availability. Hosted by Ravid Shwartz Ziv and Allen Roush; the full conversation is available on YouTube.
Keywords
More from this series
- Tiny Recursive Models Beat the Giants - Alexia Jolicoeur-Martineau (Microsoft)Sep 15, 2026 · 52 min
- Why Deep Learning Finally Works on Tables | Frank Hutter (Prior Labs)Aug 24, 2026 · 1 h 18 min
- Surya Ganguli: The Physics of IntelligenceAug 17, 2026 · 1 h 25 min
- Nathan Lambert: Inside Post-Training and the Open Model FightAug 8, 2026 · 1 h 15 min
- Sara Hooker on the End of Static AISep 16, 2026 · 1 h 36 min