Infrastructure for AI at Scale - With Benny Chen (Fireworks AI)

Benny Chen

Hosted by Ravid Shwartz Ziv, Allen Roush

Published Jun 24, 2026
1 h 6 min

Description

Fireworks AI co-founder Benny Chen explains the infrastructure required to serve machine-learning models at scale, from GPU procurement and kernels to runtimes and request routing. He discusses power constraints, hardware utilization, mixture-of-experts models, quantization, reinforcement-learning efficiency and speculative decoding. The conversation also covers distributed training across data centres, sparse autoencoders, the constraints on open-source interpretability research, and the relationship between open models and compute availability. Hosted by Ravid Shwartz Ziv and Allen Roush; the full conversation is available on YouTube.

Keywords

AI inferenceGPU computingmodel serving

More from this series

We use essential cookies to run the site. Optional analytics and public-page session replay help us improve World Wide. Learn more.