[Scale ML] Alok Puranik: Sequence Weighting at Scale
Jane Street
Hosted by Scale ML / MIT Computer Science and Artificial Intelligence Laboratory
Abstract
Alok Puranik (Jane Street) examines data mixing for large language models: allocating model capacity and computation among training sources. A small model’s preferred data mix may differ from that of a larger model, yet the cost of large-model training forces many decisions to depend on smaller experiments. Scaling laws can predict performance as size and computation increase only when trends remain stable or change predictably.
The talk uses experiments from Puranik’s work and frontier laboratories to examine data-mix questions that violate these assumptions, explain why small-scale results can mislead, and explore better methods for choosing training data. His research at Jane Street concerns scaling laws for sequence models in trading.