Skip to content
SeminarStarts todayMachine Learning

[Scale ML] Alok Puranik: Sequence Weighting at Scale

Jane Street

Hosted by Scale ML / MIT Computer Science and Artificial Intelligence Laboratory

· 60 minutes
Online

Abstract

Alok Puranik (Jane Street) examines data mixing for large language models: allocating model capacity and computation among training sources. A small model’s preferred data mix may differ from that of a larger model, yet the cost of large-model training forces many decisions to depend on smaller experiments. Scaling laws can predict performance as size and computation increase only when trends remain stable or change predictably.

The talk uses experiments from Puranik’s work and frontier laboratories to examine data-mix questions that violate these assumptions, explain why small-scale results can mislead, and explore better methods for choosing training data. His research at Jane Street concerns scaling laws for sequence models in trading.

Topics

data mixinglanguage-model scaling lawssequence models

We use essential cookies to run the site. Optional analytics and public-page session replay help us improve World Wide. Learn more.