DeLTA Seminar by Nathan Kallus

LLM Post-Training and Reasoning via Efficient Value-Based RL

ABSTRACT

Reinforcement learning (RL) has a newfound killer application in post-training LLMs pre-trained to predict next token to adapt to tasks like instruction following, math-problem solving, and generating content or recommendations that maximize user outcomes. But are the same RL algorithms that animated robots and conquered Atari the right ones to post-train LLMs?

In this talk I will present new value-based algorithms for post-training and for scaling test-time compute that leverage both the unique structure of autoregressive LLMs and recent advances on increasing efficiency by changing the Q-learning loss function. I will show how (and argue why) these new algorithms achieve state-of-the-art performance on frontier math reasoning tasks with smaller models and at a fraction of test-time FLOPs.

BIOGRAPHY

Nathan Kallus is an Associate Professor at the Cornell Tech campus of Cornell University in NYC and Director of Machine Learning and Inference Research at Netflix. Nathan's research interests include causal machine learning, sequential and dynamic decision making, optimization under uncertainty, and algorithmic fairness.

Follow this link to participate on Zoom

Join the DeLTA community

You can subscribe to the DeLTA Seminar mailing list by sending an empty email to delta-seminar-join@list.ku.dk.
Online calendar
DeLTA Lab page