Log in

Cambridge users (raven) details

Other users details

No account? details

Information on

Subscribing to talks details

Finding a talk details

Adding a talk details

Disseminating talks details

Help and Documentation details

Posterior sampling via autoregressive generation

Add to your list(s) Download to your calendar using vCal

Kelly Zhang (Imperial College London)
Friday 29 November 2024, 14:00-15:00
Centre for Mathematical Sciences MR12, CMS.

If you have a question about this talk, please contact Qingyuan Zhao.

Uncertainty quantification remains a critical challenge when using deep learning models, particularly in complex decision-making settings. We propose a new framework for learning bandit algorithms from massive historical data, by combining classical ideas from multiple imputation with autoregressive generative sequence modeling. We demonstrate our approach in a cold-start recommendation problem where, first, we use historical data to pretrain an autoregressive model to predict sequences of repeated feedback/rewards (e.g., responses to news articles shown to different users over time). In learning to make accurate predictions, the model implicitly learns an informed prior based on rich action features (e.g., article headlines) and how to sharpen beliefs as more rewards are gathered (e.g., clicks as each article is recommended). At decision-time, the algorithm autoregressively samples (imputes) a hypothetical sequence of rewards for each action and chooses the action with the largest average imputed reward. Far from a heuristic, our approach is an implementation of Thompson sampling (with a learned prior), a prominent active exploration algorithm. We prove our pretraining sequence loss directly controls online decision-making performance, and we demonstrate our framework on a news recommendation task where we integrate end-to-end fine-tuning of a pretrained language model to process news article headline text to improve performance.

This talk is part of the Statistics series.

This talk is included in these lists:

Note that ex-directory lists are not shown.

Log in

Information on

Posterior sampling via autoregressive generation

This talk is included in these lists:

Other lists

Other talks