Log in

Cambridge users (raven) details

Other users details

No account? details

Information on

Subscribing to talks details

Finding a talk details

Adding a talk details

Disseminating talks details

Help and Documentation details

Natural Experiments in NLP and Where to Find Them

Add to your list(s) Download to your calendar using vCal

Pietro Lesci, University of Cambridge
Friday 15 November 2024, 15:30-17:00
MR12, Centre for Mathematical Sciences, Wilberforce Road, Cambridge.

If you have a question about this talk, please contact Martina Scauda.

Zoom Link available upon request

In training language models, training choices—such as the random seed for data ordering or the token vocabulary size—significantly influence model behaviour. Answering counterfactual questions like “How would the model perform if this instance were excluded from training?” is computationally expensive, as it requires re-training the model. Once these training configurations are set, they become fixed, creating a “natural experiment” where modifying the experimental conditions incurs high computational costs. Using econometric techniques to estimate causal effects from observational studies enables us to analyse the impact of these choices without requiring full experimental control or repeated model training. In this talk, I will present our paper, Causal Estimation of Memorisation Profiles (Best Paper Award at ACL 2024 ), which introduces a novel method based on the difference-in-differences technique from econometrics to estimate memorisation without requiring model re-training. I will also cover the necessary econometric concepts and key literature on memorisation in language models.

This talk is included in these lists:

Note that ex-directory lists are not shown.

Log in

Information on

Natural Experiments in NLP and Where to Find Them

This talk is included in these lists:

Other lists

Other talks