Authors

Joint with Greg Leo

Abstract

Increasingly, researchers are using large language models (LLMs) to generate synthetic survey responses and behavioral data as a substitute for human samples. The justification for this approach is that appropriate prompt conditioning can recover latent patterns that LLMs learn about humans from training data. Existing evaluations of this technique typically compare LLM output to direct measurements of human behavior. However, these evaluations do not reflect the goal of this approach, which is often to make inferences about unobserved scenarios and populations. By contrast, we devise tests that can be run on LLM output alone and are based solely on structural properties implied by the assumptions underlying the technique. We show that LLMs produce data that is incoherent with respect to the statistical structure of real populations, which requires that the mean of any population lies within the range of the means of its constituent subpopulations. LLMs violate this condition at rates similar to purely random predictions. Further, we show that model outputs are subjective responses rather than objective predictions about a common (even biased) human benchmark population. In our tests the chosen model explains far more of the variation in output than can be explained by the choice of emulated demographic group. We conclude that LLM output is unreliable for inference about human populations.

Draft

LLMs Do Not Emulate Populations