Synthetic survey participants cannot substitute for sample diversity in policy
Sakshi Ghai, J. David Cummins
2026
Synthetic survey participants generated by large language models (LLMs) have recently been promoted as a means of reaching diverse and underrepresented populations in policy research. We argue that this promise rests on a mistaken equivalence between demographic labelling and meaningful representation. Because LLM training corpora disproportionately reflect Western, English-language, and third-person accounts, synthetic responses may reproduce how populations are described by others rather than how their members understand and express their own experiences.
We term this process etic-to-emic laundering: the transformation of external accounts into outputs that appear to be first-person testimony. Beyond methodological limitations, synthetic substitution may weaken incentives to invest in the infrastructure required for genuine inclusion, including translation, culturally grounded measurement, local expertise, community partnerships, and direct data collection. This displacement risks becoming self-reinforcing, as reduced collection of first-person data makes future synthetic samples still harder to validate.
We therefore argue that synthetic participants should be confined to exploratory roles, accompanied by clear provenance disclosures, and never treated as substitutes for human respondents in policy inference. Sample diversity cannot be automated: it requires sustained engagement with the people whose lives and interests policy decisions affect.