Measuring Self-Rating Bias in LLM-Generated Survey Data: A Semantic Similarity Framework for Independent Scale Mapping

Eduardo Vera Pichardo

SSRN Electronic Journal · 2026

Synthetic survey data generated by large language models (LLMs) offers a promising com plement to traditional survey research, but existing approaches suffer from a fundamental circularity: the same model family that generates text responses also maps them to numeri cal scales. We introduce Semantic Similarity Rating (SSR), a framework that decouples text generation from scale mapping by using embedding-based cosine similarity against predefined anchor statements. A pilot configuration study (N = 17, 3 domains) revealed that naturalis tic behavioral anchors outperform formal survey jargon by 29 percentage points, asymmetric embedding adds a further +6 pp, and anchor quality is the single most impactful design factor.

Cross-validation on an expanded 69-case test set across 8 domains yields 65% exact match and 91% within ±1. A direct LLM baseline (Claude Haiku 4.5, temperature = 0) achieves 87% exact match on the same benchmark, establishing that SSR’s contribution is methodological independence rather than accuracy superiority. A pre-registered circularity experiment (N = 345 generated texts, 5 personas × 69 cases) demonstrates that this in dependence matters empirically: LLM self-rating produces 4-fold compressed error variance (σ2 = 0.21 vs.

0.87 for SSR) and systematic directional bias, confirming that self-rating cir cularity is not merely a theoretical concern. We release the full calibration dataset, anchor library (15 semantic families), and framework as open-source tools.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.