An Open-Source Data Contamination Report for Large Language Models

Yucheng Li, Yunhao Guo, Frank Guérin, Chenghua Lin

2024 · 인용 15

Data contamination in model evaluation has become increasingly prevalent with the growing popularity of large language models.It allows models to "cheat" via memorisation instead of displaying true capabilities.Therefore, contamination analysis has become a crucial part of reliable model evaluation to validate results.However, existing contamination analysis is usually conducted internally by large language model developers and often lacks transparency and completeness.This paper introduces an efficient and affordable method to identify potential data contamination in LLM benchmarks.We also present an extensive data contamination report for over 15 popular large language models across six widely used multiple-choice QA benchmarks.Our experiments reveal varying contamination levels ranging from 1% to 45% across benchmarks, with the contamination degree increasing rapidly over time.Performance analysis of large language models indicates that data contamination can have significant impact on model metrics: inflated accuracy of up to 14% and 7% are observed on contaminated C-Eval and HellaSwag benchmarks, and a small increase is identified on contaminated MMLU.We also find that data contamination has grown rapidly from 2020 to 2023 and that larger models benefit more from contaminated test sets.

📄 이 논문을 인용한 Paperis 글

이 논문이 근거 목록에 올라 있는 Paperis 글입니다.