AI in EU Law: Training, Hallucinations, Memorization, and the Core Regulatory Architecture
Philipp Hacker
SSRN Electronic Journal · 2026
This chapter examines the regulatory architecture for artificial intelligence in European Union law across the three stages of the machine learning pipeline: training, model, and output. First, in the training context, it analyzes the tensions between GDPR requirements and large-scale AI training, the copyright framework under the CDSM Directive's text and data mining exceptions, and the AI Act's data governance obligations. I also offer a detailed assessment of the GDPR Omnibus proposals.
Second, at the model level, the chapter discusses the divergent rulings in Getty Images v Stability AI and GEMA v OpenAI on the question whether trained models constitute reproductions of copyrighted works. It also weighs in on the structurally parallel question whether models themselves constitute personal data. At the output level, it addresses hallucinations under the accuracy principle and proposes a strict liability regime modeled on pharmaceutical law.
The chapter identifies a structural symmetry between data protection and copyright: in both domains, legally protected content is encoded in model parameters in a manner that renders surgical excision technically near-infeasible. It argues, however, that the appropriate remedies diverge-property rules for data protection with specific exceptions, liability rules with outputbased remuneration for copyright. The chapter concludes with policy proposals for each level of the pipeline and advocates a regulatory approach that combines technology-neutral baseline protection with targeted technology-specific interventions.