No. 150: Training AI Models on Copyrighted Works under EU Law

Abstract

Training large language models and other generative AI systems requires datasets comprising hundreds of billions of documents drawn from the internet. A substantial proportion of these documents are copyrighted subject matter, raising the question of how this practice aligns with the rights of authors and other rightholders under European Union law. This thesis addresses this issue by first asking whether AI training constitutes copyright-relevant reproduction under Article 2 of the EU InfoSoc Directive, second examining to what extent the text and data mining exceptions of the EU DSM Directive allow for such use, and third evaluating how effectively the rightholder opt-out mechanism under Article 4(3) of that Directive balances the interests of AI developers and rightholders. Through doctrinal legal analysis based on the InfoSoc and DSM Directives and CJEU case law, the thesis concludes that AI training involves the right of reproduction at various stages of the training process. Without the temporary reproduction exception, a license or statutory exception is necessary. Article 3 of the DSM Directive excludes commercial AI developers, while Article 4 may only cover commercial training where lawful access exists and no valid opt-out has been exercised. This leaves data from unauthorized sources outside both provisions. Additionally, the opt-out mechanism is structurally inadequate due to the absence of technical standardization, asymmetric allocation of burdens, inability to address historical training data, and lack of dedicated enforcement. Based on these findings, the thesis concludes that EU copyright law provides a partial but inadequate framework for AI training. Drawing on recent litigation, licensing practices, and regulatory developments, the thesis proposes a set of targeted reforms, including a binding technical standard for machine-readable opt-outs, mandatory compensation for commercial text and data mining, a collective licensing scheme based on Article 12 of the DSM Directive, a harmonization clause linking the DSM Directive to the transparency obligations of the AI Act, and a regular review mechanism. The thesis argues that copyright law and AI development are not inherently irreconcilable and that the European Union has the institutional capacity to implement such a framework; the remaining question is one of political will.

Details

Author(s):
  • Violietta Hondza
Publish Date:
October 9, 2026
Publication Title:
European Union [EU] Law Working Papers
Publisher:
Stanford Law School
Format:
Working Paper
Citation(s):
  • Violietta Hondza, Training AI Models on Copyrighted Works under EU Law, EU Law Working Papers No. 150, Stanford-Vienna Transatlantic Technology Law Forum (2026).
Related Organization(s):

Other Publications By