Artificial Intelligence Training and Copyrighted Data: Reconciling Machine Learning Innovation with Exclusive Rights

Main Article Content

Prof. Sebastian R. Falk
Dr. Mirela T. Vuković

Abstract

The development of generative artificial intelligence systems has intensified legal debates concerning the use of copyrighted material in the training of machine-learning models. Modern artificial intelligence systems may require access to enormous collections of literary, artistic, musical, audiovisual, and other creative works during the training process. While such practices can contribute to technological advancement and the development of sophisticated computational systems, they may simultaneously involve acts of copying or processing that fall within the scope of copyright protection. This article examines the legal relationship between artificial intelligence training and copyright and evaluates the extent to which existing copyright doctrines can accommodate large-scale computational use of protected works. The analysis focuses on reproduction rights, text-and-data-mining exceptions, fair-use principles, licensing requirements, and the distinction between temporary computational processing and commercially exploitable outputs. Particular attention is given to the interests of authors and publishers whose works may be incorporated into training datasets without direct contractual relationships with AI developers. The article further examines whether the commercial nature of an AI system should influence the legal assessment of training activities and considers the evidentiary difficulties associated with identifying particular copyrighted works within large training datasets. Comparative analysis reveals significant differences among jurisdictions, with some legal systems providing explicit exceptions for computational analysis while others rely upon broader doctrines or emerging regulatory approaches. The article argues that a rigid application of conventional reproduction principles could unnecessarily restrict technological development, while unrestricted access to copyrighted datasets could undermine established economic incentives for creative production. It proposes a balanced framework combining clearly defined computational-use exceptions, licensing mechanisms for commercially significant datasets, transparency obligations, and appropriate safeguards for rights holders. The article concludes that the long-term relationship between artificial intelligence and copyright will require coordinated legal development capable of recognizing both technological innovation and the economic value of creative works.

Article Details

Section
Original Research Articles