USA Today Co and several newspapers it owns have sued OpenAI in federal court in Manhattan, alleging that the company infringed their copyrights by using newspaper content to train large language models. The lawsuit, filed Thursday, October 8, adds another dispute over the use of copyrighted journalism in the development of artificial intelligence.
The plaintiffs’ central allegation concerns how OpenAI developed its models, rather than simply the availability of newspaper articles online. They claim that material protected by their copyrights was used for training. Those allegations remain unproven: the filing of a complaint does not establish infringement, and the merits of the claims have not been resolved.
The distinction between an allegation and a legal finding is important in copyright cases involving AI. A court must consider the rights at issue, the conduct challenged and any applicable defenses before determining liability. The available account of this lawsuit does not identify the particular articles involved, the models at issue or the remedies the newspapers are seeking. It also does not provide OpenAI’s response to the complaint.
Large language models are developed using substantial amounts of text to learn statistical relationships in language. That training helps systems generate responses to prompts, but it has also raised questions about the legal treatment of the material used in the process. Copyright disputes can concern copying during the collection and preparation of training data, as well as the separate question of whether a system’s output reproduces protected expression. The allegation described in this case focuses on the use of newspaper content for training.
Copyright generally protects the original expression in a news article, not the underlying facts it reports. That distinction allows others to report the same events without acquiring ownership of another publisher’s wording. It does not, however, settle whether copying protected articles for a technological purpose is lawful. That question depends on the conduct involved and the applicable copyright rules.
In the United States, fair use is one framework courts can apply when evaluating the unauthorized use of copyrighted works. The analysis considers the purpose and character of the use, the nature of the original work, how much was used and the effect on the potential market for that work. It is a case-specific inquiry rather than an automatic exemption for technology companies or a guarantee that every unlicensed use amounts to infringement. The available details do not establish which defenses OpenAI will raise here.
The lawsuit places USA Today Co and its participating newspapers among publishers pursuing copyright claims over AI training. For this case, the questions remain whether the plaintiffs can establish the alleged use of their protected material and whether that conduct creates liability under copyright law. The complaint begins that legal process; it does not answer those questions.
Sources: mezha.net