A group of news publishers led by The New York Times filed a motion on July 9 in the US District Court for the Southern District of New York seeking sanctions against OpenAI. The publishers accused OpenAI of misleading the court about its ability to search copyrighted news content used to train its AI models and of deleting or compressing billions of ChatGPT conversation logs during an ongoing copyright dispute, according to medianama.com.
The publishers allege that OpenAI falsely claimed it could not search its training datasets or ChatGPT output logs for their copyrighted content, despite having conducted such searches before the lawsuit was filed. Ian Crosby, The New York Times’ lead attorney, said OpenAI lied to the court, the public, and the publishers by stating that searching ChatGPT outputs was infeasible and invasive of user privacy, while secretly performing those searches. The motion also accuses OpenAI of failing to preserve evidence by deleting or compressing conversation logs, making them unavailable for discovery.
This legal action highlights the increasing scrutiny AI companies face regarding the use of copyrighted material in training datasets and transparency in litigation. The case is part of broader copyright disputes involving AI models and content creators, raising questions about data handling and compliance. The New York Times’ motion underscores the challenges publishers face in protecting their intellectual property against AI training practices.
The court document detailing the sanctions motion was filed on July 9 in the US District Court for the Southern District of New York, marking a significant step in the ongoing legal battle between news publishers and OpenAI over AI training data and copyright issues.