Prompts, privilege & discovery
In re OpenAI, Inc. Copyright Infringement Litigation (user logs)
The district judge affirmed an order requiring protected production of 20 million de-identified consumer ChatGPT logs for sampling.
20 million logs ordered01
Case background
Material facts
Copyright plaintiffs sought a large, randomized sample from retained consumer conversation and output logs to test their theories. OpenAI proposed narrower keyword-based discovery.
02
The decision
The court’s ruling
The magistrate ordered production of 20 million de-identified logs under strict protections, denied reconsideration, and the district judge affirmed. A later protocol order addressed a still-larger data reservoir.
Why it matters
What the decision means
The order requires protected, court-supervised discovery; it does not authorize public disclosure of user chats. Model-interaction data becomes merits discovery when the requesting party establishes relevance and a sound sampling method.
Limit of the ruling. Separate preservation and protocol orders form part of the same continuing MDL. A later sanctions motion was still only an allegation at the cutoff date.
Primary reading