@ai4lawyersTikTok
Prompts, privilege & discovery

In re OpenAI, Inc. Copyright Infringement Litigation (user logs)

The district judge affirmed an order requiring protected production of 20 million de-identified consumer ChatGPT logs for sampling.

20 million logs ordered
01

Case background

Material facts

Copyright plaintiffs sought a large, randomized sample from retained consumer conversation and output logs to test their theories. OpenAI proposed narrower keyword-based discovery.
02

The decision

The court’s ruling

The magistrate ordered production of 20 million de-identified logs under strict protections, denied reconsideration, and the district judge affirmed. A later protocol order addressed a still-larger data reservoir.

Why it matters

What the decision means

The order requires protected, court-supervised discovery; it does not authorize public disclosure of user chats. Model-interaction data becomes merits discovery when the requesting party establishes relevance and a sound sampling method.
Limit of the ruling. Separate preservation and protocol orders form part of the same continuing MDL. A later sanctions motion was still only an allegation at the cutoff date.

Primary reading

Sources