Meta's copyright fight over AI training data has gained another uncomfortable layer after newly unsealed emails appeared to provide what one report called the strongest evidence yet in the case. The Ars Technica story in the packet says the material is being used by book authors who accuse Meta of training its AI systems on pirated books, and that the emails now visible in the record could strengthen their argument about how the company gathered data.
The key allegation in the source packet is that Meta torrented and seeded an 81.7 terabyte dataset containing copyrighted material. That figure is large enough to explain why the case has become so closely watched. If true, it suggests not just that a company used a lot of data, but that some of that data moved through file-sharing workflows that are hard to square with a clean licensing story. The unsealed emails, according to Ars, add internal context around that practice.
What gives the report legal weight is the phrase most damning evidence, which indicates the latest filings may be more than background color. The source excerpt says the emails complicate Meta's defense in the copyright case brought by book authors. In plain terms, the plaintiffs are trying to show that the company did not simply ingest public information or rely on clean permissions. They are arguing that the training pipeline included copyrighted books that had not been properly licensed.
The reporting is careful to keep the issue within the courtroom rather than turn it into a solved fact pattern. The packet says the emails are alleged evidence and that the claims are being discussed in a copyright case, which matters because courts still have to decide what the messages mean, how they were obtained and whether they prove the authors' central claims. Even so, the combination of torrenting, seeding and internal emails is the sort of detail that can shift a case from abstract policy debate to document-driven dispute.
That is why the story matters beyond Meta. AI companies across the industry have spent the past two years arguing about training data, fair use, licensing and whether the internet has become a free buffet for model builders. The Ars report shows how those arguments change once internal communications are put on the record. A line in an email can matter as much as a technical architecture diagram if it helps establish intent or process.
The packet does not resolve the case, and it does not say the court has ruled on the merits. What it does show is that the plaintiffs have obtained material they believe makes their claim stronger. For Meta, that means another round of legal scrutiny at a time when the company is already under pressure over how its models were built and what data they learned from.


