The Electronic Frontier Foundation has warned that publishers’ efforts to stop automated collection of their work could also prevent the Internet Archive from preserving news pages for future research. In a commentary, the digital-rights organization argued that disputes over artificial-intelligence training should be separated from the public-interest role of web archives.

EFF said The New York Times had begun using technical measures beyond the conventional robots.txt system to stop the Internet Archive from crawling its site. It added that other newspapers, including The Guardian, appeared to be moving in a similar direction. The account presented those restrictions as a threat to a record used by journalists, historians, researchers and courts, rather than merely as a limit on AI companies.

The concern arises from the changing nature of online publication. Articles can be updated, removed or replaced, and an archived copy may be the only accessible evidence of how a page appeared at an earlier point. EFF compared the Archive’s role to that of a physical library retaining newspapers, emphasizing that preservation serves a different purpose from building a commercial generative-AI product.

The group cited the scale of the Wayback Machine, saying it contains more than one trillion archived pages. It also relayed figures from Internet Archive staff indicating that Wikipedia links to over 2.6 million archived news articles across 249 languages. Those numbers were used to illustrate how blocking a preservation crawler can affect a wider ecosystem of citations and research, even when a publisher’s immediate objective is to control reuse of its own material.

Publishers have legitimate interests in how their work is copied and used, the commentary acknowledged, and some are pursuing litigation over the use of copyrighted material to train AI models. EFF nevertheless took the position that archiving and search are transformative activities protected by established fair-use principles. It pointed to court treatment of Google’s copying of books for a searchable database as an analogy for making collections discoverable. That legal interpretation is EFF’s argument; the supplied source does not resolve the separate lawsuits over AI training.

The policy distinction is central to the warning. A broad technical block may not identify whether a crawler belongs to a commercial AI company, a nonprofit archive or another research tool. EFF’s view is that cutting off the Archive would impose a lasting public cost without settling the underlying copyright dispute. If pages disappear before they can be preserved, it argued, later changes in law or publisher policy cannot reconstruct the missing historical record.