Web Archiving for Journalists When News Sites Block the Wayback Machine (2026)
For most of this century, journalists had a quiet assumption baked into their workflow: if a page mattered, the Wayback Machine probably had it. That assumption is now failing in both directions at once. Publishers are locking the archive out, and the material journalists themselves rely on keeps disappearing faster than any crawler can catch it.
This guide covers what changed, what the public archive can still do for a newsroom, and the archiving workflow for the sources your reporting actually depends on. It pairs with our general guide to Wayback Machine limitations; this one is about the journalist’s side of the problem.
The blocking wave, in numbers
Nieman Lab has been tracking the shift through 2026. In January, it reported that 241 news websites disallowed at least one Internet Archive crawler, with major publishers including The New York Times, The Guardian, and USA Today Co. limiting access over fears that AI companies might scrape the archive’s copies for training data. By May, the count had grown: “More than 340 local news sites across the United States are now limiting the Internet Archive’s ability to access and preserve their stories,” with many owned by five of the seven largest local news publishers in the country, including USA Today Co., McClatchy, Advance Local, MediaNews Group, and Tribune Publishing.
The Internet Archive’s own framing is telling. Mark Graham, director of the Wayback Machine, calls the Wayback Machine “collateral damage caught up in the conflict between AI companies and publishers.” Nobody in that conflict is trying to hurt journalists. But the effect lands on you anyway: for a growing share of the news, especially local news, there is no public snapshot to point to.
The irony is sharp, because journalism leans on the archive constantly. The Internet Archive notes that more than 100 news articles every month reference, cite, or rely on material preserved by the Wayback Machine.
Your own published work is rotting too
The decay is not only about sources. It eats published journalism from the inside. Harvard researchers John Bowers, Clare Stanton, and Jonathan Zittrain examined 553,693 New York Times articles containing roughly 2.28 million hyperlinks, and analyzed the 1.64 million deep links among them. Their finding, published with Columbia Journalism Review in 2021: “25 percent of all links were completely inaccessible,” and the rot compounds with age, from 6 percent of links made in 2018 to 72 percent of links made in 1998. More than half of articles containing deep links had at least one dead one.
That is the paper of record, with professional infrastructure behind it. The hyperlink you put in a story today points at content someone else controls, and on current evidence, its odds of surviving a decade are poor.
What disappears, and who deletes it
The past two years supplied case studies at every scale. In early 2025, public health pages and datasets came down from CDC and other federal sites until a federal judge ordered HHS, CDC, and FDA to restore them; the material was offline while reporters needed it most. In June 2024, Paramount shut down MTV News’ website, and as TheWrap reported, its archives, “more than 20 years of writing and reporting, have apparently been deleted entirely outside of independent archives.” Officials stealth-edit pages; corporations sunset whole archives; NewsDiffs, launched in 2012 to track exactly those quiet edits at major outlets, remains online as a historical archive but no longer appears to actively track new changes.
The pattern for a working journalist: the material you saw is a moment, not a fixture. If the story depends on it, the moment has to be preserved by someone whose incentives you control. That is you.
What the public archive still can and cannot do
None of this is an argument against the Wayback Machine, which remains one of the great public goods of the internet. It is an argument for knowing its edges. Save Page Now, the on-demand tool, is honest about them in the Internet Archive’s own help pages: “this method only saves a single page, not the whole site,” it “does not save any of the outlinks,” and “some sites prohibit crawling,” which after the blocking wave now includes hundreds of news sites. The Archive has also historically honored exclusion requests, and its own 2017 policy post describes the consequence: sites can become unavailable retroactively, and “we receive inquiries and complaints on these ‘disappeared’ sites almost daily.”
Use the Wayback Machine for history you did not know you would need. For the source in front of you right now, the one your story quotes, make your own capture.
A source-archiving workflow for reporters
Archivists who work with newsrooms have been saying this for years; ICIJ’s archivist put it plainly: “If you want a record of what you’ve created in the news media, you need to take action.” The same goes for what you cite. The workflow:
1. Capture at the moment of reporting. The instant a source page matters, capture it, before you email the press office for comment. Pages change fastest right after someone realizes a journalist is looking.
2. Capture the exact URL you will cite. Not the homepage, not your feed. The permalink, as rendered, with its date and context visible.
3. Use a capture that documents itself. A copy on your laptop proves little when a subject disputes what their page said. The capture should record its own URL, timestamp, and process, and fingerprint the files with cryptographic hashes, so your editor, your lawyer, and your critics can verify the copy is unaltered without taking your word for it. That is what makes the difference between “the reporter’s screenshot” and a record that survives a legal challenge.
4. Store it somewhere that outlives the story cycle. Corrections, lawsuits, and follow-ups arrive years later. A capture parked in a vendor’s cloud has the vendor’s lifespan; a capture on your hard drive has your laptop’s.
5. Still feed the public archive. Save Page Now costs nothing and builds the commons. Your own capture is the copy you control; the public copy helps everyone else. These are complements, not competitors.
Where Permavault fits
Permavault is built for step 3 and step 4 in one motion. Paste the URL and a neutral automated system captures the full page as it rendered, records the URL and timestamp, fingerprints every file with cryptographic hashes at capture, and stores the result on a permanent decentralized network of roughly 300 independent nodes, funded by a long-term storage endowment. The capture and its proof are designed to remain retrievable and verifiable independent of any vendor. Including us. No crawler schedule, no blocklist, no single organization whose policies can close the door later.
Each capture is $4.99. The price is on the page. An optional Certificate of Authenticity from $9 documents the capture process. The Legal tier adds a qualified electronic timestamp from Disig a.s., an EU-listed qualified trust service provider, applied to the signed capture manifest, plus an independent Bitcoin-anchored timestamp and a declaration template designed to support authentication under FRE 902(13) and 902(14). Under eIDAS Article 41, a qualified electronic timestamp carries a presumption of the accuracy of its date and time in EU courts.
The public archive is being fenced off from the news, and the news is rotting on its own schedule. The version of record for your reporting is the version you captured.
This article is general information, not legal advice. Publisher blocking figures and platform behaviors reflect published reporting and the Internet Archive’s own documentation as of July 2026.
Need a web page preserved exactly as it exists right now?
Capture it with Permavault