Permavault › Web Archiving for Evidence and Citation

Web Archiving for Evidence and Citation

Courts have accepted Wayback Machine captures, generally with an affidavit from the Internet Archive or by judicial notice. The limit is not the law, it is coverage: the archive only holds what its crawler happened to reach, and for a growing share of the web it now holds nothing.

Published July 11, 2026 · Updated September 3, 2026

Web archiving covers two jobs that look similar and are not. One is preserving the past: a public archive crawls the web continuously so that history exists at all. The other is preserving a specific page, right now, because your brief, story, or paper depends on what it says today.

The Wayback Machine is unmatched at the first job and structurally unsuited to the second. That distinction is the whole of this page, and getting it wrong is how a citation quietly becomes unsupported three years later.

Pages do not stay put

In 2014, Harvard researchers Lawrence Lessig, Jonathan Zittrain, and Kendra Albert checked every web link cited in United States Supreme Court opinions. Roughly half, 49.9 percent, no longer led to the material the Court had cited. Not obscure blog posts in obscure cases: sources relied on in the reasoning of the highest court in the country, gone or changed.

A decade later the problem has compounded. Pew Research Center’s 2024 study found that 38 percent of web pages that existed in 2013 were no longer accessible ten years on, and a quarter of all pages that existed at any point between 2013 and 2023 had already vanished by the study date. On Wikipedia, 54 percent of articles contain at least one dead link in their references. About one in five government and news pages contains a broken link right now.

Two failure modes hide inside those numbers, and the second is worse.

Link rot is the obvious one: the URL returns an error, the page is gone. Annoying, visible, at least honest.

Reference rot is quieter. The URL still works, but the content is no longer what you cited. Among Supreme Court links that still technically resolved, the Harvard team found only 76 percent still contained the originally cited material. A page that changed under your citation is worse than a dead one, because it can now say something different while appearing to support you. A reader checking your source finds a current page, not the page you read, and has no way to know the difference.

For a brief, that is an authority that no longer says what you quoted. For a journalist, a source that now contradicts the story. For a researcher, a methods page or dataset that silently changed after publication.

Is a Wayback Machine capture admissible as evidence?

Wayback captures are used in litigation regularly and courts have accepted them, but not on their own. The reliable route is a witness.

In United States v. Gasperini, 894 F.3d 482 (2d Cir. 2018), the government’s Internet Archive screenshots were admitted because “the government presented testimony from the office manager of the Internet Archive, who explained how the Archive captures and preserves evidence of the contents of the internet at a given time.” That witness also compared the proffered screenshots against the Archive’s own records. The Second Circuit contrasted Novak v. Tucows, Inc., 330 F. App’x 204 (2d Cir. 2009), where the same kind of screenshots were properly excluded because the proponent “offered no testimony explaining its provenance.” The Third Circuit took the same approach in United States v. Bansal, 663 F.3d 634 (3d Cir. 2011), and the Seventh Circuit excluded archive screenshots for want of such a witness in Specht v. Google Inc., 747 F.3d 929 (7th Cir. 2014).

Notice what carries the weight. It is not the archive’s reputation; it is a witness explaining a specific process, obtained for your case, plus a snapshot that happens to exist for the date in dispute. Take either away and there is nothing to certify. The Internet Archive knows this: its own help pages tell users how to request affidavits for “certified records for use in legal proceedings.”

Judicial notice: the courts are split

Where there is no witness, parties often ask the court to take judicial notice of the archived page under Rule 201. Whether that works depends on where you are.

Against. In Weinhoffer v. Davie Shoring, Inc., 23 F.4th 579 (5th Cir. 2022), the Fifth Circuit reversed a district court for doing exactly this, holding “a private internet archive falls short of being a source whose accuracy cannot reasonably be questioned as required by Rule 201.” The sentence that matters most for anyone weighing archives against their own captures is this one: exhibits derived from these sources “are not inherently or self-evidently reliable in the same way as documents designated as self-authenticating by Rule 902.”

For. In Valve Corp. v. Ironburg Inventions Ltd., 8 F.4th 1364 (Fed. Cir. 2021), the Federal Circuit noted that district courts have judicially noticed Wayback contents “as facts that can be accurately and readily determined from sources whose accuracy cannot reasonably be questioned,” and said simply: “We agree.” The D.C. Circuit followed Valve in New York v. Meta Platforms, Inc., 66 F.4th 288 (D.C. Cir. 2023).

Read the two together and the split is narrower than it looks. In Valve the archived article’s publication date had been independently confirmed by a patent examiner, and in Meta the court noted other sources corroborated the archived text. The corroboration is doing the work. Weinhoffer, where the archived page was the only proof of the term in dispute, is the case that tells you what happens without it.

So the honest answer is conditional, and the conditions are where it usually breaks:

A Wayback capture is genuinely useful evidence when it exists and covers the point. It is not a preservation strategy, because you cannot make it exist on demand for the page in front of you.

Where the public archive falls short

The blocking problem got serious in 2026. More than 340 local news outlets now limit the Internet Archive’s access to their journalism, according to Nieman Lab’s reporting in May 2026, up from 241 in January. The publishers include some of the largest chains in the country: USA Today Co., McClatchy, Advance Local, Tribune Publishing, and Alden Global Capital subsidiaries. The New York Times has also limited the Archive’s access.

The motive is not hostility to preservation. Publishers fear AI companies scraping archived content for training data, and archived copies are a side door. But the effect is the same regardless of motive: for a growing share of the news, there is no Wayback snapshot to point to.

You do not control the crawl. The Wayback Machine archives what its crawlers reach, when they reach it. A page can be created, changed, and deleted between snapshots. If the version that matters existed for six hours on a Tuesday, the odds the crawler saw that exact version are not in your favor. On-demand saving helps when you think of it, but it is subject to the same blocks and capture limits.

Captures are often incomplete. Dynamic pages, JavaScript-heavy layouts, images served from third-party hosts, and interactive elements frequently do not archive faithfully. What replays may be a partial rendering of what a visitor actually saw, and for evidence, the missing element is often the one that matters.

Availability can change retroactively. The Archive has historically honored exclusion requests, and pages can become unavailable after you cited them. A snapshot you relied on is not guaranteed to stay reachable at that URL.

It is one institution. The Internet Archive is a single nonprofit that has spent recent years fighting major lawsuits and absorbing a serious 2024 security breach. It has weathered all of it so far, and the web is lucky it did. But a preservation strategy that depends entirely on one organization’s continued existence and policies is a strategy with a single point of failure.

Why saving a PDF is not an archive either

Saving a local copy is better than nothing, and for private working files it is fine. As a citation or evidence practice it has two gaps. Your reader cannot see your local copy, so the citation still points at a URL that will rot. And when the content is contested, a copy on your own machine proves little: there is no independent record of when it was made, or that it matches what was actually online.

The scholarly world’s institutional answer is Perma.cc, built by the Harvard Library Innovation Lab: it archives the cited page and gives you a stable link, held in library custody. For citation permanence in scholarship it is a genuinely good system, and if your institution participates, use it. Its model is trust in the custodian. A reader relies on the library consortium holding the copy faithfully, rather than on a proof they can check themselves.

What a capture you control has to prove

For work that may be challenged rather than merely followed, permanence is only half the answer. The other half is proof. A capture worth citing should be:

Permanent by design. Stored somewhere engineered to outlive companies and subscription payments, not on one vendor’s cloud or one person’s hard drive.

Fixed at a moment in time. The record shows exactly what the page said on the date you cited it, with the date written by the capture process rather than typed in by hand.

Independently verifiable. Cryptographic fingerprints, computed at capture, let anyone confirm the copy is unaltered without trusting you, your firm, or the archive that holds it. This is the same property that anchors a chain of custody for digital evidence, and you can check it yourself.

Complete. A full-fidelity capture of the page as it rendered, not a text scrape that drops the chart your argument depends on.

The right division of labor

Use the Wayback Machine for what it is unmatched at: history you did not know you would need. What did this company’s homepage claim in 2019? What did this page say before you ever got involved? No tool you run today can capture yesterday, and the Internet Archive is a public good worth supporting for exactly that reason.

Make your own capture for what you know matters now. The page in front of you that your story, case, or paper depends on deserves a capture you control: made at the exact moment you need it, complete as rendered, fingerprinted at capture, and stored somewhere that does not depend on any single organization’s crawler, policies, or survival.

The two are complements. Neither substitutes for the other.

Where Permavault fits

Permavault is built for the second job. Paste the URL and a neutral automated system captures the full page as it rendered, fingerprints every file with cryptographic hashes at capture, and stores the result on a permanent decentralized network of roughly 300 independent nodes, funded by a long-term storage endowment. You choose the moment. The capture is complete as rendered, timestamped by the process, and comes with a verifiable record anyone can check, without trusting us or any archive.

That gives you citations designed not to rot: cite the original URL for convention, and add the capture as the record of what it said. When someone checks your source in five years, what you cited is what they see, verifiable byte for byte.

$4.99 per capture. The price is on the page. No subscriptions, no sales calls, no annual contracts.

Blocked outlets are the reminder that archives you do not control can close to you, and half of what the Supreme Court cited is already gone or changed. Whatever you are writing, the pages you rely on are on the same clock. Capture them while they still say what you read.

This article is general information, not legal advice for any specific matter. Admissibility depends on the facts, the jurisdiction, and the judge.

Go deeper

The guides in this cluster, each covering one part of the problem in detail.

Need a web page preserved exactly as it exists right now?

Capture it with Permavault