The Internet Archive and Digital Preservation: Why Saving the Web Matters

The Website That Doesn’t Exist Anymore

The average lifespan of a webpage is approximately 100 days according to research conducted on web content longevity. The link that worked last year produces a 404 error today; the news article that was cited in an academic paper is gone from its original URL; the product documentation that users relied on disappeared when the company shut down; the government report that informed policy decisions was removed when the administration changed. The web that feels permanent and always-available is, in practice, significantly more ephemeral than the paper archives that libraries have preserved for centuries.

The Internet Archive and its Wayback Machine (web.archive.org) exist to address this ephemerality: a non-profit organisation that has been crawling, archiving, and providing free public access to snapshots of websites since 1996, accumulating over 800 billion pages across multiple crawls over nearly three decades. It’s the closest thing to a permanent library of the web that exists, maintained by an organisation whose explicit mission is long-term digital preservation.

What the Wayback Machine Contains and How to Use It

The Wayback Machine’s interface is straightforward: enter a URL, and it shows the calendar of dates when that URL was crawled and archived, with the ability to view the page as it appeared at any specific historical date. For a heavily crawled page (major news sites, popular websites, frequently changing government pages), the archive may contain hundreds or thousands of snapshots across many years. For a niche or recently archived page, coverage may be sparser.

The most common uses of the Wayback Machine: recovering content from pages that have been deleted or changed (the original text of an article that was subsequently edited, a price that was changed, a policy that was updated), verifying what a website said at a specific date for research or legal purposes, accessing documentation for discontinued products and services, and general historical research into how web content has evolved over time.

What the Archive Preserves Beyond Websites

The Internet Archive’s scope extends well beyond website snapshots. Its collections include millions of books available for digital lending (scanned physical books in copyright that can be borrowed like a library, with ongoing legal dispute about the legality of this model), millions of audio recordings including concerts, radio broadcasts, and historical audio documents, millions of video recordings including television news clips and film archives, and software and video games from the MS-DOS and early console eras that are available to run in-browser through emulation.

The software preservation aspect is particularly valuable for computing history: the ability to run programs that were distributed in the 1980s and 1990s in browser-based emulators provides access to the actual computing experience of those eras rather than descriptions of it. Historical research about computing culture, interface design, software capabilities, and user experience is possible through actually running the software rather than reading about it — a preservation approach with no physical equivalent.

The Challenges of Digital Preservation

The Internet Archive faces legal, financial, and technical challenges that illuminate the difficulty of long-term digital preservation. The Controlled Digital Lending legal challenge (publishers suing the Archive over its digital book lending programme) threatens part of the Archive’s collection and has resulted in ongoing litigation. Funding a project that aims to preserve everything permanently requires ongoing donations without a traditional revenue model. And the technical challenge of maintaining hundreds of petabytes of archival data, keeping it accessible, and migrating it as storage technologies evolve is genuinely difficult engineering at scale.

The specific legal landscape around digital preservation is unsettled in ways that affect what the Archive can preserve and what it can make publicly accessible. Material that’s in copyright cannot be fully publicly accessible in the same way as public domain material; the legal framework for digital archiving access rights differs from the framework for physical library lending in ways that haven’t been fully resolved in US or international law.

How to Support Digital Preservation

Individuals can contribute to digital preservation in several ways. Financial support to the Internet Archive (archive.org/donate) directly supports the largest and most comprehensive general web archiving effort. Using and sharing the Archive when it’s useful — its institutional relevance to researchers, journalists, and the public is partly demonstrated by its usage metrics, which inform its importance to funders and policymakers.

The Wayback Machine’s Save Page Now feature (available at web.archive.org/save) allows anyone to submit a URL for immediate archiving — useful for preserving pages you want to ensure are archived before they might disappear. Browser extensions (Archive-It, Wayback Machine extension) can automatically archive pages you visit or allow easy one-click archiving of pages worth preserving. These individual actions contribute to the distributed effort of preserving the web that no single institution or budget can fully accomplish alone.

Related articles

Share article

Latest articles