Developer API Reference

Wayback Machine CDX API Downloader & Query Tutorial

The CDX Server API is the internal index engine behind Archive.org. Instead of scraping slow HTML calendar views, you can query the CDX API directly to retrieve clean JSON lists of all historical snapshots, status codes, and MIME types. Here is how it works.

Automate CDX Parsing with ReDrop

Skip managing pagination tokens and JSON filtering. ReDrop queries the CDX cluster automatically and packages your clean website in 42 seconds.

Try Online Downloader

Example Python CDX Extraction Script

import requests

domain = "example.com"
cdx_url = (
    f"https://web.archive.org/cdx/search/cdx?"
    f"url={domain}/*"
    f"&output=json"
    f"&filter=statuscode:200"
    f"&filter=mimetype:text/html"
    f"&collapse=urlkey"
)

response = requests.get(cdx_url, timeout=30)
data = response.json()

header = data[0]
snapshots = data[1:]
print(f"Total unique HTML snapshots discovered: {len(snapshots)}")

Related Technical Guides

Download Any Domain Directly from the CDX Index

Scan your target domain with ReDrop. Let our engine query CDX indexes, clean PBN traces, and export ready-to-host archives.