Developer API Reference
Wayback Machine CDX API Downloader & Query Tutorial
The CDX Server API is the internal index engine behind Archive.org. Instead of scraping slow HTML calendar views, you can query the CDX API directly to retrieve clean JSON lists of all historical snapshots, status codes, and MIME types. Here is how it works.
Automate CDX Parsing with ReDrop
Skip managing pagination tokens and JSON filtering. ReDrop queries the CDX cluster automatically and packages your clean website in 42 seconds.
Example Python CDX Extraction Script
import requests
domain = "example.com"
cdx_url = (
f"https://web.archive.org/cdx/search/cdx?"
f"url={domain}/*"
f"&output=json"
f"&filter=statuscode:200"
f"&filter=mimetype:text/html"
f"&collapse=urlkey"
)
response = requests.get(cdx_url, timeout=30)
data = response.json()
header = data[0]
snapshots = data[1:]
print(f"Total unique HTML snapshots discovered: {len(snapshots)}")Related Technical Guides
Download Any Domain Directly from the CDX Index
Scan your target domain with ReDrop. Let our engine query CDX indexes, clean PBN traces, and export ready-to-host archives.