Wayback Machine Scraper - URL Snapshot History avatar

Wayback Machine Scraper - URL Snapshot History

Pricing

from $0.28 / 1,000 snapshot scrapeds

Go to Apify Store
Wayback Machine Scraper - URL Snapshot History

Wayback Machine Scraper - URL Snapshot History

Get the full Internet Archive snapshot history for any URL: capture timestamps, snapshot links, status codes, MIME types and sizes. Filter by date range and match type. No API key, no browser. Independent tool, not affiliated with the Wayback Machine.

Pricing

from $0.28 / 1,000 snapshot scrapeds

Rating

0.0

(0)

Developer

Scrape Sage

Scrape Sage

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Share

Disclaimer: This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Internet Archive or any of its subsidiaries. All trademarks mentioned are the property of their respective owners. "the Wayback Machine" is referenced only to describe the publicly available website this Actor collects data from.

Get the full Internet Archive (Wayback Machine) snapshot history for any URL: every capture timestamp, snapshot link, HTTP status, MIME type, content digest and size. Filter by date range and match type (exact page, prefix, host, or whole domain). No API key, no browser.

What you get per snapshot

FieldMeaning
originalUrl / targetThe archived URL and the query it came from
timestamp / capturedAtCapture time (raw YYYYMMDDhhmmss and ISO)
snapshotUrlDirect link to the archived copy
statusCode / mimeTypeHTTP status and content type at capture
digest / lengthContent digest (to spot changes) and byte size

Input

{ "urls": ["example.com"], "matchType": "domain", "fromDate": "20200101", "maxSnapshotsPerUrl": 500 }
  • URLs - one per line. Match type: exact, prefix (path), host, or domain (host + subdomains).
  • Date range - fromDate / toDate as YYYYMMDD. Collapse duplicates - keep only distinct content versions.
  • Import from a file - paste a list, or link a public .txt/.csv, a Google Sheet/Drive link, or an Apify key-value-store record.
  • Output fields - trim every record to exactly the columns you need.

Leave everything empty and the run returns a small free sample so you can see the shape first.

Reliability

Reads the official Internet Archive CDX API with patient backoff. A URL with no snapshots, or an import that cannot be read, bills $0 and says why.

Honest limits

  • The CDX API is slow and rate-limits by IP. Heavily-archived URLs (major domains) can take tens of seconds, and very large domain queries may hit archive.org's rate limit - if a URL comes back empty with a "did not respond" note, space out the run or narrow the URL and re-run. You are never charged for a failed lookup.
  • domain/host matches can return very large result sets - raise the run memory and timeout for those.

Pricing

$0.0005 per snapshot on the FREE tier (tiered pricing lowers it with volume) - priced low because snapshot counts are high per URL. Only snapshots actually saved are billed; empty and failed lookups cost nothing.

Output views

  • Snapshots - original URL, captured date, snapshot link, status, MIME and size.

Use with AI assistants (MCP)

Available through the Apify MCP server - an agent can pull a URL's capture history, find when a page changed, or recover the last-archived version in one call.

Agent-ready: autonomous payments (x402 & Skyfire)

This actor is agent-ready - AI agents can discover it, run it, and pay for it autonomously, with no Apify account and no human in the loop. It uses pay-per-event pricing and limited permissions, so it qualifies for Apify's agentic-payment standards:

  • x402 - an open, HTTP-native payment protocol. Agents pay per run in USDC on the Base network directly through the Apify MCP server - no account, no API key.
  • Skyfire - agent-to-service payments for fully autonomous AI-agent workflows.

Building an AI agent, MCP tool, or autonomous data pipeline? This scraper is ready to plug in and pay as it goes.

Automate & schedule

Run this Actor on autopilot and pull results into your own stack:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'MY_APIFY_TOKEN' });
const run = await client.actor('scrapesage/wayback-machine-scraper').call({
"urls": [
"example.com"
],
"maxSnapshotsPerUrl": 20
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`Got ${items.length} records`);

Integrate with any app

Connect the dataset to thousands of apps - no code required:

  • Make - multi-step automation scenarios.
  • Zapier - push new records straight into your CRM or spreadsheet.
  • Slack - get notified when a scheduled run finds something new.
  • Google Drive / Sheets - auto-export every run to a spreadsheet.
  • Airbyte - pipe results into your data warehouse.
  • GitHub - trigger runs from commits or releases.

More scrapers from scrapesage

Related Actors in the same category:

FAQ

How is this Actor billed? Pay-per-event: you pay only for the results it delivers, with no monthly rental and no start fee. The per-event price is shown on the Pricing tab.

Can I schedule it and get results automatically? Yes - create a Schedule and add a webhook or an integration (Google Sheets, Slack, Make, Zapier) to push each run's dataset wherever you need it.

Which export formats are available? Every run's dataset can be downloaded as JSON, CSV, Excel (XLSX), XML, HTML or RSS from the Apify Console or the API.

Can I run it from code or an AI agent? Yes - through the Apify API and client libraries, or from Claude, ChatGPT and other assistants via the Apify MCP server.

Is it legal to scrape the Wayback Machine? This Actor collects publicly available data only. You are responsible for using the output in compliance with applicable laws (including data-protection law where personal data is involved) and the source's terms. See the Disclaimer below.

Is this an official the Wayback Machine tool? No. It is an independent, third-party Actor with no affiliation to, endorsement by or sponsorship from the Internet Archive. See the Disclaimer below.

Disclaimer

This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Internet Archive or any of its subsidiaries. All trademarks mentioned are the property of their respective owners.

"the Wayback Machine" and any related marks are the property of their respective owners and are used here only in a descriptive, nominative sense - to identify the publicly accessible website from which this Actor collects data. This Actor is not an official the Wayback Machine product, is not authorised or certified by the Internet Archive, and does not distribute the Wayback Machine software. It collects only publicly available information; you are responsible for ensuring your use of that data complies with applicable laws, regulations and the terms of the source website.

Need help?

Open an issue on the Actor's Issues tab, or visit the Apify help center. Feature requests are welcome - this Actor is actively maintained.