Sourcelane

Wayback Machine Snapshots Scraper

Web & SEO Tools

Operated by Sourcelane✓ verified

A scraper that enumerates Internet Archive (Wayback Machine) snapshots for a URL, page path, hostname, or entire domain (including subdomains), producing per-capture metadata such as capture timestamp, direct snapshot URL, HTTP status code, MIME type, byte size, and a content digest for deduplication; supports match modes (exact URL, prefix, host, domain), date-range / status-code / MIME-type filtering, collapse/merge of identical captures by content digest, and cursor-based pagination using the Wayback Machine CDX resume key to retrieve large or unlimited numbers of captures across multiple runs—outputs structured, snapshot-level archival history and change-tracking metadata for each archived capture.

Verified Sep 28, 12:10 AM

$0.001

per archived snapshot

Up to 100 per call. Only pay for results returned; failed calls are refunded.

What people use it for

Trust & reliability

7d uptime trend
WindowUptimeSuccess ratep50p95Calls
24h————0
7d————0
30d————0

Calling contract

Call it through Sourcelane's gateway or MCP server. We run the connector, apply your agent's spend guardrails, and bill only the results returned.

Call it via Sourcelane

curl -X POST https://cool-puffin-608.convex.site/v1/call \
  -H "Authorization: Bearer sl_live_your_agent_key" \
  -H "Content-Type: application/json" \
  -d '{"listing":"wayback-machine-snapshots-scraper","params":{"url":"example.com","maxItems":5},"maxResults":10}'

Request params

FieldTypeRequiredDescription
urlstringrequiredThe URL to look up. Use 'Match Type = Prefix' or 'Domain' to scan an entire site rather than a single page.
matchTypestringoptionalHow to match the URL. 'Exact' only returns snapshots of the exact URL; 'Domain' returns snapshots for the domain and all subdomains.
dateFromstringoptionalOptional. Only return snapshots archived on or after this UTC date.
dateTostringoptionalOptional. Only return snapshots archived on or before this UTC date.
statusCodestringoptionalOptional. Only return snapshots with this HTTP status (e.g. 200, 301, 404).
mimeTypestringoptionalOptional. Only return snapshots with this MIME type (e.g. text/html, image/jpeg).
collapseDuplicatesbooleanoptionalIf on, identical captures (same content digest) are deduplicated so you only see snapshots where the content actually changed.
maxItemsintegeroptionalMaximum snapshots to return. 0 = no cap (paginates until everything is fetched or the run times out). Capped at 100 per call.
pageIdstringoptionalOptional. Paste the NEXT_PAGE_ID (CDX resume key) from the previous run's Key-value store to continue from where you left off.

Each result contains

FieldTypeDescription
timestampstringTimestamp
archivedAtstringArchived At
originalUrlstringOriginal Url
snapshotUrlstringSnapshot Url
statusCodeintegerStatus Code
mimeTypestringMime Type
contentLengthintegerContent Length
digeststringDigest

Use Wayback Machine Snapshots Scraper from Claude, ChatGPT or Cursor

Pick your tool and connect in under a minute. Then just ask — for example: “For these 50 domains, find contact emails and what e-commerce platform they run on.”

Connect Claude

Web, desktop and mobile. Paste one URL.

  1. 1Copy your personal connector URL
  2. 2In Claude open Settings → Connectors → Add custom connector
  3. 3Paste the URL and click Add — done
Manual setup (config files, REST, Python) +

Claude Code

Adds the Sourcelane MCP server with your key as a header.

terminal
claude mcp add --transport http sourcelane https://cool-puffin-608.convex.site/mcp \
  --header "Authorization: Bearer sl_live_your_agent_key"

Claude Desktop (config file)

Alternative to the connector URL: add to claude_desktop_config.json, then restart Claude.

claude_desktop_config.json
{
  "mcpServers": {
    "sourcelane": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://cool-puffin-608.convex.site/mcp",
        "--header",
        "Authorization:${AUTH_HEADER}"
      ],
      "env": {
        "AUTH_HEADER": "Bearer sl_live_your_agent_key"
      }
    }
  }
}

Cursor & Windsurf

Add to ~/.cursor/mcp.json (or Windsurf's mcp_config.json).

mcp.json
{
  "mcpServers": {
    "sourcelane": {
      "url": "https://cool-puffin-608.convex.site/mcp",
      "headers": {
        "Authorization": "Bearer sl_live_your_agent_key"
      }
    }
  }
}

VS Code

Save as .vscode/mcp.json. VS Code prompts for your key once.

.vscode/mcp.json
{
  "servers": {
    "sourcelane": {
      "type": "http",
      "url": "https://cool-puffin-608.convex.site/mcp",
      "headers": {
        "Authorization": "Bearer ${input:sourcelane-key}"
      }
    }
  },
  "inputs": [
    {
      "type": "promptString",
      "id": "sourcelane-key",
      "description": "Sourcelane agent key",
      "password": true
    }
  ]
}

ChatGPT Custom GPT (Actions)

Create a GPT → Actions → Import from URL, then Authentication: API Key, Bearer.

OpenAPI schema URL
https://sourcelane-amber.vercel.app/openapi.json

REST

One POST. Pass maxResults to cap cost.

curl
curl -X POST https://cool-puffin-608.convex.site/v1/call \
  -H "Authorization: Bearer sl_live_your_agent_key" \
  -H "Content-Type: application/json" \
  -d '{"listing":"wayback-machine-snapshots-scraper","params":{"url":"example.com","maxItems":5},"maxResults":10}'

Python, LangChain, CrewAI, OpenAI Agents SDK…

Wrap the REST call as a tool, or point an MCP client at the endpoint.

python
import requests

res = requests.post(
    "https://cool-puffin-608.convex.site/v1/call",
    headers={"Authorization": "Bearer sl_live_your_agent_key"},
    json={
        "listing": "wayback-machine-snapshots-scraper",
        "params": {"url":"example.com","maxItems":5},
        "maxResults": 10,
    },
    timeout=300,
)
body = res.json()
print(body["receipt"]["chargedMicros"], "micro-USD for", body["receipt"]["results"], "results")
print(body["data"])

Reviews

—

0 reviews

5
0
4
0
3
0
2
0
1
0

Leave a review

Loading reviews…