Providers

Common Crawl

The crawl indexes of Common Crawl. No key. The newest index of each of the last five years queried one by one.
IDcommoncrawl03 / 7index.commoncrawl.org
Source

Common Crawl

  • provider: "commoncrawl"
  • index.commoncrawl.org

Crawl indexes, one per year for the last five. A broken index skips, the rest answer.

Format
CDX text
Paging
newest index of each of the last five years
Dates
firstSeen · lastSeen
Key
none

Access

Loadawait create("commoncrawl")
CLIurls discover example.com -p commoncrawl
Keynothing, no key and no account
Endpointhttps://index.commoncrawl.org/collinfo.json + one CDX index per year

What it knows

Common Crawl publishes a few crawls a year, each with its own CDX index. Together they cover a big slice of the public web, and every row has a timestamp.

How it asks

First collinfo.json, the list of all crawl indexes. Then, for each of the last five calendar years, the newest index whose id names that year, queried for *.domain as CDX text. Only index URLs on the configured host are followed, so a strange entry in collinfo.json can't send your request somewhere else.

Traps

  • The index frontend throws 502 and 504 now and then. One failing index is skipped and the other years still answer. A caller abort still stops everything.
  • From some networks index.commoncrawl.org refuses connections entirely. The provider's tests run on mocked HTTP for that reason, and a live failure there usually means the network, not the code.
  • Five years back is the whole range. Older crawls aren't asked.