Providers
Common Crawl
The crawl indexes of Common Crawl. No key. The newest index of each of the last five years queried one by one.
Source
Common Crawl
- provider: "commoncrawl"
- index.commoncrawl.org
Crawl indexes, one per year for the last five. A broken index skips, the rest answer.
- Format
- CDX text
- Paging
- newest index of each of the last five years
- Dates
- firstSeen · lastSeen
- Key
- none
Access
- Load
await create("commoncrawl") - CLI
urls discover example.com -p commoncrawl - Key
nothing, no key and no account - Endpoint
https://index.commoncrawl.org/collinfo.json + one CDX index per year
What it knows
Common Crawl publishes a few crawls a year, each with its own CDX index. Together they cover a big slice of the public web, and every row has a timestamp.
How it asks
First collinfo.json, the list of all crawl indexes. Then, for each of the last five calendar years, the newest index whose id names that year, queried for *.domain as CDX text. Only index URLs on the configured host are followed, so a strange entry in collinfo.json can't send your request somewhere else.
Traps
- The index frontend throws 502 and 504 now and then. One failing index is skipped and the other years still answer. A caller abort still stops everything.
- From some networks
index.commoncrawl.orgrefuses connections entirely. The provider's tests run on mocked HTTP for that reason, and a live failure there usually means the network, not the code. - Five years back is the whole range. Older crawls aren't asked.