Guide

Filters and scope

Default host scope. URL prefixes and globs. Extensions and query keys and substrings and dates. What each filter keeps and why the limit counts only survivors.

A busy domain has tens of thousands of archived URLs, and a good share of them is noise: font files, cache busters, templates somebody forgot to render. Filters cut that down inside each source, as the URLs arrive.

That detail matters. The collector applies every rule before it counts, so limit: 100 means 100 URLs that passed, not 100 raw rows with 90 of them thrown away afterwards.

Default scope: the domain and its subdomains

Without any option a URL stays only when its host is the domain you asked for or a subdomain of it. Ask for mozilla.org and developer.mozilla.org stays, mozilla.org.evil.example goes.

Sources don't always stay on topic. A search index can hand back a page on another host that merely touched yours, and the host check is the safety net. noScope: true turns it off when you want everything a source has.

URL scope: prefixes and globs

urlScope keeps URLs that match at least one pattern. urlOutScope drops URLs that match any. Both look at the whole URL, scheme included, not at the hostname.

ts
await wayback.discover("nuxt.com", {
  urlScope: ["https://nuxt.com/docs"],
  urlOutScope: ["*/_fonts/*"],
});

Is a pattern a prefix or a glob?

Without * it's a prefix with a boundary. https://nuxt.com/docs keeps https://nuxt.com/docs and https://nuxt.com/docs/getting-started, but not https://nuxt.com/docsearch. The next character after the prefix has to be /, ? or #.

With * it's a glob over the whole URL, and * matches anything, including nothing. *nuxt.com/docs* catches both schemes and www. in one go. Matching ignores case.

Wayback writes plenty of old URLs as http://…:80/. A prefix that starts with https:// won't see them. A glob that starts with * will.

Substrings

match keeps a URL that contains at least one of the strings. filter drops a URL that contains any. Both ignore case. Quick and blunt, and often all you need.

ts
await urlscan.discover("example.com", { match: ["api", "admin"], filter: ["logout"] });

Extensions and query keys

ext: ["js", "json", "bak"] keeps URLs whose path ends in one of those extensions. The extension comes from the last path segment, so /app.js?v=3 counts as js.

hasQuery: true keeps URLs that still have a query after tracking keys are gone. ?utm_source=x alone doesn't count, ?id=7 does. Handy when you're hunting for parameters.

Dates

from and to bound the capture date, both ends included. They take archive digits or an ISO date: from: "2019" starts on the first second of 2019, to: "2019" ends on the last one. 20190315 and 2019-03-15 work too.

A bad bound is rejected before any request goes out.

What happens to URLs without a date?

They stay. AlienVault OTX, URLScan and VirusTotal don't report when they saw a URL, and a date filter doesn't punish them for it. If you want dated URLs only, ask the archives: Wayback, Arquivo.pt, Vefsafn, Common Crawl.

Order of the checks

For each URL: date window, then dedup, then host scope, urlScope, urlOutScope, extension, query keys, and finally filter and match. The first URL of a kind wins, later duplicates only widen its firstSeen and lastSeen.

From the CLI

Every option has a flag. Lists are comma separated.

shell
urls discover nuxt.com -p wayback \
  --url-scope 'https://nuxt.com/docs' \
  --url-out-scope '*/_fonts/*' \
  --ext js,json --has-query \
  --from 2023 --to 2024

urls discover --help lists the short aliases.