Discovering URLs
Three ways to ask. They differ in who picks the source and in what a failure does.
Ask one source by name
import { create } from "@agntn/urls";
const urlscan = await create("urlscan");
const hits = await urlscan.discover("example.com", { limit: 100 });
The registry key is the name you pass everywhere: create(), the CLI's -p, the provider argument of every agent tool. Providers lists all of them.
create() also takes a config object: apiKey, baseUrl for your own mirror of the endpoint, and timeout in milliseconds.
const virustotal = await create("virustotal", { apiKey: process.env.VT_KEY, timeout: 20_000 });
Prefer a direct import? Each source is its own subpath: @agntn/urls/providers/wayback exports Wayback.
Let the library pick
import { discoverWithFallback } from "@agntn/urls";
const { provider, result } = await discoverWithFallback("example.com", { limit: 50 });
Which source goes first?
VirusTotal goes first when VIRUSTOTAL_API_KEY is set. Without it the first try is AlienVault OTX, the keyless default.
When does it move on?
Only when the failure is about access: a missing or rejected key, a payment wall, a rate limit, an operation the source doesn't serve. Then it tries the next source in registry order.
A network error or a bad domain doesn't fall through. It's thrown right away, because retrying the same broken request on every other source just hides the problem.
Ask every source at once
import { discoverAll } from "@agntn/urls";
const outcomes = await discoverAll("example.com", { limit: 20 });
for (const outcome of outcomes) {
if (outcome.error) console.warn(outcome.provider, outcome.error.message);
else console.log(outcome.provider, outcome.result.length);
}
Every source runs in parallel and reports on its own. Everybody talks at once and nobody waits for the slowpoke: one source timing out doesn't take the others down with it, and you see exactly who failed and why.
The limit applies per source, not to the whole comparison. Deduplication is per source too, so the same URL from Wayback and from Arquivo.pt shows up twice, once under each name. That's on purpose: two sources agreeing on a URL is information.
What a record looks like
interface DiscoveredUrl {
url: string; // as the source wrote it, first spelling wins
source: string; // registry key of the source
input: string; // the domain you asked about
reference?: string; // the query URL that returned it
ext?: string; // path extension, no dot
queryKeys?: string[]; // query keys left after tracking keys are dropped
firstSeen?: string; // earliest capture the walk read, ISO 8601
lastSeen?: string; // latest capture the walk read, ISO 8601
}
Only the archives with a CDX index give dates: Wayback, Arquivo.pt, Vefsafn and Common Crawl. AlienVault OTX, URLScan and VirusTotal leave firstSeen empty.
A walk that limit ends early hasn't read the captures after the cut, so firstSeen and lastSeen describe what it read, not the source's whole memory of that URL.
How deduplication decides two URLs are the same
normalizeUrl lowercases scheme and host, drops the fragment, sorts the query, strips tracking keys like utm_source and fbclid, and trims a trailing slash. http://example.com:80/ and http://example.com/ are one URL. http:// and https:// are two, and so are example.com and www.example.com.
How much comes back
Without a limit the library and CLI run a source to the end, up to its own page safeguard. AlienVault OTX stops after 20 pages, URLScan and VirusTotal after 50. The published ceiling is MAX_DISCOVER_RESULTS, 100 000 URLs, and a bigger limit is clamped to it.
The agent tools are stricter on purpose. See Agents.
Cancelling
const controller = new AbortController();
setTimeout(() => controller.abort(), 10_000);
const hits = await wayback.discover("example.com", { signal: controller.signal });
The signal reaches every request a source makes. An abort surfaces as your own reason, not as a network error.
When something fails
Failures are UrlsError or one of its subclasses: AuthError, RateLimitError, PaymentError, HTTPError, NotFoundError, InvalidInputError, UnknownProviderError. Query strings that look like secrets are redacted from every message before you see it.
Bad input fails before any request goes out. create() rejects an unknown key, and discover() rejects a domain it can't turn into a hostname.