Guide

Discovering URLs

One source by name. The automatic pick with fallback. Or every source side by side with its own results and errors.

Three ways to ask. They differ in who picks the source and in what a failure does.

Ask one source by name

ts
import { create } from "@agntn/urls";

const urlscan = await create("urlscan");
const hits = await urlscan.discover("example.com", { limit: 100 });

The registry key is the name you pass everywhere: create(), the CLI's -p, the provider argument of every agent tool. Providers lists all of them.

create() also takes a config object: apiKey, baseUrl for your own mirror of the endpoint, and timeout in milliseconds.

ts
const virustotal = await create("virustotal", { apiKey: process.env.VT_KEY, timeout: 20_000 });

Prefer a direct import? Each source is its own subpath: @agntn/urls/providers/wayback exports Wayback.

Let the library pick

ts
import { discoverWithFallback } from "@agntn/urls";

const { provider, result } = await discoverWithFallback("example.com", { limit: 50 });

Which source goes first?

VirusTotal goes first when VIRUSTOTAL_API_KEY is set. Without it the first try is AlienVault OTX, the keyless default.

When does it move on?

Only when the failure is about access: a missing or rejected key, a payment wall, a rate limit, an operation the source doesn't serve. Then it tries the next source in registry order.

A network error or a bad domain doesn't fall through. It's thrown right away, because retrying the same broken request on every other source just hides the problem.

Ask every source at once

ts
import { discoverAll } from "@agntn/urls";

const outcomes = await discoverAll("example.com", { limit: 20 });

for (const outcome of outcomes) {
  if (outcome.error) console.warn(outcome.provider, outcome.error.message);
  else console.log(outcome.provider, outcome.result.length);
}

Every source runs in parallel and reports on its own. Everybody talks at once and nobody waits for the slowpoke: one source timing out doesn't take the others down with it, and you see exactly who failed and why.

The limit applies per source, not to the whole comparison. Deduplication is per source too, so the same URL from Wayback and from Arquivo.pt shows up twice, once under each name. That's on purpose: two sources agreeing on a URL is information.

What a record looks like

ts
interface DiscoveredUrl {
  url: string; // as the source wrote it, first spelling wins
  source: string; // registry key of the source
  input: string; // the domain you asked about
  reference?: string; // the query URL that returned it
  ext?: string; // path extension, no dot
  queryKeys?: string[]; // query keys left after tracking keys are dropped
  firstSeen?: string; // earliest capture the walk read, ISO 8601
  lastSeen?: string; // latest capture the walk read, ISO 8601
}

Only the archives with a CDX index give dates: Wayback, Arquivo.pt, Vefsafn and Common Crawl. AlienVault OTX, URLScan and VirusTotal leave firstSeen empty.

A walk that limit ends early hasn't read the captures after the cut, so firstSeen and lastSeen describe what it read, not the source's whole memory of that URL.

How deduplication decides two URLs are the same

normalizeUrl lowercases scheme and host, drops the fragment, sorts the query, strips tracking keys like utm_source and fbclid, and trims a trailing slash. http://example.com:80/ and http://example.com/ are one URL. http:// and https:// are two, and so are example.com and www.example.com.

How much comes back

Without a limit the library and CLI run a source to the end, up to its own page safeguard. AlienVault OTX stops after 20 pages, URLScan and VirusTotal after 50. The published ceiling is MAX_DISCOVER_RESULTS, 100 000 URLs, and a bigger limit is clamped to it.

The agent tools are stricter on purpose. See Agents.

Cancelling

ts
const controller = new AbortController();
setTimeout(() => controller.abort(), 10_000);

const hits = await wayback.discover("example.com", { signal: controller.signal });

The signal reaches every request a source makes. An abort surfaces as your own reason, not as a network error.

When something fails

Failures are UrlsError or one of its subclasses: AuthError, RateLimitError, PaymentError, HTTPError, NotFoundError, InvalidInputError, UnknownProviderError. Query strings that look like secrets are redacted from every message before you see it.

Bad input fails before any request goes out. create() rejects an unknown key, and discover() rejects a domain it can't turn into a hostname.