ID@agntn/urlsv0.1.0
Every URL a domain
left behind.
Web archives and threat intel indexes behind one TypeScript API. Ask what they've seen for a domain, cut the noise with filters, and give your agent the same tools over MCP. The site itself never hears about it.
- Sources
- 7
- 6 keyless
- Agent page
- 100 URLs
- per source
- Tools
- 2
- 4 hosts
$ pnpm add @agntn/urlsnuxt.com
155 unique URLs from 4 of 7 sources, at most 60 each. Nobody visited the site to get them.
Paths
- nuxt.com/_fonts47
- nuxt.com/__nuxt_content14
- nuxt.com/__og-image__11
- nuxt.com/.nuxt8
- nuxt.com/6
- 61 more69
- URLs
- 155
- Answered
- 4 of 7
- Seen
- 2004 → 2026
- In 2+ sources
- 15
Sources
- AlienVault OTXHTTP 0 from https://otx.alienvault.com/api/v1/indicators/domain/nuxt.com/url_list?page=2
- Arquivo.pt60 URLs, more behind the limit
- Common CrawlHTTP 0 from https://index.commoncrawl.org/collinfo.json
- URLScan5 URLs
- Vefsafn60 URLs, more behind the limit
- VirusTotalAuthentication failed for virustotal: Set VIRUSTOTAL_API_KEY or pass apiKey in config
- Wayback Machine60 URLs, more behind the limit
Keep the URLs you came for
A popular domain has tens of thousands of archived URLs, and a good share of them is fonts, tracking junk and templates somebody forgot to render. Filters run inside every source as the URLs stream in, so a limit counts what you kept, not what you threw away. Try the chips, it's the real collector running on the sample above.
- Default scope keeps the domain and its subdomains, nothing else
- Prefix or glob over the full URL, extension, query keys, substrings
- Dates bound what an archive saw, undated hits stay in
- Unique
- 155 of 185 records
- Kept
- 155
Kept
- arquivohttps://nuxt.com/
- arquivohttp://nuxt.com/
- arquivohttps://www.nuxt.com/
7 sources, one record shape
CDX text, CDX NDJSON, paged JSON, search cursors. Each provider turns its own format into the same record: the URL, which source saw it, and when, if the source keeps dates. Modules load lazily, so asking Wayback alone never pulls in the other 6. Only VirusTotal insists on a key.
Loadawait create("wayback")
| Format | What it knows | ||
|---|---|---|---|
| AlienVault OTX | JSON pages | Threat intel URL lists. Keyless, but the public endpoint gets grumpy under load. | no key |
| Arquivo.pt | CDX NDJSON | The Portuguese web archive. CDX with dates, and a hard cap per request. | no key |
| Common Crawl | CDX text | Crawl indexes, one per year for the last five. A broken index skips, the rest answer. | no key |
| URLScan | JSON search pages | Pages people scanned. Answers without a key at small volumes, more with one. | key optional |
| Vefsafn | CDX NDJSON | The Icelandic web archive. Ignores limit and sends everything, the collector stops it. | no key |
| VirusTotal | JSON v3 pages | URLs VirusTotal has seen for the domain. No key, no answer. | key required |
| Wayback Machine | CDX text | The Internet Archive's CDX index. One streamed answer with dates, subdomains included. | no key |
Two tools for your agent
An agent doesn't need tens of thousands of URLs in its context. It gets a bounded page per source and a flag saying there's more, so it can narrow the call instead of drowning in it. Every host runs the same executor, so the answer is the same wherever you plug it in.
- urls_discover and urls_providers on MCP, AI SDK, Pi and OMP
- 100 URLs per source unless the agent asks for more
- hasMore says when the limit cut the list, truncated when the source did
urls_discover
A page per source with its count and a flag when there was more. A failed source says why.
- urls
- 185
- pages
- 3 with hasMore · 3 with an error
Start with one command
One install gives you the library, the urls CLI and the MCP server. Pick a source, hand it a domain and get records back, not a blob to parse again.
- PinPre-1.0 (v0.1.0), so pin exact versions.
- Keys6 sources answer with nothing set. VirusTotal insists on a key.
- LeadsA hit means a source saw the URL once. Whether it still loads is your call.
First call
$ pnpm add @agntn/urlsimport { create } from "@agntn/urls";const wayback = await create("wayback");const hits = await wayback.discover("example.com", { limit: 20 });hits[0].url; // "http://example.com:80/"