Skip to the tool
URLExtractor

Disclaimer

What extraction does and does not establish about an address, and the limits of the guidance published here.

Last updated

Extraction is not verification

This is the single most important thing on this page.

When a tool here returns an address, it is telling you: this string appears in the source you gave me, and it parses as a URL. That is all.

It is not telling you that:

  • the address resolves, or that anything is served from it;
  • the destination is safe, legitimate, or free of malware;
  • the address points where the surrounding text claims — link text can say one thing while the href says another;
  • the content behind it is accurate or current;
  • you are permitted to use, copy or republish anything found there.

Nothing on this site opens, scans or checks a destination. Treat every extracted address the way you would treat a link from an unknown source: look at it before you click it, and pay attention to the hostname.

Specific limits by tool

Website crawl. Follows links, so it cannot find pages nothing links to, pages behind a login, or navigation that only exists once JavaScript runs. It is bounded by your page and depth limits and by robots.txt. It is not an exhaustive site audit and does not claim to be.

Sitemap. A sitemap is a list a site chose to publish. Addresses in it can be stale, and pages that exist can be missing from it. It is not proof of existence or of indexing.

Image OCR. Optical character recognition on a picture. Small, blurry or compressed text will be misread, which is why the output is editable. A confidence score is the engine's certainty about characters and nothing more. If the picture shows link text rather than the address, the address is not in the file at all.

PDF. Reads link annotations and, optionally, URL-shaped visible text. An address that the document wrapped across two lines will come back split, because that is how it is stored.

YouTube. The live modes use the official YouTube Data API and can only reach public content. They do not download media.

Domain extractor. Uses the Public Suffix List for registrable domains. The list is maintained by volunteers and changes over time; a very recently delegated suffix may not be in the bundled copy.

The guides

Code examples and spreadsheet formulas are run against the environment named alongside them before publication. Software changes, so an example that worked when written may need adjusting later. Test anything you are about to run against your own data.

The guides describe techniques. They are not legal advice about what you may collect, store or publish.

Availability

This is a free service with no uptime guarantee. Limits and defaults may change.

Reporting a problem

If something here is wrong, tell us. Corrections are made promptly and the guide's modification date is updated when the change is material.

Share this pageWhatsAppXLinkedInFacebook