Dead link finder
Give us a starting address. Our crawler walks the site page by page, tries every link and image it meets, and brings back the ones that fail along with the pages they’re on.
Links break without anyone touching them
When you add a link, you’re counting on someone else’s server, or on a file you might rename next year. Your page only stores the address. The target moves, a domain doesn’t get renewed, an image is deleted from the media library, and your page still looks exactly the same in your editor. The only person who finds out is the one who clicks.
This slow decay is called link rot, and it never stops. A 2024 study by the Pew Research Center found that 38% of the web pages that existed in 2013 could no longer be reached ten years later. To know where your site stands, somebody has to try every link. That’s a miserable job for a person and an easy one for a program.
How the crawl proceeds
- We fetch the start page and read
robots.txt. Paths it disallows are left alone. - Every
<a href>and<img src>is collected, including lazy-loaded images indata-src.mailto:,tel:andjavascript:links are skipped, and the#fragmentis dropped. - Links to pages on the same host go into the queue to be read in turn.
example.comandwww.example.comcount as one site. Other subdomains don’t. - Everything else (other sites, PDFs, images, downloads) is only tested. We start with a
HEADrequest, which asks for the status without the body, and fall back to a normalGETif the server refusesHEAD. - Redirects are followed by hand, up to five hops. A sixth hop, or an address that redirects to itself, is reported as a loop.
The crawl stops after 100 pages read or 600 requests. It opens at most four connections to a host at a time, waits six seconds for an answer, and tests each address once no matter how often it’s linked. When a path shows up with more than five different query strings, the extra variants are ignored. Otherwise a calendar or a filter page would eat the whole budget.
Reading the four lists
- Broken links
- The target answered with an error, or didn’t answer.
404and410mean the page is gone.500to504point to a failing server and may be temporary. “Domain name not found” often means the whole site has closed. “Invalid SSL certificate” means browsers show a warning before the page. - Broken images
- Same causes, but what the visitor sees is an empty frame or a broken-picture icon.
- Internal redirects
- Links to your own pages that only work through a redirect. Nothing is broken, but each click costs an extra round trip, and somebody may remove the redirect one day.
- To check by hand
- The target answered
401,403,429or999, so it refuses robots or wants a login. LinkedIn, Facebook and many sites behind bot protection do this. You’ll have to open the link yourself to settle it.
Each row lists up to five pages where the address appears, and the failures repeated most often come first. A dead link in a footer shows up on every page, and you fix it in one place.
Repairing what the crawl finds
- Your own page moved. Correct the link, then add a permanent redirect for people arriving from elsewhere. Apache:
Redirect 301 /old-page/ /new-page/. nginx:location = /old-page/ { return 301 /new-page/; }. - Another site moved its page. Search that site for the new address. If the content is gone, link to a copy on the Wayback Machine (
web.archive.org) or remove the link. - Missing image. Upload it again under the same name, or edit the page.
- Internal redirect. Replace the address in the link with the final one shown in the report.
We only see links that are in the HTML the server sends. Menus built by JavaScript, CSS background images, srcset variants, videos and scripts aren’t tested, and neither are pages behind a login. Anchors aren’t verified either, so a link to /page#prices passes if /page exists. And a site that answers 200 for missing pages hides its dead links from every crawler. The soft 404 test detects that case.
Questions people ask
Do broken links hurt SEO?
A few dead links won’t lower a site’s ranking. What they cost you is something else. A broken internal link stops crawlers from reaching the page behind it, the value that link passed along is lost, and visitors who hit an error tend to leave.
Why is a link reported as broken when it opens in my browser?
The target doesn’t treat our crawler the way it treats you. It may block automated requests, answer too slowly for the six-second limit, serve a certificate your browser already has an exception for, or need a cookie you hold. Open the link in a private window to compare.
Does the crawler respect robots.txt, and how do I block it?
Yes. It identifies itself as MissWrenchBot and follows the group written for that name, or the * group when there is none. To keep it out, add User-agent: MissWrenchBot followed by Disallow: /. There’s more on the crawler page.
My site has more than 100 pages. What can I do?
The crawl reads pages in the order it discovers them, starting from the address you give. Run it several times from different entry points, say the blog index and then the product catalog, to cover different parts of the site.
How long are the results kept?
24 hours. While a crawl runs, the address in your browser bar changes to a link to that crawl. You can bookmark it or send it to a colleague until it expires.
What is the difference between a 404 and a 410?
404 Not Found says nothing exists at the address right now. 410 Gone says the page was removed on purpose and won’t come back. For a link on your site, the repair is the same.