Soft 404 test

A page that doesn’t exist should answer with status 404. Plenty of sites answer 200 instead, or quietly send the visitor to the home page. This test asks your server for two addresses that can’t exist and reports what it says.

We ask for two addresses that were invented on the spot. No result is stored.

Or try github.com, wikipedia.org

Every response carries a number the visitor never sees

When a server answers a request, the first line of its reply is a three-digit status code. The page comes after. People read the page, and software reads the number. 200 means “here is what you asked for”, 301 means “it moved, look over there”, 404 means “nothing lives at this address”, and 410 means “it was removed on purpose”.

A soft 404 is what you get when the two disagree. The text on screen says the page is missing, or the visitor quietly lands on the home page, and all the while the number says 200. Search engines then keep crawling, and sometimes indexing, addresses that lead nowhere. Link checkers report every dead link as healthy. Uptime monitors stay green while a whole section of the site is gone.

What the test requests and how it rules

Our server builds two addresses from random characters, so they can’t match a real page. One is shaped like a folder (/probe-check-k7m2x9qa4t/) and the other like a file (/h3n8v2c5w7pq.html). We use two shapes because a web server often answers file-like requests itself and hands everything else to the CMS. Redirects are followed and the code of every hop is kept. If the site doesn’t answer over HTTPS at all, the test falls back to plain HTTP.

What came backVerdict
404 or 410 at the address requestedCorrect.
A redirect that ends on a 404 pageWarning. Crawlers understand it, but the address the visitor typed is replaced in the browser bar.
200 at the address requestedSoft 404.
A redirect to the home page (/, /en, /index.php) that answers 200Soft 404.
A redirect to any other page that answers 200Soft 404.
500 and aboveError. Something crashes on unknown addresses.
401 or 403Warning. Usually a firewall rule answering before the site does.

The error page gets a review too

When the code is right, the test reads the HTML of the error page and looks for what a lost visitor would need:

  • A way home. Any link to the root of the site or of a language version counts, including a clickable logo.
  • An explanation. The title, main heading or text has to contain something like “404”, “not found”, “does not exist” or “sorry”. Fewer than 15 words of content, with header, menu and footer left out, is noted as thin.
  • A search field. If there isn’t one, you get a remark. It doesn’t count as a fault.
  • A <title> and a lang attribute, so browser tabs and screen readers have something to announce.
  • Weight. More than 500 KB of HTML is flagged, because bots request missing addresses all day long.

The stock page of nginx, Apache, IIS or LiteSpeed is recognized and reported as such. It has no logo and no links, so visitors assume they’ve left your site. You’re also told when the two test addresses get clearly different pages, which means the web server and the CMS each have their own.

Turning a soft 404 into a real one

Apache
ErrorDocument 404 /404.html. Give a local path. With a full URL, Apache sends a redirect and the final code becomes 200.
nginx
error_page 404 /404.html;. Remove any =200 placed between 404 and the path.
WordPress
The core already answers 404. Look for a redirection or SEO plugin with an option such as “redirect 404 to home page” and switch it off.
Single-page apps
A catch-all rewrite to index.html answers 200 for everything. Render unknown routes on the server with a 404 status, or let the host do it (Netlify and Cloudflare Pages serve a root 404.html with the right code).

Check your change from a terminal with curl -I https://example.com/no-such-page, which prints the status line first.

Two random addresses at the root don’t cover a whole site. A shop, forum or blog installed in a subfolder can behave differently, so test an invented address there by hand. We don’t run scripts, so an error message drawn by JavaScript is invisible here. And a real page that’s merely empty, which search engines also label “soft 404”, can’t be detected from outside.

Questions people ask

What is a soft 404 error?

A response that tells people the page is missing and tells software everything is fine. It usually takes one of two forms: a “not found” page served with status 200, or an unknown address redirected to the home page.

Is redirecting all 404 pages to the home page bad for SEO?

Yes. Google treats a blanket redirect to the home page as a soft 404, so it passes no value, and visitors don’t understand why they landed there. Only redirect an old address to the page that replaced it, with a 301, and let the rest answer 404.

Should a deleted page return 404 or 410?

Both are correct. 410 Gone says the removal is deliberate and permanent, and search engines tend to drop the address slightly sooner. Use 404 when you’re not sure the page ever existed.

Do 404 errors hurt my rankings?

No. Missing addresses are a normal part of the web and aren’t held against a site. What does cost you is a link on your own pages that points to one, or a popular old page that now answers 404 when it could redirect to its successor.

Why does Search Console report a soft 404 on a page that exists?

Google also uses the label for real pages that look empty: a category with no products, a search with no results, a page whose content failed to load. Add content, or return 404 if the page has no reason to exist.

How can I see the status code of a page myself?

Open the browser developer tools, select the Network tab and reload. The first line shows the code. From a terminal, curl -I followed by the address does the same.