What Actually Happens When You Get a 404
Not all missing pages behave the same way, and the difference between a real 404 and a "soft" one is invisible in the browser but very visible to a crawler. It is also the most common self-inflicted SEO problem on otherwise healthy sites.
What 404 means in the protocol
404 Not Found is the server’s answer to a specific question: do you have a resource at
this path? The answer is no, and it is the only correct answer when that is true.
The response body is then irrelevant to the protocol. Whatever you return — a styled page, a short sentence, a photograph — the status code still says the resource does not exist. This separation is the key to the problem described below.
Why the browser hides the distinction
A user requesting a URL that does not exist receives your custom 404 page. So does a user who requests a URL that exists but has been deliberately configured to return a 404 response. To the browser, both look like a well-designed page.
The difference is entirely in the status line, and entirely invisible without a tool.
Soft 404s
A soft 404 is when a server returns 200 OK for a page that does not exist — either the
homepage, a search results page, or a near-empty template that renders successfully.
This happens more often than it sounds, because some frameworks and CMS configurations render a page for any path and let routing decide what to show, without the framework ever setting a failing status. The classic pattern is a serverless or single-page-app deployment where every path returns the application shell with a 200, and the client-side router decides to show a not-found view.
Why it is a problem:
It consumes crawl budget indefinitely. Every unique URL a crawler discovers returns 200, so the crawler keeps coming back, fetching a new shell each time. A site generating infinite non-existent URLs in its index is a site whose crawl budget is being spent on nothing.
It dilutes the signal. If any URL a crawler finds returns 200 with content, it is a candidate for indexing. Your own internal links to a misspelled path, an external link with a typo, and a scraper’s guessed URLs all become index candidates.
It is harder to detect than it should be. Because it looks correct in a browser, soft 404s survive casual QA indefinitely. They typically surface in a crawl report, in search-console coverage, or in an index bloat investigation.
The fix is not cosmetic: the server must set a real 404 status for paths that do not resolve, regardless of what the client-side router intends to display.
Other statuses that look like 404s
Worth knowing because they are confused with 404s regularly:
410 Gone means the resource existed and has been permanently removed. It is more explicit than 404 and lets a crawler drop the URL faster. Some site owners avoid it because it feels harsher, but it is the honest status for a deliberately retired page.
403 Forbidden means the server knows the resource and refuses access. Some crawlers treat a 403 on a page that used to be public similarly to a 404, but it is not the correct status if the content exists.
301 or 302 means the resource moved. If you redirect everything to the homepage, you have produced a soft 404 by another route — see below.
500 is a server error and should never appear in a crawl of healthy pages.
When redirecting is right, and when it is wrong
Redirect when the content moved. If an article is genuinely at a new URL and the old one
is permanently gone, a 301 is correct. It transfers any accumulated links and signals the
move.
Redirect when an equivalent page exists. A category that was renamed has a legitimate successor.
Do not redirect everything to the homepage. This is the most common bad pattern. It
produces a soft 404: a 200 response, real content, but content unrelated to the request.
A crawler that requests a mistyped URL is told the site has an unlimited supply of relevant
pages, which is precisely the signal that makes a site’s quality harder for a search engine
to judge.
Do not redirect a removed page to an unrelated article. It is technically a redirect and practically it is a dead end, because the new page does not answer the original request.
Finding out what is actually 404ing
Server access logs are the ground truth. Beyond those:
- Search-console coverage reports separate “Not found” from pages that are indexed or discovered, which distinguishes a genuine 404 from a URL that has never been seen.
- A crawler run of your own site will find internal broken links immediately. These are the highest-priority kind, because they are links you control.
- Server-side error monitoring captures 404s from real traffic, which is where links from other sites show up.
When you do look, sort the 404 paths into three groups:
Internal links to pages that do not exist. Your fault and urgent. A broken internal link wastes crawl budget and strands content you built.
URLs that used to exist. Someone linked to them and the content moved or was removed. Decide, per URL, whether a redirect to the successor is honest. If not, a 404 is the right answer.
URLs that never existed. Misspellings, scanner traffic, and generated guesses. Return 404 and do nothing else. A large volume of this is normal and not a problem.
What a custom 404 page is for
A custom 404 page has exactly two jobs, and neither is SEO.
Tell the reader what happened, plainly. “That page does not exist” is enough. An HTTP error code alone — a bare “404” with no context — is a dead end.
Offer a way forward. A search box, links to your main sections, and your most recent articles. This is where a custom page earns its place: a reader who mistyped a URL is momentarily willing to be redirected to something useful.
What a custom 404 should not do: redirect to the homepage automatically, because it hides the failure from both the reader and the crawler; or attempt to do a search for the requested path on the server, which produces a 200 with search results and is another soft 404.
Note also that the 404 page itself should carry noindex, since it is not content, and
should not appear in your sitemap.
The whole picture
404 handling is one of the clearer examples of a technical decision with no clever answer: return the status code that matches reality, keep real content at real URLs, redirect only when content has honestly moved, and make sure the paths your own site links to resolve.
Get that right and 404s become a non-issue. Get it wrong — usually by defaulting everything to 200 or to a homepage redirect — and the result is crawl budget spent on URLs that will never be content.
Topics
Frequently asked questions
Is a 404 bad for my SEO?
No. 404 is the correct response for content that does not exist, and returning it is better than masking the problem. What matters is volume and cause: a handful of 404s from external links is harmless, while a large share of your pages returning 404 indicates a broken publishing pipeline or an internal linking problem.
Should I redirect deleted pages to the homepage?
Generally no. Redirecting every missing URL to the homepage creates a soft 404 and tells search engines the site produces replacement content for anything, which erodes trust in the signal. If there is a genuinely relevant replacement, redirect to it. Otherwise return 404 and let the custom error page help readers.
How do I find out why pages are returning 404?
Your server access logs are the ground truth. Look at the status column for 404 responses, group by requested path, and separate the URLs that were never real from the ones that used to work. The second group is the urgent one, because those had inbound links and now have none.
Sources and references
- RFC 9110 — HTTP Semantics: 404 Not Found — Internet Engineering Task Force, accessed 2026-09-02
- Inspect a URL — Google Search Central — Google, accessed 2026-09-02