Crawler view

Most bots never draw your page. They take the HTML your server sends, keep the words and leave. This check shows those words in the order a bot meets them, and flags what’s missing.

One request for the page and one for its robots.txt. Neither is stored.

Or try wikipedia.org, react.dev

Every page exists in two copies

The first copy is the HTML document that leaves your server. The second is what a browser builds from it once it has loaded the style sheets, run the scripts and fetched whatever those scripts ask for. People only ever see the second copy.

Bots mostly work from the first. Googlebot does render pages with a headless Chrome, but in a later pass that can lag behind and can time out. Bingbot renders less. The crawlers that feed AI assistants, and the preview fetchers of chat and social apps, generally don’t run JavaScript at all. If your text only exists in the second copy, a good part of your machine audience is getting an empty page.

This check requests the first copy and stops there. Nothing gets executed and no style sheet is applied.

How to read the plain-text view

We walk the page body in source order and print it as text. Scripts, styles, templates and inline SVG drawings are skipped. The rest is marked like this:

#, ##, ###
A heading and its level. Try reading only these lines. They should work as a table of contents for the page.
Underlined words
A link. Hover it to see the address it leads to.
[Image: …]
The alt text of a picture. An image with no alt attribute is shown as missing, with its file name. Buttons and form fields appear in brackets the same way.
Text marked as hidden
Words that are in the HTML but carry the hidden attribute or an inline style such as display:none, visibility:hidden, font-size:0 or a large negative text-indent.
Text marked noscript
Content inside <noscript>, which only bots and visitors without JavaScript receive.

Below the text you get the numbers: HTML size, count of visible words, share of the HTML that is text, scripts, links and images. A text share between 10 and 30% is typical of a hand-built or CMS page. At 1 or 2%, the document is mostly code.

What triggers each finding

FindingRule applied
Content only after JavaScriptFewer than 60 visible words, together with one of: an empty application container (#root, #app, #__next, #__nuxt, app-root and similar), three or more external scripts, over 50 KB of inline script, or over 150 KB in the main script files.
Very little textFewer than 150 visible words.
Page asks not to be indexednoindex or none in a robots meta tag or an X-Robots-Tag header.
Google not allowedThe page path matches a Disallow rule that applies to Googlebot in robots.txt.
More hidden than visible textHidden words outnumber visible words, with a floor of 80.
Links bots cannot followhref="#", href="javascript:…", or an onclick with no href.
No level 1 heading, images without altCounted directly in the HTML.

The last block applies your robots.txt to this one address for ten crawlers: Googlebot, Bingbot, Google-Extended, GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, CCBot and Applebot-Extended. A block on Googlebot counts as an error. A block on an AI crawler is just reported, since that’s a choice you’re entitled to make. The AI bot access check covers a longer list for the whole site.

Getting your text into the first copy

If the check found an empty container, the page is a client-side application. The server sends a shell and the browser fills it in. Here are the remedies, from most to least thorough:

  1. Render on the server. Next.js, Nuxt, SvelteKit, Remix and Angular can all produce full HTML for each request and then attach the scripts to it. It’s usually a configuration choice, not a rewrite.
  2. Generate static pages at build time. For content that rarely changes, tools such as Astro, Gatsby or the static export of the frameworks above write a complete HTML file per page.
  3. Prerender for bots. A prerendering service keeps a rendered snapshot and serves it to crawlers. It works, but now you have two versions to look after.

On a classic CMS such as WordPress or Drupal the HTML already holds the text, and the findings are smaller: a price shown only inside a picture, a menu whose links are onclick handlers, or a phone number in a banner image with an empty alt.

Hidden text is detected only from the hidden attribute and inline styles. Text hidden by a class in an external style sheet counts as visible here, because we don’t load style sheets. Collapsed menus, tabs and accordions are normal, and search engines don’t penalize them.

Questions people ask

Does Google run JavaScript when it crawls a page?

Yes, but as a second step. Google first indexes the HTML it downloads, then queues the page for rendering in a headless Chrome. That step can come later, and it may not wait for slow scripts or for content loaded on scroll or click.

Do ChatGPT, Claude and Perplexity crawlers run JavaScript?

As a rule, no. GPTBot, ClaudeBot, PerplexityBot and CCBot download the HTML and read its text. A page that needs scripts to show its content gives them little or nothing.

Is hidden text bad for SEO?

Not when it’s part of the interface. Menus, tabs, accordions and dialogs are indexed normally. What search engines treat as spam is text hidden from visitors on purpose to add keywords, such as white text on a white background.

What is a good text-to-HTML ratio?

There’s no target and it isn’t a ranking factor. Pages built with care often land between 10 and 30%. A very low figure is only a hint that the document carries a lot of markup or inline script for the amount of content.

Why does the view show almost nothing for my site?

Your server sends an application shell and the content gets added in the browser. Open the page source (Ctrl+U or Cmd+Option+U) and search for a sentence from your page. If it isn’t there, bots that skip JavaScript can’t see it either.