Screen reader preview
To a screen reader, a page is one long sentence. Enter an address to get that sentence as a transcript, play it with a voice from your device and see which items get announced without a name.
From two dimensions to one
With your eyes you take in a page all at once: logo top left, menu across, article in the middle, sidebar to the right. Speech can’t do “all at once”. A screen reader such as NVDA, JAWS, VoiceOver or TalkBack walks through the HTML in source order and speaks one item at a time, with the item’s role in front of its text: “heading level 2, Pricing”, “link, Contact us”, “edit, Email, required”.
Nobody sits through all of that from start to finish. Experienced users navigate by type. One key jumps to the next heading, another lists every link on the page, another goes straight to the main landmark. So for them, the role and the name of each element are the whole interface. A button that looks obvious on screen but has no text gets announced as “button”, and that’s all they hear.
How the transcript is built
Our server downloads the page once and walks through the <body> the way assistive software does. The page <title> comes first, since it’s the first thing spoken when a page opens. After that, every element becomes a line:
- Landmarks: banner, navigation, main, complementary, content info, search, plus named regions and forms.
- Headings with their level, links with their destination, buttons with their state (expanded, collapsed, pressed).
- Form fields with their type and state: “check box, not checked”, “combo box”, “edit, email, required”.
- Images, frames, lists with their item count and tables with their row and column count.
We work out each element’s name in the order browsers use: aria-labelledby, then aria-label, then the native source (alt, <label>, text content), then title. Anything carrying hidden, aria-hidden="true" or an inline display:none is left out, and so is the content of a closed <details>. Images with alt="" are skipped and counted as decorative.
Press Listen and your browser’s speech engine reads the lines. Roles are spoken in English. The content is spoken with a voice that matches the page’s lang attribute when your device has one, and the voice changes for passages marked with their own lang.
What the flags mean
| Flag | What a listener hears | Fix |
|---|---|---|
| no alt | “Unlabeled image”, sometimes followed by the file name | Add alt text, or alt="" if the image is decoration |
| empty link, no name | “Link” or “button” alone | Visible text, an alt on the icon, or aria-label |
| no label | “Edit” with no question | <label for="…"> tied to the field’s id |
| placeholder only | A hint that may not be read at all | A real label next to the field |
| vague text, repeated text | “Read more” five times, each going somewhere different | Name the destination in the link |
| skips a level, empty | An outline with holes | Go from h2 to h3, never straight to h4 |
| no title, no headers | “Frame”; table values without their column | title on the <iframe>; <th> cells |
A red flag means the item has no name at all. An amber flag means it has one, but the name doesn’t help. The summary also reports a missing page language, a missing <h1> and a missing main landmark, and tells you how many items come before the main content. If that number is high, every visit to every page starts with the same long recital of your menu.
Three ways to use the filters
- Select Headings and read only that list. It should work as a table of contents for the page.
- Select Links. Each line has to make sense without the sentence around it.
- Select Fields and buttons, close your eyes and press Listen. You should be able to tell what each control wants from you.
Select any line to restart reading from there. The speed slider gives you a feel for the pace regular users pick, which is often well above normal speech.
An approximation, with known gaps
The transcript comes from the HTML as the server sent it. We don’t run JavaScript or apply style sheets, so content added by script is missing and content hidden by a CSS class still gets read. The listening time is an estimate based on about 170 words per minute. Very long pages stop at 1,200 items. Every real screen reader also has its own vocabulary and shortcuts, so test with at least one of them before a release. NVDA is free on Windows, and VoiceOver ships with macOS and iOS.
Questions people ask
Is this the same as testing with NVDA or VoiceOver?
No. It shows the order, roles and names a screen reader works from, and that’s where most problems are. It doesn’t reproduce interactive behavior such as menus opening, live announcements or focus moving into a dialog. Use it to find problems quickly, then confirm with a real screen reader.
Why does the voice read my page with the wrong accent?
Either the page doesn’t declare a language with <html lang="…">, or your device has no voice installed for that language. The status line above the transcript says which voice is in use. Real screen readers behave the same way.
Why is some text on my page missing from the transcript?
It was probably added by JavaScript after the page loaded, or it sits inside a closed <details> element or an aria-hidden container. We only read the HTML the server delivers.
What should a screen reader announce for a logo?
The name of the organization and, if the logo is a link, where it goes: alt="Acme, home page". Don’t write “logo” alone, and don’t leave the alt out, or the link gets announced by its file name or address.
How many links before the main content is too many?
There’s no fixed limit, because a skip link and a <main> element let people jump past them in one keystroke. Without those two, anything beyond a handful of items gets tiring by the third page.