Garbled text fixer
When é turns into é and an apostrophe into ’, the letters usually aren’t lost. Something read them with the wrong character table. Paste the text and we read it again the right way.
Repaired text
repaired best guess, verify gone, type it again · Point at a highlighted word, or tap it, to see the original characters.
The repair runs in your browser. Your text is not sent anywhere.
One letter, two ways to read it
There are no letters in a file. There are numbers called bytes, and an encoding is the table that says which number stands for which character. Almost everything written today uses UTF-8, where plain letters take one byte, accented letters take two and typographic signs such as curly quotes or the euro sign take three.
Windows-1252 is an older table where every character is exactly one byte. Hand UTF-8 text to a program that reads it with that older table and each byte comes out as a character of its own. The two bytes of é (C3 A9) show up as à and ©. The damage has a name borrowed from Japanese, mojibake, and you’ve probably met it in one of these forms.
| Written | Displayed | Where you meet it |
|---|---|---|
| é è | é è | café, crème |
| ü ö ß | ü ö ß | Zürich, Köln, Straße |
| ñ | ñ | señor, España |
| ã ç | ã ç | São Paulo, ação |
| ø å | ø Ã¥ | København, Århus |
| ’ “ — | ’ “ — | don’t, quoted speech, dashes |
| € | € | prices |
How the repair is done
The tool walks through your text and stops at every run of characters from the upper half of Windows-1252. It turns each one back into the byte it came from, then checks whether those bytes form a legal UTF-8 character.
Legal means a lead byte between C2 and F4 followed by the right number of continuation bytes (80 to BF), no over-long forms, no surrogate values. If they do, the run is replaced by the character it encodes. If they don’t, nothing changes. That test is also what keeps correctly written text safe. The é in a healthy “café” is a single byte in Windows-1252, and a single byte can never pass for a UTF-8 pair.
Text that was damaged, saved and damaged again shows longer chains such as é. The tool repeats the pass up to five times, peeling off one layer each time, until the text stops changing.
Reading the three highlights
- Repaired
- The bytes decoded cleanly, so these changes are safe. Point at the word to see what it replaced.
- Best guess, verify
- One byte was missing and the tool filled the gap. There are two cases.
Ãfollowed by a space is read asà, because the second byte of that letter (A0) is also the non-breaking space and often gets flattened to an ordinary one. A loneâ€is read as a closing curly quote, because its third byte (9D) has no character in Windows-1252 and tends to vanish. The tool only makes these guesses when the text shows other clear signs of mojibake. - Gone, type it again
- A
�, or a question mark stuck between two letters as inr?sum?. Some program replaced the character before saving, and there’s no byte left to decode. The tool marks the spot and you supply the letter.
A stray  in front of a space is the first half of a non-breaking space. The tool removes it and shows the space that remains with a faint mark, so you can tell it from a normal one.
Stop it where it starts
Repairing pasted text only treats the symptom. Somewhere upstream, one program is writing UTF-8 and another hasn’t been told. These are the usual places.
- A CSV opened in Excel
- Double-click a CSV and Excel assumes the old Windows table. Go through Data, From Text/CSV instead, and set File origin to
65001: Unicode (UTF-8). When exporting, pick “CSV UTF-8 (Comma delimited)”. - A MySQL or MariaDB database
- Columns declared
latin1that hold UTF-8 bytes, or a dump imported with the wrong--default-character-set. Tables and the connection should both useutf8mb4. - A web page
- Put
<meta charset="utf-8">first in the<head>and have the server sendContent-Type: text/html; charset=utf-8. - An email or a text file
- For mail, the header
Content-Type: text/plain; charset=UTF-8. For files, save as UTF-8 from the editor’s “Save with encoding” menu.
Converting a whole database isn’t the same job as fixing a paragraph. Take a full backup, run the conversion on a copy and compare rows before you touch the live site. A conversion applied to text that was already correct damages it.
What it can’t recover
The tool handles one accident, which happens to be by far the most common: UTF-8 read as Windows-1252 or Latin-1. Any language is covered, Greek and Cyrillic included, as long as that was the mix-up. Other pairs are out of scope, such as Mac Roman (where é appears as √©), Windows-1251 and Shift-JIS. The guess that turns à plus a space into à suits Italian, Portuguese and other languages that use that letter as a word. In any other text, check it.
Questions people ask
Why does é show up as é?
The text was saved as UTF-8, where é is two bytes, then opened as Windows-1252, where every byte is its own character. The two bytes get displayed as à and ©. Nothing is lost at that stage, which is why it can be reversed.
What does ’ mean in my text?
It’s a curly apostrophe (’). That sign takes three bytes in UTF-8, and each one got displayed separately. The same accident turns opening quotes into “ and long dashes into —.
Why is there a  before spaces or before £?
The non-breaking space is the two bytes C2 A0 and the pound sign is C2 A3. Read with the wrong table, the first byte turns into a visible Â. The tool removes it and keeps the character that follows.
Can it fix question marks and black diamonds (�)?
No. Those appear when a program couldn’t represent a character and replaced it before saving, so the original bytes are gone. The tool highlights each one and you retype it.
How do I fix garbled characters across a whole WordPress site?
Check that DB_CHARSET in wp-config.php is utf8mb4 and that the tables don’t use a latin1 collation. If it’s only a few pages, paste each text here and put the repaired version back. For the whole database, have the tables converted on a copy first, with a backup in hand.
Is the text I paste sent to a server?
No. A script in your browser does the decoding, and the page makes no request with your text.