What UTF-8 is
UTF-8 stores every Unicode character as one to four bytes. The 128 ASCII characters keep their single byte (A = 41), while everything else becomes a lead byte followed by continuation bytes that start with the bits 10: é is C3 A9, € is E2 82 AC, and an emoji such as 😀 is F0 9F 98 80. It is the encoding used by almost every web page, JSON document and modern file.
Decoding UTF-8 bytes
Paste UTF-8 bytes as hex pairs (C3 A9 or c3a9), percent-encoding (%C3%A9), \xC3\xA9 escapes, 0x-prefixed values, 8-bit binary or decimal numbers such as 72 105, and the text they encode is shown. The format is detected automatically, or you can force it. By default a sequence that is not valid UTF-8 is rejected rather than guessed; tick Replace invalid sequences to decode leniently, the way a browser does, with � marking each bad byte.
| Input | Detected as | Text |
|---|---|---|
C3 A9 | hex | é |
%E2%82%AC | percent-encoded | € |
\xF0\x9F\x98\x80 | escapes | 😀 |
72 105 32 226 130 172 | decimal | Hi € |
Encoding text to UTF-8
Choose UTF-8 encode to see the bytes behind your text in the format you need: spaced or compact hex, decimal, binary, %XX for URLs, \xNN for C, Python or shell strings, or 0x values for an array literal. A table lists every character with its code point and bytes, which makes it easy to spot a hidden non-breaking space or a lookalike letter.
UTF-8 to ASCII
Some old systems, file names and payment or SMS gateways only accept 7-bit ASCII. UTF-8 to ASCII transliterates instead of just deleting: accented letters lose their accents, ligatures and special letters are spelled out (ß → ss, æ → ae), typographic quotes, dashes and ellipses become their ASCII equivalents and symbols such as € become EUR. Characters with no ASCII equivalent become ? (or are removed, or kept) and are counted. Café “déjà vu” becomes Cafe "deja vu".
Fix garbled text (mojibake)
When UTF-8 bytes are read as Windows-1252 or Latin-1, every accented letter turns into two or three junk characters: Café shows up as Café and a dash as –. Fix garbled text reverses that mistake by turning the characters back into their original bytes and decoding them as UTF-8.
Common questions
How do I decode UTF-8 bytes to text?
Paste the bytes with UTF-8 decode selected. Hex, percent-encoded, \x, 0x, binary and decimal bytes are recognised automatically; E2 82 AC decodes to €.
How do I see the UTF-8 bytes of a character?
Choose UTF-8 encode and type the text. The bytes appear in the format you pick, and a table shows each character’s code point and bytes; é is C3 A9.
How do I convert UTF-8 text to ASCII?
Choose UTF-8 to ASCII. Accents are removed, special letters and symbols are spelled out, smart quotes and dashes are straightened, and anything left over becomes ?.
Why do I get an invalid UTF-8 error?
The bytes do not form a valid UTF-8 sequence — for example a lone E9, which is Latin-1 é rather than UTF-8. Check the source encoding, or tick Replace invalid sequences to decode anyway.
How do I fix text that shows é instead of é?
Paste it into Fix garbled text. It re-reads the characters as the original bytes and decodes them as UTF-8, so Café becomes Café.
Is my text uploaded?
No. Encoding and decoding use the browser’s built-in TextEncoder and TextDecoder; nothing leaves the page.