UTF-8 Converter: Encode, Decode, Convert to ASCII

See the UTF-8 bytes behind any text, decode a list of UTF-8 bytes back into characters, transliterate Unicode text to plain ASCII, or repair garbled “é”-style text. Everything runs in your browser.

What UTF-8 is

UTF-8 stores every Unicode character as one to four bytes. The 128 ASCII characters keep their single byte (A = 41), while everything else becomes a lead byte followed by continuation bytes that start with the bits 10: é is C3 A9, € is E2 82 AC, and an emoji such as 😀 is F0 9F 98 80. It is the encoding used by almost every web page, JSON document and modern file.

Decoding UTF-8 bytes

Paste UTF-8 bytes as hex pairs (C3 A9 or c3a9), percent-encoding (%C3%A9), \xC3\xA9 escapes, 0x-prefixed values, 8-bit binary or decimal numbers such as 72 105, and the text they encode is shown. The format is detected automatically, or you can force it. By default a sequence that is not valid UTF-8 is rejected rather than guessed; tick Replace invalid sequences to decode leniently, the way a browser does, with � marking each bad byte.

InputDetected asText
C3 A9hexé
%E2%82%ACpercent-encoded€
\xF0\x9F\x98\x80escapes😀
72 105 32 226 130 172decimalHi €

Encoding text to UTF-8

Choose UTF-8 encode to see the bytes behind your text in the format you need: spaced or compact hex, decimal, binary, %XX for URLs, \xNN for C, Python or shell strings, or 0x values for an array literal. A table lists every character with its code point and bytes, which makes it easy to spot a hidden non-breaking space or a lookalike letter.

UTF-8 to ASCII

Some old systems, file names and payment or SMS gateways only accept 7-bit ASCII. UTF-8 to ASCII transliterates instead of just deleting: accented letters lose their accents, ligatures and special letters are spelled out (ß → ss, æ → ae), typographic quotes, dashes and ellipses become their ASCII equivalents and symbols such as € become EUR. Characters with no ASCII equivalent become ? (or are removed, or kept) and are counted. Café “déjà vu” becomes Cafe "deja vu".

Fix garbled text (mojibake)

When UTF-8 bytes are read as Windows-1252 or Latin-1, every accented letter turns into two or three junk characters: Café shows up as Café and a dash as –. Fix garbled text reverses that mistake by turning the characters back into their original bytes and decoding them as UTF-8.

Common questions

How do I decode UTF-8 bytes to text?

Paste the bytes with UTF-8 decode selected. Hex, percent-encoded, \x, 0x, binary and decimal bytes are recognised automatically; E2 82 AC decodes to €.

How do I see the UTF-8 bytes of a character?

Choose UTF-8 encode and type the text. The bytes appear in the format you pick, and a table shows each character’s code point and bytes; é is C3 A9.

How do I convert UTF-8 text to ASCII?

Choose UTF-8 to ASCII. Accents are removed, special letters and symbols are spelled out, smart quotes and dashes are straightened, and anything left over becomes ?.

Why do I get an invalid UTF-8 error?

The bytes do not form a valid UTF-8 sequence — for example a lone E9, which is Latin-1 é rather than UTF-8. Check the source encoding, or tick Replace invalid sequences to decode anyway.

How do I fix text that shows é instead of é?

Paste it into Fix garbled text. It re-reads the characters as the original bytes and decodes them as UTF-8, so Café becomes Café.

Is my text uploaded?

No. Encoding and decoding use the browser’s built-in TextEncoder and TextDecoder; nothing leaves the page.