UTF-8 Hex to Text Converter

UTF-8 Hex to Text Converter. Paste your bytes into UTF-8 Hex Input and the UTF-8 Hex to Text Converter instantly decodes your Text Output. For decoding character codes captured from a programming exercise or API response, the decimal to text converter reconstructs the original message.

Paste a string of hex bytes into this UTF-8 hex to text converter and it hands back the readable text those bytes encode, whether that is plain English, accented letters or an emoji. It is a free online tool that works in your browser, so you can paste your hex data from a log, a database dump or a debugging session and see the decoded text with no registration. Because UTF-8 stores every Unicode character in one to four bytes, a trustworthy decoder has to read each multi-byte sequence correctly, and the sections below show you exactly how that works so you can verify any result by hand.

What Is a UTF-8 Hex to Text Converter?

A hex to text converter takes numbers written in base 16 and turns them into the characters they stand for. When the characters follow the UTF-8 standard, you need a UTF-8 aware decoder rather than a simple lookup, because one character can span several bytes. This page's hex to utf8 converter does that work for any text you throw at it: English, Greek, Hindi, Chinese, symbols and emoji. Use the hex to octal converter when a legacy Unix tool expects octal but your data is stored in hexadecimal.

Understanding Hexadecimal and Base-16 Values

Hexadecimal uses sixteen digits, 0-9 and A-F, where A is 10 and F is 15. Each hex digit holds exactly four binary bits, so two digits describe one byte, which is why programmers prefer base-16 notation over long strings of ones and zeros. The byte E2 is simply 226 in decimal and 11100010 in binary. Hex values like these are the everyday shorthand for raw data in memory dumps, network captures and file editors, and the same notation is what you see whenever an encoding problem forces you to look at bytes.

UTF-8 Encoding, Unicode and Code Points

Unicode assigns every character a number called a code point, and UTF-8 encoding is the rulebook that turns a code point into one to four bytes. It is a variable-width character encoding, so common letters stay short while rarer symbols take more room. The first 128 code points match ASCII exactly, which makes UTF-8 backward-compatible with ASCII and explains why plain English hex often decodes the same under either scheme.

Hexadecimal to UTF8 Conversion Step by Step

You do not need to understand the byte math to use the tool, but knowing the input rules saves time when a conversion looks wrong. Three layouts are widely used, and a good decoder accepts all of them, even when they are combined in one box.

Accepted Input Formats

  • Space-separated values: pairs split by single spaces, such as 4A 6F 73 C3 A9.
  • 0x prefix: each byte starts with 0x, the style copied from C, Java and JavaScript source files.
  • Continuous string: one unbroken run of digits like 4A6F73C3A9, common in database exports and URLs.
  • Encoding choice: the converter assumes UTF-8 encoding; data written in another encoding needs that scheme selected before decoding.
  • Mixed formats: any combination of the above, with line breaks, commas or tabs between bytes.

Using the Hex to UTF-8 Converter

  1. Paste your hex data into the input box, using any prefixed format or plain pairs.
  2. Press the convert button, or watch the output update as you type.
  3. Read the decoded text in the result panel and copy it where you need it.
  4. If a replacement symbol appears, check for a missing digit or a byte that was cut off.

The tool strips whitespace and prefixes first, joins the digits into bytes, and then decodes the byte sequence as UTF-8. Odd-length input means a digit is missing, so that is the first thing to check. Large hex strings and several lines at once are fine, which makes batch processing of a whole export a single paste.

How UTF-8 Encoding and Decoding Work Byte by Byte

Every UTF-8 byte announces its role in its leading bits, so a decoder never has to guess where one character ends and the next begins. That self-describing design is the reason the format took over the web and is the thing worth learning if you ever debug raw bytes.

Leading Bits and Continuation Bytes

A single byte beginning with 0 is an ASCII character. A byte beginning with 110 starts a sequence of two bytes, 1110 starts a sequence of three bytes, and 11110 starts a sequence of four bytes, so a character takes anywhere from 1-4 bytes. Every continuation byte that follows carries the binary prefix 10, which is how a decoder recognises that it is still inside the same character.

Leading Bits and Continuation Bytes
Code point rangeBytesByte pattern
U+0000 to U+007F10xxxxxxx
U+0080 to U+07FF2110xxxxx 10xxxxxx
U+0800 to U+FFFF31110xxxx 10xxxxxx 10xxxxxx
U+10000 to U+10FFFF411110xxx 10xxxxxx 10xxxxxx 10xxxxxx

Rebuilding a Code Point From a Byte Sequence

To decode a three-byte character, strip the prefix bits and join what is left. With the lead byte written as \(b_1\) and the two continuation bytes as \(b_2\) and \(b_3\), the code point is:

$$\text{code point} = (b_1 \mathbin{\&} \texttt{0x0F}) \times 2^{12} + (b_2 \mathbin{\&} \texttt{0x3F}) \times 2^{6} + (b_3 \mathbin{\&} \texttt{0x3F})$$

Two-byte characters use \((b_1 \mathbin{\&} \texttt{0x1F}) \times 2^{6} + (b_2 \mathbin{\&} \texttt{0x3F})\), and four-byte characters add a fourth term with \(2^{18}\) for the lead byte. The result is the code point, which the decoder then maps to a glyph.

Waterfall chart showing hex bytes E2, 9C and 93 adding up to code point 10,003, the check mark U+2713, in UTF-8 decoding
How three hex bytes rebuild one Unicode code point.

UTF-8 Decoding Worked Example: Santiago ✓ 🚀

Take the 17 hex bytes 53 61 6E 74 69 61 67 6F 20 E2 9C 93 20 F0 9F 9A 80. The first nine bytes (53 through 20) are all below 80, so each is one ASCII character and together they spell "Santiago ". The next byte, E2, is 11100010, so it opens a three-byte character:

\((\texttt{E2} \mathbin{\&} \texttt{0F}) \times 4096 = 8192\), \((\texttt{9C} \mathbin{\&} \texttt{3F}) \times 64 = 1792\) and \(\texttt{93} \mathbin{\&} \texttt{3F} = 19\), which add up to 10003, or U+2713, the check mark ✓. After another space (20), F0 is 11110000, a four-byte lead, and the bytes F0 9F 9A 80 resolve to U+1F680, the rocket 🚀. The final text is Santiago ✓ 🚀: 12 characters stored in 17 bytes.

Stacked column chart comparing 12 decoded characters with the 17 UTF-8 bytes that encode them in the Santiago example
Twelve characters, seventeen bytes: the check mark and rocket take 7 of them.
UTF-8 Decoding Worked Example: Santiago ✓ 🚀
CharacterCode pointUTF-8 hex bytesBytes
ñU+00F1C3 B12
ΩU+03A9CE A92
₹U+20B9E2 82 B93
✓U+2713E2 9C 933
🚀U+1F680F0 9F 9A 804

Hex to ASCII Converter vs UTF-8 Decoder

A hex to ascii converter maps each byte straight to one of 128 characters, so it breaks the moment a byte is 80 or higher. A utf-8 decoder reads the same bytes but groups them by their leading bits first. For pure English the output is identical, for example the classic ASCII sample 48 65 6C 6C 6F, but with international characters and accented characters only the UTF-8 route gives correct text. Use hex to ascii when you know the data is limited to the basic set, and a UTF-8 decoder for everything else.

Latin-1, UTF-16 and Other Encodings

The same bytes mean different things in different schemes. Read C3 A9 as Latin-1 and you get two odd symbols; read it as UTF-8 and you get a single é. UTF-16 and UTF-32 use fixed 2- or 4-byte units, so their hex dumps are longer and often contain many 00 bytes. If your result is full of nulls or boxes, you are probably decoding a different character encoding than the one that produced the data. Match the encoding of the source first, and the right decoding follows; an encoding mismatch causes more garbled output than corrupt bytes ever do. Chinese, Japanese and Korean (CJK) text is a good test, since those characters usually need three bytes each and fail loudly under the wrong encoding.

Dumbbell chart comparing bytes per character in UTF-8 and UTF-16 for a letter, accented letters, a check mark and an emoji
UTF-8 is smaller for ASCII letters; UTF-16 can be smaller for symbols.

Related Conversions: Binary, Decimal and Base64

  • Binary: expand each hex digit to four bits to see the leading bits of a UTF-8 byte.
  • Decimal: code often prints raw byte arrays as 0-255 values; convert each to a two-digit hex pair before pasting them here.
  • Base64: a text-safe wrapper around the same bytes; unwrap it to bytes first, then paste the result as hex into this converter.
  • Code points: the U+ numbers Unicode assigns, the step between the pasted hex and the characters you see.
  • UTF8 to hex: the reverse direction, handled by a UTF-8 to hex converter whose encoding step turns text into bytes.
  • Percent-encoding: the %XX escapes in URLs are the same hex bytes, so URL decoding is UTF-8 decoding with a different wrapper.

Tracing a Truncated Gift Note Through UTF-8 Decoding

A support engineer is chasing a ticket: a customer typed "Zażółć 🎁" into the gift note field of an order, but the confirmation email shows only "Zażółć". Suspecting a charset problem, the engineer runs SELECT HEX(gift_note) on the orders table and pastes the result into the converter:

5A 61 C5 BC C3 B3 C5 82 C4 87 20

That is 11 bytes, and decoding them as UTF-8 returns Zażółć with a trailing space, 7 characters. Reading the bytes explains why the accents survived: C5 BC is ż (U+017C), C3 B3 is ó (U+00F3), C5 82 is ł (U+0142) and C4 87 is ć (U+0107). Each is a two-byte sequence, comfortably inside what the column can store.

Next comes the hex captured in the application log for the same request, which is 15 bytes long: the same 11 bytes followed by F0 9F 8E 81. The lead byte F0 is 11110000, so those four bytes form a single character, U+1F381, the gift emoji. The database stopped right before it.

Tracing a Truncated Gift Note Through UTF-8 Decoding
SourceBytesCharactersLast character
Application log158🎁 (U+1F381)
HEX(gift_note) from the table117space (U+0020)

MySQL's legacy utf8 charset holds at most three bytes per character, and the check against that limit is where a four-byte emoji gets cut. That is the named cause, so the next action is specific: run ALTER TABLE orders CONVERT TO CHARACTER SET utf8mb4 on a staging copy, resubmit the same note, and confirm that HEX(gift_note) now returns all 15 bytes. Once it does, decoding that dump shows the full "Zażółć 🎁" and the ticket can close.

Fixing Malformed Hex and Invalid Characters

Most failed decodes come from the input, not the tool. Good error handling flags the problem instead of guessing, and reliable UTF-8 decoding never swallows a bad byte silently, and the patterns below cover nearly every case you will hit.

Invalid Characters and Odd-Length Input

Anything outside 0-9 and A-F, apart from separators and prefixes, is an invalid character in hex input. A stray letter G, a pasted smart quote or an odd count of digits all stop the conversion. Remove them and retry. If the hex is valid but the output shows replacement symbols, the bytes themselves are malformed UTF-8, usually because a multi-byte sequence was truncated.

Overlong Sequences and the Byte Order Mark

  • Overlong encodings: a character stored in more bytes than needed, such as C0 80 for null. Strict decoders reject these because they were once used to slip past security filters.
  • BOM: the byte order mark EF BB BF at the start of a file is legal but unnecessary, and it can show up as a stray symbol before your text.
  • Invalid lead bytes: values from C0, C1 and F5 to FF never appear in valid UTF-8.
  • Truncated tails: a lead byte with too few continuation bytes after it means the data was cut off.

Where a Hex to Text Converter Earns Its Place

Developers, analysts and administrators reach for a hex to text converter more often than you might expect. These are the situations where it saves real time.

Web Development and API Responses

In web development, hex-escaped strings show up in API responses, JSON and XML payloads and HTML entities. Decoding the bytes shows whether a mangled character was already broken when it left the server or was damaged by the browser, and UTF-8 decoding at each hop tells you which hop introduced the damage. It also supports internationalization work, since you can confirm that text in many languages survives the trip.

Database and Storage Debugging

Garbled text in databases is a classic encoding bug. When MySQL rejects an emoji, converting the stored value to hex reveals whether it is a four-byte sequence that a three-byte charset cannot hold. The fix is the utf8mb4 character set. Checking the hex of a column before and after a migration is the quickest way to prove nothing changed in storage.

Forensics, Logs and System Administration

Forensics analysts and system administration teams decode hex dumps, network packets and log files to read the text hiding in binary captures. This is also routine data analysis when a vendor hands over a raw export with no readable fields. Because every byte is shown, nothing is silently reinterpreted along the way, and the encoding of each field stays visible. A mislabeled encoding in a log header is one of the most common surprises, and the hex view settles it. Strict UTF-8 decoding of the raw bytes then proves what the text really says.

UTF-8 Decoding in Python, JavaScript and PHP

For programming use, the same job takes a line or two of code. In Python, bytes.fromhex(...) followed by .decode() turns hex into a string, and the decode() call lets you choose how errors are handled. In JavaScript, the TextDecoder API decodes a Uint8Array in the browser, while Node.js uses the Buffer class, where Buffer.from(hex, 'hex') builds the bytes and a Buffer toString call decodes them. In PHP, hex2bin followed by mb_convert_encoding does it. These routines are handy to know, but a browser-based tool is faster for a one-off check.

Privacy and Security in Browser-Based Decoding

Hex dumps sometimes contain tokens, names or internal messages, so privacy matters. A tool that decodes locally in your browser never sends the bytes to a server, which means nothing is stored or logged. That makes it suitable for sensitive text during security reviews, and it works the same on a mobile device as on a desktop. Because UTF-8 is an open standard that every operating system, editor and piece of software supports, the result is also portable: decoded text pasted into another editor behaves the same everywhere, and the shared rules give developers real interoperability between systems.

Whenever you need to convert hex to utf8 and be sure of the answer, decode a short test string first, compare it with the byte table above, and then paste your real files. And when you want to go back, a utf8 to hex run reproduces the original bytes, which is the quickest way to confirm a round trip.