Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

UTF-8 decoding converts a sequence of bytes into Unicode text; UTF-8 encoding performs the reverse operation. Valid UTF-8 uses one to four bytes for each Unicode scalar value. In JavaScript, use TextDecoder for bytes-to-text and TextEncoder for text-to-bytes. When data is malformed, choose replacement characters (U+FFFD) or a fatal error instead of silently accepting invalid sequences.

What UTF-8 encoding and decoding mean

Text in a program is represented as Unicode scalar values. UTF-8 is the byte encoding used to store or transmit those values; it is not a second character set. An encoder maps scalar values to bytes, and a decoder maps valid byte sequences back to scalar values, as specified by the WHATWG Encoding Standard.

ASCII remains unchanged: the character A is byte 0x41. Characters outside ASCII use multiple bytes. UTF-8 permits one to four bytes for values from U+0000 through U+10FFFF, excluding the UTF-16 surrogate range U+D800–U+DFFF. RFC 3629 defines the legal sequence forms and rejects overlong encodings and surrogate values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Unicode value Example UTF-8 bytes Length
U+0041 A 41 1 byte
U+00E9 é C3 A9 2 bytes
U+4E2D 中 E4 B8 AD 3 bytes
U+1F600 😀 F0 9F 98 80 4 bytes

The first byte indicates the sequence length. Following bytes must be continuation bytes in the range 80–BF, and additional range rules prevent values above U+10FFFF, overlong forms, and encoded surrogates.

How do I decode UTF-8 text?

You need the original bytes, not a string that has already been misinterpreted using the wrong character set. Decode those bytes once with a UTF-8 decoder, then use the resulting Unicode text.

JavaScript in browsers and modern runtimes

const bytes = new Uint8Array([0xE2, 0x82, 0xAC]); // €
const decoder = new TextDecoder("utf-8");
const text = decoder.decode(bytes);
console.log(text); // €

TextDecoder accepts a Uint8Array, ArrayBuffer view, or similar byte source and returns a JavaScript string. The utf-8 label is case-insensitive; using it explicitly documents the intended encoding.

Reject malformed input with fatal mode

const decoder = new TextDecoder("utf-8", { fatal: true });
try {
  const text = decoder.decode(new Uint8Array([0xE2, 0x28, 0xA1]));
  console.log(text);
} catch (error) {
  console.error("Invalid UTF-8", error);
}

Without fatal: true, the WHATWG algorithm uses replacement behavior and inserts U+FFFD (displayed as �) for decoding errors. Fatal mode reports failure instead. Not every wrapper exposes both policies, so check the API you use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decode a file in Node.js

import { readFile } from "node:fs/promises";

const bytes = await readFile("message.bin");
const text = new TextDecoder("utf-8", { fatal: true }).decode(bytes);
console.log(text);

If the file is known to be text, Node.js also provides readFile("file.txt", "utf8"). Use the byte-oriented form when you need to validate errors, inspect a BOM, or distinguish UTF-8 from another encoding before decoding.

Decode in Python

data = bytes([0xE2, 0x82, 0xAC])
print(data.decode("utf-8"))       # €

try:
    print(b"xe2(xa1".decode("utf-8"))
except UnicodeDecodeError as exc:
    print("Invalid UTF-8:", exc)

Python’s default errors="strict" raises UnicodeDecodeError. Choose errors="replace" only when losing or marking damaged bytes is acceptable.

How do I encode text as UTF-8?

JavaScript

const text = "Café 😀";
const bytes = new TextEncoder().encode(text);
console.log([...bytes]);
// [67, 97, 102, 195, 169, 32, 240, 159, 152, 128]

TextEncoder always emits UTF-8. To save the result in a browser, create a Blob; to send it with fetch, pass the Uint8Array as the request body.

Python

text = "Café 😀"
data = text.encode("utf-8")
print(data)

with open("message.txt", "wb") as file:
    file.write(data)

Command line with cURL

cURL does not guess the encoding of arbitrary input. Create UTF-8 bytes first, then upload them while declaring the media type:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
printf 'Café 😀n' | curl -X POST 
  -H 'Content-Type: text/plain; charset=utf-8' 
  --data-binary @- https://example.com/endpoint

The server must also interpret the declared charset correctly. A header cannot repair bytes that were already encoded incorrectly.

Streaming and split multi-byte sequences

Network responses and files often arrive in chunks. A multi-byte character may be split between chunks, so decoding each chunk independently can produce replacement characters. Use a streaming decoder that retains incomplete sequences.

const decoder = new TextDecoder("utf-8");
let output = "";
output += decoder.decode(firstChunk, { stream: true });
output += decoder.decode(secondChunk, { stream: true });
output += decoder.decode(); // flush buffered bytes and finish

In a browser stream, call decode with stream: true for every non-final chunk and call it once without that option at the end. If the final flush encounters an incomplete sequence, replacement or fatal behavior applies according to the decoder configuration.

What is the UTF-8 BOM?

The UTF-8 byte-order mark (BOM) is the three-byte signature EF BB BF, representing U+FEFF at the start of a byte stream. UTF-8 has no big-endian or little-endian byte order, so this mark is not needed to choose byte order. The Unicode Consortium FAQ describes it as an encoding signature.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The WHATWG standard’s normal UTF-8 decode operation consumes an initial BOM. The decode-without-BOM operation preserves it, so behavior depends on the specific API or option. A preserved U+FEFF can appear as an invisible leading character and may break formats that require the first bytes to be an ASCII token, such as a Unix shebang.

const bytes = new Uint8Array([0xEF, 0xBB, 0xBF, 0x41]);
console.log(new TextDecoder("utf-8").decode(bytes)); // A

When a parser rejects a file before its first visible character, inspect the first three bytes for this signature and verify whether the parser expects it.

Why does UTF-8 show �?

The replacement glyph means a decoder encountered malformed input while using replacement behavior. Common causes include:

  • The byte sequence was truncated in transit or at a chunk boundary.
  • Bytes were produced in Windows-1252, ISO-8859-1, UTF-16, or another encoding but labeled UTF-8.
  • An invalid continuation byte, overlong form, or surrogate encoding was supplied.
  • Data was decoded once and then incorrectly decoded a second time.

Switching decoders cannot recover information that was never present. Obtain the original bytes, confirm the producer’s encoding, and decode exactly once. Use fatal mode or a strict equivalent while diagnosing corruption so the error location is not hidden by U+FFFD.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validity, security, and interoperability

Do not accept arbitrary unknown bytes as UTF-8 merely because they contain some ASCII. RFC 3629 warns that naive acceptance of invalid sequences, including overlong encodings, can create security problems when one component rejects bytes but another reinterprets them. Validate at the boundary, reject prohibited sequences, and keep the same decoding rules across proxies, application code, logs, and databases.

For new protocols and formats, use the utf-8 label and specify UTF-8 explicitly. The WHATWG standard requires UTF-8 for new formats and identifies it as the appropriate interchange encoding for Unicode. Legacy formats may still require another encoding; determine that from the protocol or file specification rather than guessing.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting checklist

Text is garbled but no error appears

  • Capture the raw bytes before any conversion.
  • Check the HTTP Content-Type charset, file metadata, or protocol specification.
  • Compare a known ASCII prefix and a non-ASCII character in hexadecimal.
  • Decode once with a strict UTF-8 setting; if it fails, test the documented source encoding instead of forcing UTF-8.

A character is broken only when data is streamed

Keep decoder state between chunks using streaming mode. Do not call a fresh decoder for every network read, and always flush the final chunk.

The first field or command is rejected

Inspect for EF BB BF. Confirm whether your decoder consumes the BOM and whether the receiving format permits a leading U+FEFF.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python raises UnicodeDecodeError

This is strict error reporting, not proof that the file is useless. Find the byte offset from the exception, verify the producer’s encoding, and choose replacement handling only if your application can tolerate marked data loss.

JavaScript displays replacement characters

Use new TextDecoder("utf-8", { fatal: true }) during diagnosis, log the bytes, and check for truncation or a mislabeled legacy encoding. Do not pass a JavaScript string to a byte decoder and expect the original bytes to reappear.

Or skip the browser setup

If your goal is to capture a web page containing decoded text rather than build a browser pipeline, ScreenshotNeo returns a screenshot or PDF from one request. It accepts cookie banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Use the API as documented at ScreenshotNeo docs:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

There is a free allowance of 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Is UTF-8 the same as Unicode?

No. Unicode defines scalar values and characters; UTF-8 is one byte encoding of those values.

Can every four-byte sequence be UTF-8?

No. It must satisfy continuation-byte and range rules, and values above U+10FFFF or in the surrogate range are invalid.

Should I always remove a BOM?

No. Follow the file or protocol specification and the behavior expected by your parser. A BOM can be useful as a signature but is not required for UTF-8 byte order.

Frequently Asked Questions

Can UTF-8 represent emoji?

Yes. Emoji such as 😀 are encoded as four UTF-8 bytes when their Unicode scalar value is valid.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does UTF-8 decoding change the original bytes?

A successful decode maps bytes to text; encoding that text again should reproduce the same canonical UTF-8 bytes, while malformed input requires an explicit error or replacement policy.

The Bottom Line

Use UTF-8 explicitly, preserve decoder state for streams, validate malformed input, and treat BOM handling as API-specific. Strict decoding is the fastest way to discover whether corruption comes from the bytes or from an incorrect encoding assumption.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.