Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk3 min

How to Count Unicode Characters, Emojis, and Words Correctly in JavaScript

JavaScript String.length counts UTF-16 code units. Choose code-point iteration, grapheme segmentation, word segmentation, or byte measurement to match the task.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JavaScript’s String.length counts UTF-16 code units—not necessarily Unicode code points, visible characters, or words. Use it for JavaScript’s string-indexing unit; use code-point iteration, Intl.Segmenter, or a separate byte measurement when your task calls for a different unit.

What does String.length count?

JavaScript strings are represented as sequences of UTF-16 code units. The length property returns the number of those units. A Unicode code point in the Basic Multilingual Plane typically uses one code unit; a supplementary code point uses a pair called a surrogate pair, so one code point can contribute two to length. MDN explains this distinction in its documentation for String.length.

As an Amazon Associate I earn from qualifying purchases.

"A".length;  // 1
"😄".length; // 2

That result is not a JavaScript error: it reflects the string’s UTF-16 representation. It simply may not answer a question about how many code points or user-perceived characters the text contains.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which counting method should you use?

Need Method What it counts
JavaScript string indexing or UTF-16 units text.length UTF-16 code units; a supplementary code point counts as two.
Unicode code points [...text].length Code points, keeping a valid surrogate pair together, but not combining marks or multi-code-point emoji sequences.
Approximate user-perceived characters Intl.Segmenter with granularity: "grapheme" Grapheme clusters, which can keep combining marks and joined emoji sequences together.
Words in a locale-sensitive text Intl.Segmenter with granularity: "word" Segments marked as word-like by the segmenter.
Encoded storage or transmission size Measure bytes using the required encoding Bytes, not characters or grapheme clusters.

How do you count Unicode code points?

String iteration treats a valid surrogate pair as one code point, so spreading a string into an array and reading its length gives a code-point count:

const codePointCount = (text) => [...text].length;

This is useful when the requirement specifically concerns code points. It does not count user-perceived characters in every case. For example, a letter followed by a combining accent can contain two code points while appearing as one character. Likewise, many emoji—including joined sequences and flags—are formed from multiple code points.

MDN’s JavaScript String guide discusses code points, surrogate pairs, and emoji sequences. For default rules defining grapheme, word, and sentence boundaries, see the Unicode Consortium’s Unicode Text Segmentation, UAX #29, Unicode 18.0.0, Version 49 (2026-09-01).

How do you count user-perceived characters, including emoji?

Use Intl.Segmenter with grapheme granularity when a limit or display reports approximate user-perceived characters. A grapheme cluster may consist of several code points, such as a base letter with a combining mark or an emoji joined to another emoji.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const graphemeSegmenter = new Intl.Segmenter("en", {
  granularity: "grapheme",
});

const graphemeCount = (text) =>
  [...graphemeSegmenter.segment(text)].length;

Choose a locale appropriate to the content or application, and check that the target JavaScript runtime supports the needed Intl.Segmenter behavior. MDN describes grapheme-level segmentation with Intl.Segmenter as useful for counting characters. Grapheme clusters are a practical approximation of user-perceived characters, not a measure of rendered width: fonts and layout can make clusters occupy different amounts of space.

How do you count words reliably?

Splitting on whitespace is a quick approximation for text that uses spaces between words, but it is not a general word-counting rule. Punctuation and writing systems that do not separate words with spaces can make whitespace splitting misleading. Use word segmentation and count only segments whose isWordLike property is true:

const wordSegmenter = new Intl.Segmenter("en", {
  granularity: "word",
});

const wordCount = (text) =>
  [...wordSegmenter.segment(text)]
    .filter((part) => part.isWordLike).length;

Set the locale to suit the text or application rather than assuming one segmentation choice is correct for every language. The MDN internationalization guide explains word segmentation and why whitespace splitting has limits.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What if you need a byte limit or display width?

Neither text.length nor a grapheme count gives the number of bytes in an encoded string. If a database, protocol, or API imposes a byte limit, measure the text in the encoding that limit specifies. Similarly, grapheme count does not calculate visual width; rendered width depends on presentation as well as the text’s segmentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.