For the usual beginner definition—one word per whitespace-separated token—use len(text.split()). Python collapses runs of whitespace for this operation, so repeated spaces, tabs, and newlines do not create extra tokens. Choose a different method only if your application defines words differently.
Count whitespace-separated words
This is a practical default for ordinary prose and simple scripts:
text = "Python makes text processing approachable."
word_count = len(text.split())
print(word_count) # 5
Called without an argument, str.split() treats runs of whitespace as separators and omits empty strings at the beginning and end. Punctuation remains attached: in "Hello, world!", the two resulting tokens are "Hello," and "world!". See the Python documentation for str.split().
Choose what “word” means for your task
Python does not impose one universal definition of a word. The counting rule should match the needs of your program or the editorial standard you are following.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
| Rule | Code | What it counts |
|---|---|---|
| Whitespace-delimited tokens | len(text.split()) |
Tokens separated by whitespace; punctuation stays attached. |
| Runs of regex word characters | len(re.findall(r'w+', text)) |
Runs of Unicode alphanumeric characters and underscores, including numbers and names such as snake_case. |
| Parts separated by non-word characters | sum(bool(part) for part in re.split(r'W+', text)) |
Non-empty runs of characters that match Python’s w rule. |
Count runs of word characters
Use this when punctuation should separate matches rather than remain attached to whitespace-delimited tokens:
import re
text = "Python makes text-processing approachable."
word_count = len(re.findall(r"w+", text))
print(word_count) # 5
Python’s Unicode w includes Unicode alphanumeric characters and the underscore. This means numbers and identifiers such as snake_case count as matches. The behavior is defined by the regular-expression syntax documentation.
Rank #2
Split at non-word characters without counting empty results
re.split(r'W+', text) splits on runs of characters outside the w category. Because splitting can leave empty strings at the edges, count only non-empty parts:
import re
text = "Hello, world!"
parts = re.split(r"W+", text)
word_count = sum(bool(part) for part in parts)
print(word_count) # 2
This is not a language-aware punctuation rule. Apostrophes and hyphens are non-word characters, so a contraction or hyphenated expression can be divided into multiple matches; underscores, by contrast, remain part of a word-character run. Python’s b boundary is likewise defined by transitions between w and W, or a string edge—not by a universal linguistic definition.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWhat Unicode and language choice change
For Unicode string patterns, Python’s default s matches Unicode whitespace as defined by str.isspace(), not just the ordinary ASCII space, tab, and newline. The default regex shorthand classes are Unicode-aware for str; adding re.ASCII makes w, W, b, B, d, D, s, and S ASCII-only. Details are in the Python regex documentation.
Whitespace splitting is still only a chosen approximation for many editorial standards and languages. If the required count depends on how a language treats compounds, apostrophes, or scripts that do not conventionally separate words with spaces, specify that rule or use a tokenizer designed for the language. The Python rules above do not supply that external standard.
Quick Recap
Best Value
Avoid common counting mistakes
- Do not default to
text.split(" "). An explicit single-space separator does not mean “split on any run of whitespace.” Usetext.split()when whitespace is your separator rule. - Do not expect
split()to remove punctuation. It separates on whitespace only; punctuation remains in each token. - Do not count every result of
re.split(). It can return empty strings at the edges. Count non-empty results instead. - Do not treat a regex boundary as a universal word definition.
wandWfollow Python’s character categories, which may not match your publication’s counting convention.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




