Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For a quick tag removal, use PHP’s strip_tags(). It removes HTML and PHP tags but does not validate malformed markup and must not be treated as an XSS defense. When you need predictable text, block breaks, links, or other structure, parse the document with a DOM API instead. PHP 8.4 adds DomHTMLDocument::createFromString(), which follows HTML5 parsing rules; older code commonly uses DOMDocument::loadHTML(), whose parser follows HTML 4 rules.

Choose the conversion method first

Need Recommended approach Important limitation
Remove tags from a trusted, simple fragment strip_tags() Malformed tags can remove more text than expected; it does not protect against XSS.
Extract text while handling nesting, links, or block boundaries DOM parsing and a text-walk routine You must decide how paragraphs, lists, whitespace, and links should appear in the output.
Parse according to modern browser HTML5 rules DomHTMLDocument::createFromString() on PHP 8.4+ Requires PHP 8.4 or newer.
Support older PHP installations DOMDocument::loadHTML() Uses an HTML 4 parser, so its tree can differ from a browser’s HTML5 tree; it is not a sanitizer.

Quick conversion with strip_tags()

The simplest implementation is a direct string transformation:

<?php
$html = '<p>Hello <strong>world</strong>.</p>';
$text = strip_tags($html);

echo $text; // Hello world.

By default, every tag is removed. You can allow selected tags by passing a second argument:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<?php
$html = '<p>First</p><p>Second</p>';
$text = strip_tags($html, '<p>');

Allowing tags does not convert them to line breaks or otherwise make the result plain text; it leaves those tags in the returned string. For plain text output, remove all tags and apply your own formatting rules.

Whitespace and entities

strip_tags() does not define a universal plain-text layout. HTML such as adjacent paragraphs may become adjacent words unless the source contains whitespace. If the input contains entities such as &amp; or &nbsp;, decode them explicitly when that matches your requirements:

<?php
$html = '<p>Fish &amp; chips</p>';
$text = html_entity_decode(strip_tags($html), ENT_QUOTES | ENT_HTML5, 'UTF-8');
$text = preg_replace('/[ t]+/', ' ', $text);
$text = trim($text);

Be careful with non-breaking spaces: converting every whitespace character to an ordinary space may be desirable for search indexing but undesirable for typographic output. Decide whether to preserve, normalize, or remove them.

Why this is not a security filter

The PHP manual warns that strip_tags() should not be used to prevent XSS. It does not validate HTML, and malformed or partial tags can cause more content to be removed than you intended. If the resulting text is inserted into an HTML response, escape it for that output context, for example with htmlspecialchars($text, ENT_QUOTES | ENT_SUBSTITUTE, 'UTF-8'). If your actual goal is to retain safe HTML rather than produce text, use a dedicated HTML sanitizer instead of either conversion method.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Structured extraction with DOMDocument

A DOM parser is a better fit when you need to preserve meaningful boundaries. The following routine walks a parsed tree, inserts newlines around common block elements, turns list items into lines, and decodes character references through the DOM’s text nodes.

<?php
function htmlToPlainTextLegacy(string $html): string
{
    $dom = new DOMDocument('1.0', 'UTF-8');

    // Suppress parser diagnostics for fragments that are not full documents.
    $previous = libxml_use_internal_errors(true);
    $dom->loadHTML(
        '<meta charset="utf-8">' . $html,
        LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD
    );
    libxml_clear_errors();
    libxml_use_internal_errors($previous);

    $blockTags = [
        'address' => true, 'article' => true, 'aside' => true,
        'blockquote' => true, 'br' => true, 'div' => true,
        'dl' => true, 'fieldset' => true, 'figcaption' => true,
        'figure' => true, 'footer' => true, 'form' => true,
        'h1' => true, 'h2' => true, 'h3' => true, 'h4' => true,
        'h5' => true, 'h6' => true, 'header' => true, 'hr' => true,
        'li' => true, 'main' => true, 'nav' => true, 'ol' => true,
        'p' => true, 'pre' => true, 'section' => true, 'table' => true,
        'tr' => true, 'ul' => true
    ];

    $walk = function (DOMNode $node) use (&$walk, $blockTags): string {
        if ($node instanceof DOMText) {
            return $node->nodeValue;
        }
        if (!($node instanceof DOMElement)) {
            $out = '';
            foreach ($node->childNodes as $child) {
                $out .= $walk($child);
            }
            return $out;
        }

        $tag = strtolower($node->tagName);
        if (in_array($tag, ['script', 'style', 'noscript', 'template'], true)) {
            return '';
        }

        $out = '';
        foreach ($node->childNodes as $child) {
            $out .= $walk($child);
        }
        if (isset($blockTags[$tag])) {
            $out .= "n";
        }
        return $out;
    };

    $text = $walk($dom);
    $text = preg_replace("/[ tx0Bf]+/u", ' ', $text);
    $text = preg_replace("/ *n */u", "n", $text);
    $text = preg_replace("/n{3,}/u", "nn", $text);
    return trim($text);
}

echo htmlToPlainTextLegacy('<h1>Title</h1><p>One <em>small</em> paragraph.</p>');

This example is deliberately a policy, not a claim that PHP supplies one canonical HTML-to-text format. Add or remove block names to match your application. For example, you might preserve table cells with tabs, prefix unordered list items with a bullet, or keep the contents of pre untouched instead of collapsing its spaces.

What loadHTML() actually parses

DOMDocument::loadHTML() accepts strings that do not form a complete, well-formed document. However, PHP documents that it uses an HTML 4 parser: “The parsing rules of HTML 5, which are what modern web browsers use, are different.” Consequently, implied elements, malformed nesting, and some modern markup can produce a DOM tree unlike the one a browser builds. The function also cannot safely be used to sanitize HTML. Use it for compatibility when you understand that parsing difference, not as a security boundary.

HTML5 parsing in PHP 8.4 and later

PHP 8.4 introduced DomHTMLDocument::createFromString(). It creates a DomHTMLDocument using the HTML5 parsing model, making it the appropriate parser when browser-conforming handling matters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<?php
if (PHP_VERSION_ID < 80400) {
    throw new RuntimeException('This example requires PHP 8.4 or newer.');
}

$html = '<article><h2>News</h2><p>Updated &amp; ready.</p></article>';
$document = DomHTMLDocument::createFromString($html);

function html5Text(DomNode $node): string
{
    $skip = ['script', 'style', 'noscript', 'template'];
    $blocks = ['article', 'blockquote', 'br', 'div', 'h1', 'h2', 'h3',
               'h4', 'h5', 'h6', 'li', 'p', 'pre', 'section', 'tr'];

    if ($node instanceof DomText) {
        return $node->data;
    }
    if ($node instanceof DomElement && in_array(strtolower($node->localName), $skip, true)) {
        return '';
    }

    $out = '';
    foreach ($node->childNodes as $child) {
        $out .= html5Text($child);
    }
    if ($node instanceof DomElement && in_array(strtolower($node->localName), $blocks, true)) {
        $out .= "n";
    }
    return $out;
}

$text = html5Text($document);
$text = preg_replace('/[ t]+/u', ' ', $text);
$text = preg_replace('/ *n */u', "n", $text);
echo trim($text);

Keep the PHP-version check when this code is distributed across mixed environments. If you support PHP versions before 8.4, route those installations to the DOMDocument implementation and document the HTML 4 parsing behavior.

Formatting decisions you must make

Block boundaries

HTML’s visual layout is not automatically represented in a text string. Treat headings, paragraphs, list items, and line breaks as boundaries if readers must be able to scan the result. Avoid inserting a newline after every element: nested blocks can produce excessive blank lines, which is why the examples normalize runs after traversal.

Links

A DOM walk returns the anchor’s visible text, not its URL. If the plain-text consumer needs destinations, detect <a> elements and append a representation such as Label (https://example.test). Do this deliberately; blindly appending every href can expose tracking URLs or javascript schemes.

Images and alternative text

Images have no text node. For accessibility-oriented output, replace an img element with its alt value when present. Decorative images should usually be omitted. Do not assume that an image URL is useful plain text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Lists and tables

For lists, add a marker before each li if the distinction matters. For tables, choose a delimiter such as a tab or vertical bar and emit a row break after each tr. There is no single correct representation for every downstream format.

Whitespace and Unicode

Normalize only what your consumer permits. Search indexes often benefit from collapsing spaces and blank lines; legal, code, or preformatted content may require preservation. Always specify UTF-8 when converting entities or writing files, and use ENT_SUBSTITUTE when malformed byte sequences must not trigger warnings or produce invalid output.

Security and reliability checklist

  • Do not use either strip_tags() or loadHTML() as an XSS mitigation.
  • Escape the final text for the context in which it is rendered, especially when placing it back into HTML.
  • Reject or limit unexpectedly large inputs before parsing to control memory and CPU use.
  • Remove script, style, and other non-content elements explicitly when walking a DOM.
  • Test malformed nesting, missing closing tags, entities, Unicode, nested lists, tables, and empty elements.
  • Keep parser warnings out of user output by handling libxml errors as shown in the legacy example.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common results

Paragraphs run together

strip_tags() removes markup but does not invent separators. Use a DOM walk that adds newlines for p, headings, and list items, or preprocess known separators before stripping.

Text disappears around malformed markup

This is an expected risk of strip_tags(). Parse the input with a DOM API and inspect the resulting tree, or reject invalid input rather than relying on tag stripping.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The output differs from browser text

On PHP versions using DOMDocument::loadHTML(), the HTML 4 parser may build a different tree from a browser. Upgrade to PHP 8.4 and use DomHTMLDocument::createFromString() when HTML5 behavior is required.

Accented characters are corrupted

Ensure the source is UTF-8, specify UTF-8 in entity decoding, and provide a charset declaration when loading fragments with DOMDocument. Do not apply byte-oriented regular expressions to multibyte text.

Sanitized output is still unsafe

Conversion and sanitization are different jobs. Escape text at output, and use a purpose-built sanitizer when the requirement is safe, limited HTML rather than plain text.

Or skip the browser setup

If the HTML first comes from a live page and your real task is obtaining a clean capture rather than parsing a stored string, ScreenshotNeo provides a single HTTP request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server can let Claude, Cursor, or another MCP client call take_screenshot, get_page_info, or capture_pdf.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for the 63 capture options, including full-page lazy-image loading, CSS selectors, device and retina settings, custom JavaScript, cookies and headers, waiting rules, PDF output, caching, signed links, asynchronous jobs, bulk capture, and usage reporting. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Which implementation should you ship?

  • Use strip_tags() for a small, trusted fragment where exact layout is unimportant.
  • Use a DOM traversal when paragraph breaks, lists, links, images, or tables matter.
  • Use DomHTMLDocument::createFromString() on PHP 8.4+ when browser-like HTML5 parsing is important.
  • Use DOMDocument::loadHTML() only with its HTML 4 parsing and non-sanitizing limitations understood.

Frequently Asked Questions

Does strip_tags() remove JavaScript safely?

It is a tag-removal function, not a security control. Escape output and use a sanitizer when untrusted HTML must remain safe.

Can PHP convert HTML to Markdown with these APIs?

No. These approaches produce text; Markdown conversion requires separate rules for headings, links, emphasis, lists, and other constructs.

Which PHP version contains DomHTMLDocument?

The HTML5-oriented DomHTMLDocument::createFromString() API was added in PHP 8.4.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.