Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The quickest reliable conversion for a local file is pandoc -f html -t markdown input.html. Pandoc reads the HTML, builds a document model, and writes Markdown. If conversion belongs inside an application, use Turndown in JavaScript or markdownify (or html-to-markdown) in Python. The right choice depends on where your HTML lives, which Markdown dialect you need, and whether tables, images, metadata, whitespace, or unsupported HTML must be preserved.

Choose a converter before you write code

HTML is a browser document model; Markdown is a family of text syntaxes. Headings, paragraphs, links, emphasis, and lists map well, but CSS layout, interactive controls, scripts, forms, embedded media, and arbitrary attributes do not have one universal Markdown representation. Decide what “conversion” means for your project first:

  • One file or a batch of documents: Pandoc is a command-line converter with explicit input and output formats.
  • JavaScript or a browser application: Turndown accepts an HTML string or a DOM node.
  • Python application or script: markdownify provides a direct function; html-to-markdown exposes additional output and whitespace controls.
  • Rendered web pages: obtain the final HTML (including any client-side rendering) before passing it to a converter. A static source file and the DOM seen by a browser can differ substantially.

Always inspect the generated Markdown with the renderer that will consume it. CommonMark, GitHub Flavored Markdown, and tool-specific dialects differ in tables, task lists, footnotes, raw HTML, and extensions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convert an HTML file with Pandoc

The basic command

Pandoc’s official manual gives this form:

pandoc -f html -t markdown input.html

-f (or --from) names the input format and -t (or --to) names the writer format. Pandoc can often infer formats from file extensions, but stating both makes a script unambiguous. The output is written to standard output, so redirect it to a file:

pandoc -f html -t markdown input.html -o output.md

Pandoc describes itself as “a Haskell library for converting from one markup format to another, and a command-line tool that uses this library.” Its reader/writer architecture parses HTML into an intermediate representation, then emits the selected Markdown variant. Filters can modify that representation when a straight conversion is not enough. See the Pandoc User’s Guide for format and extension details.

Choose a Markdown flavor explicitly

“Markdown” is not one exact syntax. Pandoc supports several Markdown readers and writers, including variants with extensions. If the destination has a documented dialect, select it rather than relying on defaults. For example, a workflow targeting GitHub may need GitHub-compatible tables or task-list syntax, while a publishing system may require raw HTML to remain available for constructs Markdown cannot express.

HTML that has no native representation can be retained as raw HTML or simplified, depending on the writer and its extensions. Review links, nested lists, code blocks, tables, and image paths in the result instead of assuming visual similarity means semantic equivalence.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convert a web page in the browser

The Pandoc project publishes a browser application at Pandoc in the browser. The page states that Pandoc WASM runs in the browser and that data is not transmitted to the server. Treat that as the application’s stated behavior, not as a general privacy audit. The project’s examples and demonstrations are collected at Pandoc Demos.

Use Turndown in JavaScript

Install and convert an HTML string

Turndown is a JavaScript tool for converting HTML into Markdown. Its README documents npm installation, browser use, and inputs that can be strings or DOM elements, documents, and fragments.

npm install turndown
const TurndownService = require('turndown');

const turndownService = new TurndownService();
const html = `
  <article>
    <h1>Release notes</h1>
    <p>Version <strong>2.0</strong> is available.</p>
    <ul><li>Faster imports</li><li>New API</li></ul>
  </article>`;

const markdown = turndownService.turndown(html);
console.log(markdown);

For browser code, pass a DOM node instead of serializing it yourself:

const markdown = turndownService.turndown(document.querySelector('article'));

Use this approach when conversion is part of an existing JavaScript pipeline—for example, sanitizing or transforming an article before saving it. The resulting text still needs review for tables, custom elements, embedded widgets, and any HTML that your Markdown renderer handles specially. Consult the Turndown README for its documented options and extensions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convert HTML in Python

Direct conversion with markdownify

The markdownify package converts an HTML string through a Python function. A minimal script is:

from markdownify import markdownify as md

html = """
<h1>Release notes</h1>
<p>Version <strong>2.0</strong> is available.</p>
<ul><li>Faster imports</li><li>New API</li></ul>
"""

markdown = md(html)
print(markdown)

Install it with pip in the environment that runs your script, then use its documented options when you need to strip selected tags or restrict which tags are converted. Those controls are useful when navigation, advertisements, or other page chrome is mixed into the source. See the package documentation on PyPI.

When html-to-markdown is a better fit

The html-to-markdown Python API documents conversion to Markdown, Djot, or plain text. Depending on enabled options, the result can include metadata, document structure, table data, inline images, and warnings. Its API also documents errors for HTML parsing failures and invalid UTF-8.

from html_to_markdown import convert

html = "<h1>Release notes</h1><p>Version 2.0</p>"
markdown = convert(html)
print(markdown)

Use the API reference at html-to-markdown’s Python documentation to select the exact function signature and options for the version you install. Its whitespace setting has two important modes: a normalized mode that collapses consecutive whitespace and a strict mode that preserves source whitespace. Choose normalized output for ordinary prose; choose strict handling when spacing is meaningful to your downstream process.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTML features that need a deliberate decision

Tables

Simple rows and cells can become a Markdown table, but complex tables with rowspans, colspans, nested blocks, or layout-only cells may not map cleanly. Decide whether to simplify the table, keep raw HTML, or extract table data separately. The html-to-markdown API is useful when your application needs structured table information in addition to Markdown text.

Images and links

Markdown normally stores a destination URL and optional alt text; it does not embed the image bytes. Check whether relative URLs remain valid from the Markdown file’s new location. For links, verify that HTML entities, fragments, query strings, and titles survived conversion.

Code, scripts, and styles

Code blocks should retain their language hint when the source provides one, but JavaScript, CSS, forms, and interactive widgets are not Markdown content. Remove them, preserve them as raw HTML, or process them in a separate extraction step. Do not expect a converter to reproduce browser behavior.

Whitespace and entities

HTML collapses many spaces during rendering; source whitespace and displayed whitespace are therefore different concepts. Compare normalized and strict modes where available, and inspect non-breaking spaces, entities, line breaks, and preformatted text. A visually identical page can produce different Markdown depending on whether the converter reads source text or a rendered DOM.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Metadata and document chrome

<title>, Open Graph tags, navigation, cookie notices, and footer links are not automatically useful in the article body. Extract metadata separately if your application needs it, and remove repeated site chrome before conversion. For a web page, obtain the content area rather than blindly converting the entire document.

A practical workflow for reliable output

  1. Identify the source. Determine whether you have a local file, an HTML string, or a browser DOM after JavaScript has run.
  2. Define the destination dialect. Record the renderer or platform that will consume the Markdown and its requirements for tables, raw HTML, links, and extensions.
  3. Isolate content. Remove navigation, ads, consent dialogs, scripts, and duplicated headers before conversion when they are not part of the document.
  4. Run the smallest conversion first. Start with Pandoc’s explicit -f html -t markdown command or the library call for your runtime.
  5. Review semantic hotspots. Check headings, list nesting, code fences, tables, image paths, links, whitespace, and non-ASCII text.
  6. Render and compare. Render the Markdown with the destination engine and compare meaning—not only appearance—to the original HTML.
  7. Automate regression checks. Keep representative fixtures containing nested lists, tables, images, entities, and raw HTML. Re-run them after upgrading a converter or changing options.

Command-line, JavaScript, and Python options compared

Option Best starting point Input interface Useful distinctions
Pandoc Files, batches, and broader document workflows Command line; explicit -f/-t Supports HTML and multiple Markdown variants; filters and raw-HTML behavior are documented in its manual
Turndown JavaScript or browser code HTML string or DOM node Fits an existing JS pipeline; inspect unsupported or custom elements
markdownify Simple Python conversion HTML string and function call Documents tag-stripping and tag-selection options
html-to-markdown Python workflows needing more than plain text Python API Documents Markdown, Djot, and plain text output, metadata/structure fields, table and image data, warnings, and whitespace modes

Or skip the browser setup

If your starting point is a live site and you would otherwise configure a browser just to capture a clean visual reference, ScreenshotNeo makes a screenshot or PDF through one GET request. It is not an HTML-to-Markdown converter, so you still need HTML extraction and one of the converters above for Markdown; use it when a rendered-page snapshot is also part of your workflow.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Before capture it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting conversion problems

“pandoc: command not found”

Pandoc is not installed or is not on the shell’s PATH. Install it using the package method appropriate for your operating system, reopen the terminal, and verify that the pandoc command resolves before rerunning the conversion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The output is empty or mostly navigation

The input may be a shell page whose meaningful content is inserted by JavaScript, or you may have converted the entire site chrome. Save the post-render DOM or isolate the article element, then convert that HTML instead of the initial response.

Tables or layout look wrong

Markdown tables cannot represent every HTML layout. Simplify merged cells, retain selected raw HTML, or use a structured table result when your Python workflow requires fidelity. Confirm what the destination renderer supports.

Images are broken

Inspect relative URLs after moving the Markdown file. Resolve them against the original HTML base URL or copy the assets into a stable location; a converter normally writes references, not image files.

Characters or whitespace changed

Check that the input is valid UTF-8, preserve a correct document encoding, and choose normalized versus strict whitespace deliberately. Parsing errors and invalid UTF-8 are documented failure cases for the html-to-markdown API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Raw HTML appears in the Markdown

That is often intentional: some HTML constructs have no Markdown equivalent. Either keep the raw element for a renderer that permits it, or add a targeted transformation/filter. Do not globally strip raw HTML if it carries required semantics such as a table or a disclosure element.

FAQ

Can I convert HTML to Markdown without installing software?

Yes. The Pandoc WASM application runs in a browser and states that it does not transmit data to its server. For repeatable builds or large batches, a local CLI or library is easier to automate.

Which converter should a JavaScript developer choose?

Start with Turndown when the HTML is already a string or DOM node in JavaScript. Use Pandoc separately when you need its broader document-format workflow.

Will CSS styling be preserved?

No. Markdown represents document structure and inline semantics, not arbitrary CSS layout. Preserve selected raw HTML or keep the original stylesheet when styling is a requirement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What if I need Markdown plus metadata?

Use a conversion API that exposes metadata or structure fields, such as the documented options in html-to-markdown, and store that data separately from the Markdown body.

Frequently Asked Questions

Can I convert HTML to Markdown without installing software?

Yes. The Pandoc WASM application runs in a browser and states that it does not transmit data to its server. For repeatable builds or large batches, a local CLI or library is easier to automate.

Which converter should a JavaScript developer choose?

Start with Turndown when the HTML is already a string or DOM node in JavaScript. Use Pandoc separately when you need its broader document-format workflow.

Will CSS styling be preserved?

No. Markdown represents document structure and inline semantics, not arbitrary CSS layout. Preserve selected raw HTML or keep the original stylesheet when styling is a requirement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What if I need Markdown plus metadata?

Use a conversion API that exposes metadata or structure fields, such as the documented options in html-to-markdown, and store that data separately from the Markdown body.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.