What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To convert a webpage into an editable Word document, fetch its HTML, use Beautiful Soup to select and clean the content you want, then map that content into a python-docx document and save it as .docx. Those are separate jobs: the parser helps decide what the page says; python-docx creates Word paragraphs, headings, lists, tables and pictures. The result is a structured document, not a guaranteed visual replica of the webpage.
What this workflow can—and cannot—preserve
A webpage is HTML that may include visible content, navigation, scripts, styles and material assembled by JavaScript. A Word file is a document with its own paragraphs, styles, tables and page layout. Converting between them means choosing which HTML structures to keep and mapping them into Word structures; it is not simply saving the page in a different file extension.
Beautiful Soup represents HTML as a tree of Python objects, which lets you select content and remove unwanted nodes. The python-docx library creates and updates Microsoft Word .docx documents. The projects document APIs for headings, paragraphs, list styles, tables, pictures, opening and saving documents. Neither API promises that every webpage will look the same in Word as it does in a browser.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall- Usually practical to retain: readable text, heading hierarchy, basic lists and tables, and images you can download and insert.
- Requires a deliberate choice: which page region is the article, whether links should remain clickable, and how much formatting to reproduce.
- Not automatically retained by text extraction: clickable hyperlinks, CSS layout, and content that only appears after JavaScript runs.
This method creates .docx, the Word format supported by the library for Word 2007 and later. The library does not open legacy Word .doc files; if you need that format, make a separate conversion step with a suitable tool.
#1 Best Overall
Install the Python packages
Use a virtual environment if you want the dependencies for this script isolated from other Python projects. These commands install the HTTP client, HTML parser and Word-document library:
python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4 python-docx
The script below uses Python’s built-in html.parser through Beautiful Soup. You do not need a browser or a separate HTML parser package for this version. If your project already uses lxml, Beautiful Soup can use that parser instead; parser choice and page markup can affect how malformed HTML is interpreted.
Fetch the page and create a DOCX
Save this as webpage_to_docx.py. Replace PAGE_URL with a page you are permitted to retrieve. It selects an <article> when present, then falls back to the page body, removes common non-article elements, and converts headings, paragraphs, lists, tables and downloadable images. It also adds clickable links for anchors in paragraphs and list items. The link text is kept; a hyperlink inside a table cell is emitted as its visible text in this example.
Rank #2
from io import BytesIO
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
from docx import Document
from docx.shared import Inches
from docx.oxml import OxmlElement
from docx.oxml.ns import qn
PAGE_URL = "https://example.com/article"
OUTPUT_FILE = "webpage.docx"
TIMEOUT_SECONDS = 30
def add_hyperlink(paragraph, label, target):
"""Append a clickable hyperlink to a python-docx paragraph."""
relationship_id = paragraph.part.relate_to(
target,
"http://schemas.openxmlformats.org/officeDocument/2006/relationships/hyperlink",
is_external=True,
)
hyperlink = OxmlElement("w:hyperlink")
hyperlink.set(qn("r:id"), relationship_id)
run = OxmlElement("w:r")
properties = OxmlElement("w:rPr")
color = OxmlElement("w:color")
color.set(qn("w:val"), "0563C1")
properties.append(color)
underline = OxmlElement("w:u")
underline.set(qn("w:val"), "single")
properties.append(underline)
run.append(properties)
text = OxmlElement("w:t")
text.text = label
run.append(text)
hyperlink.append(run)
paragraph._p.append(hyperlink)
def append_inline_text(paragraph, node, base_url):
"""Keep readable inline text and make ordinary links clickable."""
for item in node.descendants:
if getattr(item, "name", None) == "a" and item.get("href"):
label = item.get_text(" ", strip=True)
if label:
add_hyperlink(paragraph, label, urljoin(base_url, item["href"]))
elif isinstance(item, str):
# Anchor text is added as a hyperlink above; don't add it twice.
if not item.parent or item.parent.name != "a":
paragraph.add_run(item)
def add_table(doc, html_table):
rows = html_table.find_all("tr")
grid = [[cell.get_text(" ", strip=True)
for cell in row.find_all(["th", "td"], recursive=False)]
for row in rows]
grid = [row for row in grid if row]
if not grid:
return
columns = max(len(row) for row in grid)
table = doc.add_table(rows=len(grid), cols=columns)
table.style = "Table Grid"
for row_index, row in enumerate(grid):
for column_index, value in enumerate(row):
table.cell(row_index, column_index).text = value
def add_image(doc, image_url, session):
"""Insert a reachable image; skip it if retrieval or decoding fails."""
try:
response = session.get(image_url, timeout=TIMEOUT_SECONDS)
response.raise_for_status()
doc.add_picture(BytesIO(response.content), width=Inches(5.5))
except (requests.RequestException, ValueError, OSError) as error:
print(f"Skipped image {image_url}: {error}")
session = requests.Session()
response = session.get(PAGE_URL, timeout=TIMEOUT_SECONDS)
response.raise_for_status()
response.encoding = response.apparent_encoding
soup = BeautifulSoup(response.text, "html.parser")
# Remove elements that are usually not article content. Inspect and adjust
# these selectors for the site you are converting.
for node in soup.select("script, style, template, nav, footer, aside"):
node.decompose()
article = soup.select_one("article") or soup.body or soup
doc = Document()
for element in article.find_all(
["h1", "h2", "h3", "h4", "h5", "h6", "p", "ul", "ol", "table", "img"],
recursive=True,
):
# A list is processed once, including its direct list items.
if element.name in {"ul", "ol"}:
for item in element.find_all("li", recursive=False):
paragraph = doc.add_paragraph(style="List Number" if element.name == "ol" else "List Bullet")
append_inline_text(paragraph, item, PAGE_URL)
continue
# Don't process list items again: their parent list handled them.
if element.name in {"h1", "h2", "h3", "h4", "h5", "h6"}:
text = element.get_text(" ", strip=True)
if text:
level = 0 if element.name == "h1" else min(int(element.name[1]), 9)
doc.add_heading(text, level=level)
elif element.name == "p":
paragraph = doc.add_paragraph()
append_inline_text(paragraph, element, PAGE_URL)
elif element.name == "table":
add_table(doc, element)
elif element.name == "img" and element.get("src"):
add_image(doc, urljoin(PAGE_URL, element["src"]), session)
doc.save(OUTPUT_FILE)
print(f"Saved {OUTPUT_FILE}")
Run it with python webpage_to_docx.py. On success, it writes webpage.docx in the current directory. If the page is large or has many images, retrieval can take longer than the page HTML request alone because each image is fetched separately.
What to inspect in the result
- Open the DOCX in Word or another compatible editor and check that the selected content is the intended article, not a navigation-heavy fallback.
- Check list nesting, table width, image size and heading levels. The example converts direct list children and does not reproduce complex nested-list indentation or table styling.
- Check links and image rights. Relative links are resolved against the page URL; the example fetches images referenced by
src, but does not handle every lazy-loading attribute or responsive-image pattern.
How the HTML-to-Word mapping works
Headings and paragraphs
HTML heading elements become Word heading styles through add_heading(). A top-level h1 is assigned level 0, which is the document-title style in python-docx; h2 and h3 become heading levels 2 and 3. Paragraphs become separate Word paragraphs rather than one long block of text. This keeps the document easier to navigate and restyle.
Lists
Use Word’s List Bullet and List Number paragraph styles rather than inserting bullet characters into plain text. The example handles each list’s direct li children. More elaborate pages can nest lists, mix paragraphs and other blocks inside list items, or use custom list markup; extend the traversal if that structure matters.
Tables
The example creates a Word table with the same number of rows and the maximum number of cells found in any row, then copies cell text. It does not copy CSS, merged-cell relationships, column widths or visual formatting. For a page where table structure is important, inspect the HTML and explicitly handle header rows, spans and layout rather than assuming text extraction is enough.
Images
python-docx accepts a local path or a file-like object for a picture. The example downloads a referenced image into memory and inserts it at a fixed width of 5.5 inches. Image URLs may be relative, which is why it uses urljoin(). Some sites use srcset, lazy-load attributes, authentication or anti-hotlinking rules; the code does not infer or bypass those mechanisms.
Links
Getting visible text from HTML does not recreate a clickable link automatically. The helper in the script adds a hyperlink relationship for anchors inside paragraphs and list items. It deliberately does not attempt to reproduce link styling beyond a basic color and underline, and it does not add hyperlink relationships in table cells.
Choose a retrieval method that fits the page
The example uses requests because it is enough for many pages whose content is already present in the returned HTML. Keep retrieval separate from parsing: that makes it easier to add the authentication, retries, rate-limit handling or timeout policy that the specific site permits and needs.
| Page condition | Practical approach | Trade-off |
|---|---|---|
| Content is in the HTTP response HTML | Fetch with an HTTP client, then parse with Beautiful Soup. | Simple and controllable; page-specific selectors still need inspection. |
| Article content appears only after client-side JavaScript runs | Use a browser-rendering step to obtain rendered HTML, then pass that HTML through the same parsing and DOCX stages. | More operational complexity than a direct HTTP request. Beautiful Soup parses HTML; it does not execute page JavaScript. |
| You need close visual fidelity to CSS or the original page layout | Consider a browser or document-conversion engine designed to render pages. | That can add deployment complexity; this Python mapping approach prioritizes editable Word structure, not pixel parity. |
The latter trade-offs are engineering considerations, not benchmark results. The parser and document library do not guarantee a conversion success rate or a specific visual match.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Customize content selection and document output
Pick the right content container
The selector article is only a useful starting point. Some sites use an element such as main or a class specific to the article body; others have no reliable semantic container. Inspect the page’s markup and replace soup.select_one("article") with the selector that identifies the content you want. The fallback to body is convenient, but may include unrelated page regions.
Best Value
Remove boilerplate carefully
The script removes scripts, styles, templates, navigation, footers and asides before selecting the article. Beautiful Soup’s parser model lets you remove nodes, and its documentation notes that with html.parser or lxml, contents of script and style elements are generally not treated as human-visible text. Navigation, cookie banners, sidebars and other boilerplate still depend on the target site’s markup and may need additional, site-specific selectors. Avoid removing a region until you have checked that it does not contain content you want.
Adjust styles and output destination
Use Word styles, paragraph formatting and page settings to control the document’s appearance after mapping. You can also open an existing .docx with Document(path) and append or update content, or save a document to a file-like stream such as BytesIO. A stream-based output is useful when a web service needs to return DOCX bytes rather than write a file to disk.
Run the workflow as a service
For an application that converts user-submitted URLs, keep the stages separate: validate and retrieve the permitted page, select and parse its content, build the DOCX, then return the bytes. python-docx can save to a filename or a file-like object, so the document can be written into BytesIO before responding with the appropriate DOCX content type. Set timeouts on network requests and handle HTTP errors explicitly; do not let arbitrary inputs turn the converter into an unrestricted network fetcher. Respect the target site’s access rules, authentication requirements and rate limits.
For repeatable results, record which URL was processed and surface failures from each stage separately: retrieval, parsing, image downloads and document saving. This makes a network failure distinguishable from a valid response whose markup simply did not match your content selector. Add retry behavior only where it is appropriate for the site and request type.
Troubleshooting
- The output contains navigation or the wrong text: the page may lack an
articleelement, so the script usedbody. Inspect the HTML and select the actual content container; refine the removal selectors for that site. - The article is blank or missing interactive content: the server response may not contain content added by JavaScript, or the chosen selector may not match. Inspect
response.textand the parsed tree. If content is rendered client-side, add a browser-rendering stage rather than expecting Beautiful Soup to execute scripts. - The request fails or hangs: inspect the HTTP status, network access and timeout. The example calls
raise_for_status(), so unsuccessful HTTP responses stop the script instead of quietly producing an empty document. Handle authentication or other permitted access requirements in the retrieval layer. - Images are absent: check whether the page uses
data-src,srcsetor another lazy-load convention rather thansrc, and whether the resolved image URL is reachable. The example logs a skipped-image message when retrieval or decoding fails. - Links appear as plain text: the helper creates hyperlinks in paragraphs and list items only. Table-cell links are text-only in this example; add corresponding relationships there if clickable table links are required.
- Tables or lists look different in Word: the converter transfers basic cell text and list styles, not the original CSS or every nested structure. Extend the mapping for the specific markup and layout you need to preserve.
- Saving raises an error: check that the destination directory exists and is writable, and that the filename ends in
.docx. This API is for DOCX, not legacy.docinput or output.
Or skip the browser setup
If your real goal is a clean visual screenshot or PDF of a webpage—not an editable Word document—ScreenshotNeo offers a one-request capture. It does not convert webpages to DOCX, so use the Python workflow above when you need Word structure.
For a screenshot, for example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for request options. Equivalent Python and Node.js calls are:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted as a visitor and more than 60 known consent platforms, newsletter popups and chat widgets can be removed before capture; each step can be turned off.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed. Response headers identify the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_infoandcapture_pdftools for AI agents and MCP clients. - The free plan includes 1,000 shots per month without a card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan.
Sign up for 1,000 free screenshots a month with no card.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

