Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →To extract a URL’s HTML, fetch the page and save the server’s response body. For a one-off check, use your browser’s View Source command; for repeatable work, use curl, Wget, or Python Requests. If you need content that appears only after JavaScript runs, the original HTML response may not contain it: inspect the page’s live DOM and network requests, then reproduce the request or use a browser that renders the page.
What “extract the HTML” means
A URL identifies a resource. For a typical webpage, an HTTP GET request asks the server for that resource, and the response body may contain an HTML document. Saving that body gives you the markup the server delivered for that request; it does not necessarily give you every element you later see in a browser.
There are two related things people call “page source”:
- Response HTML: the document returned by the server for your request. This is what command-line tools and ordinary HTTP libraries download.
- Live DOM: the browser’s current document tree after it has parsed the response, run scripts, and possibly fetched additional data. The browser’s Elements panel shows this current tree.
Start by deciding which one you need. Use response HTML to inspect server-rendered markup, save a page for analysis, or build a simple scraper. Use the live DOM or the data requests behind it when the text or elements you need are added after the initial response.
#1 Best Overall
View the source in a browser
For a quick, one-off inspection, open the page and choose the browser’s View Source command. This displays the source document delivered for the page rather than the current DOM shown in the developer tools’ Elements panel. You can search the source for a phrase or element and copy the relevant markup.
When the question is “why is this content missing?”, open developer tools and inspect both Elements and Network. If an element appears in Elements but not in View Source, a script may have inserted it or loaded its data after the initial page request. In Network, look for XHR or fetch requests made during page load; the useful content may be in one of those responses rather than the original HTML.
Download one page with curl or Wget
curl
Use curl’s GET request to save the response body to a file. The -L option follows redirects, which is useful when a URL redirects to its final page.
curl -L "https://example.com" -o page.html
Replace the example URL with the page you can access. Open page.html in a text editor or browser to inspect the returned markup. The file contains the response body, not the response headers.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRank #2
For a quick look at headers and the body together, use -i:
curl -i -L "https://example.com"
To request headers only, use -I:
curl -I -L "https://example.com"
A HEAD request asks for headers without downloading the document body, so it is useful for checking metadata but not for extracting the HTML. If you need the markup, use GET.
Wget
Wget can also save one page to a named file:
wget -O page.html "https://example.com"
For one URL, this is a straightforward alternative to curl. Wget also offers recursive retrieval that follows HTML and CSS references such as href, src, and CSS url() values. Recursion can turn a single-page request into a much larger download, so set a depth, restrict the domain, and choose an output directory before using it to retrieve linked resources.
Fetch and save HTML with Python Requests
Requests gives you the decoded response text, raw response bytes, and response metadata. Install the package if it is not already available with python -m pip install requests. Then save the decoded text using the response’s encoding:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →import requests
url = "https://example.com"
r = requests.get(url, timeout=20)
r.raise_for_status()
html = r.text
print(html)
with open("page.html", "w", encoding=r.encoding or "utf-8") as f:
f.write(html)
The timeout prevents the request from waiting indefinitely, and raise_for_status() makes an HTTP error visible instead of letting an error response pass silently as though it were the page you wanted. A successful HTTP response still does not guarantee that the body is the intended document: inspect the status and content type, and check whether the response is an error page, a login page, or JSON.
Use r.text when you want decoded text. Requests handles response decoding and exposes the selected encoding as r.encoding. If you need the original bytes for later decoding or exact byte-level handling, use r.content instead. Inspect r.headers for metadata such as the content type. Requests also supports cookies, redirects, SSL verification, and timeouts, which can matter when you need to reproduce a legitimate browser request.
Parse the saved HTML with Beautiful Soup
Fetching and parsing are separate jobs: Requests retrieves the response; Beautiful Soup turns markup into a navigable tree. Install the parser package with python -m pip install beautifulsoup4, then pass it the HTML string:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
print(soup.title.get_text(strip=True) if soup.title else "No title")
for link in soup.select("a[href]"):
print(link.get("href"))
The example uses Python’s standard-library html.parser. You can choose lxml when that dependency is available and speed matters, or html5lib when browser-like recovery from malformed markup is useful. Different parsers can build different trees from malformed documents. If the result needs to be reproducible, record which parser you used rather than assuming every parser interprets broken markup identically.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
When the downloaded HTML differs from the browser
Check the response before changing your parser
First compare the saved response with the browser’s View Source. If the content is absent from both, the server may not have sent it in the initial document. If it appears in View Source but not in your parsed output, check the parser and the selector you use. If the page you saved is a login screen, challenge page, error document, or JSON, the request did not return the intended HTML.
Look for JavaScript-loaded data
When the browser shows information that is missing from the original response, inspect the Network panel for XHR or fetch calls. A script may request the content separately and then insert it into the page. If you can access the underlying data endpoint and are authorized to use it, reproduce the relevant request with its actual method, URL, headers, and body. The browser’s “Copy as cURL” option, when available, can help you identify what the browser sent; remove unnecessary headers and credentials before adapting the request for a script.
If reproducing the request is impractical because the page depends on JavaScript execution, use a browser-rendering workflow and extract from the rendered page. A plain HTTP client downloads a response; it does not run the page’s scripts. Do not assume that adding a delay to an ordinary Requests call will make it execute JavaScript.
Match request details only when appropriate
Some pages return different responses depending on the request’s headers, cookies, or authentication. Compare those details with the browser request and reproduce only what the task requires and you are permitted to access. A copied browser request may include temporary session credentials; do not expose those values in shared scripts, logs, or source control.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBest Value
Use Scrapy to inspect a crawler’s response
If you are building with Scrapy and want to see exactly what it receives for a URL, run:
scrapy fetch --nolog https://example.com > response.html
Open response.html and compare it with the browser’s View Source. If the result differs, examine the request headers and user-agent, then identify the browser request that supplies the missing content. Scrapy’s fetch command is useful for inspecting the response from Scrapy’s perspective; it does not make a JavaScript-rendered browser DOM appear in the response.
Choose the method that matches the job
| Method | Best for | What you get | JavaScript execution |
|---|---|---|---|
| Browser View Source | One-off inspection | The source document delivered for the page | No; inspect the live DOM separately |
| curl or Wget | Saving a response from the command line | The HTTP response body, with redirect and request options available | No |
| Python Requests | Repeatable retrieval and further processing in Python | Decoded text, raw bytes, and response metadata | No |
| Beautiful Soup | Finding elements and links in downloaded markup | A parsed tree built from the HTML you provide | No; it parses markup rather than rendering a page |
| Scrapy fetch | Inspecting what a Scrapy request receives | Scrapy’s response for the URL | Not by itself |
| Browser-rendering workflow | Pages whose needed content depends on scripts | A rendered page or DOM, depending on the workflow | Yes |
Common problems and fixes
- The command fails or returns an unexpected page: confirm the URL includes
https://, follow redirects, and inspect the status and content type. The response might be an error document, login page, or non-HTML result rather than the target page. - The saved file has no useful markup: confirm you used GET, not a headers-only HEAD request. Then compare the saved response with View Source to see whether the server sent the content in the first place.
- The browser shows content that the file lacks: compare View Source with Elements, then check Network for XHR or fetch responses that supply the missing data. Use an authorized underlying request or a rendering-capable browser workflow.
- Python appears to succeed but returns an error page: call
raise_for_status()and inspect the status, headers, and body. A page can return HTML even when that HTML is an error or login screen. - Characters look corrupted: inspect the response’s encoding and use Requests’ decoded
textor rawcontentas appropriate. Preserve the response encoding when writing decoded text to a file. - A parsed element or tree looks wrong: verify the input HTML and selector first. For malformed markup, try a different Beautiful Soup parser and note which one produced the result.
- Access depends on a session: inspect the browser’s request for relevant cookies or authentication, and reproduce only the required details when you have authorization. Keep secrets out of logs and shared code.
Or skip the browser setup
If your actual goal is a screenshot or PDF rather than the HTML source, ScreenshotNeo can capture a page through one GET request. It is not an HTML extractor: the response is a screenshot in PNG, JPEG, or WebP, or a PDF. The API can also remove cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off.
For example, this cURL call saves a WebP screenshot of a page:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for the request options. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers indicate the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.
Quick Recap
Practical checks before you rely on an extraction
- Save the URL you requested and the final response you received so you can reproduce the result.
- Keep the status and content type alongside the saved body when you need to diagnose an unexpected response.
- State whether your data came from the original HTML response or a rendered DOM; they are not interchangeable.
- For a crawl, set boundaries before following links so a one-page task does not grow into an unintended download.
- For malformed markup, retain the parser choice with your extraction code so later results can be compared consistently.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




