Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

In Selenium, read an input’s original hint with get_dom_attribute("placeholder") and its current contents with get_property("value"). Those calls return Python text, not the HTTP response’s original bytes. If you see � or mojibake, first determine whether the damage is already in the DOM or was introduced by your terminal, log file, or export step. The fix depends on that boundary.

What “placeholder value” means in HTML

The HTML Standard defines placeholder as a short hint shown when a control has no value. It is not the text the user has typed. The live value is held by the input element’s value property. See the WHATWG input specification.

What you need Where it lives Selenium Python call
Hint originally declared in markup placeholder content attribute get_dom_attribute("placeholder")
What is currently in the field Live DOM value property get_property("value")

An input can retain a placeholder attribute while its value changes as the user types. Conversely, JavaScript can change the live value without changing the original attribute. Treating the two as interchangeable is the most common reason a scraper appears to return the “wrong” string.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read the attribute and live value explicitly

Install Selenium and use a locator that matches the actual page. The following example opens a page, waits for a named field, and prints both representations safely:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)

try:
    driver.get("https://example.com/form")
    field = WebDriverWait(driver, 20).until(
        EC.presence_of_element_located((By.NAME, "search"))
    )

    placeholder_hint = field.get_dom_attribute("placeholder")
    current_value = field.get_property("value")

    print("placeholder:", repr(placeholder_hint))
    print("current value:", repr(current_value))
finally:
    driver.quit()

repr() exposes leading or trailing whitespace and makes some invisible characters easier to spot. It does not convert encodings or repair damaged text. If a script fills the field after page load, wait for the condition that represents completion (for example, a non-empty value) before reading it:

from selenium.webdriver.support.ui import WebDriverWait

value = WebDriverWait(driver, 20).until(
    lambda d: d.find_element(By.NAME, "search").get_property("value") or False
)
print(repr(value))

For visible text in a label, paragraph, or other element, inspect that element’s text or DOM structure instead. An input’s .text is not a reliable way to obtain its placeholder or value.

Why get_attribute() can confuse the result

The current Selenium Python binding documents get_attribute() as property-first: it returns a property when one exists and falls back to an attribute of the same name. That behavior is convenient when you do not care which layer supplied the value, but it is ambiguous for fields where markup and live state differ. Use the explicit methods when semantics matter:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • get_dom_attribute("placeholder") asks for the attribute declared in HTML.
  • get_property("value") asks the browser for the current DOM property.
  • get_attribute("value") may return the property rather than the original markup attribute.

The method definitions are in the Selenium 4.49.0 Python WebElement API. Selenium’s guides on element information and finding elements show the related WebElement and locator patterns.

Diagnose a “non-UTF-8” or mojibake result

A Selenium string has already passed through the browser’s HTML parser. You are not holding the original response byte sequence. Therefore, separate two cases before changing code.

Case 1: The DOM itself contains damaged characters

Inspect the value in DevTools or print its repr(). If the browser DOM already shows �, sequences such as é, or other unexpected characters, investigate how the document bytes were decoded. HTML parsing follows the applicable character-encoding declaration and the browser’s encoding algorithm. Check, in this order:

  1. Inspect the response’s Content-Type charset in the browser network panel.
  2. Look for an early <meta charset="..."> declaration. The HTML encoding-declaration rules describe how it is used.
  3. Check whether the server declares a legacy encoding while the page is actually encoded differently, or vice versa.
  4. Compare the same page in “view source” and the parsed DOM. A mismatch indicates decoding or parsing rather than a Selenium accessor problem.

The HTML parsing rules are described in the WHATWG parsing specification, and the algorithms for UTF-8 and legacy encodings are in the WHATWG Encoding Standard. Correct the server or document declaration when you control the page. If you do not, record the actual encoding and decode the source correctly before creating the page, rather than trying random conversions after Selenium has returned a string.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Case 2: The DOM is correct, but your output is damaged

If DevTools and the value returned in Python contain the expected characters but a terminal, log, CSV, or database does not, the problem is downstream. Configure that destination explicitly. For a UTF-8 text file:

text = field.get_property("value")
with open("value.txt", "w", encoding="utf-8", newline="") as output:
    output.write(text)

On a console, verify the process and terminal encoding; on Windows, a UTF-8 capable terminal or an explicit output encoding may be required. For CSV files, pass encoding="utf-8" (or the encoding required by the receiving application) when opening the file. Do not “fix” a correct Python string by repeatedly applying .encode(...).decode(...); that often creates a second layer of corruption.

Handle pages that use a legacy encoding

“Non-UTF-8” is not one universal replacement. It could mean Windows-1252, ISO-8859-1, Shift_JIS, or text that was already mis-decoded. Selenium’s WebDriver API does not expose the original network bytes through get_property(). Use the browser to determine what the DOM contains, then address the source or output boundary.

  • If the page is under your control, send a correct HTTP charset and matching HTML declaration.
  • If it is third-party content, preserve the DOM string as Selenium returned it and document the page’s declared encoding.
  • If you separately download the HTML for analysis, obtain the response bytes and decode them once with the verified charset before parsing.
  • If a replacement character (�) is already present, the original byte may be unrecoverable; changing the Python string’s encoding cannot reconstruct it.

Common failure modes and fixes

Symptom Likely cause Fix
None for the placeholder The element has no placeholder attribute, or the locator found a different element. Inspect the element in DevTools, verify the locator, and check get_dom_attribute("placeholder") is None intentionally.
Placeholder expected, but user text is returned get_attribute() selected a property or you read value. Use get_dom_attribute("placeholder").
Empty value despite visible text The field is empty and only its hint is displayed, or JavaScript has not populated it yet. Read the placeholder separately and wait for the application’s populated state.
Replacement characters in Python The DOM was decoded with the wrong charset, or the source already contained U+FFFD. Check response headers, the meta declaration, and the DOM/source boundary; do not blindly re-encode.
Correct Python output but broken file Writer or consumer uses a different encoding. Open the destination with an explicit encoding and test the receiving application.
StaleElementReferenceException A framework replaced the input after you located it. Wait for the replacement and locate the element again immediately before reading.
Element not found Wrong selector, iframe, shadow DOM, or content loaded later. Use an explicit wait, switch to the correct iframe, or use the component’s documented shadow-root access.

Make the read reliable in automation

Wait for the right state

presence_of_element_located only means the node exists. For a value populated by JavaScript, wait for a predicate that checks the property. For a placeholder injected later, wait until get_dom_attribute("placeholder") is not None.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Capture evidence at the failure boundary

When diagnosing a report, save the URL, locator, repr() of both strings, the browser and driver versions, and the page’s response charset. Comparing those four items (markup attribute, live property, DOM display, and exported output) identifies where corruption begins without guessing.

Do not mutate the browser’s value unnecessarily

Reading is side-effect free. Avoid JavaScript such as element.value.encode(...) in the browser unless you have a documented reason; it can hide whether the page or your diagnostic code introduced the problem.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF rather than DOM-level extraction, ScreenshotNeo provides a website screenshot API and MCP server. A GET request captures a URL without maintaining your own Selenium driver:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all options. Cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed before capture; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python and Node.js alternatives for a screenshot request

When a team already uses Python or Node.js for surrounding automation, the same endpoint can be called directly:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));

ScreenshotNeo supports full-page and element captures, device presets or custom viewports, retina scale, PDF controls, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous webhooks, bulk capture, usage reporting, and an OpenAPI specification. Those features complement Selenium when you need rendered evidence instead of a field’s exact DOM string.

Frequently Asked Questions

Does a placeholder become the input’s value when I click the field?

No. A placeholder remains a hint while the control is empty. Read the live value property to know what the user or script entered.

Can Selenium convert a Windows-1252 page to UTF-8 for me?

The browser decodes the document before WebDriver returns a Python string. Verify the page’s declared encoding and fix the source or output layer; there is no safe universal conversion after corruption.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does the browser display a character that my log cannot show?

The DOM and Python value can be correct while the terminal, file writer, or consuming application uses another encoding. Configure that output boundary explicitly.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.