Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The supported way to automate Yandex text searches is the Yandex Search API, not an improvised scraper for consumer SERP pages. The API accepts REST, gRPC, or the Yandex AI Studio SDK requests. In this guide you will authenticate correctly, submit the same search from Python and Node.js, decode the synchronous response, choose XML or HTML deliberately, and handle pagination, localization, deferred jobs, and changing response fields.
API retrieval versus scraping Yandex’s public SERP
“Scrape” often means extracting data automatically, but two different activities are being mixed together:
- API retrieval: your program sends a documented request to Yandex Search API and parses its response.
- SERP scraping: your program fetches the consumer search-results page and reverse-engineers its HTML.
Use the first approach for a production integration. The old Yandex.XML license page says that service became void on November 1, 2024, and describes automated requests to Yandex Search by other means as prohibited without prior approval. That is a warning about the legacy service, not permission to copy its examples. Check the current Search API terms, access requirements, limits, and pricing before deployment. A robots.txt rule on a site you own controls crawlers visiting that site; it does not authorize automated requests to Yandex Search.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →What you need before writing code
Authentication and IAM permissions
Every API request must be authenticated. A user or federated account sends an IAM token in a Bearer header and must also provide the folder ID. A service account can authenticate with an IAM token or an API key in the Authorization header and can use its own folder. Grant the account the search-api.webSearch.user role.
#1 Best Overall
Keep tokens, API keys, and folder IDs in environment variables or a secret manager. Never commit them to a repository, paste them into client-side JavaScript, or log an Authorization header.
Choose the search context
Decide the language and geography before comparing results. The API lists Russian, Turkish, international, Kazakh, Belarusian, and Uzbek search types. Region filtering is supported for Russian and Turkish search types. Record the selected search type, region, family filtering, sort settings, and page size with each job so another person can reproduce the intended context.
Request limits and response formats
queryTextis limited to 400 characters.- The documented maximum is 250 results per query; this is not an unlimited or guaranteed stable snapshot.
- XML is the default response format. HTML is available when your parser needs page-like elements such as ads or quick responses.
- Synchronous responses put XML or HTML in Base64-encoded
rawData; decode it before parsing.
Important Search API options
| Purpose | REST field | What to decide |
|---|---|---|
| Query and page | queryText, page |
Keep the query at 400 characters or fewer and request pages explicitly. |
| Language and geography | searchType, region, l10n |
Use a search type and localization that match the audience; region applies to Russian and Turkish types. |
| Filtering | familyMode, fixTypoMode |
Choose family-safe filtering and typo correction rather than relying on defaults. |
| Ranking | sortMode, sortOrder |
Make relevance or another supported ordering explicit. |
| Grouping and volume | groupMode, groupsOnPage, docsInGroup |
Control grouped results and stay within the format-specific valid ranges. |
| Billing context and output | folderId, responseFormat, resultsWithin |
Supply the required folder and select XML or HTML intentionally. |
REST uses CamelCase field names. gRPC uses snake_case equivalents. Fields can be absent, and response content can change without prior notice, so parsers must tolerate missing nodes and unknown additions.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Python: submit a synchronous REST query
The following example uses the standard library, so it needs no third-party package. Set YANDEX_IAM_TOKEN and YANDEX_FOLDER_ID in the process environment. The exact REST endpoint and request envelope should match the current Yandex Search API documentation and your account configuration.
Rank #2
import base64
import json
import os
from urllib.request import Request, urlopen
API_URL = os.environ["YANDEX_SEARCH_API_URL"]
token = os.environ["YANDEX_IAM_TOKEN"]
folder_id = os.environ["YANDEX_FOLDER_ID"]
payload = {
"queryText": "site:example.com accessibility",
"searchType": "SEARCH_TYPE_RU",
"familyMode": "FAMILY_MODE_MODERATE",
"page": 0,
"fixTypoMode": "FIX_TYPO_MODE_ON",
"sortMode": "SORT_MODE_BY_RELEVANCE",
"sortOrder": "SORT_ORDER_DESC",
"groupsOnPage": 10,
"docsInGroup": 1,
"folderId": folder_id,
"responseFormat": "RESPONSE_FORMAT_XML",
}
request = Request(
API_URL,
data=json.dumps(payload).encode("utf-8"),
headers={
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
},
method="POST",
)
with urlopen(request, timeout=60) as response:
envelope = json.loads(response.read())
raw_data = envelope.get("rawData")
if not raw_data:
raise RuntimeError(f"No rawData in response: {envelope.keys()}")
result_bytes = base64.b64decode(raw_data)
result_text = result_bytes.decode("utf-8")
print(result_text)
For XML, pass result_text to an XML parser and check each element before reading it. For HTML, parse it with an HTML parser and expect ads, quick responses, and other page elements in addition to organic results. Do not assume that an XML field has an HTML equivalent.
Node.js: the same request with fetch
Node.js 18 or later includes fetch. This example sends JSON, decodes Base64, and prints the UTF-8 payload. Set YANDEX_SEARCH_API_URL, YANDEX_IAM_TOKEN, and YANDEX_FOLDER_ID before running it.
const apiUrl = process.env.YANDEX_SEARCH_API_URL;
const token = process.env.YANDEX_IAM_TOKEN;
const folderId = process.env.YANDEX_FOLDER_ID;
if (!apiUrl || !token || !folderId) {
throw new Error('Set YANDEX_SEARCH_API_URL, YANDEX_IAM_TOKEN and YANDEX_FOLDER_ID');
}
const payload = {
queryText: 'site:example.com accessibility',
searchType: 'SEARCH_TYPE_RU',
familyMode: 'FAMILY_MODE_MODERATE',
page: 0,
fixTypoMode: 'FIX_TYPO_MODE_ON',
sortMode: 'SORT_MODE_BY_RELEVANCE',
sortOrder: 'SORT_ORDER_DESC',
groupsOnPage: 10,
docsInGroup: 1,
folderId,
responseFormat: 'RESPONSE_FORMAT_XML'
};
const response = await fetch(apiUrl, {
method: 'POST',
headers: {
Authorization: `Bearer ${token}`,
'Content-Type': 'application/json'
},
body: JSON.stringify(payload)
});
if (!response.ok) {
throw new Error(`Yandex API ${response.status}: ${await response.text()}`);
}
const envelope = await response.json();
if (!envelope.rawData) throw new Error('Response has no rawData');
const resultText = Buffer.from(envelope.rawData, 'base64').toString('utf8');
console.log(resultText);
Use a maintained XML or HTML parser after the decode step. Check HTTP status before parsing; an error document is not a search result.
Free tools Windows power users keep installed
One-click scans. No signup required.
Equivalent cURL request for debugging
cURL is useful for isolating credentials and payload problems before investigating application code.
curl --fail-with-body "$YANDEX_SEARCH_API_URL"
-H "Authorization: Bearer $YANDEX_IAM_TOKEN"
-H "Content-Type: application/json"
--data '{
"queryText":"site:example.com accessibility",
"searchType":"SEARCH_TYPE_RU",
"familyMode":"FAMILY_MODE_MODERATE",
"page":0,
"folderId":"'"$YANDEX_FOLDER_ID"'",
"responseFormat":"RESPONSE_FORMAT_XML"
}'
Pagination, grouping, and deferred processing
Pages are bounded
Increment page only while the response contains usable groups or documents, and stop at the documented 250-result ceiling. Store the query settings alongside each page. Ranking can change between requests, so do not describe a multi-page crawl as an immutable snapshot.
Groups are not individual documents
groupMode, groupsOnPage, and docsInGroup affect how results are clustered. A page containing ten groups may contain a different number of documents. Test your parser against both grouped and ungrouped responses.
Deferred jobs
The API also supports deferred mode. Instead of result data, the initial call returns an operation object. Save its operation ID, poll or track that ID, and read the response only after done becomes true. Implement a timeout, backoff, and a final failure state; do not busy-loop indefinitely.
Recommended Free Tools
Defensive parsing and operational reliability
- Treat every result field as optional. Missing titles, snippets, links, or groups should produce a partial record, not a crashed batch.
- Preserve the raw decoded payload for debugging, subject to your privacy and retention rules.
- Use connect and total request timeouts, bounded retries for transient transport failures, and no automatic retry for authentication or validation errors.
- Rate-limit your own workers and monitor status codes, latency, empty-result rates, and deferred-operation failures.
- Pin parser behavior with fixtures because Yandex warns that response content may change without prior notice.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 Unauthorized | Expired/malformed token or missing Bearer prefix | Mint a current IAM token and send Authorization: Bearer TOKEN. |
| 403 Forbidden | Account lacks search-api.webSearch.user or folder access |
Grant the role and verify the folder ID belongs to the requesting identity. |
| 400 validation error | Wrong enum, field spelling, page size, or overlong query | Use REST CamelCase, valid enum values, format-specific ranges, and no more than 400 characters. |
| Parser sees gibberish | rawData is still Base64 |
Base64-decode it, then decode UTF-8 before XML/HTML parsing. |
| Empty or missing nodes | No matches, grouping differences, or an API response change | Handle optional fields and log the raw response for inspection. |
| Results differ by country or run | Different search type, region, localization, ranking, or changing index | Set those parameters explicitly and treat rankings as time-sensitive. |
Or skip the browser setup
If your actual requirement is a visual capture of a web page rather than structured Yandex result data, ScreenshotNeo provides a one-request screenshot API and MCP server. It is not a replacement for Search API result fields, but it avoids maintaining browser automation for page images.
For example, this captures a page as WebP:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for options. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I keep using old Yandex.XML tutorials?
Only as historical reference. The license page itself says it became void on November 1, 2024. Base a new integration on the current Search API documentation and its terms.
Should I request XML or HTML?
Choose XML for a compact, structured parser. Choose HTML only when your application needs page elements such as ads or quick responses, and parse it as a different schema.
Does the API guarantee the same ranking tomorrow?
No. Geography, settings, and index changes can alter results, and the documentation warns that response content may change without prior notice.
Frequently Asked Questions
Can I keep using old Yandex.XML tutorials?
Only as historical reference. The license page itself says it became void on November 1, 2024. Base a new integration on the current Search API documentation and its terms.
Best Value
Should I request XML or HTML?
Choose XML for a compact, structured parser. Choose HTML only when your application needs page elements such as ads or quick responses, and parse it as a different schema.
Does the API guarantee the same ranking tomorrow?
No. Geography, settings, and index changes can alter results, and the documentation warns that response content may change without prior notice.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →The Bottom Line
For a maintainable Yandex integration, authenticate the documented Search API, make language and region explicit, decode rawData, and parse defensively. Treat direct consumer-SERP scraping and legacy Yandex.XML examples as separate, potentially restricted methods.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

