Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The supported way to automate Yandex text searches is the Yandex Search API, not an improvised scraper for consumer SERP pages. The API accepts REST, gRPC, or the Yandex AI Studio SDK requests. In this guide you will authenticate correctly, submit the same search from Python and Node.js, decode the synchronous response, choose XML or HTML deliberately, and handle pagination, localization, deferred jobs, and changing response fields.

API retrieval versus scraping Yandex’s public SERP

“Scrape” often means extracting data automatically, but two different activities are being mixed together:

  • API retrieval: your program sends a documented request to Yandex Search API and parses its response.
  • SERP scraping: your program fetches the consumer search-results page and reverse-engineers its HTML.

Use the first approach for a production integration. The old Yandex.XML license page says that service became void on November 1, 2024, and describes automated requests to Yandex Search by other means as prohibited without prior approval. That is a warning about the legacy service, not permission to copy its examples. Check the current Search API terms, access requirements, limits, and pricing before deployment. A robots.txt rule on a site you own controls crawlers visiting that site; it does not authorize automated requests to Yandex Search.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What you need before writing code

Authentication and IAM permissions

Every API request must be authenticated. A user or federated account sends an IAM token in a Bearer header and must also provide the folder ID. A service account can authenticate with an IAM token or an API key in the Authorization header and can use its own folder. Grant the account the search-api.webSearch.user role.

Keep tokens, API keys, and folder IDs in environment variables or a secret manager. Never commit them to a repository, paste them into client-side JavaScript, or log an Authorization header.

Choose the search context

Decide the language and geography before comparing results. The API lists Russian, Turkish, international, Kazakh, Belarusian, and Uzbek search types. Region filtering is supported for Russian and Turkish search types. Record the selected search type, region, family filtering, sort settings, and page size with each job so another person can reproduce the intended context.

Request limits and response formats

  • queryText is limited to 400 characters.
  • The documented maximum is 250 results per query; this is not an unlimited or guaranteed stable snapshot.
  • XML is the default response format. HTML is available when your parser needs page-like elements such as ads or quick responses.
  • Synchronous responses put XML or HTML in Base64-encoded rawData; decode it before parsing.

Important Search API options

Purpose REST field What to decide
Query and page queryText, page Keep the query at 400 characters or fewer and request pages explicitly.
Language and geography searchType, region, l10n Use a search type and localization that match the audience; region applies to Russian and Turkish types.
Filtering familyMode, fixTypoMode Choose family-safe filtering and typo correction rather than relying on defaults.
Ranking sortMode, sortOrder Make relevance or another supported ordering explicit.
Grouping and volume groupMode, groupsOnPage, docsInGroup Control grouped results and stay within the format-specific valid ranges.
Billing context and output folderId, responseFormat, resultsWithin Supply the required folder and select XML or HTML intentionally.

REST uses CamelCase field names. gRPC uses snake_case equivalents. Fields can be absent, and response content can change without prior notice, so parsers must tolerate missing nodes and unknown additions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python: submit a synchronous REST query

The following example uses the standard library, so it needs no third-party package. Set YANDEX_IAM_TOKEN and YANDEX_FOLDER_ID in the process environment. The exact REST endpoint and request envelope should match the current Yandex Search API documentation and your account configuration.

import base64
import json
import os
from urllib.request import Request, urlopen

API_URL = os.environ["YANDEX_SEARCH_API_URL"]
token = os.environ["YANDEX_IAM_TOKEN"]
folder_id = os.environ["YANDEX_FOLDER_ID"]

payload = {
    "queryText": "site:example.com accessibility",
    "searchType": "SEARCH_TYPE_RU",
    "familyMode": "FAMILY_MODE_MODERATE",
    "page": 0,
    "fixTypoMode": "FIX_TYPO_MODE_ON",
    "sortMode": "SORT_MODE_BY_RELEVANCE",
    "sortOrder": "SORT_ORDER_DESC",
    "groupsOnPage": 10,
    "docsInGroup": 1,
    "folderId": folder_id,
    "responseFormat": "RESPONSE_FORMAT_XML",
}

request = Request(
    API_URL,
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "Authorization": f"Bearer {token}",
        "Content-Type": "application/json",
    },
    method="POST",
)

with urlopen(request, timeout=60) as response:
    envelope = json.loads(response.read())

raw_data = envelope.get("rawData")
if not raw_data:
    raise RuntimeError(f"No rawData in response: {envelope.keys()}")

result_bytes = base64.b64decode(raw_data)
result_text = result_bytes.decode("utf-8")
print(result_text)

For XML, pass result_text to an XML parser and check each element before reading it. For HTML, parse it with an HTML parser and expect ads, quick responses, and other page elements in addition to organic results. Do not assume that an XML field has an HTML equivalent.

Node.js: the same request with fetch

Node.js 18 or later includes fetch. This example sends JSON, decodes Base64, and prints the UTF-8 payload. Set YANDEX_SEARCH_API_URL, YANDEX_IAM_TOKEN, and YANDEX_FOLDER_ID before running it.

const apiUrl = process.env.YANDEX_SEARCH_API_URL;
const token = process.env.YANDEX_IAM_TOKEN;
const folderId = process.env.YANDEX_FOLDER_ID;

if (!apiUrl || !token || !folderId) {
  throw new Error('Set YANDEX_SEARCH_API_URL, YANDEX_IAM_TOKEN and YANDEX_FOLDER_ID');
}

const payload = {
  queryText: 'site:example.com accessibility',
  searchType: 'SEARCH_TYPE_RU',
  familyMode: 'FAMILY_MODE_MODERATE',
  page: 0,
  fixTypoMode: 'FIX_TYPO_MODE_ON',
  sortMode: 'SORT_MODE_BY_RELEVANCE',
  sortOrder: 'SORT_ORDER_DESC',
  groupsOnPage: 10,
  docsInGroup: 1,
  folderId,
  responseFormat: 'RESPONSE_FORMAT_XML'
};

const response = await fetch(apiUrl, {
  method: 'POST',
  headers: {
    Authorization: `Bearer ${token}`,
    'Content-Type': 'application/json'
  },
  body: JSON.stringify(payload)
});

if (!response.ok) {
  throw new Error(`Yandex API ${response.status}: ${await response.text()}`);
}

const envelope = await response.json();
if (!envelope.rawData) throw new Error('Response has no rawData');
const resultText = Buffer.from(envelope.rawData, 'base64').toString('utf8');
console.log(resultText);

Use a maintained XML or HTML parser after the decode step. Check HTTP status before parsing; an error document is not a search result.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Equivalent cURL request for debugging

cURL is useful for isolating credentials and payload problems before investigating application code.

curl --fail-with-body "$YANDEX_SEARCH_API_URL" 
  -H "Authorization: Bearer $YANDEX_IAM_TOKEN" 
  -H "Content-Type: application/json" 
  --data '{
    "queryText":"site:example.com accessibility",
    "searchType":"SEARCH_TYPE_RU",
    "familyMode":"FAMILY_MODE_MODERATE",
    "page":0,
    "folderId":"'"$YANDEX_FOLDER_ID"'",
    "responseFormat":"RESPONSE_FORMAT_XML"
  }'

Pagination, grouping, and deferred processing

Pages are bounded

Increment page only while the response contains usable groups or documents, and stop at the documented 250-result ceiling. Store the query settings alongside each page. Ranking can change between requests, so do not describe a multi-page crawl as an immutable snapshot.

Groups are not individual documents

groupMode, groupsOnPage, and docsInGroup affect how results are clustered. A page containing ten groups may contain a different number of documents. Test your parser against both grouped and ungrouped responses.

Deferred jobs

The API also supports deferred mode. Instead of result data, the initial call returns an operation object. Save its operation ID, poll or track that ID, and read the response only after done becomes true. Implement a timeout, backoff, and a final failure state; do not busy-loop indefinitely.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Defensive parsing and operational reliability

  • Treat every result field as optional. Missing titles, snippets, links, or groups should produce a partial record, not a crashed batch.
  • Preserve the raw decoded payload for debugging, subject to your privacy and retention rules.
  • Use connect and total request timeouts, bounded retries for transient transport failures, and no automatic retry for authentication or validation errors.
  • Rate-limit your own workers and monitor status codes, latency, empty-result rates, and deferred-operation failures.
  • Pin parser behavior with fixtures because Yandex warns that response content may change without prior notice.

Common errors and fixes

Symptom Likely cause Fix
401 Unauthorized Expired/malformed token or missing Bearer prefix Mint a current IAM token and send Authorization: Bearer TOKEN.
403 Forbidden Account lacks search-api.webSearch.user or folder access Grant the role and verify the folder ID belongs to the requesting identity.
400 validation error Wrong enum, field spelling, page size, or overlong query Use REST CamelCase, valid enum values, format-specific ranges, and no more than 400 characters.
Parser sees gibberish rawData is still Base64 Base64-decode it, then decode UTF-8 before XML/HTML parsing.
Empty or missing nodes No matches, grouping differences, or an API response change Handle optional fields and log the raw response for inspection.
Results differ by country or run Different search type, region, localization, ranking, or changing index Set those parameters explicitly and treat rankings as time-sensitive.

Or skip the browser setup

If your actual requirement is a visual capture of a web page rather than structured Yandex result data, ScreenshotNeo provides a one-request screenshot API and MCP server. It is not a replacement for Search API result fields, but it avoids maintaining browser automation for page images.

For example, this captures a page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for options. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

FAQ

Can I keep using old Yandex.XML tutorials?

Only as historical reference. The license page itself says it became void on November 1, 2024. Base a new integration on the current Search API documentation and its terms.

Should I request XML or HTML?

Choose XML for a compact, structured parser. Choose HTML only when your application needs page elements such as ads or quick responses, and parse it as a different schema.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does the API guarantee the same ranking tomorrow?

No. Geography, settings, and index changes can alter results, and the documentation warns that response content may change without prior notice.

Frequently Asked Questions

Can I keep using old Yandex.XML tutorials?

Only as historical reference. The license page itself says it became void on November 1, 2024. Base a new integration on the current Search API documentation and its terms.

Should I request XML or HTML?

Choose XML for a compact, structured parser. Choose HTML only when your application needs page elements such as ads or quick responses, and parse it as a different schema.

Does the API guarantee the same ranking tomorrow?

No. Geography, settings, and index changes can alter results, and the documentation warns that response content may change without prior notice.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

For a maintainable Yandex integration, authenticate the documented Search API, make language and region explicit, decode rawData, and parse defensively. Treat direct consumer-SERP scraping and legacy Yandex.XML examples as separate, potentially restricted methods.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.