Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The most reliable way to collect Ceneo product data is to use the official API that matches your role, not to depend on fragile page selectors. Publishers can apply for Ceneo’s Partner Program API, while shops and manufacturers have a separate Business API. If you still need page HTML, build a rate-limited, failure-tolerant collector only after checking the current Ceneo terms, robots directives, and other legal requirements for the exact host and paths.

Choose the right Ceneo data route first

Your audience and intended use determine the appropriate access method. Ceneo documents two structured-data routes:

Route Intended user Useful data or operations Important constraints
Partner Program API Publishers and affiliate users who direct traffic to Ceneo Product search, product details, category records, prices, shop counts, ratings, reviews, manufacturers and product URLs Affiliate Program access, OAuth 2.0 client-credentials authorization, restricted categories, caching and possible call limits
Business API Shops and manufacturers Popular products, competitor-offer analysis, immediate offer updates and offer-position services Commercial access and method-specific pricing; current terms must be verified with Ceneo
Page HTML collection Teams with a documented, permitted need to inspect public pages Whatever is rendered in the page at capture time Markup, prices, availability and consent flows can change; permission is not established by technical accessibility

For publishers: Partner Program API

Ceneo says its Partner Program API is for publishers who send visitors to Ceneo. The documentation states that a functioning website is required when requesting a test API key. Access is limited to Affiliate Program users and requires OAuth 2.0 client-credentials authorization.

The documented GetProducts operation supports a product-name search with optional category, page size, page index, minimum and maximum price, and filters for offers in the “Buy on Ceneo” flow. Product records can include an ID, name, category, lowest and highest price, basket price where available, number of shops, rating, review count, manufacturer, popularity indicator, product URL and thumbnail URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Category records expose an IsRestricted flag. Restricted categories cannot be used in the Affiliate Program, so do not assume that a successful category lookup means you may promote those products.

For shops and manufacturers: Business API

Ceneo’s Business API is aimed at commercial sellers and suppliers. Its documented uses include finding popular products in a category, comparing competitor offers for selected product IDs, updating offers immediately and increasing an offer’s position. The page lists starting prices of 500 zł per month for popular-product data and offer updates, and 800 zł per month for competitor-offer analysis. These are page-listed starting prices, not a permanent quote; verify current commercial terms before budgeting.

What the Partner API workflow looks like

  1. Apply for the appropriate program. Confirm that your site and intended use meet Ceneo’s current Affiliate Program requirements, or request Business API access if you are a seller or manufacturer.
  2. Store credentials outside source code. Keep the client ID and client secret in environment variables or a secret manager. Never commit them to a repository or send them to a browser.
  3. Request an OAuth token. Use the token endpoint and scope specified in your Ceneo documentation. The client-credentials token expires, so your client must request a new one when it does.
  4. Call the documented resource or operation. Send the Bearer token and the supported OData parameters. Start with a small page size and a narrow query while validating your mapping.
  5. Validate category eligibility. Check IsRestricted before publishing affiliate links or including a category in an automated feed.
  6. Respect freshness and limits. Partner API query results are documented as cached for 15 minutes, although Ceneo notes that this duration may change. Quantitative limits may apply; exceeding them can produce HTTP 403.
  7. Construct attribution correctly. API product URLs do not include a Partner ID. Append the ID from your Affiliate Program account settings when creating attributed links. Do not invent an ID in code or content.

A credential-safe Python client pattern

Ceneo’s documentation supplies the exact token and resource URLs for an enrolled account. Keep those values in environment variables rather than hard-coding an unverified endpoint:

import os
import time
import requests

TOKEN_URL = os.environ["CENEO_TOKEN_URL"]
PRODUCTS_URL = os.environ["CENEO_PRODUCTS_URL"]
CLIENT_ID = os.environ["CENEO_CLIENT_ID"]
CLIENT_SECRET = os.environ["CENEO_CLIENT_SECRET"]

def get_token():
    response = requests.post(
        TOKEN_URL,
        data={"grant_type": "client_credentials"},
        auth=(CLIENT_ID, CLIENT_SECRET),
        timeout=30,
    )
    response.raise_for_status()
    payload = response.json()
    return payload["access_token"], payload.get("expires_in", 300)

def search_products(name, page_size=20, page_index=0):
    token, lifetime = get_token()
    headers = {"Authorization": f"Bearer {token}", "Accept": "application/json"}
    params = {
        "name": name,
        "pageSize": page_size,
        "pageIndex": page_index,
    }
    response = requests.get(PRODUCTS_URL, headers=headers, params=params, timeout=30)
    if response.status_code == 403:
        raise RuntimeError("Ceneo rejected the call: check limits, eligibility and token status")
    response.raise_for_status()
    return response.json()

if __name__ == "__main__":
    print(search_products("wireless headphones"))

Parameter names and casing must match the operation documented for your account. Do not infer them from an HTML form or from a different Ceneo service.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If you must inspect product-page HTML

HTML collection is a fallback, not a guaranteed interface. Before running a crawler, check the current Ceneo terms for the exact host and path, inspect the applicable robots.txt, and obtain legal advice for your jurisdiction and use case. A regulation excerpt for the Ceneo Magazine subsite prohibits bots or programs that burden or hinder that service; that wording does not establish a universal rule for every product page. Never bypass a CAPTCHA, bot check, login wall, rate limit or other access control.

Design a polite collection job

  • Use a small, explicit URL list or a documented discovery source instead of crawling the entire domain.
  • Set a descriptive User-Agent with a contact address where appropriate.
  • Apply a delay between requests, cap concurrency, and stop when the server returns 403, 429 or repeated 5xx responses.
  • Cache successful responses and avoid downloading the same page repeatedly.
  • Record status code, timestamp, final URL, response size and a content hash for auditability.
  • Collect only fields you need, minimize personal data, and define a retention period.

Runnable Python HTML collector

The following script fetches URLs supplied by you and extracts broad, configurable signals. It deliberately avoids claiming that a particular Ceneo selector is permanent. Update selectors only after inspecting the current page and confirming that doing so is permitted.

#!/usr/bin/env python3
import hashlib
import json
import sys
import time
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup

HEADERS = {
    "User-Agent": "ProductDataResearch/1.0 (+mailto:[email protected])",
    "Accept-Language": "pl-PL,pl;q=0.9,en;q=0.8",
}


def fetch(url):
    parsed = urlparse(url)
    if parsed.scheme not in {"http", "https"}:
        raise ValueError(f"Unsupported URL scheme: {url}")
    response = requests.get(url, headers=HEADERS, timeout=30)
    if response.status_code in {403, 429}:
        raise RuntimeError(f"Access denied or rate limited ({response.status_code})")
    response.raise_for_status()
    return response


def parse(response):
    soup = BeautifulSoup(response.text, "html.parser")
    title = soup.title.get_text(" ", strip=True) if soup.title else None
    canonical = soup.find("link", rel="canonical")
    json_ld = []
    for node in soup.find_all("script", type="application/ld+json"):
        try:
            json_ld.append(json.loads(node.string or node.get_text()))
        except json.JSONDecodeError:
            continue
    return {
        "title": title,
        "canonical": canonical.get("href") if canonical else None,
        "json_ld": json_ld,
        "sha256": hashlib.sha256(response.content).hexdigest(),
        "bytes": len(response.content),
    }

for index, url in enumerate(sys.argv[1:]):
    try:
        response = fetch(url)
        print(json.dumps({"url": url, "status": response.status_code, "data": parse(response)}, ensure_ascii=False))
    except Exception as exc:
        print(json.dumps({"url": url, "error": str(exc)}, ensure_ascii=False))
    if index + 1 < len(sys.argv[1:]):
        time.sleep(3)

Run it with python ceneo_pages.py https://your-authorized-url.example. Replace the example URL with a page you are allowed to access. JSON-LD can be a useful structured-data source, but its presence, fields and accuracy must be verified on each page.

Equivalent cURL and Node.js fetches

curl --fail --location --max-time 30 --user-agent "ProductDataResearch/1.0 (+mailto:[email protected])" "https://your-authorized-url.example" -o page.html
const url = process.argv[2];
if (!url) throw new Error('Pass an authorized URL');
const res = await fetch(url, {
  headers: { 'User-Agent': 'ProductDataResearch/1.0 (+mailto:[email protected])' },
  signal: AbortSignal.timeout(30000)
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
console.log(JSON.stringify({ url, bytes: Buffer.byteLength(html), html }));

Parsing, normalization and change detection

Separate acquisition from extraction

Save the raw response, metadata and parser version before transforming fields. This lets you reprocess a page when markup changes without downloading it again, subject to your retention and terms obligations.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Normalize prices carefully

Keep the original displayed string alongside a normalized decimal value and currency. Distinguish a lowest offer, a basket price, a range and a promotional price. Never treat a missing price as zero.

Track product identity

Prefer a documented product ID or canonical URL. Names are not stable identifiers: spelling, storage capacity and bundle contents can differ while titles look similar. Store the capture timestamp and source URL with every record.

Detect breakage

Monitor sudden changes in response size, title frequency, JSON-LD presence, extracted-field counts and HTTP status. A parser that returns empty strings without raising an error can silently corrupt a catalog, so require minimum validation before accepting a record.

Freshness, performance and cost decisions

The Partner API’s documented 15-minute cache means it is unsuitable for claims that require second-by-second market prices. The Business API’s popular-products operation is documented as current within 20 minutes and can return up to 100 popular products; that interval applies to that method, not automatically to every field or endpoint.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTML collection shifts the cost to your infrastructure: bandwidth, retries, storage, browser rendering and maintenance when markup changes. A plain HTTP client is cheaper than a browser but cannot reproduce JavaScript-rendered content. A browser is heavier and more likely to encounter consent banners, bot checks and timeouts. Whichever route you choose, measure successful records rather than raw requests and back off on errors.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

Symptom Likely cause Fix
401 or 403 from the API Expired token, wrong credentials, missing program eligibility, restricted category or call limit Renew the token, verify account permissions and category status, then reduce request volume. Do not retry a persistent authorization failure indefinitely.
429 or repeated 5xx responses Rate limiting or temporary service failure Use exponential backoff with a maximum retry count, honor server guidance and pause the job.
HTML contains a challenge or consent wall Bot protection or an interstitial response Stop; do not bypass it. Use an approved API or request permission and an authorized integration.
Parser returns no price Markup changed, price is JavaScript-rendered, or the item has no current offer Inspect a permitted sample, add validation and preserve the raw response. Do not substitute zero or an unrelated price.
Values disagree between API and page Different cache times, offer scope or currency context Record source and timestamp, explain the distinction to users and choose one authoritative route for each metric.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. A single request can capture an authorized Ceneo page without you managing a browser:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.ceneo.pl -o shot.webp

See the ScreenshotNeo documentation for request options. Before capture, it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Can I scrape every Ceneo product page?

No blanket permission or prohibition is established here. Check the current terms and robots directives for the exact host and path, and obtain advice for your use case before collecting at scale.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is the Partner API real-time?

No. Ceneo documents a 15-minute cache for query results and says the duration may change.

Which API should a seller use?

The Business API is the route documented for shops and manufacturers needing competitor offers, popular-product data or offer updates.

Frequently Asked Questions

Do API product URLs already contain my affiliate ID?

No. Ceneo’s Partner API documentation says product URLs do not include a Partner ID. Add the ID from your Affiliate Program account settings when constructing attributed links.

Can I use restricted categories in affiliate content?

No. The Partner API exposes an IsRestricted flag, and Ceneo states that restricted categories cannot be used in its Affiliate Program.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How many products can the popular-products operation return?

Ceneo’s developer documentation says that operation can return up to 100 popular products in a category.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.