What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can scrape public AliExpress product pages with Python, but the right method depends on what the server sends back. First inspect a normal HTTP response: if it contains the product fields you need, requests and BeautifulSoup are simpler and lighter. If the response is an incomplete JavaScript shell, use Playwright to render the page and inspect the resulting HTML. For sustained or authorized collection, consider AliExpress’s Open Platform API or a managed crawling service instead of assuming browser automation will always work.

This guide shows how to check access rules, test one public product URL, collect visible product information defensively, and respond to missing data or blocking. AliExpress markup and access conditions can change, so treat extracted fields as best-effort observations—not a stable data feed.

Before you scrape: choose public data and check access

Keep the task narrowly scoped to information presented on public product pages, such as a title, displayed price, rating, order count, store name, shipping information, product URL, and image URL. Do not scrape account pages, order histories, private information, or anything that requires circumventing a login or access control.

Before making requests, read AliExpress’s current terms and its robots.txt. RFC 9309 says that when a crawler successfully downloads a robots.txt file, it must follow the parseable rules. A robots file is not a substitute for permission or the site’s terms; it is one important access signal. If the page or path is disallowed, do not crawl it.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Anti-bot defenses can interrupt automated access. Use a low request rate, add jitter and backoff, and stop if you receive a challenge or repeated blocking responses. A proxy, a different user agent, or a CAPTCHA workflow is not a reason to evade a restriction.

Choose the collection method

Method Best fit Strength Limitation
Requests + BeautifulSoup A small test or a page whose needed values are in the HTTP response Lightweight and straightforward It cannot execute page JavaScript; fields populated in the browser may be absent.
Playwright A public product page that needs browser rendering Runs JavaScript and provides request/response diagnostics Uses more resources and can still encounter blocking.
AliExpress Open Platform API Authorized, structured access Documented HTTP request, signature, and JSON/XML response workflow Requires access and credentials and remains subject to platform terms.
Managed crawling service Teams that need rendering or crawling infrastructure Can outsource browser and network plumbing Has cost and vendor dependence; it does not remove the need to verify authorization and terms.

Check robots.txt before fetching product pages

Python’s urllib.robotparser can test whether a user agent is allowed to fetch a URL and can expose crawl-delay or request-rate values when the file supplies them. This small check is a starting point, not a legal determination. The Python documentation describes can_fetch(useragent, url) as a test of permission under the site’s robots.txt rules.

from urllib.parse import urlsplit
from urllib.robotparser import RobotFileParser


def robots_allows(url: str, user_agent: str) -> bool:
    parts = urlsplit(url)
    robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
    parser = RobotFileParser(robots_url)
    parser.read()

    delay = parser.crawl_delay(user_agent)
    rate = parser.request_rate(user_agent)
    print("robots URL:", robots_url)
    print("crawl delay:", delay)
    print("request rate:", rate)
    return parser.can_fetch(user_agent, url)


product_url = "https://www.aliexpress.com/item/EXAMPLE.html"
if not robots_allows(product_url, "MyResearchBot"):
    raise SystemExit("robots.txt does not allow this URL")

Replace the example URL with a public product URL you are permitted to inspect. If robots.txt cannot be read or parsed reliably, do not treat that uncertainty as approval to crawl; pause and resolve the access question first. For production code, also apply request timeouts, conservative pacing, and an explicit stop condition for errors or challenges.

Test a regular HTTP response first

Install the lightweight dependencies with python -m pip install requests beautifulsoup4. Fetch one page, then inspect its status, final URL, content type, and a short excerpt. A successful status code does not prove the response contains the product information you want.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

url = "https://www.aliexpress.com/item/EXAMPLE.html"
response = requests.get(
    url,
    headers={"User-Agent": "Mozilla/5.0 (compatible; ProductResearch/1.0)"},
    timeout=30,
)
print("status:", response.status_code)
print("final URL:", response.url)
print("content type:", response.headers.get("content-type"))
print(response.text[:1500])
response.raise_for_status()

The user agent here identifies an ordinary browser-like request; it is not a way to gain access or bypass a challenge. Review the returned HTML for the actual values you need. Product information may be present in structured metadata or meta tags, or only after scripts execute. If you see a challenge, an error page, a redirect to an access page, or a page shell without product details, do not repeatedly hammer the URL.

Parse fields defensively with Requests and BeautifulSoup

There is no guaranteed selector that will remain valid across AliExpress page changes, locales, or experiments. The example below reads widely used metadata when present, checks JSON-LD for Product data, and saves the raw HTML for debugging. It deliberately reports missing values as None rather than fabricating them. A field absent from the response is not proof that the product lacks that attribute.

import json
from datetime import datetime, timezone
from pathlib import Path

import requests
from bs4 import BeautifulSoup

URL = "https://www.aliexpress.com/item/EXAMPLE.html"
HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; ProductResearch/1.0)"}


def meta_content(soup, *names):
    for name in names:
        tag = soup.find("meta", attrs={"property": name}) or soup.find(
            "meta", attrs={"name": name}
        )
        if tag and tag.get("content"):
            return tag["content"].strip()
    return None


def json_ld_products(soup):
    """Collect JSON-LD Product objects when the page exposes them."""
    found = []

    def visit(value):
        if isinstance(value, list):
            for item in value:
                visit(item)
        elif isinstance(value, dict):
            kind = value.get("@type")
            kinds = kind if isinstance(kind, list) else [kind]
            if "Product" in kinds:
                found.append(value)
            for key in ("@graph", "mainEntity"):
                if key in value:
                    visit(value[key])

    for script in soup.find_all("script", attrs={"type": "application/ld+json"}):
        raw = script.string or script.get_text()
        try:
            visit(json.loads(raw))
        except (TypeError, json.JSONDecodeError):
            continue
    return found


response = requests.get(URL, headers=HEADERS, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
Path("aliexpress-product.html").write_text(response.text, encoding="utf-8")

products = json_ld_products(soup)
product = products[0] if products else {}
offers = product.get("offers", {})
if isinstance(offers, list):
    offers = offers[0] if offers else {}

result = {
    "retrieved_at_utc": datetime.now(timezone.utc).isoformat(),
    "requested_url": URL,
    "final_url": response.url,
    "status": response.status_code,
    "title": product.get("name") or meta_content(
        soup, "og:title", "twitter:title"
    ) or (soup.title.get_text(" ", strip=True) if soup.title else None),
    "price": offers.get("price") or meta_content(soup, "product:price:amount"),
    "currency": offers.get("priceCurrency") or meta_content(
        soup, "product:price:currency"
    ),
    "rating": (product.get("aggregateRating") or {}).get("ratingValue"),
    "image": product.get("image") or meta_content(soup, "og:image"),
    "json_ld_product_found": bool(products),
}
print(json.dumps(result, ensure_ascii=False, indent=2))

Metadata names and JSON-LD fields are conventions, not a promise that AliExpress provides them on every page. Some values requested by researchers—such as orders sold, store details, or shipping—may not be represented in those fields. If present only in a rendered page, inspect the HTML you actually receive and add a parser for observed, public content. Keep raw responses and retrieval timestamps so you can tell whether a parser change or a page change explains a missing value. Avoid treating displayed prices as normalized comparisons: currency, selected variation, destination, discounts, and shipping can affect what a visitor sees.

Use Playwright when the useful fields need JavaScript

Install Playwright and its browser runtime with python -m pip install playwright, then python -m playwright install chromium. Start with a single allowed page and wait for the page to settle. The example captures rendered HTML and reuses the same metadata parser pattern; it does not assume a particular AliExpress CSS class.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
from bs4 import BeautifulSoup
from playwright.async_api import async_playwright

URL = "https://www.aliexpress.com/item/EXAMPLE.html"

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        page.on("requestfailed", lambda request: print(
            "request failed:", request.url, request.failure
        ))
        page.on("response", lambda response: print(
            "HTTP", response.status, response.url
        ) if response.status >= 400 else None)
        await page.goto(URL, wait_until="domcontentloaded", timeout=60000)
        await page.wait_for_timeout(3000)
        html = await page.content()
        title = await page.title()
        await browser.close()

    soup = BeautifulSoup(html, "html.parser")
    print("page title:", title)
    print("HTML characters:", len(html))
    print("JSON-LD scripts:", len(soup.select(
        'script[type="application/ld+json"]'
    )))
    with open("aliexpress-rendered.html", "w", encoding="utf-8") as f:
        f.write(html)

asyncio.run(main())

The fixed three-second wait is only an example, not a timing guarantee. Prefer waiting for a specific product field or a known page-ready condition after inspecting the current page. Do not wait indefinitely for network idle: analytics or other long-lived requests can prevent it. Playwright’s request API can help diagnose request headers, responses, failures, and sizes; use that for troubleshooting rather than to evade restrictions. If a challenge or block appears, stop and reassess access.

Build a crawl that fails safely

  • Limit the scope. Start with a list of URLs you have a legitimate reason and permission to collect; do not follow arbitrary links across the site.
  • Keep concurrency and rate low. Add randomized pauses between requests and avoid bursts from the same IP. Follow any robots crawl delay or request rate where supplied.
  • Back off on transient failures. Retry only a small number of times for transient network/server errors, using exponential backoff plus jitter. Do not retry challenges, access-denied responses, or repeated blocks as if they were ordinary outages.
  • Set stop conditions. Stop a run when challenge pages, unexpected redirects, repeated failures, or a sharp change in response content appears.
  • Store provenance. Save the source URL, final URL, retrieval timestamp, response status, parser version, and raw HTML where your retention policy allows.
  • Validate output. Check that a title and at least one expected product field exist before accepting a record. Keep missing fields null and flag records for review rather than silently substituting values.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When the official API or a managed service makes more sense

Alibaba’s AliExpress Open Platform documentation describes an HTTP API workflow: populate parameters, generate a signature, assemble and send a request, then interpret JSON or XML. That is the better direction when you need structured access and can obtain the required authorization and credentials. Follow the platform’s current documentation and terms; an API does not grant permission beyond its scope.

A managed crawling API can reduce the work of operating browsers or network infrastructure, but it adds a vendor, cost, and dependency. It also does not make unauthorized collection acceptable. Compare the method’s data coverage and authorization requirements against your exact use case before moving a recurring workload.

Troubleshooting common failures

  • The title is present but price or ratings are missing. The page may not expose that value in the initial HTML or the structured metadata. Inspect a rendered response with Playwright; do not assume a selector from another page or locale applies.
  • Requests returns a thin shell. The useful content may be client-rendered. Confirm by examining the response body, then use a browser only if the public page is accessible and your collection is permitted.
  • You receive a redirect, challenge, or access-denied page. Record the status and final URL, stop repeated requests, and review the site rules and authorization. Do not build a bypass around the challenge.
  • A selector suddenly returns nothing. Page markup can change. Preserve the raw HTML and timestamp, inspect what changed, and update parsing only for fields that are actually present.
  • Runs are slow or unstable. Browser rendering consumes more resources than an HTTP fetch. Limit concurrent browser contexts, reuse a browser for a small bounded batch where appropriate, set timeouts, and stop rather than multiplying retries when failures indicate blocking.
  • Values disagree between runs. Check whether the page changed by locale, currency, destination, variation, or promotion. Store context with each record and avoid presenting a displayed value as a universal price.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a structured AliExpress product-data scraper: its response is a screenshot or PDF rather than parsed product fields. It can be useful when the deliverable is a visual record of a public page. One GET request captures a page; cookie banners, popups, and chat widgets are removed before capture, while bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo and the API documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.aliexpress.com/item/EXAMPLE.html -o shot.webp

Replace the example URL with the public page you are allowed to capture and provide your API key. This produces a visual snapshot, not extracted fields such as a machine-readable price or rating. Sign up for 1,000 free screenshots a month with no card.

Frequently Asked Questions

Is scraping AliExpress legal?

There is no blanket yes-or-no answer for every location and use. Check the current AliExpress terms, applicable law, robots.txt rules, and whether you have authorization for the data and purpose. Avoid private account or personal data and do not bypass access restrictions.

Does AliExpress have an API?

Alibaba documents an AliExpress Open Platform API workflow using HTTP requests, signatures, parameters, and JSON or XML responses. Availability and access depend on the platform’s current requirements; consult its current Open Platform documentation before building against it.

Can I use a screenshot instead of scraping product fields?

Only if a visual record is enough. A screenshot preserves appearance, but it is not a structured product-data response that your program can reliably query for fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.