October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AliExpress

How to Scrape AliExpress Search Pages (Safely, with Pagination and JavaScript Handling)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrape an AliExpress search page as a bounded data-collection job: build the public keyword URL, fetch each page conservatively, extract repeated product-card fields, detect challenge or empty responses, paginate only within a hard limit, and deduplicate by product URL or ID. If the HTML is only a JavaScript shell, inspect embedded JSON or use a headless browser for those pages. For production or commercial collection, obtain written permission or use an approved API or managed crawler first.

Start with permission and a precise data contract

AliExpress Terms of Use state: “Systematic retrieval of Site Content from the Sites to create or compile, directly or indirectly, a collection, compilation, database or directory (whether through robots, spiders, automatic devices or manual processes) without written permission from AliExpress.com is prohibited.” The same terms restrict copying, downloading, republishing, selling, or commercially exploiting site content. Treat this as a permission boundary, not a prompt to bypass anti-bot controls. Review applicable law, robots guidance, rate limits, and any written authorization before a production run.

Write down exactly what the collector is allowed to store. A useful product record contains:

  • Identity: canonical product URL and product ID when available.
  • Display fields: title, price text, rating text, and order-count text.
  • Provenance: original keyword, page number, retrieval timestamp, HTTP status, response length, and parser version.
  • Run controls: maximum pages, delay policy, and the reason the run stopped.

Keeping query and page metadata makes later changes explainable. It also prevents a blank challenge page from being mistaken for a valid page with zero products.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Construct a public search URL and bounded pagination

The wholesale-style search route uses a hyphenated keyword and a page query parameter. Normalize whitespace, create one URL per page, and keep a hard maximum even if the site appears to offer more results.

from urllib.parse import urlencode
import re

def search_url(keyword: str, page: int) -> str:
    slug = re.sub(r"s+", "-", keyword.strip())
    query = urlencode({"SearchText": slug, "page": page})
    return f"https://www.aliexpress.com/wholesale?{query}"

print(search_url("usb c hub", 1))

Do not assume that a successful HTTP response means a successful scrape. Stop or quarantine a page when it contains a CAPTCHA, bot-check language, an interstitial challenge, an unexpectedly tiny body, or an abrupt item-count change. A normal empty result and a blocked response are different outcomes and should be recorded separately.

A conservative Python collector

The following example uses requests and Beautiful Soup. AliExpress markup changes, so the selectors are deliberately defensive rather than tied to one private class name. Inspect a current, authorized response and adjust the candidate selectors for your run.

import json
import re
import time
from datetime import datetime, timezone
from urllib.parse import urlencode, urljoin, urlparse

import requests
from bs4 import BeautifulSoup

BASE = "https://www.aliexpress.com/wholesale"
HEADERS = {
    "User-Agent": "authorized-research-bot/1.0 ([email protected])",
    "Accept-Language": "en-US,en;q=0.8",
}
CHALLENGE_WORDS = (
    "captcha", "robot check", "verify you are human", "security verification",
    "access denied", "unusual traffic"
)


def make_url(keyword, page):
    slug = re.sub(r"s+", "-", keyword.strip())
    return f"{BASE}?{urlencode({'SearchText': slug, 'page': page})}"


def challenged(text):
    sample = text[:200_000].lower()
    return any(word in sample for word in CHALLENGE_WORDS)


def first_text(node, selectors):
    for selector in selectors:
        found = node.select_one(selector)
        if found:
            value = found.get("content") or found.get_text(" ", strip=True)
            if value:
                return value
    return None


def extract_cards(html, source_url):
    soup = BeautifulSoup(html, "html.parser")
    cards = []
    seen = set()
    # Prefer explicit product IDs; fall back to links that look like item URLs.
    candidates = soup.select("[data-product-id], [data-productid]")
    if not candidates:
        candidates = []
        for link in soup.select('a[href*="/item/"]'):
            parent = link
            for _ in range(4):
                if parent.parent:
                    parent = parent.parent
            candidates.append(parent)

    for card in candidates:
        link = card.select_one('a[href*="/item/"]') if hasattr(card, "select_one") else None
        if not link and getattr(card, "name", None) == "a":
            link = card
        if not link:
            continue
        href = urljoin(source_url, link.get("href", ""))
        parsed = urlparse(href)
        canonical = f"{parsed.scheme}://{parsed.netloc}{parsed.path}" if parsed.path else href
        if not canonical or canonical in seen:
            continue
        seen.add(canonical)
        cards.append({
            "product_id": card.get("data-product-id") or card.get("data-productid"),
            "url": canonical,
            "title": first_text(card, ["[title]", "h1", "h2", "h3", "[class*=title]"]),
            "price": first_text(card, ["[class*=price]", "[class*=Price]"]),
            "rating": first_text(card, ["[class*=rating]", "[class*=Rating]"]),
            "orders": first_text(card, ["[class*=order]", "[class*=Order]"]),
        })
    return cards


def collect(keyword, max_pages=5, delay=2.0):
    session = requests.Session()
    session.headers.update(HEADERS)
    all_rows, seen_urls, run_log = [], set(), []

    for page in range(1, max_pages + 1):
        url = make_url(keyword, page)
        retrieved = datetime.now(timezone.utc).isoformat()
        try:
            response = session.get(url, timeout=30)
            html = response.text
            state = "ok"
            rows = []
            if response.status_code != 200:
                state = f"http_{response.status_code}"
            elif challenged(html) or len(html) < 10_000:
                state = "challenge_or_shell"
            else:
                rows = extract_cards(html, url)
                if not rows:
                    state = "no_cards_or_rendered"
            run_log.append({"page": page, "url": url, "retrieved_at": retrieved,
                            "status": response.status_code, "bytes": len(response.content),
                            "state": state, "cards": len(rows)})
            for row in rows:
                if row["url"] not in seen_urls:
                    row.update({"query": keyword, "page": page,
                                "retrieved_at": retrieved, "parser_version": "1.0"})
                    seen_urls.add(row["url"])
                    all_rows.append(row)
            if state in {"challenge_or_shell", "http_403", "http_429"}:
                break
            if page > 1 and not rows:
                break
        except requests.RequestException as exc:
            run_log.append({"page": page, "url": url, "retrieved_at": retrieved,
                            "state": "request_error", "error": str(exc)})
            break
        time.sleep(delay)
    return {"items": all_rows, "pages": run_log}

if __name__ == "__main__":
    result = collect("usb c hub", max_pages=3, delay=2.0)
    print(json.dumps(result, ensure_ascii=False, indent=2))

This script records a page-level status even when extraction fails. The fallback link test is only a starting point: a card may contain several links, and a broad ancestor can merge neighboring products. Validate required fields such as title and price, then tighten the card boundary after inspecting an authorized sample.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When results are embedded or rendered by JavaScript

Inspect embedded data first

View the raw response, search for product-like JSON keys, and check script elements before launching a browser. Embedded state is usually cheaper and easier to reproduce than rendering. Parse only the fields you need, validate types, and keep the original response for debugging when permission and retention rules allow it.

Use Scrapy for repeatable pipelines

Scrapy selectors support both CSS and XPath extraction, with get() for one value and getall() for many. A minimal extraction callback can look like this:

def parse(self, response):
    for card in response.css('[data-product-id], article'):
        href = card.css('a[href*="/item/"]::attr(href)').get()
        yield {
            "url": response.urljoin(href) if href else None,
            "title": card.css('h1::text, h2::text, h3::text, [class*=title]::text').get(),
            "price": card.css('[class*=price]::text, [class*=Price]::text').get(),
            "rating": card.css('[class*=rating]::text, [class*=Rating]::text').get(),
            "orders": card.css('[class*=order]::text, [class*=Order]::text').get(),
            "query": self.keyword,
            "page": self.page,
        }

Use a retry policy and scheduler, but do not retry a challenge indefinitely. Inconsistent responses can indicate target-server blocking or another target-side problem; record the evidence and stop rather than increasing pressure.

Use Playwright only for pages that need a browser

Headless rendering gives you the browser-visible result but costs more CPU, runs more slowly, and exposes the job to more challenge points. Keep it as a targeted fallback. Playwright’s Python API also permits custom selector engines through query and queryAll functions registered before page creation; ordinary CSS locators are sufficient for most first passes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
from playwright.async_api import async_playwright

async def rendered_products(url):
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto(url, wait_until="domcontentloaded", timeout=60_000)
        try:
            await page.wait_for_selector('a[href*="/item/"]', timeout=20_000)
        except Exception:
            pass
        cards = await page.locator('a[href*="/item/"]').evaluate_all("""links => links.map(a => ({
            url: a.href,
            title: a.getAttribute('title') || a.textContent.trim()
        }))""")
        await browser.close()
        return cards

# asyncio.run(rendered_products("https://www.aliexpress.com/wholesale?SearchText=usb-c-hub&page=1"))

Wait for a meaningful selector or a bounded delay, not an arbitrary long sleep. If the page still has no products, classify it as a shell, challenge, or failure and preserve that state.

Pagination, stopping, and deduplication

There are two common pagination models:

  • Page parameter: increment page from one through your configured maximum. Stop on a valid page with no cards, a repeated page signature, a challenge, or an authorization-defined limit.
  • Offset and limit: carry offset and limit, then stop at the reported total or when a page returns fewer items than requested. Use this only when the endpoint or approved API documents those parameters.

Deduplicate on a canonical product URL or product ID, not on title text. Keep the first query/page occurrence and retain every page-level log. A simple page signature can be a hash of sorted product IDs; repeated signatures often mean pagination is no longer advancing.

Choose the least complex collection method

Approach Best use Main trade-offs
Direct HTTP plus parser Static or embedded-data responses and low-volume experiments Fast and inexpensive, but fails when content is client-rendered or challenged
Scrapy selectors Repeatable crawls with structured pipelines and retries Strong extraction and scheduling model; rendering and target blocking still require handling
Playwright Results that appear only after JavaScript execution High browser fidelity, with more CPU, slower runs, and greater challenge exposure
Managed crawling API Hosted rendering, proxying, retries, and datasets Less infrastructure to operate, but adds service cost, vendor dependency, and program-term checks

For a small authorized experiment, start with direct HTTP and prove that the required fields are present. Move only the pages that need rendering to Playwright. A managed service can make operations simpler, but it does not remove the need to verify permission and partner terms.

Reliability, performance, and data quality

  • Bound work: set maximum pages, a request timeout, a browser timeout, and a total run deadline.
  • Be conservative: use a deliberate delay, avoid unnecessary parallel requests, and honor written rate limits.
  • Version parsers: store a parser version with each record so selector changes are auditable.
  • Validate fields: reject or quarantine records missing a product URL; treat missing price, rating, or orders as nullable rather than inventing values.
  • Keep raw evidence: where allowed, retain a short response sample or hash alongside the parsed record.
  • Monitor drift: alert on sudden changes in response length, card count, challenge markers, or duplicate rate.
  • Separate costs: direct requests mainly consume bandwidth and compute; browser rendering consumes substantially more CPU and wall-clock time; managed APIs add per-request or subscription charges.

Troubleshooting common failures

Symptom Likely cause Fix
HTTP success but zero products JavaScript shell, changed markup, or a challenge page Save the response, inspect scripts for embedded data, check challenge markers, then use a targeted browser fallback.
Every page repeats the same products Ignored page parameter, redirect, cached response, or blocked pagination Log the final URL, compare response hashes, verify that the page value changes, and stop if signatures repeat.
Titles exist but prices are empty Price is split across nested elements or loaded after initial HTML Inspect the card subtree, add a narrowly scoped selector, or wait for the price element in Playwright.
403, 429, CAPTCHA, or robot check Permission, rate, or anti-bot response Do not bypass it. Stop, reduce scope, confirm authorization, and use an approved API or managed route if available.
Duplicate products across pages Ranking changes, repeated cards, or noncanonical tracking URLs Canonicalize scheme, host, path, and product ID; deduplicate while retaining query/page provenance.
Browser timeout Slow resources, a stalled page, or a challenge flow Use bounded timeouts, wait for a specific selector, capture diagnostics, and classify the page instead of retrying forever.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo can capture an authorized search URL through one request, including full-page images or PDFs. Its cleanup step accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo API documentation for parameters and response details. The following call captures one search page; replace the URL with the authorized query you need:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.aliexpress.com/wholesale?SearchText=usb-c-hub&page=1 -o shot.webp

There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan, and yearly billing gives two months free.

Create a free ScreenshotNeo account to try the one-call capture.

FAQ

Should I treat a zero-result page as valid data?

Only after the response passes your challenge, shell, status, and minimum-size checks. Otherwise store it as a failed or indeterminate page state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What is the safest key for deduplication?

Use a canonical product URL or product ID. Titles and prices change and are not stable identities.

When should a browser be the default?

Use it only when authorized results are absent from HTML and embedded state. Direct parsing is easier to bound, reproduce, and monitor.

Frequently Asked Questions

Should I treat a zero-result page as valid data?

Only after the response passes your challenge, shell, status, and minimum-size checks. Otherwise store it as a failed or indeterminate page state.

What is the safest key for deduplication?

Use a canonical product URL or product ID. Titles and prices change and are not stable identities.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should a browser be the default?

Use it only when authorized results are absent from HTML and embedded state. Direct parsing is easier to bound, reproduce, and monitor.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.