Scrape an AliExpress search page as a bounded data-collection job: build the public keyword URL, fetch each page conservatively, extract repeated product-card fields, detect challenge or empty responses, paginate only within a hard limit, and deduplicate by product URL or ID. If the HTML is only a JavaScript shell, inspect embedded JSON or use a headless browser for those pages. For production or commercial collection, obtain written permission or use an approved API or managed crawler first.
Start with permission and a precise data contract
AliExpress Terms of Use state: “Systematic retrieval of Site Content from the Sites to create or compile, directly or indirectly, a collection, compilation, database or directory (whether through robots, spiders, automatic devices or manual processes) without written permission from AliExpress.com is prohibited.” The same terms restrict copying, downloading, republishing, selling, or commercially exploiting site content. Treat this as a permission boundary, not a prompt to bypass anti-bot controls. Review applicable law, robots guidance, rate limits, and any written authorization before a production run.
Write down exactly what the collector is allowed to store. A useful product record contains:
- Identity: canonical product URL and product ID when available.
- Display fields: title, price text, rating text, and order-count text.
- Provenance: original keyword, page number, retrieval timestamp, HTTP status, response length, and parser version.
- Run controls: maximum pages, delay policy, and the reason the run stopped.
Keeping query and page metadata makes later changes explainable. It also prevents a blank challenge page from being mistaken for a valid page with zero products.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Construct a public search URL and bounded pagination
The wholesale-style search route uses a hyphenated keyword and a page query parameter. Normalize whitespace, create one URL per page, and keep a hard maximum even if the site appears to offer more results.
from urllib.parse import urlencode
import re
def search_url(keyword: str, page: int) -> str:
slug = re.sub(r"s+", "-", keyword.strip())
query = urlencode({"SearchText": slug, "page": page})
return f"https://www.aliexpress.com/wholesale?{query}"
print(search_url("usb c hub", 1))
Do not assume that a successful HTTP response means a successful scrape. Stop or quarantine a page when it contains a CAPTCHA, bot-check language, an interstitial challenge, an unexpectedly tiny body, or an abrupt item-count change. A normal empty result and a blocked response are different outcomes and should be recorded separately.
A conservative Python collector
The following example uses requests and Beautiful Soup. AliExpress markup changes, so the selectors are deliberately defensive rather than tied to one private class name. Inspect a current, authorized response and adjust the candidate selectors for your run.
import json
import re
import time
from datetime import datetime, timezone
from urllib.parse import urlencode, urljoin, urlparse
import requests
from bs4 import BeautifulSoup
BASE = "https://www.aliexpress.com/wholesale"
HEADERS = {
"User-Agent": "authorized-research-bot/1.0 ([email protected])",
"Accept-Language": "en-US,en;q=0.8",
}
CHALLENGE_WORDS = (
"captcha", "robot check", "verify you are human", "security verification",
"access denied", "unusual traffic"
)
def make_url(keyword, page):
slug = re.sub(r"s+", "-", keyword.strip())
return f"{BASE}?{urlencode({'SearchText': slug, 'page': page})}"
def challenged(text):
sample = text[:200_000].lower()
return any(word in sample for word in CHALLENGE_WORDS)
def first_text(node, selectors):
for selector in selectors:
found = node.select_one(selector)
if found:
value = found.get("content") or found.get_text(" ", strip=True)
if value:
return value
return None
def extract_cards(html, source_url):
soup = BeautifulSoup(html, "html.parser")
cards = []
seen = set()
# Prefer explicit product IDs; fall back to links that look like item URLs.
candidates = soup.select("[data-product-id], [data-productid]")
if not candidates:
candidates = []
for link in soup.select('a[href*="/item/"]'):
parent = link
for _ in range(4):
if parent.parent:
parent = parent.parent
candidates.append(parent)
for card in candidates:
link = card.select_one('a[href*="/item/"]') if hasattr(card, "select_one") else None
if not link and getattr(card, "name", None) == "a":
link = card
if not link:
continue
href = urljoin(source_url, link.get("href", ""))
parsed = urlparse(href)
canonical = f"{parsed.scheme}://{parsed.netloc}{parsed.path}" if parsed.path else href
if not canonical or canonical in seen:
continue
seen.add(canonical)
cards.append({
"product_id": card.get("data-product-id") or card.get("data-productid"),
"url": canonical,
"title": first_text(card, ["[title]", "h1", "h2", "h3", "[class*=title]"]),
"price": first_text(card, ["[class*=price]", "[class*=Price]"]),
"rating": first_text(card, ["[class*=rating]", "[class*=Rating]"]),
"orders": first_text(card, ["[class*=order]", "[class*=Order]"]),
})
return cards
def collect(keyword, max_pages=5, delay=2.0):
session = requests.Session()
session.headers.update(HEADERS)
all_rows, seen_urls, run_log = [], set(), []
for page in range(1, max_pages + 1):
url = make_url(keyword, page)
retrieved = datetime.now(timezone.utc).isoformat()
try:
response = session.get(url, timeout=30)
html = response.text
state = "ok"
rows = []
if response.status_code != 200:
state = f"http_{response.status_code}"
elif challenged(html) or len(html) < 10_000:
state = "challenge_or_shell"
else:
rows = extract_cards(html, url)
if not rows:
state = "no_cards_or_rendered"
run_log.append({"page": page, "url": url, "retrieved_at": retrieved,
"status": response.status_code, "bytes": len(response.content),
"state": state, "cards": len(rows)})
for row in rows:
if row["url"] not in seen_urls:
row.update({"query": keyword, "page": page,
"retrieved_at": retrieved, "parser_version": "1.0"})
seen_urls.add(row["url"])
all_rows.append(row)
if state in {"challenge_or_shell", "http_403", "http_429"}:
break
if page > 1 and not rows:
break
except requests.RequestException as exc:
run_log.append({"page": page, "url": url, "retrieved_at": retrieved,
"state": "request_error", "error": str(exc)})
break
time.sleep(delay)
return {"items": all_rows, "pages": run_log}
if __name__ == "__main__":
result = collect("usb c hub", max_pages=3, delay=2.0)
print(json.dumps(result, ensure_ascii=False, indent=2))
This script records a page-level status even when extraction fails. The fallback link test is only a starting point: a card may contain several links, and a broad ancestor can merge neighboring products. Validate required fields such as title and price, then tighten the card boundary after inspecting an authorized sample.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsWhen results are embedded or rendered by JavaScript
Inspect embedded data first
View the raw response, search for product-like JSON keys, and check script elements before launching a browser. Embedded state is usually cheaper and easier to reproduce than rendering. Parse only the fields you need, validate types, and keep the original response for debugging when permission and retention rules allow it.
Use Scrapy for repeatable pipelines
Scrapy selectors support both CSS and XPath extraction, with get() for one value and getall() for many. A minimal extraction callback can look like this:
def parse(self, response):
for card in response.css('[data-product-id], article'):
href = card.css('a[href*="/item/"]::attr(href)').get()
yield {
"url": response.urljoin(href) if href else None,
"title": card.css('h1::text, h2::text, h3::text, [class*=title]::text').get(),
"price": card.css('[class*=price]::text, [class*=Price]::text').get(),
"rating": card.css('[class*=rating]::text, [class*=Rating]::text').get(),
"orders": card.css('[class*=order]::text, [class*=Order]::text').get(),
"query": self.keyword,
"page": self.page,
}
Use a retry policy and scheduler, but do not retry a challenge indefinitely. Inconsistent responses can indicate target-server blocking or another target-side problem; record the evidence and stop rather than increasing pressure.
Use Playwright only for pages that need a browser
Headless rendering gives you the browser-visible result but costs more CPU, runs more slowly, and exposes the job to more challenge points. Keep it as a targeted fallback. Playwright’s Python API also permits custom selector engines through query and queryAll functions registered before page creation; ordinary CSS locators are sufficient for most first passes.
Rank #3
import asyncio
from playwright.async_api import async_playwright
async def rendered_products(url):
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto(url, wait_until="domcontentloaded", timeout=60_000)
try:
await page.wait_for_selector('a[href*="/item/"]', timeout=20_000)
except Exception:
pass
cards = await page.locator('a[href*="/item/"]').evaluate_all("""links => links.map(a => ({
url: a.href,
title: a.getAttribute('title') || a.textContent.trim()
}))""")
await browser.close()
return cards
# asyncio.run(rendered_products("https://www.aliexpress.com/wholesale?SearchText=usb-c-hub&page=1"))
Wait for a meaningful selector or a bounded delay, not an arbitrary long sleep. If the page still has no products, classify it as a shell, challenge, or failure and preserve that state.
Pagination, stopping, and deduplication
There are two common pagination models:
- Page parameter: increment
pagefrom one through your configured maximum. Stop on a valid page with no cards, a repeated page signature, a challenge, or an authorization-defined limit. - Offset and limit: carry
offsetandlimit, then stop at the reportedtotalor when a page returns fewer items than requested. Use this only when the endpoint or approved API documents those parameters.
Deduplicate on a canonical product URL or product ID, not on title text. Keep the first query/page occurrence and retain every page-level log. A simple page signature can be a hash of sorted product IDs; repeated signatures often mean pagination is no longer advancing.
Choose the least complex collection method
| Approach | Best use | Main trade-offs |
|---|---|---|
| Direct HTTP plus parser | Static or embedded-data responses and low-volume experiments | Fast and inexpensive, but fails when content is client-rendered or challenged |
| Scrapy selectors | Repeatable crawls with structured pipelines and retries | Strong extraction and scheduling model; rendering and target blocking still require handling |
| Playwright | Results that appear only after JavaScript execution | High browser fidelity, with more CPU, slower runs, and greater challenge exposure |
| Managed crawling API | Hosted rendering, proxying, retries, and datasets | Less infrastructure to operate, but adds service cost, vendor dependency, and program-term checks |
For a small authorized experiment, start with direct HTTP and prove that the required fields are present. Move only the pages that need rendering to Playwright. A managed service can make operations simpler, but it does not remove the need to verify permission and partner terms.
Reliability, performance, and data quality
- Bound work: set maximum pages, a request timeout, a browser timeout, and a total run deadline.
- Be conservative: use a deliberate delay, avoid unnecessary parallel requests, and honor written rate limits.
- Version parsers: store a parser version with each record so selector changes are auditable.
- Validate fields: reject or quarantine records missing a product URL; treat missing price, rating, or orders as nullable rather than inventing values.
- Keep raw evidence: where allowed, retain a short response sample or hash alongside the parsed record.
- Monitor drift: alert on sudden changes in response length, card count, challenge markers, or duplicate rate.
- Separate costs: direct requests mainly consume bandwidth and compute; browser rendering consumes substantially more CPU and wall-clock time; managed APIs add per-request or subscription charges.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| HTTP success but zero products | JavaScript shell, changed markup, or a challenge page | Save the response, inspect scripts for embedded data, check challenge markers, then use a targeted browser fallback. |
| Every page repeats the same products | Ignored page parameter, redirect, cached response, or blocked pagination | Log the final URL, compare response hashes, verify that the page value changes, and stop if signatures repeat. |
| Titles exist but prices are empty | Price is split across nested elements or loaded after initial HTML | Inspect the card subtree, add a narrowly scoped selector, or wait for the price element in Playwright. |
| 403, 429, CAPTCHA, or robot check | Permission, rate, or anti-bot response | Do not bypass it. Stop, reduce scope, confirm authorization, and use an approved API or managed route if available. |
| Duplicate products across pages | Ranking changes, repeated cards, or noncanonical tracking URLs | Canonicalize scheme, host, path, and product ID; deduplicate while retaining query/page provenance. |
| Browser timeout | Slow resources, a stalled page, or a challenge flow | Use bounded timeouts, wait for a specific selector, capture diagnostics, and classify the page instead of retrying forever. |
Or skip the browser setup
ScreenshotNeo can capture an authorized search URL through one request, including full-page images or PDFs. Its cleanup step accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for parameters and response details. The following call captures one search page; replace the URL with the authorized query you need:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.aliexpress.com/wholesale?SearchText=usb-c-hub&page=1 -o shot.webp
There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan, and yearly billing gives two months free.
Create a free ScreenshotNeo account to try the one-call capture.
FAQ
Should I treat a zero-result page as valid data?
Only after the response passes your challenge, shell, status, and minimum-size checks. Otherwise store it as a failed or indeterminate page state.
What is the safest key for deduplication?
Use a canonical product URL or product ID. Titles and prices change and are not stable identities.
Best Value
When should a browser be the default?
Use it only when authorized results are absent from HTML and embedded state. Direct parsing is easier to bound, reproduce, and monitor.
Frequently Asked Questions
Should I treat a zero-result page as valid data?
Only after the response passes your challenge, shell, status, and minimum-size checks. Otherwise store it as a failed or indeterminate page state.
What is the safest key for deduplication?
Use a canonical product URL or product ID. Titles and prices change and are not stable identities.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →When should a browser be the default?
Use it only when authorized results are absent from HTML and embedded state. Direct parsing is easier to bound, reproduce, and monitor.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




