Start with access, not code. If you need Udemy Business catalog metadata and your organization is an eligible customer or partner, use Udemy’s documented GraphQL Courses API and Search API. If you manage courses you teach, the authenticated Instructor API may be the correct route. Only when an authorized public-page workflow requires data that is missing from the initial HTML should you render the page with JavaScript, such as Puppeteer. Udemy’s current terms and your account agreement determine what access is permitted; the available documentation does not establish blanket permission to scrape the public marketplace.
Choose the data route before writing a scraper
Define the fields and purpose first. Typical catalog fields include a course title, public URL, rating, review count and instructor name. Do not collect learner-specific or account data unless your integration is explicitly authorized to do so.
| Route | Best fit | Access and evidence | Important limitation |
|---|---|---|---|
| Udemy Business GraphQL Courses API and Search API | Catalog metadata for an eligible Business integration | Udemy documents course metadata queries and search for Business customers, partners and provisioned organizations. | Login, subscription, credentials and an organization-specific agreement may be required. It is not an anonymous public-marketplace endpoint. |
| Udemy Instructor API v1 | Instructor-owned or taught-course workflows | Authenticated REST API over HTTPS with JSON responses, pagination and documented throttling. | It is not an open API for arbitrary courses. |
| Browser rendering with Puppeteer | A permitted page whose required fields appear only after JavaScript executes | Puppeteer is a relevant Node.js scraping tool and Udemy course instruction recommends API-first access. | No current Udemy selector, endpoint, rendering behavior or successful scrape is established here. |
Business catalog access
Udemy describes its GraphQL Courses API as “the next generation and evolution to the traditional courses API.” Its support material also says the legacy Courses API is one for which “we will not be releasing any new functionality.” Treat those statements as product documentation, not as authorization to use an endpoint outside your agreement. Ask your Udemy Business administrator or partner contact which API, scopes and environments are enabled.
Instructor workflows
The Instructor API reference documents API-client bearer authentication, HTTPS, JSON, pagination and a throttle of 100 requests per 10 seconds. That limit applies to the documented Instructor API, not automatically to every Udemy API. The documented Course model includes title, URL, rating, number of reviews, publication time and visible instructors. Keep bearer tokens on your server, never in browser JavaScript or a public repository.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors#1 Best Overall
Do not use obsolete Affiliate API examples
Udemy’s Affiliate API v2 reference states: “Access to the Affiliate API on Udemy has been discontinued since 1/1/2025.” Do not copy old Affiliate API endpoints into a new integration. A current affiliate-program process, if available to you, is separate from that discontinued API and its terms were not established by the available documentation.
Confirm permission for public pages
Before extracting a public course page, read the current Udemy terms, robots guidance and any agreement governing your account. The available sources do not answer whether a particular public-page scrape is allowed. Use only access you are authorized to use, keep collection to the minimum fields needed, and provide a way to stop the job if Udemy or the rights holder objects.
- Write down the fields, URLs and business purpose.
- Exclude login-only, learner-specific and personal data unless your authorization explicitly covers it.
- Use conservative request rates, caching and a small validation sample.
- Record retrieval timestamps because ratings, prices, instructors and publication status can change.
Inspect the normal response before launching a browser
JavaScript rendering is a fallback, not a default. Fetch one authorized URL with a normal HTTP client and inspect the response body. Look for the title, structured data and other fields your use case needs. Also inspect the rendered DOM manually in a browser. If the required value is already in the response, a browser adds cost and failure modes without adding data.
Rank #2
If a value is absent from the initial response but appears after scripts run, browser automation is justified. The relevant course guidance says: “Always check for a public API before web scraping, then use a request to fetch JSON data; only resort to automated browsers like Puppeteer as a last option.” That is instructional advice, not a statement about Udemy’s current page architecture.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Render an authorized page with Puppeteer
The example below is deliberately selector-neutral. Because current Udemy markup was not verified, replace the sample selectors with ones you confirm in your target page and test for missing values. It records the rendered HTML and extracts values from visible text or structured data without claiming a particular Udemy selector.
Install Node.js dependencies
mkdir udemy-renderer
cd udemy-renderer
npm init -y
npm install puppeteer
Complete JavaScript example
const puppeteer = require('puppeteer');
const target = process.argv[2];
if (!target) {
console.error('Usage: node scrape.js https://example.test/course');
process.exit(1);
}
(async () => {
const browser = await puppeteer.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox']
});
try {
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
await page.setUserAgent('AuthorizedCatalogClient/1.0');
const response = await page.goto(target, {
waitUntil: 'domcontentloaded',
timeout: 45000
});
if (!response) throw new Error('Navigation returned no response');
await page.waitForFunction(() => document.body && document.body.innerText.trim().length > 0,
{ timeout: 15000 });
const result = await page.evaluate(() => {
const text = (selector) => {
const node = document.querySelector(selector);
return node ? node.textContent.trim() : null;
};
const jsonLd = [...document.querySelectorAll('script[type="application/ld+json"]')]
.map(node => { try { return JSON.parse(node.textContent); } catch { return null; } })
.filter(Boolean);
return {
title: document.title || null,
canonical: document.querySelector('link[rel="canonical"]')?.href || location.href,
bodyTextSample: document.body.innerText.slice(0, 2000),
jsonLd,
// Replace these with selectors verified on your permitted target page.
courseTitle: text('[data-course-title]'),
rating: text('[data-course-rating]'),
reviewCount: text('[data-review-count]'),
instructor: text('[data-instructor]')
};
});
console.log(JSON.stringify({
status: response.status(),
url: page.url(),
data: result
}, null, 2));
} finally {
await browser.close();
}
})();
Run it with node scrape.js https://your-authorized-url.example/course. A successful navigation status does not prove that the fields exist. Treat null values as a schema-change signal, not as permission to guess selectors.
Make waits content-based
A fixed sleep can be too short on a slow page and wasteful on a fast one. Prefer a condition tied to the field you need, for example page.waitForSelector('[data-course-title]', {timeout: 15000}) after you have verified that selector. If the page uses a different structure, wait for a stable text condition or a known application state. Do not wait indefinitely.
Handle pagination and volume
For API routes, follow the documented pagination fields and stop when the service indicates there are no more results. For page rendering, queue a small number of URLs, reuse a browser process, open a fresh page per job, and cap concurrency. Cache authorized results with a timestamp and retry only transient navigation failures. Never try to defeat a CAPTCHA, bot check, login wall or access control.
Validation, reliability and data quality
- Compare with the visible page: Check a small sample manually and confirm that title, rating, review count and instructor values match what an ordinary visitor can see.
- Preserve provenance: Store URL, retrieval time, HTTP status and parser version beside each record.
- Expect missing fields: Courses can have no visible rating, no review count or multiple instructors. Represent absent values as null rather than zero or an invented string.
- Detect changes: Alert when a previously populated field becomes null, when the page title changes unexpectedly or when response status patterns shift.
- Protect credentials: Keep API tokens in a secret manager or environment variables, restrict logs, and use HTTPS.
- Respect limits: The 100-requests-per-10-seconds figure belongs to the documented Instructor API. Do not apply it to GraphQL, Search or public-page requests without current documentation.
Troubleshooting
The API is unavailable or returns an authorization error
Confirm that you are using the API associated with your account type. Business catalog APIs require the relevant Business provisioning and agreement; the Instructor API requires supported instructor credentials and scopes. Check the current reference and ask your administrator rather than switching to an undocumented endpoint.
Rank #4
The browser returns an empty page
Check the HTTP status, final URL, redirects and response body. A login wall, consent interstitial, bot check, timeout or blocked resource can all produce an apparently blank result. Do not automate a challenge or bypass access controls. Fix authorization or stop the job.
Fields are null
First determine whether the value is genuinely absent, loaded later, available only to a signed-in user, or represented in structured data rather than visible text. Inspect the rendered DOM and JSON-LD on an authorized sample, then update selectors with tests and a fallback parser.
Navigation times out
Use a bounded timeout, capture diagnostics such as final URL and status, retry transient failures with backoff, and reduce concurrency. A timeout is not evidence that a longer delay will solve the problem.
Best Value
Results change between runs
Course metadata is mutable. Store retrieval timestamps, avoid treating ratings or review counts as permanent identifiers, and record the raw value used for each export.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
For screenshots rather than structured course-field extraction, ScreenshotNeo provides a single-call API and an MCP server for AI agents. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; bot checks, blank pages, failed loads and cache hits are not billed, with the result identified by response headers. Its MCP tools include take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for parameters. A free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
When each approach is the right one
- Choose Business GraphQL/Search when you need catalog metadata and your organization is provisioned.
- Choose the Instructor API for supported instructor-owned workflows.
- Choose Puppeteer only when an authorized page truly requires JavaScript execution and no suitable API supplies the fields.
- Choose a screenshot service when the deliverable is a visual capture, not a structured catalog dataset.
Frequently Asked Questions
Can Puppeteer access every Udemy course?
No. Puppeteer only automates a browser; it does not grant permission, bypass authentication or make restricted data available. Your authorization and current Udemy terms still control access.
Recommended Free Tools
Is the Instructor API a replacement for the Business catalog APIs?
No. The Instructor API is documented for instructor workflows, while the Business GraphQL Courses API and Search API target eligible Business catalog integrations.
What should I do if a course has no rating or instructor value?
Store a null or equivalent missing value, preserve the retrieval timestamp, and avoid converting absence into zero or an inferred value.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




