Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

First check whether the job source permits your intended collection and offers an official API or partner integration. If it does, use that route rather than scraping its pages. For a permitted, server-rendered listing, Python’s requests and Beautiful Soup can fetch and parse HTML; use a browser tool such as Playwright only when the page needs JavaScript and the site allows browser automation. The example below extracts structured JobPosting data where available, writes records to CSV, and leaves site-specific selectors and pagination for you to configure rather than pretending one selector works everywhere.

Choose a permitted way to get the listings

Job postings are published on employer sites, job boards, and applicant-tracking systems. Their pages and access rules differ, so begin with the source—not with a scraper. Read its current terms and any applicable robot-exclusion rules, and confirm that your planned collection, storage, and use are allowed. A page being publicly viewable does not by itself establish permission to collect or reuse it.

Prefer an official API or partner integration when available

Indeed documents developer APIs for jobs, candidates, employers, and search integrations at Indeed’s developer documentation. Its Job Sync API is a GraphQL API for ATS partners to create, update, expire, and check the status of job postings; it is not a general-purpose permission slip to scrape Indeed pages (Indeed Job Sync API). Indeed’s developer agreement also restricts copying, redistribution, unauthorized purposes, permanent database creation, algorithmic query generation, and attempts to bypass access limits (Indeed Developer Agreement).

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LinkedIn documents a Job Posting API integration process that includes approval and vetting; it should not be read as approval for general job-page scraping (LinkedIn Job Posting API Terms). LinkedIn’s crawling terms prohibit automated crawling and indexing without express permission and require permitted crawling to follow authorized paths and robot-exclusion restrictions (LinkedIn Crawling Terms). Its recruiter help says third-party software, crawlers, bots, browser plug-ins, and scripts that scrape or automate activity are not permitted on its services (LinkedIn prohibited software guidance). Do not treat a technical workaround as authorization.

Match the tool to the page and scope

  • Requests + Beautiful Soup: A straightforward fit for a small, permitted collection when the HTML response contains the listing data. Beautiful Soup parses HTML; it does not execute page JavaScript.
  • Scrapy: Consider it when the permitted task spans many pages and benefits from request queues, retry handling, and item pipelines. It does not make a prohibited crawl permissible.
  • Playwright or Selenium: Use browser automation only if the source permits it and the content genuinely requires JavaScript rendering. Browser automation adds browser setup and resource use; it is not a reason to evade a block or access control.
  • Official API: Prefer documented fields and pagination when the source offers an API or partner route you are authorized to use.

The Python scraping reference Web Scraping with Python covers tools including Beautiful Soup, Scrapy, Selenium, and Requests. The choice among them still depends on the source’s permitted access method and how its pages are delivered.

Plan the fields before you collect

A useful record is more than a title. Decide what you need and how you will represent missing or changing values before building a crawl. Common fields include:

  • Job title and employer
  • Location, including remote or hybrid wording when the source provides it
  • Canonical posting URL or source-specific job ID
  • Description and employment type
  • Salary or compensation fields, only when shown or exposed in an authorized API
  • Publication or update time, when provided
  • Source name or listing URL, plus the time you retrieved the record

Keep the source’s original salary text or structured value where useful; normalize currency, period, and location carefully. Do not infer a salary from a range elsewhere or turn an absent value into zero. Preserve the retrieval timestamp separately from the employer’s publication date. For deduplication, use a stable source ID when available, otherwise a canonical URL; titles alone are not unique.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a conservative one-page Python scraper

This example requests one page, looks first for Schema.org JobPosting JSON-LD, and then tries CSS selectors you configure for that particular permitted site. It writes one CSV row per extracted posting. JSON-LD properties vary, and many pages do not expose them, so inspect an allowed page and adjust the fallback selectors. The script does not solve permission, authentication, JavaScript rendering, or pagination automatically.

Install the dependencies

Use Python 3 and install the two packages in your active environment:

python -m pip install requests beautifulsoup4

Save and run the script

Save as scrape_jobs.py. Set START_URL to a listing page you are authorized to collect. Set the CSS selector constants to match its markup; leave them empty if the page provides usable JobPosting JSON-LD. The example uses a descriptive user agent, a request timeout, and a delay between pages. Start with one URL.

import csv
import json
import time
from datetime import datetime, timezone
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

START_URL = "https://example.com/jobs"
OUTPUT_CSV = "jobs.csv"
REQUEST_DELAY_SECONDS = 3
MAX_PAGES = 1  # Increase only within the source's permitted scope.

# Configure these after inspecting an allowed listing page.
CARD_SELECTOR = ""
TITLE_SELECTOR = ""
EMPLOYER_SELECTOR = ""
LOCATION_SELECTOR = ""
DESCRIPTION_SELECTOR = ""
NEXT_PAGE_SELECTOR = "a[rel='next']"

HEADERS = {
    "User-Agent": "JobResearchBot/1.0 (contact: [email protected])"
}
FIELDS = [
    "job_title", "employer", "location", "description", "employment_type",
    "salary", "date_posted", "job_url", "job_id", "source_url", "retrieved_at"
]


def text_of(node):
    return " ".join(node.stripped_strings) if node else ""


def jsonld_objects(value):
    """Yield dictionaries from JSON-LD objects, including @graph and lists."""
    if isinstance(value, list):
        for item in value:
            yield from jsonld_objects(item)
    elif isinstance(value, dict):
        graph = value.get("@graph")
        if graph:
            yield from jsonld_objects(graph)
        else:
            yield value


def is_job_posting(item):
    kind = item.get("@type", [])
    kinds = kind if isinstance(kind, list) else [kind]
    return any(str(k).rsplit("/", 1)[-1] == "JobPosting" for k in kinds)


def salary_text(item):
    salary = item.get("baseSalary") or item.get("estimatedSalary") or ""
    if not salary:
        return ""
    if isinstance(salary, str):
        return salary
    if not isinstance(salary, dict):
        return json.dumps(salary, ensure_ascii=False)
    value = salary.get("value", "")
    currency = salary.get("currency", "")
    if isinstance(value, dict):
        amount = value.get("value", "")
        minimum = value.get("minValue", "")
        maximum = value.get("maxValue", "")
        unit = value.get("unitText", "")
        if minimum != "" or maximum != "":
            amount = f"{minimum}-{maximum}"
        return " ".join(str(part) for part in (currency, amount, unit) if part != "")
    return " ".join(str(part) for part in (currency, value) if part != "")


def from_jsonld(item, page_url, retrieved_at):
    employer = item.get("hiringOrganization") or ""
    if isinstance(employer, dict):
        employer = employer.get("name", "")
    location = item.get("jobLocation") or ""
    if isinstance(location, list):
        location = "; ".join(location_text(place) for place in location)
    else:
        location = location_text(location)
    identifier = item.get("identifier") or ""
    if isinstance(identifier, dict):
        identifier = identifier.get("value", "")
    url = item.get("url", "")
    return {
        "job_title": item.get("title", ""),
        "employer": employer,
        "location": location,
        "description": item.get("description", ""),
        "employment_type": item.get("employmentType", ""),
        "salary": salary_text(item),
        "date_posted": item.get("datePosted", ""),
        "job_url": urljoin(page_url, url) if url else page_url,
        "job_id": identifier,
        "source_url": page_url,
        "retrieved_at": retrieved_at,
    }


def location_text(place):
    if not isinstance(place, dict):
        return str(place)
    address = place.get("address", {})
    if isinstance(address, str):
        return address
    if not isinstance(address, dict):
        return ""
    return ", ".join(str(address.get(key, "")) for key in
                      ("addressLocality", "addressRegion", "addressCountry")
                      if address.get(key))


def parse_page(html, page_url):
    soup = BeautifulSoup(html, "html.parser")
    retrieved_at = datetime.now(timezone.utc).isoformat()
    jobs = []

    for script in soup.select('script[type="application/ld+json"]'):
        try:
            parsed = json.loads(script.string or script.get_text())
        except (json.JSONDecodeError, TypeError):
            continue
        for item in jsonld_objects(parsed):
            if is_job_posting(item):
                jobs.append(from_jsonld(item, page_url, retrieved_at))

    if jobs:
        return jobs, soup

    if not all((CARD_SELECTOR, TITLE_SELECTOR)):
        raise RuntimeError(
            "No JobPosting JSON-LD found. Configure CARD_SELECTOR and TITLE_SELECTOR "
            "for this site's permitted HTML, or use its authorized API."
        )

    for card in soup.select(CARD_SELECTOR):
        title_node = card.select_one(TITLE_SELECTOR)
        title_link = title_node.find("a", href=True) if title_node else None
        jobs.append({
            "job_title": text_of(title_node),
            "employer": text_of(card.select_one(EMPLOYER_SELECTOR)) if EMPLOYER_SELECTOR else "",
            "location": text_of(card.select_one(LOCATION_SELECTOR)) if LOCATION_SELECTOR else "",
            "description": text_of(card.select_one(DESCRIPTION_SELECTOR)) if DESCRIPTION_SELECTOR else "",
            "employment_type": "",
            "salary": "",
            "date_posted": "",
            "job_url": urljoin(page_url, title_link["href"]) if title_link else page_url,
            "job_id": "",
            "source_url": page_url,
            "retrieved_at": retrieved_at,
        })
    return jobs, soup


def main():
    session = requests.Session()
    session.headers.update(HEADERS)
    url = START_URL
    seen_urls = set()
    records = []

    for page_number in range(MAX_PAGES):
        if not url or url in seen_urls:
            break
        seen_urls.add(url)
        response = session.get(url, timeout=(5, 20))
        response.raise_for_status()
        jobs, soup = parse_page(response.text, response.url)
        records.extend(jobs)

        next_link = soup.select_one(NEXT_PAGE_SELECTOR) if NEXT_PAGE_SELECTOR else None
        url = urljoin(response.url, next_link["href"]) if next_link and next_link.get("href") else ""
        if url and page_number + 1 < MAX_PAGES:
            time.sleep(REQUEST_DELAY_SECONDS)

    # Deduplicate on a canonical posting URL where available, else source ID.
    unique = {}
    for row in records:
        key = row["job_url"] or row["job_id"]
        if key:
            unique.setdefault(key, row)

    with open(OUTPUT_CSV, "w", newline="", encoding="utf-8-sig") as file:
        writer = csv.DictWriter(file, fieldnames=FIELDS)
        writer.writeheader()
        writer.writerows(unique.values())
    print(f"Wrote {len(unique)} unique postings to {OUTPUT_CSV}")


if __name__ == "__main__":
    main()

In the script, replace [email protected] with a real contact address if you operate the collector. The HTML fallback deliberately leaves salary, employment type, and publication date blank: add selectors or documented field mappings only after confirming where the source actually provides those values. If the page has no next link, the run ends after the first page. Set MAX_PAGES conservatively and only if following those pages is allowed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Adapt the extraction without making the data less reliable

Inspect one response before scaling up

Check the returned HTML and the CSV for a single page. Look for duplicate or empty records, relative links, encoded descriptions, and whether the fields match what a person sees. JSON-LD is convenient but may be missing, incomplete, stale, or describe a different page state. A selector can also match navigation or promotional text rather than the posting. Validate a small sample against the source before collecting more.

Handle pagination and duplicates deliberately

The sample follows a link matching a[rel="next"]; many sources use a different link, a cursor, a page number, or an API-specific pagination model. Change the selector only after inspecting the permitted page structure. For API pagination, follow its documented cursor or next-page field instead of guessing page numbers. Keep a set of visited page URLs to avoid loops, and deduplicate listings with the source’s stable ID or canonical URL. A job can be reposted at a new URL, so decide whether that is a new record or an update for your use case.

Keep raw provenance and normalize conservatively

For auditability, retain the source URL and retrieval timestamp, and consider storing the raw response or a permitted subset of response metadata separately from the cleaned CSV. Normalize whitespace and location formats consistently, but keep the original value when normalization could discard meaning. Record salary currency and unit when present; a figure such as “70,000” is ambiguous without them. Treat missing values as missing, not as zero or an inferred fact.

When the page needs a browser

Requests receives the server’s HTTP response; it does not run JavaScript. If a permitted page fills its listing only after scripts run, first check for an authorized API or embedded data source. If the site permits browser automation, Playwright or Selenium can render the page before you inspect it. Configure the browser to wait for a meaningful listing selector rather than relying on a fixed long sleep, and still use a conservative rate and bounded page scope. If the source blocks the request, presents a bot check, or changes its access rules, stop rather than trying to defeat the control.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

  • HTTP 403, 429, or a bot-check page: The source is refusing or limiting access. Stop the run, review current rules, and use an authorized API or request permission; do not rotate identities or automate a challenge to get around the response.
  • Timeout or connection error: Check the URL and network, keep finite timeouts, and retry only transient failures with a bounded backoff. Repeated retries against an unresponsive host can worsen load; pause if failures continue.
  • CSV has headers but no jobs: The page may have no JobPosting JSON-LD, selectors may not match, or the content may be JavaScript-rendered. Inspect the response you actually received, then configure selectors, use an authorized API, or use permitted browser rendering.
  • Wrong title or employer repeated on every row: The selectors are too broad or are targeting a shared page element. Narrow them to each job card and inspect a few extracted records before increasing scope.
  • Descriptions contain markup or odd whitespace: JSON-LD descriptions may contain HTML. If you clean them, strip tags with an HTML parser and preserve text without changing the source meaning; keep raw values if you need traceability.
  • Duplicate or missing postings: Use a source ID or canonical job URL for deduplication, check pagination behavior, and distinguish a repeated page from a reposted job. Do not use a title as the sole key.

Performance, reliability, and cost

For a small collection, a sequential script with one request at a time is easier to inspect and less likely to overload a site than parallel requests. Add retries only for temporary network failures, limit their count, and avoid treating access denials as retryable. Scrapy can organize larger permitted crawls with queues, retries, and item pipelines, but more throughput increases operational responsibility and does not improve permission or data quality.

HTML scraping depends on markup that can change without notice; an API’s documented schema is generally a more stable contract when you have access to it. Monitor status codes, missing-field rates, duplicate rates, and extraction errors. If those change abruptly, pause and inspect before writing more data. Store only what you need, protect any personal or sensitive information, and check source terms for retention and redistribution restrictions. There is no universal cost or success rate: hosting, browser execution, storage, API access, and maintenance depend on your implementation and source.

Or skip the browser setup

For a permitted page that you need to save visually rather than extract into job fields, ScreenshotNeo is a screenshot API and MCP server. It is not a job-posting scraper and its image output does not replace structured extraction. It can capture the page without configuring a local browser; its cleanup accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture, with each cleanup step switchable. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can use its MCP server tools, including take_screenshot, get_page_info, and capture_pdf.

Python example (save the response body as a WebP screenshot):

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/jobs"},
    timeout=90,
)
open("jobs-page.webp", "wb").write(r.content)

See the ScreenshotNeo API documentation for request options and response details. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Frequently Asked Questions

Can I scrape job postings from a page behind a login?

Do not assume that being able to log in grants permission to automate collection. Check the source’s terms and obtain explicit authorization for the intended access and use; do not use the script to bypass authentication or access controls.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.