October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AI job scraper

How to Scrape Job Postings with an AI Job Board Scraper (Legally and Reliably)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The safe way to scrape job postings with AI is to start with permission, not a crawler. Use an official API, partner feed, publisher plug-in, or an employer career page whose terms allow collection. Then keep the permitted records, normalize them into a fixed schema, let AI extract only what the text actually says, and retain evidence and confidence for review. A page being publicly visible is not blanket permission to copy it.

Start with authorization and a defined purpose

Before writing a scraper, document why you need the data and who has authorized it. Your source record should include:

  • Platform or employer site and the exact API, feed, plug-in, or page scope.
  • The account, client, or contract that grants access.
  • Permitted fields, geographic scope, request limits, retention period, and deletion contact.
  • Whether you may display, transform, redistribute, or create a persistent database from the records.

Prefer an official integration. Indeed documents APIs for jobs, candidates, and employers, a Publisher JavaScript Plugin, a Partner Console, and a partner-application path. Its Developer Agreement says API access is granted only after acceptance of the relevant documentation. The agreement also restricts copying or creating permanent databases of user or job-seeker content unless expressly permitted, algorithmic queries that replace human input, bypassing limits, and using the APIs to build a competing product. Treat those provisions as engineering requirements: request the correct scope, minimize fields, obey quotas, and obtain written clarification before storing or redistributing data.

LinkedIn’s Recruiter Help says it does not permit third-party software such as crawlers, bots, browser plug-ins, or browser extensions that scrape, copy, or automate activity on LinkedIn. Its Job Posting API terms require developer and application vetting, client authorization, data-rights and privacy compliance, security controls, and deletion of certain stored data. Microsoft’s current API overview says LinkedIn is not accepting new partnerships for that Job Posting API and directs applicants to Apply Connect. An unaffiliated scraper should therefore not crawl LinkedIn pages; use an approved access route or another source.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
San Jamar 1178902 Cutting Board Scraper/Refinisher Tool
  • CUTTING BOARD REFINISHER: The San Jamar Cutting Board Refinisher helps remove grooves and cuts from plastic cutting board surfaces
  • REPLACEABLE BLADES: Easy-to-replace blades allow you to use this refinisher countless times to keep cutting boards in top shape
  • HARD TO BREAK: The refinisher is durable and designed to reduce breakage, allowing it to withstand even the toughest grooves
  • SECURE GRIP: Grooved handle provides essential grip while refinishing
  • REDUCES BOARD REPLACEMENT: Helps minimize the need for frequent cutting board replacements through continued refinishing

Do not confuse public visibility with permission

Terms of service, API agreements, privacy law, copyright, database rights, and contract terms can all apply to a publicly reachable page. Robots directives and rate limits are useful operational signals, but they do not replace authorization. If the source owner cannot explain your permitted use, stop and ask for written approval.

Design the collection pipeline

A durable AI job-board scraper separates collection, parsing, enrichment, and governance. The following sequence prevents an extraction model from silently turning an uncertain page into a fact.

  1. Choose the source and legal basis. Record the authorization, fields, territory, quota, retention rule, and deletion process before the first request.
  2. Collect only permitted records. Use the approved API or feed. For a first-party career page, check its terms, robots directives, rate limit, and whether the employer authorized reuse. Save the original URL or API identifier and retrieval timestamp.
  3. Normalize into a stable schema. Map different source names into consistent fields while preserving the raw permitted payload separately when storage is allowed.
  4. Parse deterministic values first. Rules are more reliable for dates, URLs, salary ranges, locations, identifiers, and obvious employment types.
  5. Apply AI to ambiguous language. Use a model for skill normalization, seniority classification, entity extraction, duplicate-similarity suggestions, and natural-language search. Require a source span and confidence for every extracted value.
  6. Validate and review. Reject records without a canonical URL or employer. Flag contradictory compensation or location values and route low-confidence or legally sensitive cases to a human.
  7. Deduplicate and expire. Prefer a stable source ID. Otherwise combine canonical URL, employer, title, location, and posting date. Re-check freshness and remove or mark expired records according to the source’s retention rules.
  8. Protect the data. Encrypt credentials and stored data, restrict staff access, log API calls, honor deletion requests, and keep candidate or member data out of the dataset unless the agreement explicitly permits it.
  9. Monitor change. Track parser failures, schema drift, HTTP errors, quota use, duplicate rate, extraction confidence, and deletion SLA. Pause a source when its terms or API status changes.

Use a schema that preserves evidence

A fixed schema makes feeds comparable without erasing what the employer actually wrote. Keep normalized values beside the original text and the span that supports each value.

Field What to store Validation rule
source Platform or employer feed name Required; must match an authorized source.
source_id Stable API or publisher identifier Preferred deduplication key.
canonical_url Original job or application URL Required; preserve retrieval timestamp.
title Job title as posted Do not let AI rewrite the source title.
employer Legal or displayed employer name Required; flag missing or conflicting values.
location City, region, country, or “remote” text Keep the raw phrase and normalized geography.
remote_status On-site, hybrid, remote, or not stated “Not stated” is safer than an inference.
employment_type Full-time, part-time, contract, internship, and so on Store source wording when categories are unclear.
compensation_text Verbatim salary or benefits text Never discard currency, period, or qualification.
salary_min, salary_max, salary_period Parsed range and annual, monthly, hourly, or other period Flag impossible ranges and missing currency.
skills Normalized skills plus supporting text spans Do not add skills absent from the posting.
seniority Entry, mid, senior, lead, executive, or not stated Store the phrase that justified the class.
posted_at, retrieved_at Source date and your retrieval time Keep timezone and distinguish “updated” from “posted.”
expiry_status Active, expired, unknown, or removed Set through a re-check, not an AI guess.

Deterministic parsing before AI enrichment

Use regular expressions and controlled vocabularies for values with a clear surface form. Parse a salary range into minimum, maximum, currency, and period while retaining the original sentence. Canonicalize URLs by removing only parameters your agreement allows you to remove. Normalize obvious location abbreviations with a maintained geography table. Keep “salary not listed,” “competitive,” and “up to $X” distinct instead of forcing them into a numeric range.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Then ask an AI model to classify or extract ambiguous content. Give it the job text, your allowed labels, and a strict output schema. Require each field to contain value, confidence, and evidence (an exact source span). If the model cannot find evidence, it must return not stated. Store the model name and version, prompt version, timestamp, and any reviewer correction.

AI tasks that fit this use case

  • Map “distributed,” “home-based,” and similar wording to a remote-status label only when the posting supports it.
  • Normalize “Java Script” and “JS” to a controlled skill vocabulary while retaining the original phrase.
  • Classify seniority from explicit phrases such as “staff,” “principal,” or “junior,” with “not stated” for ambiguous copy.
  • Suggest possible duplicates for human confirmation; do not delete records solely on a similarity score.
  • Answer search questions over your normalized, authorized dataset without inventing missing fields.

AI uses to avoid

Do not infer protected traits, personality, health, age, ethnicity, gender, or criminal history. Do not use an extraction model as a hiring or candidate-ranking decision maker. Indeed’s AI and Automated Employment Decision Tools FAQ identifies discrimination, systems that infringe legal rights, biometric identification without consent, criminal-offense prediction, and exploitation of vulnerabilities among prohibited practices. Keep job-posting extraction separate from candidate evaluation unless a documented, legally reviewed process exists.

A minimal authorized-feed implementation in Python

The script below demonstrates collection and deterministic normalization from a feed you are authorized to call. It writes newline-delimited JSON and preserves the original record. Install requests, set AUTHORIZED_FEED_URL, and provide the feed token only if the agreement requires one.

import json
import os
import re
from datetime import datetime, timezone
from urllib.parse import urlparse

import requests

FEED_URL = os.environ["AUTHORIZED_FEED_URL"]
TOKEN = os.environ.get("AUTHORIZED_FEED_TOKEN")
OUT = os.environ.get("OUT", "jobs.ndjson")

salary_re = re.compile(
    r"(?P[$€£])s*(?P[d,]+)(?:s*[-–]s*(?P[d,]+))?s*(?Pper year|annually|/year|per hour|/hour)?",
    re.I,
)

def first(record, *names):
    for name in names:
        value = record.get(name)
        if value not in (None, ""):
            return value
    return None

def parse_salary(text):
    if not text:
        return None
    match = salary_re.search(text)
    if not match:
        return None
    clean = lambda value: int(value.replace(",", "")) if value else None
    return {
        "currency": match.group("currency"),
        "min": clean(match.group("min")),
        "max": clean(match.group("max")) or clean(match.group("min")),
        "period": (match.group("period") or "not stated").lower(),
        "evidence": match.group(0),
    }

def normalize(raw):
    description = first(raw, "description", "body", "job_description") or ""
    url = first(raw, "canonical_url", "url", "apply_url")
    parsed = urlparse(url) if url else None
    return {
        "source": first(raw, "source") or FEED_URL,
        "source_id": first(raw, "source_id", "id", "job_id"),
        "canonical_url": url,
        "employer": first(raw, "employer", "company", "company_name"),
        "title": first(raw, "title", "job_title"),
        "location": first(raw, "location", "job_location"),
        "remote_status": first(raw, "remote_status", "workplace_type") or "not stated",
        "employment_type": first(raw, "employment_type", "type") or "not stated",
        "compensation_text": first(raw, "salary", "compensation"),
        "salary": parse_salary(first(raw, "salary", "compensation")),
        "posted_at": first(raw, "posted_at", "date_posted", "published_at"),
        "retrieved_at": datetime.now(timezone.utc).isoformat(),
        "description": description,
        "host": parsed.netloc if parsed else None,
        "raw_permitted_record": raw,
    }

headers = {"Accept": "application/json"}
if TOKEN:
    headers["Authorization"] = f"Bearer {TOKEN}"
response = requests.get(FEED_URL, headers=headers, timeout=30)
response.raise_for_status()
payload = response.json()
records = payload if isinstance(payload, list) else payload.get("jobs", [])

with open(OUT, "w", encoding="utf-8") as handle:
    for raw in records:
        job = normalize(raw)
        if not job["canonical_url"] or not job["employer"]:
            continue
        handle.write(json.dumps(job, ensure_ascii=False) + "n")

print(f"Wrote {OUT}")

Add your approved AI provider after this step by sending only the permitted description and schema, then merging the returned values with their evidence spans and confidence. Do not place credentials in the prompt or commit them to source control.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Freshness, deduplication, and operational controls

Deduplication

Use the source ID whenever it is stable. If it is absent, compare canonical URL, employer, title, normalized location, and posting date. Similarity should create a review suggestion, not an automatic deletion, because one employer may legitimately publish several requisitions with nearly identical text.

Expiry

Re-check according to the source’s permitted cadence. A removed API record, an explicit closing date, or a page that now reports no vacancy is stronger evidence than an old retrieval timestamp. Mark unknown when you cannot verify status, and delete or retain it only as the agreement allows.

Performance and reliability

  • Use the source’s quota and backoff guidance; never parallelize past its stated limit.
  • Cache permitted responses for the allowed TTL and use stable IDs to avoid repeat work.
  • Keep collection and AI enrichment in separate queues so a model outage does not cause duplicate API calls.
  • Record HTTP status, latency, response size, parser version, and quota remaining.
  • Use idempotent writes keyed by source and source ID, with retries only for transient failures.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When an employer has authorized visual collection from a first-party career page, ScreenshotNeo can provide a screenshot or PDF through one request. It is not a way around a platform’s terms or a substitute for an approved job API. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed. Its MCP server gives Claude, Cursor, and other MCP clients the tools take_screenshot, get_page_info, and capture_pdf.

Relevant controls include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/orientation/page ranges, custom CSS and JavaScript, click-before-capture, selector hiding, waits for a selector, delay, or network idle, ad/tracker/request/resource blocking, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, a chosen cache TTL, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL (see the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every feature is on every plan: Free includes 1,000 shots per month with no card; Starter is $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; and Business $249 for 1,000,000. Yearly billing gives two months free.

Create a free ScreenshotNeo account to use the 1,000 monthly screenshots without a card.

Troubleshooting common failures

HTTP 401 or 403 from a feed

Usually the token, scope, account authorization, or agreement is wrong. Confirm the approved endpoint and required Authorization format with the provider; do not rotate through accounts or attempt to bypass the response.

HTTP 429 or quota exhaustion

Reduce concurrency, honor the documented retry interval, cache permitted responses, and request a higher approved quota. Never evade limits with proxy pools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Missing salary or location

The posting may not state it, or the feed may expose a different field. Keep the raw compensation text, return “not stated” when evidence is absent, and send contradictory values to review.

Duplicate jobs after an update

Prefer the source ID. If none exists, canonicalize the URL and compare employer, title, location, and posting date. Keep separate records when the requisition IDs differ.

AI output contains an invented skill

Reject any value without an exact evidence span. Lower the model’s allowed vocabulary, require structured output, and route low-confidence records to a reviewer.

A first-party page is blank or blocked in a screenshot

Check whether the page requires consent, a login, a region, or a bot challenge. Use ScreenshotNeo’s wait, cookie, header, viewport, and network-idle controls only when you are authorized to access the page; a bot check is a reason to stop, not to defeat it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to choose a source

Decision axis Official API or partner feed General crawling
Authorization Usually clearer, but approval and narrower fields may apply. Often ambiguous and dependent on page terms and contract.
Freshness Defined update behavior and identifiers when documented. Depends on polling and page changes.
Maintenance Schema changes are communicated through documentation. Layouts, scripts, and anti-automation controls can break parsers.
Storage and deletion Agreement states what may be retained or redistributed. You must establish rights before building a database.
Operations Quotas, authentication, and support channels are explicit. Higher risk of blocks, legal disputes, and unplanned load.

For most products, the authorized API or feed is the defensible default. Use a crawler or screenshot only for a first-party page where the owner has permitted that collection and you can meet its rate, privacy, and deletion requirements.

FAQ

Should raw job descriptions be stored forever?

No. Retain raw text only for the period and purpose your authorization allows, and attach a deletion workflow to the source record.

What should a reviewer see when AI is uncertain?

Show the normalized value, confidence, exact supporting span, model and prompt versions, and the original permitted record so the reviewer can accept, correct, or reject it.

Can an AI scraper replace a human application workflow?

No. Automated collection should not bypass application controls, replace required human input, or make employment decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Should raw job descriptions be stored forever?

No. Retain raw text only for the period and purpose your authorization allows, and attach a deletion workflow to the source record.

What should a reviewer see when AI is uncertain?

Show the normalized value, confidence, exact supporting span, model and prompt versions, and the original permitted record so the reviewer can accept, correct, or reject it.

Can an AI scraper replace a human application workflow?

No. Automated collection should not bypass application controls, replace required human input, or make employment decisions.

The Bottom Line

Build the scraper around authorization, evidence, and deletion—not around bypassing a job board. An approved feed plus deterministic parsing and reviewable AI enrichment gives you cleaner data with less legal and operational risk.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Bestseller No. 1
San Jamar 1178902 Cutting Board Scraper/Refinisher Tool
San Jamar 1178902 Cutting Board Scraper/Refinisher Tool
SECURE GRIP: Grooved handle provides essential grip while refinishing
$41.18

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.