Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The best scraping format is the one your next system can consume reliably. Use JSONL for large or incremental crawls, CSV for flat records and spreadsheet or SQL workflows, JSON for nested API-style data, and XML when an integration contract requires hierarchy or namespaces. Scrapy includes all four, plus Python-oriented Pickle and Marshal exporters, while pandas can read and write CSV, JSON, HTML and XML.

Choose the format from the downstream consumer

Scraping itself does not determine the output format. Decide what will read the data, whether records are flat or nested, how much data must be processed at once, and where the feed will be stored. Scrapy feed exports support local files, FTP, Amazon S3 and standard output, so destination and serialization can be selected together. For example, JSONL in object storage suits an appendable batch pipeline; CSV on a local or FTP destination suits a fixed-schema exchange.

Format Best fit Shape and processing Main trade-off
JSON Nested API-style interchange Objects and arrays; commonly one complete document Many parsers need the whole document, making incremental processing awkward
JSONL Large, streaming or incremental feeds One JSON object per line; process records independently Less convenient when a consumer expects one conventional JSON document
CSV Spreadsheets, SQL loads and flat analysis Rows with a header and fixed columns Nested or repeated values require flattening or a join policy
XML Hierarchical or enterprise contracts Nested elements, attributes and namespaces More verbose than JSON and usually tied to a specific schema
Pickle Controlled Python-only handoff Python object serialization Weak cross-language portability; never load untrusted files
Marshal Controlled Python-runtime use Python-oriented serialization Runtime and language compatibility are limited

Scrapy documents these format keys as json, jsonlines, csv, xml, pickle and marshal in its Feed exports documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JSON: flexible records with a document-sized cost

JSON preserves nested objects and arrays without forcing a relational design. It is a natural interchange format for an API, such as an item containing a product, seller object and list of offers. Scrapy’s JsonItemExporter writes scraped items as a JSON structure, commonly an array of objects.

The limitation is incremental parsing. Scrapy’s exporter documentation warns that many JSON parsers do not support incremental parsing well. A very large array may therefore require a parser to hold or scan the complete document before useful records are available. Use JSON when the receiving contract expects one document or when the feed is moderate; choose JSONL when records should be consumed as they arrive.

Scrapy command

scrapy crawl products -O products.json

The -O option overwrites the target. Use -o when appending is appropriate for your workflow, and validate how your pipeline handles duplicate or previously exported records.

JSON Lines (JSONL): the default for big or continuing crawls

JSONL places one complete JSON value on each line. A worker can append a record, a stream processor can resume at a line boundary, and a failed job can often be restarted without reparsing a giant array. These properties make JSONL the strongest general default for large, incremental or partitioned feeds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy command

scrapy crawl products -o products.jsonl -t jsonlines

Each line should be independently valid JSON. Keep newline characters inside string values escaped, and use a stable item structure so downstream jobs do not need format-specific exceptions.

When JSONL is not ideal

A consumer that insists on a single JSON array may reject JSONL, and some business users find line-delimited objects less approachable than a spreadsheet. In that case, export JSONL as the durable landing format and transform it into the consumer’s required document later.

CSV: excellent for flat, known columns

CSV is a practical handoff to analysts, spreadsheets and database bulk loaders. It is compact and easy to inspect, but it represents rows and columns rather than arbitrary trees. Before exporting, decide how to flatten nested objects and repeated fields. You might create columns such as seller_name and seller_id, or emit a separate child file keyed by an item ID. Do not silently stringify an array if another system expects one value per row.

Keep the header stable

Scrapy’s CsvItemExporter supports FEED_EXPORT_FIELDS (or a per-feed fields setting) to control selected columns, names and order. Define this list explicitly when files are exchanged between teams; otherwise a newly encountered field can change the header or make historical files inconsistent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
# settings.py
FEED_EXPORT_FIELDS = [
    "item_id", "title", "price", "currency", "url"
]

# command
scrapy crawl products -O products.csv

Choose an encoding and delimiter that match the receiving application. Quote fields containing commas, line breaks or quote characters, and test import with real text such as accents and non-Latin scripts.

XML: use hierarchy when the contract requires it

XML remains appropriate for integrations that require nested elements, attributes, namespaces or an established XML schema. Scrapy’s XmlItemExporter is built in. XML can represent relationships that would be awkward in CSV, but the receiving contract normally dictates element names, ordering, namespace declarations and validation rules. Confirm those rules before you start a crawl rather than trying to infer them from the first output file.

scrapy crawl products -O products.xml

Pickle and Marshal: only inside a trusted Python boundary

Pickle and Marshal are available in Scrapy, but they are Python-oriented serialization choices, not general interchange formats. Use them only when producer and consumer runtimes are controlled and the trust boundary is clear. Never deserialize files supplied by an untrusted party: Python object deserialization can execute attacker-controlled behavior. For a cross-language service, choose JSON, JSONL, CSV or XML instead.

Scrapy feed settings that affect every format

Set the format explicitly

You can configure a feed in FEEDS so the path, format and storage are version-controlled:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
# settings.py
FEEDS = {
    "exports/%(name)s-%(time)s.jsonl": {
        "format": "jsonlines",
        "encoding": "utf8",
    },
}

Use feed-specific settings when one crawl needs multiple deliveries, such as JSONL for a data lake and CSV for an analyst. Scrapy supports feed-specific encoding and indentation settings; indentation is implemented for JSON and XML exporters. Pretty printing improves inspection but increases file size and write volume, so leave it off for high-volume feeds.

Separate serialization from storage

A format does not determine where the bytes live. Scrapy can write to the filesystem, FTP, Amazon S3 or stdout. A common production pattern is JSONL partitions in object storage, then a job that validates, deduplicates and loads them into analytics tables. For a small run, stdout can make a Unix pipeline convenient:

scrapy crawl products -o - -t jsonlines | gzip > products.jsonl.gz

Using the exports with pandas

pandas exposes top-level readers and DataFrame writer methods for common interchange formats. The documented I/O API includes read_csv/to_csv, read_json/to_json, read_html/to_html and read_xml/to_xml. read_html parses HTML tables into DataFrames.

CSV workflow

import pandas as pd

df = pd.read_csv("products.csv")
df.to_csv("products-clean.csv", index=False)

JSON and JSONL workflow

import pandas as pd

nested = pd.read_json("products.json")
stream = pd.read_json("products.jsonl", lines=True)
stream.to_json("products-normalized.jsonl", orient="records", lines=True)

For very large JSONL files, process chunks or use a streaming-oriented engine before constructing one DataFrame. A DataFrame is still an in-memory object, so changing the wire format does not remove pandas’ memory requirements.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

XML and HTML tables

xml_df = pd.read_xml("products.xml")
tables = pd.read_html("page-with-tables.html")

XML element paths, namespaces and repeated nodes may require additional normalization after reading. Preserve an identifier so child records can be joined back to their parent.

A practical decision checklist

  1. Identify the consumer. An API contract, analyst, SQL loader and Python-only job have different needs.
  2. Measure structure. If items contain arrays or nested objects, prefer JSON/JSONL or design a deliberate relational flattening.
  3. Estimate volume and restart needs. For continuous or very large crawls, select JSONL and partition files by date or job.
  4. Freeze the schema where required. Configure FEED_EXPORT_FIELDS for CSV and version your field definitions.
  5. Choose destination and encoding together. Test filesystem, FTP, S3 or stdout delivery with the actual consumer.
  6. Validate representative records. Check nulls, Unicode, delimiters, nested arrays, duplicate IDs and malformed pages before scaling out.
  7. Document evolution. Record exporter settings, field names, date, URL and any flattening rules alongside each batch.

Common failure modes and fixes

CSV columns shift between runs

Cause: fields are discovered in a different order or new fields appear. Fix: set FEED_EXPORT_FIELDS, keep names stable and version intentional schema changes.

A JSON parser runs out of memory

Cause: a large document is being loaded as one array. Fix: export JSONL, read line by line or use chunked processing, and partition the crawl.

Nested data disappears in CSV

Cause: a tree was forced into a row without a flattening policy. Fix: create explicit scalar columns, a child table/file, or switch to JSONL.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

XML is rejected by an enterprise endpoint

Cause: incorrect namespace, element order, encoding or schema version. Fix: obtain the endpoint’s XSD or example payload, configure the exporter to match it, and validate before delivery.

Unicode or commas corrupt a spreadsheet import

Cause: encoding, delimiter or quoting mismatch. Fix: export UTF-8, use the recipient’s expected delimiter, quote fields correctly and test with multilingual and multiline values.

Appending creates duplicates

Cause: an append export was retried without an idempotency key. Fix: include a stable source ID, deduplicate downstream and use overwrite or job-specific paths when reruns must be isolated.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Capture a source page for audit or review

If your scraping workflow needs visual evidence of the page that produced a record, you can capture a URL separately from the data export. A browser-based capture must manage consent dialogs, popups, delayed content and failed loads itself.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

ScreenshotNeo provides a one-request website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. AI agents can use its take_screenshot, get_page_info and capture_pdf MCP tools.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for options such as full-page and element captures, device presets, PDF output, custom CSS or JavaScript, waits, blocked resources, headers, cookies, geolocation, caching, signed links, asynchronous webhooks and bulk capture of up to 100 URLs per call. The Free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Can I convert JSONL to CSV later?

Yes. Read each JSONL record, define how nested and repeated values should flatten, then write a controlled column list. Keeping JSONL as the raw export preserves information that a CSV conversion might discard.

Is JSONL officially different from JSON?

JSONL is a line-delimited convention in which each line is a separate JSON value. It is not one JSON array document, so the reader must support line-oriented input.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I export HTML instead?

HTML is useful when the deliverable is a rendered page or table, not as a general-purpose Scrapy item exporter choice. For tabular HTML, pandas can parse tables with read_html; retain structured JSONL or CSV when the data itself is the primary artifact.

Frequently Asked Questions

Which format should I use for a crawl that may be interrupted?

Use JSONL and write independent records or partitions so a restart can continue without reparsing one large JSON document.

What is the safest format to exchange with another programming language?

JSON, JSONL, CSV or XML are broadly interoperable. Pickle and Marshal should remain inside a controlled Python environment.

How do I preserve nested products while also serving spreadsheets?

Keep JSON or JSONL as the lossless source, then generate a documented flat CSV view with stable columns and explicit handling for arrays.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.