Choose an export format that matches the program or person that will use your scraped records, then configure your scraper to write that format. With Scrapy, the quickest route is a command such as scrapy crawl myspider -O results.json: it runs the spider and writes its items to a JSON file. Scrapy also supports CSV, JSON Lines, and XML. The exact command and options differ across scraper frameworks and hosted services, so treat the Scrapy examples below as Scrapy-specific rather than universal.
Choose a file format before exporting
The destination of your data is the best guide to format choice. A spreadsheet user usually needs CSV; an application that consumes structured records may prefer JSON; a large or incremental record stream is often a good fit for JSON Lines. XML is available when a downstream system specifically expects it.
| Format | Good fit | Important consideration |
|---|---|---|
CSV (.csv) |
Spreadsheets and systems that expect rows with consistent columns. | CSV has a fixed header. Decide which fields to export and their order, especially if some scraped items omit fields. Nested objects and arrays do not naturally form columns; flatten them or deliberately encode them as text. |
JSON (.json) |
Structured records, including records with nested fields, and many application workflows. | A conventional JSON export is treated as a document. Consumers may need to load the whole document to parse it, and appending another run with the wrong mode can make the file invalid. |
JSON Lines (.jsonl) |
Incremental exports, record-by-record processing, and stream-like workflows. | Each line is a separate JSON value. Consumers must read one JSON value per line rather than expect a single JSON array or object. |
XML (.xml) |
A recipient or integration that specifically requires XML. | Choose it for compatibility with that consumer; do not assume it is interchangeable with JSON or CSV. |
Scrapy’s feed-export documentation lists JSON, JSON Lines, CSV, and XML serializers and explains that the format can be inferred from the output file extension. The exporter documentation also recommends JSON Lines for large exports where parsing an entire JSON document is a poor fit. These are exporter behaviors, not a promise that every scraper supports the same formats.
Save a Scrapy spider’s output to a file
In Scrapy, a spider yields items—records containing the fields you extracted. Feed export writes those items to a file. You need a Scrapy project with a runnable spider; replace myspider with the spider’s actual name and choose an output filename and extension that match the format you want.
#1 Best Overall
- Find the spider name. From the project directory, run
scrapy listto list available spiders. Use the name shown for the spider you intend to run. - Pick the format. For example, use
results.jsonfor JSON,results.csvfor CSV,results.jsonlfor JSON Lines, orresults.xmlfor XML. - Run the spider and export its items. For a fresh JSON file, run
scrapy crawl myspider -O results.jsonin the Scrapy project directory. - Inspect the file. Check that it exists, contains the expected records, and has usable field names and values. Open CSV in a spreadsheet or inspect JSON and JSON Lines with a text editor or the parser used by the next stage of your workflow.
For a different format, keep the command pattern and change the extension, for example scrapy crawl myspider -O results.csv or scrapy crawl myspider -O results.jsonl. Scrapy can infer the serializer from a supplied extension. If you need exporter settings beyond the extension-based defaults—for example, controlling the fields and order in a CSV—configure the feed explicitly in the Scrapy project.
Fresh output versus appending
Uppercase -O overwrites an existing output file; lowercase -o appends to it. That distinction matters when a scheduled job writes to the same path repeatedly. Use overwrite when each run should create a new snapshot. Use append only when you intend to combine runs and the chosen format remains valid under append behavior.
Appending to ordinary JSON can produce invalid JSON, because a second run’s document is not automatically merged into the first document. JSON Lines is a more suitable append format: each exported item is written as its own JSON value on a line, which supports incremental writing and record-by-record reading. Still check whether repeated runs will add duplicate records; append mode does not deduplicate your data.
Make CSV columns predictable
CSV is convenient only when the rows have a consistent shape. Scraped sites often expose optional fields, and an item may not contain every field another item has. Decide which columns your recipient expects and set a stable field list and order in Scrapy’s feed configuration. That prevents output shape from depending on which fields happened to appear first or in a particular run.
Before sending the file downstream, check how missing values appear and whether text containing commas, quotes, or line breaks opens as intended in the recipient’s tool. If a record contains nested data, define a transformation: flatten it into named columns, or serialize that value into a text field if the consumer can handle it. CSV does not preserve nested objects as a structured hierarchy by itself.
Export from a hosted scraper run
If the scraper runs on a hosted service rather than in your own Scrapy project, downloading results is a separate workflow. The Scrapy.io dataset API documents downloadable JSON, CSV, and JSON Lines responses, along with pagination options. That is specific to Scrapy.io’s dataset API; do not assume another hosted scraper provides the same endpoint, formats, or pagination behavior. Consult the service’s own run or dataset documentation for its download method and limits.
Check the output and diagnose common failures
A successful spider run does not by itself guarantee that the file is useful. Verify the export at the point where you will consume it, and distinguish a crawl problem from a serialization or file-path problem.
- No file appears: Confirm you ran the command from the Scrapy project directory, that the spider name is correct, and that the process reached completion. Check the terminal output for a spider startup error, filesystem permission problem, or different output path than expected.
- The file exists but has no records: The export contains items yielded by the spider, not every page it visits. Check that the spider’s parsing logic actually yields items and that the crawl reached pages that should contain data.
- The file is replaced unexpectedly: Uppercase
-Ooverwrites. Use lowercase-oonly when appending is intended, and choose a format that supports the resulting structure. - The appended JSON will not parse: Ordinary JSON is not automatically merged when a later run is appended. Export incremental records as JSON Lines, or overwrite and produce a fresh JSON document for each run.
- CSV has missing or inconsistent-looking columns: Specify the desired CSV fields and their order in the feed configuration. Decide how absent values and nested fields should be represented before relying on the file in another application.
- A downstream parser rejects the file: Confirm that the consumer expects the selected format. A JSON Lines file is not a single JSON array; a CSV file is not a nested object format. Validate against the receiving tool’s expected schema and field names.
- A later run contains duplicates: Append mode adds records; it does not establish that records are unique. Choose a fresh-output workflow or implement deduplication using an appropriate record key in your pipeline.
Or skip the browser setup
If what you need alongside your extracted records is a screenshot of a page, ScreenshotNeo captures a URL directly; it is not a replacement for exporting the records your spider extracted. One GET request saves a screenshot response to a file. See the ScreenshotNeo API documentation for parameters and response details.
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.
Cost, reliability, and repeatable runs
A local Scrapy export does not require a separate hosted dataset download step, but you are responsible for where the file is written, how long it is kept, and how it is transferred to its consumer. Use a unique filename or a deliberate overwrite policy for recurring runs. If the output is important, retain a known-good prior export until you have checked the new one.
For large output, consider whether the consumer can process records incrementally. Scrapy’s documentation recommends JSON Lines when whole-document JSON parsing is not a good fit. CSV can be practical for tabular consumers, but its fields should remain stable; a changing set of item fields can complicate downstream imports. No format choice fixes incomplete crawling, duplicate records, or changing source-page content, so validate both the crawl result and the exported structure.
For scheduled or shared workflows, record which spider and run produced each file, and check completion before downstream jobs read it. If the next system expects a specific schema, make that contract explicit—field names, order, missing-value handling, and whether a run replaces or extends prior data—rather than relying on whichever defaults happen to work for one run.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Frequently Asked Questions
What is the difference between JSON and JSON Lines?
A conventional JSON export is a structured document; JSON Lines stores a separate JSON value on each line, which is convenient for incremental and record-by-record processing.
Can I save scraped data directly from a website in my browser?
This article’s commands apply to Scrapy spiders and hosted scraper exports, not a browser’s generic save-page feature. A browser page save is not the same as exporting the structured items produced by a scraper.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




