Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
CSV

How to Save Web Scraper Data to a File

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an export format that matches the program or person that will use your scraped records, then configure your scraper to write that format. With Scrapy, the quickest route is a command such as scrapy crawl myspider -O results.json: it runs the spider and writes its items to a JSON file. Scrapy also supports CSV, JSON Lines, and XML. The exact command and options differ across scraper frameworks and hosted services, so treat the Scrapy examples below as Scrapy-specific rather than universal.

Choose a file format before exporting

The destination of your data is the best guide to format choice. A spreadsheet user usually needs CSV; an application that consumes structured records may prefer JSON; a large or incremental record stream is often a good fit for JSON Lines. XML is available when a downstream system specifically expects it.

Format Good fit Important consideration
CSV (.csv) Spreadsheets and systems that expect rows with consistent columns. CSV has a fixed header. Decide which fields to export and their order, especially if some scraped items omit fields. Nested objects and arrays do not naturally form columns; flatten them or deliberately encode them as text.
JSON (.json) Structured records, including records with nested fields, and many application workflows. A conventional JSON export is treated as a document. Consumers may need to load the whole document to parse it, and appending another run with the wrong mode can make the file invalid.
JSON Lines (.jsonl) Incremental exports, record-by-record processing, and stream-like workflows. Each line is a separate JSON value. Consumers must read one JSON value per line rather than expect a single JSON array or object.
XML (.xml) A recipient or integration that specifically requires XML. Choose it for compatibility with that consumer; do not assume it is interchangeable with JSON or CSV.

Scrapy’s feed-export documentation lists JSON, JSON Lines, CSV, and XML serializers and explains that the format can be inferred from the output file extension. The exporter documentation also recommends JSON Lines for large exports where parsing an entire JSON document is a poor fit. These are exporter behaviors, not a promise that every scraper supports the same formats.

Save a Scrapy spider’s output to a file

In Scrapy, a spider yields items—records containing the fields you extracted. Feed export writes those items to a file. You need a Scrapy project with a runnable spider; replace myspider with the spider’s actual name and choose an output filename and extension that match the format you want.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Find the spider name. From the project directory, run scrapy list to list available spiders. Use the name shown for the spider you intend to run.
  2. Pick the format. For example, use results.json for JSON, results.csv for CSV, results.jsonl for JSON Lines, or results.xml for XML.
  3. Run the spider and export its items. For a fresh JSON file, run scrapy crawl myspider -O results.json in the Scrapy project directory.
  4. Inspect the file. Check that it exists, contains the expected records, and has usable field names and values. Open CSV in a spreadsheet or inspect JSON and JSON Lines with a text editor or the parser used by the next stage of your workflow.

For a different format, keep the command pattern and change the extension, for example scrapy crawl myspider -O results.csv or scrapy crawl myspider -O results.jsonl. Scrapy can infer the serializer from a supplied extension. If you need exporter settings beyond the extension-based defaults—for example, controlling the fields and order in a CSV—configure the feed explicitly in the Scrapy project.

Fresh output versus appending

Uppercase -O overwrites an existing output file; lowercase -o appends to it. That distinction matters when a scheduled job writes to the same path repeatedly. Use overwrite when each run should create a new snapshot. Use append only when you intend to combine runs and the chosen format remains valid under append behavior.

Appending to ordinary JSON can produce invalid JSON, because a second run’s document is not automatically merged into the first document. JSON Lines is a more suitable append format: each exported item is written as its own JSON value on a line, which supports incremental writing and record-by-record reading. Still check whether repeated runs will add duplicate records; append mode does not deduplicate your data.

Make CSV columns predictable

CSV is convenient only when the rows have a consistent shape. Scraped sites often expose optional fields, and an item may not contain every field another item has. Decide which columns your recipient expects and set a stable field list and order in Scrapy’s feed configuration. That prevents output shape from depending on which fields happened to appear first or in a particular run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before sending the file downstream, check how missing values appear and whether text containing commas, quotes, or line breaks opens as intended in the recipient’s tool. If a record contains nested data, define a transformation: flatten it into named columns, or serialize that value into a text field if the consumer can handle it. CSV does not preserve nested objects as a structured hierarchy by itself.

Export from a hosted scraper run

If the scraper runs on a hosted service rather than in your own Scrapy project, downloading results is a separate workflow. The Scrapy.io dataset API documents downloadable JSON, CSV, and JSON Lines responses, along with pagination options. That is specific to Scrapy.io’s dataset API; do not assume another hosted scraper provides the same endpoint, formats, or pagination behavior. Consult the service’s own run or dataset documentation for its download method and limits.

Check the output and diagnose common failures

A successful spider run does not by itself guarantee that the file is useful. Verify the export at the point where you will consume it, and distinguish a crawl problem from a serialization or file-path problem.

  • No file appears: Confirm you ran the command from the Scrapy project directory, that the spider name is correct, and that the process reached completion. Check the terminal output for a spider startup error, filesystem permission problem, or different output path than expected.
  • The file exists but has no records: The export contains items yielded by the spider, not every page it visits. Check that the spider’s parsing logic actually yields items and that the crawl reached pages that should contain data.
  • The file is replaced unexpectedly: Uppercase -O overwrites. Use lowercase -o only when appending is intended, and choose a format that supports the resulting structure.
  • The appended JSON will not parse: Ordinary JSON is not automatically merged when a later run is appended. Export incremental records as JSON Lines, or overwrite and produce a fresh JSON document for each run.
  • CSV has missing or inconsistent-looking columns: Specify the desired CSV fields and their order in the feed configuration. Decide how absent values and nested fields should be represented before relying on the file in another application.
  • A downstream parser rejects the file: Confirm that the consumer expects the selected format. A JSON Lines file is not a single JSON array; a CSV file is not a nested object format. Validate against the receiving tool’s expected schema and field names.
  • A later run contains duplicates: Append mode adds records; it does not establish that records are unique. Choose a fresh-output workflow or implement deduplication using an appropriate record key in your pipeline.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If what you need alongside your extracted records is a screenshot of a page, ScreenshotNeo captures a URL directly; it is not a replacement for exporting the records your spider extracted. One GET request saves a screenshot response to a file. See the ScreenshotNeo API documentation for parameters and response details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Learn about ScreenshotNeo or sign up for 1,000 free screenshots a month with no card.

Cost, reliability, and repeatable runs

A local Scrapy export does not require a separate hosted dataset download step, but you are responsible for where the file is written, how long it is kept, and how it is transferred to its consumer. Use a unique filename or a deliberate overwrite policy for recurring runs. If the output is important, retain a known-good prior export until you have checked the new one.

For large output, consider whether the consumer can process records incrementally. Scrapy’s documentation recommends JSON Lines when whole-document JSON parsing is not a good fit. CSV can be practical for tabular consumers, but its fields should remain stable; a changing set of item fields can complicate downstream imports. No format choice fixes incomplete crawling, duplicate records, or changing source-page content, so validate both the crawl result and the exported structure.

For scheduled or shared workflows, record which spider and run produced each file, and check completion before downstream jobs read it. If the next system expects a specific schema, make that contract explicit—field names, order, missing-value handling, and whether a run replaces or extends prior data—rather than relying on whichever defaults happen to work for one run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

What is the difference between JSON and JSON Lines?

A conventional JSON export is a structured document; JSON Lines stores a separate JSON value on each line, which is convenient for incremental and record-by-record processing.

Can I save scraped data directly from a website in my browser?

This article’s commands apply to Scrapy spiders and hosted scraper exports, not a browser’s generic save-page feature. A browser page save is not the same as exporting the structured items produced by a scraper.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.