October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk7 min

Python JSON: Working with Data Files

Save and load Python lists and dictionaries with json.dump and json.load, use UTF-8 encoding, store multiple records safely, and diagnose JSON file errors.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To save a Python list or dictionary to a file, open the file in text mode with encoding="utf-8" and pass it to json.dump(). To load it later, open the file and pass it to json.load(). Each file should hold one JSON document. If you have many records, put them in one list or store them in JSON Lines format. Calling json.dump() repeatedly on the same file does not produce a valid file.

The basic write and read workflow

The standard library’s json module needs no installation. The following script writes a dictionary to a file and reads it back:

As an Amazon Associate I earn from qualifying purchases.

import json

record = {"name": "Ada", "active": True}

with open("record.json", "w", encoding="utf-8") as f:
    json.dump(record, f, ensure_ascii=False, indent=2)

with open("record.json", "r", encoding="utf-8") as f:
    loaded = json.load(f)

print(loaded["name"])  # Ada

Three parts of this pattern matter. The with block closes the file even if an error occurs. The encoding argument is set explicitly rather than left to the platform default. The indent=2 option makes the saved file readable by people, at the cost of a few extra bytes of whitespace. Compact output is the default when indent is omitted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

dump, dumps, load, and loads

The module has two families of functions. The ones ending in s work with strings, and the others work with file objects. Use the file versions when the data lives in a file.

Function Input Output Typical use
json.dump(obj, fp) Python value and a writable file object Writes JSON text to fp Saving to a file
json.dumps(obj) Python value Returns a str Sending JSON in a request or logging it
json.load(fp) Readable file object Python value parsed from the whole document Reading a file
json.loads(s) str, bytes, or bytearray Python value Parsing JSON already in memory

Because json.dump() writes str values, the file must accept text. A file opened with "wb" raises TypeError. Open it with "w" and an encoding instead.

Encoding: why UTF-8 and what ensure_ascii changes

The Python tutorial’s “Input and Output” section states: “JSON files must be encoded in UTF-8.” It recommends passing encoding="utf-8" whenever you open a JSON text file. Skipping this argument can cause a file to be written in one encoding and read in another, and the mismatch often shows up only with non-English text.

The ensure_ascii option controls how non-ASCII characters appear in the output:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Setting Input value Text written to the file Notes
ensure_ascii=True (default) "Zoë" "Zoë" Output is pure ASCII and safe in any encoding, but harder to read
ensure_ascii=False "Zoë" "Zoë" Characters are written directly, so the file must be saved as UTF-8

The default escaping is a reasonable choice when you control neither the reader nor the encoding. When you open the file with encoding="utf-8", ensure_ascii=False gives human-readable output without any loss.

Why repeated dump() calls produce invalid JSON

This is the most common mistake with JSON files. The json module reference puts it directly: “Unlike pickle and marshal, JSON is not a framed protocol, so trying to serialize multiple objects with repeated calls to dump() using the same fp will result in an invalid JSON file.”

import json

with open("bad.json", "w", encoding="utf-8") as f:
    json.dump({"name": "Ada"}, f)
    json.dump({"name": "Grace"}, f)

with open("bad.json", "r", encoding="utf-8") as f:
    json.load(f)  # raises json.JSONDecodeError: Extra data

The file contains {"name": "Ada"}{"name": "Grace"}. Each call writes a complete document, but nothing marks where one ends and the next begins. json.load() reads one document and then finds unexpected trailing content.

Storing many records in a file

There are two reliable ways to store several records. Choose based on whether the records are read as a single unit or processed one at a time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Option 1: one list in one document

If the records belong together, put them in a list and call json.dump() once:

import json

records = [{"id": 1, "name": "Ada"}, {"id": 2, "name": "Grace"}]

with open("records.json", "w", encoding="utf-8") as f:
    json.dump(records, f, ensure_ascii=False, indent=2)

with open("records.json", "r", encoding="utf-8") as f:
    loaded = json.load(f)

The whole file is loaded into memory at once, which is fine for files of moderate size.

Option 2: JSON Lines, one object per line

For logs, exports, or large collections that you process one record at a time, write one JSON object per line. Each line is a separate document, so you can read the file line by line without loading all of it:

import json

records = [{"id": 1, "name": "Ada"}, {"id": 2, "name": "Grace"}]

with open("records.jsonl", "w", encoding="utf-8") as f:
    for record in records:
        f.write(json.dumps(record) + "n")

with open("records.jsonl", "r", encoding="utf-8") as f:
    loaded = [json.loads(line) for line in f if line.strip()]

This works because json.dumps() without indent emits no line breaks inside a record. Do not use indent with this format, because each record must fit on one line.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate and format from the command line

The json module provides a command-line tool for checking and reformatting JSON. The current reference documents python -m json, and python -m json.tool remains supported for backward compatibility. It can read from standard input, write to standard output, accept input and output file names, sort keys, and control indentation.

# Validate and pretty-print to the terminal
python -m json record.json

# Pretty-print with sorted keys and four-space indentation into a new file
python -m json.tool --sort-keys --indent 4 record.json formatted.json

# Read from standard input
cat record.json | python -m json

# Check each line of a JSON Lines file
python -m json --json-lines records.jsonl

When the input is invalid, the tool reports an error instead of printing formatted output. Use it to find the line that breaks a file before you debug your Python code.

Handling errors when loading JSON

Invalid JSON raises json.JSONDecodeError, a subclass of ValueError. The exception provides msg, lineno, and colno attributes, which are useful for showing where the problem is:

import json

try:
    with open("record.json", "r", encoding="utf-8") as f:
        data = json.load(f)
except json.JSONDecodeError as exc:
    print(f"Invalid JSON at line {exc.lineno}, column {exc.colno}: {exc.msg}")

Catch JSONDecodeError specifically. Catching every exception, or every ValueError, hides other problems. A missing file raises FileNotFoundError, and a file that is not valid UTF-8 raises UnicodeDecodeError before any JSON parsing begins.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting common failures

Symptom Likely cause Fix
JSONDecodeError: Extra data Several dump() calls wrote multiple documents into one file Dump one list, or switch to JSON Lines
TypeError when calling json.dump() The file was opened in binary mode Open with "w" and encoding="utf-8"
UnicodeDecodeError on load The file is not encoded in UTF-8 Open with the encoding that was actually used, or re-save the file as UTF-8
Non-ASCII text appears as u sequences ensure_ascii is at its default of True Pass ensure_ascii=False when the file is UTF-8
Dictionary keys changed type after a round trip JSON object keys are always strings, so integer keys become strings Store keys as strings from the start
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Objects that JSON cannot store directly

JSON natively represents objects, arrays, strings, numbers, booleans, and null. Python lists and dictionaries with those kinds of values work directly. Class instances, dates, and sets do not. Passing them to json.dump() raises TypeError unless you supply a conversion strategy.

The most direct approach is a default function, which the encoder calls for any object it cannot serialize:

import json
from datetime import date

class Task:
    def __init__(self, name, due):
        self.name = name
        self.due = due

def encode(obj):
    if isinstance(obj, Task):
        return {"name": obj.name, "due": obj.due.isoformat()}
    raise TypeError(f"Cannot serialize {type(obj).__name__}")

with open("tasks.json", "w", encoding="utf-8") as f:
    json.dump([Task("Write draft", date(2026, 10, 9))], f, default=encode, indent=2)

Loading does not reverse this automatically. Your code must read the dictionary back and construct the Task object, including converting the date string with date.fromisoformat().

Reading untrusted input safely

The json module reference warns that parsing malicious input may consume considerable CPU and memory, and recommends limiting the size of the data. This is a resource-exhaustion risk. It is not the same as the code-execution risk associated with pickle, which is covered below. The module does not set a limit for you, so choose one based on the largest file your application should accept:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import os

MAX_BYTES = 5 * 1024 * 1024  # application-chosen limit

def load_limited(path):
    if os.path.getsize(path) > MAX_BYTES:
        raise ValueError(f"{path} exceeds {MAX_BYTES} bytes")
    with open(path, "r", encoding="utf-8") as f:
        return json.load(f)

JSON or pickle

Both modules serialize Python data, but they suit different situations. The choice depends mainly on who reads the data and where it comes from.

Criterion JSON pickle
Interoperability A common interchange format read by many applications and languages Specific to Python
Data shape Objects, arrays, strings, numbers, booleans, and null; other types need conversion code Can represent many Python objects directly
Trust Parsing does not execute code, but untrusted input still needs size limits and error handling The Python tutorial warns that deserializing malicious pickle data can execute code; never load untrusted pickle data
Readability Text you can open and edit Binary data

For configuration files, exported data, or anything shared with other systems, use JSON. Reserve pickle for trusted, Python-only data that you produced yourself and that must preserve Python types.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.