Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsArchiveBox stores captured pages and extractor output beneath the data directory’s archive/ tree. To find what is consuming space, measure that tree and inspect its snapshot directories with your operating system’s disk-usage tools; the cited ArchiveBox documentation does not describe a built-in command that sorts captures by size. Once you identify a snapshot, remove it through ArchiveBox’s supported CLI or UI—not by deleting its files directly.
Where ArchiveBox stores its data
ArchiveBox’s data or output directory contains both its index and archived files. The root commonly includes index.sqlite3 and ArchiveBox.conf; archived snapshots and extractor results live beneath archive/. A snapshot can contain an index.jsonl, an index.html, and outputs such as wget/warc/, ytdlp/media/, or git/. The exact layout can vary by release. See the ArchiveBox Usage documentation.
Current snapshots are sharded under paths of the form archive/users/<user>/snapshots/<date>/<domain>/<uuid>/, according to the Security Overview. Do not assume every installation uses an older flat directory layout.
Find which directories are using the space
1. Confirm the real data path and mount
Check your ArchiveBox configuration and deployment to find the actual output/data directory (often configured with OUTPUT_DIR). If ArchiveBox runs in Docker, identify the host path mounted into the container. Also check whether archive/ is a separate bind mount, network share, or other filesystem: the full or nearly full filesystem may not be the container’s root filesystem.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
- Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
2. Measure the archive tree with operating-system tools
These are general shell commands, not ArchiveBox features. Replace /path/to/data with the actual data directory on the machine or mount where the files reside.
du -sh /path/to/data
du -sh /path/to/data/archive
To compare the immediate subdirectories beneath the archive tree on a typical GNU/Linux host, use:
du -h --max-depth=1 /path/to/data/archive | sort -h
To sort the immediate subdirectories from largest to smallest, use:
du -h --max-depth=1 /path/to/data/archive | sort -hr
These GNU options are not portable to every Unix-like system; consult that system’s du and sort help or manual pages if they are unsupported. Permission errors can leave totals incomplete, so run the measurement as an account that can read the archive tree. In Docker or on network storage, measure the mounted host path as well as the in-container path.
Rank #2
- Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
3. Narrow the search to a snapshot
Drill into the largest date or domain directories, then compare the individual snapshot directories. Use ArchiveBox’s list or UI to match a directory to its URL or snapshot identity before removing anything. Extracted media can account for a large share of a snapshot, but the contents depend on which extractors ran.
The documentation cited here describes the layout, but not a built-in per-snapshot size report or “largest captures” command. Treat filesystem measurements as the way to locate large directories; do not mistake du output for an ArchiveBox feature.
How much disk space does ArchiveBox use?
The ArchiveBox project gives a broad estimate of roughly 1 GB to 50 GB per 1,000 snapshots, attributing much of the variation to video and audio capture and the YTDLP_MAX_SIZE limit. This is a project estimate, not a predictable per-snapshot allowance. The project repository does not make it a guarantee for a particular collection.
The ArchiveBox Usage wiki also includes an anecdote of about 1 GB for 1,000 articles, gathered on a single-threaded i5 system with a 50 Mbps connection; its author explicitly says results vary. That is not a general benchmark. Page content, media, extractor configuration, and filesystem behavior all affect actual use.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Rank #3
- Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Remove a known large capture safely
Before deletion
- Verify the exact URL or snapshot in ArchiveBox, especially if several snapshots exist for the same site.
- Back up the data first if the capture or index matters; deletion is irreversible.
- Check your installed version’s command syntax with
archivebox helpor the relevant CLI help. CLI behavior and paths can change between releases.
Remove it through ArchiveBox
For a known URL, the Security Overview documents this command:
archivebox remove --yes 'https://example.com/page'
Replace the example URL with the target URL. The documented behavior is to delete matching Snapshot rows and schedule their snapshot directories for cleanup through ArchiveBox’s normal state-machine path. The legacy --delete flag is accepted for CLI compatibility but does not change that behavior. For an interactive UI workflow, the Usage documentation describes using the snapshot’s Delete action; it warns that deletion cannot be undone.
Do not use rm -rf on a snapshot directory as the normal removal method. The index and output files are related state, and removing files behind ArchiveBox’s back can leave them inconsistent. Reserve manual filesystem intervention for a recovery procedure specific to your version, with a backup and verified database state.
Check whether space was actually reclaimed
After ArchiveBox processes cleanup, measure the data directory and the relevant mount again using the same filesystem tools. If the reported free space does not change, check whether the files were on a different mount than expected, whether cleanup is still pending, and whether ArchiveBox’s non-root account has permission to remove files. On NFS, SMB, or FUSE storage, verify server-side UID/GID mappings and ACLs as well as local permissions.
Rank #4
- Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Removing a snapshot’s output is not necessarily complete erasure of its history. Imported URL lists under sources/, operational records under logs/, and an external search backend may still contain traces. If you need privacy erasure rather than disk recovery, identify and address those stores separately, while following any retention requirements that apply to your archive.
Reduce future storage growth
Choose which extractors to run
Disable extractors you do not need. In particular, media extraction can substantially change storage requirements. The trade-off is reduced archival completeness: a smaller archive may not preserve media or other outputs you would otherwise want.
Separate index storage from bulk archive storage
ArchiveBox’s storage guidance recommends keeping the SQLite index on reliable local storage, while the bulk archive/ directory can be placed on a slower HDD or suitable network filesystem. This can make large archives more practical without moving the index to storage that may be less reliable for database access. Confirm that the ArchiveBox user can create and remove files on the target mount.
Use retention only as a deliberate deletion policy
The DELETE_AFTER setting can remove Crawls, Snapshots, ArchiveResults, and Process rows along with their on-disk outputs after the configured duration. The Configuration wiki says 0, an empty string, or None disables automatic deletion by default; the most-specific setting wins across global, persona, crawl, and snapshot levels. Retention is destructive and irreversible, so establish its scope and test your backup and recovery process before enabling it. See ArchiveBox Configuration.
Best Value
- [Upgraded Version] - This external hard drive features a mirrored logo stripe combined with a striped anti-slip design, and the rounded corners of the casing make it easier to grip. The stripes also have a heat dissipation function, ensuring stable and fast data transfer.
- 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
- 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
- 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
- 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
Consider compression or deduplication carefully
ArchiveBox mentions filesystem compression and deduplication options such as ZFS or BTRFS, and tools such as fdupes or rdfind. These are system-level approaches, not automatic ArchiveBox cleanup controls. Their benefits depend on the content and filesystem, and tools that manipulate duplicate files do not inherently understand ArchiveBox’s application state.
Or skip the browser setup
If your actual goal is to capture a clean website screenshot rather than preserve an ArchiveBox collection, ScreenshotNeo offers a one-call screenshot API. For example, using the supplied cURL pattern with a target URL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for setup and options. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 screenshots. Learn about ScreenshotNeo or sign up free for 1,000 screenshots a month with no card.
Common problems
- The size total seems too small: The command may lack permission to read all directories, or you may be measuring the container filesystem instead of the host-mounted archive. Re-run with suitable permissions and check the actual mount.
durejects an option: The example’s--max-depthflag is GNU-specific. Use your operating system’s documented equivalent or inspect directories one level at a time.- Removing a URL does not free space immediately: ArchiveBox schedules directory cleanup through its normal state-machine path. Check the application’s processing state, then measure again on the filesystem that stores the archive.
- Deletion fails on network storage: Check the ArchiveBox process’s UID/GID and the share’s ACLs or permissions; container-level ownership alone may not grant server-side removal rights.
- The disk is still full after a snapshot is removed: Confirm that the large files were in that snapshot, that the correct mount was measured, and that other stores such as logs, imported sources, or an external search backend are not the remaining source of usage.
ArchiveBox’s documentation was consulted as of October 3, 2026; the Usage wiki was edited September 20, 2026. Because CLI syntax, settings, and paths may change by release, check the documentation and CLI help for the version you run.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




