Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use the SEC’s public JSON APIs for filing metadata and standard XBRL facts, then fetch the underlying filing when you need narrative text, exhibits, custom tags or a specific section. The APIs require no API key. A reliable scraper preserves the CIK, accession number, source URL, retrieval time and parsing version alongside every normalized record.
This guide shows a practical Python implementation, equivalent cURL and Node.js requests, coverage boundaries, update handling, rate limits, schema design and troubleshooting.
Choose the SEC source that matches your data
| Source | Use it for | Important boundary |
|---|---|---|
| Submissions API | Company filing history and metadata | Recent data contains at least one year of filings or the latest 1,000 filings, whichever is greater; older history is referenced in additional files. |
| Company Facts API | Standard taxonomy facts for one company | It does not represent every narrative disclosure, custom taxonomy fact or exhibit. |
| Company-concept and frames endpoints | A single taxonomy tag, or cross-company calendar-aligned frames | Frames can use calendar periods that differ from a company’s fiscal periods. |
| EDGAR filing archives | Complete filing documents, exhibits and custom extraction | You must download, parse, version and validate the documents yourself. |
| Managed extractor such as SEC-API.io | Vendor-hosted downloads, section extraction and XBRL-to-JSON conversion | It is paid; confirm current coverage, terms and pricing before depending on it. |
The SEC describes its REST APIs and JSON output on its EDGAR APIs page. It states: “These APIs do not require any authentication or API keys to access.”
Prerequisites and access discipline
- Use a zero-padded, ten-digit CIK, such as
0000320193. - Send a descriptive
User-Agentidentifying your application and a contact address. - Keep requests efficient. SEC Developer Resources guidance reviewed March 10, 2025 limits each user to no more than 10 requests per second, regardless of the number of machines used. Excessive traffic can lead to IP blocking; unclassified bots are not allowed. Recheck the current guidance before deployment.
- Cache responses and use nightly bulk archives when acquiring data at scale. The SEC calls
companyfacts.zipandsubmission.zipthe most efficient way to fetch large amounts of API data.
Build a clean JSON scraper in Python
1. Fetch submissions and company facts
The script below writes one traceable output file. It keeps raw fact context instead of flattening units or periods and distinguishes absent data from zero.
#1 Best Overall
- You will get: the package contains 25 file cabinet dividers guides, a total of 5 sets, each set contains 5 colors, including rose, blue, green, orange and yellow, and each color has 5; there are 5 self-adhesive waterproof stickers containing letters and numbers, which can be used according to your own needs.
- Material: The top tab file guides is made of high-quality polypropylene, which is soft and durable, waterproof and wear-resistant, such as easy to clean, soft and not easy to break.
- Easy to use: The file dividers with tabs for file cabinet is very thin, easy to use, occupies no space, and easy to find. The A-Z top tab file guides sticker can be pasted as needed, or you can write the name you want to classify with a pen.
- Suitable color and size: The size of the alphabet dividers for file cabinet drawers is 30 x 25.4 cm/11.8 x 10 inches, which is applicable to the general file size, and the color is easier to distinguish, so that you can easily and quickly find the required documents in the file cabinet, saving your energy and time.
- Beautiful and versatile: The tab polypropylene guides has smooth and tidy edges, elegant appearance and high applicability. It can be used to classify filing cabinets, learning materials, customer materials, and notes to improve your efficiency.
import json
import time
from datetime import datetime, timezone
from pathlib import Path
import requests
CIK = "0000320193" # Apple; replace with a ten-digit CIK
UA = "filings-scraper/1.0 [email protected]"
BASE = "https://data.sec.gov"
s = requests.Session()
s.headers.update({"User-Agent": UA, "Accept-Encoding": "gzip, deflate"})
def get_json(path):
r = s.get(BASE + path, timeout=30)
r.raise_for_status()
return r.json()
retrieved_at = datetime.now(timezone.utc).isoformat()
submissions = get_json(f"/submissions/CIK{CIK}.json")
facts = get_json(f"/api/xbrl/companyfacts/CIK{CIK}.json")
recent = submissions["filings"]["recent"]
filings = []
for i, accession in enumerate(recent["accessionNumber"]):
filings.append({
"cik": CIK,
"accession_number": accession,
"form_type": recent["form"][i],
"filed_at": recent["filingDate"][i],
"period_of_report": recent["reportDate"][i],
"primary_document": recent["primaryDocument"][i],
"source_url": f"https://www.sec.gov/Archives/edgar/data/{int(CIK)}/{accession.replace('-', '')}/{recent['primaryDocument'][i]}",
"retrieved_at": retrieved_at,
"source_kind": "submission metadata",
"parser_name": "custom-sec-json",
"parser_version": "1.0"
})
output = {
"entity": {"cik": CIK, "name": submissions.get("name")},
"filings": filings,
"facts": facts,
"retrieved_at": retrieved_at,
"warnings": []
}
Path("sec_company.json").write_text(json.dumps(output, indent=2), encoding="utf-8")
print(f"Wrote {len(filings)} filings to sec_company.json")
Install the only dependency with python -m pip install requests. The archive URL is constructed from the accession number without hyphens and the SEC-provided primary document name; retain that URL even if you later extract text into another format.
2. Extract a useful fact without losing context
Company Facts is organized as taxonomy, tag and unit. A value can have an instant date, a start/end period, a fiscal year, a form and an accession number. Iterate through those fields rather than assuming every value is annual revenue.
facts = output["facts"]["facts"]
revenue = facts.get("us-gaap", {}).get("Revenues", {})
for unit, values in revenue.get("units", {}).items():
for value in values:
print({"unit": unit, **value})
Tags and taxonomy names vary. If a company extends US-GAAP with a custom taxonomy, that fact may not appear in Company Facts; retrieve the filing and parse its inline XBRL or use a section/XBRL extraction service.
Equivalent requests with cURL and Node.js
cURL
curl -H "User-Agent: filings-scraper/1.0 [email protected]"
"https://data.sec.gov/submissions/CIK0000320193.json"
-o submissions.json
curl -H "User-Agent: filings-scraper/1.0 [email protected]"
"https://data.sec.gov/api/xbrl/companyfacts/CIK0000320193.json"
-o companyfacts.json
Node.js 18+
const fs = require('node:fs/promises');
const headers = { 'User-Agent': 'filings-scraper/1.0 [email protected]' };
async function get(path) {
const r = await fetch(`https://data.sec.gov${path}`, { headers });
if (!r.ok) throw new Error(`${r.status} ${r.statusText}`);
return r.json();
}
(async () => {
const cik = '0000320193';
const [submissions, facts] = await Promise.all([
get(`/submissions/CIK${cik}.json`),
get(`/api/xbrl/companyfacts/CIK${cik}.json`)
]);
await fs.writeFile('sec_company.json', JSON.stringify({ cik, submissions, facts }, null, 2));
})();
Design a durable JSON envelope
The SEC does not define one universal normalized schema. A practical envelope keeps identity, provenance and semantic context together:
Rank #2
- Ample Size and Quantity: with an appreciable size of about 9.8 x 11.7 inches, these alphabetical file dividers are large enough to cater for the organization of various forms of data, documents, and charts; The package includes 50 dividers in 12 assorted colors, providing sufficient quantity for individual usage as well as sharing with classmates, friends, colleagues, and others
- Durable Alphabetical Dividers: manufactured from 500g heavyweight cardboard, the A-Z alphabet file dividers are robust, sturdy, and not easily susceptible to deform, breaking or deformation; They can endure long term usage, making them a nice choice for organizing your important documents through time
- Unique and Stylish Design: these alphabetical dividers are finished in the charming fresh color palette, providing a stylish alternative to universal file folders; Comprising of 12 different beautiful colors, these dividers not only function to keep your files organized but also present a pleasing aesthetic value to your workspace
- Enhanced Efficiency and Convenience: each alphabetical file organizer is equipped with 1/5 cut top tabs preprinted with A-Z, allowing for easy classification, reduced searching time, and elevated productivity; They are designed for desktop and drawer filing, hence promoting a tidy and organized environment
- Extensive Applications: the A-Z tab dividers are not only suitable for office use but also for classrooms and study spaces; They can be applied to categorize or organize various materials including work documents, study materials, or even recipes; Their versatile nature brings about convenience and organization to your life
cik,accession_number,form_type,filed_atandperiod_of_reportsource_url,retrieved_atandsource_kind(submission metadata, XBRL fact, full filing or extracted section)factsas an array retaining taxonomy, tag, unit, value, instant or start/end dates and filing contextparser_name,parser_versionand normalization or validation warnings
Store the raw SEC value and its unit. Do not convert a missing fact to zero, merge facts from different periods, or discard accession numbers. Keep retrieval timestamps because filings can be corrected or removed after acceptance.
Freshness, corrections and bulk acquisition
The SEC reports typical processing delays of less than one second for Submissions and under one minute for XBRL APIs, with longer delays possible during peak filing periods. These are typical service behaviors, not guarantees. Bulk ZIP files are updated nightly at approximately 3:00 a.m. Eastern Time.
EDGAR guidance explains that accepted filings can be removed or corrected for reasons including a wrong filer, duplicate submission, unreadable content or sensitive information. Run an incremental job that records the last retrieval time, re-fetches changed records, and keeps an audit log of replacements. Do not treat an accession number alone as proof that the content can never change.
Free tools Windows power users keep installed
One-click scans. No signup required.
When the public APIs are not enough
Full text, exhibits and sections
Download the filing archive when you need Management’s Discussion and Analysis, risk factors, exhibits, tables outside standard facts or exact filing language. Parse HTML or inline XBRL locally and associate every extracted section with its filing URL and accession.
Rank #3
Custom taxonomy facts
Company Facts aggregates non-custom taxonomies such as US-GAAP, IFRS, DEI and SRT for the filing entity as a whole. Company-specific extensions require filing-level extraction.
Managed extraction
SEC-API.io documents filing downloads, 10-K/10-Q/8-K section extraction and XBRL-to-JSON conversion. Its extractor accepts a filing URL and item code; unsupported item/form combinations can return errors. Its pricing page accessed September 29, 2026 lists 100 free API calls, Personal & Startups at $49 per month billed annually or $55 month-to-month, and Business Internal Use at $199/$239 on those billing bases. Prices, limits and licenses can change, so verify them before purchase.
Performance and reliability checklist
- Throttle below the 10-request-per-second ceiling and add exponential backoff for transient 429 or 5xx responses.
- Use a persistent cache keyed by URL and response date; avoid downloading unchanged archives.
- Retry with bounded attempts and record failures rather than silently dropping filings.
- Validate that CIK, accession, form and filing date arrays have equal lengths.
- Validate units and period fields before loading facts into a warehouse.
- Use bulk ZIPs for broad historical acquisition, then APIs for near-current updates.
- Keep raw responses so a parser upgrade can be replayed without another SEC download.
Troubleshooting
403 or blocked requests
Cause: missing or vague User-Agent, excessive traffic or bot-like behavior. Fix: identify your application and contact, slow the per-user rate, cache responses and follow the current SEC access guidance.
404 for a CIK endpoint
Cause: the CIK is not ten digits or contains a typo. Fix: zero-pad it and verify the issuer identity in the response.
Rank #4
- Keep important documents safe: A document organizer designed to protect papers from getting lost. Store birth certificates, social security cards, wills, tax forms, insurance policies, titles & more in one secure place.
- Easy to organize and find: Folders with pockets and a table of contents help track where documents live, while 33 hand-illustrated labels show what to save. Acid-free materials protect your papers for years to come.
- Fits documents of various sizes: This document binder includes 3 vertical and 3 horizontal envelopes for 8.5 x 11 inch papers, plus 4 half-size envelopes for smaller keepsakes and important details.
- Practical and easy to use: An important document folder organizer with a front pouch that provides a quick landing space for papers before filing, making it easy to stay organized as documents come in.
- Premium quality, timeless style: Made with custom-dyed cloth, reinforced edges, and acid-free paper for long-term durability. An elegant file organizer designed to beautifully complement your office or living room décor.
A tag is missing from Company Facts
Cause: the issuer used a custom taxonomy, the fact is narrative, or the requested tag is not applicable. Fix: inspect the filing’s inline XBRL and document the taxonomy and context.
Numbers appear duplicated
Cause: multiple forms, units, comparative periods or amended filings. Fix: filter by form and period, retain accession and amendment status, and never deduplicate on value alone.
Section extraction fails
Cause: the provider does not support that item code or form combination. Fix: download the filing and extract it locally, or consult the provider’s supported item list.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, useful when you need a visual capture of an SEC filing page or search result instead of parsing JSON. It accepts a URL in one request and can return PNG, JPEG, WebP or PDF. Cookie banners, popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are not billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.sec.gov -o shot.webp
See the ScreenshotNeo API documentation for capture options and headers. Create a free ScreenshotNeo account to start with 1,000 screenshots per month and no credit card.
Frequently Asked Questions
Do SEC data.sec.gov APIs need an API key?
No. The SEC states that its public data APIs do not require authentication or API keys.
Can Company Facts replace downloading a 10-K?
No. It covers selected standard-taxonomy facts, not complete narrative text, exhibits or every custom-tagged disclosure.
Should I use API JSON or ZIP archives for a large backfill?
Use the nightly companyfacts.zip and submission.zip archives for broad acquisition, then APIs for incremental updates.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

