Automate SEC EDGAR extraction by using the right source for each job: the public SEC JSON APIs for submissions history and standardized XBRL facts, and the original filing documents for narrative text, exhibits, custom-tagged data, or filing-specific context. These public APIs do not require an API key. A reliable pipeline resolves a company to its CIK, retains filing identifiers with every result, respects SEC access limits, and checks extracted values against their source.
Choose the SEC source that matches the data you need
There is no single endpoint that contains every useful detail from every filing. Treat SEC JSON as a structured starting point, not a universal replacement for the filing itself. The SEC describes its public APIs as JSON services without authentication or API keys; these are separate from authenticated EDGAR Next filer APIs used for filer account tasks and submissions.
| Need | SEC route | What to account for |
|---|---|---|
| Find an issuer’s filings | Submissions API | CIK-addressed history includes recent filings; follow referenced historical files for older records. |
| Retrieve standardized financial facts for an issuer | Companyfacts or companyconcept | SEC aggregation covers non-custom taxonomies and facts applying to the filing entity as a whole. It is not a complete substitute for filing-specific context. |
| Compare one fact across issuers and periods | Frames API | Frames are calendar-aligned; inspect reporting dates because issuer fiscal calendars differ. |
| Extract narrative, exhibits, custom tags, or context | Filing archive and document | Retrieve the underlying filing and parse it according to its document structure; validate results against the source. |
| Backfill many issuers or years | SEC bulk ZIPs and indexes | Bulk downloads can reduce individual requests, but check whether their fields and refresh schedule suit the job. |
| Manage a filer account or submit a filing | EDGAR Next filer APIs | These are distinct, authenticated filer tools and are not needed for public filing extraction. |
Resolve the company and enumerate its filings
Use a CIK, not a company name alone
The Central Index Key (CIK) is the SEC’s unique filer identifier. Resolve the issuer to its CIK before making requests; company names and ticker symbols can be ambiguous or change. SEC API paths use a ten-digit, zero-padded CIK, such as CIK0000320193.json. The submissions endpoint pattern is https://data.sec.gov/submissions/CIK##########.json.
Read recent submissions and follow history files
The submissions response contains recent filing arrays, including form type, filing date, accession number, and primary document. Filter these arrays for the form and date range you need. If the target is outside the recent window, inspect the response’s referenced historical submission files and retrieve the relevant one; do not assume the first response contains the issuer’s entire filing history.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
Keep the accession number and primary document name with each selected record. Those values distinguish an exact filing from another filing by the same company and provide the identifiers needed to retrieve and audit the original document.
Retrieve structured financial facts with the XBRL APIs
Choose companyfacts or companyconcept for issuer-level facts
Use companyfacts when you want the issuer’s available standardized facts together, or companyconcept when you need a particular taxonomy concept. Preserve the taxonomy, tag, unit, period, accession or source filing, and any dimensional or contextual information returned. For example, a dollar amount and a share count are not interchangeable merely because both are numeric; units and periods are essential to interpretation.
The SEC’s described company APIs aggregate facts from non-custom taxonomies when they apply to the filing entity as a whole. A custom extension tag, a fact that applies only to a segment, or a fact whose meaning depends on surrounding filing text may not be represented in the way your task needs. When the required value is absent or context-sensitive, go back to the filing rather than treating absence from companyfacts as proof that the company did not report it.
Use frames carefully for cross-company comparisons
The frames route can help retrieve one fact across issuers on a comparable calendar basis. But a calendar frame is not automatically the issuer’s exact fiscal period: the SEC selects facts by closest calendrical fit, and reporting dates can vary. Retain the dates and the source accession, then verify that the observations really represent comparable periods before calculating growth or ranking companies.
Recommended Free Tools
Rank #2
Download the filing when structured data is not enough
For narrative passages, exhibits, custom-tag context, or a traceable source copy, retrieve the filing document from the EDGAR archive using its accession number and document name. SEC indexes identify the company, form, CIK, filing date, and file path; accession numbers identify accepted submissions. The submission record’s primary document name helps locate the main filing, while an index can identify other documents in the submission.
Build your parser around the documents you actually encounter and validate its output. SEC documentation describes access routes, not a universal HTML parser or guaranteed extraction library. Filing documents can contain tables, inline XBRL markup, footnotes, exhibits, and formatting that a simple text scrape can misread. Store the document identity and accession next to every extracted field so a reviewer can navigate back to the source and check the surrounding text.
Python example: retrieve filings and facts
This example uses a known, zero-padded CIK, requests the submissions JSON, selects matching recent filings, and fetches companyfacts. Install the dependency with python -m pip install requests. Replace the CIK and descriptive User-Agent with your own values. It intentionally does not pretend to parse every filing format.
import time
import requests
CIK = "0000320193" # Replace with the issuer's 10-digit CIK
BASE = "https://data.sec.gov"
HEADERS = {
"User-Agent": "Example Research [email protected]",
"Accept-Encoding": "gzip, deflate",
}
session = requests.Session()
session.headers.update(HEADERS)
last_request = 0.0
def get_json(url):
global last_request
# This process stays below the SEC's 10 requests/second per-user guideline.
delay = 0.2 - (time.monotonic() - last_request)
if delay > 0:
time.sleep(delay)
response = session.get(url, timeout=30)
last_request = time.monotonic()
response.raise_for_status()
return response.json()
submissions_url = f"{BASE}/submissions/CIK{CIK}.json"
submissions = get_json(submissions_url)
recent = submissions["filings"]["recent"]
wanted_form = "10-K"
for i, form in enumerate(recent["form"]):
if form == wanted_form:
print({
"form": form,
"filing_date": recent["filingDate"][i],
"accession_number": recent["accessionNumber"][i],
"primary_document": recent["primaryDocument"][i],
})
facts_url = f"{BASE}/api/xbrl/companyfacts/CIK{CIK}.json"
facts = get_json(facts_url)
print("Entity:", facts["entityName"])
print("Taxonomies:", ", ".join(facts.get("facts", {}).keys()))
The loop selects only the recent array. To cover older years, read the submission response’s filings.files references and fetch the relevant history JSON before applying the same form/date filtering. For a concept-level fact request, use the companyconcept route described by SEC API documentation; for a cross-issuer calendar-aligned fact, use frames and inspect its dates. A bulk backfill may be better served by the SEC’s submissions or companyfacts ZIP than by issuing many individual requests.
Rank #3
Access discipline, freshness, and data quality
Identify and pace the client
SEC developer guidance last reviewed March 10, 2025 set a maximum of 10 requests per second per user, across all machines, and calls for efficient, identifiable access. Include a meaningful User-Agent that identifies your application and provides a contact address. The example uses 5 requests per second in one process; multiple workers or services sharing your identity must share a coordinated limit, rather than each assuming it owns the full allowance. Recheck current SEC guidance before deploying because access policy can change.
For production, add bounded retries with exponential backoff for transient network or server failures, cache responses where appropriate, and monitor status codes, response times, and incomplete jobs. Do not retry a denied or malformed request indefinitely. Avoid unclassified crawling: request only the files your workflow needs and use bulk resources for broad acquisition where practical.
Do not assume every endpoint is instantly current
As described on the SEC API page last reviewed April 8, 2025, submissions processing is typically under a second and XBRL processing typically under a minute, but delays may be longer during peak filing times. Those are typical processing times, not availability or freshness guarantees. The companyfacts and submissions bulk ZIPs are republished nightly at approximately 3:00 a.m. ET, so they are not substitutes for a current individual lookup when timing matters.
SEC-accessible filing data can be corrected or removed after acceptance, and indexes incorporate updates on their rebuild schedules. Keep enough source metadata to reconcile stored records with updated source indexes; avoid treating an accepted filing record as permanently immutable.
Rank #4
Validate periods, units, and provenance
- Preserve units and reporting periods with numeric facts; never compare raw values without confirming both.
- For frames, compare reported dates and fiscal periods rather than assuming a calendar frame exactly matches every issuer’s fiscal year.
- For document-derived values, retain accession and document identity, then verify the value in its source context.
- Record retrieval time and response status so delayed or corrected upstream data can be diagnosed.
Put a server between a browser and data.sec.gov
The SEC states that data.sec.gov does not support CORS. A browser page calling the API directly from a different origin may therefore be blocked even when the endpoint is public. A common architecture is a server-side retrieval service that makes SEC requests, applies shared pacing and caching, and returns only the fields the front end needs. The server still has to follow SEC access guidance; moving the request does not remove the per-user limit.
Troubleshooting common extraction failures
- No filing appears in the recent response: check that the CIK is correct and zero-padded, then inspect the historical submission-file references. The recent arrays are not necessarily the full history.
- A desired metric is missing from companyfacts: determine whether it uses a custom taxonomy, applies only to a segment, or depends on filing context. Retrieve the source filing and inspect its facts and text.
- A value looks wrong across companies: check taxonomy, tag, unit, dimensions, reporting period, and frame dates. Calendar alignment does not guarantee fiscal-period equivalence.
- Requests fail or are slowed: verify the endpoint and CIK, send a meaningful User-Agent, reduce aggregate request rate across all machines, and use bounded backoff for transient failures.
- A browser fetch is blocked: route the request through a server-side service because data.sec.gov does not support CORS.
- A stored filing differs from the current index: account for post-acceptance corrections or removals and reconcile records against updated SEC source data.
Or skip the browser setup
ScreenshotNeo is not an SEC filing parser and does not replace the JSON APIs or archive documents. If your workflow also needs a rendered visual reference of a public SEC page, its screenshot API can capture that page separately; use the API to extract and validate filing data. One GET request produces an image or PDF, and the service offers an MCP server for AI agents.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.sec.gov/ -o shot.webp
See the ScreenshotNeo documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for ScreenshotNeo free.
Keep the extraction auditable
A robust EDGAR pipeline keeps the original filing identifiers alongside its parsed output, chooses structured APIs only where their scope fits, and makes its request behavior visible and controlled. For broad history, assess bulk ZIPs; for filing-specific meaning, return to the source document. That separation makes it easier to detect missing facts, interpret periods correctly, and reproduce how a value was obtained.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsFrequently Asked Questions
Do I need an SEC API key to read public filings?
No. The SEC says its public data.sec.gov APIs do not require authentication or API keys.
Can I use the public SEC APIs to submit a filing?
No. Public extraction endpoints and authenticated EDGAR Next filer APIs serve different purposes; submission actions belong to the filer APIs.
Can a screenshot replace an SEC filing parser?
No. A screenshot records a rendered page visually; it does not provide the structured filing metadata, XBRL facts, or validated document extraction described in this workflow.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

