Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThe safest way to automate government-data retrieval is to use the agency’s documented API first, then an official bulk file or download, and only then permitted HTML retrieval. Before scheduling requests, read the service’s terms, the dataset’s Access and Use Information, authentication requirements and rate limits. A public webpage is not, by itself, permission to scrape.
Choose the official interface before writing a scraper
Start at the authoritative agency or portal page for the dataset. Record the publisher, dataset name or identifier, update schedule and the page that defines access conditions. Then look for these interfaces in order:
| Method | Use it when | Checks before automating |
|---|---|---|
| Official API | The service documents endpoints and supports the filters, fields and update cadence you need. | Authentication, terms, quota, pagination, response format and API version. |
| Bulk extract or direct file | You need a large, stable snapshot or the publisher provides a ready-made CSV or JSON file. | Format, update schedule, file size, license and whether incremental files exist. |
| HTML retrieval | No suitable structured interface exists and the page permits automated access. | Terms, robots.txt, authentication, crawl limits, page stability and technical controls. Never bypass a block. |
Data.gov documents APIs for dataset search and metadata. Some federal programs also publish bulk CSV or JSON, including the Site Scanning Program. The best method depends on freshness, completeness, query flexibility, stability and operational effort—not on whether a page happens to be easy to download.
Confirm permission and data-use conditions
Read service terms and dataset notices
Data.gov says that, in most cases, U.S. federal data available through it is free and without restriction, but it directs users to each dataset’s Access and Use Information. Non-federal datasets can carry different licenses. Treat those dataset-specific terms as controlling.
#1 Best Overall
Policies can prohibit automation even when information is visible in a browser. SAM.gov’s terms state: “Automated data gathering, web scraping tools are prohibited and, if detected, will result in the associated account(s) being denied access to SAM.gov via Login.gov.” That restriction applies to SAM.gov; it is not a rule for every government website.
Use robots.txt as guidance, not authorization
Digital.gov explains that robots.txt communicates crawler instructions, while noting that bad bots may ignore them. Check it, but also read the terms, API documentation and any authentication or anti-automation notice. Robots.txt does not grant permission or override an explicit prohibition.
Handle restricted data conservatively
- Do not automate around login challenges, CAPTCHAs, bot checks, paywalls or other technical controls.
- Do not reuse a personal account’s session or credentials outside the documented process.
- Ask the agency for an approved feed when the terms are unclear.
- Keep a copy of the terms and the date you reviewed them for your project record.
Build a reproducible retrieval workflow
- Identify the source. Record the agency, dataset identifier, publisher and canonical API or download location.
- Select the interface. Prefer an API for incremental queries, a bulk file for repeatable snapshots, or HTML only when permitted and necessary.
- Read current documentation. Note required parameters, pagination, field definitions, update cadence, authentication and version changes.
- Obtain credentials correctly. Create the documented API key or account, store secrets in environment variables or a secret manager, and never commit them to source control.
- Retrieve in bounded batches. Set explicit page sizes, timeouts and a maximum number of pages. Save the raw response before transforming it.
- Respect quotas. Inspect response headers, pause on throttling and use exponential backoff with a cap.
- Validate. Check HTTP status, schema, required fields, record counts, duplicate identifiers and date ranges.
- Preserve provenance. Store retrieval time, endpoint or file URL, parameters, dataset publication/version details and every transformation step.
- Schedule against the publisher’s cadence. Polling more often than the source changes wastes quota and can create unnecessary load.
API retrieval: a safe Python pattern
The following client is deliberately endpoint-neutral. Replace API_ENDPOINT and parameter names with those in the target agency’s documentation. It handles pagination, timeouts, transient failures and a bounded request count while keeping the key out of source code.
Rank #2
- FIND ANY PAPER IN SECONDS: Color-coded tabs and a blank label sheet let you sort up to 24 categories by class, client, or month, then flip straight to what you need. Write-and-erase tabs make relabeling instant when projects change.
- BUILT FOR A FULL SCHOOL YEAR: Tear-resistant covers, acid-free construction, and an oversized coil spine hold heavy paper loads without splitting or distorting. Two elastic straps lock everything shut so nothing slides out in a backpack or work bag.
- STANDARD PAGES SLIDE RIGHT IN: Each of the clear pockets fits 8.5 x 11 inch sheets without bending corners. Push papers all the way to the back edge and they stay flat every time you close the cover.
- REPLACES A BINDER AND NOTEBOOK: Works as a teacher binder, an IEP organizer for teachers, or a homeschool organization hub without hole-punching a single page. Slip syllabi, report cards, or lesson plans in and carry one item instead of three.
- EXTRAS ALREADY INCLUDED: A clear zippered utility pouch holds pens, note cards, and stencils. The customizable front cover has a non-glare overlay, and a clear back pocket lets you see loose items at a glance.
import json
import os
import time
from datetime import datetime, timezone
import requests
ENDPOINT = os.environ["API_ENDPOINT"]
API_KEY = os.environ.get("API_KEY")
MAX_PAGES = 100
PAGE_SIZE = 100
session = requests.Session()
params = {"page": 1, "page_size": PAGE_SIZE}
if API_KEY:
params["api_key"] = API_KEY
records = []
for _ in range(MAX_PAGES):
for attempt in range(5):
try:
response = session.get(ENDPOINT, params=params, timeout=30)
if response.status_code == 429:
delay = min(60, 2 ** attempt)
time.sleep(delay)
continue
response.raise_for_status()
break
except requests.RequestException:
if attempt == 4:
raise
time.sleep(min(60, 2 ** attempt))
payload = response.json()
batch = payload.get("results", payload if isinstance(payload, list) else [])
if not isinstance(batch, list):
raise ValueError("Unexpected response shape")
records.extend(batch)
next_url = payload.get("next") if isinstance(payload, dict) else None
if not next_url:
break
params = {} # use the documented next link exactly
metadata = {
"retrieved_at": datetime.now(timezone.utc).isoformat(),
"endpoint": ENDPOINT,
"record_count": len(records),
}
with open("records.json", "w", encoding="utf-8") as f:
json.dump({"metadata": metadata, "records": records}, f, indent=2)
Some APIs return a cursor rather than a page number; follow the documented cursor field. If a response supplies a complete next link, use it instead of guessing query parameters. Add schema validation appropriate to the dataset before loading records into a database.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Equivalent command-line and Node.js requests
cURL
curl --fail-with-body --retry 4 --retry-delay 2
-H "Accept: application/json"
-G "$API_ENDPOINT"
--data-urlencode "api_key=$API_KEY"
--data-urlencode "page=1"
--data-urlencode "page_size=100"
-o page-1.json
Use the agency’s documented authentication parameter or header; remove api_key if the service uses another method. Do not put a real key in shell history on shared systems.
Node.js
const endpoint = process.env.API_ENDPOINT;
const key = process.env.API_KEY;
const url = new URL(endpoint);
url.searchParams.set('page', '1');
url.searchParams.set('page_size', '100');
const headers = { Accept: 'application/json' };
if (key) headers.Authorization = `Bearer ${key}`;
const res = await fetch(url, { headers, signal: AbortSignal.timeout(30000) });
if (!res.ok) throw new Error(`${res.status}: ${await res.text()}`);
const data = await res.json();
console.log(JSON.stringify(data));
Rate limits, retries and reliability
Data.gov’s undated live guidance lists 1,000 requests per hour for a personal API key. Its DEMO_KEY is limited to 30 requests per IP per hour and 50 per IP per day. The api.data.gov developer manual describes a default limit of 1,000 requests per hour per API key, while warning that limits can vary by service. Check the target service’s response headers and documentation rather than assuming these figures apply everywhere.
Rank #3
- Great way to organize and store vital tax records
- Instruction sheet/checklist and preprinted labels included
- 12 pockets plus one large pocket in back provides ample storage
- Protective flap and elastic cord closure
- Contains 10% recycled content, 10% post-consumer material
- Honor
Retry-Afterwhen supplied. - Use exponential backoff with jitter for 429 and temporary 5xx responses.
- Do not retry authentication errors, malformed requests or explicit denials.
- Cache unchanged pages and use incremental date or identifier filters when offered.
- Set a maximum runtime and page count so a changed pagination response cannot run forever.
- Log status code, request ID, page/cursor, elapsed time and retry count without logging secrets.
Policies are service-specific. The UK National Archives, for example, publishes a limit of 3,000 requests in any five-minute period for its website and catalogue data. That figure is not a general government-site rule.
Bulk files and downloads
A bulk CSV or JSON extract is often the most reliable choice for a full refresh. Verify the file’s update schedule, encoding, delimiter, schema, compression, size and license. Download to a temporary name, verify the transfer, then rename atomically so downstream jobs never read a partial file. Keep the original file and checksum when reproducibility matters. If incremental files or change dates exist, use them instead of repeatedly downloading the entire archive.
When permitted HTML retrieval is unavoidable
Use a session with a descriptive user agent, conservative concurrency, caching and a clear stop condition. Review robots.txt and crawl-delay guidance, but treat the terms and access controls as decisive. Parse stable semantic elements rather than brittle screen coordinates, and expect markup, pagination and labels to change. Store the page URL, retrieval time and parser version with each record. Stop when the service blocks, denies or asks you to stop; do not rotate identities or evade controls.
Rank #4
- ENHANCED ORGANIZATION: Organize your paperwork with this letter-sized (10.25” x 11.75”) document organizer with 24 pockets and 12 dividers; our pocket organizer is a great choice for school supplies college folders with pockets and bible study supplies
- EFFORTLESS SORTING: This plastic folder organizer with 24 pockets provides ample space to sort and categorize your materials, ensuring easy access and efficiency; 1/3-cut reusable write & erase tabs provide three positions for convenient labeling and easy identification
- PRACTICAL DESIGN: The slash pockets can hold up to 25 sheets each; the spiral-bound design allows the office supply organizer to lay flat for convenience and rotate 360° for easy viewing; tear-resistant and water-resistant poly cover material ensures long-lasting durability
- COLOR-CODED ORGANIZATION: The 12 colorful dividers in six colors boldly split up subjects while the clear front pocket allows you to customize your organizer with a cover sheet; keep essentials in the zippered pouch for quick access
- PVC AND ACID FREE: This organizer reflects our commitment to environmental responsibility; it's acid-free and PVC-free, making it safe for long-term document storage
Validate and monitor the pipeline
- Compare record counts with the publisher’s stated totals where available.
- Reject impossible dates, missing primary identifiers and duplicate keys.
- Alert on schema changes, empty responses, sudden count drops and repeated throttling.
- Keep raw and normalized layers so a parser change can be replayed.
- Recheck terms, API versions, limits and update schedules periodically because they can change without your code changing.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Missing, expired or incorrectly placed credentials; access is restricted. | Follow the documented authentication flow, check account permissions and do not bypass the restriction. |
| 429 | Quota or burst limit exceeded. | Honor Retry-After, slow concurrency, cache results and request a higher documented quota if available. |
| Empty or partial data | Wrong page, cursor, filter or date range. | Inspect pagination metadata, test a small unfiltered request and validate the response schema. |
| HTML instead of JSON | Redirect, login page or endpoint mismatch. | Check final URL and Content-Type, then use the documented API endpoint. |
| Parser breaks | Page markup changed. | Prefer an API or bulk file; otherwise update selectors and add fixture-based tests. |
| Timeouts or 5xx | Transient service or network failure. | Use bounded retries with backoff, shorter pages and a resumable checkpoint. |
Or skip the browser setup
If your workflow needs a visual record of a government page, ScreenshotNeo can capture it without you maintaining a headless browser. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before the shot. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and whether it was billed. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for capture options such as full-page or CSS-selector shots, waiting rules, custom headers, cookies, JavaScript, blocking, device presets, PDFs, signed links, asynchronous jobs and bulk capture. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I automate any public government page?
No. Public visibility does not establish permission. The service’s terms, dataset license, authentication rules and technical controls determine what is allowed.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesShould I use an API key in a scheduled job?
Yes, when the service documents key-based access. Store it in a secret manager or environment variable, restrict its permissions and rotate it according to the agency’s guidance.
Best Value
- NOT A FLIMSY IMPORT: Doctor Stuff's 11pt Orange File Folders are USA Made, featuring a heavyweight design with 30% more paper weight compared to competitors that import. Durability, longevity and resilience in busy office environments.
- MEDICAL FILE ORGANIZATION: Our sturdy, full-cut end tab medical file folders are designed for shelf filing, ensuring easy access to crucial information. Long lasting reliability for healthcare and other filing professionals.
- LOOKS AND FEELS LIKE A FOLDER: American manufactured means that we use more paper and less air - 100 plain 11pt folders weigh 7.7 lbs compared to 5.9 lbs for imported competitors. They feel like real folders.
- PACKAGE INCLUDES: A box of 100 orange chart folders. Our durable folders will effectively organize 8½”x11” files and ideal for legal, healthcare, educational government and others that value quality.
- TRUSTED BY PROFESSIONALS: Doctor Stuff is synonymous with excellence in organizational supplies. Our Orange end tab file folders with prongs are designed to meet the exacting standards of professionals who require the best in document management and security.
How often should a job run?
Match the publisher’s update cadence and your data requirement. More frequent polling rarely improves a dataset that changes daily or weekly.
What should I retain for an audit?
Keep the source and dataset identifier, retrieval timestamp, request parameters, publication/version details, raw response or file, terms reviewed and transformation code version.
Frequently Asked Questions
Can I automate any public government page?
No. Public visibility does not establish permission. The service’s terms, dataset license, authentication rules and technical controls determine what is allowed.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Should I use an API key in a scheduled job?
Yes, when the service documents key-based access. Store it in a secret manager or environment variable, restrict its permissions and rotate it according to the agency’s guidance.
How often should a job run?
Match the publisher’s update cadence and your data requirement. More frequent polling rarely improves a dataset that changes daily or weekly.
What should I retain for an audit?
Keep the source and dataset identifier, retrieval timestamp, request parameters, publication/version details, raw response or file, terms reviewed and transformation code version.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

