Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use GitHub’s API for supported, permissioned collection, and reserve HTML scraping for a documented, policy-compliant need. An AI agent should identify the exact endpoint, authenticate with the least privilege, follow the response’s pagination links, respect primary and secondary limits, and verify every result before it writes code, opens an issue, or changes a repository. GitHub’s REST API is an HTTP request made from a method and path, with endpoint-specific headers, parameters, and (for some methods) a request body. Start with the official getting-started guide and endpoint reference rather than guessing URLs.
API collection and website scraping are different
GitHub’s acceptable-use policy defines scraping as automated extraction from the service by a bot or web crawler. It explicitly says, “Scraping does not refer to the collection of information through our API.” API calls are governed by GitHub’s API terms; downloading and parsing website HTML is governed by the applicable website policies, terms, privacy requirements and other agreements. That distinction does not make every automated use permissible.
The policy lists examples such as research using public, non-personal information when resulting publications are open access, and archival use. It prohibits uses including spam and selling personal information, and requires compliance with GitHub’s Privacy Statement. Check the current Acceptable Use Policies, Terms of Service, repository licenses, privacy obligations and your customer agreements before deployment. No source can determine the legal status of every purpose or jurisdiction.
Choose the operation and endpoint first
Write the agent’s task as a specific operation: list repositories for an organization, read issues, inspect a file, create a pull request, or update a label. The endpoint reference tells you the HTTP method, path, required headers, authentication, query parameters and body.
#1 Best Overall
- GET retrieves a resource.
- POST creates one.
- PATCH updates selected properties.
- PUT replaces a resource or collection.
- DELETE removes a resource.
Use the documented endpoint instead of scraping a page that happens to display the same data. A client library changes syntax, not permissions, policy obligations or rate limits. GitHub’s getting-started documentation covers curl, GitHub CLI and JavaScript; the examples below use curl, Python and Node.js so the same request can be tested outside an agent.
Authenticate with the smallest useful permission
Most requests should send Accept: application/vnd.github+json, an explicit X-GitHub-Api-Version, and a valid User-Agent. GitHub states: “All API requests must include a valid User-Agent header.” The current documentation example uses API version 2026-03-10; confirm the supported version when you implement.
Authentication requirements are endpoint-specific. GitHub recommends a fine-grained personal access token for personal use when possible, and a GitHub App for organizational or on-behalf-of-user integrations. In Actions, use the built-in GITHUB_TOKEN when it is suitable and set workflow permissions explicitly. Treat tokens as passwords: keep them in a secret manager or environment variable, never in prompts, logs, source control or browser code. A read-only agent should not receive write permissions.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Minimal curl request
export GITHUB_TOKEN='replace-me'
curl --fail-with-body
-H 'Accept: application/vnd.github+json'
-H 'X-GitHub-Api-Version: 2026-03-10'
-H 'User-Agent: my-github-agent/1.0'
-H "Authorization: Bearer $GITHUB_TOKEN"
'https://api.github.com/repos/octocat/Hello-World/issues?state=open&per_page=100'
Use the endpoint’s documented authentication scheme. If an endpoint permits public unauthenticated access, that may still be a poor production choice because the published limit is much lower.
Rank #2
Fetch every page, not just the first response
List endpoints commonly return a partial result. GitHub’s pagination example returns 30 issues by default even though its example repository has more than 1,600 open issues. The response’s Link header can include next, prev, first and last URLs. Follow the returned next URL rather than constructing page numbers yourself. Most endpoints allow a maximum per_page of 100, but defaults and maxima vary.
Python pagination loop
import os, time, requests
TOKEN = os.environ["GITHUB_TOKEN"]
url = "https://api.github.com/repos/octocat/Hello-World/issues"
params = {"state": "open", "per_page": 100}
headers = {
"Accept": "application/vnd.github+json",
"X-GitHub-Api-Version": "2026-03-10",
"User-Agent": "my-github-agent/1.0",
"Authorization": f"Bearer {TOKEN}",
}
items = []
while url:
response = requests.get(url, headers=headers, params=params, timeout=30)
response.raise_for_status()
items.extend(response.json())
# Parameters are already present in the Link URLs returned by GitHub.
url = response.links.get("next", {}).get("url")
params = None
time.sleep(0.2) # Keep a serial, moderate request pace.
print(f"collected {len(items)} issues")
Record the endpoint, retrieval time, page URLs or cursors, item count and any termination reason in the agent’s internal result. That provenance lets a later step distinguish a complete traversal from a one-page sample.
Octokit’s helper
import { Octokit } from "@octokit/rest";
const octokit = new Octokit({ auth: process.env.GITHUB_TOKEN });
const issues = await octokit.paginate(octokit.rest.issues.listForRepo, {
owner: "octocat", repo: "Hello-World", state: "open", per_page: 100
});
Octokit’s paginate() helper follows supported paginated responses. It does not remove the need to handle errors, limits or permissions.
Recommended Free Tools
Rate limits and resilient request handling
GitHub’s published primary REST limits, reviewed in the documentation on September 29, 2026, are 60 requests per hour for unauthenticated requests to public data and 5,000 requests per hour for authenticated users. These are current published limits, not permanent guarantees. Search endpoints have more restrictive limits, GraphQL has separate accounting, and secondary limits can apply to any integration.
- Read
x-ratelimit-remaining,x-ratelimit-resetand, when present,retry-afteron every response. - If remaining is zero, wait until the reset time. If
retry-afteris supplied, wait that duration. - For a secondary limit without either indicator, wait at least one minute, then increase the delay exponentially after repeated failures.
- Stop after a bounded number of retries. Continuing while limited can lead to an integration ban or suspension.
Prefer webhooks when an event model fits instead of frequent polling. If polling is necessary, request only fields you need, send authenticated conditional requests with ETag or If-Modified-Since, and keep requests serial. GitHub says a correctly authorized conditional GET returning 304 Not Modified does not count against the primary rate limit. Cache results with a clear freshness policy and do not use multiple tokens to evade limits; GitHub’s API terms prohibit sharing tokens for that purpose.
Give an AI agent a safe GitHub tool boundary
Expose narrowly scoped functions rather than an unrestricted HTTP client. A useful read-only tool might accept owner, repo, state and a bounded limit, validate those values, call the documented endpoint, and return both data and provenance. Keep mutations separate.
Recommended controls
- Declare whether a tool is read-only or mutating in its schema and system prompt.
- Grant only the endpoint permissions required; use a GitHub App installation permission or fine-grained token that excludes unrelated repositories.
- Show the target owner, repository, branch and proposed change to a human before consequential mutations.
- Require the agent to quote the source URL, retrieval time and completeness status in its answer.
- Validate JSON schema, repository identity, file paths and allowed operations before execution.
- Use a dry-run mode for issue, branch, release and pull-request actions.
GitHub’s Terms of Service warn that AI output can be inaccurate, incomplete or non-functional and may resemble third-party code, including code under open-source licenses. The terms state: “You are responsible for reviewing, testing, and validating any Output before use.” Apply that review before merging code or publishing generated analysis.
REST, GraphQL, polling and webhooks
| Choice | Use it when | Important constraint |
|---|---|---|
| REST | The endpoint reference directly matches the resource or action. | Follow its pagination and endpoint-specific permissions. |
| GraphQL | You need a tailored shape across related resources and understand GraphQL’s cost model. | It has separate limits and query complexity considerations. |
| Polling | No suitable event exists and a bounded freshness interval is acceptable. | Use conditional requests and a conservative interval. |
| Webhooks | You need event-driven updates for supported events. | Secure delivery and process retries and duplicate events. |
When HTML scraping is unavoidable
First check whether the needed data has a REST or GraphQL endpoint. If it does, use that interface. If a documented, permitted purpose still requires HTML, identify the exact pages, avoid login circumvention, honor applicable robots, privacy and contractual requirements, throttle requests, cache responses, and collect the minimum data. Do not infer that a page is public merely because a browser can view it, or that public visibility authorizes bulk extraction.
HTML layouts change without API-version guarantees. Build selectors with tests, detect login pages and consent interstitials, cap concurrency, and stop on repeated errors. Never collect personal data for spam or sell it, and review repository licenses before redistributing code or content.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common failures and fixes
401 or 403 responses
Check the token, expiration, endpoint permissions, repository access and authorization header. A valid token with insufficient fine-grained permissions still fails. A 403 can also indicate a rate or secondary limit; inspect headers before retrying.
422 validation errors
Read the response body’s field-level message. Typical causes are an invalid owner or repository, unsupported parameter combination, missing required body field or a duplicate operation. Correct the request instead of retrying unchanged.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Only the first 30 or 100 records appear
Your code probably ignored the Link header or stopped at a configured cap. Follow next until it disappears, and report any deliberate cap to the agent and user.
Best Value
429 or secondary-limit behavior
Honor retry-after; otherwise wait at least a minute, back off exponentially and stop after bounded retries. Reduce parallelism and polling frequency.
Agent proposes an unsafe change
Separate read and write tools, display the exact diff and target, run tests, and require human approval. Do not let a model turn an ambiguous instruction into an irreversible mutation.
Or skip the browser setup
If your agent also needs a rendered website image—for example, to inspect a README demo or capture a report—ScreenshotNeo provides a one-request screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
Free tools Windows power users keep installed
One-click scans. No signup required.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://github.com/octocat/Hello-World -o shot.webp
See the ScreenshotNeo API documentation for options such as full-page capture, CSS selectors, device presets, custom headers and cookies, JavaScript, blocking resources, PDFs, signed links, asynchronous jobs and bulk capture. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Operational checklist
- Endpoint and method match the documented operation.
- Token type and permissions are minimal and stored as secrets.
- Required
Accept, API-version andUser-Agentheaders are present. - Pagination follows returned links and records completeness.
- Primary, search and secondary limits are monitored.
- Conditional requests, webhooks or caching reduce unnecessary polling.
- Scraping purpose, privacy, licenses and agreements have been reviewed.
- Agent output is validated before any consequential action.
Frequently Asked Questions
Can an AI agent use GitHub’s API without a personal access token?
Some public-data requests can be unauthenticated, but GitHub publishes a much lower primary limit for them. Use authentication when the endpoint and purpose justify it, with the least privilege needed.
Should I use REST or GraphQL for an agent?
Use the interface that matches the data shape and permissions. REST is usually simpler for one documented operation; GraphQL can reduce round trips for related fields but has separate limits and query-cost concerns.
Does following GitHub links in a browser count as scraping?
The policy’s definition concerns automated extraction. A human browsing is different, but an automated browser that extracts pages can fall within the scraping definition; consult the live policy and your agreements.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

