Free tools Windows power users keep installed
One-click scans. No signup required.
Web scraping can support social-media OSINT, but a post being publicly visible does not by itself authorize automated collection. Start by checking the platform’s terms, robots.txt, API or export rules, permissions, and rate limits; use an official API or permitted export when it can answer the question. Then collect only what you need, preserve original URLs and timestamps, hash and retain the original artifacts, and document how you gathered and changed the data. If you cannot establish permission or a lawful basis, do not scrape.
What social-media scraping means in an OSINT investigation
Social-media scraping is automated extraction of information from a social platform or its web pages for an open-source investigation. It may collect posts, profile details, timestamps, links, or other visible information, depending on the permitted access method and what the platform exposes. It is not the same as manually viewing a page, and it is not automatically authorized just because a person can see the page in a browser.
Meta describes scraping as automated collection from a website or interface and distinguishes authorized activity, such as search-engine crawling, from unauthorized activity that violates its terms. That distinction matters: the method, permission, purpose, and platform rules all affect whether a particular collection is appropriate. An investigator should not treat public visibility as a blanket grant to automate access, compile personal information, or retain it indefinitely.
OSINT is also not synonymous with scraping. It is the broader practice of analyzing information that is openly available or otherwise lawfully accessible. An official API, a platform-provided export, a search interface, a browser capture, or independent public records may be better suited to a given question than a scraper.
#1 Best Overall
Check permission and privacy obligations before collecting
Before writing a script or running a collection tool, review the platform’s current terms of service, robots.txt, API documentation, export options, access permissions, and stated rate limits. These are distinct controls: compliance with one does not cancel another. Robots.txt communicates a publisher’s preferences to crawlers; terms and API rules may impose separate conditions, while CAPTCHAs, authentication, paywalls, and other technical restrictions indicate that access is limited.
Google says it honors open-web standards such as robots.txt. Treat that as an important publisher signal, not as a complete legal analysis or a substitute for the platform’s terms. Follow explicit exclusions, rate limits, and API conditions as well. Do not bypass authentication, CAPTCHAs, paywalls, or technical blocks to obtain data. If a collection is blocked or permission is uncertain, stop and seek an authorized route rather than trying to evade the control.
Public does not mean unrestricted
Privacy obligations can apply to information visible to the public. Canadian privacy regulators say organizations need a lawful basis and transparency, and should obtain consent where required. CNIL’s focus sheet dated January 5, 2026 says publicly accessible personal-data scraping generally rests on legitimate interest but requires additional measures to protect people’s rights and freedoms. The EDPB’s 2026 guidance update addresses GDPR legal bases and special-category data. These sources describe regulatory positions, not a universal permission rule for every jurisdiction or case.
Before collection, identify the relevant geography, the people whose data may be involved, and the legal regime that applies to your organization and investigation. Document the purpose and lawful-basis assessment, minimize sensitive or irrelevant attributes, restrict access, set retention and deletion rules, and provide notice or a contact route when required. If the investigation could involve sensitive personal data, children, vulnerable people, or a high-impact decision, obtain appropriate legal and privacy review before proceeding.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRobots.txt and platform rules are not interchangeable
- Terms of service: Check restrictions on automated access, reuse, account activity, and data retention.
- Robots.txt: Respect the site’s crawler directives and any more specific published exclusions.
- API or export terms: Confirm permitted fields, use, rate limits, historical access, and attribution requirements.
- Access controls: Do not defeat CAPTCHAs, logins, paywalls, or technical restrictions.
- Privacy rules: Establish a lawful basis and safeguards independently of whether the page is publicly visible.
A defensible social-media OSINT collection workflow
- Define the question and scope. Write down the investigative purpose, target platform or pages, date range, geography, relevant identifiers, and the minimum data needed to answer the question. Set exclusions at the outset, especially for sensitive or unrelated personal information.
- Check the permitted access route. Review current terms, robots.txt, API documentation, export options, account permissions, and rate limits. Prefer an official API or explicitly permitted export when it provides the needed fields and time coverage.
- Run a small, documented pilot. Test the smallest collection that can show whether the chosen route works. Record the query, page or account context, collection time, tool and version, and any access errors. Do not scale a test until its permission and privacy assumptions are clear.
- Set conservative request behavior. Identify the crawler transparently in its user-agent, batch requests where permitted, pace them to the platform’s rules, and implement error logging and backoff for temporary failures. AWS gives one request every 10–15 seconds for small or medium sites and 1–2 requests per second for larger sites where explicit permission exists as operational examples, not universal legal limits. The platform’s own rules take precedence.
- Preserve originals and provenance. Save the original URL, capture time, page or post identifier, original downloaded artifact, hash, and notes about the collection method. Record account or page context and the relevant query so another reviewer can understand what was accessed.
- Separate collected material from analysis. Keep raw captures unchanged. Store annotations, translations, classifications, redactions, and deduplication separately, documenting each transformation and the rules used. This makes it possible to distinguish what the source showed from what the analyst inferred.
- Control access and retention. Limit access to people who need it, maintain an access record when appropriate, apply the retention schedule, and delete material when the lawful purpose or retention period ends.
- Corroborate and report limits. Check important claims against independent sources. Note missing pages, inaccessible posts, edits, deletions, collection gaps, and uncertainty; a captured post establishes what was visible to the collector at a particular time, not necessarily that its claims are true.
Choose the collection method that fits the question
Compare methods by coverage, freshness, reproducibility, rate limits, privacy risk, terms compliance, evidence integrity, cost, and collaboration needs. No method is best for every investigation. Official APIs commonly provide structured fields and a clearer access contract, but may limit historical depth or available data. Browser or page capture can preserve what an investigator saw, but may be affected by dynamic rendering, account permissions, terms, and differences between visits.
| Approach | Useful when | Trade-offs to assess |
|---|---|---|
| Official API | You need structured fields and the platform offers an authorized endpoint for the task. | Check field coverage, historical depth, quotas, freshness, and permitted downstream use. |
| Platform export | The platform provides an export that covers the relevant account or data scope. | Check export contents, time span, format, provenance, and whether it answers the investigative question. |
| Permitted crawler or scraper | Automated page collection is explicitly allowed and an API or export is insufficient. | Assess terms, robots.txt, rate limits, privacy risks, dynamic pages, failure handling, and reproducibility. |
| Browser capture | You need to preserve the page as it appeared to an authorized viewer at a specific moment. | It records a presentation rather than guaranteeing structured coverage; rendering, account state, and later changes can affect repeatability. |
Use a capture-focused tool when provenance and review of what a person could see are central. Hunchly says it automatically collects URLs, timestamps, and hashes of pages visited and makes full-page captures of sites, searches, and social media; it also describes tagging and searching captures and preparing packages with an audit trail. Hunchly reports investigators and researchers in 84 countries, a vendor figure for which its product page does not provide a methodology or denominator. Verify current availability, terms, and fit for your case before adopting it.
Rank #3
Maltego’s official documentation describes an investigation platform with Maltego Search, Graph, Cases, Data, Monitor, Evidence, and Hunchly integration. It says the platform can support real-time social-data monitoring, OSINT searches, evidence gathering, and analysis of complex cases. Consider it when relationship mapping, monitoring, or team case management is central; a capture-focused tool may be a closer fit when the primary need is preserving provenance.
Preserve social-media evidence so another person can verify it
A screenshot alone can be useful, but it is stronger when accompanied by the source URL, collection time, page or post identifier, original artifact, and a cryptographic hash. Keep a contemporaneous collection note that identifies the tool and version, query, account or page context, and any settings that could affect what appeared. Store raw captures separately from commentary, and log changes such as redaction, conversion, extraction, or deduplication.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →- URL and identifier: Preserve the exact source address and the platform’s page, post, or content identifier where available.
- Time: Record the capture timestamp, including the time zone or UTC offset.
- Original artifact: Retain the downloaded page, export, or capture in its original form and keep a working copy for analysis.
- Integrity: Compute and record a hash for each original artifact; do not imply that a hash proves the source’s truth, only that the file can later be compared for changes.
- Chain of custody: Note who collected, transferred, accessed, or transformed the evidence and when.
- Context and limits: Preserve enough page or account context to explain what was viewed, and document inaccessible content, edits, deletions, or uncertainty.
Do not silently replace an original capture with a cleaned or annotated version. If a package must be redacted or reformatted for sharing, preserve the original securely and explain what was changed in the derivative. For legal proceedings or formal investigations, follow the receiving organization’s evidence-handling requirements and consult counsel where needed.
Rank #4
Or skip the browser setup
If your task is to preserve a permitted public page as a screenshot or PDF rather than scrape structured social data, ScreenshotNeo is a screenshot API and MCP server, not a substitute for a platform API or authorization to collect. A single GET request can return a PNG, JPEG, WebP, or PDF. For example, using an authorized target URL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/public-page -o shot.webp
See the ScreenshotNeo documentation for request parameters. Cookie banners and consent prompts are accepted or removed before capture, and known newsletter popups and chat widgets can be removed; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Free includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. These capture features do not establish that collection is permitted: check the source’s rules and your lawful basis first. Sign up for 1,000 free screenshots a month, with no card.
Common failure modes and practical fixes
- Access is denied or a CAPTCHA appears: Stop automated access. Check whether the platform provides an authorized API, export, or approved account workflow. Do not attempt to bypass the control.
- The scraper returns incomplete or inconsistent fields: Dynamic rendering, changing page layouts, permissions, or API limitations may affect coverage. Compare the output with a permitted manual view, document gaps, and use an authorized structured source if available.
- Requests are throttled or fail repeatedly: Confirm the applicable limit, reduce request pace, use permitted batching, and implement backoff and error logging. Do not increase concurrency to evade a limit.
- A page changes or disappears: Preserve each authorized capture promptly with URL, time, identifier, artifact, and hash. State that later deletion or editing could not be independently reconstructed from the capture alone.
- Two analysts cannot reproduce the same result: Compare query, account context, time, tool version, locale or access conditions, and transformations. Keep the raw artifact and record these details so the difference can be explained.
- The collection contains more personal data than necessary: Stop expanding the dataset, narrow the query, restrict access, and apply the documented minimization, retention, and deletion rules. Reassess the lawful basis if the purpose or data scope has changed.
Cost, reliability, and reporting considerations
API access, exports, permitted crawling, and capture tools have different costs and operational limits; compare them against the investigation’s required coverage and retention needs rather than assuming that scraping is the cheapest route. Reliability is not just whether a script completes: consider missing pages, changed layouts, rate limiting, account state, stale or edited content, and whether a second reviewer can interpret the preserved record.
Best Value
For each material conclusion, separate the observation from the inference. State when and how the content was collected, what could not be accessed, whether the content later changed, and which independent sources support or contradict it. Avoid presenting a captured social post as proof of identity, authorship, or truth without corroboration appropriate to the claim.
Frequently Asked Questions
Does a screenshot prove who authored a social-media post?
No. A capture documents what appeared at a source page when it was recorded; authorship and authenticity require separate corroboration.
Can an OSINT team collect data from a logged-in account?
Only where the account’s access rights, platform rules, and applicable privacy requirements permit the collection. Login access does not by itself authorize automation or broader reuse.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




