Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web scraping can turn information on public webpages—such as prices, product availability, reviews, hiring notices or store locations—into observations an investor can analyze. It is a collection method, not proof that data is accurate, representative, legally usable or predictive. Use it to test a defined investment hypothesis, check the results against independent evidence, and review access, privacy and professional-conduct obligations before collecting or relying on the data.

What alternative data and web scraping mean in investment research

Alternative data is information used in investment analysis that sits outside traditional company filings, audited financial statements and standard market-data feeds. CFA Institute groups examples into individual, business and physical-world data. Web scraping is one way—not the only way—to collect some information published on websites. It does not turn every webpage into a reliable investment dataset.

Category Examples What scraping might capture
Individual data Social media, blogs, product reviews, web-search trends and cellphone-location data Publicly visible review counts, ratings or discussion themes. These are not necessarily representative of all customers or users.
Business data Card transactions, store visits and bills of lading Some related public information, such as store locations or shipping notices. Transaction and visit datasets are often collected or supplied through other methods; do not assume scraping can reproduce them.
Physical-world data Satellite observations of agriculture, rig activity, traffic, shipping and mining Webpages that publish observations or summaries. The underlying signal may come from imagery or sensors rather than webpage content.

The FCA has described regulators using automated collection, including web scraping of publicly available websites, for market monitoring and risk analysis. That illustrates a possible use; it is not blanket permission for every investor to collect every site or field.

What a scraped signal can—and cannot—tell you

A webpage can expose changes before they appear in conventional reports: a product going out of stock, a price changing, reviews accumulating, a new location appearing, hiring activity shifting, or shipping-related information being published. These observations may help formulate or challenge a thesis. They should be treated as complementary inputs and compared with company disclosures, filings and market data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
The Psychology of Money: Timeless lessons on wealth, greed, and happiness
  • Ideal for Gifting
  • Ideal for a bookworm
  • Compact for travelling

CFA Institute’s 2024 guidance describes unstructured information as accounting for up to 90% of data. This is a broad characterization of unstructured information, not a claim that 90% of scraped pages contain useful signals, or that any particular source predicts returns.

Before collection, state what decision the data could change. For example, a researcher might ask whether listed prices for a defined set of products changed during a particular period, and whether those changes corroborate a thesis about pricing pressure. That is more testable than collecting a large volume of pages and searching afterward for a pattern. A visible price also does not establish completed sales, unit volumes, margins or the experience of customers who cannot be observed on that page.

Plan a defensible collection before scraping

  1. Write the hypothesis and decision scope. Specify the securities or companies in scope, the signal, the observation period, the investment horizon and what result would count against the hypothesis. Decide whether the intended use is exploratory research, a published recommendation or an input to trading; the governance burden can differ.
  2. Assess the source. Record who owns or publishes the page, how the information is produced, how often it changes, how far back it is available and what population it covers. Ask who is missing: a retailer’s website, for example, does not necessarily represent all sellers or customers.
  3. Review access conditions. Read the site’s terms of service and robots.txt, and use an official API where one is available. Consider whether the proposed collection method is permitted in the relevant circumstances. Limit request rate and frequency to avoid undue server load. A page being publicly viewable does not settle these questions.
  4. Minimize sensitive information. Exclude personally identifiable or otherwise sensitive information unless the organization has established a lawful basis and an appropriate governance process. Review the collection itself, not only what will eventually be stored or analyzed.
  5. Design reproducible records. Preserve timestamps, source URLs, raw captures or hashes, parser versions, transformations and exceptions. Separate the observed page content from the analyst’s interpretation. This makes it possible to investigate whether a later result came from the source, the parser or a change in methodology.
  6. Set collection limits and failure rules. Define the pages and fields in scope, the collection cadence, timeouts and what to do when a page is unavailable or changes layout. Do not silently treat a blocked request, missing page or parsing failure as a real zero or a negative signal.

A minimal Python example for an authorized public page

The following example makes one request and records a page title, visible text and collection timestamp. It is a starting point for documenting observations, not a ready-made investment signal or a way to evade access controls. Before running it, review the specific site’s terms and robots.txt, confirm the collection is appropriate for your use, and substitute a page you are authorized to access. It does not log in, solve CAPTCHAs, evade blocks or crawl a site.

Install the dependencies with python -m pip install requests beautifulsoup4, then save this as capture_page.py:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import sys
from datetime import datetime, timezone
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup

if len(sys.argv) != 2:
    raise SystemExit("Usage: python capture_page.py https://example.com/permitted-page")

url = sys.argv[1]
parsed = urlparse(url)
if parsed.scheme != "https" or not parsed.netloc:
    raise SystemExit("Provide a valid HTTPS URL")

response = requests.get(
    url,
    headers={"User-Agent": "InvestmentResearchCollector/1.0 (contact: [email protected])"},
    timeout=(5, 20),
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
for element in soup(["script", "style", "noscript"]):
    element.decompose()

record = {
    "collected_at_utc": datetime.now(timezone.utc).isoformat(),
    "requested_url": url,
    "final_url": response.url,
    "http_status": response.status_code,
    "page_title": soup.title.get_text(" ", strip=True) if soup.title else None,
    "visible_text": soup.get_text(" ", strip=True),
}
print(json.dumps(record, ensure_ascii=False, indent=2))

Run it as python capture_page.py https://example.com/permitted-page, replacing the example with the permitted target. The output is one observation. For repeated collection, add an explicit schedule, rate limits, retention rules, failure logging and versioned parsing; do not turn this one-request example into an unbounded crawler. A page’s visible text may differ from content rendered by JavaScript, and one successful response does not establish coverage or representativeness.

Or skip the browser setup

For screenshot evidence rather than structured text extraction, ScreenshotNeo is a website screenshot API and MCP server. A screenshot can help preserve how a page appeared at capture time, but it does not validate the page’s claims or replace a structured data pipeline. The API call below returns a screenshot for an allowed target; see the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners as a visitor and removes 60+ known consent platforms, newsletter popups and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and responses identify page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.

Validate the observations before using them

Validation asks whether the collected data measures the thing the thesis says it measures, whether it is stable enough to compare over time, and whether any apparent relationship survives a fair test. Separate data-quality checks from investment-performance testing: a clean parser is not proof of a useful signal, and a promising historical association is not proof of a live edge.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Compare independent evidence. Check a sample of observations against another source and, where relevant, company disclosures or filings. Investigate discrepancies rather than selecting the source that best supports the thesis.
  • Measure missingness. Track unavailable pages, absent fields, blocked requests and collection gaps. Missingness may be systematic—for example, some pages or products may be more likely to fail—and can distort apparent trends.
  • Look for revisions and layout drift. Pages can be edited after collection, and site redesigns can break parsers or change how values are presented. Preserve raw captures or hashes and parser versions so that a change in output can be distinguished from a change in the underlying phenomenon.
  • Check selection and survivorship bias. A chosen set of pages may overrepresent popular products, surviving companies or accessible locations. Define the universe before interpreting a trend and document additions, removals and exclusions.
  • Test for collection artifacts. Bot checks, partial page loads, consent overlays or other unexpected content can look like valid observations. Flag anomalies and exceptions; do not treat them as genuine business changes.
  • Control leakage in historical tests. Use only information that would have been available at the time of each simulated decision. Respect publication and collection timestamps, account for revisions, and keep the test separate from choices made after seeing outcomes.
  • Paper-test before live use. A historical or paper-trading test can reveal operational problems and help assess a hypothesis. Do not claim performance unless a test was actually run, and do not assume a test result guarantees future results.

Legal, privacy and professional-conduct considerations

There is no general legal conclusion that follows merely from a page being public or from using an automated request. The answer depends on the site, the access method, the data fields, authentication, the purpose and the countries involved. Terms of service, robots.txt, privacy obligations, intellectual-property issues and other applicable rules may all matter. For a real deployment, get focused legal and compliance review based on those facts; this article is not jurisdiction-specific legal advice.

Best Value

For investment professionals, data provenance and analytical judgment also matter. CFA Standard V(A) calls for diligence, a reasonable and adequate basis, and reasonable inquiries into sources and data accuracy; its stated principle is that members and candidates must use reasonable care and judgment to achieve and maintain independence and objectivity in their professional activities. Standard I(B) requires independence and objectivity. Standard I(C) prohibits misrepresentation, including unattributed quotations, copied research, unsourced charts and reused algorithms.

For firms producing or distributing investment research, FCA COBS 12.2 addresses conflicts of interest, information barriers, disclosures for non-independent research and restrictions on trading ahead of unpublished research. The precise obligations depend on jurisdiction, authorization, audience and distribution model. Firms should assess their own circumstances rather than treating a general summary as a compliance determination.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Compare sources and vendors on the whole research cost

Do not choose a dataset solely because it is large, fast or easy to acquire. Compare candidates against the investment question and operational constraints. A source with frequent updates but little history may be unsuitable for a long historical test; a broad source with unclear provenance may be difficult to defend or reproduce.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Criterion Questions to ask
Thesis relevance Does the field measure the proposed mechanism, or only correlate loosely with it?
Coverage and representativeness Which companies, products, geographies or users are included and excluded?
Latency and historical depth When does an observation become available, how often does it update, and how far back does comparable history extend?
Provenance and reproducibility Can the source, collection method, transformations and exceptions be documented and independently retraced?
Revision behavior Can observations change after publication, and can prior values be recovered?
Privacy and terms risk Are the fields and collection method appropriate under the site’s conditions and the applicable governance process?
API quality and rate limits Is an official API available, and are its limits and failure behavior compatible with the research schedule?
Total cost and reliability Include acquisition, storage, engineering, monitoring, legal review and the work needed to handle outages or schema changes.
Market access and crowding How many other participants can access similar data and apply similar methods? Shared approaches can increase correlation and herding risk, a concern discussed by the IMF in relation to common alternative-data and AI approaches.

Common failure modes and practical fixes

  • The output suddenly becomes empty. Check the HTTP status, final URL, page response and parser version. The site may have changed its layout, returned an access challenge or served a page whose content is rendered differently. Record the exception; do not convert it to a zero-valued signal.
  • Values jump after a parser update. Compare raw captures and parser versions, then verify a sample manually. A changed extraction rule can create an artificial trend even when the underlying page did not change.
  • Repeated requests are blocked or slow. Reduce frequency and request volume, review access conditions and look for an official API. Do not attempt to evade bot checks or CAPTCHAs.
  • Historical results look unusually strong. Recheck timestamps, revisions, selection rules, survivorship bias and information leakage. Make sure each observation was available at the simulated decision time and that the hypothesis was not chosen after inspecting the result.
  • Different sources disagree. Investigate differences in definitions, coverage, update cadence and revisions. Preserve the disagreement and state which source supports which conclusion rather than silently combining incompatible measures.
  • A screenshot is mistaken for structured data. A visual capture preserves appearance; it does not itself provide a validated table of prices, counts or product states. Use a permitted structured source and a documented extraction and validation process when analysis depends on fields.

Document the result in an investment note

For each signal used, record the source and access method, collection dates, covered universe, fields, transformation and parser version. State relevant gaps, revisions, selection choices, validation checks and unresolved uncertainty. Explain how much the signal influenced the conclusion and how it compares with primary disclosures and other evidence. If research is distributed, document attribution and applicable conflict controls, and do not present a scraped observation as an independently verified fact when it is not.

The defensible outcome may be that a source is unsuitable: permission is unclear, coverage is biased, history is inadequate, or the signal does not add reliable information. Recording that decision is part of sound research, not a failed project.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.