Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A good scraper input schema is a clear contract between the person starting a run and the scraper that will execute it. It should ask for the information the scraper genuinely needs, give sensible defaults for routine behavior, and reject values that cannot work. In Apify, the schema can also drive input validation and the Actor’s human-facing form; other frameworks may use different syntax and may not generate a form at all.

The examples below use Apify’s Actor input-schema conventions where noted. Apify’s schema resembles JSON Schema but has extensions and differences, so validate it with Apify’s tooling rather than assuming a generic JSON Schema validator will accept it unchanged.

Start with the scraper’s contract, not a list of possible settings

Before writing fields, describe what a caller must tell the scraper for one useful run. Then separate that information from implementation details that the scraper can decide itself. A schema is not a dump of every internal option: each exposed field creates work for callers and becomes part of the interface that API clients, schedules, and other automation may depend on.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a typical site scraper, the caller-controlled inputs might be:

  • Where to start: one or more URLs, if there is no meaningful target that the scraper can select on its own.
  • How much to collect: an optional maximum page or item count, with a conservative default.
  • Site-specific controls: a search term, category, date range, or pagination choice only when the target workflow genuinely needs it.

These are design examples, not a universal field set. Apify’s crawler example, for instance, uses an array of start URLs and a page function as required inputs for that example. Your schema should reflect your own scraper’s execution model.

Keep caller choices separate from scraper mechanics

Expose a value when a caller can reasonably choose it and the choice changes the run in a useful way. Keep internal selectors, retry implementation, and parsing strategies out of the public input unless users truly need to control them. This keeps the form understandable and reduces the chance that a harmless refactor breaks existing callers.

Choose required fields, defaults, and prefills deliberately

Three settings that can look similar in a form have different meanings: requiredness, a default, and a prefill. Treating them as interchangeable is a common source of confusing APIs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Required means the run cannot proceed without it

Mark a field required when omission leaves the scraper with no reasonable way to do the job. A start URL is a good candidate if the scraper has no built-in target. Avoid making optional tuning controls required just to ensure they appear in a form.

Default means the scraper has a usable value when the caller omits one

Use a default for routine behavior the implementation can choose sensibly, such as a crawl limit. In Apify’s documented behavior, omitted fields receive their defaults whether a run is started through the API, CLI, scheduler, or UI. The caller can still provide a different value when appropriate.

Prefill is an example in the UI, not an API default

Apify’s prefill is intended to demonstrate a field value in the interface and make testing convenient. It is UI-only; it does not mean an API caller that omits the field has supplied that value. Use a prefill when an example is useful but there is no reasonable default for every run. If the scraper actually needs the value, use requiredness or a real default instead.

Define types and constraints that match real requirements

For each property, choose a type, a user-facing title, and a description that says what the caller should enter and what the scraper will do with it. Apify’s documented input types include string, array, object, boolean, and integer. Its field settings can also express defaults, prefills, examples, and validation messages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Constraints are useful when they enforce an actual requirement—not merely because the schema format offers them. Apify documents string patterns and length bounds, enumerations, array limits, and nested object schemas. For example, a maximum number of start URLs can protect a run from an accidental oversized input; an enumeration is appropriate when the scraper truly supports a closed set of modes.

Illustrative Apify schema

This abbreviated example shows the shape of an Actor input schema. The names and limits are illustrative; choose values that fit your scraper rather than copying them blindly.

{
  "schemaVersion": 1,
  "title": "Catalog scraper input",
  "type": "object",
  "properties": {
    "startUrls": {
      "title": "Start URLs",
      "type": "array",
      "description": "Catalog pages to crawl.",
      "editor": "requestListSources",
      "items": {
        "type": "object",
        "properties": {
          "url": { "type": "string", "format": "uri" }
        },
        "required": ["url"]
      }
    },
    "maxPages": {
      "title": "Maximum pages",
      "type": "integer",
      "description": "Stop after this many pages.",
      "default": 20,
      "minimum": 1,
      "maximum": 500
    },
    "sortOrder": {
      "title": "Sort order",
      "type": "string",
      "enum": ["newest", "price-low-to-high"],
      "default": "newest"
    }
  },
  "required": ["startUrls"]
}

The example illustrates intent, not a guarantee that every generic schema keyword or editor works identically across frameworks. Apify’s specification documents schema version 1 and a 500 kB maximum input-schema file size; those are Apify platform facts, not general limits for scraper schemas. Confirm keyword and editor support in the platform’s own validator and documentation before publishing a schema.

Use messages that help callers fix errors

A constraint should be paired with a useful explanation. “Maximum 500 pages” tells a caller what to change; a generic “invalid input” does not. Descriptions should clarify expected formats, whether multiple values are accepted, and how defaults affect behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make the generated form reflect the data

Where a framework generates a form from its schema, select an editor that fits the value. Apify documents UI choices such as a URL-list editor for start URLs, a select for a closed set of choices, and a code editor for code-valued fields. Use descriptions as help text, and group advanced controls into sections if the platform supports that. These are Apify interface capabilities, not portable field names that every scraper framework will recognize.

A form should make the normal path obvious. Put common inputs first, make optional tuning clearly optional, and do not use a select for a list that changes frequently or is effectively open-ended. Prefills can make a test easy, but label examples so users do not mistake them for universal production values.

Decide whether unknown input fields should be accepted

Strictness is an API design choice. Apify documents root and nested-object additionalProperties behavior as permissive by default. Set it to false when an unrecognized field should fail validation before execution, rather than being silently ignored or mistaken for a supported option.

Strict validation catches misspellings early, but tightening a published schema can break older clients that send extra fields. Before changing permissive behavior, consider API integrations, scheduled runs, and callers that may already depend on the interface. Apify states that input failing its validation is rejected before the Actor starts, which is useful precisely because the failure occurs before the scraper spends time doing work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find the real inputs for JavaScript-driven pages

A page that displays data after JavaScript runs does not automatically require a browser-rendering switch in every scraper schema. First identify how the page obtains that data. Inspect browser network activity and look for the request that returns the records or page content.

Prefer the data request when it is practical

Scrapy’s version 2.1.0 documentation describes reproducing the request that supplies the data. Depending on the site, that can mean matching the HTTP method and URL and also providing a request body, headers, or form parameters. If the response already contains structured data, retrieving it directly may avoid rendering an entire page.

Do not assume that copying the URL alone is sufficient. The request may depend on parameters, a body, or headers visible in the browser’s network panel. Add a schema field for a caller-controlled value only when that value needs to vary between runs; keep fixed request details inside the implementation where possible.

Use rendering when the task actually needs a browser

If reproducing the underlying request is impractical—or the required output is a browser-visible artifact such as a screenshot—JavaScript rendering or a headless browser may be appropriate. This is a different execution strategy, with added browser setup and resource use. Do not expose a generic “render JavaScript” toggle unless users have a meaningful reason to select between strategies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If the specific output you need is a clean screenshot rather than extracted records, ScreenshotNeo provides a screenshot API and MCP server. A single GET request can return an image or PDF; it does not replace a scraper that needs to parse and collect page data.

For example, save a screenshot of the target page as WebP with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for the request options. Cookie banners, popups, and chat widgets are removed before capture; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Test the contract before publishing it

Validation should be part of schema development, not an assumption made after a form looks right. Apify recommends using its platform validation because its Actor schema has extensions and differences from generic JSON Schema. Test representative inputs through the same start paths your users will use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Smallest valid run: provide only the required fields and verify that defaults produce the expected behavior.
  • Boundary cases: test minimum and maximum values, empty arrays if permitted, and values just outside declared constraints.
  • Wrong types: try a string where an integer is expected and confirm the error identifies the relevant field.
  • Unknown fields: verify whether extras are rejected or accepted according to the schema’s strictness.
  • Nested values: test missing required properties inside objects as well as malformed entries in arrays.
  • Real callers: check API, CLI, scheduled, and UI launches if those are supported routes for the Actor.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common schema and crawl failures

A caller omits a field and the run still fails

Check whether the field has a real default or only a UI prefill. In Apify, a prefill is an example shown in the interface; it does not supply a value to an API call. Make the field required, add a true default, or handle omission in the implementation.

A generic schema validator rejects an Apify schema

Apify’s Actor input schema resembles JSON Schema but includes platform-specific extensions and differences. Validate it with Apify’s tooling and verify the supported schema version and editor options instead of treating generic validator behavior as authoritative.

An unexpected field does not cause an error

Apify’s documented root and nested-object behavior is permissive by default. If unknown properties should fail fast, set additionalProperties to false at the relevant object level. Check compatibility before making this change to an interface callers already use.

The scraper returns no records from a page that looks populated

The data may arrive through a separate browser request rather than the initial page response. Inspect network activity, identify the response that contains the records, and reproduce its method, URL, and any required body, headers, or form parameters. If the task depends on what a browser renders, assess a rendering approach instead.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validation rejects a value users believe is valid

Compare the submitted value with the declared type, bounds, pattern, enumeration, and nested requirements. Then decide whether the constraint represents a true execution limit. If the requirement has changed, update both the constraint and the user-facing explanation, and validate the revised schema in the target platform.

Balance ease of use, validation, and compatibility

Schema design is a set of trade-offs rather than a contest to maximize the number of fields or constraints.

Design choice Benefit Cost or risk
Require only essential inputs Callers can start a useful run without learning every option. The implementation must choose sensible behavior for omitted optional values.
Set realistic defaults Routine runs need less configuration and API callers get predictable values. A default that does not fit a caller’s target can produce an unintended run unless it is clearly described.
Add bounds and enumerations Invalid values fail early and closed choices are easy to understand. Overly narrow constraints can reject legitimate use cases or require schema changes later.
Reject unknown properties Misspelled or unsupported options are less likely to pass unnoticed. Existing clients that send extra properties may break when strictness is introduced.
Use direct data requests When the response contains the needed data, a browser may not be necessary. The request can require details beyond its URL, such as a method, body, headers, or form parameters.
Render the page in a browser Useful when the browser-visible result itself is required or the data request is impractical. It adds a browser execution path and should not be exposed as a needless universal option.

Review the schema as a public interface whenever the scraper’s behavior changes. Adding an optional field is usually easier for callers to absorb than changing the meaning of an existing field or making previously accepted input invalid. Keep names stable, document behavior, and test the cases that callers can actually submit.

Frequently Asked Questions

Does every web scraper need a generated input form?

No. A schema can define and validate the input contract even when a framework does not generate a UI. Form editors and generated-interface capabilities depend on the framework.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should a scraper schema include authentication credentials?

Only if callers need to supply them for a supported use case. The cited schema guidance establishes field and validation design, not a credential-storage policy; follow your platform’s security guidance for handling secrets.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.