Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape JavaScript-rendered content with PhantomJS, create a webpage object, load the URL with page.open, verify that the callback reports success, and read the rendered DOM inside page.evaluate. Return only strings, numbers, booleans, arrays, or plain objects from the page context. Because the load callback does not prove that an application’s later asynchronous requests have finished, extract only after the page reaches a site-specific ready state.

What PhantomJS actually does

PhantomJS is a scriptable headless WebKit browser. Unlike an HTTP client that receives only the initial HTML response, it executes the page’s scripts and exposes the resulting DOM to your scraper. That makes it possible to read headings, cards, prices, links, and other elements inserted after the first response.

This approach is primarily a maintenance technique for existing PhantomJS jobs. The project README says development is suspended, the GitHub repository was archived on May 30, 2023, and the project wiki describes the 2.x branch as deprecated and no longer maintained. See the archived repository and official wiki. Do not select PhantomJS as the default for a new scraper without accepting that maintenance risk.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimal working scraper

Install PhantomJS 2.1 (the project’s latest stable line), save this as scrape.js, and run phantomjs scrape.js. Replace the URL and selectors with those for the page you own or are permitted to access.

var webpage = require('webpage');
var page = webpage.create();
var url = 'https://example.com';

page.open(url, function (status) {
  if (status !== 'success') {
    console.log('Could not load page: ' + status);
    phantom.exit(1);
    return;
  }

  var result = page.evaluate(function () {
    var heading = document.querySelector('h1');
    var links = document.querySelectorAll('a');
    var hrefs = [];

    for (var i = 0; i < links.length; i += 1) {
      hrefs.push({
        text: (links[i].innerText || links[i].textContent || '').trim(),
        href: links[i].href
      });
    }

    return {
      title: document.title,
      heading: heading ? (heading.innerText || heading.textContent || '').trim() : '',
      links: hrefs
    };
  });

  console.log(JSON.stringify(result));
  phantom.exit();
});

page.open “opens the url and loads it to the page”; its callback receives the page status, normally success or fail, as documented in the page.open reference. A successful callback means the load event completed, not that every framework request or delayed component is ready.

Wait for the application, not an arbitrary delay

Single-page applications often render a shell first and populate it later. If you call page.evaluate immediately, you can receive an empty list even though a human would see data a moment later. Define readiness in terms of the page itself:

  • Choose a selector that exists only when the required data is present, such as a results container, a non-empty table body, or a “loaded” marker.
  • If the site exposes a stable state attribute or text marker, use that as the signal instead of guessing a number of milliseconds.
  • Extract only the fields needed for your record once that signal is present.
  • Keep the readiness rule specific to the target application. There is no universal fixed sleep that is reliable for every site.

The official API documentation establishes what page.open reports, but it does not define a universal application-ready condition. Build that condition from the target site’s DOM or application state, and treat a missing signal as a scrape failure rather than silently storing an empty record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use page.evaluate as a serialization boundary

The page.evaluate reference defines the method as evaluating a function in the context of the web page. The function can use document, CSS selectors, computed text, and other browser-side APIs. Only JSON-serializable arguments and return values cross back to the PhantomJS script.

Values that are safe to return

  • Strings, numbers, booleans, and null.
  • Arrays containing those values.
  • Plain objects whose properties contain serializable values.

Values that do not cross the boundary

  • DOM nodes such as an Element or NodeList.
  • Functions, closures, and objects containing functions.
  • Browser objects that cannot be represented as JSON.

Convert elements to the fields you need inside the evaluated function. In the example, each link becomes a plain object with text and href; returning the anchor elements themselves would not produce usable data.

Extraction patterns for rendered pages

One element with a fallback

var value = page.evaluate(function () {
  var node = document.querySelector('[data-price]');
  return node ? (node.textContent || '').trim() : null;
});

Returning null for a missing required element lets the outer script distinguish “not found” from an empty string.

Repeated cards or rows

var rows = page.evaluate(function () {
  var cards = document.querySelectorAll('.product-card');
  var output = [];

  for (var i = 0; i < cards.length; i += 1) {
    var card = cards[i];
    var name = card.querySelector('.name');
    var price = card.querySelector('.price');
    output.push({
      name: name ? (name.textContent || '').trim() : '',
      price: price ? (price.textContent || '').trim() : ''
    });
  }

  return output;
});

Attributes and links

Read attributes inside the page context, then return their string values. Use the browser-resolved href when you need an absolute link, or getAttribute('href') when the original relative value matters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep output small and explicit

Returning a compact record is easier to validate and serialize than returning a large section of markup. If you need raw HTML for a particular fragment, return element.outerHTML as a string and process it outside the page context.

Printing, logging, and process exit

Serialize the returned object with JSON.stringify in the PhantomJS context, as shown above. A console.log executed inside page.evaluate is page-context logging; it does not automatically appear in the PhantomJS process output. If you need browser-console messages, configure the page’s onConsoleMessage handler and forward them explicitly. Always call phantom.exit() on both success and failure paths so batch jobs do not remain running.

Diagnose the common failure modes

Symptom Likely cause Correction
status is fail The URL could not be loaded successfully. Log the status, stop the job, and verify the URL and the target’s availability before retrying.
The page loads but arrays are empty Extraction ran before asynchronous content was inserted, or the selector no longer matches. Inspect the rendered DOM, choose a selector that signals readiness, and update the selector to match the current markup.
Text is blank although an element exists The visible value is in a child node, attribute, or a different element than expected. Check textContent/innerText, inspect the relevant attribute, and return a normalized string.
Returned data cannot be serialized The evaluated function returned a DOM node, closure, function, or another non-JSON value. Map the value to primitives or plain arrays/objects before returning it.
Debug messages are missing console.log was called inside the page context. Return diagnostic data, or wire onConsoleMessage to the outer script.
Human-visible content never appears The page may require an interaction, authentication flow, or a challenge that this legacy browser cannot complete. Confirm that the required state is reachable in PhantomJS; otherwise migrate the job to a maintained browser automation stack or use a rendering service.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reliability and maintenance decisions

When keeping PhantomJS is defensible

  • You have a stable legacy script, fixed selectors, and a controlled target.
  • The deployment already packages PhantomJS 2.1 and you can monitor failures.
  • You can define a deterministic, page-specific readiness signal.

When migration is the safer choice

  • The target frequently changes its JavaScript, markup, or authentication flow.
  • You need ongoing security and browser compatibility updates.
  • Your scraper depends on modern browser behavior that the deprecated WebKit engine does not provide.
Decision factor Existing PhantomJS job New or expanding system
Maintenance status Suspended project; archived repository Prefer a maintained browser automation tool
Porting effort No immediate rewrite Plan to port selectors and extraction logic early
Readiness handling Must be implemented per site Choose tooling with reliable, documented state waits
Deployment Keep the existing PhantomJS runtime reproducible Evaluate the runtime, browser, and update process before production

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server when you need a rendered page image or PDF rather than a custom DOM data pipeline. One GET request returns PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports the result in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

For the complete parameter list, including full-page capture, lazy-image loading, CSS-selector element capture, dark mode, device presets, custom viewports, retina scale, PDF paper and page-range controls, custom CSS or JavaScript, clicks, hidden selectors, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification, see the ScreenshotNeo documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = require('fs');
fs.writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

The free plan includes 1,000 screenshots each month with no card required. Paid plans start at $5 for 3,000 screenshots, and every feature is available on every plan. Create a free ScreenshotNeo account to try the call.

Practical checklist

  1. Confirm that PhantomJS is appropriate for this legacy workload and record the runtime version.
  2. Open the URL and handle a fail status before attempting extraction.
  3. Identify a DOM selector or state marker that proves the required asynchronous content is ready.
  4. Use page.evaluate to convert only the needed fields into JSON-compatible values.
  5. Validate required fields, serialize the result, and exit with a non-zero status on failure.
  6. Monitor selector and readiness failures, because PhantomJS itself is no longer maintained.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.