Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
GPT Image 2

Screenshot and Image Generation API Quick Start (OpenAI, with Webpage Capture)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the OpenAI Images API when the result you need is a generated or edited picture; use Responses or Chat Completions when the result is an explanation of an image or screenshot. If you need a pixel-accurate screenshot of a live webpage, use a browser automation tool or a screenshot service such as ScreenshotNeo instead. This guide shows the setup, runnable Python, JavaScript, and cURL examples, base64 decoding, editing, transparency, analysis, model migration, and failure recovery.

Choose the API surface first

OpenAI exposes three useful paths for this workflow. Selecting the path before writing code avoids trying to make an image-generation endpoint perform visual analysis, or trying to make a text endpoint return a file.

Goal Use What comes back
Create a new image or edit an existing one Images API (client.images.generate or client.images.edit) An image result containing base64 data that your program decodes and saves
Ask a model to inspect an image, including a screenshot Responses API with an image input A text response describing or reasoning about the image
Have image analysis produce a conversational text answer Chat Completions with an image input Text in the assistant message
Generate through an agent-style request Responses API with the image-generation tool An image_generation_call containing base64-encoded output

The official images and vision guide explains the analysis split, while the image-generation guide documents the generation tool and its controls.

Prerequisites and a safe first request

  1. Create an API key. Keep it on the server or in a protected shell environment, never in browser JavaScript or a mobile application. Export it as OPENAI_API_KEY; the official Developer Quickstart shows the supported SDK installation flow.
  2. Install an SDK. Python uses pip install openai; JavaScript or TypeScript uses npm install openai.
  3. Pick GPT Image 2 for new work. The current prompting reference documents GPT Image 2 generation and editing. GPT Image 1.5 is deprecated and scheduled to shut down on December 1, 2026; GPT Image 1 is deprecated and scheduled to shut down on October 23, 2026. Test your prompts and visual requirements on GPT Image 2 before either deadline.

A minimal environment check is:

export OPENAI_API_KEY="your_key_here"
python -c "from openai import OpenAI; print('SDK ready')"

Generate an image and save the base64 result

Python

This script requests a PNG with a transparent background. PNG or WebP is required for transparency; JPEG cannot carry a transparent background.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import base64
from openai import OpenAI

client = OpenAI()
result = client.images.generate(
    model="gpt-image-2",
    prompt=(
        "A clean, front-facing product illustration of a red mechanical keyboard, "
        "soft studio lighting, no brand marks, centered composition"
    ),
    size="1024x1024",
    quality="high",
    output_format="png",
    background="transparent",
)

image_bytes = base64.b64decode(result.data[0].b64_json)
with open("keyboard.png", "wb") as file:
    file.write(image_bytes)
print("Wrote keyboard.png")

The important detail is result.data[0].b64_json: it is base64 text, not a ready-to-display file. Decode it as bytes and write those bytes in binary mode.

JavaScript (Node.js)

import OpenAI from "openai";
import { writeFile } from "node:fs/promises";

const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const result = await client.images.generate({
  model: "gpt-image-2",
  prompt: "A clean, front-facing product illustration of a red mechanical keyboard, soft studio lighting, no brand marks, centered composition",
  size: "1024x1024",
  quality: "high",
  output_format: "png",
  background: "transparent"
});

await writeFile("keyboard.png", Buffer.from(result.data[0].b64_json, "base64"));
console.log("Wrote keyboard.png");

cURL

curl https://api.openai.com/v1/images/generations 
  -H "Authorization: Bearer $OPENAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "gpt-image-2",
    "prompt": "A clean, front-facing product illustration of a red mechanical keyboard, soft studio lighting, no brand marks, centered composition",
    "size": "1024x1024",
    "quality": "high",
    "output_format": "png",
    "background": "transparent",
    "response_format": "b64_json"
  }' > response.json
jq -r '.data[0].b64_json' response.json | base64 --decode > keyboard.png

The CLI also returns JSON rather than writing an image automatically. Extract data.0.b64_json and pipe it through a base64 decoder; the current OpenAI CLI guide notes that image commands do not yet provide native --output support.

Edit an existing image

Pass an input image to client.images.edit and describe both the requested change and what must remain untouched. GPT Image 2 processes image inputs at high fidelity; the prompting reference says to omit input_fidelity.

import base64
from openai import OpenAI

client = OpenAI()
with open("original.png", "rb") as source:
    result = client.images.edit(
        model="gpt-image-2",
        image=source,
        prompt=(
            "Replace only the background with a pale blue gradient. "
            "Keep the object, its proportions, labels, and colors unchanged."
        ),
        output_format="png",
    )

with open("edited.png", "wb") as output:
    output.write(base64.b64decode(result.data[0].b64_json))

For an API request outside the SDK, send the image as multipart form data to the image-edit endpoint and decode the returned b64_json value the same way. Keep input files small enough for your application’s upload limits and validate the returned bytes before publishing them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control size, quality, format, and action

  • size: choose the output dimensions supported by the model; flexible sizes are documented for GPT Image 2.
  • quality: select the quality level appropriate to draft previews or final assets.
  • output_format: choose PNG, JPEG, or WebP. Use PNG or WebP when an alpha channel is required.
  • background: request transparent for cut-outs and compositing, or a solid/opaque background when transparency is unnecessary.
  • action: with the image-generation tool, auto lets the model choose between generation and editing; generate and edit force the operation.

Do not assume a file extension proves transparency. Open the output with an image library and verify that it has an actual alpha channel before using it in a compositor or as a product asset.

Analyze a screenshot or image

Generation and analysis are different jobs. For analysis, send an image input to Responses and ask for the text you need. Keep the model name in an environment variable so you can select a vision-capable model available to your project.

import os
from openai import OpenAI

client = OpenAI()
response = client.responses.create(
    model=os.environ["OPENAI_VISION_MODEL"],
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": "List every visible form label and identify any alignment or contrast problem."},
            {"type": "input_image", "image_url": "https://example.com/page.png"}
        ]
    }]
)
print(response.output_text)

The same image can be supplied as base64 data when it is not publicly reachable. Chat Completions is a suitable alternative when you specifically want the result as a conversational assistant message. For agent workflows, the Responses image-generation tool returns an image_generation_call; extract its base64 result and decode it exactly as in the generation examples.

Write prompts that are testable

  • Subject: name the object or scene and its important attributes.
  • Composition: state viewpoint, placement, aspect, background, and lighting.
  • Style: identify the visual treatment without relying on an ambiguous single adjective.
  • Constraints: specify what must not change, especially during edits.
  • Text: provide exact copy, then inspect the rendered letters yourself.

OpenAI recommends checking that required text is accurate and legible, identities and labels remain intact, edits changed only the requested area, and transparent output contains a real alpha channel. Treat those checks as acceptance tests, not optional visual polish.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Capture a live webpage screenshot

An OpenAI image request creates or analyzes raster content; it does not replace a browser renderer for a live URL. A do-it-yourself route is to launch a headless browser, wait for the page, and save the rendered viewport or full page. The exact wait condition should match your site: static pages can wait for load, while client-rendered pages should wait for a selector that proves the content is present.

npm install playwright
npx playwright install chromium
// capture.mjs
import { chromium } from "playwright";

const browser = await chromium.launch();
const page = await browser.newPage({ viewport: { width: 1440, height: 900 }, deviceScaleFactor: 1 });
await page.goto(process.argv[2], { waitUntil: "networkidle" });
await page.screenshot({ path: "page.png", fullPage: true });
await browser.close();
node capture.mjs https://example.com

This approach gives you control, but you must maintain the browser binary, consent dialogs, popups, authentication, lazy-loaded images, bot checks, timeouts, and retries yourself.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

Use the ScreenshotNeo API documentation for all options. A minimal call is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Its 63 options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus arbitrary viewports, retina scale, PDF paper size/margins/landscape/page ranges, HTML or CSS to image, custom JavaScript, click-before-capture, hidden selectors, waits for a selector/delay/network idle, blocking ads/trackers/requests/resource types, custom headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, user-selected cache TTLs, signed links for public <img> tags, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, an OpenAPI specification, and compatibility with parameter names used by other screenshot APIs.

Every feature is included on every plan. The Free plan includes 1,000 shots per month without a card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients, so an AI agent can capture pages without custom browser orchestration. Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting and reliability checks

“The output is not an image file”

You saved the JSON response or base64 text directly. Read data[0].b64_json, decode it, and write binary bytes.

Transparent pixels appear solid

Request background: "transparent", select PNG or WebP rather than JPEG, and verify the decoded file’s alpha channel.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

An edit changes more than requested

State the unchanged elements explicitly, describe the exact region to modify, and inspect labels, identities, and proportions after every edit.

Text in the image is wrong

Use exact quoted copy, keep the composition simple, and perform a legibility check before distribution; do not treat generated lettering as automatically proofread.

Old model calls fail during migration

Move integrations to GPT Image 2 ahead of the GPT Image 1 and 1.5 shutdown dates, then compare representative prompts and edits for visual differences.

A webpage capture is blank or incomplete

Wait for a meaningful selector rather than an arbitrary delay, account for lazy loading, and record whether a consent dialog, bot check, timeout, or network failure prevented a clean render. A screenshot service can report these outcomes separately from billable captures.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Privacy and production handling

Keep API keys server-side, redact sensitive screenshots before sending them for analysis, and set explicit retention and access rules for generated files. OpenAI states, “By default, we never train on customer API data,” while also noting that image inputs and outputs remain subject to API usage policies; read the April 23, 2025 announcement and current policy terms before processing regulated content.

For dependable pipelines, log the model, prompt, size, quality, output format, request identifier, decode result, and post-generation validation outcome. Cache only when the same source and parameters are intentionally reusable, and retry transient network failures with bounded backoff rather than duplicating every request blindly.

Quick decision checklist

  • Need a new illustration or an edit? Use Images API and decode b64_json.
  • Need an explanation of a screenshot? Use Responses or Chat Completions with an image input.
  • Need a browser-rendered URL? Use Playwright or ScreenshotNeo.
  • Need transparency? Request a transparent background and PNG/WebP, then verify alpha.
  • Need a current model? Start with GPT Image 2 and plan migration away from deprecated GPT Image 1 and 1.5.

Frequently Asked Questions

Can I return a URL instead of base64 image data?

The documented quick-start flow returns base64 image data; your application can store the decoded bytes and publish its own URL or object-storage reference.

Should a screenshot be sent as a public URL?

A public URL is convenient for a Responses image input, but private applications can send base64 image data instead; protect any URL that exposes sensitive content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I record to reproduce an image request?

Record the model, complete prompt, input image version, size, quality, background, output format, and the validation result.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.