DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
Codex

Generating Video, PDFs, and Images with the Codex MCP Server

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use codex mcp-server when you want an MCP-compatible client to call Codex over a local stdio connection. The server provides the agent connection; the actual asset workflow depends on the capability you invoke: Codex document skills create or revise formatted PDFs, image-generation tools create or edit images, and the Videos API creates an asynchronous video job that you later download. Choose the Codex App Server instead when your client needs richer thread state, streaming progress, or diff updates.

This guide shows how to choose the transport, connect a client, design prompts and references for each asset type, handle the asynchronous video lifecycle, and review generated files safely.

What the Codex MCP server actually provides

Running codex mcp-server exposes Codex as a tool server for any MCP client that supports stdio servers. OpenAI describes the integration as: “Run codex mcp-server and connect from any MCP client that supports stdio servers.” The MCP process is therefore the bridge between your client and Codex; it is not a separate PDF renderer, image model, or video encoder.

Asset creation happens in the capability you call:

  • PDF: ask Codex to create or revise a formatted document with its document skills, provide the source content and layout constraints, then inspect the generated file.
  • Image: call an image-generation tool to create new pixels or edit an existing image. The Responses API image tool supports PNG, WebP, and JPEG output, quality values of low, medium, high, or auto, sizes of 1024×1024, 1024×1536, and 1536×1024, and optional partial-image streaming.
  • Video: submit a prompt and optional reference image to the Videos API, monitor the asynchronous job, and download the completed video, normally as MP4.

Read the transport overview in OpenAI’s Codex harness article before wiring a production client.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose MCP or the Codex App Server

Both approaches can put Codex inside a larger tool workflow, but they expose different control surfaces.

Decision axis codex mcp-server Codex App Server
Connection Local stdio MCP server Application-oriented integration
Session semantics Callable tools and the MCP protocol Richer thread lifecycle, streaming progress, and diff updates
Best fit An existing MCP client such as an agent host that can launch stdio servers A client that needs first-class conversation state and live updates
Scope Only the operations exposed through MCP endpoints Broader application session behavior

Use MCP when portability and a narrow callable interface matter more than advanced session state. Use the App Server when users need progress events, durable threads, or structured diffs while a long task runs.

Start the MCP connection

  1. Install and authenticate Codex in the environment where the MCP client will launch it. Keep credentials out of prompts, source control, and generated assets.
  2. Run the server command:
    codex mcp-server
  3. Register the command in your MCP client. Most clients accept a server name, executable, and argument list. A representative configuration is:
{
  "mcpServers": {
    "codex": {
      "command": "codex",
      "args": ["mcp-server"]
    }
  }
}

The exact configuration key varies by client, but the executable and argument must resolve to the command above. Start with a harmless request such as “List the files in the working directory and explain what you found.” Confirm that the client can receive a response before asking it to create an asset.

Generate a formatted PDF

For PDFs, treat Codex as a document agent rather than promising a particular PDF library. Give it the source material, page and typography constraints, and an explicit inspection step.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A reliable PDF prompt

Create a formatted PDF briefing from report.md.
Requirements:
- A4 portrait pages with 20 mm margins
- Title page, table of contents, numbered headings, and page numbers
- Keep tables together where possible; avoid widows and orphans
- Use accessible heading order and descriptive link text
- Export the final PDF to ./out/briefing.pdf
After creating it, inspect the rendered pages for clipping, missing fonts,
widow/orphan problems, and broken links. Report any corrections made.
  1. Provide the source content and any logos or reference files the document skill may use.
  2. State the paper size, orientation, margins, hierarchy, table behavior, and accessibility requirements explicitly.
  3. Ask Codex to export to a known path and then inspect the generated file, not just the source markup.
  4. Open the PDF in a viewer and check representative pages at 100% zoom. Verify page breaks, font substitution, image resolution, links, and text extraction before publishing.

The Codex app’s document skills are intended for reading, creating, and editing PDF files with professional formatting and layouts; the MCP server only supplies the agent connection. See OpenAI’s Codex app announcement for that workflow.

Generate or edit an image

Image generation is a pixel-producing operation. You can ask an MCP-connected agent to call an image capability, or call the Responses API image-generation tool directly from your application. GPT Image 1 and GPT Image 1 mini are listed as image-generation models.

Prompt structure

  • Subject and action: describe what must be visible, not only a theme.
  • Composition: specify framing, camera angle, focal point, background, and negative space for text.
  • Style and constraints: give medium, palette, lighting, aspect ratio, and content exclusions.
  • References and edits: identify the input image, what must remain unchanged, and exactly what should change.
  • Output: select PNG, WebP, or JPEG; choose a supported size and quality level.

Direct Responses API example with cURL

curl https://api.openai.com/v1/responses 
  -H "Authorization: Bearer $OPENAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "gpt-image-1",
    "input": "Create a clean 16:9 hero illustration of a developer reviewing an MCP workflow, with empty space on the left for a headline.",
    "tools": [{
      "type": "image_generation",
      "size": "1536x1024",
      "quality": "high",
      "output_format": "png"
    }]
  }'

The image tool can stream partial previews, which is useful when an agent needs to show progress or refine the composition over multiple turns. For edits, attach the source image through the client’s supported input mechanism and state which pixels must be preserved. Check the current Responses API reference for the request shape supported by your account and SDK version.

Python example

import os
import requests

payload = {
    "model": "gpt-image-1",
    "input": "Create a square product illustration of a secure MCP tool chain, flat vector style, blue and charcoal palette.",
    "tools": [{
        "type": "image_generation",
        "size": "1024x1024",
        "quality": "medium",
        "output_format": "webp"
    }]
}
response = requests.post(
    "https://api.openai.com/v1/responses",
    headers={
        "Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}",
        "Content-Type": "application/json",
    },
    json=payload,
    timeout=180,
)
response.raise_for_status()
print(response.json())

Save the returned image data according to the response object your SDK exposes. Do not assume that a text response contains the binary asset itself; inspect the output items and persist the image with the format you requested.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generate a video as an asynchronous job

The Videos API is not a synchronous “wait for a file” call. Submit a job, retain its identifier, poll or retrieve status, and download the completed content. The documented models are sora-2 and sora-2-pro. Clip lengths are 4, 8, or 12 seconds. Supported sizes are 720×1280, 1280×720, 1024×1792, and 1792×1024.

Submit a job

curl -X POST https://api.openai.com/v1/videos 
  -H "Authorization: Bearer $OPENAI_API_KEY" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "sora-2",
    "prompt": "A slow camera move across a developer dashboard showing an MCP workflow, neutral lighting, no readable brand names.",
    "seconds": "8",
    "size": "1280x720"
  }'

Store the returned job ID. If you use an input-reference image, send it with the request format accepted by your API client and describe how the motion should preserve the reference composition.

Poll and download in Python

import os
import time
import requests

headers = {"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"}
create = requests.post(
    "https://api.openai.com/v1/videos",
    headers={**headers, "Content-Type": "application/json"},
    json={
        "model": "sora-2",
        "prompt": "A 4-second macro shot of colored paper cards assembling into a workflow diagram.",
        "seconds": "4",
        "size": "1024x1024",
    },
    timeout=60,
)
create.raise_for_status()
job = create.json()
job_id = job["id"]

while True:
    status = requests.get(
        f"https://api.openai.com/v1/videos/{job_id}",
        headers=headers,
        timeout=60,
    )
    status.raise_for_status()
    data = status.json()
    state = data.get("status")
    if state in {"completed", "failed", "canceled"}:
        if state != "completed":
            raise RuntimeError(f"Video job ended with status: {state}")
        break
    time.sleep(5)

content = requests.get(
    f"https://api.openai.com/v1/videos/{job_id}/content",
    headers=headers,
    timeout=180,
)
content.raise_for_status()
with open("output.mp4", "wb") as file:
    file.write(content.content)

Node.js submission and retrieval

const key = process.env.OPENAI_API_KEY;
const headers = {
  Authorization: `Bearer ${key}`,
  'Content-Type': 'application/json'
};

const create = await fetch('https://api.openai.com/v1/videos', {
  method: 'POST',
  headers,
  body: JSON.stringify({
    model: 'sora-2-pro',
    prompt: 'A wide establishing shot of a calm studio where an agent assembles a PDF, image, and video pipeline.',
    seconds: '12',
    size: '1792x1024'
  })
});
if (!create.ok) throw new Error(await create.text());
const job = await create.json();

let status;
do {
  await new Promise(resolve => setTimeout(resolve, 5000));
  const response = await fetch(`https://api.openai.com/v1/videos/${job.id}`, {
    headers: { Authorization: `Bearer ${key}` }
  });
  if (!response.ok) throw new Error(await response.text());
  status = await response.json();
} while (!['completed', 'failed', 'canceled'].includes(status.status));

if (status.status !== 'completed') throw new Error(`Video ended: ${status.status}`);
const file = await fetch(`https://api.openai.com/v1/videos/${job.id}/content`, {
  headers: { Authorization: `Bearer ${key}` }
});
if (!file.ok) throw new Error(await file.text());
const fs = await import('node:fs/promises');
await fs.writeFile('output.mp4', Buffer.from(await file.arrayBuffer()));

Use the Videos API reference for the current status fields and upload options. In production, persist job IDs, retry transient status requests with backoff, and make downloads idempotent so a worker restart does not create duplicate jobs.

Expose a focused asset workflow through your own MCP server

If you are building a team tool, expose narrow tools such as create_briefing_pdf, generate_image, and start_video_job instead of one unrestricted “run anything” function. OpenAI’s plugin guidance recommends tools that complete recognizable user goals. The official TypeScript SDK is @modelcontextprotocol/sdk; the Python SDK package is mcp. Streamable HTTP is available when the server must be reached over a network rather than local stdio. See the MCP server building guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Each tool should validate inputs, limit file paths, cap prompt and reference sizes, return a job ID for video work, and report structured errors. Keep secrets in the server environment, not in tool arguments. For documentation lookups, OpenAI provides a read-only MCP server at https://developers.openai.com/mcp; the Codex CLI can add it with:

codex mcp add openaiDeveloperDocs --url https://developers.openai.com/mcp

Design handoff and collaborative review

Generated images often need a design review before they become product assets. OpenAI describes the Figma MCP Server as connecting Codex directly to Figma and tools such as Figma Make and FigJam. A practical flow is to generate a draft image, place it in a reviewable design file, record the prompt and model settings, and ask a designer to approve or request an edit. Keep the original prompt, references, and exported binary together so revisions remain traceable.

Safety, permissions, and release checks

Codex runs sandboxed by default, with approval modes and network controls. OpenAI says OpenTelemetry logs can include prompts, tool approval decisions, tool execution results, MCP server usage, and network allow-or-deny decisions. Treat those logs as potentially sensitive.

  • Least privilege: grant only the directories, network destinations, and tools required for the job.
  • Approval gates: require confirmation before publishing, sending files externally, or executing custom scripts.
  • Input hygiene: remove secrets and personal data from prompts and reference images; verify licenses for supplied material.
  • Output review: inspect PDFs visually, check image dimensions and transparency, and verify video duration, resolution, audio, and frame-level artifacts.
  • Provenance: retain prompt, model, tool settings, source references, timestamps, and reviewer identity with every released asset.
  • Network policy: allowlist remote MCP endpoints and API hosts; deny arbitrary outbound requests from document or media tools.

Read OpenAI’s safety guidance for Codex before granting a server access to production repositories or publishing systems.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance and cost controls

  • Choose the smallest image size and quality that satisfies the use case, then upscale or regenerate only when review requires it.
  • Use short video clips for storyboards and reserve 12-second or larger-resolution jobs for approved shots.
  • For asynchronous video, queue jobs, persist IDs, and poll with exponential backoff instead of keeping an interactive request open.
  • Cache deterministic PDF inputs and image prompts where your policy permits; store hashes so retries do not create duplicate work.
  • Separate preview and publication pipelines. Preview assets can use lower quality; publication assets should pass the full review checklist.

Troubleshooting

The client cannot start Codex

Check that codex is on the MCP client’s PATH, that the client launches the command in the intended working directory, and that the argument is exactly mcp-server. Run the command manually and inspect stderr before changing the client configuration.

The MCP connection starts but tools time out

Reduce the first request to a file listing or short explanation. Long media jobs should return a video job ID rather than blocking the MCP call. For networked custom servers, verify streamable-HTTP reachability and firewall rules.

The PDF looks correct in source but wrong when opened

Inspect the exported PDF itself. Missing fonts, clipped tables, image downsampling, and incorrect page breaks are export problems, not prompt problems alone. Add explicit font, margin, and page-break constraints and regenerate.

The image output has the wrong format or dimensions

Set the supported size, quality, and output format explicitly, then verify the returned asset metadata. Do not infer dimensions from the prompt text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The video never reaches completion

Persist the job ID and query its status directly. Check for a failed or canceled terminal state, validate the model, seconds, size, prompt, and reference image, and retry only after recording the error. Avoid submitting a second job while the first is still active.

Or skip the browser setup

If your workflow only needs a clean screenshot of a generated document, image, or web page, ScreenshotNeo is the first option to try: it removes consent banners, popups, and chat widgets before capture, bills only clean shots, and has a $5 paid plan for 3,000 shots.

One GET request returns PNG, JPEG, WebP, or PDF. The API also supports full-page captures, CSS-selector elements, device presets, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, signed links, asynchronous jobs, webhooks, bulk capture, and an MCP server for AI agents.

See the ScreenshotNeo API documentation for all options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can the Codex MCP server itself render a video file?

No. MCP supplies the callable connection. A video-capable tool or the Videos API creates the media, and your workflow downloads and stores the result.

Should a video request stay open until rendering finishes?

No. Treat video creation as a job lifecycle: submit, persist the identifier, poll or retrieve status, then download the completed content.

When is a custom networked MCP server preferable to local stdio?

Use streamable HTTP when multiple clients or hosts must reach one controlled service. Use local stdio when the client and Codex run together and you want the smallest deployment surface.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should be retained for an audit of generated assets?

Keep the prompt, model and tool settings, source references, output file, timestamps, and approval decision together.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.