October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AI agents

How to Build an AI Agent from Scratch: A Bounded, Testable Python Example

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a useful first agent by combining four parts: a model that reasons, instructions that define its job, one narrowly scoped tool, and an application loop with explicit stop conditions. This guide builds a website-checking agent in Python, shows the direct API loop, compares SDK and managed-runtime choices, and adds the validation, permissions, testing and failure handling needed before giving an agent real access.

What you are building

The example agent accepts a URL, uses one capture_page tool to open it in a sandboxed browser, and returns a concise report. The model may request the tool, inspect its result, or answer that it cannot complete the task. Your application—not the prompt—owns tool execution, limits the number of turns, validates arguments and decides when the run ends.

This is the practical definition of an agent: an LLM augmented with tools, and optionally retrieval or memory, operating inside a controlled cycle. A plain model call can generate text; an agent adds the ability to choose an approved action, receive its result and continue until a defined exit condition. OpenAI describes the model, tools and instructions as the core components in its practical guide to building agents. Anthropic presents the same pattern as an augmented LLM and recommends using the simplest system that meets the task.

Step 1: Define a narrow job and a contract

Write the contract before writing prompts or code. A good first task has a clear input, a finite action set and an observable result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Input: one fully qualified HTTP or HTTPS URL and the user’s question about it.
  • Allowed action: open that URL in a browser and save one screenshot.
  • Output: page title, final URL, viewport dimensions, screenshot path and any loading error.
  • Must not do: log in, submit forms, follow user-supplied credentials, run arbitrary JavaScript, visit a second domain or modify external data.
  • Stop conditions: a final answer, a tool error, invalid arguments or six model turns.

Keep the first version single-purpose. If the task can be completed with a fixed sequence and programmatic checks, a prompt chain may be easier to inspect than an open-ended agent loop.

Step 2: Choose how much orchestration you own

Direct API calls, an SDK and a managed runtime are control choices, not interchangeable product tiers. Compare them against your workflow.

Approach Run-loop control Implementation effort State and tools Best fit
Direct model API You write the loop, state transitions, validation and stop rules. Highest initial effort, clearest behavior. Your application executes tools and stores state. Short, fixed workflows, custom approvals and strict audit requirements.
Agent SDK The SDK can manage turns, tool execution, guardrails, handoffs, sessions and tracing. Less repeated plumbing; you learn the SDK’s abstractions. SDK facilities plus your own integrations. Teams that want reusable orchestration without building every primitive.
Managed runtime More session and orchestration infrastructure is provided for you. Fastest path to a larger deployment, with less low-level control. Runtime owns more persistence and execution responsibility. Open-ended, multi-step applications where operational infrastructure is the main burden.

OpenAI documents these alternatives—including the Responses API, Agents SDK and managed Agents API—in its agents documentation. Start with direct control for this tutorial so every model decision and tool call is visible.

Step 3: Create the smallest Python project

  1. Create an isolated environment. Run python -m venv .venv, activate it, then upgrade pip with python -m pip install --upgrade pip.
  2. Install dependencies. Run pip install openai playwright, followed by playwright install chromium to install the browser used by the tool.
  3. Set credentials. Export your provider’s API key as OPENAI_API_KEY and choose a model available to your account in MODEL. Never place keys in the prompt, source repository or tool arguments.
  4. Create agent.py. The complete loop below validates the URL, limits browser behavior and stops after six turns.

The vendor-specific Agents SDK Python quickstart shows a project and virtual-environment setup and the openai-agents package. You can use that SDK later; the code here deliberately keeps orchestration in your application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 4: Write instructions that are operational

Instructions should define role, boundaries, tool-selection rules and the shape of a good answer. Avoid vague directions such as “be helpful.” In this example, the system message says to use the tool only for the supplied URL, never claim a visual fact that the tool did not return, and finish with a short report or an explicit limitation.

Step 5: Implement one narrow tool

A tool is ordinary application code with real permissions. The function below accepts only HTTP(S) URLs, rejects credentials and writes a screenshot to a local directory. In production, add domain allowlists, an outbound proxy, resource limits and a temporary filesystem.

import json
import os
from pathlib import Path
from urllib.parse import urlparse

from openai import OpenAI
from playwright.sync_api import sync_playwright

MODEL = os.environ.get('MODEL')
if not MODEL:
    raise RuntimeError('Set MODEL to a model available in your account')

client = OpenAI()
OUTPUT_DIR = Path('shots')
OUTPUT_DIR.mkdir(exist_ok=True)

TOOLS = [{
    'type': 'function',
    'name': 'capture_page',
    'description': 'Open one public HTTP(S) URL and save a full-page PNG screenshot.',
    'parameters': {
        'type': 'object',
        'properties': {
            'url': {'type': 'string', 'description': 'The page to open'},
            'filename': {'type': 'string', 'description': 'A simple PNG filename'}
        },
        'required': ['url', 'filename'],
        'additionalProperties': False
    }
}]

def capture_page(url: str, filename: str) -> dict:
    parsed = urlparse(url)
    if parsed.scheme not in {'http', 'https'} or not parsed.netloc:
        raise ValueError('Only absolute HTTP(S) URLs are allowed')
    if parsed.username or parsed.password:
        raise ValueError(' URLs containing credentials are not allowed')
    safe_name = Path(filename).name
    if not safe_name.endswith('.png'):
        safe_name += '.png'
    destination = OUTPUT_DIR / safe_name
    with sync_playwright() as playwright:
        browser = playwright.chromium.launch(headless=True)
        page = browser.new_page(viewport={'width': 1440, 'height': 900})
        try:
            response = page.goto(url, wait_until='networkidle', timeout=30000)
            page.screenshot(path=str(destination), full_page=True)
            return {
                'requested_url': url,
                'final_url': page.url,
                'title': page.title(),
                'status': response.status if response else None,
                'screenshot': str(destination)
            }
        finally:
            browser.close()

def run_agent(task: str) -> str:
    messages = [
        {'role': 'system', 'content': (
            'You are a bounded website-checking agent. Use capture_page only for '
            'the URL in the user request. Never invent page facts. Return a concise '
            'report with the final URL, title, HTTP status and screenshot path. '
            'If the tool fails, explain the failure and stop.'
        )},
        {'role': 'user', 'content': task}
    ]
    for turn in range(6):
        response = client.responses.create(
            model=MODEL,
            input=messages,
            tools=TOOLS
        )
        calls = [item for item in response.output
                 if item.type == 'function_call']
        if not calls:
            return response.output_text
        for call in calls:
            try:
                arguments = json.loads(call.arguments)
                result = capture_page(**arguments)
            except Exception as exc:
                result = {'error': str(exc)}
            messages.append(call.model_dump())
            messages.append({
                'type': 'function_call_output',
                'call_id': call.call_id,
                'output': json.dumps(result)
            })
    raise RuntimeError('Maximum turns reached without a final answer')

if __name__ == '__main__':
    request = input('URL and question: ')
    print(run_agent(request))

The model interface in this example follows the Responses API pattern: send conversation input and tool definitions, inspect returned function calls, execute approved code, append the result, and call the model again. If your provider uses different request or response objects, replace only the client.responses.create section; keep validation, permissions and stop rules in your application.

Step 6: Understand the run loop

  1. Build the current message list and expose only the tools this run needs.
  2. Call the model.
  3. If it returns a final response, return it to the user.
  4. If it requests a function, parse JSON arguments and validate them with normal code.
  5. Execute the function with least-privilege credentials and a timeout.
  6. Append the tool result, then call the model again.
  7. Stop on a final response, tool error, policy violation or maximum turn count.

OpenAI’s practical guide describes every orchestration approach as needing a “run,” typically implemented as a loop that operates until an exit condition. That exit condition is a design requirement, not an optional optimization.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 7: Run and evaluate it

Set MODEL and OPENAI_API_KEY, then run python agent.py. Try a normal public page, a redirect, a page that returns an error status, an invalid URL and a page that never reaches network idle. Record whether the agent selected the tool, whether arguments passed validation, whether the returned observation was used correctly and whether the final answer stayed within the contract.

Keep a small evaluation set in version control. Include successful cases, malformed inputs, denied domains, timeouts and misleading page titles. Review model traces and tool logs, then improve descriptions and instructions before adding another tool. Anthropic’s guidance, Building Effective AI Agents, emphasizes testing and choosing the right level of complexity rather than maximizing sophistication.

Safety and reliability controls

Authentication and authorization

Authenticate users in your application and authorize each tool call separately. Use service accounts with only the permissions required for the task. A prompt is not a security boundary.

Argument and output validation

Validate types, ranges, URL schemes, hostnames and filenames before execution. Treat tool output as untrusted data. Before displaying or persisting a model-generated result, check required fields and reject unsupported claims.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Isolation and approvals

Run browser, code or file tools in a sandbox with network, CPU, memory and wall-clock limits. Require a human approval step before sending messages, changing records, spending money or performing other consequential actions.

Observability and recovery

Log run IDs, model responses, tool arguments, tool results, latency, errors and the stop reason without logging secrets. Make retries explicit and bounded; retrying a non-idempotent action can duplicate it. Preserve enough state to reproduce a failure.

State, memory and retrieval

Do not add a database merely because agents are described as having memory. Persist only information needed across turns, such as a task ID or approved preferences, and define retention and deletion rules. Retrieval is useful when the agent must consult documents, but retrieved text still needs source boundaries and injection-resistant handling. Start with in-memory state for one request, then add durable state when evaluations show a real requirement.

When one agent is not enough

Maximize a single agent’s capability before splitting it. Multiple agents can help when responsibilities are genuinely different or one agent repeatedly chooses the wrong tool, but they add handoffs, coordination overhead, duplicated context and a question of who owns the final response. Measure task quality and failure rates before introducing specialists. A deterministic programmatic chain is often preferable for a known sequence; an agent loop fits work whose number of steps is not known in advance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, cost and operational trade-offs

  • Bound turns and context: every additional model call increases latency and can compound an earlier mistake.
  • Keep tools focused: narrow schemas reduce argument errors and make traces easier to audit.
  • Set timeouts: browser navigation, network requests and model calls need separate deadlines.
  • Control concurrency: cap simultaneous runs so a burst cannot exhaust browser workers or API quotas.
  • Cache safely: cache only observations that remain valid for the task and do not contain sensitive data.
  • Track unit economics: record model calls, input and output tokens where available, browser time and storage per successful run.

Autonomy can raise cost and compound errors, so wider permissions should follow evidence from sandboxed tests rather than optimism.

Troubleshooting common failures

The model never calls the tool

Check that the tool is included in the request, its description states when it should be used, and the user input contains a valid URL. Avoid exposing several overlapping tools while debugging.

Invalid function arguments

Inspect the raw arguments, tighten the JSON schema and validate again in Python. Never pass unvalidated arguments directly to a browser, shell or database.

Browser launch or timeout errors

Run playwright install chromium, verify the process has a writable shots directory, and test the URL manually. Increase the navigation timeout only after checking DNS, TLS, redirects and pages that never become idle.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The loop repeats forever

Enforce the maximum-turn counter, return tool errors as observations, and include an instruction to stop after an error. Log each call to identify whether the model is receiving the previous tool result.

Secrets appear in logs

Redact authorization headers, cookies, API keys and personally identifying values before storing traces. Keep secrets in environment variables or a secret manager, not in prompts or screenshots.

Results are plausible but wrong

Add representative counterexamples, require the agent to cite only returned fields, and inspect traces. Improve the tool contract and checks before increasing model autonomy.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your agent only needs a clean website image, ScreenshotNeo provides a single HTTP request instead of maintaining Playwright workers. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and whether the request was billed. It also offers an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo documentation for parameters and authentication. The same endpoint can return PNG, JPEG, WebP or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every feature is included on every plan. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. You can call the endpoint from your tool function, inspect its X-Page-Verdict and X-Billed headers, and pass the resulting image or status back to the agent. Create a free ScreenshotNeo account to get started.

FAQ

Can I build an agent without an SDK?

Yes. A direct model API plus an application-owned loop is enough. An SDK becomes valuable when repeated concerns such as sessions, tracing, handoffs and guardrails outweigh the benefit of keeping every mechanism local.

Does an agent need a vector database?

No. Add retrieval only when the task requires information outside the current prompt and tool results. A narrow agent can work with no persistent memory or database.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How many tools should the first agent have?

One is a strong default. A single, well-described tool makes selection errors and permission boundaries observable; add another only after tests show the first design is insufficient.

What should a production approval record contain?

Store the user identity, requested action, exact validated arguments, timestamp, approving person or policy, tool result and final outcome. Exclude secrets and unnecessary personal data.

Frequently Asked Questions

Can I build an agent without an SDK?

Yes. A direct model API plus an application-owned loop is sufficient; use an SDK when managed sessions, tracing, handoffs and guardrails justify its abstractions.

Does an agent need a vector database?

No. Add retrieval or durable memory only when evaluations show the task needs information beyond the current conversation and tool results.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How many tools should a first agent have?

Start with one narrowly scoped tool so selection, validation and permissions are easy to inspect.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.