Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesBuild a useful first agent by combining four parts: a model that reasons, instructions that define its job, one narrowly scoped tool, and an application loop with explicit stop conditions. This guide builds a website-checking agent in Python, shows the direct API loop, compares SDK and managed-runtime choices, and adds the validation, permissions, testing and failure handling needed before giving an agent real access.
What you are building
The example agent accepts a URL, uses one capture_page tool to open it in a sandboxed browser, and returns a concise report. The model may request the tool, inspect its result, or answer that it cannot complete the task. Your application—not the prompt—owns tool execution, limits the number of turns, validates arguments and decides when the run ends.
This is the practical definition of an agent: an LLM augmented with tools, and optionally retrieval or memory, operating inside a controlled cycle. A plain model call can generate text; an agent adds the ability to choose an approved action, receive its result and continue until a defined exit condition. OpenAI describes the model, tools and instructions as the core components in its practical guide to building agents. Anthropic presents the same pattern as an augmented LLM and recommends using the simplest system that meets the task.
Step 1: Define a narrow job and a contract
Write the contract before writing prompts or code. A good first task has a clear input, a finite action set and an observable result.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
- Input: one fully qualified HTTP or HTTPS URL and the user’s question about it.
- Allowed action: open that URL in a browser and save one screenshot.
- Output: page title, final URL, viewport dimensions, screenshot path and any loading error.
- Must not do: log in, submit forms, follow user-supplied credentials, run arbitrary JavaScript, visit a second domain or modify external data.
- Stop conditions: a final answer, a tool error, invalid arguments or six model turns.
Keep the first version single-purpose. If the task can be completed with a fixed sequence and programmatic checks, a prompt chain may be easier to inspect than an open-ended agent loop.
Step 2: Choose how much orchestration you own
Direct API calls, an SDK and a managed runtime are control choices, not interchangeable product tiers. Compare them against your workflow.
| Approach | Run-loop control | Implementation effort | State and tools | Best fit |
|---|---|---|---|---|
| Direct model API | You write the loop, state transitions, validation and stop rules. | Highest initial effort, clearest behavior. | Your application executes tools and stores state. | Short, fixed workflows, custom approvals and strict audit requirements. |
| Agent SDK | The SDK can manage turns, tool execution, guardrails, handoffs, sessions and tracing. | Less repeated plumbing; you learn the SDK’s abstractions. | SDK facilities plus your own integrations. | Teams that want reusable orchestration without building every primitive. |
| Managed runtime | More session and orchestration infrastructure is provided for you. | Fastest path to a larger deployment, with less low-level control. | Runtime owns more persistence and execution responsibility. | Open-ended, multi-step applications where operational infrastructure is the main burden. |
OpenAI documents these alternatives—including the Responses API, Agents SDK and managed Agents API—in its agents documentation. Start with direct control for this tutorial so every model decision and tool call is visible.
Step 3: Create the smallest Python project
- Create an isolated environment. Run
python -m venv .venv, activate it, then upgrade pip withpython -m pip install --upgrade pip. - Install dependencies. Run
pip install openai playwright, followed byplaywright install chromiumto install the browser used by the tool. - Set credentials. Export your provider’s API key as
OPENAI_API_KEYand choose a model available to your account inMODEL. Never place keys in the prompt, source repository or tool arguments. - Create
agent.py. The complete loop below validates the URL, limits browser behavior and stops after six turns.
The vendor-specific Agents SDK Python quickstart shows a project and virtual-environment setup and the openai-agents package. You can use that SDK later; the code here deliberately keeps orchestration in your application.
Recommended Free Tools
Step 4: Write instructions that are operational
Instructions should define role, boundaries, tool-selection rules and the shape of a good answer. Avoid vague directions such as “be helpful.” In this example, the system message says to use the tool only for the supplied URL, never claim a visual fact that the tool did not return, and finish with a short report or an explicit limitation.
Step 5: Implement one narrow tool
A tool is ordinary application code with real permissions. The function below accepts only HTTP(S) URLs, rejects credentials and writes a screenshot to a local directory. In production, add domain allowlists, an outbound proxy, resource limits and a temporary filesystem.
Rank #2
import json
import os
from pathlib import Path
from urllib.parse import urlparse
from openai import OpenAI
from playwright.sync_api import sync_playwright
MODEL = os.environ.get('MODEL')
if not MODEL:
raise RuntimeError('Set MODEL to a model available in your account')
client = OpenAI()
OUTPUT_DIR = Path('shots')
OUTPUT_DIR.mkdir(exist_ok=True)
TOOLS = [{
'type': 'function',
'name': 'capture_page',
'description': 'Open one public HTTP(S) URL and save a full-page PNG screenshot.',
'parameters': {
'type': 'object',
'properties': {
'url': {'type': 'string', 'description': 'The page to open'},
'filename': {'type': 'string', 'description': 'A simple PNG filename'}
},
'required': ['url', 'filename'],
'additionalProperties': False
}
}]
def capture_page(url: str, filename: str) -> dict:
parsed = urlparse(url)
if parsed.scheme not in {'http', 'https'} or not parsed.netloc:
raise ValueError('Only absolute HTTP(S) URLs are allowed')
if parsed.username or parsed.password:
raise ValueError(' URLs containing credentials are not allowed')
safe_name = Path(filename).name
if not safe_name.endswith('.png'):
safe_name += '.png'
destination = OUTPUT_DIR / safe_name
with sync_playwright() as playwright:
browser = playwright.chromium.launch(headless=True)
page = browser.new_page(viewport={'width': 1440, 'height': 900})
try:
response = page.goto(url, wait_until='networkidle', timeout=30000)
page.screenshot(path=str(destination), full_page=True)
return {
'requested_url': url,
'final_url': page.url,
'title': page.title(),
'status': response.status if response else None,
'screenshot': str(destination)
}
finally:
browser.close()
def run_agent(task: str) -> str:
messages = [
{'role': 'system', 'content': (
'You are a bounded website-checking agent. Use capture_page only for '
'the URL in the user request. Never invent page facts. Return a concise '
'report with the final URL, title, HTTP status and screenshot path. '
'If the tool fails, explain the failure and stop.'
)},
{'role': 'user', 'content': task}
]
for turn in range(6):
response = client.responses.create(
model=MODEL,
input=messages,
tools=TOOLS
)
calls = [item for item in response.output
if item.type == 'function_call']
if not calls:
return response.output_text
for call in calls:
try:
arguments = json.loads(call.arguments)
result = capture_page(**arguments)
except Exception as exc:
result = {'error': str(exc)}
messages.append(call.model_dump())
messages.append({
'type': 'function_call_output',
'call_id': call.call_id,
'output': json.dumps(result)
})
raise RuntimeError('Maximum turns reached without a final answer')
if __name__ == '__main__':
request = input('URL and question: ')
print(run_agent(request))
The model interface in this example follows the Responses API pattern: send conversation input and tool definitions, inspect returned function calls, execute approved code, append the result, and call the model again. If your provider uses different request or response objects, replace only the client.responses.create section; keep validation, permissions and stop rules in your application.
Step 6: Understand the run loop
- Build the current message list and expose only the tools this run needs.
- Call the model.
- If it returns a final response, return it to the user.
- If it requests a function, parse JSON arguments and validate them with normal code.
- Execute the function with least-privilege credentials and a timeout.
- Append the tool result, then call the model again.
- Stop on a final response, tool error, policy violation or maximum turn count.
OpenAI’s practical guide describes every orchestration approach as needing a “run,” typically implemented as a loop that operates until an exit condition. That exit condition is a design requirement, not an optional optimization.
Step 7: Run and evaluate it
Set MODEL and OPENAI_API_KEY, then run python agent.py. Try a normal public page, a redirect, a page that returns an error status, an invalid URL and a page that never reaches network idle. Record whether the agent selected the tool, whether arguments passed validation, whether the returned observation was used correctly and whether the final answer stayed within the contract.
Keep a small evaluation set in version control. Include successful cases, malformed inputs, denied domains, timeouts and misleading page titles. Review model traces and tool logs, then improve descriptions and instructions before adding another tool. Anthropic’s guidance, Building Effective AI Agents, emphasizes testing and choosing the right level of complexity rather than maximizing sophistication.
Safety and reliability controls
Authentication and authorization
Authenticate users in your application and authorize each tool call separately. Use service accounts with only the permissions required for the task. A prompt is not a security boundary.
Argument and output validation
Validate types, ranges, URL schemes, hostnames and filenames before execution. Treat tool output as untrusted data. Before displaying or persisting a model-generated result, check required fields and reject unsupported claims.
Isolation and approvals
Run browser, code or file tools in a sandbox with network, CPU, memory and wall-clock limits. Require a human approval step before sending messages, changing records, spending money or performing other consequential actions.
Observability and recovery
Log run IDs, model responses, tool arguments, tool results, latency, errors and the stop reason without logging secrets. Make retries explicit and bounded; retrying a non-idempotent action can duplicate it. Preserve enough state to reproduce a failure.
State, memory and retrieval
Do not add a database merely because agents are described as having memory. Persist only information needed across turns, such as a task ID or approved preferences, and define retention and deletion rules. Retrieval is useful when the agent must consult documents, but retrieved text still needs source boundaries and injection-resistant handling. Start with in-memory state for one request, then add durable state when evaluations show a real requirement.
When one agent is not enough
Maximize a single agent’s capability before splitting it. Multiple agents can help when responsibilities are genuinely different or one agent repeatedly chooses the wrong tool, but they add handoffs, coordination overhead, duplicated context and a question of who owns the final response. Measure task quality and failure rates before introducing specialists. A deterministic programmatic chain is often preferable for a known sequence; an agent loop fits work whose number of steps is not known in advance.
Performance, cost and operational trade-offs
- Bound turns and context: every additional model call increases latency and can compound an earlier mistake.
- Keep tools focused: narrow schemas reduce argument errors and make traces easier to audit.
- Set timeouts: browser navigation, network requests and model calls need separate deadlines.
- Control concurrency: cap simultaneous runs so a burst cannot exhaust browser workers or API quotas.
- Cache safely: cache only observations that remain valid for the task and do not contain sensitive data.
- Track unit economics: record model calls, input and output tokens where available, browser time and storage per successful run.
Autonomy can raise cost and compound errors, so wider permissions should follow evidence from sandboxed tests rather than optimism.
Troubleshooting common failures
The model never calls the tool
Check that the tool is included in the request, its description states when it should be used, and the user input contains a valid URL. Avoid exposing several overlapping tools while debugging.
Invalid function arguments
Inspect the raw arguments, tighten the JSON schema and validate again in Python. Never pass unvalidated arguments directly to a browser, shell or database.
Browser launch or timeout errors
Run playwright install chromium, verify the process has a writable shots directory, and test the URL manually. Increase the navigation timeout only after checking DNS, TLS, redirects and pages that never become idle.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →The loop repeats forever
Enforce the maximum-turn counter, return tool errors as observations, and include an instruction to stop after an error. Log each call to identify whether the model is receiving the previous tool result.
Secrets appear in logs
Redact authorization headers, cookies, API keys and personally identifying values before storing traces. Keep secrets in environment variables or a secret manager, not in prompts or screenshots.
Results are plausible but wrong
Add representative counterexamples, require the agent to cite only returned fields, and inspect traces. Improve the tool contract and checks before increasing model autonomy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your agent only needs a clean website image, ScreenshotNeo provides a single HTTP request instead of maintaining Playwright workers. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and whether the request was billed. It also offers an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
See the ScreenshotNeo documentation for parameters and authentication. The same endpoint can return PNG, JPEG, WebP or PDF:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every feature is included on every plan. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. You can call the endpoint from your tool function, inspect its X-Page-Verdict and X-Billed headers, and pass the resulting image or status back to the agent. Create a free ScreenshotNeo account to get started.
FAQ
Can I build an agent without an SDK?
Yes. A direct model API plus an application-owned loop is enough. An SDK becomes valuable when repeated concerns such as sessions, tracing, handoffs and guardrails outweigh the benefit of keeping every mechanism local.
Does an agent need a vector database?
No. Add retrieval only when the task requires information outside the current prompt and tool results. A narrow agent can work with no persistent memory or database.
How many tools should the first agent have?
One is a strong default. A single, well-described tool makes selection errors and permission boundaries observable; add another only after tests show the first design is insufficient.
What should a production approval record contain?
Store the user identity, requested action, exact validated arguments, timestamp, approving person or policy, tool result and final outcome. Exclude secrets and unnecessary personal data.
Frequently Asked Questions
Can I build an agent without an SDK?
Yes. A direct model API plus an application-owned loop is sufficient; use an SDK when managed sessions, tracing, handoffs and guardrails justify its abstractions.
Does an agent need a vector database?
No. Add retrieval or durable memory only when evaluations show the task needs information beyond the current conversation and tool results.
Free tools Windows power users keep installed
One-click scans. No signup required.
How many tools should a first agent have?
Start with one narrowly scoped tool so selection, validation and permissions are easy to inspect.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




