Free tools Windows power users keep installed
One-click scans. No signup required.
A runnable MCP loop has two halves. The MCP half discovers tools on a server and calls them. The model half, which belongs to your LLM provider’s API, decides whether to call a tool and with what arguments. This guide builds both halves in Python. The same loop runs over stdio (a local subprocess) or Streamable HTTP (a server listening on a port). The model is behind a small adapter, so you can run the whole thing without an API key and then plug in the provider you use.
What the loop does
- Start or connect to an MCP server.
- Ask the MCP client for the tool definitions: name, description and input schema.
- Give those definitions to your model provider in that provider’s tool format.
- If the model requests a tool, call it through MCP with the model’s arguments.
- Return the MCP result to the model as a tool result. Repeat until the model answers in plain text.
MCP standardizes how an application provides context and capabilities to an LLM. It does not replace the API you use to talk to the model. The SDK gives you discovery (list_tools()) and invocation (call_tool()). Mapping tool definitions into a provider’s request format, and results into its tool-result format, is orchestration code that you write. That is why this article keeps that layer behind an adapter instead of showing one vendor’s syntax.
Version and setup
The official MCP Python SDK documentation describes v2 as the stable release line and requires Python 3.10 or newer. Install it with uv add "mcp[cli]" or pip install "mcp[cli]". The [cli] extra provides the mcp development command.
The code below uses the long-standing ClientSession plus stdio_client sequence from the SDK’s simple-tool example, together with FastMCP. That is the v1.x API, so pin below v2:
Recommended Free Tools
#1 Best Overall
pip install "mcp[cli]>=1.28,<2"
The v1 line is in maintenance mode. The v2 client guide instead describes a context-managed Client object, and it reports the error flag as is_error. The v1 result object used below exposes it as isError. Do not mix v1 imports with v2 examples. If you target v2, check the SDK migration guide first. Only the connection code changes, and the loop in the middle stays the same.
stdio vs Streamable HTTP
| Axis | stdio | Streamable HTTP |
|---|---|---|
| Process arrangement | Host launches the server as a subprocess | Server listens independently on HTTP |
| Connection input | Command and arguments (StdioServerParameters) |
MCP endpoint URL, e.g. http://localhost:8000/mcp |
| Typical role | Local development, desktop-host style | Separately running or deployed service |
| Operational boundary | One local process relationship | Network endpoint, so deployment and access controls matter |
| SDK guidance | Default transport | Current HTTP transport |
SSE is the older HTTP transport. The SDK run guide says Streamable HTTP superseded it in the 2025-03-26 protocol revision. Use SSE only for compatibility with an existing server. The SDK’s own framing is that your only decision is the transport, meaning how the bytes between server and client move. The tools and the loop don’t change.
With stdio, the protocol runs over stdin and stdout. Never print() to stdout in a stdio server, because that corrupts the protocol stream. Send diagnostics to stderr.
Rank #2
Step 1: the MCP server
Save as server.py. mcp.run() blocks for the server’s lifetime and defaults to stdio. For Streamable HTTP the endpoint path defaults to /mcp, on host 127.0.0.1 and port 8000. The entry-point guard keeps import-based tools from starting the server accidentally.
import sys
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("demo")
@mcp.tool()
def add(a: int, b: int) -> int:
"""Add two integers."""
print("add called", file=sys.stderr) # stderr is safe on stdio
return a + b
if __name__ == "__main__":
transport = sys.argv[1] if len(sys.argv) > 1 else "stdio"
mcp.run(transport=transport) # "stdio" or "streamable-http"
For HTTP, start it separately with python server.py streamable-http. For stdio you don’t start it yourself, because the client launches it.
Step 2: the model adapter (the provider-specific part)
The loop needs one function: given the conversation and the tool list, return either text or tool requests. Real providers each define their own request, tool-declaration and tool-result shapes, so this is the only code you rewrite when you change provider. The version below is a scripted stand-in that asks for add once and then reports the result. It lets you verify the MCP plumbing end to end.
from dataclasses import dataclass, field
@dataclass
class ToolCall:
id: str
name: str
arguments: dict
@dataclass
class ModelTurn:
text: str = ""
tool_calls: list = field(default_factory=list)
async def scripted_model(messages, tools):
last = messages[-1]
if last["role"] == "tool":
if last["is_error"]:
return ModelTurn(text="The tool failed: " + last["content"])
return ModelTurn(text="The answer is " + last["content"])
assert any(t["name"] == "add" for t in tools)
return ModelTurn(tool_calls=[ToolCall("call_1", "add", {"a": 2, "b": 3})])
To use a real model, replace scripted_model with a function that does four things:
- Translates
tools(name, description,input_schema) into the provider’s tool declaration format. - Translates
messages, including thetoolentries, into the provider’s message and tool-result format. - Sends the request.
- Maps any tool-use response back into
ToolCallobjects. The model chooses the tool and the arguments. Your code only executes the choice.
Check the provider’s current documentation for those shapes. They change independently of MCP. Some agent SDKs, such as the OpenAI Agents SDK, can connect to MCP servers themselves, which removes this hand-written layer at the cost of a less visible loop.
Step 3: the transport-independent loop
The loop takes an initialized session. call_tool() returns content meant for the model, optional structured content for your application, and an error indicator. The loop keeps the error flag attached to the result, so a failed tool is never presented to the model as a success.
def render(result) -> str:
parts = [b.text for b in result.content if getattr(b, "type", "") == "text"]
return "n".join(parts) or "(no text content)"
async def run_loop(session, model, prompt, max_steps=5):
listed = await session.list_tools()
tools = [
{"name": t.name, "description": t.description or "", "input_schema": t.inputSchema}
for t in listed.tools
]
messages = [{"role": "user", "content": prompt}]
for _ in range(max_steps):
turn = await model(messages, tools)
if not turn.tool_calls:
return turn.text
messages.append({"role": "assistant", "content": turn.text,
"tool_calls": turn.tool_calls})
for call in turn.tool_calls:
result = await session.call_tool(call.name, call.arguments)
messages.append({
"role": "tool",
"call_id": call.id,
"is_error": bool(result.isError), # is_error in v2
"content": render(result),
})
raise RuntimeError("Model kept requesting tools; stopped after max_steps")
max_steps matters because a model can request tools indefinitely. If you use structured content in your own code, read it from the result rather than parsing the text.
Step 4: run it over stdio
Save as run_stdio.py next to the other files, with the adapter and loop code pasted above the main function or imported from a module. The client launches server.py as a subprocess, so no separate server is needed.
import asyncio
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
async def main():
params = StdioServerParameters(command="python", args=["server.py"])
async with stdio_client(params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
print(await run_loop(session, scripted_model, "What is 2 + 3?"))
asyncio.run(main())
Expected output: The answer is 5, plus add called on the server’s stderr.
Best Value
Step 5: run it over Streamable HTTP
Start the server in one terminal with python server.py streamable-http. Then run this client in another:
import asyncio
from mcp import ClientSession
from mcp.client.streamable_http import streamablehttp_client
async def main():
async with streamablehttp_client("http://localhost:8000/mcp") as (read, write, _):
async with ClientSession(read, write) as session:
await session.initialize()
print(await run_loop(session, scripted_model, "What is 2 + 3?"))
asyncio.run(main())
The output is identical. Only the connection block changed, which is the point of keeping the loop independent of transport. Because this server now listens on a network port, think about who can reach it before moving beyond 127.0.0.1.
Troubleshooting
- Garbled or failing stdio startup: something wrote to stdout. Move prints to stderr.
- HTTP connection refused: the server isn’t running, or the URL differs from the defaults (
127.0.0.1:8000, path/mcp). - Import errors on
mcp.clientorFastMCP: you probably installed v2 while following v1 code. Pinmcp>=1.28,<2or follow the migration guide. - Tool error treated as success: check that your adapter passes the error flag through in the provider’s tool-result format. Many providers have a dedicated way to mark a result as an error.
- Infinite tool calls: keep a step limit, and give tools precise descriptions and schemas so the model picks correctly.
Which transport should you use?
Use stdio while developing and for local tools that a host launches itself. Use Streamable HTTP when the server must run separately or be deployed, and plan access controls for it. Either way, the loop above stays the same.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




