Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To connect Ollama to tools exposed by an MCP server, your Python application must bridge two interfaces: use the MCP Python SDK to discover and call tools, then translate those tools into Ollama function definitions and dispatch the tool calls Ollama returns. The official documentation describes these pieces separately; it does not provide a single general-purpose client that handles the entire bridge. The example below shows the integration pattern and flags the version-specific details you must verify before deploying it.
What the Ollama–MCP client does
Ollama handles model chat and returns tool calls selected by the model. MCP handles the connection to a tool server: the client opens a session, lists available tools, and invokes a named tool with arguments. Your Python application joins the two:
- Connect to an MCP server and discover its tools.
- Convert each MCP tool name, description, and JSON input schema into an Ollama function definition.
- Send the user’s message and those definitions to Ollama.
- For each tool call Ollama returns, check that the tool is allowed, then invoke it through MCP.
- Append the tool results to the conversation and ask Ollama to continue.
This is an integration pattern assembled from the documented Ollama and MCP interfaces, not a claim that either project ships a ready-made, universal bridge. Ollama describes tool support in its tool support documentation; the MCP Python SDK documents client sessions and tool operations in its Python SDK and client guide.
Prerequisites and versions
Python and packages
Use Python 3.10 or newer for the combined example: the Ollama Python library documents Python 3.8+, while the current stable MCP Python SDK v2 requires Python 3.10+. Install both packages:
#1 Best Overall
python -m pip install ollama "mcp[cli]"
Ollama documents the ollama package and its async client in the Ollama Python library. The MCP SDK installation options and runtime requirements are in the MCP Python SDK documentation.
Keep MCP SDK generations consistent
The MCP Python SDK currently identifies v2 as its stable release line. Its v1 documentation is maintained separately and advises projects remaining on v1 to pin below v2; it gives mcp>=1.28,<2 as an example. Do not combine imports or examples from the v1 and v2 documentation without checking which version you installed. See the SDK documentation for v2 and the v1.x documentation for the maintenance line.
Run the two services independently
Ollama is the model endpoint; the MCP server is the tool endpoint. Start or locate each separately. Local Ollama normally serves its API at http://localhost:11434/api. For MCP, the SDK supports stdio, Streamable HTTP, and SSE. The high-level client accepts a URL for Streamable HTTP or subprocess parameters for a local stdio server. Use the transport that matches how your MCP server is deployed, not the one chosen for Ollama.
Build the basic client
The code is an integration outline, not a verified drop-in application. In particular, check the installed SDK’s import and typed-result shapes, the Ollama response serialization method, and the selected model’s tool-calling support. Replace the MCP URL and model name with values for your setup.
import asyncio
import ollama
from mcp import Client
async def main():
# This URL is an example Streamable HTTP MCP endpoint.
async with Client("http://localhost:8000/mcp") as mcp:
# A simple server may return all tools at once. See pagination below
# if the server returns a next-page cursor.
page = await mcp.list_tools()
mcp_tools = {tool.name: tool for tool in page.tools}
ollama_tools = [
{
"type": "function",
"function": {
"name": tool.name,
"description": tool.description or "",
"parameters": tool.input_schema,
},
}
for tool in mcp_tools.values()
]
messages = [
{"role": "user", "content": "Use the available tools to answer my question."}
]
response = ollama.chat(
model="<tool-capable-model>",
messages=messages,
tools=ollama_tools,
)
messages.append(response.message.model_dump(exclude_none=True))
for call in response.message.tool_calls or []:
name = call.function.name
if name not in mcp_tools:
raise ValueError(f"Model requested an undiscovered tool: {name}")
result = await mcp.call_tool(name, call.function.arguments)
text_result = "n".join(
block.text for block in result.content if hasattr(block, "text")
)
if result.is_error:
text_result = f"Tool reported an error: {text_result}"
messages.append(
{"role": "tool", "tool_name": name, "content": text_result}
)
final = ollama.chat(
model="<tool-capable-model>",
messages=messages,
tools=ollama_tools,
)
print(final.message.content)
asyncio.run(main())
The calls to ollama.chat are synchronous in this outline, while the MCP client operations are awaited. The Ollama Python library also documents an async client if you need asynchronous model requests. Its response and tool-call fields are covered in the library examples and the Ollama API documentation. The MCP operations are described in the client guide.
Rank #2
How to make the bridge robust
Collect every page of tools
The compact example reads one page. MCP servers can return a pagination cursor, so a production client should continue listing until the response has no next-page cursor, then build the Ollama tool list from the complete collection. Follow the pagination fields in the installed SDK version’s client guide.
Preserve schemas and validate arguments
Map the MCP tool’s advertised name, description, and JSON input schema to the Ollama function definition. Optional metadata can differ among tools, so do not assume every tool has identical fields. Validate the model-produced arguments against the advertised input schema before calling MCP. A schema guides the model; it is not an authorization boundary.
Handle tool errors and content deliberately
MCP tool results include content and an error indicator. Make error results visible to the model rather than presenting a failed call as a successful answer. The example joins text blocks as a simple starting point; real servers can return other content types, so adapt conversion to the content you intend to pass to the model. Bound the size of tool output to avoid flooding the conversation with large results.
Support multiple calls and another tool turn
The example dispatches every call in the first response, then asks Ollama once more for a final answer. A model may request more tools after seeing those results. A complete client should run a controlled loop: send the current history, dispatch any returned calls, append their results, and continue until the assistant responds without tool calls or an application limit is reached. Preserve the assistant message containing the calls and a corresponding tool result for each execution.
Choose the MCP transport
| Transport | Connection pattern | When it fits |
|---|---|---|
| stdio | The client launches a local server subprocess using server parameters. | Useful when the MCP server runs locally as a child process. |
| Streamable HTTP | The client connects to a server URL. | Useful for a URL-based MCP deployment. |
| SSE | Supported by the SDK as another MCP transport. | Use when the server and deployment are set up for SSE. |
The MCP Python SDK supports these transports; the appropriate one depends on the server’s setup. Keep the MCP address or subprocess configuration distinct from the Ollama inference host. Details are in the MCP Python SDK documentation and its client guide.
Choose local or hosted Ollama
| Mode | Endpoint | Authentication |
|---|---|---|
| Local Ollama | Normally http://localhost:11434/api |
No hosted API key is needed for local requests. |
| Ollama hosted API | https://ollama.com |
Use an Authorization: Bearer <OLLAMA_API_KEY> header. |
The Ollama Python library documents local and hosted client configuration, and Ollama’s API introduction explains the endpoint and authentication distinction. The documented setup establishes where requests go and whether a key is needed; it does not establish a general cost, latency, privacy, or quality ranking between the two modes. Keep hosted credentials in environment or secret-management configuration, not committed source or browser code.
Model choice and streaming
Confirm the model can call tools
Tool calling is model-specific; do not assume every Ollama model emits tool calls. Ollama’s tool-support article lists models it supported at publication. Its May 28, 2025 post on streaming responses with tool calling named Qwen 3, Devstral, Qwen2.5 and Qwen2.5-Coder, Llama 3.1, Llama 4, and others. That is a dated list, not a guarantee about every current model or later model release; check the selected model’s current capabilities.
Add streaming after the non-streaming flow works
Ollama’s SDKs do not stream by default; streaming is enabled with stream=True. In a streamed tool turn, accumulate chunks before dispatching calls and preserve the assembled assistant turn in the conversation history. Tool-call data can arrive across chunks, so treating an individual chunk as a complete call can produce incomplete arguments. See Ollama’s streaming documentation and its May 28, 2025 tool-calling streaming announcement.
Ollama’s May 2025 post says a context window of 32k or more may improve tool calling, but presents that as anecdotal and notes that longer contexts use more memory. Treat it as an observation, not a measured benchmark or universal requirement.
Security and operational checks
- Allowlist calls: only dispatch names discovered from the current MCP session. A model’s tool call is a request, not permission to run arbitrary code.
- Validate inputs: check arguments against each tool’s schema and apply the MCP server’s own permission checks.
- Limit output: bound tool content and avoid sending secrets from the MCP environment into model context unless that is intentional.
- Report failures: pass bounded, explicit error results back through the conversation rather than silently treating failures as success.
- Close resources: use the MCP client’s async context manager so sessions, HTTP resources, and launched child processes follow the SDK lifecycle.
- Separate credentials: keep hosted Ollama API keys server-side and outside source control.
Troubleshooting
Import or client-constructor errors
Likely cause: code and installed MCP SDK belong to different API generations. Fix: check the installed version and use matching v2 or pinned v1 documentation; do not mix their imports and examples.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesNo tool calls appear
Likely cause: the selected model does not support tool calling, the tools were not passed in the request, or the model answered directly. Fix: confirm current model support, inspect the request’s tools definitions, and handle responses with no tool calls as normal assistant answers.
The server is reachable but the tool list is empty or incomplete
Likely cause: the wrong MCP endpoint or transport was selected, or the response has additional pages. Fix: verify that the MCP server is running and its configured URL or stdio command is correct; follow any pagination cursor until the tool list is complete.
A tool call fails validation or the server returns an error
Likely cause: arguments do not match the advertised input schema, required server-side permissions are missing, or the tool itself failed. Fix: validate arguments before dispatch, inspect the MCP error indicator and bounded result content, and return the failure as an error result instead of claiming success.
The model cannot connect
Likely cause: the Ollama host is unavailable or the client is pointed at the wrong endpoint. Fix: distinguish the local Ollama API endpoint from the MCP server address; for hosted access, use the documented hosted endpoint and bearer key configuration.
Recommended Free Tools
Streaming produces malformed calls
Likely cause: the application dispatches a partial streamed tool call. Fix: accumulate the stream into the complete assistant response before parsing arguments or invoking MCP.
Best Value
Or skip the browser setup
If a workflow also needs website screenshots as a tool, ScreenshotNeo provides a website screenshot API and MCP server for developers. Its API can return a screenshot or PDF from one GET request; see the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
For Python, use the same endpoint with requests:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
ScreenshotNeo accepts cookie or consent banners as a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card required.
Frequently Asked Questions
Can Ollama and the MCP server use different transports?
Yes. The MCP transport connects your Python client to the tool server; it is separate from the Ollama model endpoint.
Does this example require Ollama Cloud?
No. It can use a local Ollama server; hosted API access is an alternative.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

