Large Language ModelsGenerate imagesGenerate videos
Ollama MCP Server: How to Connect Local Models With a Bridge
Build an Ollama MCP server bridge in both directions: a Python client that lets local models call MCP tools, and a server that exposes Ollama to Claude Desktop and Claude Code. Includes hardware sizing, working code, config files, and fixes for the three bugs that break most first attempts.
Your laptop can already run a capable language model, but most of your tools have no idea it exists. An Ollama MCP server setup fixes that. A small bridge sits between Ollama, which serves models at localhost:11434, and the Model Context Protocol (MCP), the standard AI apps use to find and call tools. Build it once and a local model can read your files, query a database, or search your notes. Flip the direction and Claude Desktop or Claude Code can hand cheap, private jobs to the GPU under your desk.
This article builds both directions with working Python code: a bridge client that lets Ollama models use any MCP server, and a bridge server that exposes Ollama as a set of tools to any MCP client. You also get memory numbers, config files, and the three bugs that eat most of the debugging time.
What an Ollama MCP Bridge Does
Ollama and MCP solve different halves of one problem. Ollama downloads open-weight models and serves them over a plain HTTP API. MCP, which Anthropic published in late 2024, standardizes how an AI application talks to tools, files, and data through small programs called servers. Ollama's API speaks chat messages and tool definitions. It does not speak MCP, and MCP clients have no idea where your models live. The bridge translates between the two.
Two Ways to Build a Bridge
"Bridge" means two different programs depending on who needs what, so pick the direction before you write any code.
Direction
MCP client
MCP server
Typical use
Ollama uses tools
Your bridge script, wrapping an Ollama model
Filesystem, database, or search server
A local assistant that reads your notes
Ollama as a tool
Claude Desktop, Claude Code, Cursor
Your bridge script, wrapping Ollama
Offloading private or cheap tasks to a local model
The first direction gives a local model hands. The second gives a cloud assistant a local colleague. Most people end up building both, because the second one takes about 30 lines.
A quick note on transport. The examples here use stdio, where the client starts the server as a child process and talks to it over standard input and output. That is the simplest option for a personal machine. If you want one bridge shared by several apps or several computers on your network, run it over the Streamable HTTP transport instead and keep it behind your own firewall. The tool code stays the same, and only the last line of the server changes.
How a Tool Call Travels
Whichever direction you pick, a tool call follows the same loop:
The bridge connects to an MCP server and asks for its tool list with tools/list.
It rewrites each tool schema into the format Ollama expects.
It sends the user's question and those tool definitions to a local model.
The model answers with a tool_calls entry instead of text.
The bridge runs that call on the MCP server with tools/call.
The result goes back to the model as a tool message.
Steps 3 to 6 repeat until the model replies with plain text.
💡 The model never touches your disk. It only requests an action. Your bridge decides whether to run it, which makes the bridge the right place for allow lists, confirmation prompts, and logging.
What You Need First
Hardware That Works
Model weights have to fit in GPU memory (or unified memory on a Mac) with room left for context. These figures are approximate, for 4-bit quantized builds:
Model size
Memory for weights
Comfortable setup
3B to 4B
2 to 3 GB
Any recent laptop
7B to 8B
5 to 6 GB
8 GB GPU or 16 GB unified memory
14B
9 to 10 GB
12 GB GPU
20B
13 to 15 GB
16 GB GPU or 24 GB unified memory
32B
19 to 21 GB
24 GB GPU
Models that spill into system RAM still run, but token speed drops hard. A bridge makes several model calls per question, so speed matters more than size. An 8B model that fits entirely on the GPU usually beats a 32B model that does not.
Models That Support Tool Calls
Not every model can ask for a tool. Ollama checks the model's chat template, and a model without tool support returns an error when you pass tools. Filter the Ollama library by the tools tag, or start from this shortlist:
Ollama tag
Size on disk
Why use it
llama3.1:8b
about 4.9 GB
Reliable default for first tests
qwen3:8b
about 5.2 GB
Strong at multi-step tool use, slower when thinking is on
mistral-nemo
about 7.1 GB
Long context, handles many tools
gpt-oss:20b
about 14 GB
The open-weight GPT OSS 20B, built with tool use in mind
Install the Pieces
Activate a virtual environment first (source .venv/bin/activate on macOS and Linux, .venv\Scripts\activate on Windows), then run:
# 1. Pull a tool-capable model and confirm Ollama is serving
ollama pull llama3.1:8b
curl http://localhost:11434/api/tags
# 2. Install the two Python packages the bridge needs
pip install ollama mcp
# 3. Check Node, because the example MCP server runs through npx
node --version
If curl returns a JSON list of your models, Ollama is ready. If the connection is refused, start it with ollama serve.
Build the Bridge Client in Python
The client is one file. It launches an MCP server as a child process over the stdio transport, reads its tools, and runs the loop from the previous section. The example server is the official filesystem server, pointed at a notes folder.
Convert MCP Tools to Ollama Format
An MCP tool carries a name, a description, and an inputSchema written in JSON Schema. Ollama's function-calling format wants the same three pieces, wrapped in a function object. The conversion is mostly a rename:
Most schemas pass through untouched. If a server ships exotic constructs such as anyOf or $ref, flatten them before sending. Small models handle flat schemas far better.
Write the Tool Call Loop
import asyncio
import ollama
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
MODEL = "llama3.1:8b"
NOTES_DIR = "/home/me/notes"
MAX_TURNS = 8
server_params = StdioServerParameters(
command="npx",
args=["-y", "@modelcontextprotocol/server-filesystem", NOTES_DIR],
)
def to_ollama_tool(tool):
return {
"type": "function",
"function": {
"name": tool.name,
"description": tool.description or "",
"parameters": tool.inputSchema,
},
}
async def ask(question: str) -> str:
async with stdio_client(server_params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
listed = await session.list_tools()
tools = [to_ollama_tool(t) for t in listed.tools]
client = ollama.AsyncClient()
messages = [{"role": "user", "content": question}]
for _ in range(MAX_TURNS):
response = await client.chat(model=MODEL, messages=messages, tools=tools)
message = response.message
messages.append(message)
if not message.tool_calls:
return message.content
for call in message.tool_calls:
result = await session.call_tool(
call.function.name, dict(call.function.arguments)
)
text = "\n".join(
block.text for block in result.content if block.type == "text"
)
messages.append(
{"role": "tool", "tool_name": call.function.name, "content": text}
)
return "Stopped: the model kept calling tools."
if __name__ == "__main__":
print(asyncio.run(ask("Which markdown files in my notes folder mention invoices?")))
Four details deserve attention:
MAX_TURNS is a safety rail. A confused model can call the same tool forever, and a counter turns that into a clear failure.
dict(call.function.arguments) matters because Ollama returns arguments as a mapping that the MCP session expects as a plain dictionary.
Text blocks only. MCP results can include images and embedded resources. This version keeps the text and ignores the rest.
tool_name on the tool message tells the model which call the result belongs to, which keeps parallel calls from getting mixed up.
Run It Against Real Files
Save the file as bridge_client.py, change NOTES_DIR to a real folder, and run python bridge_client.py. A healthy run makes a directory listing or search call first, then one or two file reads, then answers in plain text. To watch the model choose tools, add print(call.function.name, call.function.arguments) at the top of the inner loop.
💡 Windows tip: if Python cannot launch npx, use npx.cmd as the command. The -y flag lets npx install the filesystem server on first run without asking.
Before you point the bridge at anything sensitive, decide what the model may do. The filesystem server only reaches the folders you pass on its command line, so give it a notes folder rather than your home directory. For tools that write, delete, or send data, add a confirmation step inside the loop: print the call and ask for a yes before session.call_tool runs. Thirty seconds of friction beats an 8B model deciding that a cleanup is a good idea.
Expose Ollama as an MCP Server
Now flip the direction. Instead of a local model using other people's tools, you publish the local model as a tool. Any MCP client can then call it for drafting, summarizing, or classifying text that should never leave your machine.
A Small Server in Python
The official Python SDK includes FastMCP, which builds the tool schema from your type hints and the description from your docstring:
import sys
import ollama
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("ollama-bridge")
client = ollama.AsyncClient()
DEFAULT_MODEL = "llama3.1:8b"
@mcp.tool()
async def list_local_models() -> list[str]:
"""List the models installed in the local Ollama instance."""
listed = await client.list()
return [m.model for m in listed.models]
@mcp.tool()
async def ask_local_model(
prompt: str, model: str = DEFAULT_MODEL, temperature: float = 0.2
) -> str:
"""Send a prompt to a local Ollama model and return its reply.
Use it for private text or cheap drafts that should stay on this machine."""
print(f"ask_local_model: {model}", file=sys.stderr)
response = await client.chat(
model=model,
messages=[{"role": "user", "content": prompt}],
options={"temperature": temperature},
)
return response.message.content
if __name__ == "__main__":
mcp.run(transport="stdio")
Write the docstring for the caller, not for yourself. The client model reads it to decide when your tool is worth calling, so "use it for private text" does real work.
Connect It to Claude Desktop
Open claude_desktop_config.json. On Windows it lives in %APPDATA%\Claude, and on macOS in ~/Library/Application Support/Claude. Add the server under mcpServers:
Point command at the virtual environment's Python, not the system one. Claude Desktop does not activate your environment, and a system interpreter without the mcp package fails silently. Quit and reopen the app, and the two tools appear in the tools menu.
Connect It to Claude Code
Claude Code registers servers from the terminal:
claude mcp add ollama-bridge -- /path/to/.venv/bin/python /path/to/ollama_bridge.py
claude mcp list
Everything after the double dash is the launch command. Once claude mcp list shows the server as connected, ask Claude Code to "summarize this log with the local model" and watch it call ask_local_model.
Fix the Problems You Will Hit
Three problems account for most failed first attempts. All three have short fixes.
Stdout Breaks Stdio Servers
A stdio MCP server sends JSON-RPC messages through stdout. One stray print() injects text into that stream, and the client drops the connection or reports a parse error. The symptom is a server that connects, then disconnects within a second.
The fix is a habit: log to stderr (print(..., file=sys.stderr)) or to a file, and never write anything else to stdout. That includes progress bars and warnings printed by imported libraries.
Small Models Ignore Tools
An 8B model given 25 tools often answers from memory or calls the wrong one. Four changes help, roughly in order of impact:
Send fewer tools. Filter the list down to the three to six that fit the question.
Rewrite descriptions. "Read the contents of one file by absolute path" beats "File reader".
Lower the temperature to 0.1 or 0.2 for tool selection.
Step up a size. Moving from 3B to 8B fixes more tool-calling failures than any prompt trick.
Context Windows Fill Fast
Tool definitions and tool results both consume context, and a single large file read can push the question out of the window. Ollama's default context is small, so raise it explicitly and trim results before they reach the model:
response = await client.chat(
model=MODEL,
messages=messages,
tools=tools,
options={"num_ctx": 8192},
)
text = text[:4000] # trim large tool results before appending them
Higher num_ctx values use more memory for the attention cache, so raise them in steps. Also set OLLAMA_KEEP_ALIVE to a longer value such as 30m. Ollama unloads idle models after five minutes by default, and every reload adds seconds to the first reply.
Local Models or Hosted Models
A bridge does not force a choice. It lets you route each job to the cheapest place that can do it well.
Local models win when:
the text is private, such as contracts, health notes, or source code under NDA
the job repeats thousands of times, so per-token pricing adds up
you work offline or on a network you do not control
Hosted models win when:
your GPU has under 8 GB of memory
the task needs a model above 30B parameters
you need an answer in seconds on a cold start
In practice a hybrid works best. Let the local model handle first drafts, classification, and anything that touches private files, then escalate the hard 10 percent to a larger hosted model. Because both sit behind the same MCP interface, your assistant can choose between ask_local_model and a hosted alternative with nothing more than a clearer tool description.
Hosted options on PicassoIA make useful yardsticks. Run the same prompt on a local 8B model and on one of these, and you see exactly what the extra size buys:
If you want to test gpt-oss:20b before you download 14 GB, run the same open-weight model in the browser:
Open the GPT OSS 20B page in the Large Language Models collection.
Type your prompt in the Prompt field. A good test is the docstring you plan to give ask_local_model, with the question "Would you know when to call this tool?"
Leave Temperature at its default of 0.1 for precise, repeatable output. Raise it for brainstorming.
Keep Max Tokens at 2048 for long answers, or lower it for short replies.
Adjust Top P, Presence Penalty, and Frequency Penalty only if the output loops or feels repetitive.
Run it, tweak, and run again. The model page lists unlimited generations, so iterating costs nothing.
💡 Compare the hosted answer with your local run. If they match closely, the local setup is doing its job and you can stop paying for the hosted call.
Pair the Bridge With Image Generation
Once ask_local_model exists, it can do more than summarize logs. A local model is a cheap, private way to draft image prompts, and that is where a bridge meets visual work.
Add a third tool that takes a topic and returns a 60-word photographic prompt: subject, setting, light direction, lens, and film look. Your assistant calls it, and you paste the result into a text-to-image model. P-Image is a quick option for drafts. FLUX 2 Pro suits a second pass when a prompt needs more detail.
A simple routine works well:
Ask the local model for three prompt variations on one subject.
The same pattern works for thumbnails, product shots, and blog headers. The bridge keeps the drafting step free and private, and the hosted models handle the heavy rendering.
Make Your Own Images on PicassoIA
You now have a working bridge in both directions: a local model that can use tools, and a local model that other apps can call. The next step is to put it to work on something you can see.
Take the prompt-drafting idea and try it today. Ask your local model for a photographic prompt, open Picasso IA, and generate your first image with P-Image or FLUX 2 Pro. Change the lens, the light direction, or the setting, and run it again. When a frame works, turn it into a short video. Browse every available model in the full PicassoIA catalog and start experimenting with Picasso IA.