Large Language ModelsGenerate imagesGenerate videos

Ollama MCP Server: How to Connect Local Models With a Bridge

Build an Ollama MCP server bridge in both directions: a Python client that lets local models call MCP tools, and a server that exposes Ollama to Claude Desktop and Claude Code. Includes hardware sizing, working code, config files, and fixes for the three bugs that break most first attempts.

Ollama MCP Server: How to Connect Local Models With a Bridge
Cristian Da Conceicao
Founder of Picasso IA

Your laptop can already run a capable language model, but most of your tools have no idea it exists. An Ollama MCP server setup fixes that. A small bridge sits between Ollama, which serves models at localhost:11434, and the Model Context Protocol (MCP), the standard AI apps use to find and call tools. Build it once and a local model can read your files, query a database, or search your notes. Flip the direction and Claude Desktop or Claude Code can hand cheap, private jobs to the GPU under your desk.

This article builds both directions with working Python code: a bridge client that lets Ollama models use any MCP server, and a bridge server that exposes Ollama as a set of tools to any MCP client. You also get memory numbers, config files, and the three bugs that eat most of the debugging time.

What an Ollama MCP Bridge Does

Ollama and MCP solve different halves of one problem. Ollama downloads open-weight models and serves them over a plain HTTP API. MCP, which Anthropic published in late 2024, standardizes how an AI application talks to tools, files, and data through small programs called servers. Ollama's API speaks chat messages and tool definitions. It does not speak MCP, and MCP clients have no idea where your models live. The bridge translates between the two.

Low-angle view of a mossy stone arch bridge crossing a misty river at dawn

Two Ways to Build a Bridge

"Bridge" means two different programs depending on who needs what, so pick the direction before you write any code.

DirectionMCP clientMCP serverTypical use
Ollama uses toolsYour bridge script, wrapping an Ollama modelFilesystem, database, or search serverA local assistant that reads your notes
Ollama as a toolClaude Desktop, Claude Code, CursorYour bridge script, wrapping OllamaOffloading private or cheap tasks to a local model

The first direction gives a local model hands. The second gives a cloud assistant a local colleague. Most people end up building both, because the second one takes about 30 lines.

A quick note on transport. The examples here use stdio, where the client starts the server as a child process and talks to it over standard input and output. That is the simplest option for a personal machine. If you want one bridge shared by several apps or several computers on your network, run it over the Streamable HTTP transport instead and keep it behind your own firewall. The tool code stays the same, and only the last line of the server changes.

How a Tool Call Travels

Whichever direction you pick, a tool call follows the same loop:

  1. The bridge connects to an MCP server and asks for its tool list with tools/list.
  2. It rewrites each tool schema into the format Ollama expects.
  3. It sends the user's question and those tool definitions to a local model.
  4. The model answers with a tool_calls entry instead of text.
  5. The bridge runs that call on the MCP server with tools/call.
  6. The result goes back to the model as a tool message.
  7. Steps 3 to 6 repeat until the model replies with plain text.

💡 The model never touches your disk. It only requests an action. Your bridge decides whether to run it, which makes the bridge the right place for allow lists, confirmation prompts, and logging.

What You Need First

Hardware That Works

Extreme close-up of a graphics card with two black fans installed inside an open PC case

Model weights have to fit in GPU memory (or unified memory on a Mac) with room left for context. These figures are approximate, for 4-bit quantized builds:

Model sizeMemory for weightsComfortable setup
3B to 4B2 to 3 GBAny recent laptop
7B to 8B5 to 6 GB8 GB GPU or 16 GB unified memory
14B9 to 10 GB12 GB GPU
20B13 to 15 GB16 GB GPU or 24 GB unified memory
32B19 to 21 GB24 GB GPU

Models that spill into system RAM still run, but token speed drops hard. A bridge makes several model calls per question, so speed matters more than size. An 8B model that fits entirely on the GPU usually beats a 32B model that does not.

Models That Support Tool Calls

Not every model can ask for a tool. Ollama checks the model's chat template, and a model without tool support returns an error when you pass tools. Filter the Ollama library by the tools tag, or start from this shortlist:

Ollama tagSize on diskWhy use it
llama3.1:8babout 4.9 GBReliable default for first tests
qwen3:8babout 5.2 GBStrong at multi-step tool use, slower when thinking is on
mistral-nemoabout 7.1 GBLong context, handles many tools
gpt-oss:20babout 14 GBThe open-weight GPT OSS 20B, built with tool use in mind

Install the Pieces

Activate a virtual environment first (source .venv/bin/activate on macOS and Linux, .venv\Scripts\activate on Windows), then run:

# 1. Pull a tool-capable model and confirm Ollama is serving
ollama pull llama3.1:8b
curl http://localhost:11434/api/tags

# 2. Install the two Python packages the bridge needs
pip install ollama mcp

# 3. Check Node, because the example MCP server runs through npx
node --version

If curl returns a JSON list of your models, Ollama is ready. If the connection is refused, start it with ollama serve.

Build the Bridge Client in Python

The client is one file. It launches an MCP server as a child process over the stdio transport, reads its tools, and runs the loop from the previous section. The example server is the official filesystem server, pointed at a notes folder.

Convert MCP Tools to Ollama Format

An MCP tool carries a name, a description, and an inputSchema written in JSON Schema. Ollama's function-calling format wants the same three pieces, wrapped in a function object. The conversion is mostly a rename:

def to_ollama_tool(tool):
    return {
        "type": "function",
        "function": {
            "name": tool.name,
            "description": tool.description or "",
            "parameters": tool.inputSchema,
        },
    }

Most schemas pass through untouched. If a server ships exotic constructs such as anyOf or $ref, flatten them before sending. Small models handle flat schemas far better.

Write the Tool Call Loop

import asyncio

import ollama
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client

MODEL = "llama3.1:8b"
NOTES_DIR = "/home/me/notes"
MAX_TURNS = 8

server_params = StdioServerParameters(
    command="npx",
    args=["-y", "@modelcontextprotocol/server-filesystem", NOTES_DIR],
)


def to_ollama_tool(tool):
    return {
        "type": "function",
        "function": {
            "name": tool.name,
            "description": tool.description or "",
            "parameters": tool.inputSchema,
        },
    }


async def ask(question: str) -> str:
    async with stdio_client(server_params) as (read, write):
        async with ClientSession(read, write) as session:
            await session.initialize()
            listed = await session.list_tools()
            tools = [to_ollama_tool(t) for t in listed.tools]

            client = ollama.AsyncClient()
            messages = [{"role": "user", "content": question}]

            for _ in range(MAX_TURNS):
                response = await client.chat(model=MODEL, messages=messages, tools=tools)
                message = response.message
                messages.append(message)

                if not message.tool_calls:
                    return message.content

                for call in message.tool_calls:
                    result = await session.call_tool(
                        call.function.name, dict(call.function.arguments)
                    )
                    text = "\n".join(
                        block.text for block in result.content if block.type == "text"
                    )
                    messages.append(
                        {"role": "tool", "tool_name": call.function.name, "content": text}
                    )

            return "Stopped: the model kept calling tools."


if __name__ == "__main__":
    print(asyncio.run(ask("Which markdown files in my notes folder mention invoices?")))

Four details deserve attention:

  • MAX_TURNS is a safety rail. A confused model can call the same tool forever, and a counter turns that into a clear failure.
  • dict(call.function.arguments) matters because Ollama returns arguments as a mapping that the MCP session expects as a plain dictionary.
  • Text blocks only. MCP results can include images and embedded resources. This version keeps the text and ignores the rest.
  • tool_name on the tool message tells the model which call the result belongs to, which keeps parallel calls from getting mixed up.

Run It Against Real Files

Over-the-shoulder view of a developer typing on a laptop at a sunlit cafe table

Save the file as bridge_client.py, change NOTES_DIR to a real folder, and run python bridge_client.py. A healthy run makes a directory listing or search call first, then one or two file reads, then answers in plain text. To watch the model choose tools, add print(call.function.name, call.function.arguments) at the top of the inner loop.

💡 Windows tip: if Python cannot launch npx, use npx.cmd as the command. The -y flag lets npx install the filesystem server on first run without asking.

Before you point the bridge at anything sensitive, decide what the model may do. The filesystem server only reaches the folders you pass on its command line, so give it a notes folder rather than your home directory. For tools that write, delete, or send data, add a confirmation step inside the loop: print the call and ask for a yes before session.call_tool runs. Thirty seconds of friction beats an 8B model deciding that a cleanup is a good idea.

Expose Ollama as an MCP Server

Now flip the direction. Instead of a local model using other people's tools, you publish the local model as a tool. Any MCP client can then call it for drafting, summarizing, or classifying text that should never leave your machine.

Top-down view of a small silver mini PC, router, and notebook sketches on an oak desk

A Small Server in Python

The official Python SDK includes FastMCP, which builds the tool schema from your type hints and the description from your docstring:

import sys

import ollama
from mcp.server.fastmcp import FastMCP

mcp = FastMCP("ollama-bridge")
client = ollama.AsyncClient()
DEFAULT_MODEL = "llama3.1:8b"


@mcp.tool()
async def list_local_models() -> list[str]:
    """List the models installed in the local Ollama instance."""
    listed = await client.list()
    return [m.model for m in listed.models]


@mcp.tool()
async def ask_local_model(
    prompt: str, model: str = DEFAULT_MODEL, temperature: float = 0.2
) -> str:
    """Send a prompt to a local Ollama model and return its reply.
    Use it for private text or cheap drafts that should stay on this machine."""
    print(f"ask_local_model: {model}", file=sys.stderr)
    response = await client.chat(
        model=model,
        messages=[{"role": "user", "content": prompt}],
        options={"temperature": temperature},
    )
    return response.message.content


if __name__ == "__main__":
    mcp.run(transport="stdio")

Write the docstring for the caller, not for yourself. The client model reads it to decide when your tool is worth calling, so "use it for private text" does real work.

Connect It to Claude Desktop

Open claude_desktop_config.json. On Windows it lives in %APPDATA%\Claude, and on macOS in ~/Library/Application Support/Claude. Add the server under mcpServers:

{
  "mcpServers": {
    "ollama-bridge": {
      "command": "C:\\tools\\ollama-bridge\\.venv\\Scripts\\python.exe",
      "args": ["C:\\tools\\ollama-bridge\\ollama_bridge.py"],
      "env": { "OLLAMA_HOST": "http://127.0.0.1:11434" }
    }
  }
}

Point command at the virtual environment's Python, not the system one. Claude Desktop does not activate your environment, and a system interpreter without the mcp package fails silently. Quit and reopen the app, and the two tools appear in the tools menu.

Connect It to Claude Code

Claude Code registers servers from the terminal:

claude mcp add ollama-bridge -- /path/to/.venv/bin/python /path/to/ollama_bridge.py
claude mcp list

Everything after the double dash is the launch command. Once claude mcp list shows the server as connected, ask Claude Code to "summarize this log with the local model" and watch it call ask_local_model.

Fix the Problems You Will Hit

Close-up of a technician's hand plugging a connector into a network patch panel full of colored cables

Three problems account for most failed first attempts. All three have short fixes.

Stdout Breaks Stdio Servers

A stdio MCP server sends JSON-RPC messages through stdout. One stray print() injects text into that stream, and the client drops the connection or reports a parse error. The symptom is a server that connects, then disconnects within a second.

The fix is a habit: log to stderr (print(..., file=sys.stderr)) or to a file, and never write anything else to stdout. That includes progress bars and warnings printed by imported libraries.

Small Models Ignore Tools

An 8B model given 25 tools often answers from memory or calls the wrong one. Four changes help, roughly in order of impact:

  • Send fewer tools. Filter the list down to the three to six that fit the question.
  • Rewrite descriptions. "Read the contents of one file by absolute path" beats "File reader".
  • Lower the temperature to 0.1 or 0.2 for tool selection.
  • Step up a size. Moving from 3B to 8B fixes more tool-calling failures than any prompt trick.

Context Windows Fill Fast

Tool definitions and tool results both consume context, and a single large file read can push the question out of the window. Ollama's default context is small, so raise it explicitly and trim results before they reach the model:

response = await client.chat(
    model=MODEL,
    messages=messages,
    tools=tools,
    options={"num_ctx": 8192},
)

text = text[:4000]  # trim large tool results before appending them

Higher num_ctx values use more memory for the attention cache, so raise them in steps. Also set OLLAMA_KEEP_ALIVE to a longer value such as 30m. Ollama unloads idle models after five minutes by default, and every reload adds seconds to the first reply.

Local Models or Hosted Models

Wide low-angle view down a long aisle of black server racks in a data center

A bridge does not force a choice. It lets you route each job to the cheapest place that can do it well.

Local models win when:

  • the text is private, such as contracts, health notes, or source code under NDA
  • the job repeats thousands of times, so per-token pricing adds up
  • you work offline or on a network you do not control

Hosted models win when:

  • your GPU has under 8 GB of memory
  • the task needs a model above 30B parameters
  • you need an answer in seconds on a cold start

In practice a hybrid works best. Let the local model handle first drafts, classification, and anything that touches private files, then escalate the hard 10 percent to a larger hosted model. Because both sit behind the same MCP interface, your assistant can choose between ask_local_model and a hosted alternative with nothing more than a clearer tool description.

Hosted options on PicassoIA make useful yardsticks. Run the same prompt on a local 8B model and on one of these, and you see exactly what the extra size buys:

ModelGood for
Llama 4 Scout InstructFast drafts and summaries
DeepSeek R1Step-by-step reasoning on hard questions
Qwen3.7-PlusText generation plus image input
Granite 4.1 8BChat and code at a small-model size

Use GPT OSS 20B on PicassoIA

If you want to test gpt-oss:20b before you download 14 GB, run the same open-weight model in the browser:

  1. Open the GPT OSS 20B page in the Large Language Models collection.
  2. Type your prompt in the Prompt field. A good test is the docstring you plan to give ask_local_model, with the question "Would you know when to call this tool?"
  3. Leave Temperature at its default of 0.1 for precise, repeatable output. Raise it for brainstorming.
  4. Keep Max Tokens at 2048 for long answers, or lower it for short replies.
  5. Adjust Top P, Presence Penalty, and Frequency Penalty only if the output loops or feels repetitive.
  6. Run it, tweak, and run again. The model page lists unlimited generations, so iterating costs nothing.

Close-up of a woman's hands typing on a slim laptop at a bright white desk

💡 Compare the hosted answer with your local run. If they match closely, the local setup is doing its job and you can stop paying for the hosted call.

Pair the Bridge With Image Generation

Once ask_local_model exists, it can do more than summarize logs. A local model is a cheap, private way to draft image prompts, and that is where a bridge meets visual work.

Add a third tool that takes a topic and returns a 60-word photographic prompt: subject, setting, light direction, lens, and film look. Your assistant calls it, and you paste the result into a text-to-image model. P-Image is a quick option for drafts. FLUX 2 Pro suits a second pass when a prompt needs more detail.

A designer standing before a wall of pinned photographic prints in a bright studio

A simple routine works well:

  1. Ask the local model for three prompt variations on one subject.
  2. Generate all three and keep the best frame.
  3. Animate the winner with an image-to-video model from the PicassoIA model catalog.

The same pattern works for thumbnails, product shots, and blog headers. The bridge keeps the drafting step free and private, and the hosted models handle the heavy rendering.

Make Your Own Images on PicassoIA

A smiling young man leaning back with a laptop in a golden hour home studio

You now have a working bridge in both directions: a local model that can use tools, and a local model that other apps can call. The next step is to put it to work on something you can see.

Take the prompt-drafting idea and try it today. Ask your local model for a photographic prompt, open Picasso IA, and generate your first image with P-Image or FLUX 2 Pro. Change the lens, the light direction, or the setting, and run it again. When a frame works, turn it into a short video. Browse every available model in the full PicassoIA catalog and start experimenting with Picasso IA.

Share this article