Large Language ModelsGenerate imagesGenerate videos

LM Studio MCP: mcp.json Example and Best Servers for Web Search

A working LM Studio mcp.json you can paste today, explained entry by entry, plus a side by side look at the best MCP servers for web search: Brave Search, Tavily, Exa, DuckDuckGo and Fetch. You get exact package names, tool names, token settings and the safety checks to run first.

LM Studio MCP: mcp.json Example and Best Servers for Web Search
Cristian Da Conceicao
Founder of Picasso IA

Your local model is fast, private and sharp, and it still believes it is living in the month its training data stopped. Ask LM Studio's chat window about last week's release and you get a confident guess. The fix is one JSON file. Add an MCP server for web search to mcp.json, and the same model starts calling a search tool, reading real pages, and answering with sources instead of memory.

This article gives you a working LM Studio MCP mcp.json example you can paste today, then compares the five web search servers people reach for most: Brave Search, Tavily, Exa, DuckDuckGo and the official Fetch server. Every config below comes from each project's own documentation, so the package names, environment variables and tool names are the real ones.

A person's hands resting on a laptop at a walnut desk in late afternoon light

The Stale Knowledge Problem

A language model is a snapshot. Whatever it knew on the day training ended is all it will ever know, no matter how long you leave it running on your desk. That is fine for rewriting an email or explaining a regex. It is painful for anything that moves: library versions, prices, release notes, the error message from last month's update.

Pasting web pages into the chat by hand works, but it burns your context window and your patience. A search tool lets the model decide when it needs fresh facts, fetch them, and fold them into the answer.

What MCP Does in Practice

The Model Context Protocol (MCP) is a standard way for an app to connect a model to outside tools. LM Studio acts as the host: it starts or connects to each server listed in your config, asks what tools they offer, and passes those tool descriptions to the model. When the model decides to call one, LM Studio runs it and feeds the result back into the chat.

💡 You write no code for this. A search server is a small program someone else maintains. Your only job is telling LM Studio how to launch it.

MCP support arrived in LM Studio 0.3.17, and the app follows Cursor's mcp.json notation. That matters because most server READMEs show a Cursor or Claude Desktop snippet, and those snippets paste in with little or no change.

Where mcp.json Lives

Inside the app, open the Program tab in the right-hand sidebar, click Install, then choose Edit mcp.json. The built-in editor is the safest route because you see mistakes as you type.

Top-down view of an open laptop with a code editor on a light oak desk

If you would rather edit the file directly, these are the paths:

SystemPath to the file
macOS~/.lmstudio/mcp.json
Linux~/.lmstudio/mcp.json
Windows%USERPROFILE%/.lmstudio/mcp.json

After you save, check the Program tab: each server in the file should appear there with its tools. Some MCP authors also publish an Add to LM Studio button that writes the entry for you, which is handy when a server has a long config.

A Working mcp.json Example

Every entry sits inside one top-level object called mcpServers. Each server gets a name you choose, followed by either a url (a remote server) or a command (a local process).

Remote Server Entry

The official LM Studio docs use the Hugging Face MCP server as the remote example:

{
  "mcpServers": {
    "hf-mcp-server": {
      "url": "https://huggingface.co/mcp",
      "headers": {
        "Authorization": "Bearer <YOUR_HF_TOKEN>"
      }
    }
  }
}

Three details matter here. url points at the hosted server. headers carries your credential. And the placeholder in angle brackets has to be replaced with a real token before the server will connect.

Local Server With npx

Local servers use command, args and env. This is Brave's own snippet for the Brave Search MCP server:

{
  "mcpServers": {
    "brave-search": {
      "command": "npx",
      "args": ["-y", "@brave/brave-search-mcp-server", "--transport", "stdio"],
      "env": {
        "BRAVE_API_KEY": "YOUR_BRAVE_TOKEN"
      }
    }
  }
}

npx ships with Node.js, so install that first. The -y flag skips the confirmation prompt that would otherwise hang a background process, and --transport stdio tells the server to talk to LM Studio through standard input and output. The env block hands your token to the server as an environment variable, which is how nearly every search server expects to receive it.

Several Servers in One File

A realistic file mixes sources: Brave for broad search, Fetch for reading a page in full, DuckDuckGo as a fallback that needs no account.

{
  "mcpServers": {
    "brave-search": {
      "command": "npx",
      "args": ["-y", "@brave/brave-search-mcp-server", "--transport", "stdio"],
      "env": {
        "BRAVE_API_KEY": "YOUR_BRAVE_TOKEN"
      }
    },
    "fetch": {
      "command": "uvx",
      "args": ["mcp-server-fetch"]
    },
    "ddg-search": {
      "command": "uvx",
      "args": ["duckduckgo-mcp-server"]
    }
  }
}

uvx comes from the uv Python tool, so the two Python servers need uv installed, just as the Node server needs Node.

💡 Two commas cause more broken configs than anything else: a missing comma between servers and a trailing comma after the last one. JSON allows neither, and it has no comments either.

Over-the-shoulder view of a developer editing a configuration file on a large monitor

Search servers do three different jobs: find links, return clean page text, and read one specific URL. The best setups pair a finder with a reader instead of asking one server to do everything.

A compact mini PC and a router on a wooden shelf in a home office corner

Brave Search MCP Server

Brave is the default pick for most people. The package is @brave/brave-search-mcp-server, the token goes in BRAVE_API_KEY, and the tool list is broad: brave_web_search, brave_local_search, brave_video_search, brave_image_search, brave_news_search, brave_summarizer, brave_place_search and brave_llm_context.

Brave returns raw ranked results (links and snippets) and lets your local model do the reading. That keeps responses small, which matters when your context window is limited. The repository notes that Pro plans add extras such as additional snippets and full local search.

Tavily MCP Server

Tavily leans toward extraction. Its tools are tavily-search, tavily-extract, tavily-map and tavily-crawl, so the model can search, pull the text of a page, map a site or crawl it. The local entry looks like this:

{
  "mcpServers": {
    "tavily-mcp": {
      "command": "npx",
      "args": ["-y", "tavily-mcp@latest"],
      "env": {
        "TAVILY_API_KEY": "your-tavily-token"
      }
    }
  }
}

Tavily also runs a remote server at https://mcp.tavily.com/mcp/?tavilyApiKey=<your-token>. The local version above is the better habit, because it keeps your token out of a URL.

Exa MCP Server

Exa is the shortest setup of the five, since it is hosted. There is nothing to install, so no Node and no Python:

{
  "mcpServers": {
    "exa": {
      "url": "https://mcp.exa.ai/mcp"
    }
  }
}

The default tools are web_search_exa, which returns search results with clean content, and web_fetch_exa, which reads a page as markdown. Exa accepts a token as a URL parameter or in an Authorization header. The header is the tidier choice.

DuckDuckGo MCP Server

No account, no token. Run uvx duckduckgo-mcp-server and you have search, fetch_content and expand_link. It is the quickest way to prove your whole setup works before you sign up for anything.

The trade-off is rate limits. The defaults are 30 searches per minute and 20 page fetches per minute, adjustable with the DDG_SEARCH_RPM and DDG_FETCH_RPM environment variables. On an HTTP 429 response the server honors Retry-After and retries once.

Fetch MCP Server

Fetch comes from the official MCP servers repository and does one thing: it downloads a URL and converts the HTML to markdown. Its fetch tool takes a url, plus optional max_length (default 5000 characters), start_index and raw.

That start_index parameter is useful. When a page is longer than one response, the model can ask for the next slice instead of losing the end of the article. Fetch honors robots.txt for model-initiated requests by default, and its documentation warns that it can reach local and internal IP addresses, which is a real risk on a work machine.

A hand placing a sticky note onto index cards pinned to a cork board

Which Server Fits Your Setup

ServerLaunchTokenToolsBest for
Brave SearchnpxYes, BRAVE_API_KEYbrave_web_search plus seven moreGeneral web search
Tavilynpx or remoteYes, TAVILY_API_KEYtavily-search, tavily-extract, tavily-map, tavily-crawlResearch that needs page text
ExaRemote URLURL parameter or headerweb_search_exa, web_fetch_exaZero install setup
DuckDuckGouvxNonesearch, fetch_content, expand_linkFree testing
FetchuvxNonefetchReading a known URL

Three starting points suit most people:

  • One server only: pick Brave Search, or Tavily if you care more about page text than link lists.
  • No accounts at all: pair DuckDuckGo with Fetch.
  • Least installation: use the remote Exa entry and skip Node and Python entirely.

Pick a Model That Calls Tools

Not every model can call tools. A model has to be trained to emit a structured tool call, and LM Studio marks tool-capable models in its model list with a hammer badge. Pick one without it and the model will talk about searching but never actually search.

Side view of a graphics card installed inside an open PC case

Two settings help just as much as the model choice:

  • Context length: search results are long. Raise the context length when you load the model, or the results get cut off before the model reads them.
  • Model size: very small models tend to fumble multi-step tool use, searching once and answering from a half-read snippet.

How to Use GPT OSS on PicassoIA

Before you download a multi-gigabyte model, rehearse your prompts on a hosted one. GPT OSS 20B is an open-weight, 20-billion-parameter language model, and its model page offers unlimited generations with no credit caps.

One honest limit: this hosted model runs in the browser and does not connect to your local MCP servers. Use it to polish the system prompt and to see how a model summarizes pasted search results, then carry the winning prompt back to LM Studio.

  1. Open the model page. Go to the GPT OSS 20B page.
  2. Write the prompt. Try: "You are a research assistant. Using only the search results below, answer in five bullets and cite the source URL for each." Paste real search results underneath.
  3. Keep temperature low. The default is 0.1, which suits factual summaries. Raise it only when you want brainstorming.
  4. Check the limits. Max tokens defaults to 2048, top p to 1, and both penalties to 0. Add a small presence or frequency penalty if the output starts looping.
  5. Generate and compare. Run it, edit one line, run it again. Save the version that follows your format without drifting.
  6. Carry it over. Paste the winning text into your LM Studio system prompt.
SettingDefaultUse it for
Temperature0.1Precise, repeatable summaries
Top P1Output diversity
Max Tokens2048Length of the answer
Presence Penalty0Pushing toward new topics
Frequency Penalty0Cutting repeated words

A woman typing on a laptop at a sunlit cafe window table

Other hosted models worth running through the same test: GPT OSS 120B, Qwen3.7-Plus, Granite 4.1 8B and Kimi K2.6. If one of them follows your format better, you know what to look for in a local download.

Safety Rules Before You Install

A worn brass padlock hanging on a weathered wooden gate

LM Studio's own docs are blunt about this: some MCP servers can run arbitrary code, read your local files and use your network connection. Never install one from a source you do not trust. A search server is only a program running on your machine with your permissions.

Read Every Tool Call

When a model calls a tool, LM Studio shows a confirmation dialog. You can review and edit the arguments before anything runs, then allow the tool once or permanently. Stay on "allow once" for any server that touches files or the network until you trust it. Permanent permissions are managed in App Settings > Tools & Integrations.

Protect Your Tokens

Your mcp.json stores tokens in plain text, so treat the file like a password list:

  • Never commit it to Git or paste it into a forum post or screenshot.
  • Prefer env and headers over tokens inside URLs, which tend to land in logs.
  • Rotate a token immediately if it leaks.
  • Think twice before running Fetch on a machine that can reach internal admin pages.

Fixing Common Setup Problems

Hands plugging a blue ethernet cable into the back of a white router

Most failures fall into five buckets:

  1. Broken JSON. Check commas, quotes and brackets. On Windows, any file path inside the JSON needs doubled backslashes, or use forward slashes.
  2. Missing runtime. npx needs Node.js and uvx needs uv. Run node --version or uv --version in a terminal to confirm.
  3. The model ignores your tools. The model may lack tool support, or the prompt is too vague. Add one line to the system prompt: "Use the search tool for anything newer than your training data."
  4. Empty or throttled results. DuckDuckGo caps searches at 30 per minute by default, and paid APIs have their own quotas. Ask the model to search once, then read.
  5. Pages cut short. Fetch returns 5000 characters by default. Ask for the next slice with start_index, or raise max_length.

💡 Test every new server alone first. Put one entry in the file, ask a question that needs fresh facts, and watch the tool call. Add the next server only after that works.

Make Your Own Images on PicassoIA

Once your local setup is searching the web, you will want visuals for the write-ups, tutorials and notes that come out of it: a header for your blog post, a diagram background, a thumbnail for a video walkthrough. That is where PicassoIA fits in.

Try one experiment: write a single photographic prompt, such as "a developer's desk at dawn with a laptop and a notebook of hand-drawn boxes", and run it through three text-to-image models. Flux Krea Dev leans toward images that skip the usual AI look, Seedream 4.5 produces sharp, high-resolution results, and GPT Image 2 follows long, detailed prompts closely. Put the three results side by side and keep the one that fits.

PicassoIA also offers a developer API and MCP connections for its image and video models, so the same workflow can plug into your own tools later. Open PicassoIA, pick a model, and make your first image today.

Share this article