Large Language ModelsGenerate imagesGenerate videos

GitHub MCP Server: Setup, Token Usage and Rate Limits Without the Surprises

Connect the GitHub MCP server to VS Code, Claude Desktop or Cursor with OAuth or a narrow token, trim its tool definitions with toolsets and read-only mode, and handle GitHub's 5,000 per hour and 80 per minute limits without 403 or 429 surprises.

GitHub MCP Server: Setup, Token Usage and Rate Limits Without the Surprises
Cristian Da Conceicao
Founder of Picasso IA

Plug the GitHub MCP server into your editor and an AI agent can read issues, review pull requests and open branches on your behalf. It also loads a long list of tool definitions into every conversation, and it spends the request budget of whatever token you hand it. Three things decide whether the setup feels smooth or painful: how you connect, how many context tokens the server eats, and which rate limits you hit first.

Every figure below comes from GitHub's own documentation or from community measurements published in 2026, and each one is labeled as such. Token counts in particular shift between server versions, so treat them as ranges instead of promises. The same goes for any number you read about MCP tooling: check the version it was measured on before you build a budget around it.

💡 Short version: use the remote server, give it a narrow token, switch on only the toolsets you use, and back off for at least a minute when GitHub answers 403 or 429.

Remote or Local Server?

Both options expose the same GitHub tools. What changes is who runs the process and how you sign in. Pick the wrong one and you end up maintaining Docker on five laptops, or waiting on GitHub for a feature you needed yesterday.

Remote Server Basics

GitHub hosts the remote server at https://api.githubcopilot.com/mcp/. Your client points at that URL and you sign in through the browser with OAuth, or you send a personal access token in an Authorization header. There is nothing to install and nothing to update, because new tools arrive on GitHub's schedule, not yours.

The same host offers an insiders variant at https://api.githubcopilot.com/mcp/insiders for early features. Try it on a test machine before it touches your daily setup.

Local Docker Option

The local server ships as the image ghcr.io/github/github-mcp-server. Your client launches it with docker run -i --rm and talks to it over stdio. You choose the version, you control every environment variable, and you can point it at GitHub Enterprise Server with GITHUB_HOST. The price is Docker on every machine and a token sitting in a config file.

Narrow server room aisle between black racks with tidy network cables

Remote serverLocal Docker server
InstallNoneDocker image
Sign inOAuth, or a token in a headerToken in GITHUB_PERSONAL_ACCESS_TOKEN, or OAuth with a callback port
UpdatesHandled by GitHubYou pull the image
GitHub Enterprise ServerNot supportedSupported through GITHUB_HOST
Best forMost individual developersPinned versions and Enterprise Server

Setup in Five Minutes

Every snippet below comes from the project's README. Paste one, restart the client, and ask the agent to list your open pull requests as a smoke test. If the list comes back, the connection works and everything after this point is tuning.

Developer hands typing on a laptop with a hardware security token plugged in

VS Code With OAuth

The shortest path. No token to create and no secret to store.

{
  "servers": {
    "github": {
      "type": "http",
      "url": "https://api.githubcopilot.com/mcp/"
    }
  }
}

VS Code opens a browser tab, you approve access, and the credential stays in memory instead of landing in a file.

VS Code With a Token

Use this route when your client does not support the OAuth flow, or when you want a token limited to specific scopes.

{
  "servers": {
    "github": {
      "type": "http",
      "url": "https://api.githubcopilot.com/mcp/",
      "headers": {
        "Authorization": "Bearer ${input:github_mcp_pat}"
      }
    }
  },
  "inputs": [
    {
      "type": "promptString",
      "id": "github_mcp_pat",
      "description": "GitHub Personal Access Token",
      "password": true
    }
  ]
}

The password: true flag masks the value when VS Code asks for it, so the token never gets written into the file you might commit.

Claude Desktop Config

Claude Desktop starts the local Docker server and hands it the token through an environment variable:

{
  "mcpServers": {
    "github": {
      "command": "docker",
      "args": ["run", "-i", "--rm", "-e", "GITHUB_PERSONAL_ACCESS_TOKEN",
               "ghcr.io/github/github-mcp-server"],
      "env": {"GITHUB_PERSONAL_ACCESS_TOKEN": "your_token_here"}
    }
  }
}

Swap your_token_here for a real token and keep this file out of version control. Cursor accepts a structure similar to the VS Code examples.

💡 If a token and OAuth are both configured, the token wins. The project docs say GITHUB_PERSONAL_ACCESS_TOKEN takes precedence over OAuth.

Personal Access Token Scopes

Brass padlock resting on a closed laptop beside a notebook and fountain pen

A token is the one part of this setup that can hurt you. An agent holding a broad token can do anything that token allows, including mistakes you would never make by hand.

Pick the Smallest Scopes

The README recommends three scopes:

ScopeWhat it allows
repoRepository operations
read:packagesDocker image access
read:orgOrganization team access

Start with repo and add the others only when a tool fails with a permission error. If you only work in a handful of repositories, a fine-grained token restricted to those repositories is tighter still. Use a separate token for each project, so revoking one does not break the rest.

Keep Tokens Out of Git

Three habits prevent most leaks:

  • Put the token in an environment variable or a prompt input, never in a committed config file.
  • Restrict permissions on any local config with chmod 600 ~/.your-app/config.json.
  • Rotate tokens on a schedule, and revoke one the moment it shows up in a diff.

The Real Cost of Tool Definitions

Two stacks of printed pages, one tall and one thin, beside a steel ruler

An MCP client loads the definitions of every connected server into the model's context so the model knows what it can call. That payload counts against your context window on every request, whether the agent uses a tool or not. Community estimates put a single definition at roughly 300 to 600 tokens once you add the name, description and parameter schema.

GitHub's server has a lot of tools, which is why it shows up in every conversation about context bloat.

What the Numbers Say

These are community measurements published on dev.to in 2026, not GitHub figures:

MeasurementTokensToolsSource
Full tool surfaceabout 55,00093Piotr Hajdas
Full tool surface, lower countabout 42,000not statedThe Daily Agent
Default toolsets onlyabout 4,20026Ken Imoto

On a 200,000-token window, the 55,000 figure is more than a quarter of the space gone before you type a word. The default toolsets come to roughly 2%. The counts differ because people measured different server versions and toolset configurations.

Your client matters too. According to a 2026 write-up, Claude Code defers MCP tool schemas behind a tool-search step by default, and one measurement reported a 46.9% reduction from it. Cursor, Windsurf and Gemini CLI load definitions upfront, so they pay the full price.

Cut It With Toolsets

Toolsets switch whole groups of tools on or off. With no setting, the server enables context, issues, pull_requests, repos and users. The rest of the list is actions, code_quality, code_security, copilot, dependabot, discussions, gists, git, governance, labels, notifications, orgs, projects, secret_protection, security_advisories and stargazers. The special values all and default do what they say.

For the local server, set GITHUB_TOOLSETS:

docker run -e GITHUB_PERSONAL_ACCESS_TOKEN=<token> \
  -e GITHUB_TOOLSETS="repos,issues,pull_requests" \
  ghcr.io/github/github-mcp-server

For the remote server, send the X-MCP-Toolsets header with the same comma-separated list:

{
  "servers": {
    "github": {
      "type": "http",
      "url": "https://api.githubcopilot.com/mcp/",
      "headers": {
        "Authorization": "Bearer ${input:github_mcp_pat}",
        "X-MCP-Toolsets": "repos,issues,pull_requests"
      }
    }
  }
}

Two more controls fine-tune the result. GITHUB_TOOLS (or the X-MCP-Tools header) adds individual tools on top of your toolsets, for example get_gist without enabling all of gists. X-MCP-Exclude-Tools removes tools, and the docs state that excluded tools take precedence over toolsets and individual tools.

Woodworker's pegboard wall with a few chosen tools laid out on the bench

💡 Resist all. Going from about 4,200 tokens of definitions to about 55,000 buys tools you will probably never call, and you pay for them on every request.

Read-Only and Lockdown Modes

Museum visitor reading an old document through a glass display case

Two switches shrink the blast radius without touching your toolsets:

ModeLocal serverRemote serverEffect
Read-only--read-only or GITHUB_READ_ONLYX-MCP-Readonly: true, or the path /mcp/x/all/readonlyDisables every write tool, even ones you requested
Lockdown--lockdown-mode or GITHUB_LOCKDOWN_MODEX-MCP-Lockdown headerSurfaces only public-repository content from users with push access

Read-only acts as a strict filter that outranks the rest of your configuration, and disabled tools also mean fewer definitions to load. Lockdown matters when an agent reads issues or comments written by strangers, because text from people without push access is exactly where hostile instructions hide. In HTTP mode, once an operator enables lockdown on the server, the X-MCP-Lockdown header can no longer switch it off for a single request.

Rate Limits You Will Actually Hit

Aerial view of a highway toll plaza with cars queuing at the booths

The server calls GitHub's API under your identity, so the limits that matter are the ones GitHub documents for its REST API. Two layers apply: a primary limit per hour and secondary limits per minute.

Primary Limits per Hour

CallerLimit
Unauthenticated60 requests per hour
Authenticated user (token, OAuth app or GitHub App)5,000 per hour
Apps owned or approved by Enterprise Cloud organizations15,000 per hour
GitHub App installations outside Enterprise5,000, scaling up to 12,500
GITHUB_TOKEN inside Actions1,000 per hour per repository (15,000 on Enterprise Cloud)

At 5,000 an hour you get about 83 calls per minute on average. An agent that lists a repository's issues, opens each one and then reads every comment can spend hundreds of calls in a few minutes, so the hourly budget is a real constraint during big triage sessions.

Secondary Limits per Minute

Secondary limits exist to stop bursts, and agents produce bursts. GitHub documents these:

  • 100 concurrent requests at most.
  • 900 points per minute for REST endpoints.
  • 90 seconds of CPU time per 60 seconds of real time.
  • 80 content-generating requests per minute and 500 per hour.
  • 2,000 OAuth access token requests per hour.

Parallel tool calls pile up against the concurrency limit. An agent that comments on dozens of issues in a loop hits the content-generating cap of 80 per minute long before it touches the hourly budget.

Fixing 403 and 429 Errors

Tired developer at a laptop late at night under a single desk lamp

When a limit trips, GitHub answers with 403 or 429. Read the response headers before you change anything:

HeaderMeaning
x-ratelimit-limitMaximum requests per hour
x-ratelimit-remainingRequests left in the current window
x-ratelimit-usedRequests made in the current window
x-ratelimit-resetWhen the window resets, in UTC epoch seconds
x-ratelimit-resourceWhich resource the request counted against

Backoff That Works

  1. If the response carries a retry-after header, wait that many seconds.
  2. Otherwise wait until the time in x-ratelimit-reset.
  3. For secondary limits with no headers, wait at least one minute, then lengthen the delay exponentially on each new failure.

Put the rule into your agent's instructions too: when GitHub returns 403 or 429, stop and report instead of retrying. An agent that retries instantly just burns more requests against a limit that has not reset.

3 Common Mistakes

  1. Switching on every toolset. The context fills with tool definitions and the agent gets slower and less accurate at picking tools.
  2. Writing in a loop. Bulk comments, labels or issues trip the 80-per-minute content cap fast. Batch the work and pause between batches.
  3. Reusing one broad token everywhere. When it leaks or misbehaves, everything breaks at once. Give each project its own narrow token.

Put a PicassoIA Model to Work

Three developers reviewing a laptop screen together around a wooden table

The GitHub MCP server returns raw material: issue lists, diffs, comment threads. A language model turns that into decisions. Claude Sonnet 5 on PicassoIA handles multi-step coding and tool-use tasks, reads screenshots, and lets you set how hard it thinks, which makes it a good second pair of eyes on what your GitHub agent returns.

How to Use Sonnet 5 on PicassoIA

  1. Open Claude Sonnet 5 on PicassoIA.
  2. Paste the issue list or diff your agent returned into the Prompt field. Strip tokens and private data first.
  3. Set a System Prompt once, for example: "You review pull requests and flag risky changes in three bullets." Reuse it for the whole project.
  4. Choose the Effort level. low is the default and the fastest, while high or max suits a bug that touches several files.
  5. Leave Max Tokens at 8,192 unless the answer gets cut off.
  6. Attach a screenshot of the error to the Image field when text alone is not enough.

PicassoIA also exposes its own developer API and MCP connector, and the same limit logic applies. The API lives at https://api.picassoia.com/v1, takes Bearer credentials that start with pia_sk_, and runs jobs asynchronously: create a prediction, poll it, then fetch the result. Four models are reachable through the API and MCP: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video and Seedance 2.5 Lite. An account can run 5 concurrent predictions, shared across credentials and MCP connections. Same lesson as GitHub's 100 concurrent requests: the concurrency cap is the first wall a parallel agent hits.

Try It on Picasso IA

Your README, release notes and issue templates read better with a real image on top. Generate a photorealistic header with PicassoIA Image, then refine the framing or fix a detail with PicassoIA Image Editor Pro.

Three prompts worth trying today:

  • A tidy developer desk at sunrise with a laptop, a notebook and a mug, shot on 35mm film.
  • A narrow server room aisle with soft overhead light and neat cables, wide angle.
  • A quiet team review session around a wooden table, natural window light.

Pick one, generate a few variations, and see which one makes your next repository page stand out. Start creating your own images with Picasso IA and experiment until the result looks like your project.

Share this article