Large Language ModelsGenerate imagesGenerate videos

AI API Gateway Open Source: Best Options and Vercel AI Gateway

Compare the best open source AI API gateways, including LiteLLM, Bifrost, Envoy, and Helicone, with Vercel AI Gateway. See how routing, fallbacks, budgets, and logs differ, which license limits apply, and how to pick the right setup for your team in one afternoon.

AI API Gateway Open Source: Best Options and Vercel AI Gateway
Cristian Da Conceicao
Founder of Picasso IA

Every team that ships an AI feature hits the same wall within a few weeks. One provider rate limits you at 2 a.m., another renames a model, finance asks why the invoice doubled, and your codebase now carries four different SDKs. An AI API gateway fixes this by sitting between your app and every model provider, so routing, retries, budgets and logs live in one place instead of being scattered across services.

This article compares the best AI API gateway open source options with Vercel AI Gateway, the hosted choice many teams reach for first. You will see what each one does well, where each one hurts, and how to pick without running a month-long bake-off. Everything here reflects how these tools are documented by their maintainers, so check the current license and pricing pages before you commit.

What an AI Gateway Actually Does

Aerial view of a highway interchange where traffic splits toward different exits

Think of a gateway as a traffic interchange for model requests. Your application sends one request in a single format, usually the same shape as the OpenAI chat API. The gateway decides which provider and which model should answer, forwards the call, and returns a normalised response. Your code never needs to know who handled it.

That is why you will also see this layer called an LLM gateway, an LLM proxy or a unified API. The names differ, the job is the same: one door in, many providers out.

One Endpoint, Many Providers

Without a gateway, every provider brings its own SDK, authentication scheme, error codes and streaming format. Adding a second provider means a second integration, and a third means a third. With a gateway, you change a base URL and a model name. Most gateways advertise an OpenAI compatible endpoint for this reason: existing client libraries keep working, and the provider behind them can change on a Tuesday afternoon without a deploy.

The Jobs It Handles

A solid gateway handles six jobs:

  • Routing: send each request to the right provider, model or region, with load balancing across several deployments.
  • Fallbacks: switch to a backup when the first choice errors out or times out.
  • Budgets and rate limiting: cap spend and request rates per team, project or user.
  • Caching: return stored answers for repeated prompts to cut cost and latency.
  • Observability: log token usage, latency, cost and errors for every call.
  • Guardrails: filter or redact sensitive input and output before it leaves your network.

💡 Tip: You do not need all six on day one. Most teams start with routing, fallbacks and logs, then add budgets after the first surprising invoice.

Open Source or Hosted Gateway?

Close-up of ethernet cables plugged into a rack mounted network switch

This is the first fork in the road, and it matters more than any feature table.

When Open Source Wins

Open source gateways run on your own infrastructure, so prompts and responses pass through machines you control. That is the main reason regulated teams choose a self-hosted proxy server. You can read the code, patch a bug, add a provider nobody supports yet, and avoid a per-request fee from the gateway itself. The trade is operational work: you own upgrades, scaling, secrets storage and the 3 a.m. page when the proxy falls over.

Credentials deserve special care. A gateway holds provider credentials for every vendor you use, which makes it a high value target. Store them in a secrets manager rather than environment files, rotate them on a schedule, and give each internal team a virtual credential so the real provider ones never leave the gateway.

When Hosted Is Smarter

A hosted gateway removes that work. There is nothing to deploy, no database to back up and no cluster to patch. If your team is three developers shipping a product, an hour spent on a managed gateway beats a week spent tuning a proxy. The cost is dependence on a vendor's uptime, data policies and pricing, plus less room to bend the behaviour to your needs.

QuestionOpen source, self-hostedHosted service
Who runs it?Your teamThe vendor
Where do prompts travel?Inside your networkThrough the vendor
Setup timeHours to daysMinutes
Custom behaviourEdit code or pluginsSettings only
Ongoing workUpgrades, scaling, monitoringClose to none

💡 License check: a public repository does not always mean fully open source. Several gateways keep features such as single sign-on, role-based access control or audit logs under a separate commercial license. Read the license file, not just the README.

Best Open Source AI API Gateways

Developer typing in a terminal window at a standing desk in morning light

These are the projects most teams shortlist. Licenses and provider counts change often, so treat this section as a map, not a contract.

LiteLLM: Widest Provider Support

LiteLLM is usually the first name that comes up. It ships as a Python library and as a proxy server, speaks the OpenAI format, and supports a long list of providers, which makes it the fastest route to a working gateway. The core is MIT licensed. Single sign-on beyond a small number of users, role-based access control, audit logs and some moderation callbacks sit behind a separate enterprise license, so check that boundary before you plan around them. Because it runs on Python, teams with very high request volume tend to benchmark it carefully and scale it horizontally.

Portkey and Bifrost Compared

Portkey Gateway is MIT licensed and keeps fallbacks, retries and guardrails in the open source core. It suits teams that want policy controls without buying a full platform on day one.

Bifrost is written in Go, released under Apache 2.0, and advertises sub-millisecond overhead. Independent tests do not always agree with the headline number, so run your own load test with your own prompt sizes before trusting any latency claim. Its provider list is shorter than LiteLLM's, which is fine if the major vendors are all you need.

Envoy, Helicone, and Kong

Envoy AI Gateway is built on Envoy and licensed under Apache 2.0. It fits teams that already run Kubernetes and treat Envoy as their traffic layer, since model traffic then follows the same policies, tracing and rollout habits as every other service.

Rows of server racks along a clean data center aisle in one point perspective

Helicone is Apache 2.0 licensed and grew out of observability, so its logging and cost views are a strong point. Kong adds AI routing through plugins on top of its open source API gateway, which is convenient if Kong already fronts your other APIs. Check which AI plugins ship in the free edition before you rely on them.

GatewayLicenseBest fit
LiteLLMMIT core, paid enterprise tierFast start, many providers
Portkey GatewayMITGuardrails and retries
BifrostApache 2.0Low overhead, Go stack
Envoy AI GatewayApache 2.0Kubernetes platforms
HeliconeApache 2.0Cost and usage visibility
KongApache 2.0 core, some paid pluginsTeams already on Kong

Vercel AI Gateway Up Close

Vercel AI Gateway is the hosted counterpart, and the honest framing is simple: it is not open source. The AI SDK around it is open source, but the gateway itself is a managed service. You are choosing convenience over control, and for plenty of teams that is the right trade.

What You Get

One credential and one endpoint give you access to models from many providers. It exposes OpenAI compatible and Anthropic compatible endpoints, so existing clients usually need only a new base URL. On top of that you get provider routing, automatic retries, model fallbacks, spend and latency logs, and per-credential budgets.

On cost, Vercel states that it charges the provider's list price with no token markup, including when you bring your own provider credentials (often called BYOK). New teams receive a small monthly free credit, listed at $5 at the time of writing. Check the current pricing page, because credit amounts change.

Vercel also reports that across its production traffic through April 2026, automatic fallback saved roughly 3.5% of requests that had hit an error, rate limit or timeout on the first route. That figure comes from the vendor, so read it as a signal about how common provider hiccups are, not as a promise for your workload.

Where It Falls Short

  • Not self-hostable. Prompts pass through a vendor, which can rule it out under strict data policies.
  • Credit based. Prepaid credits suit small teams, but finance departments sometimes want invoices and committed spend.
  • Less customisation. You tune settings, you do not patch code.
  • Mild lock-in. Because the API is OpenAI compatible, leaving is mostly a base URL change, but your logs and budgets stay behind.

Routing, Fallbacks, and Retries

Low angle view of an air traffic control tower at dusk

Fallbacks are the feature that justifies a gateway on their own. A fallback chain lists models in order: try the primary, and if it returns a 5xx error, a 429 rate limit or a timeout, move to the next one. A sensible chain mixes providers, not just models, because a provider outage takes down every model it hosts.

A practical chain for a chat feature might look like this:

  1. Primary: Claude Sonnet 5 for quality sensitive answers.
  2. First backup: GPT 5.6 Terra on a different provider.
  3. Last resort: Gemini 3.5 Flash for speed when everything else is struggling.

Aerial view of a railway switching yard with tracks branching in many directions

Every gateway in this article can express a chain like that in configuration. What differs is how much control you get over when it triggers: by status code, by latency threshold, by content filter result, or by custom rule.

Load balancing is the quiet sibling of fallbacks. If you have two deployments of the same model, say in different regions or under different accounts, the gateway can spread traffic between them by weight and shift away from the one that is slow. This also lifts your effective rate limit, since each account has its own ceiling.

Retries Need Limits

Retries help with short lived errors, but they carry risk. A retried request can be billed twice if the first call finished late, and a retry storm can push a struggling provider further down. Keep these habits:

  • Use two or three retries at most, with exponential backoff and random jitter.
  • Set timeouts per model. A reasoning model can legitimately think for a minute, while a small chat model should answer in seconds.
  • Plan for streaming. Once the first tokens have reached the user, a mid-stream fallback cannot quietly restart the answer, so decide early how your interface handles it.

Controlling Cost and Usage

Overhead view of a desk with a notebook, calculator and coffee for tracking costs

Model bills grow quietly. A gateway turns spend into something you can set and watch.

Budgets Per Team and Project

Issue one virtual credential per team, project or customer and attach a monthly budget and a rate limit to it. When a runaway script loops, it exhausts its own budget and stops, instead of draining the shared account. Tag every request so the log answers "who spent what" without a spreadsheet exercise. Both LiteLLM and Vercel AI Gateway offer budgets, and with any self-hosted tool you should confirm whether the budget feature lives in the free edition.

Then use the gateway to match the model to the job:

  • Send classification, tagging and short summaries to a small, fast model such as Gemini 3.5 Flash or Qwen3.7-Plus.
  • Reserve top tier models for code, reasoning and customer facing text, for example Kimi K2.6 for agent style work.
  • Run open weight models such as Llama 4 Maverick Instruct, Deepseek v3.1 or GPT OSS 120B behind the same endpoint as paid APIs, which gives you a low cost tier.
  • Cache repeated prompts such as FAQ answers, and alert at 80% of a budget rather than 100%.

💡 Log with care: storing every prompt makes debugging easy, but it can also store personal data. Decide on retention and redaction rules before you switch on full body logging.

How to Choose in an Afternoon

Two engineers discussing a hand drawn diagram on a whiteboard

You do not need a month of research. You need a clear picture of your constraints and one short test.

A Simple Decision Table

Your situationStart with
Small team, want it working todayVercel AI Gateway
Prompts must stay inside your networkLiteLLM or Portkey Gateway, self-hosted
Already run Kubernetes and EnvoyEnvoy AI Gateway
Added latency per request matters mostBenchmark Bifrost against LiteLLM
Kong already fronts your APIsKong AI plugins
Cost and usage visibility comes firstHelicone

Then run the same five step test on your top two picks:

  1. Point a sandbox app at the gateway's base URL.
  2. Send a thousand realistic prompts, not toy examples.
  3. Revoke the primary provider's credentials and confirm the fallback fires.
  4. Check that logs show tokens, cost and errors per request.
  5. Measure added latency against calling the provider directly.

Three Mistakes to Avoid

  1. Choosing on benchmarks alone. A gateway that adds one millisecond in a lab can add more once logging, authentication and guardrails are switched on.
  2. Skipping the fallback test. Pull the plug on the primary provider and watch what happens. An untested chain is a guess.
  3. Treating the gateway as a single point of failure. Run at least two instances behind a load balancer, or keep a direct provider path in your code as an emergency switch.

Skip the Plumbing and Create

Designer at a wide desk reviewing a landscape photograph on a large monitor

Not every project needs a gateway. If your goal is to produce text, images or video rather than operate infrastructure, Picasso IA gives you the models directly in the browser, with no proxy to deploy and no credentials to rotate.

On the language side you can try Claude Sonnet 5, GPT 5.6 Terra and Gemini 3.5 Flash side by side, the same comparison a gateway would let you script. For pictures, run one prompt through Seedream 5 Pro, GPT Image 2 and Flux 2 Max and keep the result you like best. For motion, Seedance 2.0 turns text into video with built in audio, Veo 3.1 Fast makes quick 1080p clips, and Kling v3 Video aims for cinematic shots.

Pick a prompt you actually need, run it across three models, and compare the results. That one habit teaches more about model quality and cost than any benchmark table. Open Picasso IA, choose a model, and create your first image or video today.

Share this article