Large Language ModelsGenerate imagesGenerate videos
Hugging Face Inference API Free Tier: Limits and Pricing
Free Hugging Face accounts no longer list included monthly inference credits, PRO adds $2.00 for $9 a month, and the Hub throttles traffic in 5-minute windows. See the real numbers for credits, rate limits, billing rules and dedicated endpoint costs.
If you typed Hugging Face Inference API free tier limits and pricing into a search bar, you probably hoped for a clear monthly allowance of free calls. The picture in October 2026 is more nuanced, and plenty of forum threads and older tutorials are now out of date. The old serverless Inference API was folded into Inference Providers, and the rules for credits, billing and throttling moved with it.
This article lays out what a free account can do today, what a PRO subscription adds, how the rate limits behave, and where real costs appear once you leave the playground. Every number comes from the official pricing, rate limit and Inference Endpoints pages as they read at the time of writing, so open the billing page of your own account before you commit a budget.
What the Free Tier Gives You
The short version: a free account is a way to test an idea, not a production allowance. You get an account, an access token, the Hub, the playground, and the option to buy credits when you want to run real traffic. What you do not get, according to the current pricing page, is a monthly bucket of included inference credit.
Credits for Free Accounts
At the time of writing, the official pricing page lists no included monthly credits for free users. Free accounts can still call Inference Providers, but only after purchasing credits. Several third-party summaries still quote $0.10 per month for free accounts, so you will see that figure repeated around the web. If your own billing page shows an allowance, trust it. If it shows nothing, the allowance is gone for you.
Where PRO Fits In
PRO costs $9 per month and includes $2.00 in monthly credits. Those are general-purpose compute credits, so they also pay for Inference Endpoints, upgraded CPU and GPU hardware for Spaces (including ZeroGPU usage beyond your quota) and Jobs. They are credited every month and applied automatically before any pay-as-you-go usage is billed.
Team and Enterprise Seats
Team and Enterprise organizations receive $2.00 per seat, shared among all members. To spend that pool you must name the organization on each request, either with the X-HF-Bill-To header or with the bill_to parameter of the Python client. Admins can also set a spending limit and switch off specific providers.
Account type
Included monthly credits
Extra usage
Free
None
Yes, after buying credits
PRO ($9 per month)
$2.00
Yes, pay as you go
Team or Enterprise
$2.00 per seat
Yes, pay as you go
💡 Tip: Credits apply only to requests routed through Hugging Face. If you plug in your own provider account, that provider bills you directly and your included credits stay untouched.
How Billing Works Behind the Scenes
Hugging Face says it passes provider prices straight through with no markup. What changes is who sends the invoice, and that depends on how you make the call.
Routed Requests vs Your Own Provider
Routed by Hugging Face
Your own provider account
Who bills you
Hugging Face
The provider
Included credits apply
Yes
No
Provider account needed
No
Yes
Best for
Simplicity and one invoice
Billing control, providers not integrated
Routed requests are the default. You send your Hugging Face token to the router, which forwards the call to the provider. The router speaks an OpenAI-compatible format at router.huggingface.co/v1, so most existing OpenAI client code works after you change the base URL and the token.
Why hf-inference Looks Different
The hf-inference provider is the direct successor of the old "Inference API (serverless)". Past your included credits, you pay for compute time multiplied by the price of the underlying hardware. As of July 2025 it focuses mostly on CPU inference: embeddings, text ranking, text classification and small models of historic importance such as BERT or GPT-2. Large chat models and image generators are usually served by partner providers instead.
A Cost Example With Real Numbers
The pricing docs include a worked example. A request to FLUX.1-dev that takes 10 seconds on a GPU costing $0.00012 per second is billed $0.0012. At that rate:
$2.00 of PRO credit pays for roughly 1,666 images.
$0.10 (the allowance older posts mention) would have paid for about 83 images.
Treat this as an illustration. Real prices vary by model and provider, and your own requests may run longer than 10 seconds. A few hundred test images a month fit comfortably inside that credit, while a public app serving real users can burn through it in an afternoon.
Rate Limits and 429 Errors
Separate from billing, the Hub throttles requests in 5-minute fixed windows. The values below come from the September 2025 table on the rate limit page, and the docs flag the anonymous and free numbers as subject to change depending on platform health.
The Three Request Buckets
Hugging Face sorts traffic into three buckets:
Hub APIs: model and dataset search, repo creation, user management.
Resolvers: URLs containing /resolve/, which serve model and dataset files to libraries and apps.
Pages: the web pages on huggingface.co.
Plan
API
Resolvers
Pages
Anonymous (per IP)
500
3,000
100
Free
1,000
5,000
200
PRO
2,500
12,000
400
Team
3,000
20,000
400
Enterprise
6,000
50,000
600
All numbers are requests per 5 minutes. For organizations, limits apply to each member individually, not to the group. Individual inference providers may add limits of their own on top, so read the error body before you decide which layer said no.
Reading the RateLimit Headers
Every throttled call returns 429 Too Many Requests. Responses follow the IETF draft for rate limit headers: RateLimit reports how many requests remain and the seconds until reset, while RateLimit-Policy states the window, for example 100 requests per 5 minutes. Your billing page also shows three live gauges, one per bucket, that turn red when you cross the line.
The most common fix is almost boring: pass your HF_TOKEN everywhere. Anonymous traffic sits at the lowest limits, and a library that forgets to forward the token will burn through them fast.
A Retry Pattern That Works
From version 1.2.0, the huggingface_hub library reads the RateLimit header on a 429 and waits exactly that long before retrying, for file downloads and paginated Hub API calls. For your own inference calls, a small backoff wrapper is enough:
import os
import time
from huggingface_hub import InferenceClient
client = InferenceClient(token=os.environ["HF_TOKEN"])
def ask(prompt, retries=4):
for attempt in range(retries):
try:
reply = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V3-0324",
messages=[{"role": "user", "content": prompt}],
max_tokens=200,
)
return reply.choices[0].message.content
except Exception as err:
if "429" not in str(err):
raise
time.sleep(2 ** attempt)
raise RuntimeError("Still rate limited after retries")
Spread jobs over longer periods, prefer Resolver URLs to Hub API calls where possible, and upgrade only when the gauges prove you need it.
Speed on Shared Capacity
Free and low-cost traffic usually shares capacity with everyone else, and that shows up in response time. The pricing pages I checked publish no latency guarantee, so measure your own.
Time ten requests at different hours and record the median and the slowest call.
Pick a smaller model when a short answer is enough. Less compute means a smaller bill and a faster reply.
Cap max_tokens so a chatty model cannot run for 30 seconds on a one-line question.
Cache repeated prompts in your own database instead of paying twice.
Run bulk jobs off-peak and let a queue absorb the waiting.
A prototype can live with a slow first response. A checkout flow cannot, and that is usually the moment people start looking at dedicated hardware. Keep a small log of your timings in a spreadsheet, because a pricing change or a busy week is much easier to spot when you have a baseline.
When Inference Endpoints Pay Off
Inference Endpoints give you a dedicated machine for one model. Billing is by the minute while the endpoint is running or initializing. Paused endpoints cost nothing, while endpoints scaled to zero still count against your quota, so pause them if you need the slot back.
Hourly Rates by Hardware
These are the AWS prices from the Inference Endpoints page:
Hardware
Memory
Hourly rate
CPU x1 (1 vCPU)
2 GB
$0.033
CPU x4 (4 vCPUs)
8 GB
$0.134
GPU T4 x1
14 GB
$0.50
GPU L4 x1
24 GB
$0.80
GPU A10G x1
24 GB
$1.00
GPU A100 x1
80 GB
$2.50
GPU H200 x1
141 GB
$5.00
GCP also lists an H100 with 80 GB at $10.00 per hour.
What a Month Really Costs
Multiply the rate by hours. An endpoint left on for 730 hours costs about $24 on the smallest CPU, $365 on a T4, $730 on an A10G and $1,825 on an A100. If you run it only 8 hours a day on 22 working days (176 hours), the T4 drops to about $88 and the A10G to $176.
💡 Reality check: the $2.00 of monthly PRO credit buys about four hours of a T4 endpoint. It is a trial budget for experiments, not a hosting plan.
Picking the Right Plan
Solo Developers and Hobbyists
Stay on the free account while you prototype with small models and the playground. Buy a few dollars of credits when you need real traffic, and set a personal spending cap you are comfortable with. PRO earns its $9 when you also want the higher Hub limits (2,500 API requests per window against 1,000), the $2.00 of compute credit and faster support responses.
Small Teams and Startups
Team plans pool $2.00 per seat and allow centralized billing through the X-HF-Bill-To header, so one invoice pays for everyone's tokens. Set the spending limit on day one and disable providers you do not use. Enterprise adds an endpoint for pulling usage as a daily series broken down by member, model and provider, with requests appearing up to 2 hours after they are made.
Students and Researchers
Academic budgets are tight, so mix your tools. Use the Hub for downloads, where free users get the highest bucket at 5,000 Resolver requests per 5 minutes, run small experiments on local hardware, and buy credits only for the final evaluation run. Academia Hub organizations get Team-level limits of 3,000 API, 20,000 Resolver and 400 page requests.
Track spend before it surprises you. The Inference Providers settings page shows last month's usage by model and provider, while the billing page shows credits and the rate limit gauges. Check both once a week.
GPT OSS 120B is a 120-billion-parameter open-weight model for writing, summaries, Q&A and code. The settings below are the ones its page exposes:
Open the GPT OSS 120B page and find the Prompt field.
Type a specific instruction. Name the format, the audience and the length you want.
Keep Temperature at its default of 0.1 for focused, factual answers, and raise it when you want more varied wording.
Leave Max Tokens at 2048, the default output ceiling, or lower it for short replies.
Touch Top P, Presence Penalty and Frequency Penalty only if a long answer starts looping or sounding flat.
Run the model, read the result, then tighten the prompt and run it again.
Setting
Default
What it changes
Temperature
0.1
Focused vs varied wording
Top P
1
How broadly the model samples vocabulary
Max Tokens
2048
Length ceiling for the answer
Presence Penalty
0
Pushes the model toward new topics
Frequency Penalty
0
Reduces repeated words and phrases
💡 Tip: Use the language model to write your image prompts, then paste them into a text-to-image model. One browser tab handles the draft and the picture.
Reading pricing tables is useful, but nothing beats running a prompt and seeing what comes back. Open PicassoIA, write a short brief for a product shot, a landscape or a portrait, and let a text-to-image model render it. Then ask GPT OSS 120B to rewrite your prompt in three different moods and compare the results side by side.
You do not need a token, a credit balance or a retry loop to begin. Pick a model from the full catalog, try a few prompts, and keep the ones that surprise you. Your first image is a few minutes away.