Large Language ModelsGenerate imagesGenerate videos

Hugging Face Inference API Free Tier: Limits and Pricing

Free Hugging Face accounts no longer list included monthly inference credits, PRO adds $2.00 for $9 a month, and the Hub throttles traffic in 5-minute windows. See the real numbers for credits, rate limits, billing rules and dedicated endpoint costs.

Hugging Face Inference API Free Tier: Limits and Pricing
Cristian Da Conceicao
Founder of Picasso IA

If you typed Hugging Face Inference API free tier limits and pricing into a search bar, you probably hoped for a clear monthly allowance of free calls. The picture in October 2026 is more nuanced, and plenty of forum threads and older tutorials are now out of date. The old serverless Inference API was folded into Inference Providers, and the rules for credits, billing and throttling moved with it.

This article lays out what a free account can do today, what a PRO subscription adds, how the rate limits behave, and where real costs appear once you leave the playground. Every number comes from the official pricing, rate limit and Inference Endpoints pages as they read at the time of writing, so open the billing page of your own account before you commit a budget.

What the Free Tier Gives You

Flat lay of a notebook with handwritten pencil calculations, a calculator and black coffee beside a laptop

The short version: a free account is a way to test an idea, not a production allowance. You get an account, an access token, the Hub, the playground, and the option to buy credits when you want to run real traffic. What you do not get, according to the current pricing page, is a monthly bucket of included inference credit.

Credits for Free Accounts

At the time of writing, the official pricing page lists no included monthly credits for free users. Free accounts can still call Inference Providers, but only after purchasing credits. Several third-party summaries still quote $0.10 per month for free accounts, so you will see that figure repeated around the web. If your own billing page shows an allowance, trust it. If it shows nothing, the allowance is gone for you.

Where PRO Fits In

PRO costs $9 per month and includes $2.00 in monthly credits. Those are general-purpose compute credits, so they also pay for Inference Endpoints, upgraded CPU and GPU hardware for Spaces (including ZeroGPU usage beyond your quota) and Jobs. They are credited every month and applied automatically before any pay-as-you-go usage is billed.

Team and Enterprise Seats

Team and Enterprise organizations receive $2.00 per seat, shared among all members. To spend that pool you must name the organization on each request, either with the X-HF-Bill-To header or with the bill_to parameter of the Python client. Admins can also set a spending limit and switch off specific providers.

Account typeIncluded monthly creditsExtra usage
FreeNoneYes, after buying credits
PRO ($9 per month)$2.00Yes, pay as you go
Team or Enterprise$2.00 per seatYes, pay as you go

💡 Tip: Credits apply only to requests routed through Hugging Face. If you plug in your own provider account, that provider bills you directly and your included credits stay untouched.

How Billing Works Behind the Scenes

Low-angle view down a data center aisle between rows of black server racks with neat cable bundles overhead

Hugging Face says it passes provider prices straight through with no markup. What changes is who sends the invoice, and that depends on how you make the call.

Routed Requests vs Your Own Provider

Routed by Hugging FaceYour own provider account
Who bills youHugging FaceThe provider
Included credits applyYesNo
Provider account neededNoYes
Best forSimplicity and one invoiceBilling control, providers not integrated

Routed requests are the default. You send your Hugging Face token to the router, which forwards the call to the provider. The router speaks an OpenAI-compatible format at router.huggingface.co/v1, so most existing OpenAI client code works after you change the base URL and the token.

Why hf-inference Looks Different

The hf-inference provider is the direct successor of the old "Inference API (serverless)". Past your included credits, you pay for compute time multiplied by the price of the underlying hardware. As of July 2025 it focuses mostly on CPU inference: embeddings, text ranking, text classification and small models of historic importance such as BERT or GPT-2. Large chat models and image generators are usually served by partner providers instead.

A Cost Example With Real Numbers

The pricing docs include a worked example. A request to FLUX.1-dev that takes 10 seconds on a GPU costing $0.00012 per second is billed $0.0012. At that rate:

  • $2.00 of PRO credit pays for roughly 1,666 images.
  • $0.10 (the allowance older posts mention) would have paid for about 83 images.

Treat this as an illustration. Real prices vary by model and provider, and your own requests may run longer than 10 seconds. A few hundred test images a month fit comfortably inside that credit, while a public app serving real users can burn through it in an afternoon.

Rate Limits and 429 Errors

Close-up of a developer's hands typing at a desk with blurred code on a monitor behind

Separate from billing, the Hub throttles requests in 5-minute fixed windows. The values below come from the September 2025 table on the rate limit page, and the docs flag the anonymous and free numbers as subject to change depending on platform health.

The Three Request Buckets

Hugging Face sorts traffic into three buckets:

  • Hub APIs: model and dataset search, repo creation, user management.
  • Resolvers: URLs containing /resolve/, which serve model and dataset files to libraries and apps.
  • Pages: the web pages on huggingface.co.
PlanAPIResolversPages
Anonymous (per IP)5003,000100
Free1,0005,000200
PRO2,50012,000400
Team3,00020,000400
Enterprise6,00050,000600

All numbers are requests per 5 minutes. For organizations, limits apply to each member individually, not to the group. Individual inference providers may add limits of their own on top, so read the error body before you decide which layer said no.

Reading the RateLimit Headers

Every throttled call returns 429 Too Many Requests. Responses follow the IETF draft for rate limit headers: RateLimit reports how many requests remain and the seconds until reset, while RateLimit-Policy states the window, for example 100 requests per 5 minutes. Your billing page also shows three live gauges, one per bucket, that turn red when you cross the line.

The most common fix is almost boring: pass your HF_TOKEN everywhere. Anonymous traffic sits at the lowest limits, and a library that forgets to forward the token will burn through them fast.

A Retry Pattern That Works

From version 1.2.0, the huggingface_hub library reads the RateLimit header on a 429 and waits exactly that long before retrying, for file downloads and paginated Hub API calls. For your own inference calls, a small backoff wrapper is enough:

import os
import time
from huggingface_hub import InferenceClient

client = InferenceClient(token=os.environ["HF_TOKEN"])

def ask(prompt, retries=4):
    for attempt in range(retries):
        try:
            reply = client.chat.completions.create(
                model="deepseek-ai/DeepSeek-V3-0324",
                messages=[{"role": "user", "content": prompt}],
                max_tokens=200,
            )
            return reply.choices[0].message.content
        except Exception as err:
            if "429" not in str(err):
                raise
            time.sleep(2 ** attempt)
    raise RuntimeError("Still rate limited after retries")

Spread jobs over longer periods, prefer Resolver URLs to Hub API calls where possible, and upgrade only when the gauges prove you need it.

Speed on Shared Capacity

Tight close-up of a hand holding a brushed steel stopwatch above a wooden desk

Free and low-cost traffic usually shares capacity with everyone else, and that shows up in response time. The pricing pages I checked publish no latency guarantee, so measure your own.

  • Time ten requests at different hours and record the median and the slowest call.
  • Pick a smaller model when a short answer is enough. Less compute means a smaller bill and a faster reply.
  • Cap max_tokens so a chatty model cannot run for 30 seconds on a one-line question.
  • Cache repeated prompts in your own database instead of paying twice.
  • Run bulk jobs off-peak and let a queue absorb the waiting.

A prototype can live with a slow first response. A checkout flow cannot, and that is usually the moment people start looking at dedicated hardware. Keep a small log of your timings in a spreadsheet, because a pricing change or a busy week is much easier to spot when you have a baseline.

When Inference Endpoints Pay Off

Technician crouched beside an open rack-mounted GPU server holding a flashlight and a tablet

Inference Endpoints give you a dedicated machine for one model. Billing is by the minute while the endpoint is running or initializing. Paused endpoints cost nothing, while endpoints scaled to zero still count against your quota, so pause them if you need the slot back.

Hourly Rates by Hardware

These are the AWS prices from the Inference Endpoints page:

HardwareMemoryHourly rate
CPU x1 (1 vCPU)2 GB$0.033
CPU x4 (4 vCPUs)8 GB$0.134
GPU T4 x114 GB$0.50
GPU L4 x124 GB$0.80
GPU A10G x124 GB$1.00
GPU A100 x180 GB$2.50
GPU H200 x1141 GB$5.00

GCP also lists an H100 with 80 GB at $10.00 per hour.

What a Month Really Costs

Multiply the rate by hours. An endpoint left on for 730 hours costs about $24 on the smallest CPU, $365 on a T4, $730 on an A10G and $1,825 on an A100. If you run it only 8 hours a day on 22 working days (176 hours), the T4 drops to about $88 and the A10G to $176.

💡 Reality check: the $2.00 of monthly PRO credit buys about four hours of a T4 endpoint. It is a trial budget for experiments, not a hosting plan.

Picking the Right Plan

Solo Developers and Hobbyists

Stay on the free account while you prototype with small models and the playground. Buy a few dollars of credits when you need real traffic, and set a personal spending cap you are comfortable with. PRO earns its $9 when you also want the higher Hub limits (2,500 API requests per window against 1,000), the $2.00 of compute credit and faster support responses.

Small Teams and Startups

Four colleagues gathered around a long oak table in a bright co-working space looking at a laptop

Team plans pool $2.00 per seat and allow centralized billing through the X-HF-Bill-To header, so one invoice pays for everyone's tokens. Set the spending limit on day one and disable providers you do not use. Enterprise adds an endpoint for pulling usage as a daily series broken down by member, model and provider, with requests appearing up to 2 hours after they are made.

Students and Researchers

Student in side profile at a long wooden library table with a laptop, textbooks and a highlighter

Academic budgets are tight, so mix your tools. Use the Hub for downloads, where free users get the highest bucket at 5,000 Resolver requests per 5 minutes, run small experiments on local hardware, and buy credits only for the final evaluation run. Academia Hub organizations get Team-level limits of 3,000 API, 20,000 Resolver and 400 page requests.

Track spend before it surprises you. The Inference Providers settings page shows last month's usage by model and provider, while the billing page shows credits and the rate limit gauges. Check both once a week.

Over-the-shoulder view of a man at a cafe table watching a blurred bar chart on a laptop beside a flat white

Run the Same Models on PicassoIA

Photographer reviewing large printed mountain lake landscapes in a studio at dusk

If your goal is to use open-weight models without juggling tokens, credits and 429 errors, PicassoIA hosts many of them in the browser. The large language model category includes GPT OSS 120B, Llama 4 Scout Instruct, Qwen3 235B A22B Instruct 2507, DeepSeek v3.1, Kimi K2.6 and the compact Granite 4.1 8B.

How to Use GPT OSS on PicassoIA

GPT OSS 120B is a 120-billion-parameter open-weight model for writing, summaries, Q&A and code. The settings below are the ones its page exposes:

  1. Open the GPT OSS 120B page and find the Prompt field.
  2. Type a specific instruction. Name the format, the audience and the length you want.
  3. Keep Temperature at its default of 0.1 for focused, factual answers, and raise it when you want more varied wording.
  4. Leave Max Tokens at 2048, the default output ceiling, or lower it for short replies.
  5. Touch Top P, Presence Penalty and Frequency Penalty only if a long answer starts looping or sounding flat.
  6. Run the model, read the result, then tighten the prompt and run it again.
SettingDefaultWhat it changes
Temperature0.1Focused vs varied wording
Top P1How broadly the model samples vocabulary
Max Tokens2048Length ceiling for the answer
Presence Penalty0Pushes the model toward new topics
Frequency Penalty0Reduces repeated words and phrases

💡 Tip: Use the language model to write your image prompts, then paste them into a text-to-image model. One browser tab handles the draft and the picture.

The same account reaches the image and video categories too. For still images, try FLUX 2 Pro, Seedream 4.5, Imagen 4 Fast, P-Image or FLUX Dev, the same family used in the pricing example above. For motion, Wan 2.6 T2V, Veo 3.1 Fast and Seedance 2.0 turn a text prompt into a clip, and PicassoIA Video is listed as a free, unlimited generator from text or an image.

Create Your Own Images Today

Reading pricing tables is useful, but nothing beats running a prompt and seeing what comes back. Open PicassoIA, write a short brief for a product shot, a landscape or a portrait, and let a text-to-image model render it. Then ask GPT OSS 120B to rewrite your prompt in three different moods and compare the results side by side.

You do not need a token, a credit balance or a retry loop to begin. Pick a model from the full catalog, try a few prompts, and keep the ones that surprise you. Your first image is a few minutes away.

Share this article