GitHub Models is no longer a free API option. GitHub states that the playground, catalog, inference API, and bring-your-own-key feature were fully retired on July 30, 2026. That changes the practical shortlist for developers who need zero-cost inference for prototypes. The strongest remaining choices are not interchangeable: Gemini Developer API, Cloudflare Workers AI, and Hugging Face Inference Providers meter free use in different units, reset on different schedules, expose different model catalogs, and apply different data-handling rules.
Quick answer
After GitHub Models retired on July 30, 2026, the most useful free AI API alternatives serve different jobs. Gemini offers capable multimodal models under model-specific quotas, Cloudflare Workers AI provides 10,000 Neurons daily, and Hugging Face gives free users $0.10 monthly routed-inference credit. Choose by workload, privacy, reset period, model portability, and failure behavior—not by a misleading “free tokens” headline.
What changed on July 30, 2026
GitHub Models used to be an unusually convenient bridge between a model playground and an API authenticated with a GitHub credential. It was attractive for experiments because developers could compare models without opening several provider accounts. That option is now closed. GitHub's current documentation says GitHub Models has been fully retired since July 30, 2026, including its playground, catalog, inference API, and BYOK capability. GitHub points users toward Azure AI Foundry for a broad catalog or GitHub Copilot for workflows inside GitHub, but neither is a drop-in continuation of the former free inference API.
The retirement is more than a catalog change. It shows why a prototype should not encode “free API” as a permanent infrastructure assumption. Free access can be throttled, narrowed, renamed, moved behind billing, or discontinued. A robust application needs an adapter boundary, explicit capacity measurements, and a defined response when the allowance is exhausted.
This comparison therefore asks a more useful question than “which provider gives the most free tokens?” The providers do not all meter tokens. The correct question is: which free allowance can complete your particular workload with acceptable privacy, predictability, and migration risk?
Comparison at a glance
| Option | Free allowance documented on September 29, 2026 | Reset or budget period | Best fit | Main limitation |
|---|---|---|---|---|
| Gemini Developer API | Selected models are free of charge, subject to model- and project-level rate limits | Rate limits use RPM, TPM and RPD dimensions; exact active quotas are shown in AI Studio | Rich multimodal prototypes, long context, structured output | Quotas vary by model and tier; unpaid-service data terms require review |
| Cloudflare Workers AI | 10,000 Neurons per day | Daily at 00:00 UTC | Predictable edge or serverless inference with a hard daily envelope | Neurons are model-dependent, so they do not translate to one universal token count |
| Hugging Face Inference Providers | $0.10 monthly credit for free users, subject to change | Monthly credit | Trying many providers and open models through one client | The free budget is small and the downstream provider still matters |
| GitHub Models | Retired | None | No new workload | Playground, catalog, API and BYOK are unavailable |
The table deliberately avoids manufacturing a token equivalent. A Gemini request can include text, images, audio or video. Cloudflare prices models in Neurons, with model-specific conversion rates. Hugging Face deducts the actual provider cost from a dollar credit. Converting all three into “tokens” without choosing a model, prompt length, output length and modality would create false precision.
Gemini: broad capability, variable capacity
Google's Gemini Developer API has a genuine Free Tier for selected models. The pricing table labels supported input and output as free of charge for models such as Gemini 2.5 Flash and Flash-Lite, while other models or modalities may be paid-only. The important qualifier is “selected.” Free availability belongs to a model and usage tier, not to the mere existence of an API key.
Google applies rate limits across requests per minute, tokens per minute, and requests per day. Limits are per project, not per API key, and a request can be constrained by any applicable dimension. The published rate-limit page also directs developers to AI Studio for active limits, which matters because quotas may change and preview models may have tighter rules. Creating several keys in one project does not multiply the project's capacity.
This makes Gemini strong for capability-first experiments. Its API can cover text, images, audio, video, embeddings, structured output and tool-oriented patterns, depending on the selected model. It is a good fit when one prototype needs more than basic chat completion or when a long-context task would be awkward on a smaller serverless model.
Its free capacity is less convenient for fixed-volume planning. “Free of charge” tells you the token price, but not how many successful jobs your application can complete today. You must inspect the active project quota, choose a stable model, and measure average input, thinking and output consumption. Preview aliases also create migration risk because preview models may be deprecated faster and can carry more restrictive capacity.
Data treatment is another decision input. Google's pricing and billing documentation distinguishes unpaid and paid service treatment. The pricing tables mark Free Tier data as eligible to improve Google products for many models, while Paid Tier is marked otherwise. Google also documents regional and account exceptions in its terms. Do not infer the rule from price alone: read the current terms for the project's billing state, account type and region, and never send secrets or personal data merely because a request costs zero.
When Gemini is the best choice
Choose Gemini when model capability and modality breadth matter more than a perfectly fixed allowance, and when the data terms fit the material being processed. It is especially useful for document understanding, multimodal extraction, prototypes that need structured JSON, and evaluations that would otherwise require several specialized APIs.
Avoid making it the only backend if a daily hard quota could stop a public feature. Build retry logic for 429 responses, cap outputs, and expose a graceful “capacity unavailable” state rather than silently switching to a weaker model.
Cloudflare Workers AI: the clearest daily envelope
Cloudflare documents a free allocation of 10,000 Neurons per day for Workers AI. The allowance resets at 00:00 UTC. On the free plan, operations beyond the allowance fail; on Workers Paid, the first 10,000 daily Neurons remain included and excess usage is billed at the documented per-Neuron rate.
This is the most operationally explicit allowance in the comparison. A developer can treat 10,000 Neurons as a daily capacity budget and observe consumption in the dashboard. The catch is that a Neuron is a compute unit, not a token. Each model has its own Neuron conversion, and modalities such as image generation, speech or embeddings consume capacity differently. You must calculate jobs from the exact model's pricing row.
Workers AI fits applications already using Cloudflare Workers, Pages, AI Gateway, R2, KV, Durable Objects or Vectorize. Inference can be invoked through a Worker binding, REST, or compatible interfaces documented by Cloudflare. Keeping application logic and inference near the edge can simplify authentication and latency for small public utilities.
Cloudflare's data-usage page provides a notable privacy statement: customer content is not used to train models or improve Cloudflare or third-party services without explicit consent. It also warns that the models are third-party services with their own licenses. Customer content may be stored when the application explicitly combines inference with a storage product. That distinction is useful: the inference service's policy does not excuse an application from reviewing the model license or from controlling what it writes to storage and logs.
When Cloudflare is the best choice
Choose Workers AI when you want a predictable daily ceiling, serverless deployment, and tight integration with an edge application. It works well for embeddings, classification, summarization, transcription, or bounded chat experiences where each request has a known maximum.
Do not assume 10,000 Neurons means 10,000 requests. First select a model, then estimate Neurons per typical and worst-case job. Add a daily software budget below the provider ceiling so background work cannot consume the capacity reserved for users.
Hugging Face Inference Providers: maximum catalog flexibility, tiny free budget
Hugging Face Inference Providers is an aggregation and routing layer rather than one homogeneous inference engine. Its current pricing page grants free users $0.10 in monthly credit, subject to change. PRO users receive $2 monthly, and team or enterprise organizations receive $2 per seat. The free-user allowance is therefore best understood as evaluation credit, not as production capacity.
The strength is choice. A developer can use Hugging Face's client and route requests across supported providers, select a provider explicitly, or use automatic selection. The documentation describes automatic failover when the preferred provider is marked unavailable. Hugging Face also supports custom provider keys: the request can keep the same client shape while billing moves to the external provider.
The cost model is transparent in principle because routed usage is charged at the provider rate. It is not necessarily predictable before a model and provider are fixed. Ten cents could support many low-cost embeddings or only a small number of expensive generations. The monthly reset also behaves differently from Cloudflare's daily recovery: an accidental burst can exhaust the experiment budget for the rest of the month.
Hugging Face states that it does not store routed request bodies or responses and does not use user data for training. It keeps debugging logs for up to 30 days without user data or tokens, according to its security page. However, each downstream inference provider has its own security policy. The routing layer reduces integration friction; it does not collapse all providers into one privacy regime.
When Hugging Face is the best choice
Choose it for model discovery, portability tests, open-model comparisons, and a proof of concept that may later bring a custom provider key. It is useful when the primary question is “which model/provider combination works?” rather than “how can this run all month for zero cost?”
Do not build a public workload around the $0.10 credit. Set a provider explicitly after evaluation, record the model revision and license, and configure a hard spending boundary before enabling pay-as-you-go.
A fair capacity calculation
Free allowance should be translated into completed useful jobs, not marketing units. Use this method:
- Define one representative job, such as classifying a 700-token support message and returning 80 tokens, embedding a 1,000-token document, or transcribing five minutes of audio.
- Measure at least 30 real samples. Record the median and the 95th percentile for input size, output size, latency, retries and provider units.
- Apply the provider's current meter: Gemini quota dimensions, Cloudflare Neurons for the selected model, or Hugging Face's routed dollar cost.
- Reserve 20% to 30% for retries, bursts, schema failures and operational tests.
- Divide the remaining allowance by the 95th-percentile job cost, not the best case.
- Re-run the measurement after changing a model, reasoning level, context window, image size or provider route.
For example, if one Workers AI job consumes 40 Neurons at the 95th percentile, a 10,000-Neuron allocation is not a promise of 250 public requests. After a 25% reserve, the safe software budget is 7,500 Neurons, or about 187 such jobs. The exact number is illustrative; the model's documented conversion and your measurements must supply the real value.
For Gemini, compute three independent headrooms: RPM, TPM and RPD. A batch of long prompts can hit tokens per minute before requests per minute. A lightweight health check repeated too frequently can waste requests per day. The first exhausted dimension is the real limit.
For Hugging Face, measure dollar cost per successful job. Include failed provider calls if they are billable, and distinguish automatic routing from a pinned provider. A low median cost is not enough if occasional long outputs consume most of the monthly credit.
Architecture that survives the next retirement
A free API should sit behind an internal interface with fields your application controls: messages, modality, maximum output, timeout, response schema and safety requirements. Provider-specific model IDs, headers and error shapes belong in adapters.
Use explicit routing policy rather than hidden fallback. A fallback can change accuracy, language behavior, safety filters, data jurisdiction and license obligations. If two providers are approved for the same low-risk task, record that fact and test both. Otherwise fail closed with a clear user message.
Store capacity as observable state. Track successful jobs, rejected jobs, provider units, latency, model/version, reset time and the reason for every fallback. Alert before exhaustion. Never log raw prompts by default, particularly when users can submit confidential material.
Finally, keep a paid escape hatch even if it is disabled. That may be a billing-linked Gemini project, Workers Paid, a pinned Hugging Face provider key, or self-hosted inference. The goal is not to spend money automatically; it is to know the migration path before a free allowance changes.
Decision guide by workload
Multimodal document or media prototype
Start with Gemini if the chosen Free Tier model supports the required modality and the content is compatible with the applicable data terms. Validate the active quota in AI Studio and pin a stable model.
Small edge feature with a daily traffic pattern
Start with Workers AI. Its daily Neuron budget is easier to partition across users and background work. Select a model with documented unit pricing and enforce a per-request maximum.
Open-model evaluation or provider portability test
Start with Hugging Face Inference Providers. Use the monthly credit for evaluation, compare pinned providers, and decide where billing and data processing should ultimately live.
Existing application built on GitHub Models
Treat the endpoint as retired, not temporarily unavailable. Inventory every referenced model, capture required modalities and schemas, then run a paired migration test. GitHub's suggested Azure AI Foundry route may fit teams that want a broad managed catalog, but compare its billing and identity model separately; it is not covered as a free replacement here.
Common mistakes
The first mistake is counting API keys instead of projects or accounts. Multiple keys often share one quota boundary.
The second is comparing allowances without the selected model. Tokens, Neurons and dollars are not interchangeable.
The third is treating “free” as a privacy guarantee. Pricing and data use are separate contractual questions.
The fourth is relying on preview models or `latest` aliases without a retirement plan.
The fifth is publishing a fallback that was never evaluated. Availability is not usefulness if the substitute breaks schemas or changes task accuracy.
The sixth is designing for average consumption. Output tails, retries and bursts are what exhaust small allowances.
Limitations and verification date
This comparison was verified against official public documentation on September 29, 2026. Free allowances, models, quotas and terms can change. Gemini's exact project quotas must be checked in AI Studio. Cloudflare Neuron consumption depends on the model. Hugging Face credit is explicitly subject to change and downstream providers retain separate policies.
The article does not benchmark model quality because the catalogs do not represent one matched set of models or tasks. It compares free-access mechanics, operational predictability, routing and data-governance questions. It is not legal advice.
Final verdict
There is no single successor to GitHub Models.
Gemini is the strongest general choice for capability-rich and multimodal experiments, provided the current project quota and data terms fit. Cloudflare Workers AI offers the clearest recurring free capacity for a bounded serverless feature. Hugging Face offers the best discovery and routing flexibility, but its free-user monthly credit is an evaluation budget rather than durable production capacity.
The durable decision is architectural: measure completed jobs, pin the model, enforce a budget, expose exhaustion cleanly, and keep providers replaceable.
Sources and related FreeAI Tokens guides
Primary evidence comes from GitHub's retirement notice, Google Gemini pricing, billing and rate-limit documentation, Cloudflare Workers AI pricing and data-usage documentation, and Hugging Face Inference Providers pricing, routing and security documentation. All factual capacity statements are attributed to those official sources and dated rather than presented as permanent promises.
For adjacent resources, see Free AI APIs, free AI tiers, free GPU credits, and the FreeAI Tokens editorial hub.