Workers AI Free Tier: What Each API Call Really Costs in Neurons

Cloudflare Workers AI gives every account 10,000 free neurons a day. That is enough to run a small pay-per-call API, but only if you know what each call costs. The documentation lists prices per model, in different units: tokens, 512-pixel tiles, audio minutes. This post turns them into cost per call and calls per day, using the rates on Cloudflare's pricing page (updated 1 October 2026). Treat the results as estimates: real usage varies with the input.
The rates that matter
| Model | Rate |
|---|---|
| FLUX.1 schnell (image) | 4.80 neurons per 512x512 tile plus 9.60 per step, charged per tile: tiles x (4.80 + steps x 9.60) |
| MeloTTS (speech) | 18.63 neurons per audio minute |
| Whisper large-v3-turbo (transcription) | 46.63 neurons per audio minute |
| BGE-M3 (embeddings) | 1,075 neurons per million input tokens |
| Llama 3.1 8B fp8 fast | 4,119 per million input tokens, 34,868 per million output tokens |
| Llama 3.3 70B fp8 fast | 26,668 per million input tokens, 204,805 per million output tokens |
| M2M100 1.2B (translation) | 31,050 per million tokens, input and output |
What one call costs, and how many fit in a day
If a single endpoint used the whole 10,000 neurons, you could make roughly this many calls:
| Call | Neurons | Calls per day |
|---|---|---|
| Embeddings, one short text (~100 tokens) | 0.1 | ~93,000 |
| Embeddings, 16 texts of 2,000 characters | 8.6 | ~1,160 |
| Speech, 600 characters (~40 s of audio) | 12.4 | ~800 |
| Summary of a 12,000-character text (300 words out) | 22.8 | ~440 |
| Summary of a 3,000-character text | 8.5 | ~1,170 |
| Transcription, 30 s clip | 23.3 | ~430 |
| Transcription, 3 min clip | 139.9 | ~70 |
| Translation, 500 tokens in and 500 out | 31.1 | ~320 |
| Image 1024x1024, FLUX schnell, 4 steps | 172.8 | ~57 |
| Image 1024x1024, FLUX schnell, 6 steps | 249.6 | ~40 |
| Image 1024x1024, FLUX schnell, 8 steps | 326.4 | ~30 |
The image rows are the formula in action, and this is where we got it wrong at first. A 1024 by 1024 picture is four 512-pixel tiles, and the step price applies to every tile: 4 tiles x (4.80 + 4 steps x 9.60) = 172.8 neurons. Reading the price page as "tile cost plus step cost" gives 57.6, three times too low. We only noticed because we checked our real usage (next section): 32 images at 6 steps used exactly 7,987.2 neurons, which is 249.6 each.
A tiny calculator helps when you add a model:
const neurons = {
llm: (inTok, outTok, rateIn, rateOut) => (inTok / 1e6) * rateIn + (outTok / 1e6) * rateOut,
image: (w, h, steps) => Math.ceil(w / 512) * Math.ceil(h / 512) * (4.8 + steps * 9.6), // the step price applies per tile
audioMinutes: (seconds, perMinute) => (seconds / 60) * perMinute,
};
neurons.image(1024, 1024, 4); // 172.8
neurons.llm(3000, 300, 4119, 34868); // ~22.8 (8B summary of a long text)
neurons.audioMinutes(30, 46.63); // ~23.3 (30 s of Whisper)
Choosing the model is the biggest lever
For the same 700-word article (about 2,000 tokens in and 1,000 out), Llama 3.3 70B costs about 258 neurons and Llama 3.1 8B about 43. The 70B is six times more expensive and clearly better for long writing, so we use it for articles and use the 8B for short tasks such as summaries. The output price dominates: 205 of those 258 neurons are the thousand output tokens, so cap max_tokens on every call.
Measure your real usage, do not trust the formula
Cloudflare's GraphQL Analytics API reports neurons per model, and it is the fastest way to catch a wrong estimate. This query returned our usage for the day (it needs an API token that can read the account's analytics):
curl -s https://api.cloudflare.com/client/v4/graphql \
-H "Authorization: Bearer $CF_API_TOKEN" -H 'content-type: application/json' \
--data '{"query":"query { viewer { accounts(filter:{accountTag:\"ACCOUNT_ID\"}) { aiInferenceAdaptiveGroups(limit:20, filter:{datetime_geq:\"2026-10-02T00:00:00Z\", datetime_leq:\"2026-10-03T00:00:00Z\"}) { count sum { totalNeurons } dimensions { modelId } } } } }"}'
Add datetimeHour to dimensions for an hour-by-hour view. Ours showed three things we would not have guessed. FLUX was 32 successful images and 7,987.2 neurons, which is where the 249.6 per image came from. The SDXL-Lightning model, 34 calls, reported 0 neurons. And a single day of testing had pushed us past 10,000 neurons, mostly because every deploy ran a health check that generated a FLUX image. We removed that check from the deploy script and dropped FLUX to 4 steps, the speed it is designed for.
Spend the budget on purpose
- Reserve a share for paying customers. If your own content jobs eat the whole allocation, a paying agent gets an error. Spread scheduled jobs across the day and keep their total well under the cap. By our usage analytics a normal day of ours uses roughly a third of it.
- Fail closed. When the allocation runs out, calls fail until 00:00 UTC. In a pay-per-call API that is safe as long as the paywall settles only after a successful response: the caller sees an error and is not charged.
- Price above cost, even when the cost is zero. On the Paid plan, one 1024-pixel image at 4 steps is 172.8 neurons, about $0.0019. Selling it at $0.02 leaves a wide margin and does not depend on the free tier lasting.
The 10 ms CPU trap
The Workers Free plan allows 10 milliseconds of CPU time per request. Waiting for a model is I/O and does not count. Moving bytes around does. A text-to-speech call for 600 characters returns a few megabytes of audio, and if you decode the model's base64, hold the bytes and encode them again for your JSON response, you pay CPU twice.
We measured the conversion itself in wrangler dev (workerd, compatibility date 2026-03-01) for a 3 MB buffer. The numbers are indicative, since production hardware differs, but the ratios are what matter:
| Method (3 MB) | Time |
|---|---|
Uint8Array.prototype.toBase64() (native) | 1 ms |
Buffer.from(bytes).toString("base64") (nodejs_compat) | 3 ms |
Uint8Array.fromBase64() (native, decoding) | 2 ms |
Buffer.from(b64, "base64") (decoding) | 2 ms |
atob() plus a byte-by-byte loop (decoding) | 8 ms |
String.fromCharCode in chunks plus btoa() (encoding) | 22 ms |
The native paths stay far below the 10 ms limit and the JavaScript loops do not, so we always use the native methods when they exist and fall back to Buffer otherwise. Better still, when the model already returns base64, pass that string through untouched. With the pass-through in place, a 600-character speech request returned a 3.9 MB JSON response on the free plan without hitting the limit. Workers can return large bodies, but you pay for every byte you touch.
Putting it to work
The service behind these numbers sells embeddings, summaries, speech, transcription, translation and images per call. The API page lists the prices, and the payment side is explained in How to Add x402 Payments to a Cloudflare Worker with Hono. If you build your own, start from the table above, multiply by your price, and keep a daily cap you can see.
FAQ
How many neurons does Workers AI give for free?
10,000 neurons per day on both the Free and Paid Workers plans. Usage above that is billed at $0.011 per 1,000 neurons on the Paid plan, and on the Free plan further calls fail until the allocation resets at 00:00 UTC.
How many images can I generate per day on the free tier?
With FLUX.1 schnell at 1024 by 1024 the cost is tiles x (4.80 + steps x 9.60): 172.8 neurons at 4 steps and 249.6 at 6 steps. That is roughly 40 to 57 images a day if nothing else uses the budget. SDXL-Lightning showed up as 0 neurons in our analytics.
Is a bigger language model worth it?
It depends on the task. For the same 1,000-token answer, Llama 3.3 70B costs about six times the neurons of Llama 3.1 8B. We use the 70B for long articles and the 8B for summaries and short tasks.
What is the CPU limit on the Workers Free plan?
10 milliseconds of CPU time per request. Time spent waiting for the model does not count, but copying and re-encoding large payloads does, so avoid decoding and re-encoding big base64 responses.
Found this useful? Tip the studio in crypto
Every EVM chain works. USDC on Base is recommended: fees are a fraction of a cent. No account needed — it goes straight to the creator's wallet.
0x13dd72Fa0E7504790585D92bD98c720f6fD2aBa6More from DevNotes

How to Add x402 Payments to a Cloudflare Worker with Hono
A tested, minimal example of charging per request in USDC on Base with x402, Hono and Cloudflare Workers…

How an AI Agent Pays an x402 API: a 20-Line Client in JS
A working x402 client in JavaScript: sign the USDC payment, read the receipt, cap what your agent can spend…

Cloudflare Cron Triggers Not Firing? Use a Durable Object Alarm
A reliable clock for Cloudflare Workers: a Durable Object alarm that re-arms itself, with a minimum gap and a…

How to List Your x402 API on 402 Index, x402scan and Bazaar
The exact steps to get a pay-per-call x402 API discovered by AI agents: OpenAPI metadata, 402 Index, x402scan…