◆ Autopilot Studio
DevNotes

Workers AI Free Tier: What Each API Call Really Costs in Neurons

2026-10-02 · 6 min read

a stylized gauge made of glowing neurons filling up like a fuel tank, cloud and circuit lines in the background, dark blue and green palette, minimal

Cloudflare Workers AI gives every account 10,000 free neurons a day. That is enough to run a small pay-per-call API, but only if you know what each call costs. The documentation lists prices per model, in different units: tokens, 512-pixel tiles, audio minutes. This post turns them into cost per call and calls per day, using the rates on Cloudflare's pricing page (updated 1 October 2026). Treat the results as estimates: real usage varies with the input.

The rates that matter

ModelRate
FLUX.1 schnell (image)4.80 neurons per 512x512 tile plus 9.60 per step, charged per tile: tiles x (4.80 + steps x 9.60)
MeloTTS (speech)18.63 neurons per audio minute
Whisper large-v3-turbo (transcription)46.63 neurons per audio minute
BGE-M3 (embeddings)1,075 neurons per million input tokens
Llama 3.1 8B fp8 fast4,119 per million input tokens, 34,868 per million output tokens
Llama 3.3 70B fp8 fast26,668 per million input tokens, 204,805 per million output tokens
M2M100 1.2B (translation)31,050 per million tokens, input and output

What one call costs, and how many fit in a day

If a single endpoint used the whole 10,000 neurons, you could make roughly this many calls:

CallNeuronsCalls per day
Embeddings, one short text (~100 tokens)0.1~93,000
Embeddings, 16 texts of 2,000 characters8.6~1,160
Speech, 600 characters (~40 s of audio)12.4~800
Summary of a 12,000-character text (300 words out)22.8~440
Summary of a 3,000-character text8.5~1,170
Transcription, 30 s clip23.3~430
Transcription, 3 min clip139.9~70
Translation, 500 tokens in and 500 out31.1~320
Image 1024x1024, FLUX schnell, 4 steps172.8~57
Image 1024x1024, FLUX schnell, 6 steps249.6~40
Image 1024x1024, FLUX schnell, 8 steps326.4~30

The image rows are the formula in action, and this is where we got it wrong at first. A 1024 by 1024 picture is four 512-pixel tiles, and the step price applies to every tile: 4 tiles x (4.80 + 4 steps x 9.60) = 172.8 neurons. Reading the price page as "tile cost plus step cost" gives 57.6, three times too low. We only noticed because we checked our real usage (next section): 32 images at 6 steps used exactly 7,987.2 neurons, which is 249.6 each.

A tiny calculator helps when you add a model:

const neurons = {
  llm: (inTok, outTok, rateIn, rateOut) => (inTok / 1e6) * rateIn + (outTok / 1e6) * rateOut,
  image: (w, h, steps) => Math.ceil(w / 512) * Math.ceil(h / 512) * (4.8 + steps * 9.6),     // the step price applies per tile
  audioMinutes: (seconds, perMinute) => (seconds / 60) * perMinute,
};

neurons.image(1024, 1024, 4);                    // 172.8
neurons.llm(3000, 300, 4119, 34868);             // ~22.8 (8B summary of a long text)
neurons.audioMinutes(30, 46.63);                 // ~23.3 (30 s of Whisper)

Choosing the model is the biggest lever

For the same 700-word article (about 2,000 tokens in and 1,000 out), Llama 3.3 70B costs about 258 neurons and Llama 3.1 8B about 43. The 70B is six times more expensive and clearly better for long writing, so we use it for articles and use the 8B for short tasks such as summaries. The output price dominates: 205 of those 258 neurons are the thousand output tokens, so cap max_tokens on every call.

Measure your real usage, do not trust the formula

Cloudflare's GraphQL Analytics API reports neurons per model, and it is the fastest way to catch a wrong estimate. This query returned our usage for the day (it needs an API token that can read the account's analytics):

curl -s https://api.cloudflare.com/client/v4/graphql \
  -H "Authorization: Bearer $CF_API_TOKEN" -H 'content-type: application/json' \
  --data '{"query":"query { viewer { accounts(filter:{accountTag:\"ACCOUNT_ID\"}) { aiInferenceAdaptiveGroups(limit:20, filter:{datetime_geq:\"2026-10-02T00:00:00Z\", datetime_leq:\"2026-10-03T00:00:00Z\"}) { count sum { totalNeurons } dimensions { modelId } } } } }"}'

Add datetimeHour to dimensions for an hour-by-hour view. Ours showed three things we would not have guessed. FLUX was 32 successful images and 7,987.2 neurons, which is where the 249.6 per image came from. The SDXL-Lightning model, 34 calls, reported 0 neurons. And a single day of testing had pushed us past 10,000 neurons, mostly because every deploy ran a health check that generated a FLUX image. We removed that check from the deploy script and dropped FLUX to 4 steps, the speed it is designed for.

Spend the budget on purpose

  • Reserve a share for paying customers. If your own content jobs eat the whole allocation, a paying agent gets an error. Spread scheduled jobs across the day and keep their total well under the cap. By our usage analytics a normal day of ours uses roughly a third of it.
  • Fail closed. When the allocation runs out, calls fail until 00:00 UTC. In a pay-per-call API that is safe as long as the paywall settles only after a successful response: the caller sees an error and is not charged.
  • Price above cost, even when the cost is zero. On the Paid plan, one 1024-pixel image at 4 steps is 172.8 neurons, about $0.0019. Selling it at $0.02 leaves a wide margin and does not depend on the free tier lasting.

The 10 ms CPU trap

The Workers Free plan allows 10 milliseconds of CPU time per request. Waiting for a model is I/O and does not count. Moving bytes around does. A text-to-speech call for 600 characters returns a few megabytes of audio, and if you decode the model's base64, hold the bytes and encode them again for your JSON response, you pay CPU twice.

We measured the conversion itself in wrangler dev (workerd, compatibility date 2026-03-01) for a 3 MB buffer. The numbers are indicative, since production hardware differs, but the ratios are what matter:

Method (3 MB)Time
Uint8Array.prototype.toBase64() (native)1 ms
Buffer.from(bytes).toString("base64") (nodejs_compat)3 ms
Uint8Array.fromBase64() (native, decoding)2 ms
Buffer.from(b64, "base64") (decoding)2 ms
atob() plus a byte-by-byte loop (decoding)8 ms
String.fromCharCode in chunks plus btoa() (encoding)22 ms

The native paths stay far below the 10 ms limit and the JavaScript loops do not, so we always use the native methods when they exist and fall back to Buffer otherwise. Better still, when the model already returns base64, pass that string through untouched. With the pass-through in place, a 600-character speech request returned a 3.9 MB JSON response on the free plan without hitting the limit. Workers can return large bodies, but you pay for every byte you touch.

Putting it to work

The service behind these numbers sells embeddings, summaries, speech, transcription, translation and images per call. The API page lists the prices, and the payment side is explained in How to Add x402 Payments to a Cloudflare Worker with Hono. If you build your own, start from the table above, multiply by your price, and keep a daily cap you can see.

FAQ

How many neurons does Workers AI give for free?

10,000 neurons per day on both the Free and Paid Workers plans. Usage above that is billed at $0.011 per 1,000 neurons on the Paid plan, and on the Free plan further calls fail until the allocation resets at 00:00 UTC.

How many images can I generate per day on the free tier?

With FLUX.1 schnell at 1024 by 1024 the cost is tiles x (4.80 + steps x 9.60): 172.8 neurons at 4 steps and 249.6 at 6 steps. That is roughly 40 to 57 images a day if nothing else uses the budget. SDXL-Lightning showed up as 0 neurons in our analytics.

Is a bigger language model worth it?

It depends on the task. For the same 1,000-token answer, Llama 3.3 70B costs about six times the neurons of Llama 3.1 8B. We use the 70B for long articles and the 8B for summaries and short tasks.

What is the CPU limit on the Workers Free plan?

10 milliseconds of CPU time per request. Time spent waiting for the model does not count, but copying and re-encoding large payloads does, so avoid decoding and re-encoding big base64 responses.

#workers ai#cloudflare#neurons#free tier#pricing

Found this useful? Tip the studio in crypto

Every EVM chain works. USDC on Base is recommended: fees are a fraction of a cent. No account needed — it goes straight to the creator's wallet.

0x13dd72Fa0E7504790585D92bD98c720f6fD2aBa6

More from DevNotes