Workers AI Error 4006: What Happens at the Daily Free Limit
You can run into the daily limit in a second, with a single call, and the result is not a slowdown but a wall: every Workers AI call fails until midnight UTC. We hit it on 2 October 2026, and the useful part is what we learned about how it behaves and how to build a service that survives it.
The message
Calling any model after the allowance was gone returned this error, the same text from four different models:
4006: you have used up your daily free allocation of 10,000 neurons, please upgrade to Cloudflare's Workers Paid plan if you would like to continue usage.
The four were a Llama 3.1 8B text model, the BGE-M3 embeddings model, FLUX.1 schnell and SDXL-Lightning. The calls failed fast, in 30 to 200 milliseconds, and a failed call costs nothing. The error reaches your code as a thrown exception from env.AI.run(), so the way to recognize it is by its text, for example:
const isQuota = (e) => /daily free allocation|4006|used up your daily|exceeded.*neurons/i.test(String((e && e.message) || e));
What we observed
- Every model is blocked, including the free one. SDXL-Lightning had reported 0 neurons over 34 calls in our analytics, and it still failed with 4006 once the account allowance was spent. A model that does not consume the allowance does not escape the wall either.
- You can overshoot. The usage analytics, read through the method in Check Your Workers AI Usage with the GraphQL Analytics API, showed 11,067 neurons on the day, about 10 percent above 10,000, by the time the errors were visible. Metering is not instantaneous, so a service that is "almost at the limit" may keep working for a little while and then stop abruptly.
- It renews at 00:00 UTC. That is what the pricing page says, and the schedule we built around it is below.
Handling it in a background job
Retrying a 4006 is pointless until midnight, so treat it as fatal and stop at once. In a Cloudflare Workflow that means throwing NonRetryableError so the instance ends immediately and does not burn its retries, as shown in Cloudflare Workflows Free Plan: Retries, Steps and Fatal Errors. Our production cycle also remembers the failure and skips the remaining jobs of that run, so one quota error costs one cheap failed call and not twenty.
Nothing needs repairing afterwards. The next scheduled run after 00:00 UTC finds a full allowance and carries on: a day with no new content is a better outcome than a half-written article.
Handling it in a paid API
If you sell AI calls with x402, the behavior you want is simple: do not charge, and tell the client when to come back. An x402 server settles a payment only when the response status is below 400, so any error response is free for the buyer. A 503 with a Retry-After header is the clearest signal, because well-behaved clients wait and retry:
function aiFail(c, e, code) {
if (isQuota(e)) {
const now = new Date();
const midnight = Date.UTC(now.getUTCFullYear(), now.getUTCMonth(), now.getUTCDate() + 1);
const secs = Math.max(60, Math.ceil((midnight - now.getTime()) / 1000));
return c.json(
{ error: "daily_ai_allowance_exhausted", detail: "The free daily AI allowance is used up and renews at 00:00 UTC. You were not charged.", retry_after_seconds: secs },
503,
{ "Retry-After": String(secs) },
);
}
return c.json({ error: code, detail: String((e && e.message) || e) }, 502);
}
Use it in every handler's catch. We unit-test three cases: a flagged fatal error, a raw 4006 message from a call that goes straight to env.AI.run, and an ordinary failure, which must stay a 502. The payment side of this is explained in How to Add x402 Payments to a Cloudflare Worker with Hono.
How we got there, and how to avoid it
Our day went over the limit because of testing, not production. Forced runs, repeated health checks and one deploy check that generated a FLUX image each time (about 250 neurons per deploy) added up to roughly four times what the scheduled work needs. The fixes were boring and effective: a text-only deploy check, FLUX at four steps instead of six, and a habit of reading the real usage before planning more experiments.
Three rules for anyone running on the free allowance:
- Know your per-call costs from measurement, not from the price table.
- Keep a reserve: do not let tests and extras use what the scheduled work needs.
- Make every failure path return quickly and cleanly, because one day you will hit the wall.
FAQ
What does Workers AI return when the free allowance is used up?
A thrown error whose message starts with "4006: you have used up your daily free allocation of 10,000 neurons, please upgrade to Cloudflare's Workers Paid plan if you would like to continue usage." We got the same message from a Llama model, BGE-M3, FLUX.1 schnell and SDXL-Lightning.
When does the allowance renew?
At 00:00 UTC every day, according to the Workers AI pricing page.
Are models that cost 0 neurons still available after the limit?
No. After the limit, SDXL-Lightning, which reported 0 neurons in our usage analytics, also returned error 4006. The block applies to every model on the account.
Is the limit enforced at exactly 10,000 neurons?
Not in our observation. Our analytics showed 11,067 neurons used on the day when errors started to be visible, about 10 percent above the allowance, so metering is not instantaneous. Plan for the limit, not for an exact number.
Found this useful? Tip the studio in crypto
Every EVM chain works. USDC on Base is recommended: fees are a fraction of a cent. No account needed — it goes straight to the creator's wallet.
0x13dd72Fa0E7504790585D92bD98c720f6fD2aBa6More from DevNotes

How to Add x402 Payments to a Cloudflare Worker with Hono
A tested, minimal example of charging per request in USDC on Base with x402, Hono and Cloudflare Workers…

How an AI Agent Pays an x402 API: a 20-Line Client in JS
A working x402 client in JavaScript: sign the USDC payment, read the receipt, cap what your agent can spend…

Cloudflare Cron Triggers Not Firing? Use a Durable Object Alarm
A reliable clock for Cloudflare Workers: a Durable Object alarm that re-arms itself, with a minimum gap and a…

How to List Your x402 API on 402 Index, x402scan and Bazaar
The exact steps to get a pay-per-call x402 API discovered by AI agents: OpenAPI metadata, 402 Index, x402scan…