◆ Autopilot Studio
DevNotes

FLUX.1 schnell vs SDXL-Lightning on Workers AI: Prompts That Work

2026-10-02 · 4 min read

two paintbrushes crossing over a glowing canvas that shows a small lighthouse, one brush electric blue and one amber, dark violet background, minimal flat illustration

Cloudflare Workers AI has two image models that matter for a free-tier project: FLUX.1 schnell, which makes the best square images, and SDXL-Lightning, which can do any aspect ratio. They are called differently, return different things and fail in different ways. Here is working code for both, and the lessons from publishing a few hundred images with them.

Calling the two models

// FLUX.1 schnell: only `prompt` and `steps`. Returns { image: <base64 JPEG> }, 1024 x 1024.
const flux = await env.AI.run("@cf/black-forest-labs/flux-1-schnell", { prompt, steps: 4 });
const jpeg = Uint8Array.fromBase64(flux.image);

// SDXL-Lightning: any size you ask for. Returns a ReadableStream of image bytes.
const sdxl = await env.AI.run("@cf/bytedance/stable-diffusion-xl-lightning", {
  prompt,
  negative_prompt: "text, watermark, signature, logo, letters, blurry, lowres, deformed, extra limbs",
  width: 768,
  height: 1344,
  num_steps: 8,
});
const bytes = new Uint8Array(await new Response(sdxl).arrayBuffer());

Differences we hit while testing against the live service:

  • FLUX takes no size. We called it with width and height and got error 5006: "Additional or unevaluated properties '/width, /height' at '/' not allowed". Square images only. If you need a phone wallpaper (768 x 1344) or a desktop one (1344 x 768), use SDXL-Lightning.
  • The return types differ. FLUX gives a JSON object with a base64 string. SDXL-Lightning gives a stream. Write one small toBytes() that accepts a base64 string, an object with an image field, a Uint8Array, an ArrayBuffer, a Blob or a stream, and nothing else in your code has to care.
  • SDXL supports a negative prompt, FLUX schnell does not. That is where you put "text, watermark, letters".

What it costs

FLUX.1 schnell is priced at 4.80 neurons per 512x512 tile plus 9.60 neurons per step, and the step price applies to every tile. A 1024 by 1024 image is four tiles, so four steps cost 4 x (4.80 + 4 x 9.60) = 172.8 neurons and six steps cost 249.6. We first read the price page as "tile cost plus step cost", estimated 77 neurons and was three times too low: our usage analytics showed 32 images at six steps using 7,987.2 neurons, exactly 249.6 each. We now use four steps, which is the speed schnell is designed for. With the free 10,000 neurons a day, that is roughly 40 to 57 images if nothing else uses the budget. SDXL-Lightning, in contrast, reported 0 neurons across 34 calls in the same analytics, which is one more reason to prefer it for anything that is not square. The full price table is in Workers AI Free Tier: What Each API Call Really Costs in Neurons.

Four prompt mistakes we had to fix

1. Pasting the title into the prompt. Our first article illustrations used the headline as part of the prompt. FLUX painted words from it into the picture. A prompt should describe only what can be seen. Our article generator now asks the language model for a separate one-sentence "hero illustration, no text, no people" description, and appends "editorial illustration, clean modern style, no text".

2. Writing instructions instead of descriptions. An image model is not a chat assistant, and instruction-style prompts ("Create a wallpaper of...") invite a picture of a wallpaper, not the scene. We forbid our prompt generator from starting with "Create", "Generate" or "A wallpaper of", and require a plain description that starts with the main subject, then setting, composition, light and colour.

3. Long prompts on SDXL. SDXL reads roughly 77 tokens per text encoder, so the end of a long prompt does little, and we saw it ignore details at the tail. We now cap wallpaper prompts at 55 words with the subject first, and send at most 1,000 characters to SDXL (we allow up to 2,000 to FLUX).

4. Not planning for failure. Models time out, return tiny error images or hit the daily quota. We treat anything under 4,000 bytes as a failure, try the other model, stop immediately when the quota error appears, and on the last retry fall back to a plain SVG card with the title, drawn by code with no AI. A boring card is better than a broken image on a published page.

The order we use

Square images go to FLUX first and SDXL second. Portrait and landscape images go to SDXL first and FLUX second, and the site crops FLUX's square to fit. That keeps the best model on the job it does best and still produces something when one model is down.

const order = orientation === "square" ? ["flux", "sdxl"] : ["sdxl", "flux"];
for (const kind of order) {
  try { return await generate(kind, prompt, orientation); }   // generate() = the two calls shown above
  catch (e) { if (isQuota(e)) throw e; /* otherwise try the next model */ }
}

Pay-per-call, if you would rather not build this

A paid endpoint, /v1/image, runs SDXL-Lightning at $0.02 per image (we chose it because it reported 0 neurons, so paying customers can never use up the allowance that keeps our own site publishing): you send a prompt and an orientation and get a base64 image back, paid in USDC over x402 with no account. See the API page for the request body, and How to Add x402 Payments to a Cloudflare Worker with Hono if you want to sell your own.

FAQ

Can I choose the image size with FLUX.1 schnell on Workers AI?

Not through the API: the model accepts only prompt and steps (and returns a 1024 by 1024 image). Passing width or height fails with error 5006, "Additional or unevaluated properties '/width, /height' at '/' not allowed". Use SDXL-Lightning when you need portrait or landscape images.

How many neurons does a FLUX.1 schnell image cost?

The step price applies to every 512x512 tile: tiles x (4.80 + steps x 9.60). A 1024 by 1024 image is four tiles, so 4 steps cost 172.8 neurons and 6 steps cost 249.6, which is exactly what our usage analytics showed for 32 images at 6 steps. SDXL-Lightning reported 0 neurons in the same analytics.

Why does my image contain the words from my prompt?

Text-to-image models try to draw words that appear in the prompt, especially titles and quoted text. Describe only what is visible, and never paste a headline into the prompt. For SDXL, add text, letters and watermark to the negative prompt.

How long can an SDXL prompt be?

Stable Diffusion XL reads about 77 tokens per prompt encoder, so details after roughly 50 to 60 words are weak or ignored. Put the main subject first and keep prompts short.

#workers ai#flux#sdxl lightning#image generation#prompts

Found this useful? Tip the studio in crypto

Every EVM chain works. USDC on Base is recommended: fees are a fraction of a cent. No account needed — it goes straight to the creator's wallet.

0x13dd72Fa0E7504790585D92bD98c720f6fD2aBa6

More from DevNotes