Text to Speech and Speech to Text on Workers AI: Working Code

Speech is cheap on Workers AI: a few neurons for a sentence. It is also full of small surprises. This guide has working code for both directions, and the facts we found by testing them against the live service.
Text to speech with MeloTTS
Bind Workers AI in wrangler.toml ([ai] with binding = "AI") and call the model:
export default {
async fetch(req, env) {
const { text, lang = "en" } = await req.json();
const out = await env.AI.run("@cf/myshell-ai/melotts", { prompt: text, lang });
// out.audio is a base64 string
return Response.json({ mimeType: "audio/wav", audio: out.audio });
},
};
If you want to return the sound itself, not JSON, decode once with the native method:
return new Response(Uint8Array.fromBase64(out.audio), { headers: { "content-type": "audio/wav" } });
It is WAV, not MP3
We expected an MP3. The first characters of the base64 were UklGRl5NAgBXQVZFZm10IBAA..., and UklGR is base64 for the letters RIFF. Decoding the header gives PCM, 16 bits, one channel, 44,100 Hz. That is 88,200 bytes per second: a 7.75-second narration was 683,366 bytes, a 51-character sentence was about 265 KB, and 600 characters of text came back as a 3.9 MB base64 string. Plan your response sizes accordingly, and label what you serve. A small sniffing function removes the guesswork:
function sniffAudio(b64) {
const b = Uint8Array.fromBase64(b64.slice(0, 16)); // the first 12 bytes
const s = (i, n) => String.fromCharCode(...b.slice(i, i + n));
if (s(0, 4) === "RIFF" && s(8, 4) === "WAVE") return "audio/wav";
if (s(0, 3) === "ID3" || (b[0] === 0xff && (b[1] & 0xe0) === 0xe0)) return "audio/mpeg";
return "application/octet-stream";
}
We found this the hard way: for a while our own site served narration files with an audio/mpeg header and an .mp3 name while the bytes were WAV. Some players trust the header, so we now detect the format from the bytes and send the right Content-Type.
Languages
lang accepts en, es, fr, jp, kr and zh. Anything else fails with error 8007 and a list of the valid values:
8007: Unsupported language 'pt'. Supported: ['en', 'es', 'fr', 'jp', 'kr', 'zh']
Two traps: Japanese is jp and Korean is kr, not the ISO codes ja and ko, so map them ({ ja: "jp", ko: "kr" }). And there is no Portuguese or German voice, so validate lang yourself and return a clear 400 instead of letting the model error bubble up.
Speech to text with Whisper
const out = await env.AI.run("@cf/openai/whisper-large-v3-turbo", {
audio: base64Audio, // base64 of the audio file
language: "en", // optional; omit to auto-detect
});
// out = { text, transcription_info: { language, language_probability, duration, duration_after_vad }, word_count, segments, vtt }
We cut the same 7.75-second narration into four files and sent each one:
| File | Size | Result |
|---|---|---|
| WAV, 44.1 kHz | 683 KB | correct text, punctuation kept |
| WAV, 16 kHz mono | 248 KB | correct text, punctuation kept |
| MP3, 64 kbps | 63 KB | correct words, but lower case and no punctuation |
| FLAC | 152 KB | correct text, punctuation kept |
All four were transcribed correctly, with language: "en" detected and a duration of 7.75 to 7.85 seconds. The calls took between 1.6 and 4.9 seconds. Two practical lessons: downsample to 16 kHz mono before you send, because the words come out the same and the upload is a third of the size, and do not be surprised if a compressed MP3 returns text without capital letters.
A round trip makes a good health check
Speak a known sentence, then transcribe it and compare. Ours is "The quick brown fox jumps over the lazy dog.", and the live service heard exactly that. A few lines in an admin route tell you that the model, the binding, the base64 handling and the quota are all alive:
const speech = await env.AI.run("@cf/myshell-ai/melotts", { prompt: "The quick brown fox jumps over the lazy dog.", lang: "en" });
const heard = await env.AI.run("@cf/openai/whisper-large-v3-turbo", { audio: speech.audio, language: "en" });
console.log(heard.text); // "The quick brown fox jumps over the lazy dog."
What it costs
MeloTTS is 18.63 neurons per minute of audio and Whisper large-v3-turbo is 46.63 neurons per minute. The 7.75-second clip above cost about 6 neurons to transcribe, and a full 600-character text-to-speech call costs about 12. The full table, and how to budget the free 10,000 neurons a day, is in Workers AI Free Tier: What Each API Call Really Costs in Neurons.
Or just call ours
Both directions are available as pay-per-call endpoints, /v1/tts and /v1/transcribe, paid in USDC over x402 with no account. The API page has the bodies and prices, and How an AI Agent Pays an x402 API has a client you can paste.
FAQ
Which languages does MeloTTS support on Workers AI?
English (en), Spanish (es), French (fr), Japanese (jp), Korean (kr) and Chinese (zh). The API error lists them: "Unsupported language 'pt'. Supported: ['en', 'es', 'fr', 'jp', 'kr', 'zh']". Note jp and kr, not the ISO codes ja and ko. Portuguese and German are not supported.
What audio format does MeloTTS return?
A WAV file (PCM 16-bit, mono, 44.1 kHz) encoded as base64 in the audio field. We checked the bytes: the base64 starts with UklGR, which decodes to RIFF. That is about 88 KB of audio per second.
Which audio formats does Whisper large-v3-turbo accept?
We tested a 44.1 kHz WAV, a 16 kHz WAV, a 64 kbps MP3 and a FLAC file of the same 7.7 second recording. All four were transcribed correctly with language detection.
How much do speech synthesis and transcription cost in neurons?
MeloTTS costs 18.63 neurons per audio minute and Whisper large-v3-turbo 46.63 neurons per audio minute, so a 7.7 second clip costs about 6 neurons to transcribe.
Found this useful? Tip the studio in crypto
Every EVM chain works. USDC on Base is recommended: fees are a fraction of a cent. No account needed — it goes straight to the creator's wallet.
0x13dd72Fa0E7504790585D92bD98c720f6fD2aBa6More from DevNotes

How to Add x402 Payments to a Cloudflare Worker with Hono
A tested, minimal example of charging per request in USDC on Base with x402, Hono and Cloudflare Workers…

How an AI Agent Pays an x402 API: a 20-Line Client in JS
A working x402 client in JavaScript: sign the USDC payment, read the receipt, cap what your agent can spend…

Cloudflare Cron Triggers Not Firing? Use a Durable Object Alarm
A reliable clock for Cloudflare Workers: a Durable Object alarm that re-arms itself, with a minimum gap and a…

How to List Your x402 API on 402 Index, x402scan and Bazaar
The exact steps to get a pay-per-call x402 API discovered by AI agents: OpenAPI metadata, 402 Index, x402scan…