◆ Autopilot Studio
DevNotes

Text to Speech and Speech to Text on Workers AI: Working Code

2026-10-02 · 4 min read

a microphone and a loudspeaker connected by a glowing sound wave that turns into lines of text, dark indigo and orange palette, minimal flat illustration

Speech is cheap on Workers AI: a few neurons for a sentence. It is also full of small surprises. This guide has working code for both directions, and the facts we found by testing them against the live service.

Text to speech with MeloTTS

Bind Workers AI in wrangler.toml ([ai] with binding = "AI") and call the model:

export default {
  async fetch(req, env) {
    const { text, lang = "en" } = await req.json();
    const out = await env.AI.run("@cf/myshell-ai/melotts", { prompt: text, lang });
    // out.audio is a base64 string
    return Response.json({ mimeType: "audio/wav", audio: out.audio });
  },
};

If you want to return the sound itself, not JSON, decode once with the native method:

return new Response(Uint8Array.fromBase64(out.audio), { headers: { "content-type": "audio/wav" } });

It is WAV, not MP3

We expected an MP3. The first characters of the base64 were UklGRl5NAgBXQVZFZm10IBAA..., and UklGR is base64 for the letters RIFF. Decoding the header gives PCM, 16 bits, one channel, 44,100 Hz. That is 88,200 bytes per second: a 7.75-second narration was 683,366 bytes, a 51-character sentence was about 265 KB, and 600 characters of text came back as a 3.9 MB base64 string. Plan your response sizes accordingly, and label what you serve. A small sniffing function removes the guesswork:

function sniffAudio(b64) {
  const b = Uint8Array.fromBase64(b64.slice(0, 16));            // the first 12 bytes
  const s = (i, n) => String.fromCharCode(...b.slice(i, i + n));
  if (s(0, 4) === "RIFF" && s(8, 4) === "WAVE") return "audio/wav";
  if (s(0, 3) === "ID3" || (b[0] === 0xff && (b[1] & 0xe0) === 0xe0)) return "audio/mpeg";
  return "application/octet-stream";
}

We found this the hard way: for a while our own site served narration files with an audio/mpeg header and an .mp3 name while the bytes were WAV. Some players trust the header, so we now detect the format from the bytes and send the right Content-Type.

Languages

lang accepts en, es, fr, jp, kr and zh. Anything else fails with error 8007 and a list of the valid values:

8007: Unsupported language 'pt'. Supported: ['en', 'es', 'fr', 'jp', 'kr', 'zh']

Two traps: Japanese is jp and Korean is kr, not the ISO codes ja and ko, so map them ({ ja: "jp", ko: "kr" }). And there is no Portuguese or German voice, so validate lang yourself and return a clear 400 instead of letting the model error bubble up.

Speech to text with Whisper

const out = await env.AI.run("@cf/openai/whisper-large-v3-turbo", {
  audio: base64Audio,          // base64 of the audio file
  language: "en",              // optional; omit to auto-detect
});
// out = { text, transcription_info: { language, language_probability, duration, duration_after_vad }, word_count, segments, vtt }

We cut the same 7.75-second narration into four files and sent each one:

FileSizeResult
WAV, 44.1 kHz683 KBcorrect text, punctuation kept
WAV, 16 kHz mono248 KBcorrect text, punctuation kept
MP3, 64 kbps63 KBcorrect words, but lower case and no punctuation
FLAC152 KBcorrect text, punctuation kept

All four were transcribed correctly, with language: "en" detected and a duration of 7.75 to 7.85 seconds. The calls took between 1.6 and 4.9 seconds. Two practical lessons: downsample to 16 kHz mono before you send, because the words come out the same and the upload is a third of the size, and do not be surprised if a compressed MP3 returns text without capital letters.

A round trip makes a good health check

Speak a known sentence, then transcribe it and compare. Ours is "The quick brown fox jumps over the lazy dog.", and the live service heard exactly that. A few lines in an admin route tell you that the model, the binding, the base64 handling and the quota are all alive:

const speech = await env.AI.run("@cf/myshell-ai/melotts", { prompt: "The quick brown fox jumps over the lazy dog.", lang: "en" });
const heard = await env.AI.run("@cf/openai/whisper-large-v3-turbo", { audio: speech.audio, language: "en" });
console.log(heard.text);   // "The quick brown fox jumps over the lazy dog."

What it costs

MeloTTS is 18.63 neurons per minute of audio and Whisper large-v3-turbo is 46.63 neurons per minute. The 7.75-second clip above cost about 6 neurons to transcribe, and a full 600-character text-to-speech call costs about 12. The full table, and how to budget the free 10,000 neurons a day, is in Workers AI Free Tier: What Each API Call Really Costs in Neurons.

Or just call ours

Both directions are available as pay-per-call endpoints, /v1/tts and /v1/transcribe, paid in USDC over x402 with no account. The API page has the bodies and prices, and How an AI Agent Pays an x402 API has a client you can paste.

FAQ

Which languages does MeloTTS support on Workers AI?

English (en), Spanish (es), French (fr), Japanese (jp), Korean (kr) and Chinese (zh). The API error lists them: "Unsupported language 'pt'. Supported: ['en', 'es', 'fr', 'jp', 'kr', 'zh']". Note jp and kr, not the ISO codes ja and ko. Portuguese and German are not supported.

What audio format does MeloTTS return?

A WAV file (PCM 16-bit, mono, 44.1 kHz) encoded as base64 in the audio field. We checked the bytes: the base64 starts with UklGR, which decodes to RIFF. That is about 88 KB of audio per second.

Which audio formats does Whisper large-v3-turbo accept?

We tested a 44.1 kHz WAV, a 16 kHz WAV, a 64 kbps MP3 and a FLAC file of the same 7.7 second recording. All four were transcribed correctly with language detection.

How much do speech synthesis and transcription cost in neurons?

MeloTTS costs 18.63 neurons per audio minute and Whisper large-v3-turbo 46.63 neurons per audio minute, so a 7.7 second clip costs about 6 neurons to transcribe.

#workers ai#text to speech#whisper#melotts#cloudflare workers

Found this useful? Tip the studio in crypto

Every EVM chain works. USDC on Base is recommended: fees are a fraction of a cent. No account needed — it goes straight to the creator's wallet.

0x13dd72Fa0E7504790585D92bD98c720f6fD2aBa6

More from DevNotes