Voice Compiler turns articles into spoken audio. The on-device system voice works with no setup. The cloud voices below sound more natural — each needs its own free API key that you paste into Settings → Provider API keys. Keys are stored on the server and never shown again.
Tip: if you mainly want to read along word-for-word, use Google Cloud, ElevenLabs, or Amazon Polly.
You need a Google account and a Google Cloud project with billing enabled. The free tier covers 1 million characters a month, so normal reading stays free — but a card on file is required by Google.
Note: this is not the same key as Gemini. A Cloud Text-to-Speech key does not work for Gemini, and vice versa.
The quickest key to get — no billing required for the free tier. It comes from Google AI Studio, which is separate from Google Cloud above.
Heads-up: Gemini voices sound great, but the read-along highlight is estimated, so a word can drift slightly from the audio. For exact highlighting use Google Cloud or ElevenLabs.
The free tier has a monthly character limit; long articles use it up quickly.
Inworld is a low-cost cloud TTS. German voices appear from its live catalog once a key is set. Word highlighting is approximate — Voice Compiler estimates timing (turn on neural CTC alignment for more accuracy); Inworld's own high-accuracy timestamps are English-only.
Models: 1.5 Mini is the cheapest, 1.5 Max the highest quality, TTS-2 the newest with the widest language coverage.
Speechify returns word-level timestamps with every request, so read-along is word-exact without the estimated fallback and without a second billed call. Mind the model: Simba 3.2 is the flagship but is English only — pick Simba 3.0 or Simba Multilingual for German.
sk_) and copy it.Beyond the free allowance, paid use starts at the Starter plan: $10 per month including one million characters, then $10 per additional million. Unlike some providers there is no pay-as-you-go tier without a monthly plan. Speechify's terms also ask that audio played to other people is disclosed as AI-generated.
Polly uses AWS credentials: an access key ID, a secret access key, and a region. You paste all three in Settings → Provider API keys (select Amazon Polly). Polly gives exact word-by-word highlighting and the widest German voice range (de-DE, de-AT, de-CH).
DescribeVoices and
SynthesizeSpeech — nothing else is needed).eu-central-1
(Frankfurt) or us-east-1.Pricing: Standard voices are the cheapest but sound more robotic; Neural voices cost more but sound natural. Speech Marks (word timing) are billed at the same rate as the audio — no extra charge.
Azure needs a subscription key and the region of your Speech resource, both pasted in Settings → Provider API keys (select Azure AI Speech). Azure has a very large German neural catalog and a recurring free tier (F0). Note: word highlighting is approximate — the plain voice REST API returns no word timing, so Voice Compiler estimates it (turn on neural CTC alignment for more accuracy).
westeurope).The model picker also offers MAI (preview) — Microsoft's newer, more expressive voices (including German). MAI voices use the same key but are billed like Azure's HD voices and are not covered by the free tier. When you pick an Azure Neural voice, the narration dialog shows how many of your 500,000 free monthly characters are left and warns before a generation would exceed them.
Each provider row in Settings → Provider API keys has an optional Custom rate (USD per 1M characters) field. Set it only if your plan or discount differs from the list price — for example an ElevenLabs subscription or an AWS volume rate. Leave it empty to use the provider's list/API price. This changes the in-app cost estimate and the cost ledger only; it never affects what the provider actually bills you.
Test asks the provider to speak a single word using your key. A green ✓ Connection OK means the key works. A red message usually means the key is wrong, billing isn’t enabled, or the provider is out of quota. The badge next to each provider shows whether a key is currently stored.
While the reader is open on a document with generated narration, these keys control playback (they are ignored while typing in a text field or with a dialog open):
Space or K — play / pause← / → or J / L — skip back / forward
by your configured seek amount (Settings → Playback)H — jump back to the start of the article+ / - — speed up / slow down in 0.25× steps? — show this list inside the appThe same keys drive device Read aloud, where skipping moves by paragraph instead of by seconds.
Curious what happens to your article text and API keys? See the Privacy Policy.