If you want natural-sounding narration without sending every sentence to a paid API, Kokoro 82M TTS is one of the most practical open models to test first. The useful qualification is that “82M” does not mean “a universal ElevenLabs replacement”: Kokoro is strongest with preset voices, clean narration, and local batch jobs, while cloning, dramatic control, and managed reliability point elsewhere.
The short decision
Kokoro-82M is an open-weight text-to-speech model with 82 million parameters. The current Hugging Face model card lists v1.0 as a January 27, 2025 release with eight languages and 54 voices, under Apache 2.0 weights: official model card. The official inference library is hexgrad/kokoro.
Choose it when you need private or offline narration, preset voices are acceptable, and your workload is large enough for API charges to matter. Choose ElevenLabs when you need voice cloning, a managed studio, or expressive commercial delivery. Choose Qwen3-TTS when multilingual coverage and cloning matter more than a very small footprint.
Local setup: the path that actually works
The official examples use Python, kokoro, soundfile, PyTorch, and espeak-ng; the pipeline returns audio at 24 kHz. On Linux, the basic installation is:
pip install -q "kokoro>=0.9.4" soundfile
sudo apt-get install espeak-ng
A minimal English example is:
from kokoro import KPipeline
import soundfile as sf
pipeline = KPipeline(lang_code="a")
text = "Kokoro is a small local text to speech model."
for index, (_, _, audio) in enumerate(
pipeline(text, voice="af_heart", speed=1)
):
sf.write(f"kokoro-{index}.wav", audio, 24000)
The repository documents language codes for American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese. Japanese and Mandarin require the corresponding Misaki extras. The selected language code must match the voice.
Windows is not a one-command install: the project README sends users to the espeak-ng MSI release. Apple Silicon users can try PyTorch MPS with PYTORCH_ENABLE_MPS_FALLBACK=1. These details matter more than the model’s parameter count when a first run fails.
CPU, GPU, and real-time speed: what is verified
CPU inference is supported, but neither the official model card nor the official library gives one universal real-time factor. Speed depends on the runtime, processor, language, text splitting, and whether the first pass includes model loading.
Community measurements show why broad claims are risky. In a detailed comparison, @analogalok reported about 0.7 seconds to generate 21 seconds of audio on a GPU, with roughly 0.9 GB of VRAM in that setup. That is useful evidence of a low-memory GPU path, not a promise for every laptop.
“I can run it comfortably using my CPU and the resulting audio is very pleasant to listen to.” — @rasmus1610, X
For your own hardware, benchmark three separate numbers:
- Time from process start to the first audio segment.
- Time after the model and voice are warm.
- Audio duration divided by warm-generation time: the real-time factor.
For a batch narrator, the third number is the one that affects throughput. For an interactive assistant, first-token latency and streaming behavior matter more. Do not quote a GPU result as a CPU result, and do not call a 10-second sample “production throughput.”
Voices, languages, and the important limitations
The v1.0 model card lists eight languages and 54 voices, a substantial expansion over v0.19’s one language and 10 voices. The library’s documented language codes include:
| Code | Language |
|---|---|
a | American English |
b | British English |
e | Spanish |
f | French |
h | Hindi |
i | Italian |
j | Japanese |
p | Brazilian Portuguese |
z | Mandarin Chinese |
Kokoro uses preset voices; it is not a reference-audio voice-cloning system. It also relies on the Misaki grapheme-to-phoneme stack and espeak-ng fallback paths. Names, abbreviations, numbers, and mixed-language text deserve a test pass before a long export.
The Apache 2.0 model license is attractive for commercial deployment, but “the model is Apache 2.0” should not be read as a blanket answer for every voice sample, dataset, or third-party dependency. Keep the model card, voice files, and dependency licenses in your release checklist.
Kokoro vs ElevenLabs vs Qwen3-TTS
A blind listening comparison is only meaningful when all three systems use the same text, voice brief, normalization, output level, and evaluator pool. This article does not claim a controlled blind-listening winner; the official Kokoro materials do not publish such a test. The more defensible comparison is the workflow trade-off:
| Decision factor | Kokoro-82M | ElevenLabs | Qwen3-TTS |
|---|---|---|---|
| Deployment | Local/open weights | Hosted API and studio | Local/open model family |
| Model license/source | Apache 2.0 weights | Provider terms | Official model cards list Apache 2.0 |
| Preset voices | Yes | Large hosted voice library | CustomVoice and other variants |
| Voice cloning | No | Available on supported plans/features | Base variant supports reference cloning |
| Languages | 8 listed for v1.0 | Model-dependent; official API pricing varies by model | 10 listed languages |
| Best fit | Private batch narration | Managed, expressive production | Multilingual or cloning experiments |
| Main cost | Hardware or hosted inference | Character/credit billing | Hardware or hosting |
ElevenLabs’ official API pricing page currently lists about $0.05 per 1,000 characters for Flash/Turbo and $0.10 per 1,000 characters for v2 Multilingual and v3: API pricing. Qwen3-TTS’s official model card lists Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, plus cloning-oriented variants: Qwen3-TTS model card.
The practical result is not “Kokoro sounds better.” It is that Kokoro removes recurring API dependency and can keep text on your machine, while ElevenLabs buys convenience and Qwen3-TTS buys a wider multilingual/cloning feature set at a larger local-model footprint.
Batch narration cost: the numbers you can defend
Local Kokoro has no per-character model charge. Your marginal cost is electricity, storage, and whatever machine or cloud instance runs inference. If you already own the machine, the incremental software cost is effectively zero; if you rent a GPU or CPU, use the provider’s hourly rate multiplied by wall-clock generation time.
For an API comparison, the Kokoro model card recorded April 2025 hosted examples below $1 per million input characters, including $0.65 on Replicate and $0.80 on DeepInfra. Those are dated reference points, not current price guarantees. At the same text volume, ElevenLabs’ official rates imply:
| Workload | Kokoro local model charge | ElevenLabs Flash/Turbo | ElevenLabs v2 Multilingual/v3 |
|---|---|---|---|
| 100,000 characters | $0 locally* | $5 | $10 |
| 1,000,000 characters | $0 locally* | $50 | $100 |
| 10,000,000 characters | $0 locally* | $500 | $1,000 |
\*Excludes electricity, hardware, hosting, and engineering time. ElevenLabs figures use the official $0.05/$0.10 per 1,000-character rates and ignore plan discounts, credits, taxes, and overage rules. The table is a budgeting model, not a promise of final billing.
That makes Kokoro especially compelling for repeatable narration pipelines: subtitles, documentation, accessibility audio, and internal training clips. It is less compelling when the human cost of checking pronunciation and stitching segments outweighs the API bill.
FAQ
Can Kokoro run without a GPU?
Yes. CPU inference is supported, but the official project does not promise a universal real-time factor. Measure warm throughput on the exact CPU and language you plan to ship.
Does Kokoro clone voices?
No. Kokoro is a preset-voice system. Use a cloning-capable model such as an appropriate Qwen3-TTS variant when speaker identity is a core requirement.
Is Kokoro free for commercial use?
The Kokoro-82M model card lists Apache 2.0 weights, which is a permissive license. Review the exact voice, dependency, and distribution terms before shipping a commercial product.
Which languages does Kokoro support?
The v1.0 model card lists eight languages; the official library documents American and British English plus Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese. Language-specific packages may be required.
Is Kokoro better than ElevenLabs?
Not in every dimension. Kokoro is the better fit for local, private, preset-voice batch narration; ElevenLabs is the safer fit for managed infrastructure, voice cloning, and expressive production workflows.
The practical choice
Start with Kokoro if your acceptance test is: “Can this machine produce clean narration privately, at the required throughput, with one of the available voices?” Run that test on names, dates, numbers, long paragraphs, and the languages you actually need.
If it passes, the economics are simple: you replace recurring character charges with machine time. If it fails on pronunciation, speaker identity, or emotional direction, moving to ElevenLabs or Qwen3-TTS is not a quality defeat; it is paying for a capability Kokoro intentionally does not prioritize.
For the commercial end of the same decision, see the ElevenLabs Eleven v3 API guide.