AIREITER

Kokoro 82M TTS: Local CPU Test, Voices, and Cost

Last Updated: 2026-09-28 01:20:48

If you want natural-sounding narration without sending every sentence to a paid API, Kokoro 82M TTS is one of the most practical open models to test first. The useful qualification is that “82M” does not mean “a universal ElevenLabs replacement”: Kokoro is strongest with preset voices, clean narration, and local batch jobs, while cloning, dramatic control, and managed reliability point elsewhere.

The short decision

Kokoro-82M is an open-weight text-to-speech model with 82 million parameters. The current Hugging Face model card lists v1.0 as a January 27, 2025 release with eight languages and 54 voices, under Apache 2.0 weights: official model card. The official inference library is hexgrad/kokoro.

Choose it when you need private or offline narration, preset voices are acceptable, and your workload is large enough for API charges to matter. Choose ElevenLabs when you need voice cloning, a managed studio, or expressive commercial delivery. Choose Qwen3-TTS when multilingual coverage and cloning matter more than a very small footprint.

Local setup: the path that actually works

The official examples use Python, kokoro, soundfile, PyTorch, and espeak-ng; the pipeline returns audio at 24 kHz. On Linux, the basic installation is:

pip install -q "kokoro>=0.9.4" soundfile
sudo apt-get install espeak-ng

A minimal English example is:

from kokoro import KPipeline
import soundfile as sf

pipeline = KPipeline(lang_code="a")
text = "Kokoro is a small local text to speech model."

for index, (_, _, audio) in enumerate(
    pipeline(text, voice="af_heart", speed=1)
):
    sf.write(f"kokoro-{index}.wav", audio, 24000)

The repository documents language codes for American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese. Japanese and Mandarin require the corresponding Misaki extras. The selected language code must match the voice.

Windows is not a one-command install: the project README sends users to the espeak-ng MSI release. Apple Silicon users can try PyTorch MPS with PYTORCH_ENABLE_MPS_FALLBACK=1. These details matter more than the model’s parameter count when a first run fails.

CPU, GPU, and real-time speed: what is verified

CPU inference is supported, but neither the official model card nor the official library gives one universal real-time factor. Speed depends on the runtime, processor, language, text splitting, and whether the first pass includes model loading.

Community measurements show why broad claims are risky. In a detailed comparison, @analogalok reported about 0.7 seconds to generate 21 seconds of audio on a GPU, with roughly 0.9 GB of VRAM in that setup. That is useful evidence of a low-memory GPU path, not a promise for every laptop.

“I can run it comfortably using my CPU and the resulting audio is very pleasant to listen to.” — @rasmus1610, X

For your own hardware, benchmark three separate numbers:

  1. Time from process start to the first audio segment.
  2. Time after the model and voice are warm.
  3. Audio duration divided by warm-generation time: the real-time factor.

For a batch narrator, the third number is the one that affects throughput. For an interactive assistant, first-token latency and streaming behavior matter more. Do not quote a GPU result as a CPU result, and do not call a 10-second sample “production throughput.”

Voices, languages, and the important limitations

The v1.0 model card lists eight languages and 54 voices, a substantial expansion over v0.19’s one language and 10 voices. The library’s documented language codes include:

CodeLanguage
aAmerican English
bBritish English
eSpanish
fFrench
hHindi
iItalian
jJapanese
pBrazilian Portuguese
zMandarin Chinese

Kokoro uses preset voices; it is not a reference-audio voice-cloning system. It also relies on the Misaki grapheme-to-phoneme stack and espeak-ng fallback paths. Names, abbreviations, numbers, and mixed-language text deserve a test pass before a long export.

The Apache 2.0 model license is attractive for commercial deployment, but “the model is Apache 2.0” should not be read as a blanket answer for every voice sample, dataset, or third-party dependency. Keep the model card, voice files, and dependency licenses in your release checklist.

Kokoro vs ElevenLabs vs Qwen3-TTS

A blind listening comparison is only meaningful when all three systems use the same text, voice brief, normalization, output level, and evaluator pool. This article does not claim a controlled blind-listening winner; the official Kokoro materials do not publish such a test. The more defensible comparison is the workflow trade-off:

Decision factorKokoro-82MElevenLabsQwen3-TTS
DeploymentLocal/open weightsHosted API and studioLocal/open model family
Model license/sourceApache 2.0 weightsProvider termsOfficial model cards list Apache 2.0
Preset voicesYesLarge hosted voice libraryCustomVoice and other variants
Voice cloningNoAvailable on supported plans/featuresBase variant supports reference cloning
Languages8 listed for v1.0Model-dependent; official API pricing varies by model10 listed languages
Best fitPrivate batch narrationManaged, expressive productionMultilingual or cloning experiments
Main costHardware or hosted inferenceCharacter/credit billingHardware or hosting

ElevenLabs’ official API pricing page currently lists about $0.05 per 1,000 characters for Flash/Turbo and $0.10 per 1,000 characters for v2 Multilingual and v3: API pricing. Qwen3-TTS’s official model card lists Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, plus cloning-oriented variants: Qwen3-TTS model card.

The practical result is not “Kokoro sounds better.” It is that Kokoro removes recurring API dependency and can keep text on your machine, while ElevenLabs buys convenience and Qwen3-TTS buys a wider multilingual/cloning feature set at a larger local-model footprint.

Batch narration cost: the numbers you can defend

Local Kokoro has no per-character model charge. Your marginal cost is electricity, storage, and whatever machine or cloud instance runs inference. If you already own the machine, the incremental software cost is effectively zero; if you rent a GPU or CPU, use the provider’s hourly rate multiplied by wall-clock generation time.

For an API comparison, the Kokoro model card recorded April 2025 hosted examples below $1 per million input characters, including $0.65 on Replicate and $0.80 on DeepInfra. Those are dated reference points, not current price guarantees. At the same text volume, ElevenLabs’ official rates imply:

WorkloadKokoro local model chargeElevenLabs Flash/TurboElevenLabs v2 Multilingual/v3
100,000 characters$0 locally*$5$10
1,000,000 characters$0 locally*$50$100
10,000,000 characters$0 locally*$500$1,000

\*Excludes electricity, hardware, hosting, and engineering time. ElevenLabs figures use the official $0.05/$0.10 per 1,000-character rates and ignore plan discounts, credits, taxes, and overage rules. The table is a budgeting model, not a promise of final billing.

That makes Kokoro especially compelling for repeatable narration pipelines: subtitles, documentation, accessibility audio, and internal training clips. It is less compelling when the human cost of checking pronunciation and stitching segments outweighs the API bill.

FAQ

Can Kokoro run without a GPU?

Yes. CPU inference is supported, but the official project does not promise a universal real-time factor. Measure warm throughput on the exact CPU and language you plan to ship.

Does Kokoro clone voices?

No. Kokoro is a preset-voice system. Use a cloning-capable model such as an appropriate Qwen3-TTS variant when speaker identity is a core requirement.

Is Kokoro free for commercial use?

The Kokoro-82M model card lists Apache 2.0 weights, which is a permissive license. Review the exact voice, dependency, and distribution terms before shipping a commercial product.

Which languages does Kokoro support?

The v1.0 model card lists eight languages; the official library documents American and British English plus Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese. Language-specific packages may be required.

Is Kokoro better than ElevenLabs?

Not in every dimension. Kokoro is the better fit for local, private, preset-voice batch narration; ElevenLabs is the safer fit for managed infrastructure, voice cloning, and expressive production workflows.

The practical choice

Start with Kokoro if your acceptance test is: “Can this machine produce clean narration privately, at the required throughput, with one of the available voices?” Run that test on names, dates, numbers, long paragraphs, and the languages you actually need.

If it passes, the economics are simple: you replace recurring character charges with machine time. If it fails on pronunciation, speaker identity, or emotional direction, moving to ElevenLabs or Qwen3-TTS is not a quality defeat; it is paying for a capability Kokoro intentionally does not prioritize.

For the commercial end of the same decision, see the ElevenLabs Eleven v3 API guide.