Eleven v3 left alpha on February 2, 2026, and the model ID your code calls in the ElevenLabs API is eleven_v3. ElevenLabs and fal.ai both charge $0.10 per 1,000 characters, so unit price cannot decide between them; billing structure, concurrency, and voice access do. GA also fixed real problems: per the GA announcement, structured-text errors fell from 15.3% to 4.9% across a 27-category, 8-language benchmark. Plan around two hard limits: 5,000 characters per request, and a model its own maker still calls unsuitable for real-time use.
What the ElevenLabs Eleven v3 API gives you
The Eleven v3 API exposes four surfaces: create speech, stream speech, and separate create-dialogue and stream-dialogue endpoints for multi-speaker audio. The model speaks 70+ languages, accepts inline audio tags for emotion and delivery, and at GA became stable enough that ElevenLabs' own testing preferred it over the alpha build 72% of the time.
Before you trust v3 with structured text, check the category detail in that benchmark: chemical formulas dropped from a 45.6% error rate to 0.6%, phone numbers from 16.9% to 0.6%, and ISBNs to zero. Geographic coordinates remain the weak spot at 17.5%, and math expressions sit at 6.9% (fine for narration, risky for coordinate-heavy content).
The model table below shows where v3 sits against its siblings, per the official models reference:
| Model | Model ID | Languages | Max chars/request | Stated latency | Price / 1k chars |
|---|---|---|---|---|---|
| Eleven v3 | eleven_v3 | 70+ | 5,000 (~5 min) | ~250–300 ms (tier) | $0.10 |
| Eleven v3 Conversational | via Text-to-Dialogue WebSocket | 70+ | n/a | ~280 ms | $0.10 |
| Multilingual v2 | eleven_multilingual_v2 | 29 | 10,000 (~10 min) | ~250–300 ms | $0.10 |
| Flash v2.5 | eleven_flash_v2_5 | 32 | 40,000 (~40 min) | ~75 ms | $0.05 |
One documentation discrepancy to know about: the pricing page groups "Multilingual v2/v3" under a 40,000-character tier limit, but the models reference caps Eleven v3 at 5,000 characters per request. Treat 5,000 as the real number for v3; it is the model-specific figure, and long-form pipelines that assume 40k will fail.
Pricing: official ElevenLabs vs fal.ai
Unit pricing is identical at $0.10 per 1,000 characters, and only Flash v2.5 is cheaper in the family at $0.05/1k. The differences that matter sit in billing and parallelism:
| ElevenLabs official | fal.ai | |
|---|---|---|
| Unit price (Eleven v3) | $0.10 / 1k chars | $0.10 / 1k chars |
| Billing | Subscription plans with included credits, or pay-as-you-go | Pure usage-based; "no seat licenses, no subscriptions, no minimums" |
| Free tier | 10k Multilingual chars/mo (20k Flash) | None stated on the model page |
| Example plan | Creator $22/mo (first month $11), 220k Multilingual chars | n/a |
| Top plan | Business $990/mo, 9.9M Multilingual chars | n/a |
| Concurrency | Plan-based: 2 concurrent Multilingual requests on Free, 15–30 on Scale/Business, custom on Enterprise | Serverless queue; no per-account cap documented on the model page |
| Queue penalty | Exceed the limit and requests queue, "typically" adding ~50 ms | Queue API with webhooks for batch jobs |
| Controls (documented schema, plus user reports) | Voice by ID; stability (users report 3 positions on v3) | Voice, stability, similarity_boost, speed, language_code, apply_text_normalization, seed, timestamps, output_format |
| Startup program | 12 months free, 33M characters, higher concurrency | n/a |
The official route prices concurrency through plan tiers: a Free key making parallel v3 calls gets 2 concurrent requests, which throttles batch work hard. fal routes all jobs through a serverless queue with webhook callbacks aimed at exactly that batch case. Neither platform's pricing documentation states whether whitespace or bracketed audio tags count toward billed characters; budget as if they do.
If you only need v3 occasionally and already run other models through fal, the platform case is covered in our fal.ai review.
Audio tags: directing emotion inside the text
Audio tags are bracketed instructions placed inline in the text itself; the model reads them as performance direction, not as spoken words. A minimal example from the fal model page:
[slowly] Back then... [chuckles] we had no phones.
[whispers] Just dirt roads and [coughs] big dreams.
[sad] Then it happened.
| Category | Working examples |
|---|---|
| Emotions | [happy], [sad], [excited], [angry], [sarcastically], [nervous], [happily] |
| Delivery | [whispers], [shouting], [shouts], [slowly], [quickly], [softly] |
| Non-verbal sounds | [laughs], [chuckles], [sighs], [gasps], [coughs], [gulps], [clears throat], [applause] |
| Accents | [british accent], [southern accent], [strong canadian accent] |
Three rules from the official prompting docs govern tag behavior: no canonical tag list exists (tags are natural-language direction, and ElevenLabs says more effective tags likely exist beyond the published examples, so testing beats memorizing), SSML break tags are unsupported (pacing comes from punctuation, ellipses, and line breaks), and effectiveness varies by voice and context. That last rule is where things break:
"V3 is amazing, certainly the most human sounding" "it still feels like a Beta release to me" — u/ConcertNeat8147, both lines from the same post, r/ElevenLabs field report on v3
In the same thread, u/poundingCode, a software engineer with two decades of experience, reports that adding emotional tags can occasionally flip the speaker's accent entirely. An ElevenLabs staff account replied that "these are all things we're currently working on."
Workarounds reported by creators in that same thread (community experience, not official guidance):
- Warm-up sentence. Prepend a throwaway sentence that sets tone and pitch, then cut it in post; voices often "settle" a few seconds into a generation.
- Trailing
end.Append a line break and the word "end." after your text to stop trailing-word cutoffs, then trim it from the audio. - End sentences with ellipses. Sample cutoffs at word endings reportedly drop when sentences close on "...".
- Max the stability setting. The v3 stability control has three positions; the highest reduces the voice drift at the start of generations.
Limits that shape the integration
- The 5,000-character cap is the big one. A request tops out around 5 minutes of audio, so audiobook and long-form pipelines must chunk text: split at sentence or paragraph boundaries, not mid-clause, and keep tags inside the chunk they direct. Multilingual v2 allows 10,000 characters if you need fewer, larger calls.
- Concurrency is a plan feature on the official API. Per the models reference: Free keys get 2 concurrent Multilingual-tier requests (4 Flash), Scale and Business get 15–30, Enterprise negotiates custom limits, and queueing past the limit "typically" adds ~50 ms. The same page offers ElevenLabs' capacity heuristic: a concurrency limit of 5 supports roughly 100 simultaneous audio broadcasts in a conversational workload; the penalty is cheap per call, brutal at batch scale if you under-provision.
- fal documents no maximum input size at all. Its Eleven v3 page specifies the endpoint, parameters, and formats, but not an input ceiling, queue retention, or retry behavior. Nothing stops you from submitting long text; nothing guarantees it either.
- Timestamps split across platforms. fal's input schema has a
timestampsboolean that returns per-word alignment data, useful for subtitles and lip-sync. On the official API, v3 dialogue output ships without timestamps, a gap developers have flagged publicly (r/ElevenLabs: "v3 API, no timestamps?"); fal's parameter is the documented route for alignment data today.
- Real-time is still off the table. ElevenLabs says v3 is unsuitable for real-time and conversational use due to higher latency and variable consistency. The sanctioned options: Flash v2.5 at ~75 ms for agents, or Eleven v3 Conversational at ~280 ms over WebSocket when expressiveness must survive into an interactive context. A real-time v3 variant is on the roadmap ("we're working on a real-time version") but was not shipped as of the GA announcement.
- Commercial use. During the alpha, developers asked publicly whether v3 output was commercially usable; the GA announcement opened the model across all platforms, and fal labels its endpoint for commercial use outright.
Your first request: official SDK and fal client
The official path, from the quickstart verbatim:
pip install elevenlabs
from elevenlabs import ElevenLabs
client = ElevenLabs(api_key="YOUR_ELEVENLABS_API_KEY") # or ELEVENLABS_API_KEY env var
audio = client.text_to_speech.convert(
voice_id="JBFqnCBsd6RMkjVDRZzb", # "George"
text="[whispers] The first move is what sets everything in motion.",
model_id="eleven_v3",
output_format="mp3_44100_128",
)
The fal path uses fal-client against the fal.run endpoint:
pip install fal-client
import fal_client
result = fal_client.subscribe(
"fal-ai/elevenlabs/tts/eleven-v3",
arguments={
"text": "[whispers] The first move is what sets everything in motion.",
"voice": "Rachel", # default voice
"stability": 0.5, # 0–1, lower = more expressive variation
"timestamps": True, # per-word alignment data
},
)
print(result["audio"]["url"]) # hosted MP3 URL
Both platforms explicitly warn against shipping API keys in client-side code, so proxy requests through your own server. For streaming, the official API has a stream-speech endpoint, and fal offers fal.stream plus a queue API with webhooks for batch jobs.
Which endpoint should you call?
| Your situation | Call |
|---|---|
| You need your own cloned voices (PVC/IVC) or designed voices by ID | Official: custom voices are referenced by ID in your workspace; fal's page only documents preset voice names |
| You already pay for an ElevenLabs plan with included credits | Official: included characters are sunk cost, and each fal call is new spend |
| Batch audiobook/dubbing jobs, no subscription, webhooks on completion | fal: its pay-per-use queue and webhook flow target exactly this case |
| One API key across image, video, and TTS models in one stack | fal: v3 rides the same client as the rest of your fal models |
| Voice agents or anything latency-bound | Neither v3 TTS endpoint: use Flash v2.5 (~75 ms) or Eleven v3 Conversational (~280 ms) |
| Enterprise needs (SOC 2, HIPAA BAA, custom rate limits, IP allowlisting) | Official: documented Enterprise plan features on the pricing page; fal's model page does not state which certifications cover its endpoint |
FAQ
What is the model ID for Eleven v3?
eleven_v3, passed as model_id on the standard text-to-speech endpoints. Multi-speaker dialogue uses the separate Text-to-Dialogue endpoints rather than a model parameter.
Does the Eleven v3 API support streaming?
Yes. The official API offers a stream-speech endpoint that returns audio as it generates, and fal exposes fal.stream for the same model. Streaming does not make v3 real-time-safe; official guidance still assigns conversational use to Flash v2.5 or Eleven v3 Conversational.
What is the character limit for Eleven v3?
5,000 characters per request, roughly 5 minutes of audio, per the official models reference. The pricing page's 40,000-character figure applies to the Multilingual tier grouping, not to v3 specifically.
How much does the Eleven v3 API cost?
$0.10 per 1,000 characters on both routes: $1.00 per 10,000-character story, $100 per 1M-character audiobook. Plan credits offset official costs (Creator $22/mo includes 220k Multilingual characters); fal has no subscription.
Can I use cloned voices with Eleven v3?
On the official API, yes: professional, cloned, and designed voices are referenced by voice ID. Compatibility varies in practice; r/ElevenLabs users report existing Professional Voice Clones sometimes sounding like different speakers under v3, while PVCs whose creators flag V3 compatibility tend to hold up best. fal's schema accepts a voice name or ID but only documents preset voices, so private ElevenLabs voice IDs are undocumented territory there.
Is Eleven v3 good for real-time voice agents?
No. ElevenLabs explicitly lists real-time and conversational use as unsuitable for v3 due to higher latency and variable consistency. Use Flash v2.5 (~75 ms, 32 languages) or Eleven v3 Conversational (~280 ms, 70+ languages, audio-tag control).
Why does Eleven v3 sound different through the API than on the website?
Voice previews are not representative of what you get with your own text, a recurring complaint in the r/ElevenLabs field reports, and v3 output varies noticeably between generations. Test each candidate voice against a real script chunk, at your production stability setting, before committing.
Related reading: Qwen Audio 3 TTS API guide · Replicate pricing and alternatives · fal.ai review 2026