The verdict up front: Seed Audio 1.0 is worth trying today if you produce short-form English or Chinese audio — but through a third-party gateway, not the official API. ByteDance's model generates a whole audio scene (dialogue, background music, and sound effects) from a single prompt, and multi-speaker dialogue with built-in music mixing is what separates it from single-purpose TTS or music tools. Three constraints hold it back: output caps at two minutes per generation, it speaks only English and Chinese, and the official API is still invite-only with no published price. Here's the full breakdown.
What Seed Audio 1.0 is
Most audio tools do one job: text-to-speech, or music, or sound effects. Seed Audio 1.0 (full name Doubao-Seed-Audio 1.0) generates all three on a single timeline in one pass. Give it a prompt like "a suspense radio drama in a late-night convenience store," and it returns finished audio with spoken lines, ambient sound, and a music bed already mixed.
It runs in three modes. Text-to-audio (T2A) builds a scene from a written description. Text-and-audio-to-audio (TA2A) takes reference clips plus text, so you can steer voice and style. Straight TTS handles plain narration when you just need clean speech. You pick the mode by what you feed it.
One point is worth clearing up: is this a voice cloning tool? Seed Audio 1.0 can make one voice perform several roles and can imitate the tone of a reference clip, but the output is synthesized rather than bound to a specific real speaker's identity. Treat it as zero-shot voice styling, not a consent-free clone of a named person. It's built on ByteDance's earlier Seed-TTS and Seed-Music research, which is why speech expressiveness and music generation both land in the same model.
The specs that decide if it fits
Before anything else, check whether your use case survives the hard limits. These are the numbers that matter, cross-checked across fal and EvoLink's model listings:
| Spec | Value |
|---|---|
| Max output per generation | 120 seconds (2 minutes) |
| Text prompt limit | ~1,500–2,000 characters (sources differ; confirm in-console) |
| Reference audio | Up to 3 clips, each ≤30 seconds |
| Languages | English and Chinese only |
| Output formats | WAV, MP3, PCM, OGG Opus |
| Sample rates | 48K / 24K / 16K / 8K |
| Controls | Speed, pitch, volume (no SSML) |
| Billing model | Output-duration based |
The two-minute ceiling is the one that bites hardest. A full podcast episode or long narration means splitting the script and stitching segments, which risks tonal drift between chunks. fal's model notes point to a planned update raising single-generation length toward ten minutes, with explicit length control and broader language support — worth waiting for if long-form is your job.
Pricing
There's no official price yet. As of July 2026, Volcengine hasn't published a per-second rate for Seed Audio 1.0 — it's still invite-only on Volcano Ark, and official rates usually land at general availability. Gateways like fal and EvoLink don't hardcode a rate either; they bill by output duration (cost = generated seconds × rate), so retries count too. Ignore any token-based figures floating around — token billing doesn't fit a duration-based audio model.
How to actually access it today
There are two paths, and they're very different in practice.
The official path is Volcano Ark, where Seed Audio 1.0 is invite-only. You apply, wait for access, and use ByteDance's own console — the authoritative source once general availability and pricing land.
The practical path is a third-party gateway. fal hosts the model as bytedance/seed-audio-1.0 with no whitelist: you can hit the endpoint through its API with a standard key, and EvoLink offers similar access. This is how most people can generate with the model right now without waiting on an invite.
The calling pattern is a standard queue-based API — install the client, set your key, submit a prompt, poll for the result. If you've integrated any hosted generative model this way before, there's nothing new to learn here; our fal.ai review covers the developer experience in more detail.
A useful first test is a short multi-speaker scene that exercises dialogue, ambience, and pacing at once — for example, "a 30-second suspense radio drama in a late-night convenience store, English." Keep the first clip under 30 seconds so a failed take costs little while you learn how the model interprets your prompt.
Seed Audio 1.0 vs ElevenLabs, Stable Audio & AudioCraft
The comparison that matters isn't quality-per-clip — it's what each tool is shaped to do.
| Tool | Core strength | Scene mixing | Languages | Voice control |
|---|---|---|---|---|
| Seed Audio 1.0 | Full audio scenes (speech + music + SFX) | Yes, one pass | EN, CN | Zero-shot styling |
| ElevenLabs | Speech and voice cloning | Speech only | 30+ | Strongest cloning |
| Stable Audio | Music and sound generation | Single-track | N/A (music) | Prompt-based |
| AudioCraft | Open-source music/SFX | Single-track | N/A (music) | Prompt-based |
If you need the best pure speech or precise voice cloning across many languages, ElevenLabs still wins. If you need standalone music, Stable Audio or Meta's AudioCraft are built for that. Seed Audio 1.0's edge is combining dialogue, music, and effects into one coherent, editable scene — no other tool here does that in a single call.
Where it wins and where it falls short
Wins: Multi-speaker dialogue and one-pass music-and-effects mixing are the core draw, and the documented editing workflows — extending a clip, in-painting a section, stitching takes, per fal's usage guide — make it a production tool rather than a one-shot generator.
Falls short: The two-minute cap forces stitching for anything long. English and Chinese only rules out most global voice work. It's an offline generation model, so it's wrong for real-time voice agents. And official access is still gated behind an invite.
Is it worth using yet?
If your content is English or Chinese and lives in short-form — dubbing, audiobook snippets, short-video soundtracks, prototype audio dramas — Seed Audio 1.0 is worth trying today through a gateway. Generating the scene in one pass removes the manual step of layering speech, music, and effects yourself.
If you need real-time interaction, a language it doesn't support, or long uninterrupted audio, wait. General availability and the planned length update will change the math, and a real-time model like Qwen Audio 3 is a better fit for live use cases in the meantime. For everyone else, this is a "test on a gateway, keep your current stack running" moment rather than a full switch.
FAQ
Is Seed Audio 1.0 free?
No. There's no free official tier — the Volcano Ark API is invite-only and billed by output duration once you have access. Third-party gateways charge their own metered rates.
Can Seed Audio 1.0 clone my voice?
It can imitate the style of a reference clip and voice multiple roles, but it produces synthesized speech rather than a locked clone of a specific real person. Don't treat it as a drop-in identity-cloning tool.
What languages does Seed Audio 1.0 support?
English and Chinese only at launch. Broader multilingual support is on the roadmap for a planned update, per fal's model notes.
How long can generated audio be?
Up to 120 seconds per generation today. The planned update raises single-generation length toward ten minutes with explicit length control.
Is there an official Seed Audio 1.0 API?
Yes, on Volcengine's Volcano Ark, but it's invite-only for now. For open access without a whitelist, use a gateway like fal that hosts bytedance/seed-audio-1.0.
What's the difference between Seed Audio and Seed Music?
Seed-Music was ByteDance's earlier music-focused research; Seed Audio 1.0 folds that music capability together with speech and sound effects into one scene-generation model.