AIREITER

SeedRealtime: ByteDance's Audio-Visual Full-Duplex LLM

Last Updated: 2026-08-05 06:58:06

ByteDance's SeedRealtime, launched August 5, 2026, is a native audio-visual full-duplex LLM. Its real differentiator is not "real-time voice" (several of those exist) but one end-to-end model that jointly fuses audio and video, with no external module deciding whose turn it is. The catch for developers: at launch there is no public API or pricing, and it ships inside ByteDance's own products.

SeedRealtime launch announcement on ByteDance's Seed site, headlined "An Audio-Visual Full-Duplex LLM" and dated August 5, 2026.

What SeedRealtime Is

Built by ByteDance's Seed team, it fuses audio, video, and text in a single architecture, perceiving, understanding, deciding, and responding over continuous multimodal streams rather than shuttling data between stages.

The framing ByteDance uses is "watch, listen, and speak at once": real-time interaction in which what is heard and what is seen jointly inform each response. SeedRealtime sits in the Seed model lineup alongside Seed2.1 (the LLM family that powers Doubao), Seedance (video generation), Seedream (image generation), Seed Audio 1.0 (audio generation), and Seed GR-RL (robotics).

The SeedRealtime model card on ByteDance's Seed models page, shown alongside Seed2.1, Seedance 2.5, Seedream 5.0 Pro, Seed Audio 1.0, and Seed GR-RL.

It does not fit the usual boxes: not a TTS model (it generates speech but also understands it), not a vision model (it sees, but in service of conversation), and not merely another real-time voice bot.

Why "Full-Duplex" Is the Whole Story

"Full-duplex" means the model can listen and respond at the same time, continuously modeling who is speaking and when to chime in, the way a real conversation works. SeedRealtime's claim is that it reaches this by unifying sound, vision, timing, and expression inside one end-to-end model rather than bolting modules together.

Real-time AI interaction has been stuck between two older paradigms. Cascaded systems chain ASR (speech-to-text) to an LLM or VLM to TTS (text-to-speech); they are modular and debuggable, but each handoff adds latency and strips information, and tone, overlap, accent, and visual context tend to die at the ASR stage. End-to-end voice models remove those handoffs and sound more natural, but most still lean on an external voice-activity detector (VAD) to decide whose turn it is, which keeps them effectively half-duplex: one question, one answer.

SeedRealtime folds perception, understanding, decision-making, and expression into a single model that runs them in parallel over continuous audio-visual streams, with no external VAD deciding turns.

ApproachHow it worksDuplexTurn-takingMain weakness
Cascaded (ASR + LLM/VLM + TTS)Modules chained in sequenceHalfFixed endpointingLatency and information loss at each handoff
End-to-end + external VADOne model, external turn detectorHalfExternal VADStill one-question-one-answer
SeedRealtimeUnified audio-visual end-to-end modelFull (claimed)Internal, multimodalNewest; capabilities are vendor-described

The hard part is that video, unlike speech, has no natural pauses: the model must constantly decide which object to attend to, whose voice matters, and whether this is the moment to speak, without reacting to background noise.

The Three Capabilities

ByteDance groups SeedRealtime's behavior into three capabilities: joint audio-visual understanding, proactive interaction, and natural conversational timing. It demonstrates each with concrete scenarios rather than benchmark tables.

ByteDance's launch post for SeedRealtime, dated August 5, 2026, listing the three core breakthroughs.

Joint audio-visual understanding

The model uses the live scene to resolve ambiguous speech. When someone says "how do I do this," SeedRealtime combines the screen, gestures, gaze, and prior actions to decide what "this" is.

In ByteDance's demo, a user introduces four friends at dinner one by one, and SeedRealtime matches names to faces, then keeps each person's voice tied to their identity while the group plans a trip around conflicting preferences: one wants the beach, one hates the heat, one is allergic to seafood. At a Sichuan restaurant it reads a Chinese-only menu from the camera, recommends dishes in English, and explains why "fish-fragrant pork" contains no fish.

Proactive interaction

The model speaks up on its own when the scene changes, rather than waiting to be asked. Touring a museum, a user says "remind me when you see the bronze tiger-devouring-a-deer stand," and SeedRealtime watches the camera feed until that artifact appears, then offers the reminder plus context on the piece.

Watching someone pour whole coffee beans straight into a portafilter, it cuts in: grind them into a fine powder first, and after extraction it suggests shortening the next pull by two to three seconds based on the crema's color. Tool calls such as lookups and bookings are woven into these responses rather than handled in a separate step.

Natural conversational timing

SeedRealtime decides when to speak, pause, or stay quiet from multimodal context, and it resists false triggers. In a crowded Beijing airport it ignores a companion's offhand mention of "Old Li's flight," but when asked it recalls departure-board information that has already scrolled off screen, then goes online for the baggage-carousel location. At home, with a parent on a background phone call, it stays focused on correcting a child's English pronunciation.

ByteDance's end-to-end human evaluation reports that, versus cascaded models, SeedRealtime roughly halves audio-visual conversational pacing issues: fewer mid-sentence cutoffs, sluggish responses after pauses, and false triggers from noise, plus a higher rate of single turns completed smoothly and fully. These are vendor-run evaluations, not independent benchmarks.

SeedRealtime vs GPT-4o Realtime, Gemini Live, and Cascaded Stacks

The practical question is how SeedRealtime differs from the real-time systems developers can already call. The honest answer: most existing "realtime" APIs center on audio, while SeedRealtime's distinguishing claim is jointly modeling audio and video in full-duplex.

OpenAI's Realtime API and Google's Gemini Live consumer feature (with its developer Live API) are low-latency, real-time conversation interfaces, but they center on audio (speech in, speech out), with vision handled separately and turn-taking that still depends on endpointing or VAD. SeedRealtime's positioning is the combination: native audio-visual fusion, full-duplex timing decided inside the model, proactive prompting, and tool use folded into the conversation.

ModelNative modalitiesDuplexTurn-takingProactive / tool-usePublic API
SeedRealtimeAudio + video + text (fused)Full (claimed)Internal, multimodalYes, woven inNo (at launch)
GPT-4o Realtime APIAudio + textReal-time, VAD-managedExternal endpointingVia function toolsYes
Gemini Live / Live APIAudio + text (camera in consumer app)Real-time, VAD-managedExternal endpointingLimitedYes (Live API, audio)
Cascaded stackText (audio via ASR/TTS)HalfFixedCustomBuild-your-own

SeedRealtime's advantages come from ByteDance's own demos and evaluations; there is no published head-to-head against GPT-4o realtime or Gemini Live. Treat the comparison above as architectural positioning, not measured superiority.

Can Developers Use SeedRealtime Yet?

As of its August 5, 2026 launch, SeedRealtime is not available as a developer API. ByteDance says it is "fully rolled out," but that refers to deployment inside the company's own products, not an open endpoint anyone can call.

There is no public API, no pricing page, and no developer documentation for SeedRealtime at launch. ByteDance's commercial API arm is Volcengine (火山引擎), which hosts the Doubao model family: Seed 2.1 Pro and Turbo LLMs, Seed-ASR, Seed-TTS, and others. SeedRealtime is not listed among them at launch, and third-party relays offer no real-time audio-visual substitute.

The realistic near-term options for teams that need a real-time endpoint today remain audio-centric: OpenAI's Realtime API, Google's Live API, or a self-built cascaded stack, with SeedRealtime as the architecture to watch. ByteDance has not named a date for a Volcengine or BytePlus real-time API.

FAQ

Is SeedRealtime available to the public?

It is deployed inside ByteDance's own products as of the August 5, 2026 launch, but there is no public-facing app or endpoint named SeedRealtime that you can sign up for.

Does SeedRealtime have an API?

Not at launch. ByteDance has not released API access, pricing, or developer docs; the likely future home is Volcengine's Doubao API family, which does not yet list SeedRealtime.

How is SeedRealtime different from GPT-4o realtime?

GPT-4o's Realtime API is audio-centric: speech in, speech out. SeedRealtime's claim is jointly modeling audio and video in a single full-duplex model that also decides its own timing and acts proactively. There is no independent head-to-head yet.

Is SeedRealtime free?

There is no standalone product or API with a price tag at launch, so "free vs paid" does not yet apply; it is a capability inside ByteDance's products.

What does "full-duplex" mean for an AI model?

It means the model can listen and respond at the same time and continuously model conversational state, deciding on its own when to speak or pause, rather than locking into one-question-one-answer turns.

Related reading