Skip to main content
Released by ByteDance Seed · August 5, 2026

SeedRealtime: ByteDance's Audio-Visual Full-Duplex LLM

SeedRealtime is a native audio-visual full-duplex model from ByteDance Seed that watches, listens, and speaks at the same time — fusing audio, video, and text in one unified end-to-end architecture for real-time multimodal interaction.

Watching · Listening · Speaking
LIVE
Watching · Listening · Speaking
Which case is the bronze galloping horse in?
It just came into view — the third case on your left.
Last updated: August 2026Reading time: 8 minutes
Table of Contents
OVERVIEW

What Is SeedRealtime?

SeedRealtime is a native audio-visual full-duplex large language model built by ByteDance Seed and publicly released on August 5, 2026. Instead of chaining separate modules for speech recognition, visual understanding, language processing, and speech generation, SeedRealtime fuses audio, video, and text inside one unified architecture.

The model continuously processes audio, video, and text streams, so it can listen, watch, and respond at the same time. Perception, understanding, decision-making, and response generation all run in a single end-to-end model rather than a cascaded pipeline.

This design moves real-time AI beyond turn-based voice modes toward omni-modal natural interaction: continuous conversation over live audio and video, with timing awareness and visual context built in.

Developer
ByteDance Seed
Model type
Audio-visual full-duplex LLM
Released
August 5, 2026
Availability
Doubao App
KEY FEATURES

SeedRealtime Key Features

What a native audio-visual full-duplex model can do that cascaded pipelines cannot

Native Audio-Visual Fusion

Audio and visual streams are understood jointly in one model. Visual context helps resolve homophones and speech ambiguity, and the model tracks how a scene changes over time.

Full-Duplex Real-Time Conversation

SeedRealtime listens and speaks simultaneously over continuous multimodal streams, with streaming generation and low-latency responses instead of rigid turn-taking.

Natural Conversational Timing

The model judges when to speak, detects pauses, distinguishes unfinished speech from a completed turn, and handles interruptions at natural moments.

Multi-Speaker & Noise Handling

SeedRealtime distinguishes different speakers in multi-speaker scenes, filters background and bystander chatter, and avoids false triggers from unrelated conversations.

Visual Scene Understanding

Persistent environmental awareness: object recognition, scene change detection, gesture following, and linking what is in view right now to earlier conversation.

Proactive Interaction & Tool Calling

The model can speak up on its own when a scene changes or a target appears, and can invoke tools during responses for reading, navigation, and device-operation assistance.

ARCHITECTURE

How SeedRealtime Works

One end-to-end model instead of a listen, transcribe, think, then speak pipeline

01

Unified End-to-End Architecture

Perception, understanding, decision-making, and expression run in parallel inside a single model. With no module handoffs, SeedRealtime avoids the context loss and error accumulation of cascaded systems.

02

Continuous Audio-Visual Stream Processing

Chunked audio-visual input lets SeedRealtime consume live audio and video as continuous streams while tracking conversational state and timing in real time.

03

Low-Latency Streaming Generation

Streaming generation combined with quantization and inference optimization keeps responses fast enough for natural, low-latency full-duplex conversation.

COMPARISON

SeedRealtime vs Cascaded Multimodal Models

Why a unified full-duplex model responds faster and more naturally than a multi-stage pipeline

Why Cascaded Pipelines Fall Short

Listen
Transcribe
Think
Speak
  • Separate speech recognition, vision-language, and text-to-speech modules chained together
  • Latency accumulates at every module handoff
  • Context and information are lost between stages
  • Turn-taking depends on external voice activity detection rules

What a Unified Full-Duplex Model Changes

SeedRealtime
Watch, listen, and speak — in parallel
  • Native audio-video integration with parallel multimodal processing
  • Lower conversational latency with streaming generation
  • More natural interruption handling and response timing
  • Better speaker tracking, scene tracking, and proactive responses
DimensionCascaded PipelineSeedRealtime
ArchitectureSeparate speech recognition, vision-language, and text-to-speech modulesSingle unified end-to-end model
Audio-video understandingProcessed in separate stages; context is lost at each handoffJointly understood over continuous multimodal streams
LatencyLatency accumulates at every module handoffLow-latency streaming generation
Turn-takingExternal voice activity detection rulesBuilt-in timing awareness and interruption handling
Proactive interactionResponds only when explicitly triggeredSpeaks up when scenes change or targets appear
~50%

ByteDance reports that audio-video conversational timing issues drop by about half compared with cascaded models.

USE CASES

SeedRealtime Use Cases

Where real-time watch-listen-speak interaction matters

Real-Time Voice Assistant

A proactive, scene-aware assistant that tracks multiple speakers — recognizing names, faces, and voices around a dinner table, or ignoring unrelated chatter in a busy airport.

Live Translation & Speech-to-Speech

Continuous speech-to-speech and real-time translation over live conversation, without waiting for turn-by-turn transcription.

Video Analysis & Scene Monitoring

Persistent environmental awareness for object identification and scene understanding — like alerting you when the museum exhibit you are looking for enters the frame.

Reading, Navigation & Device Assistance

Understanding documents, charts, and pointing gestures in view — finding a section in a paper, catching a mistake while you brew coffee, or guiding device operation step by step.

GET STARTED

How to Use SeedRealtime

SeedRealtime has moved from research demo to consumer product

Step-by-Step: Real-Time Video Conversation in Doubao

  1. 1Download the Doubao App (iOS or Android), or open doubao.com in your browser.
  2. 2Start a real-time video call with the AI assistant from the conversation screen.
  3. 3Talk naturally — you can interrupt at any time, and the model keeps listening while it speaks.
  4. 4Point your camera at objects, documents, or scenes and refer to them directly, like "this" or "over there".
  5. 5Let it work proactively — when a target you asked about appears in view, the model can speak up on its own.

Try SeedRealtime in the Doubao App

SeedRealtime powers real-time voice and video conversation in the Doubao App. Start a live video call with the assistant to experience full-duplex watch-listen-speak interaction.

SeedRealtime API & Availability

ByteDance Seed has published an official model page and technical blog for SeedRealtime. A public API and open-source release have not been announced yet — check the official ByteDance Seed website for developer access updates.

FAQ

SeedRealtime FAQ

Is SeedRealtime free to use?

SeedRealtime is available through the Doubao App, which is free to download. Feature tiers and usage limits are determined by Doubao — there is no separate paid SeedRealtime product.

Does SeedRealtime have an API?

ByteDance Seed has not announced a public SeedRealtime API. The model is currently accessible through the Doubao App; watch the official ByteDance Seed website for developer access updates.

Is SeedRealtime open source?

No open-source release has been announced. ByteDance Seed has published an official model page and blog post, but no model weights or GitHub repository so far.

What is the latency of SeedRealtime?

ByteDance has not published exact latency figures. SeedRealtime uses streaming generation, quantization, and inference optimization for low-latency serving, and reports about half the conversational timing issues of cascaded models.

How is SeedRealtime different from voice mode assistants?

Typical voice modes are cascaded: listen, transcribe, think, then speak, with vision handled separately. SeedRealtime is a single end-to-end model that processes audio and video together and can speak while it listens and watches.

When was SeedRealtime released?

ByteDance Seed publicly released SeedRealtime on August 5, 2026, alongside an official model page and a technical blog post.

CONCLUSION

SeedRealtime and the Future of Multimodal Interaction

SeedRealtime marks a shift from turn-based voice modes toward always-on, omni-modal natural interaction. By fusing perception, understanding, decision-making, and speech generation in one end-to-end architecture, it removes the latency accumulation and context loss that have limited cascaded assistants.

With availability in the Doubao App and an official model page from ByteDance Seed, SeedRealtime has already moved from research demo to consumer product. The open questions — a public API, benchmarks, and open-source plans — are worth watching on the official ByteDance Seed website.