gen‑ai.news
← Back
Video

ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio

ByteDance's Seedance 2.5 generates 30-second video clips with built-in audio

ByteDance has released Seedance 2.5, a video generation model that produces synchronized video and audio together rather than treating them as separate outputs to be combined later. The clips it generates run up to 30 seconds - three times the length of what Google's Gemini Flash Omni currently offers - which puts it toward the higher end of what commercially available models can produce in a single generation.

One of the more notable aspects of Seedance 2.5 is its multimodal input support. Users can provide dozens of reference files at once, spanning still images, existing video footage, and audio recordings. That kind of flexible reference system allows for more directed outputs, giving creators and production teams a way to anchor generated content to specific visual or sonic materials rather than working purely from text prompts.

The integrated audio generation is a meaningful step. Most video AI tools either omit audio entirely or require a separate model and a manual sync step afterward. Generating both in a single pass reduces friction in the pipeline and keeps audio and video naturally timed to each other from the start, which can be difficult to achieve when the two are produced independently.

ByteDance has framed part of the appeal around advertising production, where teams often need to assemble short clips into a finished piece. If a single model can output a polished 30-second segment with usable audio, that collapses several steps in the typical workflow. Whether Seedance 2.5 delivers on that in practice - particularly for brand-specific or heavily art-directed work - will depend on how much control the reference inputs actually provide and how consistent the outputs are across a production run.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.