gen‑ai.news
← Back
Video

Vidu Q3 AI Video Model with Native Audio

Vidu Q3 AI Video Model with Native Audio

Shengshu has launched Vidu Q3, an updated AI video generation model that introduces native audio as a core capability rather than a post-processing add-on. Previous versions of Vidu focused primarily on visual output, so the integration of audio directly into the generation pipeline is a meaningful shift in the model's scope. Creators working on short-form content, social media clips, or early-stage production work can now generate video and sound together in a single pass.

Beyond audio, the Q3 release brings improvements to motion quality - addressing a persistent challenge in AI video where movement can appear unnatural or stuttered. Stronger prompt control is also highlighted, meaning the model is better at translating detailed text descriptions into accurate visual results. Both of these areas have been active pain points across the broader AI video landscape, and incremental gains here have practical consequences for how usable the output is without manual correction.

Vidu is developed by Shengshu Technology, a Beijing-based AI company that has been building in the video generation space for several years. The model competes in a field that includes tools like Runway, Kling, and Hailuo, all of which are also iterating rapidly on motion fidelity and creative control. Native audio generation, however, remains less common across the field, giving Vidu Q3 a degree of differentiation at this moment.

For users, the practical appeal of native audio is the reduction of pipeline complexity - fewer tools needed to arrive at a shareable video. Whether the audio quality holds up for professional use cases will depend on how well the model handles synchronization, sound design variety, and consistency across different prompt types. Vidu Q3 is available through the Vidu platform at vidu.com, where users can test the model's capabilities directly.

Read at shengshu →
Share:X

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.