gen‑ai.news
← Back
Video

Vidu S1

Shengshu, the Chinese AI company behind the Vidu platform, has released Vidu S1, a new iteration of its text-to-video and image-to-video generation model. The release marks a notable step forward in the company's ongoing development of its video synthesis technology, which has been evolving rapidly since Vidu's initial debut.

Vidu S1 is designed to produce higher-quality video output with improved motion consistency and more detailed rendering compared to earlier versions. These kinds of improvements - particularly around temporal coherence, where objects and characters maintain their appearance across frames - are among the most technically demanding challenges in generative video, and progress in this area tends to have an outsized impact on usability in real creative workflows.

Shengshu operates in a competitive landscape that now includes well-resourced models from companies like Runway, Kling, Hailuo, and Google. Vidu has distinguished itself in part by offering strong performance at accessible pricing, targeting both professional creators and developers building video-generation into their own applications. The S1 release appears aimed at keeping the platform competitive as the broader field continues to advance quickly.

Details on the specific technical architecture underlying Vidu S1 - such as training data scale, model size, or inference approach - have not been fully disclosed. Users can access the model through the Vidu platform at vidu.com, where the company has made the tool available for direct testing and production use.

Read at shengshu →
Share:X

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.