gen‑ai.news
← Back
Video

Alibaba's Wan3.0 generates AI videos up to 30 seconds long from text, images, and documents

Alibaba's Wan3.0 generates AI videos up to 30 seconds long from text, images, and documents

Alibaba has launched Wan3.0, the latest version of its video generation model, expanding both the length and input flexibility of AI-generated video clips. Where many competing tools cap output at a few seconds or require a single prompt type, Wan3.0 accepts text, images, PDF documents, and PowerPoint files as source material - broadening the range of practical use cases, particularly for business and presentation-oriented content.

The model supports video generation at up to 1080p resolution and can produce clips as long as 30 seconds, a duration that remains relatively uncommon among publicly available video generation systems. Alibaba has set the price at $6 per 30-second 1080p clip, positioning it as a commercial offering rather than a free research preview. The ability to ingest structured document formats like PDFs and slide decks is a notable differentiator, potentially making it easier for users to convert existing materials into video without manual reformatting.

Wan3.0 builds on Alibaba's ongoing investment in generative AI, a push that has come with significant financial costs. The company reported a 75 percent drop in quarterly profit compared to the same period a year earlier, a decline directly tied to increased AI-related capital expenditure. This pattern - trading near-term earnings for long-term AI positioning - mirrors spending trends seen at other large technology companies investing heavily in model development and infrastructure.

The broader context here is a competitive landscape in AI video generation that has grown considerably more crowded over the past year, with offerings from companies like OpenAI, Google, Runway, and others all targeting different segments of the market. Wan3.0's document-input capability and relatively long clip duration give it a distinct angle, though real-world output quality and consistency will ultimately determine how it fares against established alternatives.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Claude can now generate animated explainer videos and live data dashboards from text prompts
Video

Claude can now generate animated explainer videos and live data dashboards from text prompts

Anthropic has introduced two beta features for Claude: Dashboards, which connects to data sources like BigQuery and Snowflake to build live dashboards from text prompts, and Motion, which produces animated explainer videos from text and images. The company also expanded its existing Docs, Slides, and Design tools to all users, including those on free plans.

Worried About AI Training on Your Videos? This Platform Lets You Protect and Monetize Your Work
Video

Worried About AI Training on Your Videos? This Platform Lets You Protect and Monetize Your Work

Sinima is a newly launched video hosting platform aimed at filmmakers and creators who want more control over how their work is used in AI training. The platform assigns each uploaded file a digital fingerprint and ownership certificate, and lets creators set their own AI licensing permissions. When footage is licensed for AI training, creators keep 85% of the fee, with eligible content valued at $4,000 to $7,500 per hour.

DittoDub Launches New Native 6 AI Dubbing Model With Support for Over 100 Languages
Video

DittoDub Launches New Native 6 AI Dubbing Model With Support for Over 100 Languages

DittoDub has released Native 6, a new AI dubbing model that now serves as the default option on its platform and supports 109 languages across 126 dialect variations. The model is built around preserving a speaker's vocal character and emotional tone when translating video content. It slots into DittoDub's existing localization workflow, covering dubbed audio, subtitles, and translated metadata.