gen‑ai.news
← Back
Video

China's MiniMax H3 is the first open model to top an AI video ranking

China's MiniMax H3 is the first open model to top an AI video ranking

MiniMax, a Shanghai-based AI company, has publicly released the model weights for its H3 video generation model - making it the first open model to reach the top position on a widely referenced AI video quality ranking. Until now, the upper tier of such leaderboards has been dominated by closed, API-only systems from well-funded Western labs, making the H3 result a meaningful shift in how the open and closed ecosystems compare on video generation quality.

The release of model weights means developers, researchers, and hobbyists can download and run H3 directly, rather than accessing it through a paid API or restricted interface. This kind of open availability has practical consequences: it enables fine-tuning for specific use cases, local deployment without usage costs, and broader scrutiny of how the model actually behaves. For the generative video field, which has seen relatively few high-performing open releases compared to the image generation space, H3 fills a notable gap.

MiniMax has been building its profile steadily in the generative AI space. The company previously released video and multimodal models under its Hailuo brand, and H3 appears to represent a substantial step up in output quality. Benchmark rankings for AI video - such as those that evaluate motion consistency, prompt adherence, and visual fidelity - have increasingly become reference points for both developers choosing tools and researchers tracking progress, though they carry the usual caveats about how well leaderboard performance translates to real-world use.

The broader context here is a continuing pattern of Chinese AI labs releasing capable open models that compete with or exceed proprietary Western counterparts on benchmarks - a trend seen earlier in the language model space with releases like DeepSeek. Whether H3 holds its ranking position as other labs update their own models remains to be seen, but the release adds meaningful competition to the open generative video ecosystem at a time when that space is still maturing.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.