gen‑ai.news
← Back
Video

MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

MiniMax has released H3, a model the company describes as a general-purpose multimodal generation system. Unlike conventional text-to-video tools that treat audio as an afterthought or a separate post-processing step, H3 is designed from the ground up to ingest and output multiple modalities together - text, images, video, and audio are all handled within one unified context window, and the resulting video clip includes native stereo sound rather than mono or silence.

On the output side, H3 generates clips at 2K resolution with durations ranging from 4 to 15 seconds in integer steps. That ceiling of 15 seconds is notably longer than what many competing models currently offer at comparable resolution, and the stereo audio output is a meaningful differentiator at a time when most video generation models still produce silent footage that requires a separate audio pipeline to be useful in production.

The architecture framing - treating all modalities as one context rather than chaining specialized sub-models - has practical implications for consistency. When a model processes a reference audio clip, a style image, and a text prompt together, the resulting video is less likely to show the seams that appear when separate systems are stitched together after the fact. This approach mirrors design decisions seen in some large language models that handle mixed inputs, now applied to a generation rather than comprehension task.

MiniMax has been building out its video generation capabilities steadily, having previously released models in its Video-01 line. H3 represents a more ambitious architectural step, moving away from a video-first model with optional extras toward something closer to a native omni-modal system. Details on the underlying architecture, training data, and API availability were not fully disclosed at launch, but the model appears aimed at developers and creators who need a single endpoint capable of handling complex, mixed-modality prompts without assembling multiple tools.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Watching Roku’s AI channel is like eating from a trough
Video

Watching Roku’s AI channel is like eating from a trough

Roku has launched a 24/7 free ad-supported streaming channel dedicated entirely to AI-generated content, sourced from a startup called Fairground. The move marks one of the more visible attempts to bring generative video into mainstream living-room viewing. Whether audiences will warm to it is another question.

See what 5 builders are making with Gemini Omni
Video

See what 5 builders are making with Gemini Omni

Google's Gemini Omni lets users generate and edit video through natural conversation, and a new spotlight from the company shows how five independent builders are putting that capability to practical use. The examples range from visualizing abstract ideas to streamlining video editing workflows. Together, they offer a concrete look at how conversational video AI fits into real creative and production work.

Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0
Video

Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0

Black Forest Labs has moved FLUX 3 Video out of early access and into general availability, offering Full HD video generation with clips up to 20 seconds long. The model includes native audio output and lip-synced dialogue across more than 14 languages. According to BFL's own Elo benchmark rankings, it outperforms both Gemini Omni Flash and Seedance 2.0.