gen‑ai.news
← Back
Video

ICYMI: Black Forest Labs opens FLUX 3 Video early access

ICYMI: Black Forest Labs opens FLUX 3 Video early access

Black Forest Labs, the company behind the widely used FLUX family of image generation models, has begun accepting early access applications for FLUX 3 Video. The new model supports clips of up to 20 seconds in length and - notably - generates audio natively alongside the video, rather than relying on a separate audio pipeline bolted on after the fact.

Native audio generation in a video model is still relatively uncommon. Most video AI tools either skip audio entirely or handle it as a post-processing step, so integrating it into the core generation process positions FLUX 3 Video closer to a small group of models pushing toward fully unified audiovisual output. How well the audio quality holds up across varied prompts will be one of the key things to watch as early access users begin sharing results.

Black Forest Labs has also outlined plans for several related releases. These include an image model update under the FLUX 3 name, an "action" model variant - likely oriented toward motion-heavy or physically dynamic scenes - and an open-weight release. The open-weight commitment in particular continues a pattern the company established with earlier FLUX models, which became popular in the self-hosted and research communities precisely because weights were made publicly available.

FLUX 3 Video enters a competitive field that includes Runway, Kling, and Google's Veo, among others. Black Forest Labs built its reputation on strong image fidelity and prompt adherence with its earlier FLUX.1 models, and it will be worth seeing whether those qualities carry over into the video domain. Early access is currently limited, but the phased rollout - with open-weight models to follow - suggests the company intends FLUX 3 Video to reach a broad audience over time.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.