gen‑ai.news
← Back
Video

Your Edit Isn’t Finished Until It Sounds Right

Your Edit Isn’t Finished Until It Sounds Right

For most video editors, sound comes last. The picture gets locked, and only then does the search begin for the right music track, voiceover, or ambient effects. Adobe is pushing against that convention with three new audio capabilities inside Firefly: Generate Music, Generate Speech, and Generate Sound Effects. The idea is to make audio available as a creative tool during the edit itself, not just after it.

Generate Music produces original, fully licensed tracks based on the mood and duration of a video. An editor can try several different musical directions early in a cut to see how each one shifts the feel of a scene - something that previously required either a music library subscription or waiting until a budget was approved for composition. Generate Speech converts written scripts into narration using Adobe's own Firefly speech model or ElevenLabs voices, which can serve as a temporary voiceover track to test whether a sequence has the right pacing before a professional recording is scheduled. Generate Sound Effects builds custom audio from descriptions tied to the action and timing of a specific moment - a footstep, a whoosh, an ambient texture - giving motion designers and vloggers alike a way to prototype the sonic detail of a scene without digging through sample libraries.

What connects these tools is the emphasis on experimentation over finalization. A generated music cue might reveal that a scene is moving too quickly. A placeholder voiceover might show that a sequence needs more breathing room between cuts. Sound has always shaped how a story feels, but in traditional post-production it has rarely been available early enough in the process to influence those decisions. Firefly's audio tools are positioned to close that gap.

The audio additions are part of a wider push by Adobe to make Firefly a single environment for experimenting across different media types. Alongside its own models, Firefly now integrates generation tools from Google, Kling AI, Luma AI, OpenAI, Runway, and others, letting creators move between image, video, and audio without switching between separate applications. The underlying premise is that creative decisions rarely happen in a fixed sequence - and that having sound available from the start gives editors one more way to figure out what a project actually needs.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.