Your Edit Isn’t Finished Until It Sounds Right

For most video editors, sound comes last. The picture gets locked, and only then does the search begin for the right music track, voiceover, or ambient effects. Adobe is pushing against that convention with three new audio capabilities inside Firefly: Generate Music, Generate Speech, and Generate Sound Effects. The idea is to make audio available as a creative tool during the edit itself, not just after it.
Generate Music produces original, fully licensed tracks based on the mood and duration of a video. An editor can try several different musical directions early in a cut to see how each one shifts the feel of a scene - something that previously required either a music library subscription or waiting until a budget was approved for composition. Generate Speech converts written scripts into narration using Adobe's own Firefly speech model or ElevenLabs voices, which can serve as a temporary voiceover track to test whether a sequence has the right pacing before a professional recording is scheduled. Generate Sound Effects builds custom audio from descriptions tied to the action and timing of a specific moment - a footstep, a whoosh, an ambient texture - giving motion designers and vloggers alike a way to prototype the sonic detail of a scene without digging through sample libraries.
What connects these tools is the emphasis on experimentation over finalization. A generated music cue might reveal that a scene is moving too quickly. A placeholder voiceover might show that a sequence needs more breathing room between cuts. Sound has always shaped how a story feels, but in traditional post-production it has rarely been available early enough in the process to influence those decisions. Firefly's audio tools are positioned to close that gap.
The audio additions are part of a wider push by Adobe to make Firefly a single environment for experimenting across different media types. Alongside its own models, Firefly now integrates generation tools from Google, Kling AI, Luma AI, OpenAI, Runway, and others, letting creators move between image, video, and audio without switching between separate applications. The underlying premise is that creative decisions rarely happen in a fixed sequence - and that having sound available from the start gives editors one more way to figure out what a project actually needs.