gen‑ai.news
← Back
Multimodal

Meta launches Muse Image across its apps and previews Muse Video

Meta launches Muse Image across its apps and previews Muse Video

Meta has begun deploying Muse Image, its proprietary image generation model, across Meta AI-powered surfaces including Instagram Stories. The rollout makes prompt-based image creation available to the large existing user base of Meta's apps, rather than requiring users to seek out a standalone tool. It is one of the broader deployments of a natively developed image model by a major social platform.

Muse Image allows users to describe a scene or concept in text and receive a generated image in response. By embedding this functionality inside Instagram Stories and other Meta AI entry points, the company is positioning generative imagery as a casual, everyday feature rather than a specialist capability. This kind of integration - where creation happens inside the app a user is already in - lowers the barrier to use considerably compared with visiting a dedicated image generation service.

Alongside the Muse Image launch, Meta offered a preview of Muse Video, suggesting the underlying model family is being developed with video generation as a near-term target. Details on Muse Video's capabilities, resolution, clip length, and availability timeline were not fully disclosed at this stage, but the preview indicates Meta is working to keep pace with other labs and platforms that have been advancing text-to-video generation over the past year.

Meta has been building out its generative AI infrastructure steadily, with its open-weight Llama models forming the backbone of much of its AI work. Muse represents a parallel track focused specifically on visual media. As the company controls some of the world's largest social content platforms, its ability to distribute these tools at scale - and to gather feedback from millions of users quickly - gives it a distinct position in the generative image and video landscape.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Multimodal

Google’s new SynthID website can identify AI-generated media

Google has launched a public-facing website for SynthID, its AI watermarking technology, allowing anyone to upload and check whether an image, video, or audio clip was generated by AI. The tool extends SynthID beyond its previous developer and enterprise integrations, putting detection capabilities directly in the hands of everyday users. This marks a notable step toward accessible, practical tools for identifying synthetic media in the wild.

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model
Multimodal

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

Reka AI has released Rho-1, a 19-billion-parameter model that handles text, images, video, and robot control actions within a single neural network. Rather than routing different tasks to specialized subsystems, it processes all modalities as tokens in one shared context window. The model was trained on 320 H100 GPUs over roughly three months, using considerably less compute than many leading models today.

Uniting LED Volumes and Live Visual Chaos: Cinematographer Bradford Lipson on 'Rolling Loud'
Multimodal

Uniting LED Volumes and Live Visual Chaos: Cinematographer Bradford Lipson on 'Rolling Loud'

Cinematographer Bradford Lipson explains how he and director Jeremy Garelick used LED volume technology to blend real footage from the Rolling Loud music festival with controlled stage work, creating the illusion that actors were surrounded by tens of thousands of live concertgoers. The production shot on location at the festival for just three days before rebuilding that world on the Lux Stage at Trilith Studios. Lipson walks through the technical and artistic decisions behind matching concert