gen‑ai.news
← Back
Multimodal

Meta launches Muse Image across its apps and previews Muse Video

Meta launches Muse Image across its apps and previews Muse Video

Meta has begun deploying Muse Image, its proprietary image generation model, across Meta AI-powered surfaces including Instagram Stories. The rollout makes prompt-based image creation available to the large existing user base of Meta's apps, rather than requiring users to seek out a standalone tool. It is one of the broader deployments of a natively developed image model by a major social platform.

Muse Image allows users to describe a scene or concept in text and receive a generated image in response. By embedding this functionality inside Instagram Stories and other Meta AI entry points, the company is positioning generative imagery as a casual, everyday feature rather than a specialist capability. This kind of integration - where creation happens inside the app a user is already in - lowers the barrier to use considerably compared with visiting a dedicated image generation service.

Alongside the Muse Image launch, Meta offered a preview of Muse Video, suggesting the underlying model family is being developed with video generation as a near-term target. Details on Muse Video's capabilities, resolution, clip length, and availability timeline were not fully disclosed at this stage, but the preview indicates Meta is working to keep pace with other labs and platforms that have been advancing text-to-video generation over the past year.

Meta has been building out its generative AI infrastructure steadily, with its open-weight Llama models forming the backbone of much of its AI work. Muse represents a parallel track focused specifically on visual media. As the company controls some of the world's largest social content platforms, its ability to distribute these tools at scale - and to gather feedback from millions of users quickly - gives it a distinct position in the generative image and video landscape.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Multimodal

Harvard’s $699 startup bootcamp offers AI avatars of its instructors

Harvard Business School's new Foundry program is a $699 online startup bootcamp that uses AI avatars modeled on its instructors to give participants feedback during practice pitches and simulated board meetings. The approach moves AI-generated likenesses from novelty into a structured educational setting. It raises practical questions about how well synthetic instructor proxies can replicate the nuance of human mentorship.

New benchmark confirms AI models still perform poorly at visual perception
Multimodal

New benchmark confirms AI models still perform poorly at visual perception

A new benchmark from Moonshot AI isolates visual perception from logical reasoning in multimodal models, and the results are sobering. No tested frontier model clears 60 percent accuracy, with GPT-4o leading by only a narrow margin. The findings suggest that many errors previously attributed to faulty reasoning may actually originate much earlier, at the point of reading the image itself.

You can now turn off Google Gemini’s visible watermarks
Multimodal

You can now turn off Google Gemini’s visible watermarks

Google has added a toggle in Gemini and its AI video tool Flow that lets users remove the visible "sparkle" watermark from AI-generated images, videos, and music. Even with the visible mark turned off, content will still carry invisible SynthID watermarks and C2PA metadata. The change affects content produced by Google's Nano Banana and Omni models.