gen‑ai.news

The pulse of generative image & video AI.

Twice a week, the most important stories in image and video generation - new models, notable research, and meaningful product releases - distilled into a 2-minute read. No hype, no filler.

Free. Unsubscribe any time. No spam, ever.

Archive

Apple Unveils Its New Mac Studio With M5 Max and M5 Ultra Options Aimed for On-Device AI
Multimodal

Apple Unveils Its New Mac Studio With M5 Max and M5 Ultra Options Aimed for On-Device AI

Apple has refreshed the Mac Studio with new M5 Max and M5 Ultra chip options, bringing meaningful spec bumps across CPU, GPU, storage, and memory bandwidth. While Apple's marketing leans heavily on AI performance figures, the underlying hardware gains are relevant for video editors and other creative professionals as well. The M5 Max starts at $2,499 and the M5 Ultra at $5,499.

How does converting a video to 4K actually work?
Video

How does converting a video to 4K actually work?

Upscaling a video to 4K is more nuanced than simply increasing a file's resolution setting. Modern software uses increasingly sophisticated techniques to fill in missing detail, but the results are not always perfect. Here is a look at how the process actually works and what limitations remain.

Alibaba's Wan3.0 generates AI videos up to 30 seconds long from text, images, and documents
Video

Alibaba's Wan3.0 generates AI videos up to 30 seconds long from text, images, and documents

Alibaba has released Wan3.0, a video generation model capable of producing clips up to 30 seconds long from a range of inputs including text prompts, images, PDFs, and PowerPoint files. A 30-second clip at 1080p is priced at $6. The release comes as Alibaba reported a sharp 75 percent year-over-year drop in quarterly profit, reflecting heavy investment in AI infrastructure.

xAI Sues Photographer, Blaming Him for Sexual Images Created With Grok
Image

xAI Sues Photographer, Blaming Him for Sexual Images Created With Grok

xAI has filed a lawsuit against an Arkansas photographer accused of using the Grok AI platform to generate sexually explicit images of minors from his client photographs. The company's legal action positions the user - not the platform - as the party responsible for the resulting material. The case raises pointed questions about liability in AI-generated content involving child exploitation.

Photos of ‘Cat in the Hat Serial Killer’ Are AI-Generated, Police Say
Image

Photos of ‘Cat in the Hat Serial Killer’ Are AI-Generated, Police Say

Police in the U.K. and Ireland have had to issue public statements clarifying that images and videos depicting a "Cat in the Hat Serial Killer" are AI-generated and not real. The content, which spread widely online, showed a figure resembling the Dr. Seuss character in apparent crime-scene scenarios. Authorities were direct in their messaging, telling the public that "Cat in the Hat was not in your driveway."

No image
Multimodal

Harvard’s $699 startup bootcamp offers AI avatars of its instructors

Harvard Business School's new Foundry program is a $699 online startup bootcamp that uses AI avatars modeled on its instructors to give participants feedback during practice pitches and simulated board meetings. The approach moves AI-generated likenesses from novelty into a structured educational setting. It raises practical questions about how well synthetic instructor proxies can replicate the nuance of human mentorship.

World models that ignore human beliefs predict the wrong actions, new research shows
Video

World models that ignore human beliefs predict the wrong actions, new research shows

Current AI world models simulate physical environments but leave out a critical layer: what people believe, want, and intend. New research introduces a "Mental World Modeling" framework that adds these mental variables, and finds that even smaller models using it can outperform larger ones that ignore human psychology. The key bottleneck turns out to be modeling how physical and mental states evolve together over time.

No image
Image

Building Agentic Document Intelligence Pipelines: Creating Scientific Figures with AutoFigure

AutoFigure is a toolkit that takes text descriptions and research paper content as input and produces publication-ready scientific figures. A new tutorial from MarkTechPost walks through the full setup process, from configuring an API-backed workflow to exporting styled diagram galleries. The piece is aimed at researchers and developers looking to automate the visual documentation side of document intelligence pipelines.

No image
Image

Changelog - 8/20/26

Midjourney has pushed a round of updates to its alpha site, responding to two weeks of feedback from thousands of testers. The changelog reflects changes shaped directly by community input, both supportive and critical. Here is a look at what was addressed in this latest iteration.

Runway News | The Next Phase of Enterprise Video Generation
Video

Runway News | The Next Phase of Enterprise Video Generation

Runway's Chief Revenue Officer has distilled hundreds of enterprise conversations into five themes shaping how large organizations are approaching AI video generation. The piece covers everything from model consolidation and data sovereignty to shifting cost structures and the move toward autonomous execution.

Major YouTube creators are facing backlash for accepting AI money
Video

Major YouTube creators are facing backlash for accepting AI money

Several prominent filmmaking YouTubers, including Matti Haapoja and Sam "Kold" Kolder, have drawn criticism after posting sponsored content promoting Higgsfield's AI video platform without clearly disclosing the paid nature of those partnerships. The backlash intensified when other creators began sharing apparent screenshots of outreach from PR firms working on Higgsfield's behalf. The episode has sparked a broader conversation about transparency and trust in the creator community around AI tool

OpenAI's GPT-Image-2 can now generate images without a background
Image

OpenAI's GPT-Image-2 can now generate images without a background

OpenAI has added transparent background support to its GPT-Image-2 model, available in preview through the API. Rather than removing a background after the fact, the model bakes the alpha channel directly into the generation process. The feature is enabled with a single parameter.

Eddie AI Unveils a Specialized 9B Model Built for Private, Personalized Video Editing
Video

Eddie AI Unveils a Specialized 9B Model Built for Private, Personalized Video Editing

Eddie AI has announced a specialized 9-billion-parameter video editing model that trains privately for individual customers, keeping creative knowledge fully under their control. Built on a fine-tuned version of Qwen3.5-9B, the model achieves up to a 47% win rate against Kimi K3 on narrative arc identification - a result that puts it close to parity with a much larger general-purpose system. Its compact size also means it can run on standard GPUs and be deployed on-premises.

Subtlefakes: Slightly Altered Nonconsensual AI Images Are Taking Over X
Image

Subtlefakes: Slightly Altered Nonconsensual AI Images Are Taking Over X

A growing trend on X involves nonconsensual AI-generated images that have been subtly altered - tweaked just enough to evade automated detection while still being clearly identifiable as depicting real individuals. The technique is making moderation significantly harder and is spreading rapidly across the platform. 404 Media reports on how these so-called "subtlefakes" are changing the landscape of image-based abuse.

Exclusive: Early outputs of Muse Video model from Meta
Video

Exclusive: Early outputs of Muse Video model from Meta

Meta's Muse Video model has entered closed beta, offering early testers a look at its 10-second video generation capabilities. The model stands out for its native audio integration and reportedly strong temporal consistency across frames. These early outputs offer one of the first concrete glimpses into Meta's push into generative video.

No image
Video

Val Kilmer Resurrected by AI: ‘As Deep as the Grave’ Releases Seven Minutes of Footage Featuring Younger Version of Actor

A new independent film, "As Deep as the Grave," has released seven minutes of footage featuring a digitally recreated younger Val Kilmer, constructed using generative AI trained on family archives and the actor's past performances. Director Coerte Voorhees drew heavily from Kilmer's work in "Heat" and "Batman Forever" to build the likeness. The project raises fresh questions about AI's role in preserving - or reanimating - the images of actors after death or incapacitation.

Meta ran ads for an app promising to nudify female politicians
Video

Meta ran ads for an app promising to nudify female politicians

Meta's advertising platform ran ads promoting an app that claimed to generate non-consensual nude images of female politicians, with at least one ad featuring explicit deepfake content resembling a sitting US politician. The incident raises renewed questions about how effectively major platforms screen AI-enabled content that targets real individuals. It is the latest example of generative image tools being weaponized for harassment and political targeting.

What 3 creatives built with unlimited access to Google Flow
Image

What 3 creatives built with unlimited access to Google Flow

Google gave three creative professionals unlimited access to Flow, its AI-powered creative studio, and asked them to build full campaigns for local businesses. The results offer a grounded look at how the tool performs in real-world, deadline-driven creative work. Here is what each of them produced and what the process revealed.

Robin Williams’ Instagram account brought back to fight ‘AI abuse’
Image

Robin Williams’ Instagram account brought back to fight ‘AI abuse’

Robin Williams' children have reactivated their late father's Instagram account, framing it as a direct response to the unauthorized use of his likeness by AI. Zak, Zelda, and Cody Williams say they want the profile to serve as a reliable, authentic space for preserving his memory. The move follows Zelda Williams' earlier and vocal opposition to AI-generated recreations of the actor.

How Populous Brings the World's Most Iconic Venue Designs to Life with Runway
Video

How Populous Brings the World's Most Iconic Venue Designs to Life with Runway

Populous, the architecture firm behind landmark venues like Wembley Stadium and the Las Vegas Sphere, has begun integrating Runway's generative video tools into its design workflow. The firm reports that the technology helps teams visualize large-scale venues before construction and has reclaimed roughly two weeks of time previously spent on competition pitches. It is a practical case study in how generative AI is finding a foothold in high-stakes architectural practice.

ByteDance and Hollywood Reach Agreement Over AI Video IP Rights
Video

ByteDance and Hollywood Reach Agreement Over AI Video IP Rights

ByteDance and the Motion Picture Association have signed an agreement covering copyright protections for the company's AI video and image models. The deal marks one of the more formal arrangements between a major generative AI developer and Hollywood's central industry body. Details of the licensing and compliance framework signal a broader shift in how studios are approaching AI-generated content.

The Future of Deepfakes and the Decline of Reality (With Hany Farid)
Video

The Future of Deepfakes and the Decline of Reality (With Hany Farid)

Hany Farid, one of the foremost researchers on synthetic media and digital forensics, sits down with 404 Media to discuss the trajectory of deepfakes - from their origins to where the technology is headed. The conversation covers how detection, policy, and public perception have struggled to keep pace with the rapid advancement of generative tools. It is a grounded look at what the erosion of verifiable reality means for individuals and institutions alike.

AI video market has bounced back from Sora's false start
Video

AI video market has bounced back from Sora's false start

What began as a wave of skepticism after OpenAI's Sora failed to immediately reshape filmmaking has given way to a maturing AI video industry with real commercial footholds. Companies like Promise are embedding themselves near Hollywood studios, Netflix is integrating AI tools across hundreds of titles, and startups in the space are commanding multi-billion-dollar valuations. The business infrastructure - job roles, revenue splits, and specialized vendors - is now catching up to the technology i

New benchmark confirms AI models still perform poorly at visual perception
Multimodal

New benchmark confirms AI models still perform poorly at visual perception

A new benchmark from Moonshot AI isolates visual perception from logical reasoning in multimodal models, and the results are sobering. No tested frontier model clears 60 percent accuracy, with GPT-4o leading by only a narrow margin. The findings suggest that many errors previously attributed to faulty reasoning may actually originate much earlier, at the point of reading the image itself.

Google will now allow users to remove visible watermarks from AI content
Image

Google will now allow users to remove visible watermarks from AI content

Google is updating its AI-generated content policy to let users remove visible watermarks from images and other media produced by its tools. The invisible SynthID watermark, however, will stay embedded in the content regardless. This marks a shift in how Google balances user flexibility with its underlying approach to AI provenance tracking.

You can now turn off Google Gemini’s visible watermarks
Multimodal

You can now turn off Google Gemini’s visible watermarks

Google has added a toggle in Gemini and its AI video tool Flow that lets users remove the visible "sparkle" watermark from AI-generated images, videos, and music. Even with the visible mark turned off, content will still carry invisible SynthID watermarks and C2PA metadata. The change affects content produced by Google's Nano Banana and Omni models.

No image
Image

Google will now allow users to remove visible watermark from its AI generations

Google is giving users the option to remove the visible watermark that appears on AI-generated images and videos, responding to feedback from creators who found the branding intrusive. The change applies only to the visible mark - the invisible SynthID watermark, used to identify content as AI-generated, remains embedded regardless of the setting. It is a notable shift in how Google balances transparency with creative flexibility.

Tech Bro Scrapes Anti-AI Photo App Cara, Then Gloats About It
Image

Tech Bro Scrapes Anti-AI Photo App Cara, Then Gloats About It

Cara, a social platform built specifically to protect photographers from AI data scraping, has reportedly had its entire image library harvested by an outside actor. The incident is a pointed reminder that technical and policy-level protections do not always hold against determined scrapers. The person responsible apparently made their actions public, drawing sharp criticism from the creative community.

I looked inside an AI generated movie, and the best parts were all human
Video

I looked inside an AI generated movie, and the best parts were all human

A new short film produced with Higgsfield's AI video tools offers a rare look at what AI-assisted filmmaking actually looks and feels like in practice. The Verge went inside the production to find out where the technology helped - and where human craft still carried the weight. The result is a candid assessment of where generative video sits today.

How To Convert Or Upscale A Photo
Image

How To Convert Or Upscale A Photo

AI-powered upscaling tools have made it easier than ever to improve the resolution and clarity of older or low-quality photos. Engadget walks through some of the most accessible methods available today, from dedicated apps to built-in software features. Whether you are working with a scanned print or a compressed digital file, there are practical options that do not require professional editing skills.

Roku Created a New Channel That's Only AI Slop 24/7
Video

Roku Created a New Channel That's Only AI Slop 24/7

The Roku Channel has launched Fairground AI, a 24/7 linear channel running entirely on AI-generated content, including narrative shorts and commercials. The move raises real questions about audience appetite in Western markets, even as China's short-form AI video ecosystem already commands billions of views. Whether this is a glimpse of where streaming is headed, or a curiosity that quietly disappears, remains an open question.

Twitch streamers can now opt out from training Amazon’s AI
Multimodal

Twitch streamers can now opt out from training Amazon’s AI

Twitch has introduced an opt-out setting that lets streamers prevent their content - including streams, VODs, clips, chat logs, and channel images - from being used to train Amazon's generative AI models. The control applies to future training only and covers AI systems designed to generate or synthesize text, audio, images, or video. Non-generative AI features such as captions and safety tools are unaffected by the setting.

Insta360 Takes a Big Swing With Its Bold New X6 360° Camera Capable of 8K50p Video
Video

Insta360 Takes a Big Swing With Its Bold New X6 360° Camera Capable of 8K50p Video

Insta360 has announced the X6, its latest 360-degree action camera, featuring an 8K50p video recording capability built around a 1/1.1" sensor and a triple AI chip system. The camera also supports 6K60p, a Bullet Time mode at 8K100p, and 16K time-lapses, placing it at the higher end of the consumer 360-degree camera market. Pricing and full availability details have been released alongside the announcement.

No image
Video

Xiaomi’s MiLM Plus Releases PROVE: Perception-Aligned Object Removal Metrics RC-S and RC-T With a Real-World Video Benchmark

Xiaomi's MiLM Plus research team has proposed PROVE, a new evaluation framework for video object removal that introduces two perception-aligned metrics - RC-S and RC-T - alongside a real-world benchmark dataset. The work addresses a growing mismatch between how well modern removal models perform and how poorly existing metrics capture that performance. Standard measures like PSNR, SSIM, and LPIPS regularly rank model outputs in ways that disagree with human judgment.

No image
Video

The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model

LTX-2.5 is a new open-weights video generation model designed to run on consumer NVIDIA hardware, producing clips up to 6.8 seconds long with native multishot support. Lightricks released it with day-one ComfyUI integration, making it accessible to hobbyists and professionals working locally. The release positions capable video generation as something achievable without cloud infrastructure.

No image
Image

Google’s Gemini app surges to 1 billion users

Google has announced that its Gemini app has reached one billion users, marking a significant milestone for the company's AI assistant. Among those users, image generation has become a heavily used feature, with Gemini now producing more than 150 million images per day. Voice interaction is also prominent, with nearly two-thirds of users engaging the assistant through direct speech.

Apple could help you prove your iPhone photos aren’t deepfakes
Image

Apple could help you prove your iPhone photos aren’t deepfakes

Apple appears to be working on a feature called Apple Reference Image that would embed provenance metadata directly into photos at the moment of capture on an iPhone. The system is designed to help users demonstrate that a photo is genuine and not AI-generated. Code references for the feature have surfaced in the iOS 27 beta 5, though it is not yet live for users.

Claude will apply invisible watermarks to AI text and images
Multimodal

Claude will apply invisible watermarks to AI text and images

Anthropic has announced plans to embed invisible watermarks into text and images produced by Claude, making it easier for platforms and users to identify AI-generated content. The move is tied to European regulatory requirements around AI transparency. The changes are not yet live, but represent a formal commitment from the company.

No image
Video

Implementing a MiniMax-H3 Multimodal Video and Audio Generation Pipeline with ComfyUI APIs

A new technical guide walks through building a fully programmable MiniMax-H3 pipeline using ComfyUI as a headless backend, covering everything from hardware profiling to joint video-audio decoding. The tutorial treats ComfyUI not as a visual tool but as an API-driven inference engine for multimodal generation. It offers a practical path for developers who want reproducible, automated workflows without relying on a graphical interface.

Mark Zuckerberg doesn’t understand how to live
Image

Mark Zuckerberg doesn’t understand how to live

The Verge takes a critical look at Mark Zuckerberg's vision for AI and what it reveals about a broader cultural trend - the outsourcing of meaning, motivation, and creativity to generative tools. Using a telling anecdote about an AI-generated motivational poster, the piece questions what is lost when personal expression is handed off to a machine. It is a thoughtful provocation aimed at the assumptions quietly baked into how AI is being sold to us.

Meta's 'open source' Muse Glimmer model can run on a single computer
Image

Meta's 'open source' Muse Glimmer model can run on a single computer

Meta has released Muse Glimmer, a new open-source generative image model designed to run on a single consumer computer rather than requiring large-scale cloud infrastructure. The lighter footprint makes the model more accessible to independent developers and researchers working outside of data center environments. It marks another step in Meta's ongoing push to distribute AI capabilities beyond proprietary, server-dependent systems.

No image
Multimodal

Meta is back with Muse Glimmer: local, agentic, multimodal, and open source

Meta has released Muse Glimmer, an open-source multimodal AI system designed to run locally and operate with agentic capabilities. The model combines image understanding and generation with the ability to take multi-step actions, all without requiring a cloud connection. It marks Meta's latest push into accessible, on-device generative AI tools.

xAI launches Imagine Image 2.0 in Grok Quality Mode
Image

xAI launches Imagine Image 2.0 in Grok Quality Mode

xAI has updated its image generation tool inside Grok with the release of Imagine Image 2.0, available through the app's Quality Mode on both web and mobile. The update introduces a range of new capabilities including precise editing, smart resizing, and multi-reference image input. Workflow templates are also part of the release, aimed at streamlining repeated or complex image tasks.

Watching Roku’s AI channel is like eating from a trough
Video

Watching Roku’s AI channel is like eating from a trough

Roku has launched a 24/7 free ad-supported streaming channel dedicated entirely to AI-generated content, sourced from a startup called Fairground. The move marks one of the more visible attempts to bring generative video into mainstream living-room viewing. Whether audiences will warm to it is another question.

Adobe Is Coming for Canva With Expanded ChatGPT Integration
Image

Adobe Is Coming for Canva With Expanded ChatGPT Integration

Adobe has expanded its presence inside ChatGPT, moving beyond its initial Photoshop, Express, and Acrobat integrations to bring its entire suite of applications into OpenAI's conversational platform. The move positions Adobe more directly against Canva, which has built much of its recent growth on accessible, AI-assisted design tools. The unified plugin signals a deeper strategic bet on ChatGPT as a creative interface.

See what 5 builders are making with Gemini Omni
Video

See what 5 builders are making with Gemini Omni

Google's Gemini Omni lets users generate and edit video through natural conversation, and a new spotlight from the company shows how five independent builders are putting that capability to practical use. The examples range from visualizing abstract ideas to streamlining video editing workflows. Together, they offer a concrete look at how conversational video AI fits into real creative and production work.