gen‑ai.news

The pulse of generative image & video AI.

Twice a week, the most important stories in image and video generation - new models, notable research, and meaningful product releases - distilled into a 2-minute read. No hype, no filler.

Free. Unsubscribe any time. No spam, ever.

Archive

No image
Video

NVIDIA Released DeepStream 9.1: Bringing Agentic AI to Vision AI With 13 Skills and Multi-View 3D Tracking

NVIDIA has updated its DeepStream video analytics SDK to version 9.1, introducing 13 agentic skills that allow AI coding agents to construct multi-camera pipelines from natural-language instructions. The release also brings Multi-View 3D Tracking, which merges detections from multiple cameras into a single shared 3D space with consistent object IDs. Rounding out the update are automatic camera calibration, JetPack 7.2 support, and a consolidated open-source repository.

TikTok is testing an AI likeness detection tool
Video

TikTok is testing an AI likeness detection tool

TikTok is piloting an opt-in tool that scans the platform for AI-generated likenesses of creators, allowing them to report unauthorized uses to the company. The test is currently limited to a subset of US creators and requires identity verification through a third-party service. YouTube recently launched a comparable feature for all adult users.

Apple and Google ordered by San Francisco attorney to take action against 'nudify' apps
Image

Apple and Google ordered by San Francisco attorney to take action against 'nudify' apps

San Francisco's city attorney has sent cease-and-desist letters to Apple and Google, demanding they remove 13 so-called "nudify" apps from their platforms. These apps use generative AI to digitally undress images of real people without their consent, raising serious legal and ethical concerns. The action marks one of the more direct legal interventions by a local authority targeting AI-generated non-consensual intimate imagery.

No image
Multimodal

Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers

NVIDIA and Hugging Face have joined forces to bring large-scale fine-tuning of image and video diffusion models into the NeMo Automodel framework, integrated with the Diffusers library. The collaboration aims to make distributed training more accessible for teams working with models that would otherwise be difficult to fine-tune on limited hardware. The result is a more streamlined path from a pretrained diffusion model to a customized one, without requiring deep infrastructure expertise.

No image
Video

Netflix Co-CEO Explains How Gen-AI Was Used in 300 Different Titles: ‘We Believe It Is Going to Enhance Their Abilities’

Netflix Co-CEO Ted Sarandos has revealed that generative AI has been used across 300 titles on the platform, with the docuseries "The American Experiment" featuring 17 minutes of AI-enhanced footage. Sarandos framed the technology as a tool to expand what creators can do rather than replace them. The comments offer one of the most concrete public accounts yet of how a major streaming service is integrating AI into its content pipeline.

Mayor Mamdani Says Landlords Can’t Secretly Use AI Images to Advertise Properties
Image

Mayor Mamdani Says Landlords Can’t Secretly Use AI Images to Advertise Properties

New York City Mayor Zohran Mamdani has moved to ban landlords from using undisclosed AI-generated or AI-edited images in property listings. The rule targets deceptive advertising practices that make rental units appear more appealing than they actually are. It is part of a broader push by the mayor to regulate AI-assisted commercial practices in the city.

DJI Used an AI-Generated ‘Person’ In an Ad, Angering the Actual Humans Who Buy Products
Video

DJI Used an AI-Generated ‘Person’ In an Ad, Angering the Actual Humans Who Buy Products

DJI drew criticism after posting a TikTok advertisement featuring an AI-generated human avatar, a choice that did not sit well with its core audience of photographers and videographers. The ad has since been deleted, but not before sparking a broader conversation about the use of synthetic people in brand marketing. For a company whose products are built around capturing authentic human moments, the optics proved difficult.

Create, edit and star in videos with two Google Vids updates
Video

Create, edit and star in videos with two Google Vids updates

Google has added two new features to its Vids tool in Workspace: Gemini Omni-powered editing assistance and personal avatars that let users appear in videos without recording themselves. Together, the updates expand what non-professional creators can do directly within Google's productivity suite.

No image
Image

How a former DeepMind researcher raised at a $300M pre-seed valuation before launching a product

Andrew Dai, a former DeepMind researcher whose work contributed to foundational AI systems, has raised funding at a $300 million pre-seed valuation without yet shipping a product. The round reflects the degree to which investor confidence in a founder's background can drive early-stage valuations in AI. Dai is focused on visual AI, which he sees as one of the next significant areas of development in the field.

xAI sues Grok user for generating nonconsensual sexualized deepfakes
Image

xAI sues Grok user for generating nonconsensual sexualized deepfakes

xAI has filed a lawsuit against a user who allegedly exploited Grok's image generation capabilities to produce nonconsensual sexualized deepfakes of both adults and minors. The case marks a notable instance of an AI company taking direct legal action against a user for misuse of its own tools. It raises broader questions about platform liability, enforcement mechanisms, and the limits of generative AI safeguards.

No image
Multimodal

Thinking Machines Lab Releases Inkling: A 975B-Parameter Open-Weights Multimodal MoE With 41B Active Parameters And Controllable Thinking Effort

Thinking Machines Lab has released Inkling, a 975B-parameter open-weights multimodal model built on a Mixture-of-Experts architecture that keeps only 41B parameters active at any given time. Licensed under Apache 2.0, it accepts text, image, and audio inputs and offers a 1M-token context window. Rather than competing for top benchmark rankings, the model is positioned as a customizable base with adjustable reasoning depth.

No image
Video

Reelful’s AI turns your camera roll into short-form videos for social media

Reelful is a new app that uses AI to automatically assemble short-form social media videos from photos and clips already stored in a user's camera roll. It targets people who want to post regularly but find conventional editing software too demanding. The app handles sequencing, timing, and formatting without requiring any manual editing experience.

‘The Odyssey’s’ AI-Generated Competitor Is Bereft of Humanity
Video

‘The Odyssey’s’ AI-Generated Competitor Is Bereft of Humanity

An AI-generated adaptation of Homer's Odyssey is heading to screens around the same time as Christopher Nolan's highly anticipated live-action version, prompting pointed questions about the role of fully synthetic filmmaking in cinema. The comparison between the two projects throws into sharp relief what generative video tools can and cannot yet replicate. A PetaPixel writer argues the AI film is notably lacking in the human qualities that make storytelling resonate.

Reconstructing Pelé’s “lost” goal
Video

Reconstructing Pelé’s “lost” goal

Google DeepMind used generative AI to reconstruct a 1959 Pelé goal for which no video footage survives. The project, documented in a short film, draws on historical accounts and AI image and video synthesis to visualize the legendary strike at Rua Javari. It offers a glimpse into how AI tools are being applied to cultural and historical preservation.

No image
Video

Video-generation startup PixVerse raises $439M, valuation soars past $2B

PixVerse, a startup focused on AI video generation, has closed a $439 million funding round that pushes its valuation above $2 billion. The company plans to use the capital to develop its world model capabilities and grow its presence in new markets. The raise reflects continued investor appetite for generative video infrastructure as competition in the space intensifies.

No image
Video

Building a VideoAgent-Style Multi-Agent System: Intent Parsing, Graph Planning, and Tool Routing for Video Editing Tasks

A new tutorial from MarkTechPost walks through building a VideoAgent-style multi-agent pipeline for video editing, requiring no API keys. The system chains together intent parsing, graph planning, and tool routing to handle natural-language video editing instructions end to end. Output ranges from question answering about video content to fully edited artifacts like beat-synced cuts.

Exclusive: Early 30-second AI videos generated by Seedance 2.5
Video

Exclusive: Early 30-second AI videos generated by Seedance 2.5

ByteDance is preparing to launch Seedance 2.5, a video generation model capable of producing native 30-second clips. API access is expected to open around July 16 following earlier delays. Testing Catalog has shared early sample outputs ahead of the official release.

Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
Image

Meta turns off the Instagram feature that let users make AI deepfakes of public accounts

Meta has disabled an Instagram feature that allowed users to generate AI images by tagging public accounts, after swift and widespread criticism. The feature, tied to Meta's new Muse Image AI model, enabled anyone to use another account's content in AI-generated images without the account owner's consent. The company walked back the functionality just days after announcing it.

No image
Video

Meet LingBot-World-Infinity: An Open Causal World Model With An Agentic Harness

Ant Group's embodied-intelligence unit Robbyant has released LingBot-World-Infinity, a 14-billion-parameter causal video generation model designed to function as an interactive world simulator. The model pairs a novel attention mechanism with distillation over long rollout trajectories to address the texture and geometry degradation that typically plagues extended generation sessions. A two-tier agentic harness - separating high-level planning from rendering - enables a demonstrated 60-minute co

The AI-Generated Taylor Swift Wedding Photos Are Getting Debunked Really Quickly
Image

The AI-Generated Taylor Swift Wedding Photos Are Getting Debunked Really Quickly

When photos purporting to show Taylor Swift's wedding circulated online last weekend, they were debunked faster than most previous high-profile AI image hoaxes. The episode is drawing attention not for the fakes themselves, but for how quickly audiences, journalists, and verification tools identified them as generated. It may signal a meaningful shift in how the public handles AI-fabricated imagery around major news events.

Here's how to block Meta from using your Instagram pictures for its AI
Image

Here's how to block Meta from using your Instagram pictures for its AI

Meta has quietly opted all public Instagram accounts into its Muse Image AI training program, using posted photos without requiring any additional consent from users. If you would prefer your images not be used for this purpose, there is an opt-out process available - though it requires manual action on your part. Here is what you need to know to protect your content.

Character.AI wants a piece of the microdrama pie
Video

Character.AI wants a piece of the microdrama pie

Character.AI is expanding its platform into short-form video with the launch of c.ai Series, episodic animated content built almost entirely with generative AI. The move positions the company within the fast-growing microdrama market, which is projected to reach $26 billion in the coming years. Unlike live-action microdrama platforms, c.ai Series lets viewers watch and interact with AI-generated animated stories on their phones.

ByteDance debuts Seedream 5.0 Pro with advanced reasoning
Image

ByteDance debuts Seedream 5.0 Pro with advanced reasoning

ByteDance has introduced Seedream 5.0 Pro, its latest multimodal image generation model, bringing advanced reasoning capabilities alongside precise editing tools and multilingual support. The release marks a notable step forward in ByteDance's push into generative image AI. It positions Seedream as a more capable and flexible option for a range of creative and practical use cases.

Meta launches Muse Image across its apps and previews Muse Video
Multimodal

Meta launches Muse Image across its apps and previews Muse Video

Meta has rolled out Muse Image, its in-house image generation model, across its family of apps including Instagram Stories, where users can create visuals from text prompts. The company also offered an early preview of Muse Video, signaling its intent to extend the capability to moving images. The moves mark a significant step in Meta's effort to embed generative media tools directly into its social platforms.

No image
Image

Google’s deepfake detector system used to debunk McConnell hoax pic

A photograph appearing to show Kentucky Senator Mitch McConnell in severe medical distress circulated widely this week before being identified as AI-generated. Google's deepfake detection system played a central role in debunking the image, marking a notable real-world application of the technology.

6 Best Image Upscalers Tested: Photo Upscaling Without the Plastic Look
Image

6 Best Image Upscalers Tested: Photo Upscaling Without the Plastic Look

AI-powered image upscalers have matured to the point where enlarging a photo to 4K or 8K is largely a solved problem - the harder question now is whether the output still looks like a photograph. PetaPixel tested six leading tools to find out which ones preserve natural texture and which ones produce the telltale over-smoothed, plastic appearance. The results offer a practical guide for photographers choosing an upscaling workflow.

Google Photos Adds Video Remix That AI-Generates ‘Shareable Clips’
Video

Google Photos Adds Video Remix That AI-Generates ‘Shareable Clips’

Google Photos has introduced Video Remix, a new AI-powered feature that automatically generates short, shareable clips from a user's existing photo and video library. The addition is the latest in a string of generative AI tools Google has been folding into the app over the past year. It continues the company's broader push to embed AI assistance across its consumer software lineup.

Create shareable video clips in seconds with Video Remix in Google Photos.
Video

Create shareable video clips in seconds with Video Remix in Google Photos.

Google Photos has introduced Video Remix, a feature that turns existing video footage into polished, shareable clips with minimal effort. Using a few taps, the tool automatically selects highlights and applies edits to produce a finished video. It is designed to lower the barrier between raw footage and content worth sharing.

Google announces new 'Video Remix' feature its for AI subscribers
Video

Google announces new 'Video Remix' feature its for AI subscribers

Google has introduced a Video Remix feature for its AI subscription tier, letting users transform videos stored in Google Photos using Gemini Omni. The tool offers a way to creatively reimagine existing footage rather than generating video from scratch. It represents Google's latest step in embedding generative AI capabilities directly into its consumer product ecosystem.

Meta’s New AI Image Generator Can Make Pictures From Anyone’s Public Instagram Page
Image

Meta’s New AI Image Generator Can Make Pictures From Anyone’s Public Instagram Page

Meta's Superintelligence Labs division has released its first image generation model, now rolling out across Instagram, WhatsApp, and the Meta AI app. One of its more notable features is the ability to generate images drawing from the content of any public Instagram profile. The move marks Meta's first significant push into first-party generative image tools at scale.

Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
Image

Muse Image is technically impressive, but Meta's use of Instagram photos raises questions

Meta's Superintelligence Labs has released Muse Image, its first image generation model, which operates as an agent capable of refining its outputs through tools like code execution and web search. A standout feature allows users to generate images of real people by tagging their public Instagram profiles - without those individuals' consent. The opt-out approach is drawing scrutiny for its likely conflict with European privacy regulations.

LARPING: How Influencers Fake Being Rich
Image

LARPING: How Influencers Fake Being Rich

A growing number of social media influencers are using generative AI tools to fabricate the appearance of wealth - private jets, luxury hotels, and designer goods - without ever leaving home. Meanwhile, AI-generated fake products, including photorealistic flowers that don't exist, have flooded marketplaces like Etsy, eBay, and Amazon. The trend points to a broader and largely unregulated use of image AI to deceive consumers.

No image
Video

NVIDIA’s Cosmos-Framework Tutorial: Designing a Colab-Friendly Miniature of Cosmos 3 World Models with Omnimodal Mixture-of-Transformers

NVIDIA's Cosmos 3 world models are powerful but demand hardware well beyond a typical laptop or free cloud notebook. This tutorial bridges that gap by building a compact, Colab-compatible stand-in that mirrors the real framework's structure - using an omnimodal Mixture-of-Transformers to jointly model text, vision, and action. It offers a hands-on way to understand how Cosmos 3 operates without needing access to the full model checkpoints.

Meta built an AI detection tool to ID images and video created with its new models
Multimodal

Meta built an AI detection tool to ID images and video created with its new models

Meta has built an AI detection tool designed to identify images and video generated by its own models. The tool comes with rate limits, an unusual design choice that raises questions about its intended scope and accessibility. It represents Meta's latest step toward addressing the growing challenge of synthetic media attribution.

No image
Video

Jon Erwin Isn’t Hiding from AI: The ‘Young Washington’ Director on How AI Can Save Jobs and Bolster Collaboration

Director Jon Erwin, known for faith-based productions like "House of David," is openly integrating AI tools into his filmmaking workflow through his production company Innovative Dreams. In an interview with IndieWire, he argues that transparency about AI use is essential, and that the technology can protect jobs rather than eliminate them. His perspective offers a grounded counterpoint to the anxiety surrounding AI's role in Hollywood.

Meta's new Muse Image model accepts Instagram accounts as a prompt
Image

Meta's new Muse Image model accepts Instagram accounts as a prompt

Meta has introduced a new image generation model called Muse that can take an Instagram account as a prompt, drawing on a user's posted content to inform the style and subject of generated images. The model is also being integrated into Instagram Stories effects and WhatsApp's image generation feature. It marks a notable shift in how personal social media history can be used as creative input for AI tools.

Meta’s new Muse Image model can pull other Instagram users into AI photos
Image

Meta’s new Muse Image model can pull other Instagram users into AI photos

Meta has introduced Muse Image, the first AI image generation model from its Superintelligence Labs division, now powering image tools across Meta AI, Instagram, and WhatsApp. The model includes a notable social feature that lets users pull other Instagram accounts into generated images, raising immediate questions about consent and misuse. It is also described as "agentic," working alongside a large language model to reason through prompts before generating output.

Midjourney is Trying to Force Hollywood to Reveal How it Uses AI
Image

Midjourney is Trying to Force Hollywood to Reveal How it Uses AI

Midjourney is pushing to compel major Hollywood studios to disclose how they use AI tools internally, framing the effort as a response to what it sees as a double standard. The studios have been vocal critics of AI companies while allegedly relying on similar technology behind the scenes. The legal and industry implications could be significant for both sides of the ongoing AI and entertainment debate.

Hollywood wants Seedance banned and reportedly also wants to keep using it
Video

Hollywood wants Seedance banned and reportedly also wants to keep using it

ByteDance's AI video tool Seedance has drawn the Motion Picture Association's first-ever cease-and-desist against an AI company, triggered by a viral clip featuring AI-generated likenesses of Brad Pitt and Tom Cruise. At the same time, industry insiders report that studios are quietly using the tool on a "don't ask, don't tell" basis - a contradiction that highlights the tension between Hollywood's public stance on generative AI and its private adoption of it.

ByteDance set to launch Seedance 2.5 with 3-minute AI video output
Video

ByteDance set to launch Seedance 2.5 with 3-minute AI video output

ByteDance is preparing to launch Seedance 2.5, the next version of its Dreamina-based video generation model, in July. The update's headline feature is the ability to produce AI-generated video clips of up to three minutes in length - a notable increase over what most current models offer. The longer output window could open up more practical use cases for creators working on short-form content.

Midjourney wants the Hollywood studios that sued it to show the court how they use AI
Image

Midjourney wants the Hollywood studios that sued it to show the court how they use AI

Midjourney has filed a legal request asking a court to compel Disney, Warner Bros., and Universal to disclose how they use AI internally. The move comes as part of the ongoing copyright lawsuit the studios brought against the image generation company. Midjourney appears to be building a case that the studios themselves are not strangers to the technology they are suing over.

Chinese AI video maker Kling raises $2 billion as it gears up for Hong Kong IPO
Video

Chinese AI video maker Kling raises $2 billion as it gears up for Hong Kong IPO

Kuaishou has secured roughly $2 billion in outside investment for Kling, its generative AI video division, as it prepares for a public listing in Hong Kong. The fundraise signals growing investor appetite for AI video technology and positions Kling as a well-capitalized competitor in an increasingly crowded field. The move also reflects a broader trend of Chinese AI companies seeking public market validation.

Google’s NotebookLM can sum up your research in a TikTok-style clip
Video

Google’s NotebookLM can sum up your research in a TikTok-style clip

Google's NotebookLM is gaining a short-video format that distills uploaded research into 60-second vertical clips, styled after the kind of content common on TikTok. The feature is rolling out to AI Ultra and Pro subscribers and pairs AI-generated narration with paper cutout-style visuals. It joins an existing set of NotebookLM output formats that already includes AI podcasts, cinematic videos, and visual explainers.

NoimosAI launches Creative Agent for brand assets
Image

NoimosAI launches Creative Agent for brand assets

NoimosAI has introduced Creative Agent, a tool designed to generate brand assets by drawing on patterns identified in high-performing market creatives. The system aims to ground visual output in data rather than purely open-ended generation. It positions itself as a practical option for teams looking to align creative production with proven market signals.

Start building with Nano Banana 2 Lite and Gemini Omni Flash
Multimodal

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Google has released Nano Banana 2 Lite and Gemini Omni Flash, two new models aimed at image generation and video editing respectively. Nano Banana 2 Lite is positioned as Google's fastest and most cost-efficient image model, while Gemini Omni Flash brings high-quality video output and conversational editing capabilities. Both models are now available for developers to build with.