gen‑ai.news
← Back
Multimodal

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Google has announced the availability of two new generative media models: Nano Banana 2 Lite, a lightweight image generation model, and Gemini Omni Flash, designed for video generation and conversational editing. The releases appear targeted at developers and builders who need capable models at lower cost and latency than flagship offerings.

Nano Banana 2 Lite sits at the efficient end of Google's image model lineup, prioritizing speed and cost over raw output quality. This kind of tiered model strategy is common across the industry - it lets developers prototype quickly, run high-volume workloads affordably, and integrate image generation into products where strict budget constraints apply. The "Lite" designation signals it is a smaller, distilled variant of a larger model in the Nano Banana 2 family.

Gemini Omni Flash extends the Gemini model family into video, offering what Google describes as high-quality video generation alongside conversational editing - a workflow where users can iteratively refine video output through natural language instructions. Conversational editing has been an active area of development across the industry, as it lowers the barrier for non-technical users to make precise adjustments without learning dedicated editing tools.

Together, these two models reflect a broader push by Google to make generative media capabilities accessible at different price and performance points. With both models now open for developer access, the near-term focus will likely be on how well Nano Banana 2 Lite holds up in real-world image quality benchmarks relative to competitors, and whether Gemini Omni Flash's conversational video editing delivers consistent, controllable results in practice.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

New benchmark confirms AI models still perform poorly at visual perception
Multimodal

New benchmark confirms AI models still perform poorly at visual perception

A new benchmark from Moonshot AI isolates visual perception from logical reasoning in multimodal models, and the results are sobering. No tested frontier model clears 60 percent accuracy, with GPT-4o leading by only a narrow margin. The findings suggest that many errors previously attributed to faulty reasoning may actually originate much earlier, at the point of reading the image itself.

You can now turn off Google Gemini’s visible watermarks
Multimodal

You can now turn off Google Gemini’s visible watermarks

Google has added a toggle in Gemini and its AI video tool Flow that lets users remove the visible "sparkle" watermark from AI-generated images, videos, and music. Even with the visible mark turned off, content will still carry invisible SynthID watermarks and C2PA metadata. The change affects content produced by Google's Nano Banana and Omni models.

Twitch streamers can now opt out from training Amazon’s AI
Multimodal

Twitch streamers can now opt out from training Amazon’s AI

Twitch has introduced an opt-out setting that lets streamers prevent their content - including streams, VODs, clips, chat logs, and channel images - from being used to train Amazon's generative AI models. The control applies to future training only and covers AI systems designed to generate or synthesize text, audio, images, or video. Non-generative AI features such as captions and safety tools are unaffected by the setting.