Exclusive: Early look at the next Gemini desktop upgrade

Google’s Gemini desktop client for Mac is set to gain Voice Mode, Stream to Cursor, Omni video generation, and Spark-powered features.

Google’s Gemini desktop client for Mac is set to gain Voice Mode, Stream to Cursor, Omni video generation, and Spark-powered features.
Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.
Free. Unsubscribe any time. No spam, ever.
Alibaba's Qwen team has moved Qwen3.8-Max from preview to general availability, marking it as the most capable model in the Qwen lineup to date. The mixture-of-experts model carries 2.4 trillion parameters and supports text, image, and video input over a 1 million-token context window. Open weights are expected to follow next week.
NVIDIA and Hugging Face have joined forces to bring large-scale fine-tuning of image and video diffusion models into the NeMo Automodel framework, integrated with the Diffusers library. The collaboration aims to make distributed training more accessible for teams working with models that would otherwise be difficult to fine-tune on limited hardware. The result is a more streamlined path from a pretrained diffusion model to a customized one, without requiring deep infrastructure expertise.
Thinking Machines Lab has released Inkling, a 975B-parameter open-weights multimodal model built on a Mixture-of-Experts architecture that keeps only 41B parameters active at any given time. Licensed under Apache 2.0, it accepts text, image, and audio inputs and offers a 1M-token context window. Rather than competing for top benchmark rankings, the model is positioned as a customizable base with adjustable reasoning depth.