gen‑ai.news
← Back
Multimodal

Adobe adds AI agents to Photoshop, Premiere, and more Creative Cloud apps

Adobe adds AI agents to Photoshop, Premiere, and more Creative Cloud apps

Adobe is bringing AI agent functionality to its flagship Creative Cloud applications - Photoshop, Premiere, and others - under what the company calls a "creative agent" framework. Rather than manually executing a series of individual steps, users can describe what they want to achieve, and the agent works through the necessary operations on their behalf. The feature is also being made available through third-party AI platforms, including ChatGPT and Claude.

The practical implication is a shift in how users interact with tools they may have spent years learning. Tasks that previously required navigating menus, applying adjustments in sequence, or scripting repetitive actions can now be initiated through a conversational prompt. For professionals managing high volumes of work, that kind of delegation to an automated process can meaningfully reduce time spent on execution versus creative decision-making.

Adobe has been building out its AI capabilities steadily through its Firefly model family, which underpins many of the generative features already present in Photoshop and other apps - such as Generative Fill and text-to-image tools. The agent layer sits on top of these capabilities, adding the ability to chain multiple actions together rather than triggering them one at a time. Extending this to external AI assistants like Claude and ChatGPT suggests Adobe is also thinking about users who want to manage creative work from within general-purpose AI interfaces, rather than opening dedicated desktop applications.

The broader context is one of increasing competition among creative software platforms to offer agentic functionality. As AI models become more capable of handling sequential, context-dependent tasks, the line between "assistant" and "autonomous operator" continues to shift. Adobe's approach - tying the agent to established professional applications with existing export pipelines and file format support - positions it as an incremental addition to familiar workflows rather than a standalone product. How much creative control users retain, and how the agent handles ambiguous or conflicting instructions, will likely determine how widely the feature is adopted in professional settings.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

ByteDance launches SeedRealtime full-duplex AI model
Multimodal

ByteDance launches SeedRealtime full-duplex AI model

ByteDance has released SeedRealtime, a full-duplex multimodal model that handles audio, video, and text within a single unified system. Unlike turn-based conversational AI, it supports continuous, overlapping dialogue with natural timing and proactive responses. The model represents a step toward more fluid human-computer interaction without the stop-and-start feel of earlier voice interfaces.

Mistral releases Shieldstral for multimodal moderation
Multimodal

Mistral releases Shieldstral for multimodal moderation

Mistral has released Shieldstral, a 3-billion-parameter open-weights multimodal safety classifier designed to moderate text, images, and combined text-image content. The model is built for customizability, giving developers and organizations direct control over how moderation policies are applied. Its open-weights nature means it can be deployed and adapted without relying on a hosted API.

No image
Multimodal

Evaluating Multimodal Vision Models with Moonshot PerceptionBench Using Robust Data Loading and Automated Judging

Moonshot's PerceptionBench offers a structured way to measure how well multimodal vision models handle fine-grained visual tasks, from OCR and object counting to depth understanding and hallucination detection. A new tutorial walks through building a complete evaluation pipeline, including environment setup, dataset loading, and automated answer judging. The workflow is designed to run in Google Colab, making it accessible for researchers without dedicated infrastructure.