gen‑ai.news
← Back
Image

New AI model called "Count Anything" does exactly what it says, and that's harder than it sounds

New AI model called "Count Anything" does exactly what it says, and that's harder than it sounds

Counting objects in images sounds straightforward, but it has long been one of the more stubborn problems in computer vision. Earlier systems were typically trained for narrow domains - counting people in a crowd, for instance, or cells in a microscopy slide - and performed poorly when asked to generalize beyond those specific contexts. "Count Anything" is designed to break that constraint by accepting an open-ended text prompt as its only guidance, allowing a single model to handle a broad range of counting tasks without retraining or fine-tuning for each category.

According to the researchers, the model achieves roughly half the error rate of previous general-purpose counting systems on standard benchmarks. That is a meaningful improvement, since counting accuracy tends to degrade quickly as scenes become more complex or the target objects vary in size, occlusion, and appearance. The text-prompt approach means a user can simply describe what they want counted - "white blood cells," "people waiting in line," "cars in a parking lot" - and the model attempts to locate and tally each instance accordingly.

The underlying approach likely draws on the growing body of work that combines vision encoders with language models, allowing visual understanding to be steered by natural language descriptions rather than fixed category labels. This kind of open-vocabulary design is increasingly common in object detection and segmentation, and applying it to counting is a logical extension. The challenge is that counting demands not just identifying that something is present, but precisely localizing and distinguishing every individual instance - a harder requirement than simple classification or detection.

Despite the progress, the model has clear limits. Very dense configurations - tightly packed crowds or overlapping cells - still produce higher error rates, which is consistent with the difficulty of separating individual instances when they occlude one another heavily. Ambiguous or abstract text prompts also cause problems, since the model must interpret what the user means before it can begin counting. These limitations suggest that "Count Anything" is a solid step forward for general-purpose visual counting rather than a finished solution, and the domain will likely see continued iteration as training data and architecture choices improve.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Trump Shares Doctored Photo That Appears to Replace Official With Natalie Harp
Image

Trump Shares Doctored Photo That Appears to Replace Official With Natalie Harp

President Trump shared a doctored photo on Truth Social depicting himself alongside Chinese President Xi Jinping, with experts concluding that an official in the original image appears to have been replaced with Trump aide Natalie Harp. The incident adds to a growing pattern of manipulated images circulating at the highest levels of politics. It raises fresh questions about the role of AI-assisted photo editing in shaping public perception of diplomatic events.

You Can Edit Your Photos With Lightroom and Photoshop Inside Google Gemini
Image

You Can Edit Your Photos With Lightroom and Photoshop Inside Google Gemini

Adobe is bringing Lightroom and Photoshop editing capabilities directly into Google Gemini, allowing users to work with its tools without leaving the AI platform. The move extends Adobe's broader strategy of embedding its software into third-party AI environments, following earlier integrations with ChatGPT, Claude, Slack, and Microsoft Copilot.

No image
Image

Edit updates, thumbnail previews, and more

Midjourney has rolled out a set of interface updates on its alpha site, including live style previews that show how a prompt looks across different styles before committing to one. The update also brings thumbnail previews and refinements to the Edit workflow. These changes reflect the team's ongoing work to make the image generation and editing experience more interactive and informed.