Instagram’s AI detection is a mess (again)

Instagram's AI content labeling system is intended to give users a quick, reliable signal when an image has been created or significantly altered by generative AI. Over recent weeks, however, the feature has been behaving erratically - applying the "AI Content" label to images that were produced with conventional editing tools, and failing to catch imagery that was actually generated by AI models.
The misfires appear to stem from several different causes. A number of users have reported that the label was triggered by minor edits made in tools such as Canva, specifically its background removal feature. Background removal is a longstanding, non-generative image editing technique, and its presence as a trigger points to a broader problem: Meta's detection system may be reading low-level file metadata or processing artifacts rather than genuinely identifying AI-generated content.
This matters because content labels only work if they are applied consistently and accurately. A label that appears on an ordinary vacation photo edited in Canva trains users to ignore it. At the same time, AI-generated imagery that slips through without any label does the opposite of what the feature is supposed to accomplish. Together, the two failure modes leave the platform in a worse position than if no labeling system existed at all - users have no reliable baseline for what the tag actually means.
Meta has been under pressure from regulators, researchers, and civil society groups to improve transparency around synthetic media, and visible labels were positioned as a practical step toward that goal. The company has also been working with the Coalition for Content Provenance and Authenticity (C2PA), an industry group developing technical standards for tracking content origin through embedded metadata. Whether the current labeling problems are rooted in flawed metadata parsing, an overly broad detection model, or gaps in third-party tool coverage is not yet clear. Until the underlying logic is made more transparent and the accuracy improves, the feature risks doing more to erode trust on the platform than to restore it.