Provenance
How platforms label AI content, and what triggers the label

Quick answer
Platforms label AI content from three inputs: what the uploader discloses, provenance data found in the file (C2PA manifests and IPTC source type), and their own detection. Most require creators to disclose realistic generated or altered media. Labels on link previews are rarer than on native uploads. Publishers avoid surprises by disclosing honestly, keeping metadata accurate, and checking each platform's current help pages.
On this page
The three inputs behind a label
Creator disclosure
The uploader ticks a box or adds a label when posting. Several platforms make this mandatory for realistic synthetic media and may penalise accounts that repeatedly fail to disclose.
Metadata in the file
Industry provenance data, such as a C2PA manifest or an IPTC digital source type, tells the platform that a generative tool made or changed the file. This is the input publishers influence most.
Platform detection
Classifiers and fingerprinting look for generated media without any metadata. Accuracy varies, and detection is the usual source of labels that surprise honest creators.
How the main platforms describe their approach
| Platform | What its public policy says, in brief | Source |
|---|---|---|
| Meta (Facebook, Instagram, Threads) | Labels content when it detects industry-standard AI indicators or the user discloses it; asks users to disclose realistic video and audio | Meta newsroom |
| YouTube | Creators must disclose realistic altered or synthetic content; labels appear in the description or on the player for sensitive topics | YouTube Help |
| TikTok | Requires labels on realistic AI-generated content and can apply labels automatically from Content Credentials | TikTok Support |
| X | A synthetic and manipulated media policy that can label or limit misleading media | X Help Center |
These summaries are deliberately short. Policies are revised often, and the exact thresholds (what counts as realistic, which topics are sensitive) are defined on the platforms' own pages, which you should read before a campaign.
Links versus native uploads
Most labelling happens on media uploaded directly to a platform. When a reader shares a link, the network builds a preview from your Open Graph tags and fetches the image named in og:image; whether it inspects that image for provenance data varies by network. Because Content Credentials often vanish when images are resized, a preview image rarely carries enough data to trigger a label either way. Your page is where disclosure has to live.
Regulation is catching up
In the European Union, Article 50 of the AI Act (Regulation 2024/1689) sets transparency obligations that apply from August 2026: providers of generative systems must mark output in a machine-readable way, and deployers who publish deepfakes, or generated text meant to inform the public on matters of public interest, must disclose it, with an exception for text that has undergone human review and for which a person holds editorial responsibility. That exception is one more reason to keep a named, accountable editor in your byline.
A checklist before you publish
Before a piece with generated or heavily edited media goes out, check three things. First, that the caption or disclosure line on the page matches what actually happened. Second, that the metadata in the file says the same thing, because a platform will trust the file over your caption. Third, that the share image named in og:image is the one you intended, not an automatically generated crop of a different picture. Most labelling disputes start with one of those three drifting apart.
Avoiding false labels
- Check the metadata your editing tools write. Some editors mark any file touched by a generative feature, even a small background fill, and platforms may label the whole image.
- If a label is wrong, use the platform's appeal or edit option and keep the camera original, which is your best evidence.
- Do not strip metadata to dodge labels. It removes your own evidence and does not stop detection-based labels.
- State your policy on the page, as described in what publishers should do about AI content in feeds.
More on this topic sits in the provenance and authorship section: start with byline schema in WordPress: marking up a human author, then authorship tags for shared links: author, article:author and fediverse:creator.
Questions
Why was my real photo labelled as AI?
Usually because an editing step wrote generative metadata into the file, or a detector misread it. Appeal and keep the original.
Do link previews get AI labels?
Less often than native uploads. Preview images are usually resized copies without provenance data.
Does the EU AI Act apply to my blog?
Its transparency duties can apply to deployers publishing deepfakes or generated public-interest text in the EU, with an exception for human-reviewed text under editorial responsibility. Take advice for your situation.
Hannah Voss, Editor. Checks every guide against a working WordPress install and the networks' current documentation. Last reviewed September 2026.
Related reading
- Content Credentials on shared images: what survives a shareWhat C2PA Content Credentials are, why WordPress image sizes and social uploads often strip them, and how to keep provenance on the images you share.4 min read
- AI-generated content in social feeds: what publishers should doHow AI-generated images and text change what happens when your posts are shared, and the practical steps publishers can take: disclosure, provenance, bylines.4 min read
- Byline schema in WordPress: marking up a human authorMark up article authors with schema.org Person data in WordPress: the fields search engines read, author pages, sameAs profiles and common mistakes.4 min read
- Share buttons and GDPR: what needs consentWhich share buttons need consent under GDPR and ePrivacy rules, why plain share links usually do not, the two-click pattern, and privacy policy wording.3 min read