Responsible AI
Disclosure, AI-Generated Content, and Provenance
Design disclosures, machine-readable marking, provenance, deepfake labeling, and authenticity workflows for AI-generated or manipulated content.
By the end you can
- Explain why AI content governance combines material disclosure, provenance, marking, verification, and remedy while distinguishing origin from truth
- Distinguish Visible label, Watermark or detector, and Signed provenance
- Identify evidence that connects creation disclosure to remedy and enforcement
- Design a review that moves from classify the transformation to operate remedy
Comparison
Visible label, Watermark or detector, or Signed provenance?
Labels speak to the reader, detectors guess, and signed provenance carries a claim about origin. None of the three says the content is true.
A visible label works without special software. It can also be vague, hidden or removed. It has to be timed and presented accessibly, and it works best when tailored to interpretation risk. A watermark or detector operates at scale. It returns false positives and false negatives, it degrades under transformation, and it should never be sole proof. Signed provenance supports chain-of-custody evidence. But its metadata may arrive already stripped, it depends on issuer trust and key security, and it says nothing about whether the scene is real.
The middle column is the one where the failure has been counted. Seven widely used GPT detectors were run over 91 human-written TOEFL essays and 88 US eighth-grade essays. The average false-positive rate on the TOEFL essays was 61.3%. The eighth-grade essays were classified near-perfectly. A 2023 paper in Patterns records how far the agreement went: “All detectors unanimously identified 19.8% of the human-written TOEFL essays as AI authored, and at least one detector flagged 97.8% of TOEFL essays as AI generated.” The errors are not sprayed evenly across writers. They concentrate on non-native English writers. The cost of a detector run is paid by a particular group of people.
Watermarks fail differently, and on purpose. A regeneration attack adds Gaussian noise to an image's latent representation, then reconstructs the image with a generative model. A 2024 NeurIPS paper measured what that does: “For the resilient watermark scheme RivaGAN, our regeneration attacks successfully remove 98% of the invisible watermarks while maintaining a PSNR above 30 compared to the original images.” The paper does not stop at the measurement. It proves the attack defeats any watermark that perturbs the image within a limited L2 range, whether or not that watermark has yet been invented. “Can degrade under transformation” is a proof, not a caveat.
Visible label
Communicates directly to the audience.
- Works without special software
- Can be vague, hidden, or removed
- Needs timing and accessible presentation
- Best when tailored to interpretation risk
Watermark or detector
Attempts to identify generated content statistically.
- Can operate at scale
- May have false positives and negatives
- Can degrade under transformation
- Should not be sole proof
Signed provenance
Carries cryptographic origin and edit claims.
- Supports chain-of-custody evidence
- Metadata may be absent or stripped
- Depends on issuer trust and key security
- Does not establish truth of the scene
Visual
What was generated, and whether it matters
Disclosure states what was generated. Materiality states whether that matters. Provenance records the history, detection covers the cases where the record is gone, and remedy handles the ones that already spread.
Creation disclosure states whether content was generated, transformed, translated, or edited with AI. Materiality explains whether identity, event, speech, or evidence was significantly changed. Provenance records source, tool, credentials, and transformation history. Detection and verification supply signals and channels when metadata are missing or disputed. Remedy and enforcement cover correction, removal, appeal, attribution, and incident response.
The five stages are not interchangeable. The later ones are not spare capacity for the earlier ones. Detection exists because provenance is routinely gone by the time a file reaches a reader. Remedy exists because content spreads before any of the first four stages is consulted.
- 1
Creation disclosure
States whether content was generated, transformed, translated, or edited with AI.
- 2
Materiality
Explains whether identity, event, speech, or evidence was significantly changed.
- 3
Provenance
Records source, tool, credentials, and transformation history.
- 4
Detection and verification
Provides signals and channels when metadata are missing or disputed.
- 5
Remedy and enforcement
Supports correction, removal, appeal, attribution, and incident response.
Materiality carries the weight
A disclosure should tell the reader what role AI played, how materially the content was altered, where the material came from, and what all that changes about how to read it. Provenance records origin and transformation history. It does not certify truth, consent, legality, or harmlessness. Materiality carries most of that weight. Without it the same four words get attached to a color correction and to a fabricated quotation.
Controls can include user-facing labels, machine-readable markers, watermarking, signed provenance, content credentials, detection tools, platform policy, and verification channels. That list looks like a designer's menu. In at least one jurisdiction it is not. China wrote the split into law. The Measures for Labeling AI-Generated Synthetic Content were issued on 7 March 2025 by the Cyberspace Administration of China with three other ministries, and have been in force since 1 September 2025. Article 3 defines explicit labels — text, sound or graphics a user can plainly perceive — and implicit labels, technical marks carried in the file data. Article 5 tells the implicit layer what it has to carry: the content's attribute information, the provider's name or code, and a content number. Article 10 then protects the mark itself: “任何组织和个人不得恶意删除、篡改、伪造、隐匿本办法规定的生成合成内容标识,不得为他人实施上述恶意行为提供工具或者服务,不得通过不正当标识手段损害他人合法权益。”
So the visible layer and the machine-readable layer are two obligations, addressed to two audiences. Named fields on one side, an unlawful act on the other. How well each control holds up still varies under editing, screenshotting, recompression, adversarial removal, and missing metadata. And a screenshot defeats every machine-readable item on the list at once. No malice, nobody deleting anything. What is left is the wording a human can read. That is why the wording has to be built to carry the load, not to announce that a tool was involved.
The words chosen for a label are a control in their own right, because they are what survives when a screenshot strips everything machine-readable.
Case
Article 50(2) marks, Article 50(4) discloses, and Article 99 prices both
Regulation (EU) 2024/1689 splits the duty in two and puts a number behind it. Article 50(2) puts the machine-readable marking duty on the provider: outputs of AI systems generating synthetic audio, image, video or text must be marked in a machine-readable format and detectable as artificially generated or manipulated. Article 50(4) turns to whoever publishes — the deployer: “Deployers of an AI system that generates or manipulates image, audio or video content constituting a deep fake, shall disclose that the content has been artificially generated or manipulated.” The tool vendor marks the file. The publisher tells the audience. Neither obligation absorbs the other. A newsroom that points at its vendor's metadata has not discharged its own.
A breach of either is priced the same way. Article 99(4)(g) sets administrative fines of up to EUR 15 000 000 or 3 % of total worldwide annual turnover, whichever is higher. Two dates are already live. Article 50 became applicable on 2 August 2026. The Digital Omnibus on AI of 8 July 2026 then added an Article 111(4): providers whose systems were already on the market before 2 August 2026 have until 2 December 2026 to comply with Article 50(2).
The standard the mark travels on is older than the duty. Six companies — Adobe, Arm, the BBC, Intel, Microsoft and Truepic — founded C2PA on 22 February 2021, to standardise how origin and edit claims move with a file. Version 2.1 of the Content Credentials technical specification is dated 20 September 2024.
A screenshot of the image still strips every one of those markings. The regulation survives the screenshot and the specification survives it. The mark does not. The fine is attached to a duty whose evidence has just left the building.
Key idea
Every label moved belief; the vague ones moved nothing else
A disclosure that merely says “AI was used” may satisfy a formal checkbox while failing to say what was meaningfully altered. And a label should not imply that unlabelled content is authentic, or that labelled content is false. Both halves of that warning have now been measured. Two preregistered survey experiments, with 7,579 Americans, set process-based labels — which describe how the content was made — against harm-based labels, which describe its potential to mislead. The result, published in PNAS Nexus in 2025: “Overall, we find that all of the labels we tested significantly decreased participants' belief in the presented claims. However, in both studies, labels that simply informed participants that content was generated using AI tended to have little impact on respondents' stated likelihood of engaging with their assigned post.”
Read that as a design result rather than a curiosity. The wording most disclosure regimes ask for — this was generated using AI — shifts what a reader believes. It leaves what a reader does roughly where it was. The same paper flags the second failure directly: an “implied authenticity” risk, in which labelling a subset of content raises the credibility of everything that goes unlabelled. No marking ecosystem covers every tool and every distribution channel, so the unlabelled set is always large and always mixed. Whoever publishes has to combine provenance, disclosure, authentication, verification, education, and response. And they have to choose the words on the assumption that the words, not the metadata, are what the audience will act on.
Every label tested cut belief in the claim; the ones that only said AI was used barely moved what readers said they would do with the post.
Analogy
A chain-of-custody label on evidence
Evidence in a case file travels with a record of who collected it, who moved it, and who altered it. The chain establishes where the item came from without establishing that the scene it shows is real.
Four national security agencies put that distinction in writing and signed it jointly. On 29 January 2025 the US National Security Agency, with its Australian, Canadian and UK counterparts, published a joint information sheet on Content Credentials. Content Credentials metadata, it says, does not let a consumer decide whether content is true; it supplies only contextual information about authenticity. The boundary this lesson keeps drawing is, in that document, operational doctrine for four governments.
A physical exhibit has one custody record and one copy. A file can be re-encoded, stripped of its metadata, and forwarded a million times. The copy that reaches the reader is usually the one with no record attached. The same sheet names that as the systemic weakness: “Any processes that detach or strip metadata degrade the prospects of implementing and/or consuming Content Credentials.” It is why C2PA adds Durable Content Credentials — watermark plus fingerprint matching — as a recovery path for the copy that arrives bare. And it is why that recovery path inherits the watermark column's own problem rather than escaping it.
Provenance supports authenticity claims about origin and edits, not truth of the underlying assertion.
Example
A looped seven-second clip and a policy that only covered video
The defect of a label keyed to how content was made, rather than to what it does, has been adjudicated at platform scale. A looped seven-second clip of President Biden stayed up. The rule it had been judged under was dismantled in the same decision. On 5 February 2024 the Meta Oversight Board upheld the decision to leave the clip online, in case 2023-029-FB-UA. The Manipulated Media policy covered only video, only AI-generated or AI-altered content, and only people appearing to say words they did not say. The Board's verdict on that scope: “In its current form, the Manipulated Media policy fails to clearly specify the harms it is seeking to prevent, and the scope of its prohibitions is incoherent in both policy and technical terms.”
The recommendation is the part worth copying. The Board proposed that Meta stop removing manipulated media where nothing else is violated, and attach a label instead: the content is significantly altered and could mislead. Materiality and consequence in the label, not the production method. Meta answered on 5 April 2024. Labelling would begin in May 2024, and removals under the manipulated video policy alone would stop in July 2024.
- Scope keyed to method: The policy asked how the content was made — video, AI, fabricated speech. A clip that was none of those three fell outside a rule written for the harm it caused.
- Audience need: Readers need to know whether a real event, person, or quote was materially altered. That is why the Board asked for a label saying the content is significantly altered and could mislead, rather than one saying a tool was used.
- Provenance: Source and edit history may be available through signed metadata, but the Board's remedy did not wait on it — the label is what reaches the viewer.
- Distribution loss: Platforms can strip metadata during upload or recompression, so the machine-readable layer is missing from most copies in circulation before any policy is applied.
- Truth gap: Authentic provenance can document origin without proving the depicted claim is true. Meta's move from removal to labelling in July 2024 is the response shifting from suppressing a file to qualifying how it should be read.
Steps
Classify the transformation before the label
Classify the transformation first. What counts as material, and therefore what has to be disclosed, depends on that answer.
Step 1 classifies the transformation: generation, restoration, editing, dubbing, translation, or personalization. Step 2 assesses materiality and risk — identity, evidence, public interest, deception, vulnerability. Step 3 chooses layered signals: visible labels, metadata, provenance, watermarking, verification channels. Step 4 tests persistence and comprehension across transformations, platforms, accessibility, and audience interpretation. Step 5 operates remedy: correct, remove, attribute, appeal, notify, and learn from incidents.
Step 4 is where this lesson's numbers belong, because both halves of it have been measured by other people already. Persistence: 98% of invisible RivaGAN watermarks removed by a regeneration attack while PSNR stayed above 30, and metadata detached by any process that touches the file. Comprehension: two preregistered experiments over 7,579 respondents, in which every label cut belief and the ones that merely reported AI involvement barely moved engagement. A team that skips step 4 is not missing a formality. It is declining to find out which of those two results applies to its own workflow.
Step 1 is not a formality either. The Manipulated Media policy classified by production method and was held incoherent for it. The class it keyed on was not the class that carried the harm. A label written before the transformation is classified is a label written before anyone knows what it has to say.
1. Classify the transformation
Generation, restoration, editing, dubbing, translation, or personalization.
2. Assess materiality and risk
Consider identity, evidence, public interest, deception, and vulnerability.
3. Choose layered signals
Use visible labels, metadata, provenance, watermarking, and verification channels.
4. Test persistence and comprehension
Evaluate transformations, platforms, accessibility, and audience interpretation.
5. Operate remedy
Correct, remove, attribute, appeal, notify, and learn from incidents.
The operational conclusion — AI content disclosure and provenance
Labels and provenance are signals about origin. They survive some transformations and not others, so persistence and comprehension are both measurable — and, in this lesson, measured. Seven detectors averaged a 61.3% false-positive rate on 91 human-written TOEFL essays, a cost borne by non-native English writers. A regeneration attack removed 98% of invisible watermarks, with a proof that covers schemes not yet built. Every label tested cut belief across 7,579 respondents, while the process-based wording left engagement roughly untouched. Four signals agencies stated that Content Credentials supply contextual information about authenticity and not a verdict on truth. Article 50(4) obliges the deployer to disclose, Article 50(2) obliges the provider to mark, and Article 99(4)(g) prices a breach of either at up to EUR 15 000 000 or 3 % of total worldwide annual turnover.
None of those figures decides anything by itself. Decide which stripping, false-positive, or comprehension result would force the reviewer to redesign, restrict, remedy, or retire the system. Write that threshold down before the next measurement arrives. A number without a pre-committed threshold is an observation, not a decision.
Key takeaways
- Disclosure should explain the role and materiality of AI rather than repeat one vague label. The Meta Oversight Board held Meta's Manipulated Media policy incoherent in case 2023-029-FB-UA precisely because it keyed on how content was made instead of what it could do.
- Provenance records origin and transformation history, not truth or consent. The US, Australian, Canadian and UK cyber agencies said so jointly on 29 January 2025: Content Credentials supply contextual information about authenticity and do not let a consumer decide whether content is true.
- Visible labels, watermarks, detectors and signed metadata fail in different ways and by different amounts. Seven GPT detectors averaged 61.3% false positives on 91 human-written TOEFL essays, and a regeneration attack removed 98% of RivaGAN watermarks while keeping PSNR above 30.
- Metadata can be stripped, signals can degrade and detectors can be wrong. That is why Article 10 of China's Measures makes maliciously deleting, tampering with, forging or concealing the label unlawful, and why C2PA adds Durable Content Credentials as a recovery path.
- Unlabelled content should not be assumed authentic merely because a marker is absent. The 2025 PNAS Nexus study flags the implied-authenticity risk: labelling a subset of content raises the credibility of everything left unlabelled.
- A responsible program includes verification, correction, appeal and incident response, plus the deployer's own duty under Article 50(4) to disclose deep fakes, backed by fines of up to EUR 15 000 000 or 3 % of total worldwide annual turnover.