Skip to content
AI.info

Computer vision

Privacy, Fairness, Security, and Media Provenance

Integrate privacy, demographic evaluation, adversarial threats, access control, media provenance, and accountable product policy.

By the end you can

Vision systems can infer more than the product asks for

A camera intended to count occupancy may also capture faces, health cues, screens, badges, behavior, and bystanders; stored embeddings can enable future linkage even when raw images are deleted.

Governance starts before training: purpose limitation, field of view, resolution, retention, access, inference scope, and acceptable secondary use.

European law now draws one of these lines in statute rather than in a design review. Article 5(1)(e) of the AI Act prohibits “the placing on the market, the putting into service for this specific purpose, or the use of AI systems that create or expand facial recognition databases through the untargeted scraping of facial images from the internet or CCTV footage”. The European Commission lists that practice among the ones it bans outright.

What the line is worth is easier to see once it has been priced. Clearview AI built a database of more than 30 billion photos scraped from the internet, and turned every face in it into a unique biometric code. The people depicted did not know, and did not consent. On 3 September 2024 the Dutch Data Protection Authority fined Clearview EUR 30.5 million, with further penalties of up to EUR 5.1 million if it fails to comply with the orders attached. “Facial recognition is a highly intrusive technology, that you cannot simply unleash on anyone in the world”, said Aleid Wolfsen, chairman of the Dutch Data Protection Authority. The underlying decision is dated 16 May 2024. Clearview did not object, so it cannot appeal.

Nothing in that database required a new sensor. Every one of the 30 billion faces was already a published image, and the inference — one biometric code per person — was added afterwards by someone the depicted person never met. That is the whole distance between what a system collects and what it can be made to know.

The least risky visual feature is often the one never collected.

Visual

Privacy risk follows the visual data lifecycle

Controls should exist at every transition rather than only at storage. The last step is the one most often written as an intention rather than a procedure — and a court has already written it as a procedure.

Rite Aid ran face recognition in its stores. The FTC sued on 19 December 2023, and in February 2024 a federal judge signed the settlement. It bars Rite Aid for five years from deploying or using any facial recognition or analysis system in any retail store, retail pharmacy or online retail platform. It caps retention of biometric information at five years. And it puts a clock on what already exists: “Within forty-five (45) days after the effective date of this Order, delete or destroy all photos and videos of consumers used or collected in connection with the operation of a Facial Recognition or Analysis System prior to the effective date of this Order, and any data, models, or algorithms derived in whole or in part therefrom, and provide a written statement to the Commission, sworn under penalty of perjury, confirming that all such information has been deleted or destroyed;”

Read that against step 3 of the lifecycle. The order does not stop at the pictures. The data, models and algorithms derived from them fall under the same forty-five-day clock, and a named person has to swear to the result under penalty of perjury. That is what “verifiable deletion” costs when someone else defines it.

What produced the order is a lifecycle failure at the other end. The system ran from October 2012 to July 2020. The FTC alleged that it generated thousands of false-positive matches, and “was more likely to generate false positives in stores located in plurality-Black and Asian communities than in plurality-White communities”.

FigureProcess · 5 steps
  1. 1. Capture

    Limit field of view, resolution, duration, and bystander exposure.

  2. 2. Transform

    Crop, redact, aggregate, or process on device where feasible.

  3. 3. Represent

    Protect embeddings, metadata, labels, and derived attributes.

  4. 4. Use and share

    Restrict purpose, authorization, export, and human access.

  5. 5. Retain or delete

    Enforce schedules, legal holds, revocation, and verifiable deletion.

Comparison

Fairness questions depend on the decision being made

No single parity metric resolves all visual-product harms.

The measurements exist, and they are not close. In 2018 Buolamwini and Gebru tested three commercial gender classification systems on a benchmark balanced by gender and skin type. They reported “that darker-skinned females are the most misclassified group (with error rates of up to 34.7%)”. At the other end of the same benchmark: “The maximum error rate for lighter-skinned males is 0.8%”. One aggregate accuracy figure averages those two groups together and reports neither.

NIST ran a far larger study the following year. It put “a total of 18.27 million images of 8.49 million people through 189 mostly commercial algorithms from 99 developers”, and the December 2019 report states the result plainly: “Our main result is that false positive differentials are much larger than those related to false negatives and exist broadly, across many, but not all, algorithms tested. Across demographics, false positives rates often vary by factors of 10 to beyond 100 times. False negatives tend to be more algorithm-specific, and vary often by factors below 3.” Which error you choose to measure decides whether you see a disparity of 3 or one of 100. The same systems, the same images, two different findings.

Outcome equity is where the third column stops being a metric at all. Detroit police arrested Robert Williams in January 2020. Face recognition had flagged his driver's licence photo as a match to security video from a 2018 theft at a Shinola watch store. He was held for about thirty hours. On 28 June 2024 the parties settled Williams v. City of Detroit. Detroit agreed to pay $300,000. It also agreed to policies: no arrest on a face-recognition result alone, no photo lineup drawn solely from a face-recognition lead, training on the technology's higher misidentification rates for people of colour, and an audit of every case since 2017 in which face recognition produced an investigative lead followed by an arrest or arrest warrant. The court keeps jurisdiction to enforce the agreement for four years. Phil Mayor, senior staff attorney at the ACLU of Michigan, put it this way in the press release announcing the settlement: “Police reliance on shoddy technology merely creates shoddy investigations.”

Notice what the remedy is made of. Not a recalibrated threshold: sign-off rules, lineup rules, training, a retrospective audit back to 2017, and four years of court supervision. The error was in the model; the harm and the fix were both in the workflow around it.

FigureComparison · 3 columns

Quality parity

Do capture and preprocessing work comparably across groups and contexts?

  • Exposure and focus
  • Occlusion and device fit
  • Image-quality rejection
  • Example: selfie enrollment

Error parity

Who experiences false positives, false negatives, or localization errors?

  • Threshold-specific
  • Intersectional slices
  • Uncertainty intervals
  • Example: face verification

Outcome equity

How do workflow, access, delay, appeal, and enforcement affect people?

  • Beyond model metrics
  • Institutional context
  • Remedy matters
  • Example: benefits review

Key idea

Removing a sensitive column does not remove sensitive information

Faces, names, locations, devices, clothing, neighborhoods, and correlated context can encode protected attributes or their proxies. A model can produce disparate outcomes without ever receiving an explicit demographic field. Rite Aid's system was not given a demographic input, and the FTC still alleged that its false positives fell more heavily on stores in plurality-Black and Asian communities.

Keep protected attributes available in controlled evaluation when lawful and appropriate. Otherwise disparities may become impossible to measure. Nobody could have reported 34.7 percent against 0.8 percent, or the factor-of-100 false-positive spread, without group labels to slice by.

Fairness requires visibility into outcomes, not blindness to the variables needed for audit.

Figure

An aggregate accuracy figure averages the two groups together and reports neither.

Analogy

A building protected by more than its front-door lock

A building can be secured at the front door while delivery entrances, badge printers, maintenance contractors, and surveillance feeds go ignored. A strong front lock cannot protect every path.

A building's entrances can be walked and counted one at a time, while ML attacks can exploit data, models, sensors, software, and human workflows simultaneously. A system-level threat model is what the building argues for.

The delivery entrance has a documented name. On 8–9 March 2021 attackers reached the cloud camera platform of Verkada, and the company's own incident report describes the route: “The attackers’ vector of entry was through a misconfigured customer support server exposed to the internet. Once the attackers accessed that server, they found customer support administrator credentials and used those to log into a customer support web interface, where they accessed customer devices using internal support functionality that emulated user sessions.” The exposed machine was an internet-facing customer-support Jenkins server.

The report also counts the damage: 97 customers had camera video or images viewed, 8 had access-control badge credentials potentially exposed, 8 had wifi credentials taken, and 4,530 cameras may have been accessed. “Fifteen People Analytics searches for images of persons were performed in five organizations.” The FTC's account of the same breach puts over 150,000 live customer cameras within the attackers' reach. The fifteen searches are the sentence to sit with: the side door did not only reach the video, it reached the inference layer built on top of it.

On 30 August 2024 the FTC announced a proposed stipulated order, filed by the DOJ and effective only once the court approves it. It resolves charges that Verkada failed to require unique and complex passwords, adequately encrypt customer data and implement secure network controls. For those security failures the order requires a comprehensive information security program with third-party audits, and bars misrepresentations about Verkada's privacy and data security practices. The $2.95 million penalty in the same matter is the largest the FTC has obtained for a CAN-SPAM Act violation, and it is imposed for Verkada's separate email-marketing conduct, not for the security failures.

Model robustness is one layer of visual-system security.

Example

Attack surfaces across a computer vision product

Threats can alter inputs, training, software, or decisions. Two of the seven below have been measured in public, and the measurements are the reason to take the other five seriously.

Start with physical evasion. In 2018 Eykholt and colleagues stuck black and white stickers on a real stop sign, then filmed it from a moving vehicle. Their abstract reports both numbers: “With a perturbation in the form of only black and white stickers, we attack a real stop sign, causing targeted misclassification in 100% of the images obtained in lab settings, and in 84.8% of the captured video frames obtained on a moving vehicle (field test) for the target classifier.” They call the attack algorithm Robust Physical Perturbations (RP2). The two percentages are the same attack under two conditions, a lab number and a field number, and only one of them is the deployment number. NIST's adversarial machine learning taxonomy files the work under “physically realizable attacks”.

Then poisoning, which turns out to be a supply-chain problem wearing a data-quality costume. Dataset publishers hand out lists of URLs rather than images. NIST's taxonomy spells out what follows: “dataset publishers may provide a list of URLs to constitute a training dataset, and attackers may be able to purchase some of the domains that serve those URLs and replace the site content with their own malicious content”. Carlini and colleagues priced the attack in 2023: “By exploiting specific invalid trust assumptions, we show how we could have poisoned 0.01% of the LAION-400M or COYO-700M datasets for just $60 USD.” Ten popular datasets were vulnerable at the time. NIST's recommended defence is unglamorous — publish cryptographic hashes so downloaders can verify what they downloaded.

Bullets five and seven are therefore one attack seen from two ends: the poisoned example arrives because a dependency in the supply chain, an expired domain, changed hands for the cost of a dinner.

  • Physical evasion: Patterns, coverings, lighting, or placement reduce detection — black-and-white stickers gave 100% targeted misclassification in the lab and 84.8% of frames from a moving vehicle.
  • Digital perturbation: Modified pixels target model behavior after image ingestion.
  • Presentation attack: Photos, replays, masks, or synthetic media impersonate a live source.
  • Sensor or pipeline injection: A compromised device sends fabricated frames.
  • Training-data poisoning: Manipulated examples or labels change learned behavior — 0.01% of LAION-400M or COYO-700M for about $60.
  • Model extraction or inversion: Queries reveal model behavior or information about training data.
  • Dependency compromise: Malicious weights, libraries, or build artifacts enter the supply chain, including the expired domains behind a dataset's list of URLs.

Steps

Threat-model a vision system from attacker goal to recovery

Security tests should reflect realistic capabilities and consequences.

Step 2 lists query budget beside equipment and physical proximity, which reads like padding until someone spends the budget. In 2016 Tramèr and colleagues attacked two live commercial services through nothing but their prediction APIs: “Given these practices, we show simple, efficient attacks that extract target ML models with near-perfect fidelity for popular model classes including logistic regression, neural networks, and decision trees. We demonstrate these attacks against the online services of BigML and Amazon Machine Learning.” Suppressing confidence values does not close the attack.

NIST records the same result in its section on model extraction: “The first model stealing attacks were shown by Tramer at al. [376] on several online ML services for different ML models, including logistic regression, decision trees, and neural networks”. It classes model extraction as a stepping stone to stronger white-box attacks. Steal a copy, then attack the copy at leisure with full knowledge of its weights.

That is what turns step 4's rate limits from a formality into a control. An adversary who can only ask questions is still an adversary. The number of answers you are willing to give is part of the attack surface, and it is one of the few parts a product team sets directly.

FigureProcess · 5 steps
  1. 1. Define assets

    List identities, images, decisions, credentials, models, and operational availability.

  2. 2. Define attackers

    Specify access, knowledge, equipment, query budget, and physical proximity.

  3. 3. Map attack paths

    Trace sensor, network, preprocessing, model, API, UI, and update channels.

  4. 4. Test controls

    Evaluate authentication, provenance, rate limits, detection, fallback, and least privilege.

  5. 5. Plan response

    Define logging, containment, rollback, disclosure, remediation, and evidence retention.

Key idea

Media provenance records a history; it does not certify reality

A signed manifest can link an asset to a capture or editing process and reveal later modifications; credentials may be absent, stripped, forged outside the trust chain, or attached to a staged scene.

Use provenance to reason about origin and transformation, not to declare semantic truth; combine it with checking the source, weighing the context, and forensic analysis.

The specification says as much about itself. The C2PA's own explainer for Content Credentials states that “Provenance information alone cannot tell you whether the digital content is true, accurate or factual”, and adds that “No assumption should be made about the trustworthiness of a particular asset purely based on its usage of Content Credentials”.

NIST reached the same place from the other direction, in a November 2024 report on reducing the risks posed by synthetic content. “Metadata recorded within a file can similarly be stripped altogether, as it often is when files are shared (e.g., via social media platforms)”, the report notes, and digital content transparency “may contribute to trustworthiness but does not guarantee it, and in some cases may undermine it”. A missing credential is therefore evidence of very little. A present one is evidence about a file's history, not about the world it depicts.

Authentic provenance and truthful depiction are related but distinct claims.

Example

Practice: review a smart-building occupancy camera

The product promises energy savings by estimating room use without identifying individuals. Every item below has a precedent earlier in this lesson: a scraped database nobody consented to, a forty-five-day deletion clock that reached the derived models, a false positive that cost thirty hours of detention, and a support interface that emulated user sessions.

  • Reduce the field of view, resolution, and retention to the minimum needed.
  • List unintended inferences that remain possible from bodies, clothing, and schedules, including any embedding or code that could later link one person across rooms and days.
  • Create quality and error slices for lighting, mobility aids, clothing, and room geometry, and report false positives separately from false negatives rather than one accuracy figure.
  • Threat-model unauthorized feeds, model updates, prediction-API query volume, support interfaces that emulate user sessions, and access to embeddings.
  • Write a deletion, incident, appeal, and secondary-use policy that names a deadline, covers data and models derived from the images, and names who certifies that it happened.

Key takeaways