Ethics & Governance
AI and Privacy: The Surveillance Bargain at the Heart of the Intelligence Economy
Europe's flagship AI privacy fine was annulled in March 2026, and the largest AI payout yet was for pirated books, not personal data. Why enforcement keeps arriving after the model is built, and never reaches the weights.

Gabriele Masetti ·
The Bargain No One Signed
Every discussion of "AI and privacy" eventually reaches for the word bargain, as if somewhere a negotiation took place. It didn't. The intelligence economy was built by scraping the public web, licensing what couldn't be scraped quietly, and asking forgiveness — in court, in regulatory filings, in the occasional nine-figure settlement — only after the models were already trained and shipped.
The people whose photographs, writing, and biometric data became training fodder were not counterparties to this bargain. They were its raw material. The honest way to describe the arrangement is not a bargain but a fait accompli that regulators and courts are now, slowly and unevenly, trying to retrofit with consent.
That retrofit is failing in a specific, diagnosable way: penalties are arriving late, sized as a fraction of the value already extracted, and rarely forcing anyone to delete a trained model or unlearn what it memorized. Until enforcement can reach the model itself — not just the dataset or the practice that produced it — the surveillance bargain will keep favoring whoever scrapes first.
Scraping First, Litigating Later
The clearest illustration is The New York Times Company v. Microsoft Corporation and OpenAI, filed in December 2023 in the Southern District of New York. The Times alleges that OpenAI and Microsoft used millions of its articles without authorization to train GPT models, and that ChatGPT can reproduce Times journalism near-verbatim. Judge Sidney Stein denied the bulk of the defendants' motion to dismiss in March 2025, letting the core copyright claims proceed while narrowing some DMCA claims.
Discovery produced a long fight over ChatGPT logs — the plaintiffs wanted 120 million; Magistrate Judge Ona Wang ordered a 20-million-log sample, scrubbed of personal identifiers, in November 2025, a ruling Judge Stein affirmed in January 2026. By September 2026 all three parties had moved for summary judgment, and on 17 September the briefs were unsealed, mostly unredacted, putting more than 91,000 copies of Times, Daily News and Center for Investigative Reporting work inside OpenAI's mid-training datasets, along with a Microsoft director's verdict on the practice:
the largest theft of labor in human history — Brent Hecht, Director of Applied Science, Microsoft
Two weeks earlier the Justice Department had filed a statement of interest urging the court to hold that training on copyrighted text is fair use, citing American competitiveness and national security. It was the first time the federal government took a side in AI training-data litigation, and it took the defendants'.
No settlement has been reported. The queue has grown instead: the Seattle Times and Newsday sued both companies in early September 2026, and Times Publishing filed in the Southern District of New York on 16 September. A company that built a trillion-dollar product first is now negotiating, under subpoena and with the government arguing on its side, what it owed the people whose work it consumed.
Getty Images v. Stability AI ran the same playbook and produced a more sobering result for content owners. Getty alleged that 7.3 million of its images trained Stable Diffusion v1 and 4.4 million trained v2. But because Getty could not show the training itself happened in the UK, it dropped its primary copyright and database-right claims before closing arguments.
In its November 2025 judgment, the England and Wales High Court held that Stable Diffusion's weights are not "infringing copies" under UK copyright law — the model stores statistically trained parameters, not reproductions of the photographs — though Getty won a narrow trade-mark claim over outputs that reproduced Getty's watermark. Read together, the two cases show the emerging shape of the law: courts are far more willing to police what a model outputs than what went into training it. That asymmetry is precisely backward from a privacy standpoint, since the harm of nonconsensual data collection happens at ingestion, not at generation.
| Data point | Volume |
|---|---|
| Images Getty alleges trained Stable Diffusion v1 | 7.3 million |
| Images Getty alleges trained Stable Diffusion v2 | 4.4 million |
| ChatGPT logs plaintiffs requested (NYT case) | 120 million |
| ChatGPT log sample court ordered | 20 million |
Biometric Data Is the Sharpest Edge of the Bargain
Text and images are recoverable, in principle, from public sources. Faces are not something you can opt to stop publishing. That is why biometric data has produced the intelligence economy's largest verified penalties.
Clearview AI scraped billions of photographs from social media and the open web to build a facial-recognition database it sold to police departments and private companies. In Illinois — the one state with a statute, the Biometric Information Privacy Act (BIPA), that creates a private right of action for biometric collection without consent — the ACLU sued Clearview in 2020.
The 2022 settlement permanently bars Clearview from selling its faceprint database to most private businesses nationwide and bars sales to police in Illinois for five years, plus creates an opt-out mechanism for Illinois residents. The direct payments were modest: $50,000 to advertise the opt-out program and $250,000 in fees. The value of the settlement was structural, not financial — it constrained Clearview's business model rather than pricing the harm.
BIPA's real deterrent power showed up elsewhere. Facebook paid $650 million in 2021 to settle a BIPA class action over its Tag Suggestions face-template feature — at the time the largest privacy class-action settlement in U.S. history. Texas, which has its own biometric statute (the Capture or Use of Biometric Identifier Act), used it in 2024 to extract a $1.4 billion settlement from Meta over the same facial-recognition practice, the largest privacy settlement any state attorney general has obtained.
Both cases confirm the pattern: statutory, per-violation liability with no requirement to prove individualized harm is the only mechanism that has produced fines large enough to actually change corporate behavior, rather than simply being absorbed as overhead.
| Case | Penalty | Status |
|---|---|---|
| Anthropic / authors (Bartz) | $1.5 billion | Final approval, 20 July 2026 |
| Meta / Texas (CUBI) | $1.4 billion | Settled 2024 |
| Facebook / BIPA | $650 million | Settled 2021 |
| Clearview AI / UK ICO | £7.5 million | Contested since 2022, remitted October 2025 |
| Clearview AI / Illinois (ACLU) | $50,000 + $250,000 | Settled 2022 |
| OpenAI / Italy Garante | €15 million | Annulled, Court of Rome, 18 March 2026 |
The UK's answer for Clearview has been messier and shows the limits of GDPR when a company has no local presence. The Information Commissioner's Office fined Clearview £7.5 million in 2022 for unlawfully processing UK residents' data. Clearview appealed on jurisdictional grounds, and in 2023 the First-tier Tribunal sided with Clearview, ruling its activity fell outside GDPR's scope because it served foreign law enforcement and national-security clients.
The ICO appealed, and in October 2025 the Upper Tribunal reversed course on three of the four grounds, holding that the ICO does have jurisdiction — but sent the question of whether the penalty should stand back to a fresh First-tier Tribunal panel. In December 2025 it gave Clearview permission to take the jurisdiction point to the Court of Appeal, so the remitted hearing sits behind another appeal.
Four years on, whether Clearview has to pay is still undecided. A penalty that takes half a decade to become final, against a company that has operated without interruption throughout, is not a deterrent. It is a rounding error with legal fees attached.
The Retreat of Facial-Recognition Bans
Municipal bans on police facial recognition — San Francisco was first, in 2019, followed by roughly a dozen other cities by 2021 — have stalled rather than spread. No new city bans passed in 2022 or 2023, and New Orleans repealed its ban in July 2022. San Francisco's own police department has faced a lawsuit alleging it skirted its ban by asking neighboring agencies to run facial-recognition searches on its behalf.
The energy has shifted from outright municipal bans to state-level guardrails: by the end of 2024, roughly fifteen states had laws constraining police facial-recognition use, with Montana (2023) and Utah (2024) both adding warrant requirements. That is a real improvement over an unregulated status quo, but a warrant requirement is a procedural check, not a prohibition — it assumes the technology is legitimate as long as a judge signs off, which sidesteps the deeper question of whether biometric surveillance infrastructure should exist at consumer or municipal scale at all.
Europe's Blunter Instrument
The EU's GDPR, and Italy's Garante in particular, have shown what happens when a regulator is willing to use its most disruptive tool immediately rather than after years of litigation. On 30 March 2023, the Garante ordered OpenAI to immediately stop processing Italian users' personal data, citing the absence of a legal basis for using personal data to train ChatGPT, inadequate transparency disclosures, inaccurate outputs about real people, and no effective age verification.
OpenAI restored service on 28 April after agreeing to transparency notices, an age-verification mechanism, and a mechanism for EU residents to object to their data being used in training. The ban lasted under a month, but it forced concrete product changes across OpenAI's entire EU footprint, not just in Italy — the opposite of the years-long, jurisdiction-limited Clearview fight.
In December 2024, the Garante announced a €15 million fine against OpenAI for the same underlying training-data violations, plus an order to run an awareness campaign in the Italian press. Neither survived. On 18 March 2026 the Court of Rome annulled both on jurisdiction alone: OpenAI had opened an Irish establishment in February 2024, which under the GDPR's one-stop-shop rule made Ireland's regulator its lead supervisory authority nine months before the Italian decision was signed.
The court did not rule on whether the training was lawful. It ruled that Italy was no longer the regulator entitled to ask. Europe now has no standing fine against any frontier lab for training-data practices, and a demonstrated route to avoiding one: incorporate in a member state of your choosing while the investigation runs. The emergency shutdown, ordered in weeks, remains the only European instrument that has changed OpenAI's product.
The Machine Remembers What It Was Told Not To
Regulation aimed only at collection practices misses a second privacy problem that lives inside the trained model: memorization. Carlini et al.'s 2021 USENIX Security paper, "Extracting Training Data from Large Language Models," showed that querying GPT-2 with the right prefixes could surface memorized verbatim text, including individuals' names, phone numbers, and email addresses that had appeared in training data.
Follow-up research on production systems has found the problem scales with model size and capability rather than disappearing — larger, more capable models memorize more, and adversarial prompting techniques (special-character triggers, targeted prefix attacks, membership-inference probes) have kept pace with mitigations. Some analyses have found that newer chat-tuned models are harder to extract from per query than raw base models, but that when extraction succeeds against production systems at scale, it surfaces memorized content far more often than earlier, cruder attacks did.
In other words: alignment tuning changed the odds per attempt, not the underlying fact that these models retain fragments of the personal data they were trained on, recoverably.
This matters because it breaks the standard privacy remedy of deletion. If a person's data is memorized inside model weights rather than stored in a queryable database, there is no clean way to "remove" it short of retraining — an expensive, sometimes technically unproven operation, especially at the scale of frontier models. GDPR's right to erasure and CCPA's right to delete were written for databases, not for parameters distributed across billions of weights. Nobody has demonstrated at scale that you can reliably make a trained model "forget" a specific person's data without retraining from scratch.
Differential Privacy: Real, but Partial
The technical fix that actually works, as far as it goes, is differential privacy: adding calibrated statistical noise during training so that no single training example can be reliably recovered from the model's outputs, with the strength of the guarantee controlled by a parameter called epsilon — smaller epsilon means stronger privacy, larger epsilon means weaker. Apple applies differential privacy (with epsilon values reported in roughly the 2–8 range across different data types) to the usage telemetry it collects from iPhones and Macs.
Google trains its Gboard keyboard-prediction models with differentially private stochastic gradient descent, reporting an epsilon around 8.9 for federated learning rounds — specifically so that no individual user's typing patterns get memorized into the shared model. These are real, deployed, verifiable defenses, and they are the right direction.
They are also not what powers frontier large language models trained on scraped web text, because differential privacy at the scale and epsilon values needed for a meaningful guarantee measurably degrades model quality — which is exactly why OpenAI, Anthropic, and Google's largest generative models are not trained this way today. Differential privacy protects users of privacy-conscious products built specifically around it. It does not protect the authors, artists, and ordinary people whose text and images were scraped into a general-purpose model that was never designed with a privacy budget in mind.
Why Fines Aren't Deterring Anyone
Line up the numbers and a pattern jumps out. A trillion-dollar-plus industry has, over roughly a decade, produced one $1.5 billion settlement (Anthropic, for pirated books rather than personal data, finally approved on 20 July 2026), one $1.4 billion settlement (Meta/Texas), one $650 million settlement (Facebook/BIPA), one contested £7.5 million fine that is four years old and unpaid (Clearview/UK), and one €15 million fine that no longer exists (OpenAI/Italy).
The Anthropic settlement came closest to reaching the artifact rather than the practice: it obliged the company to destroy the pirated files it had downloaded, and left the models those files trained untouched.
Only the two biometric statutory-damages cases reached figures large enough to plausibly function as deterrence rather than a cost of doing business — and both depended on unusually aggressive state laws (Illinois's BIPA, Texas's CUBI) that most states don't have. GDPR's own ceiling — up to 4% of global annual revenue or €20 million, whichever is greater — has never been imposed on a frontier AI lab for training-data practices. The one fine that came closest was a fraction of that ceiling, and was thrown out on a procedural point before a euro of it was paid.
What an Actual Bargain Would Require
The through-line across every case here — NYT, Getty, Clearview, Facebook, Meta, Italy — is that enforcement keeps arriving after the model is built, trained, deployed, and monetized, and keeps settling for penalties calibrated to look serious in a press release rather than to threaten the business model.
A genuine bargain would require three things regulators have so far avoided: statutory per-violation damages of the BIPA/CUBI kind extended well beyond two states, so that biometric and other sensitive personal data collection carries real financial exposure before a product ships, not after; a legal right to compel retraining or deletion of a model shown to memorize identifiable personal data, so that "we can't undo it" stops being an accepted defense; and mandatory disclosure of training-data provenance before deployment, not through years of discovery litigation, so the NYT-style multi-year fight over what a model was trained on becomes unnecessary.
Differential privacy shows the industry knows how to build privacy-preserving systems when a product's business case can absorb the utility cost. What's missing is any enforcement mechanism forcing labs to absorb that cost when it's inconvenient — which, for frontier models trained on the entire scraped internet, it currently always is. Until that changes, "bargain" remains the wrong word. It is a tax, and the public has been paying it without ever being told the rate.