Skip to content
AI.info

creators

Nicholas Carlini: AI Security Research at Anthropic

Nicholas Carlini is a research scientist at Anthropic studying attacks on machine learning, and the co-author of the Carlini-Wagner adversarial attack.

Nicholas Carlini is a research scientist at Anthropic, working at the intersection of machine learning and computer security; he describes his current work as studying what bad things you could do with, or do to, language models. He holds a PhD from UC Berkeley under David Wagner and a BA in computer science and mathematics, also from Berkeley. He was a research scientist at Google Brain from 2018 to 2023 and at DeepMind from 2023 to 2025. He is best known for the Carlini-Wagner attack, one of the most widely used methods for constructing adversarial examples against neural network image classifiers, and for years of later work puncturing claimed robustness and privacy guarantees in machine learning systems. His paper “Extracting Training Data from Large Language Models” showed that verbatim training data, including personal information, could be recovered from GPT-2 through black-box queries alone, which made memorisation a live privacy problem for language models. He also co-wrote the Greedy Coordinate Gradient paper on universal, transferable jailbreak attacks against aligned language models. His own list of papers records best paper awards at IEEE S&P, EuroCrypt, USENIX Security twice and ICML three times. Away from research he writes what he calls useless code, including a tic-tac-toe game in a single printf call that won the 2020 IOCCC Best of Show.

Specialization
adversarial machine learning, AI security, training-data privacy, language model attacks
Country
United States

Work