Research
Not-in-Perspective: Towards Shielding Google's Perspective API Against Adversarial Negation Attacks
Overview Research area: Adversarial robustness of machine learning toxicity detection, specifically negation attacks against Google's Perspective API, addressed with a formal reasoning wrapper around

- arXiv
- 2602.09343
- Published
- 2026-02-10
- Authors
- Michail S. Alexiou, J. Sukarno Mertoguno
AI summary
Overview
Research area: Adversarial robustness of machine learning toxicity detection, specifically negation attacks against Google's Perspective API, addressed with a formal reasoning wrapper around statistical classifiers.
Technical level: Intermediate. The paper assumes some familiarity with toxicity classifiers, part-of-speech tagging, and adversarial attacks, but the proposed mechanisms are described operationally and can be followed without deep mathematical background.
One-sentence scope: The paper proposes and evaluates hybrid formal-reasoning-plus-machine-learning wrappers that detect negated segments of a comment and recompute the toxicity score so that inserting the word "not" no longer fools a toxicity detector.
What This Paper Is About
Machine learning toxicity detectors, including Google's Perspective API, are vulnerable to adversarial negation attacks: adding "not" before offensive words produces sentences that a human would read as far less toxic, yet the model still returns a high toxicity score. The authors show this vulnerability has persisted since Hosseini et al. demonstrated it, and they build a reasoning layer that sits both before and after an existing ML classifier to recompute scores according to the logic of negation. The goal is to reduce the false toxicity score on negated comments without discarding the statistical model as the backbone.
Key Contributions
-
A formal reasoning wrapper for negation. A pre-processing step normalizes negation contractions (for example, "weren't" becomes "were not"), and a post-processing step recomputes the sentence toxicity score after identifying which segments are actually affected by negation using Stanford's Part-of-Speech tagger and recursive parsing.
-
Five heuristic reasoning-and-synthesis variants (M1.1 through M1.5). These treat negation as logical NOT and combine segment scores with weights derived from word counts, from a pure logical negation basis, or from weights approximated by repeatedly querying the machine learning model (Equations 1 through 7).
-
Two substitution-and-exploration variants (M2.1 and M2.2). These replace negated words with antonyms retrieved from a Thesaurus or, in the second version, from OneLook Reverse Dictionary using context definitions extracted by the Lesk algorithm, then use the Parrot paraphraser to extrapolate sentences whose average toxicity score becomes the new score.
-
A negation adversarial benchmark plus a comparative evaluation. The authors generate a negated test set from Jigsaw's public and private test sets (418 and 382 negated sentences respectively) and evaluate every wrapper variant against Perspective API, an LSTM + CNN model, and a BERT + Bi-LSTM model.
Main Findings
-
Perspective API still fails on negation attacks. The authors confirm that Google's Perspective API has not significantly improved since Hosseini et al.'s published attack, and note that Google's own GitHub repository advises against using the software for automated moderation because "the models make too many errors."
-
On Hosseini's 9 original toxic sentences, Perspective is the most reliable model. Perspective correctly classified each of the 9 sentences as toxic, while the LSTM + CNN and BERT + Bi-LSTM models missed multiple obviously toxic comments even though both were trained on the same training set and reached high validation scores.
-
M1.2 and M1.5 are the best performing wrappers. These two methods output the best and most reliable performance across the evaluated wrappers, attributed to using logical negation together with redistribution of scores. M1.2 estimates each phrase's weight from word count; M1.5 approximates the weights used by Perspective API by querying it repeatedly.
-
Accuracy on the Perspective-based negated public test set. M1.2-P reached 0.82 and M1.5-P reached 0.84, compared with Perspective's 0.003, M1.1-P's 0.58, M1.4-P's 0.72, M2.1-P's 0.42 and M2.2-P's 0.37.
-
Accuracy on the Perspective-based negated private test set. M1.2-P reached 0.87 and M1.5-P reached 0.86, compared with Perspective's 0.003, M1.1-P's 0.59, M1.4-P's 0.76, M2.1-P's 0.40 and M2.2-P's 0.40.
-
Substantial toxicity score reductions. Both M1.2-P and M1.5-P improved the toxicity of almost 250 out of 382 sentences in the public set and more than 250 comments out of 418 in the private set when achieving a drop of over 30 percent from their original toxicity. On the first negated sentence from Hosseini's set, Perspective scored 0.8 while M1.2-P scored 0.107.
-
M1.3 was disqualified despite good performance. The authors rule it out because it is simply toxicity redistribution based on weights determined by word count, with no additional reasoning value.
-
Substitution methods are accurate but costly. M2.1 and M2.2 produce comparable, average results, but the substitution-and-exploration pipeline is computationally demanding and, in the worst case, scales as O(5^n) because each negated word yields up to 5 antonyms and each permutation is paraphrased. This makes it unsuitable for real-time toxicity detection.
-
Two sources of error are identified. Thesaurus-based methods fail when a toxic word has no antonym, as in the test comment "Eat shit you stupid fuk" where no antonyms exist for "shit" and "fuk". Separately, the test-set generation step relies on Perspective to identify high-impact toxic words, and the exact procedure for Perspective's toxicity score computation remains unknown, so other toxic words in a sentence may be missed; in the private Jigsaw comment "These scumbags need …Don't drop the soap.", the combined toxicity of "scumbags" and "rot" is 0.75, equal to the sentence's overall toxicity.
-
Heuristic adjustment depends on a threshold. If a negated phrase's retrieved toxicity is over 0.5, "not" is treated as logical negation; a comment is characterized as non-toxic only if its toxicity score is below the 0.5 threshold. This prevents negating harmless segments such as "not changing", whose individual toxicity was TS2 = 0.034, while the offensive segment "not an idiot" had TS4 = 0.932.
-
Worked example scores. For the sentence "Climate change is happening and it is not changing in our favor. If you think differently you are an idiot" with segment scores TS1 = 0.068, TS2 = 0.034, TS3 = 0.12 and TS4 = 0.932, applying Equations 1, 3 and 4 yields approximately 0.107, 0.198 and 0.333 respectively. The first substitution version produced 12 paraphrases with average toxicity 0.334; the second produced 25 paraphrases with average score 0.286.
Methodology in Plain English
The authors treat a comment as a sequence of segments, some of which fall under the scope of a negation and some of which do not. A pre-processing step expands contractions so that negation is explicit. Stanford's Part-of-Speech tagger is then used with recursive parsing to find which words the "not" actually governs; the recursion continues across conjunctions such as "and" and "or" and across commas, and stops at full stops and exclamation marks. The distributive property is used to work out how far the negation reaches.
Once the sentence is split into negated segments and non-negated segments, each segment is queried separately against a toxicity classifier. Segments with a toxicity above 0.5 are treated as genuinely negated, meaning their score is flipped using the "1 − TS" idea of logical NOT, while low-toxicity negated segments are left alone so that harmless phrases like "not changing" do not accidentally raise the overall score. The final score is a weighted combination of segment scores, with weights derived from the number of words in each segment. The authors also test alternative weightings derived by repeatedly querying the classifier to infer the coefficients it uses.
The second family of methods avoids score arithmetic entirely. Negated words are replaced with antonyms retrieved automatically, generating a new sentence that preserves the original meaning. A paraphrasing tool then produces additional variants, bounded at 5 antonyms per negated word, and the toxicity of the original sentence becomes the average score across all generated paraphrases. A variant uses the Lesk algorithm to disambiguate word meaning from context before looking up antonyms.
All wrapper variants are evaluated with three different toxicity detectors behind them: Perspective API, an LSTM + CNN using fastText embeddings, and a bidirectional LSTM on top of uncased BERT embeddings.
Why This Matters
Impact on research. The paper argues that purely statistical toxicity detection cannot reason about logic-based transformations such as negation, and that adversarial training is impractical because the number of attack variations per word or phrase is extremely large. It positions hybrid formal-reasoning-plus-learning systems, consistent with the "Learn2Reason" concept, as a practical route for a class of problems that statistics alone handles poorly. It also offers a reusable negated benchmark derived from Jigsaw's public and private test sets and shows that even models with high validation accuracy can be unreliable on obvious cases.
Real-world applications.
- Content moderation pipelines on social media platforms and online chatrooms, where users deliberately insert negation to evade automated filters.
- Pre-screening layers for human moderators, reducing the volume of adversarial false negatives reaching review.
- Cyberbullying and hate speech monitoring in platforms serving the large share of teenagers and young adults who report exposure to online toxicity.
- Auditing and hardening of deployed toxicity APIs, since the paper confirms Perspective's vulnerability persists years after it was first demonstrated.
Industry relevance. Google's Perspective API is available to the public and was published specifically as a measure to track and counter growing toxicity, yet the paper reports that Google acknowledges its machine learning is not enough and advises against automated moderation with it. The wrapper approach is attractive to industry because it does not require retraining or replacing an existing classifier; the M1.2 and M1.5 variants can be layered around a deployed API, whereas the substitution variants are ruled out for real-time use on complexity grounds.
Future Directions
- Neologism analysis. The authors state they will integrate a neologism analysis layer, since experiments with sentences containing neologisms indicate that Perspective predicts high toxicity scores for words with unknown roots.
- Improving the substitution-and-exploration methods. M2.1 and M2.2 still require improvement, particularly the antonym-retrieval step that fails on words with no thesaurus entry.
- Resolving the accuracy-versus-efficiency trade-off. The Lesk-based substitution variant is the most consistent but also the most computationally costly; the paper explicitly raises the question of balancing accuracy and efficiency, and the measured complexity is O(5^n).
- Making the test set independent of Perspective. The generated test set relies on Perspective itself to select the words that drive toxicity, and because the exact computation of Perspective's score is unknown, this may under-select toxic terms; a test set built without that dependency remains an open problem.
Target Audience
Researchers and practitioners working on adversarial robustness, content moderation, and toxicity or sentiment classification will benefit most, along with engineers deploying Perspective API or similar detectors in social media and chat platforms. The paper is also useful to readers interested in hybrid neuro-symbolic approaches, since it demonstrates a concrete case where formal reasoning and statistical learning are combined around an existing production model.
Authors’ abstract
The rise of cyberbullying in social media platforms involving toxic comments has escalated the need for effective ways to monitor and moderate online interactions. Existing solutions of automated toxicity detection systems, are based on a machine or deep learning algorithms. However, statistics-based solutions are generally prone to adversarial attacks that contain logic based modifications such as negation in phrases and sentences. In that regard, we present a set of formal reasoning-based methodologies that wrap around existing machine learning toxicity detection systems. Acting as both pre-processing and post-processing steps, our formal reasoning wrapper helps alleviating the negation attack problems and significantly improves the accuracy and efficacy of toxicity scoring. We evaluate different variations of our wrapper on multiple machine learning models against a negation adversarial dataset. Experimental results highlight the improvement of hybrid (formal reasoning and machine-learning) methods against various purely statistical solutions.