Research
In Quest of an Extensible Multi-Level Harm Taxonomy for Adversarial AI: Heart of Security, Ethical Risk Scoring and Resilience Analytics
Overview Research area: AI safety and ethics — specifically the conceptual foundations of harm as used in adversarial AI, cybersecurity, and risk analysis. Technical level: Intermediate. The work is c
- arXiv
- 2601.16930
- Published
- 2026-01-23
- Authors
- Javed I. Khan, Sharmila Rahman Prithula
AI summary
Overview
Research area: AI safety and ethics — specifically the conceptual foundations of harm as used in adversarial AI, cybersecurity, and risk analysis.
Technical level: Intermediate. The work is conceptual and framework-building rather than quantitative; it assumes some familiarity with ethical theory and with how harm is discussed in AI safety, but the abstract presents no mathematics, datasets, or experiments.
Scope (1 sentence): The paper proposes an extensible, ethics-grounded taxonomy that turns "harm" from a vague rhetorical term into an enumerable, structured object that can be classified, attributed to victims, and assessed for normative severity.
What This Paper Is About
Harm is invoked constantly across cybersecurity, ethics, risk analysis, and adversarial AI, yet there is no systematic or agreed-upon list of what harms actually are, and the concept itself is rarely defined precisely enough for serious analysis. Because current discourse relies on vague and underspecified notions of harm, structured qualitative assessment is effectively impossible. The paper's goal is to close that gap by building a structured, expandable taxonomy of harms that makes harm explicit, countable, and analytically tractable.
Key Contributions
- A structured, expandable taxonomy of harms grounded in an ensemble of contemporary ethical theories, designed so that harm becomes explicit and enumerable rather than implicit.
- A two-domain, eleven-category organization of 66+ distinct harm types, where each of the eleven major categories is explicitly aligned with one of eleven dominant ethical theories.
- A theory-aware taxonomy of victim entities, introducing victims as a first-class part of harm classification rather than an afterthought.
- Formalized normative harm attributes — explicitly including reversibility and duration — that the authors argue materially change how ethically severe a given harm is.
The abstract adds a design claim that distinguishes the taxonomy from a flat laundry list: it is extensible by design, but its upper levels are intentionally kept stable.
Main Findings
- No systematic harm list currently exists: The authors assert that despite harm being invoked across cybersecurity, ethics, risk analysis, and adversarial AI, there is no systematic or agreed-upon enumeration of harms.
- Harm is rarely defined precisely: The abstract claims the concept is seldom defined with the precision that serious analysis requires, and that this vagueness blocks nuanced, structured, qualitative assessment.
- The taxonomy enumerates 66+ harm types: These are organized into two overarching domains — human and nonhuman — and eleven major categories.
- Categories map onto ethical theories: Each of the eleven major categories is explicitly aligned with one of eleven dominant ethical theories, producing an "ensemble" grounding rather than reliance on a single theory.
- Severity depends on more than category: Normative harm attributes such as reversibility and duration are claimed to materially alter ethical severity — meaning two harms of the same type can differ ethically based on these attributes.
- Extensibility with stable upper levels: The framework is designed to be extended with new harm types while its top-level structure remains fixed.
- Claimed outcome: Taken together, these contributions are said to transform harm from a rhetorical placeholder into an operational object of analysis.
Note: the abstract reports no empirical validation, benchmark, dataset, score, or comparison to existing taxonomies. Any such evidence is not in the abstract.
Methodology in Plain English
The authors take a conceptual and framework-building approach rather than an experimental one. They start from an ensemble of contemporary ethical theories and use those theories as the organizing scaffold for a classification system. Harms are then enumerated and sorted into a two-level structure: two broad domains (human and nonhuman), which fan out into eleven major categories, each tied to a specific ethical theory. Within that structure, they define a separate taxonomy for victim entities, and they specify attributes of a harm — reversibility and duration are the two named — that determine how ethically serious it is. The design constraint is that the top levels stay stable while lower levels remain open to additions, so the taxonomy can grow without being rebuilt. The abstract describes the resulting artifact and its intended use; it does not describe how the taxonomy was tested or validated.
Why This Matters
Impact on research: If harm can be enumerated and structured rather than gestured at, researchers gain a shared vocabulary for comparing and reasoning about harms across AI safety, adversarial machine learning, cybersecurity, and ethics. The abstract frames this as a precondition for rigorous ethical reasoning and for long-term safety evaluation of AI systems, and notes the framework is intended to apply to other sociotechnical domains where harm is a first-order concern.
Real-world applications the abstract's framing supports (the abstract names domains but does not detail specific deployments):
- Adversarial AI evaluation: Giving safety and red-team work a structured way to name and categorize the harms an attack or model behavior could produce.
- Cybersecurity risk analysis: Replacing vague harm language in risk assessments with specific, enumerable harm types and severity attributes.
- Ethics review and governance: Providing reviewers with a checklist-like structure for identifying which harms a system might cause and to whom.
- Risk scoring and resilience work: The paper's own title points toward ethical risk scoring and resilience analytics, though the abstract does not explain how scoring or resilience metrics are constructed.
Industry relevance: Organizations building or deploying AI need defensible ways to document, compare, and prioritize harms. A shared taxonomy with explicit severity attributes (reversibility, duration) gives compliance, safety, and policy teams a common structure for internal review, incident classification, and cross-team comparison. The abstract does not present adoption evidence, tooling, or deployment results.
Future Directions
- Extending the taxonomy: Since the framework is extensible by design, an open question is how new harm types get proposed, vetted, and slotted into the stable upper levels without destabilizing the structure.
- Validation of the theory mapping: The abstract asserts alignment between eleven categories and eleven ethical theories but offers no external validation; testing whether that mapping holds up under scrutiny from ethicists is a natural next step.
- Operationalizing the attributes: Reversibility and duration are named as severity-altering attributes, but the abstract does not specify how they are measured, scaled, or combined — a clear gap to close.
- Connecting to scoring and resilience: The title promises ethical risk scoring and resilience analytics that the abstract does not describe, leaving open how the taxonomy converts into quantitative or semi-quantitative assessments.
- Application to real AI systems: Moving from a conceptual framework to demonstrated use in AI safety evaluation and other sociotechnical settings.
Target Audience
AI safety and adversarial-ML researchers looking for structured harm language; applied ethicists and AI ethics practitioners interested in theory-grounded classification; security and risk analysts who need a more precise vocabulary than "harm"; and policy, governance, and compliance professionals who need to document and compare harms across systems and domains. Readers seeking empirical results, scoring formulas, or validation studies will not find them in the abstract.
Authors’ abstract
Harm is invoked everywhere from cybersecurity, ethics, risk analysis, to adversarial AI, yet there exists no systematic or agreed upon list of harms, and the concept itself is rarely defined with the precision required for serious analysis. Current discourse relies on vague, under specified notions of harm, rendering nuanced, structured, and qualitative assessment effectively impossible. This paper challenges that gap directly. We introduce a structured and expandable taxonomy of harms, grounded in an ensemble of contemporary ethical theories, that makes harm explicit, enumerable, and analytically tractable. The proposed framework identifies 66+ distinct harm types, systematically organized into two overarching domains human and nonhuman, and eleven major categories, each explicitly aligned with eleven dominant ethical theories. While extensible by design, the upper levels are intentionally stable. Beyond classification, we introduce a theory-aware taxonomy of victim entities and formalize normative harm attributes, including reversibility and duration that materially alter ethical severity. Together, these contributions transform harm from a rhetorical placeholder into an operational object of analysis, enabling rigorous ethical reasoning and long term safety evaluation of AI systems and other sociotechnical domains where harm is a first order concern.