Research
Ethics Readiness of Artificial Intelligence: A Practical Evaluation Method
Overview Research area: AI safety and ethics — specifically, methods for embedding ethical reflection into the engineering design process of AI systems. Technical level: Intermediate. The method is co

- arXiv
- 2512.09729
- Published
- 2025-12-10
- Authors
- Laurynas Adomaitis, Vincent Israel-Jost, Alexei Grinbaum
AI summary
Overview
Research area: AI safety and ethics — specifically, methods for embedding ethical reflection into the engineering design process of AI systems.
Technical level: Intermediate. The method is conceptual and process-oriented rather than mathematically technical, but it presupposes familiarity with both AI system design and applied ethics practice.
Scope in one sentence: The paper proposes Ethics Readiness Levels (ERLs), a four-level iterative evaluation method that translates abstract ethical values into concrete prompts, checks, and controls within real AI use cases, illustrated through two case studies.
What This Paper Is About
There is a persistent gap between high-level ethical principles for AI and the day-to-day decisions engineers actually make when building systems. This paper introduces a structured way to close that gap: a four-level, iterative assessment that turns ethical values into specific, context-dependent checks embedded in a real use case.
The goal is not just to score a system's ethics, but to make ethical reflection a working part of design — something that generates concrete changes and can be tracked over time.
Key Contributions
- Ethics Readiness Levels (ERLs): A four-level, iterative method for tracking how ethical reflection is actually implemented in the design of an AI system, rather than treating ethics as a one-off review.
- A translation mechanism: A systematic way to convert high-level ethical principles into concrete prompts, checks, and controls situated within real use cases.
- A dynamic, tree-like questionnaire: An evaluation instrument built from context-specific indicators, so the assessment stays relevant to the particular technology and application domain rather than being generic.
- A dialogue and tracking function: Beyond acting as a managerial tool, ERLs are positioned as a way to structure conversations between ethics experts and technical teams, with a scoring system for tracking progress over time.
Main Findings
- The method produces design changes: The authors report that the ERL tool effectively catalyzes concrete design changes in the systems it is applied to. The abstract does not state what those specific changes were for either case study.
- It shifts mindset: Applying ERLs is said to promote a move away from narrow technological solutionism toward a more reflective, ethics-by-design approach.
- It bridges principles and practice: The framework is claimed to connect high-level ethical principles with everyday engineering work, which is the central gap the paper targets.
- Context-specificity matters: The questionnaire is built from indicators specific to the technology and domain, which the authors present as the reason the evaluation remains relevant rather than abstract.
- Two demonstration cases: The methodology is demonstrated on an AI facial sketch generator for law enforcement and on a collaborative industrial robot. The abstract gives no quantitative results, baselines, or comparison metrics for either case.
Methodology in Plain English
The authors designed a four-step, repeating assessment called Ethics Readiness Levels. Instead of asking whether a system "is ethical" in the abstract, the method asks where ethical reflection currently sits in the design process and what the next concrete step would be.
To make that practical, the researchers build a questionnaire shaped like a tree: answers lead down different branches depending on the context. The questions themselves are indicators tailored to the specific technology and the specific domain where it will be used, so a law enforcement tool and a factory robot get relevant, not generic, prompts.
The output is a score that can be revisited over time, so teams can see whether their ethics integration is actually advancing. The authors also position the process as a social one — a structured conversation between ethics specialists and the engineers building the system — not just a form to fill in.
They then ran the method on two real scenarios to show it works in practice.
Why This Matters
Impact on research: The paper offers a candidate infrastructure for operationalizing AI ethics — a way to test whether principles-based guidance actually changes engineering outcomes, and a possible template other researchers could adapt or critique.
Real-world applications:
- Law enforcement AI tools: The facial sketch generator case shows how the method applies to a high-stakes, rights-sensitive domain where ethical scrutiny is contested and consequential.
- Industrial robotics: The collaborative robot case extends the method to workplace settings where human safety, task allocation, and worker autonomy are at issue.
- AI product development generally: Any team building an AI system could in principle use the questionnaire structure to embed ethics checks into design documentation.
- Governance and audit: A scoring system that tracks progress over time could feed into internal review or external oversight processes.
Industry relevance: The method is framed as a managerial tool as much as an ethical one, which means it targets the people who allocate engineering time and sign off on designs — not only ethics committees. If it catalyzes design changes as claimed, it offers a way to make ethics-by-design a measurable workflow rather than a stated value.
Future Directions
- Validation beyond two cases: The abstract reports only two demonstrations; whether ERLs generalize across other technologies and domains is an open question.
- Scoring reliability: The abstract mentions a scoring system for tracking progress but gives no detail on how scores are calibrated or whether different evaluators would agree on them.
- From tool to outcome: The authors claim the method catalyzes design changes and shifts mindset; the abstract does not describe how those shifts are verified or sustained after the evaluation ends.
- Integration with existing processes: How ERLs would sit alongside, rather than duplicate, existing ethics review, risk assessment, and regulatory compliance work is not addressed in the abstract.
Target Audience
Ethics researchers and AI policy scholars interested in operationalizing principles; engineering managers and technical leads who need a practical way to build ethical reflection into design workflows; ethics experts and consultants who facilitate dialogues with technical teams; and regulators or auditors looking for structured, trackable evaluation approaches. Readers seeking quantitative benchmarks or comparative performance data will not find them in this abstract.
Authors’ abstract
We present Ethics Readiness Levels (ERLs), a four-level, iterative method to track how ethical reflection is implemented in the design of AI systems. ERLs bridge high-level ethical principles and everyday engineering by turning ethical values into concrete prompts, checks, and controls within real use cases. The evaluation is conducted using a dynamic, tree-like questionnaire built from context-specific indicators, ensuring relevance to the technology and application domain. Beyond being a managerial tool, ERLs help facilitate a structured dialogue between ethics experts and technical teams, while our scoring system helps track progress over time. We demonstrate the methodology through two case studies: an AI facial sketch generator for law enforcement and a collaborative industrial robot. The ERL tool effectively catalyzes concrete design changes and promotes a shift from narrow technological solutionism to a more reflective, ethics-by-design mindset.