Skip to content
AI.info

Ethics & Governance

AI in Military Applications: Ethical Boundaries and International Law

The UN deadline for a binding instrument on autonomous weapons passed in 2026 with nothing agreed. What the CCW's expired mandate, the US walk-back at REAIM and Maven's promotion mean for accountability.

AI in Military Applications: Ethical Boundaries and International Law

Gabriele Masetti ·

A twenty-second decision

In the spring of 2024, +972 Magazine and Local Call published interviews with Israeli intelligence officers describing a system called Lavender, which during the early weeks of the Israel-Hamas war flagged as many as 37,000 Palestinians as suspected Hamas or Palestinian Islamic Jihad operatives and, by extension, as potential airstrike targets.

According to the officers cited, human review of the machine's output was frequently reduced to a rubber-stamp check lasting around twenty seconds, largely to confirm that a flagged target was male, before the file moved toward a strike. A companion system known as The Gospel generated recommendations for striking structures and buildings associated with those targets.

The Israel Defense Forces has disputed aspects of this reporting and maintains that a human analyst reviews every target, but even the more cautious accounts converge on a specific and unsettling fact: the software did the selecting, and the person in the loop was optimized, procedurally and psychologically, to agree with it.

Metric Figure
Targets Lavender flagged (early weeks of war) ~37,000
Reported human review time per target ~20 seconds
States endorsing US Political Declaration (Nov 2023) 45

That is the accountability problem in military AI in miniature. It is not primarily a question about whether a machine can distinguish a combatant from a civilian, though that remains genuinely hard. It is a question about what happens to human judgment once an algorithm has already done the cognitive work and a person is asked only to ratify it under time pressure.

International law was written for a world of triggers pulled by identifiable people. It is not yet equipped for a world of decisions pre-made by statistical models and merely countersigned by an officer who may not know how the system reached its conclusion, or have realistic room to override it.

The law we already have, and where it strains

International humanitarian law does not have a blank spot where military AI is concerned. The core rules of distinction, proportionality, and precaution in attack, codified in the 1977 Additional Protocols to the Geneva Conventions and reflected in customary law, apply regardless of what technology is doing the targeting. Article 36 of Additional Protocol I additionally obligates states to conduct legal reviews of new weapons, means, and methods of warfare before fielding them, to determine whether their use would be prohibited by the Protocol or other applicable rules.

The strain shows up in implementation. Article 36 reviews were designed around physical weapons with knowable blast radii and predictable failure modes, not around machine-learning systems whose behavior can shift with retraining, whose statistical errors are probabilistic rather than mechanical, and whose reasoning is often opaque even to the engineers who built them.

The Stockholm International Peace Research Institute and the International Committee of the Red Cross have both noted that few states publish the methodology, let alone the results, of their weapons reviews, so there is no way to verify from outside that an autonomous or AI-assisted targeting system was ever tested against these standards.

National reviews are also inherently self-graded: the state fielding the system is the same state judging its own legality, with no external audit and no venue for a losing challenge before the weapon is used.

The deeper legal gap is about responsibility after the fact. IHL assigns individual criminal responsibility for war crimes and command responsibility for a commander's failure to prevent or punish them. Both doctrines presuppose that a person made, or should have made, a decision that can be traced and judged. When a targeting recommendation emerges from a model trained on patterns in intercepted communications and movement data, and a human approves it in seconds without independent verification, it becomes genuinely difficult to locate the decision that a court would need to scrutinize.

Several legal scholars, including researchers affiliated with the Lieber Institute at West Point, have described this as an emerging accountability gap: not a lawless zone, but a zone where existing law's assumptions about a traceable human decision-maker no longer hold cleanly.

Two governance tracks, moving at different speeds

Two parallel efforts are trying to close that gap, and they disagree sharply on method.

The United States has pursued a domestic-regulation-plus-diplomacy track. In January 2023 the Department of Defense issued the first major revision since 2012 of DoD Directive 3000.09, "Autonomy in Weapon Systems." The updated directive requires that autonomous and semi-autonomous weapon systems be designed to allow commanders and operators to exercise appropriate levels of human judgment over the use of force, mandates the ability to detect and disengage systems exhibiting unintended behavior, and creates a senior-level Autonomous Weapon Systems Working Group, drawing on policy, acquisition, legal, and testing offices, to review such systems before they enter formal development and again before fielding.

It is a real institutional mechanism, but it is also a self-administered one: the Pentagon reviews the Pentagon's weapons, and DefenseScoop has reported that after the directive's update, the department largely declined to disclose which specific systems have gone through the process.

The directive is now being rewritten. In June 2026 the White House issued National Security Presidential Memorandum 11, which directs a revision of 3000.09 in the name of faster AI adoption. Senator Ruben Gallego wrote to the department within weeks questioning the plan.

Washington then took this model abroad. At the Responsible AI in the Military Domain summit in The Hague in February 2023, the US unveiled a Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy, a set of non-binding commitments including maintaining human control over nuclear weapons decisions, ensuring senior review of high-consequence AI systems, and subjecting military AI to rigorous testing.

After revisions, the declaration had attracted 45 endorsing states by November 2023, including most NATO members, Japan, South Korea, and Israel, and gained adherents afterward. Notably, China has not endorsed it, and the declaration binds no one to anything enforceable — it is a statement of intent, not a treaty, with no verification regime and no penalty for a state that endorses it and then violates its own commitments.

Then the momentum reversed, and Washington led the reversal. At the third REAIM summit, held in A Coruña in early February 2026, roughly thirty-five states endorsed a twenty-point text called Pathways for Action. The United States, which launched the declaration in 2023, declined to sign.

The United Nations track has aimed higher and moved slower. Since 2014, a Group of Governmental Experts operating under the Convention on Certain Conventional Weapons has met to negotiate, by consensus, possible elements of a legal instrument on lethal autonomous weapons systems. Sessions through 2024 and 2025 produced a rolling text that, by a November 2024 draft, included provisional language on prohibiting LAWS that operate without context-appropriate human control and judgment, and on obligating states to ensure human responsibility and accountability for their use.

But CCW rules require consensus among all participating states, giving any single objector an effective veto, and the group ran out of time. Its 2026 sessions closed in early September with no binding instrument agreed and the mandate expired, moving the decision up to the Convention's Seventh Review Conference, convened for 16 to 20 November 2026 in Geneva.

In 2023 the UN Secretary-General and the president of the International Committee of the Red Cross jointly called on states to negotiate a legally binding instrument by 2026. On 25 August 2026, with that date behind them and nothing agreed, they issued the appeal again.

We are now dangerously close to crossing a moral red line: the autonomous targeting of humans by machines — António Guterres and Mirjana Spoljaric Egger, Secretary-General of the United Nations and President of the ICRC

Frustration with that pace boosted a parallel push at the UN General Assembly. The first resolution on the risks of lethal autonomous weapons systems passed in December 2023 by 152 votes in favor, 4 against and 11 abstaining. A year later the margin widened: on 2 December 2024 the Assembly adopted resolution 79/62 by 166 in favor, 3 against — Belarus, North Korea and Russia — and 15 abstaining. Both were lopsided political signals with no binding force.

The Campaign to Stop Killer Robots, a coalition of more than 270 nongovernmental organizations across some 70 countries, together with Harvard Law School's International Human Rights Clinic, has proposed a dedicated treaty built around a general obligation to maintain meaningful human control over the use of force, an outright prohibition on weapons that select and engage targets without it, and positive duties for every other targeting system. No state has yet agreed to negotiate that text.

Vote tally on the UN resolution addressing risks of lethal autonomous weapons systems.

Where the accountability actually breaks in practice

The gap between declared principle and battlefield practice is where the real risk lives, and Gaza is not the only illustration. Project Maven, the Pentagon's computer-vision and target-identification program that began in 2017, has evolved into a wide-reaching system reportedly credited with supporting US airstrike targeting in Iraq, Syria, and Yemen and with tracking hostile maritime assets in the Red Sea during 2024, with officials describing dramatic increases in the volume of targets processed per day compared with manual analysis.

Its status has since been formalised. In a memo dated 9 March 2026, Deputy Secretary of Defense Steve Feinberg directed the Palantir-built Maven Smart System to become a department-wide program of record before the end of fiscal year 2026, and wrote that the department must "establish AI-enabled decision-making as the cornerstone of our strategy for" Combined Joint All-Domain Command and Control.

Maven is described by the Pentagon as a decision-support tool rather than a weapon that fires itself, which matters legally: Directive 3000.09 and the autonomous-weapons debate are concerned with systems that select and engage targets without further human input, and a recommendation engine a human must still act on sits in a less-scrutinized category, even though its influence on the outcome can be just as large.

That category distinction is precisely the loophole exposed by the Lavender and Gospel reporting. Neither system reportedly fired a weapon; both reportedly shaped, and in practice often determined, which humans and buildings a human would then approve for a strike. If the surrounding process compresses genuine deliberation into a twenty-second confirmation, the formal presence of a human in the loop does little to preserve the substantive judgment that the laws of war require, and it does even less to preserve a clear chain of individual responsibility if the strike turns out to have hit the wrong target.

A recommendation system with sufficiently high trust, sufficiently high speed, and sufficiently low friction to override can produce the same accountability vacuum as full autonomy, without ever triggering the rules written specifically for full autonomy.

Where the line should sit

Given all this, drawing the line only at whether the machine pulls the trigger is not adequate, and policy should stop treating that distinction as the central one. Three positions follow from the evidence above.

First, meaningful human control has to be defined operationally, not aspirationally, and enforced through auditable minimums: a documented, time-bounded requirement that a human targeting decision include independent review of underlying evidence, not just confirmation of the AI's conclusion, with the review time and the evidence consulted logged and available to post-incident investigators and, in appropriate cases, to international bodies. A rubber stamp delivered in twenty seconds does not satisfy any honest reading of appropriate levels of human judgment, and directives like 3000.09 should say so explicitly rather than leaving it to officers under pressure to keep pace with the system.

Second, decision-support systems that function as de facto targeting engines — Maven-class tools, Lavender-class scoring systems — should be brought inside the same review and disclosure regime as autonomous weapons proper, because the accountability harm they can cause is not meaningfully smaller. Regulating only weapons that fire themselves while leaving unregulated the systems that decide, with high confidence and at scale, who gets fired upon, is a distinction without the difference that matters to the people on the receiving end.

Third, self-review is not accountability. Article 36's national weapons-review model and DoD Directive 3000.09's internal working group are useful administrative safeguards, but a state investigating and clearing its own weapons and processes, with no requirement to publish results, is not equivalent to independent scrutiny.

The CCW process, slow and consensus-bound as it is, remains the only forum with the legitimacy to set a shared external floor, and in November 2026 it gets one more chance to use it. The states resisting it should be named as the primary obstacle to closing the accountability gap, not treated as one reasonable position among several equally valid ones. That list now includes the United States, which wrote the voluntary declaration and then would not sign its successor at A Coruña, and China, which has joined neither.

The honest conclusion is not that military AI is inherently illegitimate or that it can be banned out of existence — battlefield commanders will keep using pattern-recognition tools because they work, and no treaty will unmake that fact. The honest conclusion is that the current arrangement, in which speed and technical sophistication are advancing far faster than the auditability and enforceability of the human oversight wrapped around them, produces exactly the kind of unaccountable harm the Geneva framework was built to prevent.

Closing that gap requires treating meaningful human control as a measurable engineering and process requirement with teeth, not a phrase for policy documents to gesture toward, and it requires the states most invested in these systems to accept external verification rather than continuing to grade their own homework.

Explore

More articles