The Human-in-the-Loop Retreat: Why Weakening AI Weapons Rules Raises the Risk of Machine Speed Escalation

The international debate over autonomous weapons reached a paradoxical moment in September 2026. After more than a decade of negotiations, the 128 states participating in the Convention on Certain Conventional Weapons agreed on possible elements for governing lethal autonomous weapons systems. That was a diplomatic achievement, but it came after states diluted safeguards that would make human control real rather than rhetorical.

Late changes removed or softened language on predictable and reliable operation, ethical considerations and human assessment of AI-generated targeting. The final text retained important principles: international humanitarian law applies, responsibility cannot be transferred to machines, and human judgement or control is needed. It nevertheless leaves governments greater room to decide how much involvement is enough while militaries accelerate the fielding of AI-enabled and robotic systems.

This retreat matters because the danger is not confined to a fully autonomous machine choosing and attacking a person. AI already sits upstream in intelligence analysis, target nomination, threat classification and weapon assignment. When those functions operate at machine speed, a nominal human decision can become a hurried confirmation of an opaque recommendation. False data, adversarial manipulation, classification error or an unexpected interaction between systems can therefore move from sensor to strike before commanders understand the context or political leaders can intervene.

Effective control should be judged by its causal power. A human must have enough information, time and authority to reject a recommendation, change the mission and stop the system. States should translate that standard into auditable requirements: bounded targets and areas, independent verification for high consequence strikes, realistic testing, protection against automation bias, secure logs, controlled updates and reliable abort mechanisms. The November 2026 Review Conference should open negotiations on a binding instrument, while NATO adopts a common assurance standard. Tightly bounded defensive automation need not allow software to set the pace of strategic escalation.

The rollback happened in the implementation details

The September agreement deserves to be understood accurately. It did not abolish human control or declare autonomous killing acceptable. The final report of the UN group of governmental experts states that international humanitarian law applies fully, that systems must operate within a responsible human chain of command and control, and that accountability cannot be transferred to machines. It also provides the first consensus-based characterization of the systems under discussion. After years of stalemate, those points could give the November Review Conference a basis for formal negotiations.

The weakness lies in the movement from principle to practice. An earlier draft described more specific features of control. It referred to human assessment of legal obligations and anticipated effects; limits on the types and number of targets, duration, geographical scope and scale of operation; authorization before material changes to mission parameters; the ability to deactivate or neutralize a system; and adequate predictability, reliability and traceability. It also addressed testing, automation bias and the preservation of human responsibility across the life cycle.

Reporting on the final negotiations indicates that the United States and Russia pressed to weaken parts of that language. References to predictable and reliable operation and to ethical considerations were removed or narrowed. A provision associated with human review of AI-generated military targets before attack was also dropped. The consensus text still contains useful safeguards, but it converts several concrete requirements into factors for states to consider. That distinction gives technologically advanced militaries more flexibility, while giving other states less certainty about what their competitors will permit.

Consensus diplomacy often produces this kind of bargain. A broad agreement can be better than another failed meeting if it starts treaty negotiations, but it can also create false closure. “Human control” is not self-executing. Unless it is connected to system design, mission limits, data quality, time for deliberation and the ability to halt an attack, governments can claim compliance while retaining only formal authorization.

The political setting makes ambiguity more dangerous. The United States has been reviewing Department of Defense Directive 3000.09, which since 2023 has required appropriate human judgement over force and rigorous verification, validation and testing. Russia argues that existing humanitarian law is sufficient and opposes a binding regime. China supports restrictions in principle while preserving room for development and use. The pressure is therefore toward flexible national interpretation precisely when shared limits are most valuable.

A human click is not meaningful control

The phrase “human in the loop” is attractive because it appears to draw a clear line: the machine recommends and a person decides. In real operations, that line can be nearly empty. A targeting interface may present a confidence score, a bounding box and a recommended weapon, while concealing the data gaps, assumptions and model behavior that produced them. If the operator receives dozens of recommendations, has seconds to respond and knows that delay could expose friendly forces, the final click may be legally significant but operationally automatic. Meaningful control requires the decision-maker to understand the recommendation and its uncertainty, have time to test it against other information and retain the authority and practical means to stop the action. Remove any condition and the person can become a liability buffer rather than an independent decision-maker.

Automation bias compounds the problem. Operators tend to trust systems that perform well in routine conditions, especially under overload. Poor interfaces can highlight a recommended action while burying uncertainty or dissenting sensor data. Command culture reinforces the tendency if operators are evaluated primarily on response time or warned that hesitation will let a target escape.

The inverse problem also matters. Demanding a fresh human decision at every instant is not always safer. A close-in defensive system intercepting incoming rockets or missiles may have only seconds to respond. Human involvement can be exercised earlier through carefully defined target classes, defended areas, operating periods, weapons release conditions and abort criteria. The relevant question is therefore not whether a person touches every engagement. It is whether people establish and retain informed, timely and causally effective control over what the system may do. That standard separates substance from ceremony. A person can be “in the loop” and still lack meaningful control. A tightly bounded defensive system can act automatically while remaining under substantial human control at the mission level. Law and policy should focus on the quality of the control relationship, the foreseeable consequences of failure and the context in which the system is used, rather than treating a button press as a universal safeguard.

Speed changes escalation before it changes accuracy

Military AI is usually justified through speed. It can fuse sensor feeds, classify objects, prioritize threats and allocate weapons faster than a human staff. In a dense drone or missile attack, that advantage can save lives. In a wider crisis, the same speed can reduce the time available to distinguish an attack from an accident, deception or system error. It can also make restraint appear costly if leaders believe an opponent's decision cycle is already faster than their own.

The result is reciprocal compression. One military automates surveillance and target generation to strike mobile systems before they move. Its opponent disperses, delegates authority and automates counterfire. Each adaptation is rational alone; together they allow ambiguous sensor events to generate rapid actions and countermoves. Political leaders may learn that a threshold has been crossed only after multiple engagements.

Accuracy does not eliminate this risk. A model can perform well and still fail in an unfamiliar setting, under electronic warfare or when an adversary manipulates its inputs. It can misread decoys, civilian vehicles or damaged equipment. Systems that behave acceptably alone may interact unpredictably when each responds to the other's output. At machine speed, consequences arrive before the cause can be diagnosed.

Escalation can be direct, accidental or inadvertent. A system might attack the wrong object. It might correctly strike a military target whose political significance was not represented in the model, such as a command node connected to strategic forces. Or it might cause an adversary to misinterpret a limited operation as preparation for a broader attack. The last category is especially difficult because technical correctness does not guarantee strategic comprehension.

Nuclear risk sits at the outer edge of this problem. Most AI-enabled conventional systems are not nuclear weapons, but they can affect early warning, command and control, air defence, cyber operations and strikes against dual-capable forces. A conventional attack recommended by AI could degrade assets that an opponent relies on for nuclear decision-making. Compressed timelines and opaque assessments make it harder to signal limited intent or correct a mistaken interpretation. Human control over the final weapon is not enough if machines have already shaped the information, options and tempo presented to the decision maker.

Decision support is the regulatory blind spot

Debate about autonomous weapons often concentrates on the final platform: a drone, loitering munition, robot vehicle or defensive interceptor that selects and engages a target. This focus can miss the architecture that precedes the shot. AI systems can search imagery, correlate identities, flag patterns of life, rank targets, estimate collateral effects and recommend weapons. A person may formally select the target while relying on a machine-generated chain that no individual has reviewed in full.

This creates a regulatory boundary problem. A decision-support system may fall outside a narrow definition of an autonomous weapon because it does not itself apply force. Yet it can determine which objects enter the target list, how urgent they appear and what information reaches the commander. If hundreds of potential targets are processed faster than a staff could independently examine them, human authorization becomes dependent on the system's framing. Ukraine has demonstrated the value of connecting sensors, software and precision weapons quickly. Major powers are building similar networks across larger theatres. The Pentagon's creation of an Autonomous Warfare Command in September 2026 is a strong signal that autonomy is moving from individual programmes toward a force-wide organizing principle.

Governance should therefore cover the whole decision chain. High-consequence target nominations need traceable sources, confidence ranges and recorded dissenting evidence. Commanders should know when a model is operating outside the conditions in which it was tested. Systems should distinguish between verified facts, probabilistic inferences and recommendations. Where the target could have strategic effects, an independent channel should confirm identity and context rather than merely reproduce the same data through another algorithm.

Logs support accountability only if investigators can reconstruct which model version, data, parameters and human decisions produced the outcome. Updates must be controlled because retraining or a new sensor can change behavior after review. Procurement contracts should guarantee access to technical records and known limitations. A military cannot outsource accountability to a provider whose system it does not understand.

Military necessity does not require unlimited autonomy

Opponents of stronger rules argue that operational contexts vary too widely for prescriptive international standards. A static ban or a single definition could indeed become obsolete as software, sensors and tactics evolve. Rules must also recognize that rapid defensive automation may be necessary against massed missiles, drones or cyber effects. None of this requires accepting open-ended autonomy.

The practical solution is to regulate risk through boundaries. Target type, area, duration, scale and environmental conditions can be restricted before activation. A system may engage incoming uncrewed aircraft inside a defensive zone but not pursue targets beyond it. An anti-materiel system may operate where civilians are absent, while autonomous targeting of people is prohibited. Communications loss or data outside validated ranges should trigger a safe state, not mission expansion. Predictability and reliability remain central even if no system can be perfectly predictable. The requirement is not mathematical certainty. It is a justified basis for commanders to anticipate the system's effects in the intended environment. Testing should include deception, sensor degradation, electronic attack, crowded civilian settings and interactions with other automated systems. Failure modes must be characterized before deployment, not discovered through combat at public expense.

The International Committee of the Red Cross has proposed a useful structure: prohibit autonomous systems whose effects cannot be sufficiently understood, predicted and explained; prohibit systems designed or used to target human beings; and impose strict limits on all others. Those limits include target types, time, geography, scale and situations of use. This approach protects a zone for legitimate defensive automation while preventing the most dangerous forms of machine selection.

Binding law and national doctrine can reinforce each other. A treaty can set prohibitions and minimum controls, while military policy specifies testing, approval and command procedures. The United States Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy, NATO's responsible-use principles and the REAIM Blueprint already contain overlapping commitments on accountability, reliability, explainability, traceability, governability and human responsibility. The gap is not a shortage of principles. It is the absence of a common mechanism that turns them into verifiable operating constraints.

The rules race is becoming a deployment race

The timing of the September compromise matters. Militaries are no longer discussing autonomy only as a future capability. AI-enabled targeting, navigation, electronic warfare and uncrewed systems are being deployed, tested and revised during active conflicts. Commercial technologies shorten development cycles, and operational software can change more quickly than traditional weapon review processes. Institutions that take years to agree on language are trying to govern systems updated in weeks.

Competition encourages permissive interpretation. If one state believes a rival is removing humans from targeting, it may fear that stronger safeguards create a disadvantage. Procurement officials may treat assurance as delay, commanders may request broader mission parameters, and industry may resist disclosure. Together these pressures create a race to the lowest acceptable constraint.

The international process can interrupt that dynamic only if it creates reciprocal confidence. The September consensus is a starting point because it records common legal principles and keeps negotiations alive. It is not an endpoint. The November Review Conference should establish a mandate to negotiate a legally binding instrument with clear prohibitions and positive obligations. Verification will be difficult, but difficulty is not an argument for leaving the most consequential decisions undefined. Progress does not have to wait for universal agreement. NATO members should align their procurement and operational policies around a common assurance framework. Systems should not enter service without documented legal review, realistic testing, human-factors assessment, cybersecurity evaluation and a clearly identified accountable commander. Allies should agree on minimum evidence for target verification and on additional review for attacks that could affect strategic command systems, civilian infrastructure or escalation thresholds. Shared exercises should test not only speed and accuracy but also the ability to pause, explain and recover from failure.

Coalitions can make standards economically influential. If major customers require auditable logs, controlled updates, operator-centered interfaces and vendor access for investigations, suppliers will build those features into their products. Export controls and assistance programmes can extend the same baseline. This would not eliminate irresponsible users, but it would prevent responsible-use principles from becoming optional extras removed under competitive pressure.

Build control into the system

The most durable safeguard is control by design. Mission planning should specify what the system may target, where, for how long and under which conditions. Interfaces should show uncertainty, conflicting inputs and the reason for a recommendation rather than only a confidence number. Operators should be trained and authorized to reject machine outputs. Communications loss, sensor disagreement or behavior outside tested limits should trigger pre-agreed restrictions or shutdown.

High-risk decisions require deliberate friction. That does not mean routing every defensive engagement through senior headquarters. It means identifying decisions whose consequences justify additional time and independent review: attacks on people, targets near civilians, infrastructure with cascading effects, command-and-control assets and objects that may be dual-capable or strategically sensitive. The fastest technically available option should not automatically become the default operational process.

Accountability also needs continuity. Someone must be responsible for approving the system, defining the mission, monitoring performance and investigating failure. Records must link those decisions to the software and data in use at the time. Legal reviews should be revisited after material changes, including new sensors, retrained models, modified target libraries or expanded operating areas. Red-team testing should include adversarial manipulation and unexpected system interaction, not only benchmark accuracy.

At the strategic level, states should create firebreaks around nuclear forces and crisis decision-making. AI generated warnings or target recommendations should never become the sole basis for an action with nuclear implications. Conventional operations near strategic command, early-warning or dual-capable assets should require higher-level review and multiple independent information sources. Communication channels between rivals remain necessary because no technical safeguard can guarantee that an adversary will interpret an automated action as intended.

These controls impose costs. They require testing infrastructure, trained personnel, slower approval and, in some cases, acceptance that a target will be missed. Those costs should be compared with the price of an erroneous strike, an uncontrolled interaction or an escalation neither side intended. A system that produces a faster decision but cannot support a defensible judgement has not solved the command problem. It has moved the risk downstream.