The Human in the Loop Can Become an Alibi for War Machines

The deepest military AI risk is a workflow that erases uncertainty before a human authorizes force.

The sharpest risk of automated military decision-making is not simply an unmonitored algorithm running wild; it is an officer positioned at the end of an opaque pipeline whose foundational reasoning has already evaporated. Research into AI-mediated executive judgment defines this structural flaw as a formation-authorisation gap—a rupture between the evidentiary grounds required to justify an operational commitment and those that remain accessible, intelligible, and challengeable when an authority finally approves it . This diagnosis recasts the conventional insistence on keeping a human in the loop. Physical presence at a terminal guarantees nothing if upstream abstractions have systematically sheared away assumptions, uncertainties, and model dependencies . Under such conditions, an authorizing human supplies institutional legitimacy rather than independent scrutiny, acting as a political shock absorber instead of a substantive control. Because none of the cited studies evaluates deployed weapons platforms, they do not quantify how regularly this collapse occurs on the battlefield; they instead reveal failure modes observed or formalized across other algorithmic environments . Their lesson for military command remains direct: accountability must attach to the entire trajectory along which evidence becomes authorizable, not just to the signature that rubber-stamps the result.

The mechanism driving this erosion is qualification attrition. Epistemically consequential qualifications degrade through successive data handoffs, even when every algorithmic filter or human analyst acts competently within their local remit . Picture a targeting workflow: ambient sensor noise hardens into discrete classifications, classifications are sorted into ranked strike options, and ranked options contract into a single actionable briefing. The ethical exposure is not merely whether each transformation was locally valid, but whether the caveats dropped along the way can be recovered under operational pressure. Critical caveats rarely disappear through explicit dissent. Rather, interface compression, procedural silos, and institutional velocity quietly translate conditional probabilities into apparent certainties—an operational realization of the dynamic described by process theory rather than an empirical audit of a specific military service . Responsibility inverts. Every contributor can demonstrate that their discrete deliverable met local guidelines, yet no individual retains access to the complete evidentiary foundation of the authorized act . The ultimate hazard is distributed moral deskilling, yielding a lethal directive that dozens helped construct but none remains equipped to challenge.

Adding an extra human checkpoint does not automatically solve the problem, because the quality of any review is bounded by what reaches the console. An evaluation of an industrial propose-verify-decide workflow showed that only 18 of 28 cases met expectations, autonomous strategy-workflow success languished at 3 out of 10, and the system correctly rejected invalid inputs in 7 of 8 instances . Failures stemmed from semantic distortion, incomplete evidence, and overlooked invalid inputs—defects capable of slipping past a streamlined dashboard . Crucially, every one of the four cases that cleared preliminary screening, generated complete evidence dossiers, and reached final engineering review received unconditional sign-off . That pattern suggests final approval often mirrors the polished construction of upstream evidence rather than any penetrating recovery of lost context . Because that investigation modeled digital-twin operations for virtual surgical-instrument sorting lines, its metrics cannot forecast weapon-system reliability; what transfers is the architectural distinction between parking a supervisor at the terminal and structuring auditable evidence throughout the pipeline . Meaningful human control demands traceable data provenance, visible uncertainty margins, and the procedural authority to halt and unroll earlier transformations—not a solitary button to confirm an algorithm's verdict.

Nor does collective deliberation provide a reliable escape, since structured communication can manufacture agreement while worsening the underlying choice. A study of distributed inference revealed that linear systems sharing identical eigenvalue and singular-value spectra generated opposite communication gains based purely on message orientation; in specific examples, changing that orientation raised decision accuracy from 72.6% to 91.2% or depressed it to 65.9% . Furthermore, shared community bias allowed higher mean individual accuracy to coexist with reduced global-vote accuracy or direct harm to unaffected groups . Imposing calibration constraints mitigated observed community harm while preserving much of the mean benefit, yet offered no ironclad safety guarantee . Applied cautiously to command environments, these findings caution that consensus among automated sensors, staff analysts, and commanding officers does not verify sound judgment. Deliberative convergence often indicates an unexamined shared framing circulating through a network rather than authentic, independent corroboration .

Automation can also disguise contentious ethical choices as objective technical parameters. In a sequential social dilemma, default strategies varied widely across eight large language models and diverged sharply from human baselines; conditioning those models on human-derived Social Value Orientation profiles, however, improved behavioral alignment by up to 70% . Higher induced prosociality scores reduced the probability that models would exit the game early, mirroring the cooperative dynamics observed in people . Yet the researchers emphasize that this reflects a capacity to map specified preference distributions onto strategic moves, not proof of internal moral preferences . The relevant governance question is never whether an automated agent holds internal values, but whose ethical priorities are encoded into its objective functions, fire thresholds, and decision trees. Multi-objective reinforcement learning highlights how volatile that moral tuning can be: while increasing fairness pressure generally favored equitable outcomes, moderate pressure produced an unexpected reversal, making responders more forgiving of substandard offers because payoff maximization and fairness goals clashed . An automated military optimizer can shift its operational behavior unpredictably when given competing ethical targets, and treating that shift as neutral model performance obscures the prior human choice of which trade-offs to force upon the machine .

Certain failure modes can be constrained through hard mathematical boundaries rather than human vigilance, but formal guardrails have inherent limits. A verified Ethereum electronic-invoice protocol formalized its lifecycle through guarded state transitions and machine-checked invariants, proving reimbursement uniqueness, face integrity, and authorization soundness under explicit cryptographic and consensus premises . Its smart contract reliably blocked duplicate, over-limit, or forged-receipt claims even when an agent's internal policy failed, demonstrating how a deterministic safety envelope can render prohibited actions mechanically unrepresentable . The military parallel is not that blockchain contracts should govern munitions, but that system designers must distinguish hard prohibitions that can be encoded as non-negotiable invariants from dynamic tactical appraisals that require situated moral evaluation. Even objective functions contain implicit value judgments: an audio bandwidth extension experiment noted that standard risk-neutral loss functions systematically smooth over rare high-frequency transients, prompting the use of discriminators engineered around tail risk and epistemic uncertainty . Though having nothing to do with warfare, that research illustrates a universal hazard: optimizing for aggregate, average-case metrics ignores the catastrophic outliers that carry the entire ethical burden of a strike decision .

Accountability must ultimately center on the preservation of active human reasoning, not the bureaucratic station of the person issuing an order. A theoretical analysis of substantive agency breaks the concept into coherent plurality, causal openness, anticipatory non-pointing, act-level singularisation, and endogenous sourcehood, while explicitly stressing that it does not prove human cognition satisfies its required bridge and certification premises . Whatever the scope of that formal framework, it clarifies why nominal human intervention can ring hollow: if options, framings, evidence sets, and execution timelines are entirely engineered upstream, the commander merely singularizes an inevitable outcome without acting as its authentic source. A more actionable standard emerges from the process theory's four pillars of judgment governance: boundary setting, interpretive challenge, reliance calibration, and authorization with answerability . Applied to armed conflict, these pillars require delineating operations that must never be delegated, establishing institutional channels for adversarial interpretation, calibrating reliance strictly to verified evidence, and holding accountable only those leaders granted the operational latitude to reshape the pipeline itself . The ethical test of military automation is not whether an officer's hand rests on the switch, but whether the entire chain of operational reasons remains accessible, intelligible, and challengeable before force is loosed.

Decisions · Articles · The Trolley Problem · Human–AI Alignment · Resonant