2,015 Israeli military personnel participated in experiments that reconstructed a real-world targeting decision-support system to measure how human decision-makers respond to AI recommendations.
Replica AI decision-support system tested with 2,015 Israeli military personnel
The study, presented under the name "Black Box Warfare: Human Judgment and Military Decision-Making in the Age of AI," deployed a high-fidelity replica of an AI decision-support system (DSS) used in military targeting. Researchers ran two experiments to test the replica's impact on combat decisions, drawing on a sample of 2,015 Israeli military personnel. The setup intentionally reconstructed both the interface and functionality of the actual system in order to observe behavior in settings designed to mimic operational decision points.
Algorithmic aversion dominates, especially when collateral damage is high
Contrary to a widely held fear of automation bias—where users uncritically follow machine recommendations—the experiments produced "strong evidence of algorithmic aversion." In other words, participants tended to discount or reject algorithmic recommendations rather than follow them. That tendency was particularly pronounced in scenarios that involved high collateral damage. The result suggests that, at least in the conditions tested, human decision-makers become more guarded with AI input as the potential human cost rises.

The cyber insurance questionnaire just landed. Now what?
SOC 2, HIPAA, insurance renewals - someone has to own security strategy. Nubivance provides fractional CISO leadership without the full-time salary.
Get a security leadExplainable AI features reduce aversion and change evaluation
The researchers tested interface variations that included "explainable AI" features and found these features altered behavior. Specifically, integrating explainable AI reduced algorithmic aversion and "promoted more thoughtful evaluations of algorithmic recommendations." In plain terms, when the system provided explanations for its outputs, users were less likely to dismiss the AI outright and more likely to weigh its recommendations together with their own judgment.
Trust is dynamic: individual predispositions, perceived stakes, and informational features
The study frames trust in military AI not as a single static quantity but as a variable shaped by three interacting factors: individual predispositions, perceived operational stakes, and the informational features of the interface. Put differently, whether an operator accepts, rejects, or interrogates an AI recommendation depends on who that operator is, how risky the decision feels, and what the system actually shows them. The authors tie these findings back to a central claim: trust in military AI is dynamic and context-dependent.
What this means for technologists, policymakers, and end users
- Technologists and security teams: Interface design choices — especially the inclusion of explainable AI elements — materially influence whether operators accept algorithmic guidance. The study provides empirical evidence that such features can reduce aversion and encourage deliberation.
- Policymakers and procurement leaders: The experiments underscore that human oversight remains central. Because algorithmic aversion increased with potential collateral damage, policy approaches that assume simple automation compliance are misaligned with observed operator behavior.
- End users and operational commanders: The findings suggest that training, individual predispositions, and presentation of information will matter as much as underlying model accuracy. Systems intended to support targeting decisions should consider how explanations are surfaced to users, not just what recommendations are produced.
The study's central contribution is empirical: it moves normative debates about AI in warfare from hypothetical risk lists to observed human choices under controlled conditions. By showing both a default aversion to algorithmic recommendations and a concrete pathway to reduce that aversion through explainability, the research highlights that the path to greater human–AI coordination is not solely technical accuracy but also the design of interfaces and information flows.
For militaries and those who design their tools, the takeaway is straightforward and pointed: if human agency matters in high-stakes decisions—as the experiments indicate—then the next wave of investment should fund the parts of AI systems that make their outputs intelligible and usable, not just more confident predictions.




