In 2025, the Center for Strategic and International Studies (CSIS) and Scale AI tested seven major foundation models across 400 diplomatic and security scenarios — and the results showed significant divergence: some models were uniquely hawkish, and every model exhibited national bias.
CSIS and Scale AI's 2025 experiment
The 2025 experiment is the clearest empirical signal that AI is not a single, neutral instrument. Seven foundation models, presented with identical international crises, produced markedly different recommendations. The test exposed both divergent judgment and national bias, underlining a central complication for allied militaries: different AI architectures will not naturally agree on courses of action, even when fed the same inputs.
September 2025 Freedom Edge exercise
Allied forces have already begun exploring practical integration. During the September 2025 Freedom Edge trilateral exercise, the United States, South Korea, and Japan demonstrated real-time data exchange by linking their respective simulation systems with AI tools. That demonstration showed technical connectivity is possible — but the CSIS/Scale AI results warn that connectivity alone does not answer the question of mutual trust in automated judgments.

Audit-ready is a season. It shouldn't be.
Evidence in spreadsheets, controls drifting between audits, frameworks multiplying on flat headcount. Nubivance runs continuous compliance on Rapid7 Cyber GRC - SOC 2, HIPAA, ISO 27001, PCI, CMMC.
End the scrambleJapan’s two-track strategy: SoftBank, Sakura Internet, Palantir, and Anduril
Tokyo has adopted a deliberate two-track approach. Before June 2026 Japan often leveraged existing U.S. systems under the alliance framework while building sovereign capacity through domestic firms such as SoftBank and Sakura Internet. Microsoft and OpenAI investments with those Japanese tech companies signaled that U.S. AI capabilities remain a plausible option on Japanese soil. By mid‑August 2026, Japanese firms including Fujitsu had adopted Palantir-based solutions, and the government was considering Palantir’s Maven Smart System as well as Anduril’s Lattice. Nikkei reported that the government “envisions bringing in multiple AI systems,” a dual strategy aimed at combining allied technology with domestic safeguards for Self‑Defense Force information.
South Korea’s Defense AI Transformation and the Philippines’ choices
South Korea declared 2026 the year of the Defense AI Transformation and is pushing an independent path. Naver launched a dedicated defense AI organization and is entering the military market with its own foundation models and sovereign AI capabilities, while SK Telecom and Hanwha Systems have joined the competition. By contrast, the Philippines — lacking large domestic AI enterprises — remains undecided and could move quickly to adopt foreign platforms such as Palantir’s if it opts for an outside supplier.
NIST’s AI Risk Management Framework as the common yardstick
As differing national choices take root, the National Institute of Standards and Technology (NIST) AI Risk Management Framework (AI RMF) emerges in the source as the likely mechanism for translating difference into interoperability. Traditional risk management frameworks cover middleware security and interoperability; the AI RMF addresses the newer dimension of model trustworthiness. It can quantify how resilient a model is to deception such as camouflage or spoofing, how accuracy degrades outside training conditions (fog, nighttime), and how reliability changes with poor sensor data quality. Rather than a binary stamp, the AI RMF produces actionable measures of a system’s strengths and breaking points — information commanders need if they are to exercise meaningful “human‑in‑the‑loop” control instead of rubber‑stamping automated outputs.
NSPM‑11 and AI Forge: U.S. policy and the technical build
Policy and programs are aligning to make verification operational. On June 5, 2026, the United States issued National Security Presidential Memorandum (NSPM‑11), anchoring national security AI policy on four pillars — adoption, adaptation, assurance, and accountability — and directing the creation of standardized Test, Evaluation, Verification, and Validation methodologies for national‑security AI. In early June 2026, DARPA, the National Science Foundation, and NIST’s Center for AI Standards and Innovation launched AI Forge, a national program focused on interpretability, control, adversarial robustness, and credible, verifiable benchmarks. With NIST centrally involved, the AI RMF is positioned to serve as the technical foundation that enables cross‑national verification.
If allies cannot trust each other’s AI, they cannot fight together. The practical path described in the source is not a single vendor or a forced standard: it is permitting diverse models but using a shared RMF to measure trustworthiness, to compare systems on the same terms, and to let human commanders set mission‑specific baselines. That yardstick — measurable, comparable, and continually applied as systems evolve — is the foundation the allies will need if coalition warfare is to remain coherent in an age of divergent military AIs.




