The Seduction of the Perfect Dogfight
Recent years have been filled with headlines about AI agents defeating experienced human fighter pilots in simulated dogfights. In events like the Defense Advanced Research Projects Agency (DARPA) AlphaDogfight Trials, an AI pilot has won decisively,
sometimes without the human pilot scoring a single hit. These demonstrations are visually spectacular and serve as powerful proof-of-concepts, showcasing an AI's ability to process data and react at superhuman speeds. The AI can execute maneuvers that are physically demanding or counterintuitive for a person, giving it a distinct advantage in a pure, one-on-one aerial battle. However, the very conditions that make these demonstrations possible are what make them poor preparation for reality. They are ideal demonstrations, conducted in sterile virtual environments where all variables are known and the rules are clear. The AI often has access to perfect information about its opponent's position and speed, something a real pilot almost never has.
Reality Is Messy and Unpredictable
A real combat mission is nothing like a clean simulation. It is a chaotic, complex, and uncertain environment. Pilots face challenges far beyond just outmaneuvering an opponent. They must contend with electronic warfare that jams sensors and communications, navigate in GPS-denied areas, and deal with unexpected hardware failures. They operate under strict, often complex, rules of engagement and must coordinate with other human and autonomous assets in a confusing battlespace. The weather can turn, and the mission objectives can change mid-flight. None of these messy, real-world factors are typically present in the high-profile AI dogfights. As a result, the AI is being trained for a reality that doesn't exist. This creates a brittle system—one that performs exceptionally well within its narrow training parameters but may fail catastrophically when faced with a situation it has never seen before.
The Need for Representative Tasks
This is why the conversation must shift from ideal demonstrations to representative tasks. A representative task is a test that mirrors the complexity and ambiguity of an actual mission. Instead of a simple dogfight, an AI pilot should be tasked with completing a multi-stage mission that includes a multitude of challenges. This could mean successfully identifying a target while under electronic attack, navigating a complex route with faulty sensors, or making a judgment call about whether to engage a target based on incomplete intelligence and strict ethical guidelines. The goal should not be to simply prove that an AI can win a fight, but to verify that it can perform its mission reliably and safely under the wide range of adverse conditions it will inevitably face. The conflict in Ukraine has already become a live testing ground for AI systems, demonstrating their use in targeting, navigation, and electronic warfare in a heavily contested environment. This real-world feedback loop is invaluable because it exposes weaknesses that no lab can predict.
Building Resilient and Trustworthy AI
Focusing on representative tasks does more than just prepare AI for the real world; it builds trust. Military leaders, and the public, cannot have confidence in an autonomous system if its only claim to fame is winning a video game, no matter how sophisticated. True trust comes from seeing a system perform consistently and predictably in situations that are difficult and messy. Developing this trust requires a new approach to testing and evaluation. It means investing in complex, multi-agent simulation environments where AI has to cooperate and compete against other intelligent systems. It means accepting that AI development is not just about writing better algorithms, but also about creating better data and more realistic testing environments. Ultimately, an AI pilot should not be judged on whether it can defeat a human, but on whether it can be a reliable teammate that enhances the capabilities and safety of the entire force.














