The three words that will decide whether robots can kill people in war
Imagine this: Two countries are at war. Country X sends a drone into a major industrial city in Country Y, aiming to take out two propane tanks. A routine sequence. But this time, the drone never reaches its targets. Instead, Country Xâs drone accidentally strikes a wall nearby, explodes, and kills three young civilians.
The three words that will decide whether robots can kill people in war
A high-stakes battle over the future of âkiller robotsâ comes down to semantics.
When Country Y eventually retrieves the droneâs remnants for intel, it finds an AI supercomputer inside that reveals something unsettling. Instead of a human deciding what to strike â the AI did.
Key takeaways
- A Russian AI-enabled drone reportedly selected its own target in Ukraine, killing three civilians â an ominous case of machines making lethal choices without direct human intervention.
- The global debate over autonomous weapons has narrowed to two competing frameworks: âmeaningful human control,â which pushes for human intervention in lethal decisions, and the Pentagonâs more flexible âappropriate human judgment.â
- OpenAI has adopted the Pentagonâs language, accepting a standard that does not require a person to decide every lethal action and leaves its practical limits to future, case-by-case military applications.
The truth is, you donât have to imagine this scene â it just happened. For the first time in the Russia-Ukraine war, as reported recently in the New York Times, three Ukrainian civilians were killed by a Russian drone, developed, designed and released by humans, that, in the end, selected its target autonomously. And this new reality is shaping up to be the future of warfare.
Thatâs because a number of countries, including those with the worldâs most consequential militaries, are rejecting the idea that human beings need always be in control of weapons in war.
Instead, as new autonomous weapons technologies become more capable and existing international legal frameworks struggle to keep up, some countries are embracing a more expansive view of human responsibility: that people can exercise enough control not by approving each strike, but by designing, testing, and setting the rules under which a given weapon operates.
This shift has produced one of the most consequential policy debates of today. And ultimately, human dominion over âkiller robotsâ â as autonomous weapons are colloquially called â could come down to a battle between two three-word phrases: âmeaningful human controlâ versus âappropriate human judgment.â
A high-stakes semantic battle
In the early 2010s, âmeaningful human controlâ emerged as an initial framework in the first international discussions on regulating an acceptable level of human involvement (or lack thereof) in deploying autonomous weapons. While the term quickly became an initial organizing principle among many states within these debates, it also drew immediate opposition from several others, such as the US and Russia.
âThe problem is what does [meaningful human control] mean?â said Lena Trabucco, an expert on AI and human control and a non-residential fellow at the Stockton Center for International Law at the Naval War College. âAnd what makes something meaningful versus not meaningful human control?â There was both a lack of consensus about what âmeaningfulâ meant and what âcontrolâ meant, she said. But the problems went beyond simple semantic ambiguity. âGetting a whole bunch of countries to agree on a standard for what âmeaningful controlâ is,â Trabucco said, became âa near impossible task.â
Ultimately, while an exact definition of meaningful human control never totally solidified, the term became intelligible enough to facilitate continued international dialogue on how to police autonomous weapons. As Trabucco explained, âWe all kind of understood what we were trying to grasp with the idea of meaningful human control, even if we didnât agree on a kind of standard for what is âmeaningfulâ or what âcontrolâ exactly means.â
Brad Boyd, retired colonel and senior military fellow at Stanfordâs Center for International Security and Cooperation, said that in this international context, âmeaningful controlâ came to be understood as human involvement at the exact moment a given weapon is fired. However, from the US perspective, that consensus definition still left a number of problems unresolved.
One of the most important holes in the definition was the issue of timing. âThe release of a weapon could theoretically be minutes, hours, days, weeks ahead of when the weapon actually strikes the target,â Boyd explained. For example, a drone can be released to sweep a designated area and remain airborne for hours, searching for anything that matches its given target criteria. Seconds, minutes, or hours might pass between the moment a person launches the drone and when the machine finds and fires upon a target.
âThis expansion of the timeline became very difficult for the construct of âmeaningful human controlâ to actually seem like it was doing what people wanted it to,â Boyd said. The tension between certain technical or engineering problems and the policy language preferred by international forums created, from the US perspective, insurmountable obstacles to making meaningful human control a truly practicable framework.
So, the US adopted its own alternative: âappropriate human judgment.â
âAutonomous and semi-autonomous weapon systems will be designed to allow commanders and operators to exercise appropriate levels of human judgment over the use of force,â a key Department of Defense directive reads. The new language effectively moved the required point of human intervention away from the moment that a weapon is fired toward oversight of a given systemâs entire life cycle â from its design, to its development, to its deployment.
There are some contexts where a government might not need as much human control over a weapon to comply with international law, Trabucco said. Thatâs where the subtle preference of âappropriateâ over âmeaningfulâ matters. âIf weâre on the high seas, maybe itâs not as necessary to meet a super high threshold of human control because thereâs not much risk to civilians or civilian property in those contexts,â she said.
âNow, in a city, an urban environment,â where the risk of civilian harm and other collateral damage is much greater, Trabucco explained, âthen thatâs where that high threshold would become important.â The US sought flexibility to determine how much human involvement it deemed necessary, based on the battlefield context in question, as opposed to having a fixed, universal standard of âmeaningfulness.â
And why the move from âcontrolâ to âjudgment?â Well, Boyd explained that âanytime we automate anything, whether itâs automating a car or automating machinery, we are trying to make it go faster, more precise, et cetera.â So, instead of insisting that humans be involved in any given part of the process, which could slow combat operations down, the US simply aimed to ensure autonomous systems behaved according to legal and ethical standards, no matter what situation they were deployed in.
Who decides how much human judgment is appropriate?
âItâs not necessarily the control that we want. What we really want is the machine to reflect our values, our laws, and our regulations,â Boyd said. âWhen humans employ our values, laws, and regulations, we call that judgment.â
Though designed to avoid setting a universal standard of human involvement as demanded by meaningful human control, appropriate human judgment is not a totally empty phrase. According to the DoD directive, the framework requires testing systems, defining operational limits, assessing likely civilian harm, training operators, setting rules of engagement, and ensuring that a system remains within its authorized mission. And, in some contexts, those requirements may institute more thorough protections built into weapons than a simplistic condition that a human be the one to pull the trigger in the end.
But the framework is not without a core puzzle of its own â who decides how much human judgment is appropriate? And what happens when the private sector, as in the ones developing such technologies, adopts this language before we have an answer?
And now â the private sector is forced to pick a side
In July, the same month of Russiaâs autonomous drone strike, OpenAI did something important not many noticed: it revised its relationship to military uses of its technology once more. Just three years ago, OpenAI maintained a total ban on âmilitary and warfareâ uses of its technology. Now, after a few quiet revisions since 2023, a new five-page policy document outlined the companyâs provisions for just that. (Disclosure: Vox Media is one of several publishers that have signed partnership agreements with OpenAI. Our reporting remains editorially independent.)
Most significantly, in explaining its basic condition for employing its technology in military contexts, particularly those involving decisions over the use of lethal force, OpenAI borrows a familiar phrase: appropriate human judgment.
The timing was not subtle. The document arrived just months after the public showdown between OpenAIâs competitor Anthropic and the Pentagon over the formerâs reservations about military applications of its technology. That fight ended with President Donald Trump demanding the immediate cessation of all Anthropic use within the government. Mere hours after Anthropic was booted, OpenAI CEO Sam Altman announced his company had struck its own deal with the government. (Disclosure: Future Perfect is funded in part by the BEMC Foundation, whose major funder was also an early investor in Anthropic; they donât have any editorial input into our content.)
But OpenAI didnât just earn a contract in the wake of the Anthropic-Pentagon showdown, it also took a lesson â aligning your policy with the governmentâs wins you favor. Or, worse, that opposing the government bears a steep price.
âAppropriate human judgment,â the policy document reads, âdoes not require a human decision on every discrete system action.â The framework, instead, requires that âhumans make informed decisions about the conditions for deployment.â
As both Trabucco and Boyd noted, the âappropriateâ level of human involvement can vary, depending on the operating environment, the type of target, a particular systemâs technical prowess, the anticipated risk to civilians, and a number of other political, economic, and strategic considerations important to a military operation. That flexibility is operationally appealing to the military, of course. But it also has a cost â safeguards against handing over total control of lethal force to machines become hard to identify, harder to measure, and hardest to enforce.
When will we know when the human-machine balance of power in war becomes âinappropriateâ? The truth is â thereâs no clear answer.
By cosigning the Pentagonâs flexible framework, OpenAI accepts that this standard has no settled meaning, that its application will be decided case-by-case behind the walls of military bureaucracy, and that the government may need room to change its mind. Itâs a choice that suggests the company is less interested in establishing clear red lines and is more receptive to the militaryâs own versatility about how AI and autonomy might be used in lethal operations.
âHuman judgment over critical decisions must be meaningful in practice, not merely formal,â the companyâs principles document reads. But, instead of drawing its own clear boundary around what its technology will and wonât do, OpenAI has accepted ambiguity as the price of partnership.
That undoubtedly makes the company a more useful ally to the Pentagon â while making it harder for the public to know where human judgment ends and machine-controlled violence begins.
That undoubtedly makes the company a more useful ally to the Pentagon â while making it harder for the public to know where human judgment ends and machine-controlled violence begins.
As the development and deployment of autonomous weapons rapidly accelerates, without many guardrails in place at all, one question is worth asking right now: before other AI labs and tech companies adopt the framework in an effort to align themselves with the US, what does appropriate human judgment truly mean? Ironically, what the phrase doesnât mean may be whatâs most consequential.
For the sake of humanity, the semantics of policing autonomous weapons is worth clarifying â or we risk totally losing control.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.