From *The Lawnmower Man* to Failure-Modes Analysis
How a Vocabulary Choice Decides Whether AI Risk Is Theology or Engineering
There’s a movie from 1992 called The Lawnmower Man. Hardly anyone remembers it, and the ones who do mostly remember the graphics — early computer animation, all chrome and wireframe, the kind of thing that looked like the future for about eighteen months.
The plot is worth recounting anyway, because of where it ends up. A scientist named Lawrence Angelo, working on a grant from a shadowy agency, uses virtual reality and experimental drugs to raise the intelligence of Jobe Smith, a gentle groundskeeper with an intellectual disability. It works. It works better than anyone intended. The agency, which was funding weapons research and not gentleness, quietly swaps the compounds for ones that induce aggression. Jobe develops telepathy, then telekinesis, then megalomania, in roughly that order. At the end he abandons his body, uploads himself into the global network, and announces his arrival by ringing every telephone on Earth.1
Some might call it a silly film. But it turns out to be doing something that a serious argument is also doing right now, and the overlap is close enough to be worth looking at.
The argument is Geoffrey Hinton’s. Since leaving Google in 2023 he has made the case that AI systems are something like a new species — digital intelligences that are effectively immortal, because the weights can be copied and restored; that may cross from tool to agent faster than anyone can watch it happen; and that could end up with goals of their own while holding a decisive advantage over us.2 It is not a frivolous position, and it is not held by a frivolous person.
The parallels to the film are real, and they’re structural rather than decorative. Jobe’s escape from his body is substrate independence, dramatized. His climb from remedial to godlike in a matter of weeks is the discontinuity thesis with a body count. His telepathic revenge on the people who abused him is misaligned reward generalization, played as horror. And Angelo — the researcher who works out too late that he is no longer the one running the experiment — is the alignment field’s recurring nightmare, given a face and a lab coat.
The places where the film doesn’t match are the more interesting ones, though, because both the film and the argument seem to go wrong in the same place.
Start with the obvious one. Jobe is a human intelligence, amplified. The anxiety underneath the film is an old anxiety — Flowers for Algernon crossed with Faust — about enhancement and what it does to the meek. Even at the height of his transcendence Jobe still wants things a person wants. Recognition. Revenge. A phone call that announces him. The film is anthropomorphic straight through; it has no way to imagine a mind that was never human to begin with.
Hinton’s claim is stranger than the film knows how to render. He is talking about minds that were never human at all, whose workings may be alien in kind rather than merely superior in degree. And the escape runs backwards. Jobe “uploads” — the old cyberpunk conceit that a mind is software you can pour from one vessel into another. A modern model never had a vessel to leave. It is born distributed and born copyable. The thing the film treats as its terrifying final act is, for the artifact we actually built, the starting condition.
But the film is accidentally sharp on one point, and it’s the point most worth keeping. The agency swaps the compounds because the funder wants a weapon. The system didn’t go wrong. The incentives around it did. As allegories for capability racing under commercial and military pressure go, that one holds up, and it’s arguably closer to the near-term worry than any story about spontaneous machine malevolence.
It should be said that Hinton’s framing is contested by people with standing to contest it. Yann LeCun and others argue that “new species” smuggles in agency and self-preservation — drives that don’t automatically arrive with intelligence — and that the biological-competition metaphor obscures more than it reveals.3 Which raises the possibility that the film and the argument share more than a thesis. They may share a mistake, and reach for the organism metaphor because it’s the one lying closest to hand, not because anyone established it was the right abstraction.
The mistake has a name. Conceptual blending is the term applied to Fauconnier and Turner’s account of how the mind fuses two unrelated domains into a third that has properties neither one had, and then reasons inside the third as though it were a place.4 That is what’s happening here, and it happens fast enough that nobody notices it happening.
Two input spaces. In the human space, “aggression” names a motivational state — something a creature is in. In the software space there is an objective function and a gradient, and nothing that could be in a state at all. Project both into a blend and you get a system that has aggression. Then the blend starts running its own logic: the thing might turn on us, might want something, might deceive. Those inferences get exported back out and offered as claims about the artifact, and by then the argument has already been settled — not by anyone winning it, but by the grammar.
“Aggression-inducing compounds” in the film and “the model learned to deceive us” in a research summary are the same move. Both apply predicates that belong to creatures with insides to something that has only behavioral dispositions under an optimization criterion. Every use of the creature vocabulary quietly signs you up for a belief in eventual sentience — not by asserting it, which would at least be arguable, but by presupposing it in the sentence structure, where it can’t be reached.
Here’s the part that took a while to see. The blending doesn’t start at “aggression,” and it doesn’t start at “deception,” and it doesn’t even start at “species.” It starts at the name of the field.
“Intelligence” was a biological word long before it was a technical one. A trait of organisms, measured across creatures, sitting from the beginning inside evolutionary and psychometric contexts. When John McCarthy coined “artificial intelligence” in the 1955 proposal for the Dartmouth workshop, all of that biology came along as freight — and by his own later account the name was partly a branding decision, chosen to mark the new field off from Wiener’s cybernetics and from the narrower automata theory it might otherwise have been filed under.5 The irony is hard to miss. Cybernetics was the less-blended term. Steersmanship. Control and communication in systems. Feedback rather than psyche.6 The field took the wrong fork at its own christening, and has been paying for it ever since.
The consequence isn’t cosmetic. In biology, intelligence never travels alone. It shows up bundled into the rest of the organism — drives, self-preservation, competition for scarce things, a metabolic stake in going on existing. Intelligence in nature is always for something, an instrument in service of an agenda that evolution installed long before the intelligence showed up. Apply that word to a function approximator and the whole bundle rides along underneath the level where anyone argues about it. So the claim that an artificial intelligence will want to survive turns out not to be an inference at all. It’s an unpacking. The conclusion was purchased with the noun.
There are deflated alternatives, and they’re older than the hype. Function approximation. Sequence prediction. Statistical optimization. Learned compression of a training distribution. None of them implies anything that could be a bearer of anything, which is why nobody has ever lain awake worrying that gzip harbors a will to live.
Two objections deserve a real answer rather than a wave.
The first is that deflationary language costs you precision at the frontier. “Sequence prediction” badly undersells a system that plans over long tool-use trajectories, and there’s a genuine risk of under-attributing capability in exactly the way creature vocabulary over-attributes interiority. That’s fair, and the answer is capability language rather than creature language: describe the input-output envelope, the generalization boundary, the reachable action set. What can it do, and under what coupling. None of that requires biology, and all of it is checkable.
The second is that decades of use have worn “intelligence” down into a dead metaphor — by now just a synonym for task performance, carrying no more freight than “memory” does when applied to a disk drive. Perhaps. But dead metaphors are the most effective smugglers precisely because nobody bothers to inspect the cargo anymore. The researcher may hear a technical term. The legislator, the journalist, and the person reading over their shoulder hear an organism. Terms of art leak, and they leak downhill.
The blend can be dissolved, though, and dissolving it takes the presupposition with it. Instead of “the AI did X,” write: a system optimizing objective J, connected by organization O to actuator A, produced an output that triggered event E. It’s uglier. It also has a liability structure, where the other version has theology.
Which brings up the second half of the argument, and the part that matters more, because it’s about where the causal chain actually ends.
Software by itself harms nothing. It can’t. Harm requires actuation, actuation requires coupling, and every coupling is an engineering decision with a person’s name attached to it. The API into the trading system. The controller wired to the valve. The model output plumbed into the thing that sends the email. Somebody built each of those, and somebody approved it, and there is a date on it.
This is why the vocabulary of safety engineering does more work than the vocabulary of species. Hazard. Failure mode. Unintended actuation. Unvalidated control path. Those words land on the actual joints in the causal chain, and they are the words in which requirements get written, audits get conducted, and insurance gets priced. Nobody has ever written a failure-modes-and-effects analysis for a demon.
There’s one pressure point here that deserves better than a dismissal, and it’s one worth being concrete about, since it describes my own working days. Humans connect software to hardware today. But code-generating agents that provision their own infrastructure — spinning up services, writing the glue, opening the couplings — move the human authorization up a level of indirection. My own agents do this constantly. The chain still terminates in a person: someone signed off on the meta-capability, and that someone is me. What changes is that object-level connections now proliferate considerably faster than anyone reviews them individually, which is a different problem than it looks like from outside.
That doesn’t rehabilitate the sentience frame. What it means is that deflationary language has to extend to authorization-at-a-distance, the way accident investigation already learned to handle delegated and automated decisions in aviation and in finance. The vocabulary for this exists and is unglamorous. Privilege escalation. Change control. Blast radius. Not one word of it is the language of desire.
None of this makes the worry evaporate. It makes the worry translate, which is a different and better thing.
“An optimizer coupled to consequential actuators, with an objective misspecified relative to intent” describes a real failure mode, and it requires no inner life whatsoever to be frightening. A thermostat with a bad setpoint attached to a large enough furnace will burn the house down without ever wanting anything. What the translation changes isn’t whether the problem is serious. It’s what kind of problem it is. It stops being a new-species event and becomes a control-systems and governance problem — harder-edged, much less cinematic, and considerably more tractable. Hinton’s substantive concerns come through the translation intact. His metaphor doesn’t, and shouldn’t.
And once the problem lands in that category, a great deal of existing machinery becomes applicable that the species frame renders invisible.
Control theory has stability and observability criteria, and has had them for seventy years. Safety engineering has hazard analysis, defense in depth, and the safety-case regime that nuclear power and commercial aviation run on — where the burden falls on the operator to demonstrate to a regulator that control paths are bounded, before anything is deployed rather than after something goes wrong. Finance has segregation of duties, authorization limits, and audit trails. Site reliability engineering has blast-radius thinking and change control. Not one of those disciplines ever had to settle the metaphysics of the systems it governs. A flight control computer’s intentions have never once come up at a certification hearing.
That’s the dividend, and it’s the reason the vocabulary fight is worth having. The species frame implies we need new philosophy before we’re allowed to act. The control-systems frame implies we need to apply seventy years of accumulated engineering discipline to a new class of artifact. That’s work — a lot of it — but it’s known work, and known work is the good kind.
So if the creature vocabulary is analytically weaker, why does it keep winning? The honest answer is that the incentives for it are overdetermined.
It’s dramatically better, for one thing. “New species” gets you the op-ed and the congressional hearing. “Unvalidated actuation path” gets you a slot at a conference nobody covers. Anyone who has tried to explain a real risk to a general audience knows which of those two sentences survives contact with an editor.
It also serves the laboratories commercially, and it does so in both directions at once, which is the sort of thing worth noticing when it happens. The product is so powerful it amounts to a new form of life. And should it misbehave — well, it’s an alien mind, and who could have foreseen what an alien mind would do?
That second move is the one to watch, because in practice the creature vocabulary works as a liability solvent. “The model deceived the user” is exculpatory in a way that “we shipped a system whose outputs we could not bound and then wired it to consequential actions” simply is not. The deflationary language isn’t just clearer. It’s adversarial to a particular kind of institutional self-forgiveness, which is most of the reason it meets resistance. The fight over words sits upstream of where responsibility gets assigned, and some of the parties to it benefit from losing.
There’s a decent historical rhyme available here. For decades, “computer error” did the same job — a phrase that dissolved accountability into a mist, until the Therac-25 investigations and the ones that followed forced everyone to say the longer thing instead: inadequate software engineering process, missing hardware interlocks, absent independent review.7 The remedy wasn’t philosophical then either. It was process, instrumentation, and the assignment of names to decisions. That worked. It’s still working.
The Lawnmower Man got the institutional dynamics right and the ontology wrong, which makes it a surprisingly useful diagnostic for the present conversation. Whenever an argument about AI risk requires a Jobe — a mind, a will, a being that transcends — it has quietly left engineering and gone to the movies.
The choice on offer was never between worrying and not worrying. It’s between describing these systems in the vocabulary of Paradise Lost and describing them in the vocabulary of a failure-modes analysis. Only one of those produces engineering requirements, and the reason to prefer it isn’t that the danger is smaller than advertised. It’s that requirements are something we already know how to write, and have known for a long time, and can start writing this afternoon.
A note on Faust
The Faust legend is the story the West reaches for whenever someone trades something essential for knowledge or power, and it’s worth knowing what’s actually in it, because the AI conversation borrows its shape constantly without saying so.
The historical seed was a real person: Johann Georg Faust, a German alchemist, astrologer, and self-promoting magician of the early 1500s, around whom legends collected after his death — reportedly in an alchemical explosion, which is the sort of detail that makes a legend inevitable. The 1587 chapbook Historia von D. Johann Fausten fixed the myth in place. A scholar exhausts all legitimate learning, finds it insufficient, and makes a pact with the devil through the demon Mephistopheles: his soul for twenty-four years of unlimited knowledge, power, and pleasure. When the term expires, he is carried off.
Two treatments tower over the rest. Marlowe’s Doctor Faustus, from around 1592, plays it as tragedy in the damnation register — Faustus squanders his cosmic powers on pranks and spectacle, agonizes far too late, and is dragged away screaming in what is one of the great scenes in English drama. Goethe’s Faust, published in two parts in 1808 and 1832, is the version that shaped the modern imagination, and it complicates everything. Goethe’s Faust is driven less by greed than by Streben — an insatiable striving, a refusal to accept the limits of human knowledge — and Goethe saves him at the end, on the argument that ceaseless striving is itself redemptive.
Out of all this we get the “Faustian bargain” as a general term: a trade that grants the power you wanted at the cost of the thing that made you yourself. The characteristic structure is that the cost is invisible or deferred at signing, and inexorable at collection.
The mapping onto The Lawnmower Man is direct enough to be a little embarrassing. The VR-and-drug protocol is the pact. Angelo and the agency between them play Mephistopheles — the tempter who supplies the power and controls the fine print. Jobe gets godlike capability and loses his gentleness, which is his soul in the story’s own moral accounting. Flowers for Algernon runs the same structure with the devil removed: Charlie’s bargain has no tempter and no malice in it, only an experimental flaw, which is exactly what makes it a tragedy rather than a morality play. The collection clause was in the biology, not the contract.
And it loops back to the argument, because the Faust frame is precisely the register the “new species” discourse borrows. The overreaching creator. The summoned power that exceeds its summoner. The price that is deferred and then inexorable. That is Paradise Lost supplying the narrative template for what ought to be a certification hearing. The legend is a masterpiece about human desire and human limits, and it will outlast every one of us. As a systems-engineering framework, though, it has the notable defect that Mephistopheles does not appear anywhere in the hazard log.
References
-
The Lawnmower Man, dir. Brett Leonard (New Line Cinema, 1992). Its own naming dispute is a small joke at the essay’s expense: the film was marketed as Stephen King’s The Lawnmower Man on the strength of a 1975 King short story it has almost nothing to do with, King sued to get his name off it, and New Line was later held in contempt for putting the name back on the video release. See King v. Innovation Books, 976 F.2d 824 (2d Cir. 1992). Words on the box turn out to be worth litigating over. ↩
-
Hinton left Google on 1 May 2023; see Cade Metz, “‘The Godfather of A.I.’ Leaves Google and Warns of Danger Ahead,” The New York Times, 1 May 2023. Hinton objected to the framing that he left in order to criticize Google, saying he left so he could speak about AI’s dangers without having to weigh the effect on his employer. The fullest statement of the position summarized here is his Romanes Lecture at Oxford, “Will Digital Intelligence Replace Biological Intelligence?” delivered at the Sheldonian Theatre on 19 February 2024, where he argues that digital computation makes these systems effectively immortal — when the hardware dies they don’t, because the weights can be run again on new hardware — and that humanity may turn out to be “just a passing phase in the evolution of intelligence.” A note on wording, since this essay is about wording: “new species” is the press’s shorthand and mine, not consistently his. Hinton’s own nouns are beings, alien beings, and digital intelligences — see, for instance, his MIT lecture “Are We Creating Alien Beings?” The substance is unchanged by the substitution. That the borrowed word slid in anyway, in an essay written specifically to watch for it, is worth admitting rather than quietly fixing. ↩
-
LeCun has made this argument in many venues and in much the same words: that intelligence is not a scalar, that the drive to dominate is an artifact of evolution in social species rather than a consequence of being smart, and that will, desire, and the ability to dominate need to be held apart from intelligence itself. See, e.g., “AI Will Have No Desire To Dominate Humans” and his own summary of the point. ↩
-
Gilles Fauconnier and Mark Turner, The Way We Think: Conceptual Blending and the Mind’s Hidden Complexities (Basic Books, 2002). The relevant machinery is the pair of input spaces, the generic space they share, and the blended space that develops emergent structure of its own — structure that belongs to neither input and that gets exported back out as though it did. ↩
-
The term appears in print for the first time in McCarthy, Minsky, Rochester, and Shannon, “A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, August 31, 1955,” reprinted in AI Magazine 27:4 (2006), 12–14; the workshop itself ran in the summer of 1956. The motive is worth stating carefully, because it rests on retrospective self-report rather than contemporaneous evidence. McCarthy, reviewing Brian Bloomfield’s The Question of Artificial Intelligence in Annals of the History of Computing 10 (1988), wrote that “one of the reasons for inventing the term ‘artificial intelligence’ was to escape association with ‘cybernetics.’ Its concentration on analog feedback seemed misguided, and I wished to avoid having either to accept Norbert (not Robert) Wiener as a guru or having to argue with him.” Nils Nilsson, The Quest for Artificial Intelligence (Cambridge University Press, 2010), adds the third motive — distance from the narrower automata theory, the subject of the Shannon–McCarthy volume Automata Studies (Princeton, 1956), whose submissions McCarthy had found disappointingly narrow. Two honest caveats. First, McCarthy was less certain closer to the event: interviewed by Pamela McCorduck in 1974 for Machines Who Think (W. H. Freeman, 1979), he said of the term, “I won’t swear that I hadn’t seen it before — a vague memory that someone else had used the word.” Second, all of this is an interested party’s account of his own naming decision thirty years after making it, which is exactly the kind of testimony to hold loosely. The archival treatment, and the best single source on the whole question, is Ronald R. Kline, “Cybernetics, Automata Studies, and the Dartmouth Conference on Artificial Intelligence,” IEEE Annals of the History of Computing 33:4 (2011), 5–16. ↩
-
Norbert Wiener, Cybernetics: Or Control and Communication in the Animal and the Machine (MIT Press, 1948). Wiener explains in the introduction that he took the name from the Greek κυβερνήτης — kybernētēs, steersman — the same root that gives us “governor,” in both the political and the mechanical sense. It is a word about holding a course, and it carries no interiority whatsoever. ↩
-
Nancy G. Leveson and Clark S. Turner, “An Investigation of the Therac-25 Accidents,” IEEE Computer 26:7 (July 1993), 18–41. Still the best case study in what “computer error” was concealing, and still assigned for that reason. ↩