The Imitation Game Was Never a Benchmark
Turing's 1950 proposal, modern LLM evaluations, and what it would mean to test whether a model reasons
Abstract. Turing did not give us a universal leaderboard for intelligence. He replaced an underspecified metaphysical question with an operational experiment. Modern debates about whether LLMs think or reason are therefore largely debates about constructs, protocols, judges, and alternative explanations.
This essay is in progress.
1. The question Turing refused to answer
Alan Turing begins Computing Machinery and Intelligence with a question that still frames arguments about AI:
“Can machines think?”
He then refuses the apparently sensible next step—defining machine and think. If those words were settled by their ordinary use, he argues, the answer would amount to a Gallup poll. That is absurd not because definitions are useless, but because this pair of definitions would decide too much before an inquiry had even begun. They would inherit every ordinary-language intuition about minds, persons, tools, consciousness, agency, and intelligence, then ask us to treat the resulting consensus as an answer.
Instead, Turing makes the move on which the essay turns:
“Instead of attempting such a definition I shall replace the question by another, which is closely related to it and is expressed in relatively unambiguous words.”
This is an operational move, not a metaphysical solution. Turing does not show that a machine which succeeds in the game thinks; nor does he prove that thinking reduces to successful performance. He changes the question into one with identifiable participants, a communication channel, a procedure, and an observable outcome—we can now see exactly what was asked, how the exchange was arranged, and what counted as success. We may still dispute what a positive result warrants; we can no longer dispute what was tested.
That distinction matters because a win in this game is neither a general theory of intelligence nor a score by which every machine can be ranked. It is the outcome of one constrained encounter under one specified protocol—and that protocol is not as settled as it looks. Turing’s game is, on the page, a contest between a man and a woman before it is ever a contest between a human and a machine; what happens when a machine “takes the part of A” is exactly the question the next section has to answer.
2. The exact imitation-game protocol
Turing’s replacement question does not begin with a machine and a human. It begins with three people:
“It is played with three people, a man (A), a woman (B), and an interrogator (C) who may be of either sex.”
The interrogator cannot see or hear A and B. Through written answers, C must decide which labelled participant is the man. A’s task is to induce the wrong identification; B’s is to help C make the right one. This is not yet the familiar picture of a person trying to distinguish a human interlocutor from a machine.
The machine enters one sentence later:
“What will happen when a machine takes the part of A in this game?”
That sentence creates a real ambiguity. Read literally, the machine takes the role of a man whose object is to persuade the interrogator that he is a woman. On that reading, the machine’s assignment is gendered mimicry. Read more generally, the initial game supplies the form of an experiment: a text-only interrogator confronts a participant whose success depends on producing behavior that defeats an attempted identification. The machine then inherits the form of the game, not the specific task of impersonating a woman.
I take the second reading to be Turing’s operative one. His own later restatement removes the gendered roles altogether:
“Are there imaginable digital computers which would do well in the imitation game?”
The earlier setup has not disappeared; it has done its work. It has shown why the exchange is text-mediated, why the interrogator must make a forced identification, and why outward performance rather than physical resemblance is at issue. But the later formulation no longer asks a machine to enact womanhood. It asks whether an imaginable digital computer can sustain the relevant sort of conversational performance.
That is why I will use behavioral indistinguishability under a conversational protocol for the target of the game. The phrase is deliberately narrower than “intelligence.” It says what the protocol observes: behavior in an exchange, assessed by an interrogator under specified constraints. It does not settle what that behavior reveals about consciousness, understanding, or origination.
3. Nine objections, grouped
Turing gives nine objections to machine intelligence. I will not repeat them in his order, because they do not all make the same demand. Some challenge what behavior can establish. Some predict that machines will simply be unable to do certain things. Some ask a different question altogether. That distinction matters more in 2026 than it did in 1950, because the world has changed most dramatically for the second kind of objection.
I group the objections into five sets: consciousness and origination, which concern the evidentiary status of behavior; disabilities and informality, which make crude capability claims; the mathematical objection; continuity and the consequences of machine intelligence, which I will name but not try to settle here; and two historically situated objections, theological argument and extrasensory perception. All nine matter. They do not all carry equal weight for this essay.
What behavior cannot establish
The Argument from Consciousness is the cleanest limit on the imitation game. Jefferson’s demand was not merely that a machine produce a sonnet or answer questions. It was that the production arise from felt thought and emotion. Turing identifies the sharpest version of the problem:
“the only way by which one could be sure that a machine thinks is to be the machine and to feel oneself thinking.”
That standard would make knowledge of other human minds impossible too. Turing therefore treats everyday attribution of thought as a convention we need in order to talk with one another. I take his more limited point: a successful imitation game can give us evidence about a system’s behavior and about our willingness to attribute thought to it. It cannot, by itself, turn private experience into a publicly measured variable.
Much has changed since 1950 without removing that limit. We can now inspect biological and artificial systems at scales Turing could not have imagined: the FlyWire consortium published a wiring diagram for an entire adult fly brain, with 139,255 neurons and 54.5 million synapses; contemporary multimodal models accept and generate across text, image, audio, and video. [Dorkenwald et al., 2024] [OpenAI, 2024] Neither a richer behavioral repertoire nor a more detailed wiring diagram supplies first-person evidence. That is not a counsel of despair. It is a boundary on the conclusion this protocol can support.
Lady Lovelace’s objection reaches the same boundary through a different word: originate.
“The Analytical Engine has no pretensions to originate anything. It can do whatever we know how to order it to perform.”
Turing’s reply is often flattened into the claim that machines can be creative. It is more careful than that. Lovelace did not have evidence that Babbage’s machine could originate; Turing argues that this absence of evidence did not establish a limit in principle. He also concedes that surprise alone will not silence the critic, because the critic can always ask whether the surprise reflects a creative act by the machine or by the observer.
The contemporary case is therefore worth taking seriously on both sides. AlphaEvolve, for example, combines Gemini models with automated evaluators and an evolutionary search loop; DeepMind reports that the system found a 48-scalar-multiplication algorithm for 4×4 complex matrix multiplication, improving on the previously best known result in that setting. [Google DeepMind, 2025] This is good evidence against a casual claim that learning systems can only repeat their prompts. It is not a clean answer to Lovelace. The result belongs to a system comprising a model, a search process, an evaluator, and a human-specified objective. Deciding where origination resides within that arrangement is the dispute, not a nuisance to be waved away.
What 1950 got wrong about capability
Turing’s Arguments from Various Disabilities collect claims that machines will never be kind, resourceful, funny, capable of learning, able to use words properly, or able to do anything genuinely new. The Argument from Informality of Behaviour makes a related move: human conduct cannot be reduced to an explicit rulebook, so humans cannot be machines. Turing separates rules of conduct, which a person can state and follow, from laws of behaviour, which may govern a person without being available as a rulebook. [Turing, 1950, pp. 448–453]
By 2026, the crude capability claims cannot survive intact. Systems can participate in extended text exchange, interpret images and audio, generate text, images, and speech, and support specialised workflows that would have looked unlike the small, fixed-purpose machines from which Turing’s contemporaries generalized. [OpenAI, 2024] This does not show that every item in Turing’s list has been met; terms such as kind, understand, and learn already contain the gap between passing and possessing. It does show that a blanket claim of machine incapacity is no longer a serious general argument.
The same care applies to informality. A model’s behavior may be difficult to predict, may not be representable as a short rulebook, and may adapt across a large range of inputs. None of that proves understanding. But neither does our inability to write down a complete list of human rules prove that the behavior is non-mechanical. Section 3 can close the crude capability question—can machines produce behavior of these kinds? often, yes—without closing the deeper question of what producing it establishes.
Objections to name, not solve
Turing’s “Heads in the Sand” objection says that the consequences of thinking machines would be too dreadful, so we should hope they are impossible. [Turing, 1950, p. 443] This is not a dead 1950 anxiety; it is recognisable in contemporary arguments about AI risk. But it concerns the consequences of machine intelligence, not the evidentiary meaning of a successful conversational performance. A system can be dangerous without passing an imitation game, and a system can pass one without settling any question about catastrophic risk. I will park the objection for that reason, not because it lacks force.
The Argument from Continuity in the Nervous System deserves the same honesty. Turing’s response is that a discrete machine need not reproduce a continuous system’s internal process exactly to give the right sort of typed answers under the conditions of the game. [Turing, 1950, pp. 451–452] That is enough to show why continuity does not automatically defeat his protocol. It is not enough to settle whether the relevant differences matter for mind; this essay will not pretend otherwise.
A 1950 landscape, and one durable mathematical point
Theological argument and extrasensory perception reveal how contingent an objection can be. Turing largely dismisses theological argument as speculative and historically unreliable [Turing, 1950, pp. 443–444]; by contrast, he treats telepathy as a serious threat to the game and proposes a telepathy-proof room [Turing, 1950, pp. 453–454]. I read that contrast as historical texture, not live evidence for or against machine intelligence.
The Mathematical Objection is different. Turing grants that formal results set limits on any particular discrete-state machine. His reply is that no proof shows humans exempt from comparable limits, and that a question defeating one machine may be answered by another. [Turing, 1950, pp. 444–445] This is not a claim that present-day LLMs solve Gödel-style limitations. Nor is it a claim that one general model should excel at every task. It is a warning against turning a limitation of one formal system into a proof of uniquely human superiority.
What changed from 1950 to 2026 is substantial: machines are no longer confined to the narrow, visibly mechanical roles that made the disability arguments plausible. What has not changed is the gap between an observed performance and a conclusion about consciousness, understanding, or origination. Turing’s objections become clearer when sorted this way. Some were bets about what machines could do, and history has overturned many of those bets. The most important ones were always questions about what behavior can prove. The next task is to ask what a passed protocol can actually justify.