Primer

The Turing Test

Alan Turing proposed sidestepping 'can machines think' with a simpler question: can a machine talk its way into being mistaken for a person?

In 1950, the mathematician Alan Turing opened a paper with the question “Can machines think?”, then immediately set it aside as too vague to answer directly, tangled up in disputes about what “machine” and “think” even mean. In its place he proposed something more concrete, which he called the imitation game and which has been known ever since as the Turing test.

The setup

A human judge holds a text-only conversation with two hidden participants, one a human, one a machine, without being told which is which. If the judge cannot reliably tell the machine from the human based on the conversation, the machine passes. Turing’s proposal was that passing this test, sustained, convincing, general conversation, is a reasonable standard for calling a machine intelligent. He wasn’t claiming the test proves a machine truly thinks in some deep metaphysical sense, only that the question of “real” thought might not be answerable at all, and that a behavioral standard is the most useful one we’re going to get.

Why this test, and not some other

Turing’s move was deliberately operational: rather than trying to define thought and then check for it, he replaced an unanswerable question with a measurable one. This makes the test easy to apply and hard to argue with on its own terms, which is exactly what makes it controversial. It measures conversational performance, not internal process, and Turing treated that substitution as basically fine. Not everyone has agreed.

The Chinese Room objection

The most direct challenge comes from our primer on the Chinese Room. Searle’s whole point was that a system can produce fluent, convincing conversational output, exactly the kind of performance that would pass a Turing test, while, on his account, understanding nothing at all. If Searle is right, passing the Turing test shows a machine can imitate the outward signs of a mind without demonstrating that anything mind-like is happening inside it. The two thought experiments were built to test the same underlying assumption, that the right behavior is sufficient evidence of the right internal state, from opposite directions.

The test can be gamed

A separate, more practical worry is that fooling a human judge doesn’t require general intelligence, only good tricks. Joseph Weizenbaum’s 1966 program ELIZA, a simple script that reflected users’ statements back as questions in the style of a therapist, convinced some users they were talking to something that understood them, despite having no model of meaning whatsoever. This became known as the ELIZA effect: people’s strong tendency to attribute understanding to a system based on surface fluency alone. If human judges are this easy to fool, passing the test may say more about human psychology than about machine intelligence.

Why it’s live again

For decades the Turing test was mostly a thought experiment, since no machine came close to sustaining a genuinely open-ended conversation. That changed abruptly with modern large language models, which now pass informal versions of the test routinely, producing fluent, context-aware, often genuinely helpful conversation. This hasn’t settled the underlying philosophical question so much as made it urgent again: fluent conversation is now cheap to produce, which means the gap Turing tried to sidestep, between behaving like something understands and something actually understanding, is no longer hypothetical. It’s the exact question this journal exists to keep asking.