In 1997, IBM's Deep Blue defeated world chess champion Garry Kasparov in a match. The achievement looked like a milestone on the road to machine intelligence: a computer had conquered a game associated for centuries with planning, calculation and human genius.
But imagine asking that same machine to perform the less glamorous part of the match — look at the board, identify the correct wooden piece, reach across the table, pick it up without knocking over its neighbors and place it precisely on another square.
The calculation was the easy part.
This upside-down relationship between human and machine difficulty is known as Moravec's paradox. Artificial intelligence has repeatedly become competent at tasks we regard as intellectually demanding while robots continue to find apparently trivial aspects of everyday physical life astonishingly difficult.
A toddler can navigate a cluttered room, recognize a familiar object from a strange angle, adjust a grip when something begins to slip and learn the consequences of bumping into a table. None of those abilities feels like advanced reasoning. Yet reproducing the entire package in a machine requires vision, touch, body awareness, prediction, control and constant adaptation operating together in real time.
Nearly four decades after Hans Moravec helped popularize the observation, modern AI has changed dramatically. The paradox has not disappeared. It has simply moved deeper into the physical world.
What Moravec actually noticed
Hans Moravec was a pioneer of mobile robotics at Carnegie Mellon University. During the 1970s and 1980s, researchers attempting to build machines that could see and move encountered a frustrating contrast.
Early AI programs had produced impressive results in highly structured intellectual domains. Computers could prove mathematical theorems, manipulate symbols and play games. Researchers therefore expected perception and physical manipulation to be relatively manageable additions.
They were not.
In a 1985 Carnegie Mellon report, Moravec described how a camera image arrives at a computer as an enormous array of numerical values. Humans immediately see people, trees, doors, screwdrivers and cups. For the machine, however, those meaningful objects do not arrive pre-labeled. They have to be extracted from the numerical pattern.
Early robot experiments therefore simplified reality dramatically. Instead of dealing with ordinary rooms, they used bright blocks against controlled backgrounds. Even then, systems far more computationally powerful than earlier programs that handled chess, geometry and calculus could struggle simply to locate and grasp a block reliably.
Moravec developed the idea further in his 1988 book Mind Children. Around the same period, AI researchers including Rodney Brooks and Marvin Minsky made related observations: capabilities humans experience as effortless can conceal enormous computational complexity.
The “paradox” is not a formal mathematical contradiction. It is an empirical reversal of intuition.
Why chess is a friendlier world than a kitchen
Chess is extraordinarily difficult for humans, but its universe is unusually polite to computers.
The board contains 64 squares. The pieces belong to known categories. Their legal moves are precisely defined. Both players can see the complete state of the game. The rules do not suddenly change because the room is dark, the queen is slippery or a pawn has been moved three millimeters off-center.
Most importantly, the objective is explicit: win the game.
A kitchen is nothing like this.
Tell a robot to “bring me that cup” and an avalanche of hidden problems appears. Which object is the cup? Is it partially hidden behind a plate? Is it ceramic, paper or flexible plastic? Is it empty or full? Where is its handle? Can the robot reach it without hitting something else? How hard should its fingers squeeze? What if the cup slips? What if someone walks between the robot and the table? What if the object looks unlike every cup in its training data?
The robot must continuously connect perception to action. Each action changes what its sensors perceive, which changes what it should do next.
That feedback loop is one reason physical intelligence is so demanding.
But modern AI can recognize cups
This is where the classic explanation of Moravec's paradox needs updating.
Modern computer vision systems can classify and detect ordinary objects with accuracy unimaginable in the 1980s. Large multimodal models can describe photographs, interpret scenes and answer questions about visual content. Saying that today's AI simply “cannot recognize a cup” is no longer accurate.
The harder problem is recognizing and manipulating the right cup robustly in the real world.
A photograph is fixed. Physical environments are not. Lighting changes, objects rotate, hands block cameras, surfaces reflect, materials deform and contact produces forces that vision alone cannot measure. A robot must distinguish not merely what an object is, but what can safely be done with it right now.
Research published in Nature Machine Intelligence in 2025 highlighted the importance of high-resolution tactile sensing for adaptive robotic grasping. A robotic hand needs feedback about pressure, contact and slip so that it can adjust its grip rather than blindly executing a precomputed motion.
Other recent systems combine vision, force sensing and large language models. One 2025 project demonstrated a robot completing multi-stage tasks including coffee making in changing environments by using an LLM for high-level planning while lower-level systems handled visual and force feedback.
That architecture itself illustrates Moravec's paradox. The machine can formulate a sophisticated plan in language, yet carrying out “pour the coffee” requires an entirely different layer of intelligence.
The million-year head start inside your body
Moravec offered an evolutionary explanation for the paradox.
Abilities such as visual perception, locomotion, balance and hand–eye coordination are ancient. Natural selection spent hundreds of millions of years refining nervous systems that could extract useful information from sensory signals and control bodies in unpredictable environments.
Explicit symbolic reasoning — mathematics, formal logic and chess-like deliberation — is evolutionarily much newer.
Humans therefore experience a strange psychological illusion. Because perception and movement happen largely outside conscious awareness, they feel simple. We notice the effort involved in solving an equation because we consciously struggle through the steps. We do not experience the enormous stream of computation required to keep our body upright while walking toward the refrigerator.
A one-year-old cannot explain inverse kinematics, estimate friction coefficients or describe visual segmentation. Yet the child's nervous system continuously solves practical versions of those problems.
This evolutionary account is influential, but it should not be mistaken for a precise scientific law predicting which task an AI will find difficult. Modern machine learning has shown that large amounts of data and computation can automate perceptual abilities once thought exceptionally resistant to machines. The basic insight survives more strongly in robotics, where bodies must interact with reality rather than merely interpret stored data.
Why physical AI cannot simply copy the language-model recipe
One reason language models improved so quickly is data. The digital world contains enormous quantities of text, images and video that can be copied, processed and used for training.
Robot experience is expensive.
A robot learning how to manipulate a drawer has to interact with a physical drawer, or a sufficiently realistic simulation of one. Motors wear out. Experiments take real time. Hardware breaks. Human demonstrations are costly to collect. A training mistake can knock an object onto the floor rather than merely produce a wrong word on a screen.
A 2026 editorial in Nature Machine Intelligence argued that this remains a defining asymmetry in AI. Systems can model and reason about the world impressively, while robustly perceiving, acting and adapting within the physical world remains an open challenge. Even mundane tasks such as opening ordinary doors across all their real-world variations can still expose weaknesses.
Researchers increasingly describe the solution as embodied intelligence or physical AI. The idea is that intelligence cannot always be separated neatly from the body performing the task. Sensors, actuators, mechanical design, materials and environmental feedback all participate in intelligent behavior.
A well-designed hand can make grasping easier before software even enters the problem. Soft materials can conform naturally to irregular objects. Tactile skin can detect contact where cameras cannot see it. Physical structure can therefore perform part of the “computation.”
Nature solved many of these problems the same way: brains evolved together with bodies.
Moravec's paradox in the age of humanoid robots
Robotics is advancing rapidly. Humanoid machines can now walk over difficult terrain, recover from disturbances, manipulate objects and execute impressive demonstrations. New tactile sensors can detect fine pressure distributions, vibration, temperature and texture. Simulation allows robots to practice enormous numbers of actions before attempting them in reality.
Yet demonstrations and dependable everyday autonomy are different standards.
A household robot must succeed not once in a carefully prepared laboratory but repeatedly among unfamiliar furniture, children, pets, clutter, liquids, fragile objects and humans who give ambiguous instructions. It has to know when it is uncertain and avoid turning a small mistake into an injury or broken object.
This is why the frontier of AI increasingly looks physical. Language and abstract planning have progressed so quickly that the bottleneck has become connecting those abilities to reliable perception and action.
The irony would have delighted Moravec. A robot may be able to explain the physics of friction, write computer code for a control algorithm and plan how to set a dinner table. Then it reaches for a glass and discovers that knowing is not the same thing as doing.
The paradox tells us something about ourselves
Moravec's paradox is ultimately as revealing about human intelligence as it is about machines.
We tend to rank abilities by how difficult they feel. Algebra seems intelligent because it requires concentration. Catching a falling object seems ordinary because the body does it before conscious thought has time to intervene.
AI reversed that hierarchy.
Computers showed that some prestigious intellectual tasks become tractable when the rules, states and objectives can be represented precisely. Robotics revealed that a skill so mundane we barely notice it — reaching for a cup — can require an extraordinary integration of perception, prediction, touch, mechanics and control.
Modern AI has narrowed the gap. Machines can now see, speak and plan in ways that would have seemed astonishing when Moravec formulated his observation. But the physical world continues to resist clean abstraction.
A chessboard tells the computer exactly where every piece is. Reality never does.