Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

AlphaGo’s Lost Art: Why Today’s LLMs Fall Short on True Reasoning

Дата публикации: 02-10-2026 21:12:16

A decade after AlphaGo blended intuition with search to defeat Lee Sedol, today's LLMs rely on next-token prediction without separate reasoning mechanisms. New research exposes missing epistemic states, unfaithful chains of thought, and absent abductive leaps. Performance improves but true reasoning remains elusive.

Основное содержимое страницы с новостью.

Ten years after AlphaGo stunned the world by defeating Go champion Lee Sedol, the machines that followed took a different path. They excel at fluent conversation and pattern matching on a massive scale. Yet something essential remains missing.

The Machinery That Won Go

Move 37 in that historic match looked like a spark of intuition. Commentators called it creative genius from the machine. But the real story, as detailed in a new MIT Technology Review analysis, points to something else. AlphaGo combined rapid intuition with deep search. The system evaluated countless future positions. It revised its assessments as it went. Neither element operated in isolation. Intuition alone would never have selected that surprising play. Brute force search alone could never have sifted the possibilities efficiently.

Today’s large language models operate differently. They predict the next token. They do it again and again. Recent gains in math and coding come from longer chains of this prediction. The model generates intermediate steps before committing to a final answer. But those steps come from the same mechanism. No separate reasoning engine kicks in. No persistent record tracks hypotheses, confidence levels, or open questions.

Three clear gaps stand out. Models hold no explicit, inspectable epistemic state. They blend knowledge and manipulation inside the same neural weights with no clean separation. And their displayed chains of thought often emerge after the fact. The model reaches an answer one way yet reports another. In medicine or engineering, the path matters as much as the destination.

But the performance looks good enough. Benchmarks climb. Headlines celebrate new reasoning models. Industry insiders watch closely. They know the difference between surface success and deeper capability. And the gap shows up repeatedly in fresh research.

Consider scientific discovery. Tom Zahavy argues in a September 2026 paper from the Proceedings of the 43rd International Conference on Machine Learning that LLMs master induction and deduction yet lack abduction. Einstein described discovery as a leap from experience to new axioms, then logical deduction from there. Modern models compress data well. They prove theorems from given premises. They cannot generate the novel explanatory hypotheses when data stays sparse. The “jump” stays out of reach. Zahavy points to General Relativity as a case study. Observational clues were limited. Pure pattern matching from existing literature would not have produced the theory.

Chain-of-thought traces, now standard in reasoning models, fail to tell the full story. A Quanta Magazine report from July 2026 highlights work showing that replacing correct reasoning traces with incorrect ones often leaves final performance untouched. Between 30 and 60 percent of thinking steps in some frontier models carry minimal causal impact on the answer. Researchers at the Santa Fe Institute found models solving analogy-like visual puzzles through surface-level shortcuts rather than generalizable logic.

Subbarao Kambhampati has grown blunt on the subject. His group presented a position at the same 2026 machine learning conference titled “Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!” The core issue is training. Next-token prediction makes producing a correct solution easier than producing an accurate description of how the model arrived at it. So the trace becomes post-hoc rationalization more often than faithful record.

Recent papers from the July 2026 ICML proceedings pile on. One shows chain-of-thought in the wild is not always faithful. Models produce coherent but contradictory arguments on simple comparison questions. Another reveals reasoning models struggle to control their own traces when instructed to avoid certain words. Compliance drops to single digits in some cases. A third demonstrates systematic failures in multi-agent settings. When information sits distributed across agents, collective accuracy collapses to 30 percent even as individual performance with full information exceeds 80 percent. Models fail to reason about what others know but have not said. They converge too soon on shared facts.

Apple’s earlier work on puzzle complexity, covered by Ars Technica, already signaled trouble. Performance falls off a cliff beyond certain thresholds. Reasoning effort increases with complexity up to a point, then drops even with token budget available. Standard models sometimes outperform reasoning variants on low-complexity tasks. Both collapse at high complexity. The illusion of thinking meets hard limits.

Yet the models improve. Scaling delivers gains. Post-training on reasoning data helps in narrow domains. Enterprises adopt them for tasks once considered out of reach. The question is not whether they deliver value. They do. The question is whether the field confuses sophisticated pattern completion with the machinery of reasoning that scientists, engineers, and strategists actually rely upon.

AlphaGo succeeded because it paired statistical intuition with explicit search and revision. Current architectures iterate prediction without that separation. They produce fluent output that mimics deliberation. But the ledger stays hidden. The beliefs stay entangled. The reported steps often rationalize a conclusion reached by other means.

Progress will continue. New benchmarks test multimodal logical reasoning across inductive, deductive, and abductive categories. They reveal persistent imbalances. Work on latent reasoning architectures raises oversight concerns because chains of thought lose their transparency when thinking moves inside the model. The field debates these limits openly now.

Insiders should watch the gap. Real reasoning in high-stakes settings demands inspectable states, separable knowledge and manipulation, and faithful accounts of process. Today’s systems approximate the outputs. They have not yet rebuilt the engine that let a machine surprise the world on a Go board a decade ago. The difference matters. And it has not closed.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1LLM уверенно называет шахматные ходы, которых на доске нет. Как я проверяю каждый ее ответ кодом05.8726-09-2026
2Can AI reason without words? A small model puts the idea to the test07.1222-09-2026
3The AI Inference Revolution Is Here015.215-09-2026
4Why Strategy Now Outweighs Software in the Global AI Contest012.1102-10-2026
5 The search for consciousness inside LLMs 01020-08-2026
6Vos retours sur les différents LLM07.9819-02-2026
7OpenAI’s new dots agent comes with a crew of friendly mascots012.4629-09-2026
8Gemini 4 Argon takes on GPT-6 Astra and Claude Opus 5.5 with aggressive pricing022.1301-10-2026
9Even AI can’t perfectly assemble Ikea furniture—yet012.2328-09-2026

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 9. Тональность: 0. Информативность: 11.14. Источник: www.webpronews.com.