40 Years of Backpropagation: Remembering the Forgotten David Rumelhart

Revisiting the 40th anniversary of backpropagation and the overlooked legacy of deep learning pioneer David Rumelhart.
Using the anniversary of the backpropagation paper as a starting point, this article focuses on David Rumelhart — an underappreciated figure in the history of deep learning. It traces his foundational work in systematizing the backpropagation algorithm and advancing Parallel Distributed Processing (PDP) theory, noting that while today's large model training still relies on this mathematical framework, public narratives tend to spotlight recent commercial breakthroughs while overlooking early theoretical builders. Rumelhart's later diagnosis with Pick's disease — a cruel irony for someone who devoted his life to studying cognition — serves as a poignant reminder to respect those who do the quiet, long-term foundational work that makes future breakthroughs possible.
A Paper That Changed the Fate of AI
Modern deep learning stands on theoretical foundations laid decades ago. As we mark the anniversary of backpropagation — the algorithm at the heart of neural network training — one name deserves to be revisited: David Rumelhart. It was he and his colleagues who, in a landmark seminal paper, systematically articulated how multi-layer neural networks could be trained through error backpropagation.
The significance of that paper lies in the viable mathematical framework it provided for training neural networks. Before it, multi-layer networks lacked effective learning methods, and neural network research had fallen into a prolonged slump. The backpropagation algorithm enabled networks to adjust weights layer by layer via gradient descent, making it genuinely possible to learn from data.
The core idea of backpropagation is to use the Chain Rule to propagate prediction errors from the output layer backward through the network, computing the partial derivative (gradient) of each weight with respect to the final loss, and then updating weights via gradient descent. This process solved the "credit assignment" problem in multi-layer networks — determining which intermediate-layer weights should be held responsible for the final error. In 1986, Rumelhart, Hinton, and Williams published "Learning representations by back-propagating errors" in Nature, formalizing this approach and experimentally demonstrating that networks could automatically learn meaningful internal representations. While this was not the first time backpropagation had been proposed (Linnainmaa and others had arrived at similar derivations earlier), its clarity of exposition and direct application to neural networks made it the foundational reference the entire field recognizes today.
Why Rumelhart
When the public thinks of deep learning's founders, the names that come to mind are usually Geoffrey Hinton, Yann LeCun, and Yoshua Bengio. Yet David Rumelhart — a pioneer at the intersection of cognitive science and neural networks — has largely been overshadowed by those who came after him.
He was not only a core author of the backpropagation paper, but also a driving force behind Parallel Distributed Processing (PDP) theory. This framework understood human cognition as the parallel collaboration of large numbers of simple units, and it profoundly shaped the development of connectionism. His work spanned psychology, cognitive science, and computer science in an attempt to use computational models to explain how humans learn, remember, and reason.
PDP theory was most fully expressed in the two-volume work Parallel Distributed Processing: Explorations in the Microstructure of Cognition (1986), edited by Rumelhart and James McClelland. The framework argued that cognitive phenomena should not be understood as serial symbol manipulation, but as the emergent coordination of activation patterns across large numbers of interconnected simple processing units. This view directly challenged the Symbolicism paradigm that dominated cognitive science at the time and gave connectionism its theoretical legitimacy. The PDP framework carries a deep conceptual legacy into today's neural networks: the idea of distributed representation — that a concept is encoded across many neurons rather than stored in a single node — is the theoretical forerunner of modern word embeddings and representation learning in large language models.
Contributions Buried by Time
In discussions on communities like Reddit, many practitioners have used the anniversary of this paper to call for a reexamination of this history. The narrative of deep learning tends to focus on recent breakthroughs and commercialization milestones, while memories of the early theoretical founders gradually fade.
In his later years, Rumelhart was diagnosed with Pick's disease, a rare neurodegenerative condition that impaired his cognitive abilities. There is a profound irony in the fact that this scientist, who spent his life studying the mechanisms of cognition, was ultimately robbed of his ability to continue working by a cognitive disease. In his honor, the cognitive science community established the Rumelhart Prize, awarded annually to scholars who have made outstanding contributions to the theoretical foundations of human cognition.
Pick's disease is a subtype of frontotemporal dementia (FTD) that primarily damages the frontal and temporal lobes of the brain, leading to personality changes, language impairment, and executive dysfunction, while early-stage memory is relatively preserved — distinguishing it from the more common Alzheimer's disease. Rumelhart was diagnosed in the early 2000s and passed away in 2011. The Rumelhart Prize, established by the Cognitive Science Society and awarded annually since 2001 with a prize of $100,000, is one of the most prestigious awards in cognitive science. Its past recipients include Geoffrey Hinton and Yoshua Bengio, forming a line of continuity connecting early cognitive science with modern AI research.
Lessons for the Current AI Boom
The explosion of generative AI has filled the industry with noise, but amid that noise we should maintain a clear understanding of where the technology comes from. The training mechanisms that today's large models depend on still have, at their mathematical core, the backpropagation framework established decades ago. Revisiting this history reminds us that every technological leap of the present stands on the shoulders of those who came before.
Commemorating Rumelhart is not mere sentimentality — it is a form of respect for scientific lineage. In an era of ever-accelerating algorithmic iteration, remembering those who paved the road helps us understand the full arc of AI's development, and reminds those who follow that true breakthroughs almost always emerge from long-term, unglamorous, and rigorous foundational research.
Conclusion
The anniversary of the backpropagation paper is an opportunity to revisit the history of AI. David Rumelhart's name may not shine as brightly as today's star researchers, but his ideas have long been woven into every neural network that gets trained. To remember him is to remember the long journey AI took from an obscure theoretical pursuit to a mainstream technology.
Related articles

Getting Started with Codex: How OpenAI's Coding Agent Differs from ChatGPT
Codex is OpenAI's AI coding agent that autonomously reads, modifies code, and runs tests. Learn how it differs from ChatGPT and why developers need to master it.

Gemini Agent Launches: Argon Model Too Powerful to Release, Weekly AI Roundup
Google launches Gemini Agent for workplace use, supports Gemini 4 Argon and Claude Opus 5.5, but Argon stays unreleased over safety concerns. Weekly AI roundup covering Odyssey 3, OpenAI revenue, and Arena's $200M raise.

Sophos Cuts Threat Response Time by 96% with OpenAI Daybreak
Sophos CTO reveals how OpenAI Daybreak and custom security agents cut MDR average response time from 38 minutes to 89 seconds—a 96% reduction. A look at the plan-execute-observe loop architecture.