How the Brain Sees and Thinks Simultaneously: A Complete Guide to Predictive Coding Circuits in Visual Cortex

How the brain's predictive coding circuits merge vision and cognition, and what this means for AI design.
This article explores how the visual cortex uses bidirectional feedforward-feedback circuits to simultaneously perceive and predict the world. It explains predictive coding theory—where the brain generates internal predictions and processes only prediction errors—and discusses its validation through visual illusions and clinical pathology. The piece then draws implications for AI, including world models, JEPA frameworks, and spiking neural networks as paths toward brain-inspired efficient vision systems.
Where Vision Meets Thought
The human brain is an astonishing information-processing machine. At every moment, we are "seeing"—receiving massive visual signals from the external world—while simultaneously "thinking"—interpreting those signals based on experience, memory, and expectation. For a long time, neuroscience generally treated "perception" and "cognition" as relatively independent processes: the eyes receive, and higher brain regions think. However, a growing body of research is challenging this traditional distinction, revealing that the brain contains specialized neural circuits that allow "seeing" and "thinking" to operate in an intertwined manner at the very same moment.
The core question is: How does the brain manage to both faithfully represent the external world and actively predict and infer about it? This is not only a fundamental question in neuroscience but also carries profound implications for the architectural design of artificial intelligence.
Limitations of the Traditional Model: The One-Way Bottom-Up Flow
The Classical Visual Processing Framework
In the classical visual neuroscience framework, information processing is described as a one-way "bottom-up" pipeline: light is converted into electrical signals at the retina, transmitted through the Lateral Geniculate Nucleus (LGN) to the primary visual cortex (V1), and then proceeds level by level to higher areas that process increasingly complex features (V2, V4, IT, etc.). Lower brain regions handle basic features like edges and orientations, while higher regions accomplish object and scene recognition.
Extended Background: Structural Details of LGN and V1 The Lateral Geniculate Nucleus (LGN) is a thalamic nucleus that serves as the primary relay station from the retina to the visual cortex. It is far from a simple signal relay—the LGN has a six-layered structure that separately processes information from both eyes and distinguishes between the Magnocellular pathway (primarily responsible for motion perception) and the Parvocellular pathway (responsible for color and detail). Notably, the LGN receives far more feedback connections than feedforward inputs from the retina, which itself suggests that "top-down" modulation is already present at the earliest stages of visual information processing. The primary visual cortex V1, located at the posterior occipital lobe, contains approximately 140 million neurons. Its orientation-selective cells were discovered by Hubel and Wiesel in the 1960s, earning them the Nobel Prize and establishing the experimental foundation for hierarchical visual processing theory.
This classical framework was further refined in subsequent research into the dual-stream model of the "ventral pathway" and "dorsal pathway." The ventral stream extends from V1 through V2 and V4 to the inferotemporal cortex (IT), primarily responsible for object recognition and shape analysis—known as the "What pathway." The dorsal stream extends from V1 through V2 and MT to the parietal lobe, primarily processing spatial location and motion information—known as the "Where/How pathway." This framework, proposed by Ungerleider and Mishkin in 1982, laid the anatomical foundation for understanding the functional segregation of visual information and revealed that the visual system is far more complex than a single pipeline—even within the "bottom-up" framework, information is processed across multiple parallel channels.
Extended Background: Clinical Neurological Validation of the Dual-Stream Model The functional dissociation between the ventral and dorsal pathways is evident not only at the anatomical level but has been powerfully corroborated in clinical neurology. Damage to the ventral pathway can cause "visual agnosia," where patients can see an object's contours but cannot identify what it is—the classic case described by neurologist Oliver Sacks in The Man Who Mistook His Wife for a Hat is precisely this type. Damage to the dorsal pathway can cause "optic ataxia," where patients can recognize objects but cannot accurately reach out to grasp them. This double dissociation reveals that "knowing what something is" and "knowing where it is/how to interact with it" are independent computational processes supported by different neural substrates, providing neuroscientific justification for "perception-action" decoupled design in AI visual systems. Furthermore, recent research has found extensive cross-connections between the ventral and dorsal pathways, with their functional boundaries being far more dynamic and interdependent than the original model described, further emphasizing the integrative nature and complexity of the visual system.
This model is intuitive and profoundly influential—it even inspired the design of early deep convolutional neural networks. However, it has a significant flaw: it cannot explain the active nature of human vision. Why can we still "understand" an image under dim, blurry, or partially occluded conditions? Why are visual illusions so persistent?
The Brain Is Not a Passive Receiver
The human visual system continuously "fills in gaps" and "guesses" about the world based on existing knowledge. When you see a cat half-hidden behind tree branches, the brain automatically completes the occluded portion—this is precisely the "top-down" prediction mechanism at work. The brain has never been merely a passive camera.
Visual illusions are perhaps the most intuitive behavioral evidence of this mechanism. Consider the classic "hollow mask illusion": when viewing the concave back of a mask, the brain still perceives it as a convex face—because the brain's prior knowledge that "faces are convex" is so strong that it overrides the actual depth information from the senses. The clearly visible white contours in the Kanizsa triangle that don't actually exist are similarly the result of the brain using predictions to fill sensory gaps. Intriguingly, patients with schizophrenia are often immune to the hollow mask illusion—researchers believe this relates to abnormal modulation of sensory input by prior knowledge in their brains, providing a pathological reverse validation of the biological reality of the prediction mechanism.
The Neural Circuits That Let the Brain See and Think Simultaneously
The Bidirectional Dialogue of Feedforward and Feedback
The true key to enabling the brain's simultaneous "thinking and seeing" lies in the massive feedback connections within the visual cortex. Research shows that the number of neural fibers projecting back from higher brain regions to lower ones often exceeds the number of feedforward connections. This means that "primary" areas like V1 not only receive input from the eyes but also continuously receive "expectation" signals from higher brain regions.
This bidirectional dialogue exhibits a precise sequential structure in the temporal dimension. Neural recording experiments show that feedback signals from higher brain regions typically arrive at V1 approximately 50 to 100 milliseconds after the feedforward wave, creating a measurable temporal signature. This means the "feedforward perceptual phase" and the "cognitive modulation feedback phase" can be experimentally distinguished. This temporal structure not only provides direct electrophysiological evidence for predictive coding theory but also suggests that AI system design could incorporate explicit iterative refinement temporal mechanisms: the first forward pass forms an initial percept, followed by multiple rounds of feedback correction that gradually converge on the optimal interpretation, rather than relying on a single feedforward pass to produce the final result.
This bidirectional dialogue constitutes a dynamic circuit: lower-level areas report actual observational information upward, while higher-level areas send down predictions based on experience; the difference between the two—the "prediction error"—becomes the core signal through which the brain continuously corrects its cognition. This is the essence of predictive coding theory.
Notably, this "top-down" modulation is deeply bound to attention mechanisms. The prefrontal cortex (PFC) and parietal attention networks can selectively enhance or suppress neural activity in specific visual areas through gain modulation, amplifying perceptual signals relevant to the current task. This process relies on broadcast-style modulation by neuromodulators such as acetylcholine (ACh)—the brain uses chemical signals to globally switch "attention modes," allowing the relative weights of prediction and error signals to be dynamically balanced.
Extended Background: Multi-Dimensional Modulation by Neuromodulator Systems The biological implementation of the brain's attention system is far more complex than the mathematical abstraction of Transformers. The prefrontal cortex modulates neural gain in the visual cortex through long-range projections, and this process depends on the coordinated action of multiple neuromodulators: acetylcholine (ACh) enhances the signal-to-noise ratio of perceptual signals, making neurons more sensitive to sensory input while reducing background noise; norepinephrine (NE) regulates the system's trade-off between "broad sampling" and "precise focusing," corresponding to the exploration-exploitation balance in cognitive science; dopamine is closely related to reward-tagging of prediction error signals, forming the neural basis of reinforcement learning. This multi-dimensional, multi-timescale modulation mechanism endows biological attention systems with global state-dependence and cross-modal dynamic adaptability that Transformers lack—features that remain among the most difficult to fully replicate in current neuromorphic AI research.
This bears certain mathematical similarities to the attention mechanism in Transformers in artificial intelligence, but the latter lacks the temporally iterative error-correction loop—a fundamental architectural distinction between the two.
Predictive Coding: A Framework Unifying Perception and Cognition
Predictive coding theory posits that the brain is essentially a "prediction machine." It continuously generates internal models of the external world and uses sensory input to test and update these models: when predictions match actual input, neural activity is "suppressed," conserving energy; when they don't match, error signals propagate upward to trigger model updates.
Extended Background: Academic Origins of Predictive Coding The modern neuroscience version of predictive coding was systematically proposed by Rao and Ballard in 1999 in Nature Neuroscience, but its intellectual roots trace back to Helmholtz's 19th-century concept of "perception as unconscious inference." This idea was introduced into cognitive psychology by Gregory and others in the mid-20th century, and after decades of theoretical accumulation, was ultimately developed by Karl Friston in the early 21st century into the "Free Energy Principle." Friston argues that all cognition and behavior of the brain can be unified as minimizing "variational free energy"—the upper bound of prediction error. This framework draws on tools from statistical physics and variational Bayesian inference, unifying "minimizing perceptual error" and "actively sampling the environment to verify predictions" as two implementation strategies of the same objective function: the former corresponds to perceptual updating, while the latter corresponds to Active Inference—where organisms act to make the world conform to their expectations rather than merely passively updating their internal models. This property of Active Inference provides a unified decision-perception theoretical foundation for embodied intelligence: robotic systems need not design "perception modules" and "decision modules" separately but can share the same objective function of minimizing prediction error, with perception and action naturally coupled. This framework applies not only to vision but has also been used to explain attention, learning, emotion, and even consciousness, and is regarded as one of the most ambitious unifying theories in contemporary computational neuroscience. At the information-theoretic level, predictive coding is highly consistent with the "differential coding" concept in data compression: transmitting only the changes rather than the complete signal, thereby dramatically reducing information redundancy—which also explains why this mechanism can deliver significant energy efficiency advantages.
This mechanism elegantly explains why "seeing" and "thinking" can merge seamlessly: the world we "see" is actually a co-construction of sensory input and brain predictions. Vision is not pure reception but an ongoing process of inference.
Deep Implications for AI Architecture
From Unidirectional Feedforward to Bidirectional Predictive Networks
Current mainstream deep learning vision models remain predominantly feedforward in architecture. Although structures like Transformers achieve dynamic information integration through attention mechanisms, they still have a fundamental gap compared to the brain's deeply coupled bidirectional predictive circuits.
The brain's design points in a clear direction: a truly efficient, robust visual system needs to deeply integrate "perception" and "prediction" rather than separating them into independent modules. The recently emerging "World Models" research is exploring precisely this direction—enabling AI systems not only to recognize current input but also to continuously predict future states.
Extended Background: Technical Progress in World Models The "World Models" concept was formally proposed by David Ha and Jürgen Schmidhuber in 2018. The core idea is to let agents learn a compressed representation of the environment internally and simulate the future "in their heads" rather than relying entirely on real-time interaction with the real environment. Recent progress in this area has been rapid: DeepMind's Dreamer series models perform planning in latent space, dramatically improving sample efficiency in reinforcement learning; Meta's V-JEPA attempts to build hierarchical predictive representations in video prediction by predicting latent representations between frames rather than pixel-level reconstruction, capturing high-level semantic patterns. Yann LeCun's proposed JEPA (Joint Embedding Predictive Architecture) framework explicitly centers on "predictive perception," seeking to replace traditional generative or discriminative paradigms—its design logic holds that rather than predicting raw pixels (computationally expensive and full of redundancy), it's better to predict semantic-level changes in abstract representation space. This closely echoes the brain's predictive coding logic of "only processing errors" and is seen as an important path toward brain-inspired visual AI. The success of OpenAI's Sora and other video generation models also confirms the viability of the world model approach from another angle: models capable of predicting world dynamics often simultaneously possess stronger perceptual understanding capabilities.
The Dual Advantages of Energy Efficiency and Robustness
The human brain consumes only about 20 watts of power yet accomplishes visual cognitive tasks far exceeding those of today's supercomputers. This remarkable energy efficiency is not the product of a single mechanism but the result of multi-level coordinated optimization. The "only process errors" mechanism from predictive coding is one important source; additionally, the analog characteristics of synaptic transmission avoid the precision waste of digital computation; myelinated axons enable high-speed, low-loss long-distance signal transmission; local cortical circuits achieve automatic gain control through excitatory-inhibitory balance (E/I balance), preventing signal over-amplification or attenuation. These mechanisms together constitute a system-level energy efficiency optimization framework far exceeding any single AI chip design.
But the brain's energy-saving design goes further—biological neurons communicate using discrete "spikes," firing action potentials only when needed, inherently possessing sparsity and event-driven characteristics. Inspired by this, Spiking Neural Networks (SNNs) are viewed as an important path toward brain-inspired computing.
Extended Background: Current Status and Challenges of Spiking Neural Networks and Neuromorphic Computing Neuromorphic computing hardware such as Intel's Loihi chip and IBM's TrueNorth are designed based on spiking neuron principles and can achieve power consumption orders of magnitude lower than traditional GPUs on certain sparse inference tasks. Intel's second-generation Loihi 2 chip already supports on-chip online learning, further enhancing its ability to adapt to dynamic environments. However, widespread application of SNNs still faces core challenges: standard backpropagation algorithms cannot be directly applied to discrete, non-differentiable spike signals, and researchers are making breakthroughs through surrogate gradient methods and temporal coding strategies. Combining the algorithmic logic of predictive coding with the hardware characteristics of SNNs is considered a frontier direction for achieving truly brain-inspired low-power AI systems—predictive coding provides algorithmic sparsity through "only computing differences," while SNNs provide hardware sparsity through "only computing when events trigger." The two are complementary, offering tremendous practical value for edge AI and embodied intelligence that pursues low power and high robustness, and represent a frontier intersection point of neuromorphic computing currently receiving attention from both academia and industry.
Rethinking What It Means to "See"
The topic of "circuits that let the brain think and see simultaneously" reminds us to reexamine a seemingly ordinary ability—seeing. Vision has never been a passive mirror reflection but rather an active construction process that integrates observation, memory, expectation, and inference. The never-ending bidirectional dialogue between feedforward and feedback in the visual cortex is precisely the physiological foundation that allows us to both faithfully perceive and flexibly understand the world.
For the future of artificial intelligence, this is not merely a fascinating scientific story but a precious blueprint from nature's long evolutionary history. Understanding how the brain "sees while thinking" may be the key to unlocking more intelligent, more efficient machine vision systems.
Key Takeaways
- Visual processing is not a one-way pipeline: The feedforward pathway from LGN to V1 to higher cortical areas is only half the story. Feedback connections often outnumber feedforward ones, with their signals arriving approximately 50-100 milliseconds after the feedforward wave, constituting a measurable, continuous bidirectional dialogue.
- Predictive coding is the core mechanism for the fusion of perception and cognition: The brain actively constructs perception by continuously generating predictions and computing errors rather than passively receiving signals; this is highly consistent with differential coding in information theory and is a key algorithmic source of the brain's extreme energy efficiency.
- Visual illusions and pathological phenomena provide behavioral and clinical validation of this mechanism: The hollow mask illusion, Kanizsa triangles, and abnormal perception in schizophrenia patients all point to the same predictive modulation system.
- The Free Energy Principle provides a unified theoretical framework: From perceptual updating to Active Inference, Friston's framework integrates perception, action, and learning into a single objective of minimizing prediction error, providing neuroscientific justification for integrated perception-decision design in embodied intelligence.
- Implications for AI architecture are concrete and far-reaching: World Models, the JEPA framework, and Spiking Neural Networks echo the brain's predictive coding design logic from the perspectives of algorithms, representations, and hardware respectively; iterative refinement temporal mechanisms offer a new paradigm for transcending single-pass feedforward inference.
- Clinical dissociation of the dual-stream model validates visual modularity: The damage patterns of the ventral pathway (What) and dorsal pathway (Where/How) are clearly corroborated in neurological cases, providing biological reference for perception-action decoupled design in AI visual systems.
- Brain energy efficiency comes from multi-level system-wide coordination: The algorithmic sparsity of predictive coding, event-driven hardware sparsity of SNNs, automatic gain control through E/I balance, and low-loss myelinated transmission together constitute an energy efficiency optimization system far exceeding any single mechanism.
- Neuromodulator systems reveal the multi-dimensional nature of attention: The coordinated modulation of acetylcholine, norepinephrine, and dopamine makes biological attention far exceed Transformers' static weight allocation, pointing toward the developmental direction of introducing dynamic state modulation mechanisms in future AI systems.
Related articles

Project Rai-chan Technical Breakdown: A Complete Guide to Building a Fully Local AI Companion
Deep dive into Project Rai-chan's tech stack: Ollama+Gemma local LLM, Unity rendering, VOICEVOX speech synthesis, and more — exploring the technical path for local AI companions.

Ollama vs OpenCode: Real-World GLM Quota Difference Analysis (380 vs 880 Requests)
Real-world comparison of Ollama vs OpenCode GLM5.2 quota consumption — from 380 to 880 requests. Analyzing context length, billing differences, and key factors for AI coding tool users.

U.S. Ban Targets Chinese Humanoid Robots: A New Front in the AI Race
Analysis of the U.S. ban on Chinese humanoid robots: data security concerns, industrial protection motives, and how the AI race extends into Physical AI and robotics hardware.