登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Steve Jurvetson
@FutureJurvetson
Co-founder of Future Ventures and DFJ, supporting passionate founders to forge a better future. Early VC investor in Tesla, SpaceX, Planet, Commonwealth Fusion.
参加 March 2010
69 フォロー中    110.9K ファン
🎶 On to the next one… predicting the next token From when I first started studying them in the 80’s, I have been fascinated by the emergent homologies between artificial and biological neural networks. From the emergence of similar hierarchical layers in vision systems to critical periods of curriculum learning, we now have evidence that our brains and our LLMs predict the next noun in a spoken sentence before hearing it. TLDR; from the Nature study below: when listening to an audiobook wired to EEG and MEG scanners, the brain activates structures for the anticipated next noun or adjective before hearing it. The brain does a form of next-token prediction. They fed the same audiobook into a Llama LLM and found the artificial neural net was doing the same thing, and like the humans, it was much better with nouns than verbs. This would come to no surprise to Jeff Hawkins, pictured here at his neuroscience institute. From his books, On Intelligence followed by A Thousand Brains, Jeff bravely presents a framework for how the brain works to produce intelligence from cells organized into ~150 thousand cortical columns. His decades of self-funded dedication to studying how the brain works affords a possibly unique and unifying perspective. In his first book, Hawkins presents a memory-prediction framework for intelligence. The neurons in the neocortex provide a vast amount of memory that learns a model of the world. These models continuously make low-level predictions in parallel across all of our senses. We only notice them when a prediction is incorrect. Higher in the hierarchy, we make predictions at higher levels of abstraction (the crux of intelligence, creativity and all that we consider being human), but the structures are fundamentally the same. If that is not mind-bending enough, in his latest book, Jeff extends the memory framework to the construct of “reference frames”. Everything we perceive is a constructed reality, a cortical consensus from competing internal models resident in many cortical columns, the amalgam of 1000 brains. Those models are updated by data streaming from the senses. But our reality resides in the models. And in particular, sequence memory predicts the next note in a melody and “is also used for language. Recognizing a spoken work is like recognizing a short melody.” With that in mind, let’s jump to summary passages from the recent study: “The brain continuously generates predictions about incoming sensory inputs, including the next word during language comprehension” Large language “models rely almost entirely on predictive processes to generate language enabling them to approach - or even surpass - the Turing-Test with unprecedented proficiency. The human brain, too, is not a passive receiver, but an active ‘prediction machine’, constantly anticipating upcoming words and events” “If language indeed represents the fundamental ability required for the development of general intelligence, chain-of-thought reasoning and abstract cognition, and if grammar naturally emerges through language usage – thereby aligning brain mechanisms with patterns observed in deep neural networks – this raises the critical question: can LLMs trained solely on next-word prediction evolve into artificial general intelligence (AGI)?” From a prior study, “even a relatively simple neural network trained only on next-word prediction can spontaneously internalize basic grammatical structures. Hence, it seems plausible that the human brain, with its approximately 100 billion neurons, could accomplish this feat through continuous language-based prediction alone.” “We observed significant pre-onset activity for nouns… in a complementary analysis, we trained a linear probe neural network on the hidden states of Llama 3.2, revealing that nouns and adjectives are more easily predicted than verbs. We discuss our findings in the context of anticipatory frameworks in artificial neural networks, highlighting potential parallels between biological and computational approaches to language anticipation.” “In the EEG, nouns, adjectives, and proper nouns elicited a significant negative peak beginning before 0 seconds, whereas verbs showed no evidence of early anticipatory activity. The MEG data also revealed a significant peak around 0 seconds in left frontal channels for nouns, but not for the other word types, suggesting early predictive mechanisms specifically associated with this word class. [And with the LLM Llama] we found that the prediction probability for the next word is highest for the word class ‘noun’, and considerably lower for the other three word classes.” “We acknowledge that this represents only a preliminary step toward understanding syntactic and semantic processing by comparing LLMs and human brain activities. We present initial evidence suggesting that different transformer blocks may functionally correspond to distinct cortical regions, though further exploration is needed.” “Our analyses show that the predictive signal in the brain seems to consist of two components: a syntactic readiness, which can be observed in temporal regions, and a semantic readiness potential, mainly located in frontal areas. These convergent findings are consistent with the Bayesian framework of the brain, suggesting that both neural and computational systems continually update their internal models by integrating prior expectations with incoming information. Taken together, this evidence highlights the presence of distinct but complementary predictive mechanisms underlying semantic and syntactic processing in language comprehension.” “The question remains whether LLMs can serve as a valid model for understanding the human brain. Large language models are constructed as layered stacks of transformer blocks that operate via self-attention rather than explicit recurrent connections, yet their repetitive structure - characterized by self-similarity and fractal organization being a universal principle in biological structures - may serve as an analogue to the recurrent transmission of signals through the arcuate fasciculus. In the human brain, the arcuate fasciculus facilitates dynamic bidirectional communication between Broca’s and Wernicke’s areas, a pathway long proposed as the neural substrate for a universal innate grammar. Although LLMs do not replicate the full complexity of biological recurrence, the iterative processing achieved by stacking transformer blocks appears to approximate the brain’s integration of syntactic and semantic cues. These observations suggest that the self-similar structure of the LLM may provide valuable insights into the neural strategies underlying anticipatory activity and integrative processing during language comprehension.” “Our study demonstrates that prediction-related readiness in language processing operates through both syntactic and semantic anticipation, as reflected in distinct pre-word onset activity captured by combined MEG and EEG recordings.” “In addition, our results suggest that LLMs provide a computational framework that approximates human prediction-related readiness, with their stacked transformer blocks potentially mirroring the recurrent interactions between Broca’s and Wernicke’s areas via the arcuate fasciculus. The self-similar organization of these transformer architectures may reflect universal hierarchical principles of cognitive processing” “The fusion of generative AI and neural data has the potential to refine cognitive computational neuroscience (CCN) and provide deeper insights into the hierarchical organization of language processing in biological and artificial systems. Ultimately, bridging neuroscience, AI and linguistic theory may not only reveal the cognitive mechanisms that govern human language, but also drive the development of artificial intelligence - bringing it closer to the way the human brain anticipates and processes language” 🧠 -------- And I have transcribed my favorite passages from Hawkins most recent book, A Thousand Brains. I revisit them to learn. Travelling without moving, as we’ll see… “The cells in your head are reading these words. Think how remarkable that is.” “If you ignore folds and creases, then the neocortex looks like one large sheet of cells, with no obvious divisions. The neocortex looks similar everywhere. Every part of the neocortex generates movement. In every region we have examined, scientists have found cells that project to some part of the old brain related to movement. The complex circuitry seen everywhere in the neocortex performs a sensory-motor task. There are no pure motor regions and no pure sensory regions.” The cortex is relatively new development by evolutionary time scales. After a long period of simple reflexes and reptilian instincts, only mammals evolved a neocortex. “At some point millions of years ago, a new piece of the brain appears that we now call the neocortex. It starts small, but then grows larger, not by creating anything new, but by copying a basic circuit over and over. As the neocortex grows, it gets larger in area but not in thickness.” Given the recency, it’s “probably not enough time for multiple new complex capabilities to be discovered by evolution, but it’s plenty of time for evolution to make more copies of the same thing.” • Vernon Mountcastle’s proposition from 1978: “All the things we associate with intelligence, which on the surface appear to be different, are, in reality, manifestations of the same underlying cortical algorithm. Darwin proposed that the diversity of life is due to one basic algorithm (evolution). Mountcastle proposed that the diversity of intelligence is due to one basic algorithm.” Beyond the evolutionary time-scale argument, the brains’ vast flexibility to accept different, even prosthetic, sensory input changes and its ability to learn many different things point to a universal framework for learning. • Cortical Columns are “the largest and most important piece of the puzzle.” They are roughly one square millimeter in size with 100K neurons. A mouse has one column per whisker. “Every cortical column is making predictions. We are not aware of the vast majority of these predictions unless the input to the brain does not match.” • Learning through movement: “The brain learns its model of the world by observing how its inputs change over time. There isn’t another way to learn. Every time we take a step, move a limb, move our eyes, tilt our head, or utter a sound, the input from our sensors change. For example, our eyes make rapid movements, called saccades, about three times a second. With each saccade, our eyes fixate on a new point in the world and the information from the eyes to the brain changes completely.” We don’t perceive any of this because we are living in the model, which is predicting the next input to come, across all the senses. “Vision is an interactive process, dependent on movement. Only by moving can we learn a model of the object.” “To avoid hallucinating, the brain needs to keep its predictions separate from reality. We are not aware of most of the predictions made by the brain unless an error occurs.” “Thoughts and experiences are always the result of a set of neurons that are active at the same time (about 2% of the total). Individual neurons can participate in many different thoughts or experiences. Everything we know is stored in the connections between neurons. Every day, many of the synapses on an individual neuron will disappear and new ones will replace them. Thus, much of learning occurs by forming new connections between neurons that were not previously connected.” • Locus of Predictions: “Oddly, less than 10% of the pyramidal cell’s synapses are in the proximal area. The other 90% are too far away to trigger a spike. For many years, no one knew what 90% of the synapses in the neocortex did. The big insight I had was that dendrite spikes are predictions. A dendrite spike occurs when a set of synapses close to each other on a distal dendrite get input at the same time, and it means the neuron had recognized a pattern of activity in some other neurons. When the pattern of activity is detected, it raises the voltage at the cell body, putting the cell into what we call a predictive state. The cell is primed to spike… and the cell spikes a little bit sooner than if it would have if the neuron was not in a predictive state.” And this inhibits other neurons from ever firing, the ones who were behind in that race. “When an input arrives that is unexpected, then neurons fire at once. If the input is predicted, then only the predictive-state neurons become active. This is a common observation about the neocortex: unexpected inputs cause a lot more activity than expected ones.” Predictions prime the pump, sub-threshold. “Predictions are not sent along a cell’s axon to other neurons, which explains why we are unaware of most of them.” “Most predictions occur inside neurons. With thousands of distal synapses, each neuron can recognize hundreds of patterns that predict when the neuron should become active. Prediction is built into the fabric of the neocortex. As few as 20,000 neurons can learn thousands of complete sequences. The sequence memory continued to work even if 30% of the neurons died or the input was noisy.” • Reference Frames: “The secret of the cortical column is reference frames. A reference frame is like an invisible, 3D-grid surrounding and attached to something” (like a map) “Predicting the next input in a sequence and predicting the next input when we move are similar problems. Our sequence-memory circuit could make both types of predictions if the neurons were given an additional input that represented how the sensor was moving.” “Most of the circuitry is there to create reference frames and track locations. The brain builds models of the world by associating sensory input with locations in reference frames. You need a reference frame to specify the relative position and structure of objects. Roboticists rely on them to plan the movements of a robot’s arm or body. Reference frames were the missing ingredient, the key to unraveling the mystery of the neocortex and to understanding intelligence. We showed that a single cortical column could learn the 3D shape of objects by sensing and moving and sensing and moving. Each cortical column must know the location of its input relative to the object being sensed. To do that, a cortical column requires a reference frame that is fixed to the object. The brain must have neurons whose activity represents the location of every object that we perceive.” “Mammals have a powerful internal navigation system. There are neurons in the old part of our brain that are known to learn maps of the places we have visited” — the hippocampus and enthorhinal cortex, organs roughly the size of a finger. “Place cells tell a rat where it is based on sensory inputs, but planning movement requires grid cells. Grid cells form a grid pattern. The two types of cells work together to create a complete model of the rat’s environment. Every time a rat enters an environment, the grid cells create a new reference frame to specify locations and plan movements.” In the new brain, these same cells and structures create models of objects instead of environments. “Every cortical column learns models of complete objects. The columns do this using the same basic method that the old brain uses to learn models of environments. It is as if nature stripped down the hippocampus to a minimal form, made tens of thousands of copies, and arranged them side by side in cortical columns. That became the neocortex. Each patch of your skin and each patch of your retina has its own reference frame in the neocortex. Your five fingertips touching a cup are like five rats exploring a box.” “Not all cortical columns are modeling objects. Language and other high-level cognitive abilities are, at some fundamental level, the same as seeing, touching, and hearing. The reference frames that are most useful for certain concepts have more than three dimensions.” • Thinking is a form of movement: “The brain arranges all knowledge using reference frames, and thinking is a form of moving. Thinking occurs when we activate successive locations in reference frames.” “A cortical column is just a mechanism that tries to discover and model the structure of whatever is causing its inputs to change” whether the structure of environments, physical objects or conceptual objects. “Reference frames are not an optional component of intelligence; they are the structure in which all information is stored in the brain. Every fact you know is paired with a location in a reference frame. Organizing knowledge this way makes the facts actionable” to “determine what actions are needed to achieve a goal.” “To recall stored knowledge, we have to activate the appropriate locations in the appropriate reference frames. Thinking occurs when the neurons invoke location after location in a reference frame, bringing to mind what was stored in each location. The succession of thoughts we experience when thinking is analogous to the succession of sensations we experience when touching an object with a finger, or the succession of things we see when we walk about a town.” • What and Where Pathways. “Your brain has two vision systems. If you follow the optic nerve as it travels from the eye to the neocortex, you will see that it leads to two parallel vision systems, called the ‘what’ visual pathway and the ‘where’ visual pathway.” If you disable one, you can identify what something is but not where, or vice versa. “Similar pathways also exist for other senses. There are what and where regions for seeing, touching, and hearing.” “Cortical grid cells in What columns attach reference frames to objects. Cortical grid cells in Where columns attach reference frames to you body.” The distinction depends on where the inputs come from. “If a cortical column gets input from the body, such as the neurons that detect the joint angles of the limbs, it will automatically create a reference frame anchored to the body.” “Your body is just another object in the world. However, unlike external objects, your body is always present. A significant portion of the neocortex — the Where regions — is dedicated to modeling your body and the space around your body.” For abstract concepts like mathematics, there are difference reference frames one could use to learn. “Part of learning is discovering what is a good reference frame, including the number of dimensions.” History can be learned on a timeline, or geographically. “They lead to different ways of thinking about history. They might lead to different conclusions and different predictions. Becoming an expert in a field of study requires discovering a good framework to represent the associated data and facts. Discovering a useful reference frame is most difficult part of learning, even though most of the time we are not consciously aware of it.” It's no surprise that the memory trick called the memory palace, is a good method for remembering a large sequential list of nouns. From fMRI studies, “the process of storing items in a reference frame and recalling them via ‘movement’ is the same.” “Nested structure and recursion are key attributes of language. Each cortical column has to be able to learn nested and recursive structure. Cortical columns create reference frames for every object they know. Reference frames are then populated with links to other reference frames. The brain models the world using reference frames that are populated with reference frames; it’s reference frames all the way down.” • The Thousand Brains Theory of Intelligence: The prevailing view of the neocortex was a hierarchy of feature detectors, from edge detectors up to face detectors. Jeff argues that each and every column is a sensory-motor system. “When the eyes saccade from one fixation point to another, some of the neurons in the V1 and V2 visual regions do something remarkable. They seem to know what they will be seeing before the eyes have stopped moving. These neurons become active as if they can see new input, but the input hasn’t yet arrived. There are connections between low-level visual regions and low-level touch regions.” Mouse vision occurs in the V1 region; it does not depend on a hierarchy of vision abstractions. “All cortical columns, even in low-level sensory regions, are capable of learning and recognizing complete objects. A column that senses only a small part of an object (e.g., from a patch of retina) can learn a model of the entire object by integrating its inputs over time.” “Learning is not a separate process from sensing and acting. We learn continuously. When a neuron learns a new pattern, it forms new synapses on one dendrite branch. The new synapses don’t affect previously learned ones on other branches. Thus, learning doesn’t force the neuron to forget or modify something it learned earlier.” It’s additive. “What a column learns is limited by its inputs. Columns in V1 can recognize letters and words in the smallest font. V1 and V2 learn models of objects, such as letters and words, but the models differ by scale.” “Knowledge of something is distributed in thousands of columns, but these are a small subset of all the columns. This is why we call it the Thousand Brains Theory: knowledge of any particular item is distributed among thousands of complimentary models. The columns are not redundant, and each is a complete sensory-motor system.” • The Solution to Sensor Fusion and the Binding Problem: “Columns vote. Your perception is the consensus the columns reach by voting.” “If you touch something with only one finger, then you have to move it to recognize the object. But if you grasp the object with your entire hand, then you can usually recognize the object at once. In almost all cases, using five fingers will require less movement than using one.” (made me think of reading Braille with multiple fingers). “Voting works across sensory modalities (sight, touch, etc.)” How? “Cells in some layers send axons long distances within the neocortex” between left and right-hand brain regions or between V1 and A1, the primary vision and auditory regions. “These cells with long-distance connections are voting. Cells that represent what object is being sensed can vote and will project broadly. Often a column will be uncertain, in which case its neurons will send multiple possibilities at the same time. Simultaneously, the column receives projections from other columns representing their guesses. The most common guesses suppress the least common ones until the entire network settles on one answer. The voting mechanism works well even if the long-distance axons connect to a small, randomly chosen subset of other columns” • The Stability of Perception with ever-changing inputs: “What we perceive is based on the stable voting neurons. We are not consciously aware of the changing activity in each column.” Roughly 98% are silent at any given time and 2% are continuously firing. Consider the experience of an optical illusion duality (like the drawing of a pair of faces or vase); you can only see one at a time, and there is a delay if you force yourself to switch. “Recognizing an object in one sensory modality leads to predictions in other sensory modalities.” • Attention: We have the perception of multiple objects in our visual field even though we can only attend to one at a time. “Attention plays an essential role in how the brain learns models. The brain can attend to smaller or larger parts of the visual field. Exactly how the brain does this is not well understood, but it involves a part of the brain called the thalamus, which is tightly connected to all areas of the neocortex. It is so intimately connected to the neocortex that I consider it an extension of the neocortex.” • Consciousness: “Neurons form a continuous memory of both our thoughts and actions. It is this accessibility of the past — the ability to jump back in time and slide forward again to the present — that gives us our sense of presence and awareness. This is the core of what it means to be conscious. If we couldn’t replay our recent thoughts and experiences, then we would be unaware we are alive.” “The neocortex does not directly control any muscles. The neocortex has to be attached to something that already has sensors and already has behaviors (the primitive brain). It does not create completely new behaviors; it learns how to string together existing ones in new and useful ways.” “Reverse engineering the brain and understanding intelligence is the most important scientific quest humans will ever undertake.”
もっと見る