The What and Why of Binding
Generate an AI Snapshot to get a quick, structured summary of this paper.
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
In attempts to formulate a computational understanding of brain function, one of the fundamental concerns is the data structure by which the brain represents information. For many decades, a conceptual framework has dominated the thinking of both brain modelers and neurobiologists. That framework is referred to here as "classical neural networks." It is well supported by experimental data, although it may be incomplete. A characterization of this framework will be offered in the next section. Difficulties in modeling important functional aspects of the brain on the basis of classical neural networks alone have led to the recognition that another, general mechanism must be invoked to explain brain function. That mechanism I call "binding." Binding by neural signal synchrony had been mentioned several times in the literature (Legéndy, 1970; 26Bair W. Koch C. Temporal precision of spike trains in extrastriate cortex of the behaving monkey.Neural Comput. 1996; 8: 44-66Crossref Google Scholar) before it was fully formulated as a general phenomenon (38Baylis G.C. Rolls E.T. Leonard C.M. Selectivity between faces in the responses of a population of neurons in the cortex in the superior temporal sulcus of the monkey.Brain Res. 1985; 342: 91-102Crossref PubMed Google Scholar). Although experimental evidence for neural synchrony was soon found, the idea was largely ignored for many years. Only recently has it become a topic of animated discussion. In what follows, I will summarize the nature and the roots of the idea of binding, especially of temporal binding, and will discuss some of the objections raised against it. Classical neural networks were developed as models of brain function. In developing these models, several questions needed to be addressed: (1) How are brain states to be interpreted as representations of actual situations? In other words, how is neural activity interpreted as a neural code, or, in computer parlance, as a data structure? (2) What is the nature of the mechanisms by which brain states are organized? (3) In what format is information laid down permanently in the brain? (4) How is memory laid down? In other words, what are the mechanisms of learning? The questions remain open, but, as we shall see, plausible answers have been offered by classical neural networks. Interestingly, in the context of the field of Artificial Intelligence, no general answers to these questions are provided, in the conviction that specific problems need specific data structures and specific algorithms (or, in our parlance, mechanisms of organization). Neuroscience and neural modeling, on the other hand, have the ambition to find general answers. It is this commitment to generality that results in the binding problem being a fundamental feature of the neural code. Before discussing the answers to the above questions that are postulated by classical neural networks, it is important to introduce an important parameter. Although not often made explicit, it is important to fix a temporal scale T, which we will refer to as the "psychological moment" (in the sense of "short period"). At times shorter than T, one speaks of mental state or brain state, whereas at times greater than T one sees a succession of states or a "state history." Whereas state history is subject to conscious scrutiny (that is, it potentially gets reflected in all modalities—memory, language, etc.), no such conscious analysis is possible below T. State history is ignored by most or all models, and the conceptual disagreement about the binding issue focuses exclusively on the definition of state, that is, on times below T. It is difficult to pin down a definite value for T, but a plausible range may be from 50 to 200 msec. Regardless of the exact value, it is important to realize that the arena for the discussion of the binding problem is at a time scale less than T. Classical neural networks are described by a deeply engrained set of concepts, often attributed to 14Alonso J.-M. Martinez L.M. Functional connectivity between simple cells and complex cells in cat striate cortex.Nat. Neurosci. 1998; 1: 395-403Crossref PubMed Google Scholar and 13Albright T.D. Stoner G.R. Visual motion perception.Proc. Natl. Acad. Sci. USA. 1995; 92: 2433-2440Crossref PubMed Scopus (76) Google Scholar but in reality much older, which give definite answers to the questions 1–4 above. (A) The neural code: neurons are taken as concrete symbols, as semantic atoms. They can be interpreted in relation to patterns and events external to the organism. Neurophysiology has provided a solid experimental basis for this statement, although some extrapolation is needed to extend it to all neurons in the brain. The interpretation of neurons as semantic atoms is generally accepted and is not the matter of much dispute. However, the following addition will have to be a focus of our discussion: (A′) A neuron has only one degree of freedom for a given psychological moment: it is either on or off (or it is on to a certain degree). Thus, the brain state is described by a list—or vector—of neural activities. In order to know what the brain is about in a given psychological moment, it is only necessary to know this vector, as well as, of course, a description of the symbolic meanings of all neurons. This interpretation of brain activity deliberately ignores the fact that actually recorded neural signals are not constant in any sense over T—or over any fixed time scale, for that matter. It is rather maintained that the observed microstate history is inconsequential for the function of the brain. (B) The mechanism of organization of brain states is based on the fluxes of excitation and inhibition, a neuron collecting incoming signals and firing when a threshold is surpassed (see, for instance, the switching rule of 25Bair W. Koch C. Precision and reliability of neocortical spike trains in the behaving monkey.in: Kluwer B.J. The Neurobiology of Computation. Academic Publishers, New York1995Google Scholar). The dynamics of the system is regulated so that it stabilizes activity within the psychological moment. (In associative memory models, this is, for instance, achieved by requiring connections between any pair of neurons to be symmetric, with the consequence that the system displays attractor dynamics.) Without this restriction, a McCulloch and Pitts system would be a general digital machine without any inherent tendency to organize, kept on track only by the force majeure of a programmer with detailed insight into the switching process. (C) Long-term memory is stored in terms of synaptic weights. (D) Long-term memory is laid down by mechanisms of synaptic plasticity, based on the statistics of neural signals, especially their temporal correlations. These postulates will be referred to as the framework for "classical neural networks." It forms the conceptual basis for a large and important part of current neurosciences, especially for the genre of brain modeling usually referred to as Neural Networks or Connectionism. Since existing neural models cover only a small range of the brain's functional repertoire, an all-important issue is whether the above framework constitutes an adequate basis from which to conquer the rest solely by constructing appropriate specific wiring diagrams and control parameters. It has been argued that the classical code of neural networks is very poor, too narrow in its possibilities to serve as a basis for an expansion of the functional range of current brain models (40Beck J. Perceptual grouping produced by changes in orientation and shape.Science. 1966; 154: 538-540Crossref PubMed Google Scholar; Fodor and Pylyshin, 1988). The underlying weakness is best illustrated by a classic example due to Frank 31Ballard D.H. Hinton G.E. Sejnowski T.J. Parallel visual computation.Nature. 1983; 306: 21-26Crossref PubMed Scopus (49) Google Scholar: imagine a specific neural network for visual recognition, which is internally structured such that it can derive four propositions and represent them by output neurons. Two neurons recognize objects, a triangle or a square, both generalizing over position. The other two indicate the position of objects in the image: in the upper half or in the lower half, both generalizing over the nature of the object (see Figure 1). When showing single objects to the network it responds adequately, e.g., with (triangle, top) or (square, bottom). A problem arises, however, when two objects are present simultaneously. If the output reads (triangle, square, top, bottom) it is not clear whether the triangle or the square is in the upper position. This is the binding problem: the neural data structure does not provide for a means of binding the proposition top to the proposition triangle, or bottom to square, if that is the correct description. In a typographical system, this could easily be done by rearranging symbols and adding brackets: [(triangle, top), (square, bottom)]. The problem with the code of classical neural networks is that it provides neither for the equivalent of brackets nor for the rearrangement of symbols. This is a fundamental problem with the classical neural network code: it has no flexible means of constructing higher-level symbols by combining more elementary symbols. The difficulty is that simply coactivating the elementary symbols leads to binding ambiguity when more than one composite symbol is to be expressed. Let's assume it was vital for an organism to trigger some action in response to a triangle if it was in an upper position but not in a lower one. The reaction would then have to be tied to the coincidence of activity in cells (triangle) and (top), which, however, would also occur if the triangle were at the bottom and a square at the top. The animal therefore would respond to a so-called false conjunction, perhaps with grave consequences. An analogous situation occurs in the brain. Correspondence between object type and position is explicit on the retinal level. Its loss on the way to the output of the circuit is due to the generalization that is taking place within the circuit: for instance, in the brain's "what" and "where" pathways in the temporal and parietal pathways of primate cortex. This problem is a general one, with implications far beyond that of vision. Imagine a mental object that is represented by the set P of neurons (refer to Figure 2) and another mental object represented by set Q (possibly overlapping with P, but that is not a point here). Now it becomes important to activate both objects in the same mental operation (when, for instance, comparing them). What would be more natural than to coactivate both sets in the same brain state? Such coactivation, however, leads to what we call the "superposition catastrophe": the two sets will merge into one, and the neural code will not express the information needed to subdivide the composite state into its components (see Figure 2). Rosenblatt's problem has a simple solution in terms of combination-coding cells. It would suffice if there existed a neuron that reacted to a triangle in the top position (or to a set P or Q of neurons). This could be realized with the help of connections from the lower, elemental levels, on which generalization has not yet taken place. However, a problem arises where appropriate combination-coding cells do not exist or cannot exist (due to the impracticality of the large numbers required, or to the previous history of the system) and where there are connection patterns that could cause confusion. The assumption that combination coding cells are available when and where required is problematic when the system to be modeled is a general purpose device. Most symbol systems have means of combining elementary symbols into more complex ones, which can then be handled as units without danger of ambiguity and which have explicit structure on the basis of which they can be compared, recognized, decomposed, and further combined to build even higher structures. That the classical code of neural networks doesn't have such means is the root of the binding problem. It is a very curious proposition that the brain, the ultimate handler of symbol systems, shouldn't have a general mechanism for combining subsymbols. Other symbol systems (such as mathematics or natural languages) suggest that more complex binding patterns may be required than just grouping a number of elements into one block with no internal structure. The visual image of an extended object may need to be represented as an array of local features that are bound together in a way that expresses the topological neighborhood relationships within the figure or even a hierarchy of object parts (6Adelson E.H. Perceptual organization and the judgment of brightness.Science. 1993; 262: 2042-2044Crossref PubMed Google Scholar); similarly, the representation of natural language structures requires binding arrangements with hierarchical structure. Why not have a purely hierarchical system of combination-coding cells? There may be good reasons that certain combination-coding cells should not exist. The more complex the combination a cell represents, the more special the context to which it refers. Any experience gained in a particular context (and recorded in terms of changed synaptic connections) should be affixed to the most general description of the situation, so it can be exploited in other contexts. If, for instance, something was associated with the appearance of a triangle that just happened to be in an upper position, it would be inefficient to affix the association only to an upper-triangle cell, for then the experience would have to be repeated and relearned for all positions of the triangle. In a similar vein, the absolute location of a speck of ink on paper is of very little relevance; what counts is its relative location in a pattern. Our environment is complex because it is combinatorial: complex objects and situations are constructed by combining simpler elements. To try to represent this complexity in a noncombinatorial way, letting single cells stand for external objects of any complexity, appears to be a terribly inefficient strategy. Any combinatorial symbol system, however, needs a mechanism to bind elements into groups. Some evidence exists that suggests that the brain does not simply use combination-coding cells to process stimuli. Psychophysical studies show that under certain conditions the brain's binding mechanism may fail, and people will report "illusory conjunctions" (see Wolfe and Cave, 1999 [this issue of Neuron]). When given enough viewing time, subjects do not commit such "false conjunction" errors. In a related vein, conjunction search experiments (36Barlow H.B. The twelfth Bartlett memorial lecture the role of single neurons in the psychology of perception.Quart. J. Exp. Psychol. 1985; 37: 121-145Crossref Google Scholar), in which subjects are asked to find an object with a specific combination of features in an array of distractors with one of those features each, show that reaction times scale linearly with the total number of elements in the display. From these types of experiments, one can conclude that relevant combinations of features are not represented by combination-coding cells and that the brain requires time to form or ascertain the correct bindings. There is a widespread opinion that classical neural networks are a universal medium with no limits to their abilities and that consequently they are not subject to the binding problem. I will address this claim in two steps. First, I will discuss whether universality suffices as a solution to the brain's problems and then I will raise doubts as to whether classical neural networks are indeed a universal medium. The idea of universality was crystallized with Alan Turing's formulation of the Turing machine and his demonstration that no effective procedure can be conceived that cannot be realized as the program of a Turing machine. Thus, any completely specified function can be realized as a program, or algorithm, run on a computer, the only limits to this being storage space and time. From this, it was extrapolated that mental processes, if only made concrete in terms of rules, could be realizable in machines. Under this view, the brain is a digital machine and has the same universality as the computer or the Turing machine, if only sufficiently many neurons are available. 25Bair W. Koch C. Precision and reliability of neocortical spike trains in the behaving monkey.in: Kluwer B.J. The Neurobiology of Computation. Academic Publishers, New York1995Google Scholar applied this idea to modeling of the nervous system, proving that any logical function can be realized as a network of threshold elements if they are appropriately connected. But what does universality buy? It can be compared to the universality of a pen and sufficiently many sheets of white paper as a universal medium for formulating novels. You still have to write them. Over time, the field of Artificial Intelligence discovered that it is not a practical task at all to write a program that emulates the capabilities of the brain. It is becoming increasingly clear that the only goal we can hope for is to establish a system that constitutes a basis for self-organization and learning, as the equivalent of a newborn, or better still, the genetic program for the development of a brain, and to let it learn from experience and from communication with others. Brain theorists realized this in the late 1950s and modified McCulloch and Pitts' networks to accommodate self-organization and learning. The resulting framework of classical neural networks has a tendency to fall into stable patterns and learns by synaptic plasticity. However, these changes may have come at a price: it is not clear whether neural networks are universal in any sense, although the community seems to have inherited the implicit belief that they are and that any brain function can be modeled on the basis of those few abstractions from the real nervous system that went into the formulation of neural networks. It is not even clear how to formulate a new universality theorem. Classical universality states, "give me a procedure and I'll tell you how to implement it." The neural network version of universality would have to be, "give me a brain problem and I will be able to implement it in classical neural networks." But how can we characterize brain problems in any general and satisfactory way? It would be foolish to argue that "this is a particular problem I have solved on the basis of classical neural networks, which proves that all of them can be solved this way." In the context of the present discussion, the particular version of universality that some critics of the binding problem uphold, is "state a concrete problem, and I will solve it with a classical neural network without ever running into a binding ambiguity." The common argument to support this posits that any binding ambiguity implicit in a stated problem can be dealt with by the concrete combination-coding cells in the model network that solves it (for example, see Riesenhuber and Poggio, 1999a [this issue of Neuron]). A concrete example of this approach is Mel and Fiser's model (1999) for recognizing words from text. The network codes text in terms of triplets of contiguous letters and is based on the statistical observation that in English no two words agree in all their letter triplets. Thus, in this special case, the model completely avoids the binding problem and ensuing compositional ambiguity that would beriddle a model based on a representation by single-letter cells only. Although such examples are meant to support a general universality claim, it is very doubtful that such a claim can ever be established. It is too easy to state problems that are far beyond the abilities of present neural network models. Just think of the task, "emulate the human ability to segment visual scenes, with all the necessary cue integration." Although the problem is old, there exists no classical neural network solution to it, and perhaps for good reason. How could binding be implemented in the brain? The basic idea of temporal binding is that signals of neurons that are to be grouped together are correlated in time. Neural signals can thus be evaluated in two ways: one of them is the classical concept of neural firing rate, in which the relevant parameter is the running average of the number of spikes arriving within any period T. The second concerns temporal correlations of signal fluctuations happening on time scales faster than T, and it is these correlations that express binding. The subdivision of the time scales above and below T is on final account arbitrary (unless one sticks to the distinction that changes slower than T are accessible to introspection while faster changes are not). Pick a scale T, and then evaluate signal fluctuations below T in terms of correlations, while calling fluctuations above T rate changes (there is, of course, a lower bound, that depends on the fastest temporal scale that can be processed by neural tissue). For the brain, T is related to the psychological moment, ill-defined as it may be, so we will take T to be of the order of 50 to 200 msec (although scales up to minutes and beyond are also of potential relevance). Throughout this discussion, I discuss signal correlations as if they were to be evaluated without taking into account relative delays. However, it may be necessary to also consider delayed coincidences as argued in 8Adelson E.H. Movshon J.A. Phenomenal coherence of moving visual patterns.Nature. 1982; 300: 523-525Crossref PubMed Google Scholar. To understand the general idea of temporal binding, the exact nature of the signal fluctuations is not relevant. There has been much discussion over the experimental evidence concerning whether oscillatory signals are or are not important in this respect, but that discussion is a side issue that shouldn't cloud the more fundamental question of whether binding is a problem and whether neural signal correlations are a solution to that problem. Although much of the discussion of signal correlations focuses on the binary case, the correlation of just two neurons, it should be emphasized that the much more relevant and important type of event concerns correlations of higher order: the simultaneous firing of larger groups of neurons. The reason is this: for correlations to play an important role in the brain it must be possible for them to be evaluated quickly and reliably. Since for any given set of impulses there may be accidental coincidences, it is vitally important to discern true correlations from noisy background. For binary correlations, this is only possible for long observation times, but for correlations of sufficiently high order, even individual coincident events can become highly significant. With regard to the plausibility of temporal binding, there are the several fundamental questions that need to be discussed: (1) How do correlations in temporal signal structure arise? Ultimately, the purpose of temporal binding is to express significant relations between data items, e.g., of causal or spatial nature; the physical interactions establishing such relations must be represented by signal correlations. Since many of these interactions are already present in the external world, temporal correlations can be imposed by external stimuli triggering the neural signals. That this happens with causally related external events is evident, but it is less often realized that the same can be due to our own bodily and eye movements, which create a stream of sensory impressions whose temporal structure expresses spatial structure of the environment. Thus, some of the signal correlations relevant for binding are already implicit in the perceptual input (see also Singer, 1999b [this issue of Neuron]). As a short aside, the argument is often raised that the use of temporal patterns for expressing binding may clash with the use of temporal patterns for other purposes, such as the representation of temporal structure as given in the external world. This is especially relevant to the auditory, language, and motor modalities. This clash may be avoided by the nervous system by recoding temporal signals into a format that does not involve rapid signal changes. Single neurons responding to and representing syllables would be an example of this. More correlations (probably the overwhelming majority) are created within the nervous system by synaptic connections. If neuron a fired neuron b, the signals of the two would be correlated (disregarding a small delay). Correlations induced by synaptic connections also signify causality. In addition, activity in neurons without connections between them but with connections from a common input can be correlated. In the Rosenblatt example, the binding problem could be solved if the neurons in V1 that are activated by a triangle in a given position pass the temporal signature of their signals on to neurons expressing shape identity on the one hand and position on the other, such that those signals came to express their common origin. (2) How are signal correlations evaluated in the brain? If two action potentials arrive at a common target neuron, their relative timing exerts a strong influence. If signals arrive simultaneously, they can cooperate to raise the neuronal membrane potential above firing threshold. If, however, they miss each other in time, so that the effect of the first impulse has decayed before the second arrives, they might both fail to fire the target neuron. Thus, neurons act as coincidence detectors and do evaluate signal correlations (1Abeles M. Role of cortical neuron integrator or coincidence detector?.Israel J. Med. Sci. 1982; 18 (a): 83-92PubMed Google Scholar, 23Assad J. Maunsell J. Neuronal correlates of inferred motion in primate posterior parietal cortex.Nature. 1995; 373: 518-521Crossref PubMed Google Scholar). The exact details of this interaction depend upon many complex factors, including membrane time constants, nonlinear effects, and dendritic geometry. In consequence, current neurophysiology cannot solidly predict the temporal resolution at which spike coincidences are evaluated; however, a likely range is 1–10 msec. If all correlations were to be evaluated globally by single neurons, a combination-coding cell would be required for each binding pattern, defeating the purpose of binding. However, complex correlation patterns created by a circuit of interconnected neurons can be evaluated by other circuits of appropriately interconnected neurons, each individual neuron checking only a small subpattern. Thus, pairs of circuits may or may not resonate with each other in terms of the correlation patterns that they produce. This point is probably most easily understood with reference to the concrete models of invariant object recognition that are discussed below. It was proposed (38Baylis G.C. Rolls E.T. Leonard C.M. Selectivity between faces in the responses of a population of neurons in the cortex in the superior temporal sulcus of the monkey.Brain Res. 1985; 342: 91-102Crossref PubMed Google Scholar) that correlation patterns are also evaluated by rapid reversible synaptic plasticity (in addition to slow plasticity). A connection that is physically present and would cause confusion in a given situation could be temporarily inactivated when activity on both the presynaptic and postsynaptic side is sensed, but is uncorrelated. Confusion could thus be suppressed in the given situation, even if the signals involved were to develop stray coincidences, until the switched-off connection returned to near its previous value on the time scale of seconds or minutes. (3) How can correlation patterns be effective on physiological timescales? If they are to play a role in the brain's function, it is mandatory that they be evaluated within short time intervals. Finding a pair of correlated neurons in a set of others firing stochastically may take unrealistically long integration times. The situation can be improved in two ways. As previously mentioned, one way relies on coincidences of high order. Even when superimposed on stochastic signals, a single event of n simultaneous spikes, with n large enough (say, 50 or 100), can be of high statistical significance. The other way relies on the suppression of accidental correlations by appropriate inhibitory circuits. Thus, if a large set of neurons needs to be subdivided into several bound subsets, inhibition between the subsets can make sure that no coincident spikes between neurons in different subsets occur at all. This is an integral part of many models, e.g., 43Bergen J.R. Adelson E.H. Early vision and texture perception.Nature. 1988; 333: 363-364Crossref PubMed Google Scholar or 49Bienenstock E. A model of neocortex.Network. 1995; 6: 179-224Crossref Scopus (101) Google Scholar. (4) How are the network patterns created that are required for the production and evaluation of significant firing patterns? Random connection patterns will neither be able to create significant firing patterns nor be able to distinguish them. Many of the arguments raised against the validity of the idea of temporal binding, e.g., 34Barlow H.B. Single units and cognition a neurone doctrine for perceptual psychology.Perception. 1972; 1: 371-394Crossref PubMed Google Scholar and 33Barbas H. Pandya D.N. Architecture and intrinsic connections of the prefrontal cortex in the rhesus monkey.J. Comp. Neurol. 1989; 286: 353-375Crossref PubMed Google Scholar( [ this issue of Neuron]), are implicitly or explicitly based on the assumption of random connectivity. If, however, the nervous system is endowed with the capacity to self-organize using synaptic plasticity of slow (46Biederman I. On the
