TL;DR: A decoding method based on quantitative receptive-field models that characterize the relationship between visual stimuli and fMRI activity in early visual areas is developed and it is suggested that it may soon be possible to reconstruct a picture of a person’s visual experience from measurements of brain activity alone.
Abstract: Recent functional magnetic resonance imaging (fMRI) studies have shown that, based on patterns of activity evoked by different categories of visual images, it is possible to deduce simple features in the visual scene, or to which category it belongs. Kay et al. take this approach a tantalizing step further. Their newly developed decoding method, based on quantitative receptive field models that characterize the relationship between visual stimuli and fMRI activity in early visual areas, can identify with high accuracy which specific natural image an observer saw, even for an image chosen at random from 1,000 distinct images. This prompts the thought that it may soon be possible to decode subjective perceptual experiences such as visual imagery and dreams, an idea previously restricted to the realm of science fiction. Recent functional magnetic resonance imaging (fMRI) studies have shown that it is possible to deduce simple features in the visual scene or to which category it belongs. A decoding method based on quantitative receptive field models that characterize the relationship between visual stimuli and fMRI activity in early visual areas has now been developed. These models make it possible to identify, out of a large set of completely novel complex images, which specific image was seen by an observer. A challenging goal in neuroscience is to be able to read out, or decode, mental content from brain activity. Recent functional magnetic resonance imaging (fMRI) studies have decoded orientation1,2, position3 and object category4,5 from activity in visual cortex. However, these studies typically used relatively simple stimuli (for example, gratings) or images drawn from fixed categories (for example, faces, houses), and decoding was based on previous measurements of brain activity evoked by those same stimuli or categories. To overcome these limitations, here we develop a decoding method based on quantitative receptive-field models that characterize the relationship between visual stimuli and fMRI activity in early visual areas. These models describe the tuning of individual voxels for space, orientation and spatial frequency, and are estimated directly from responses evoked by natural images. We show that these receptive-field models make it possible to identify, from a large set of completely novel natural images, which specific image was seen by an observer. Identification is not a mere consequence of the retinotopic organization of visual areas; simpler receptive-field models that describe only spatial tuning yield much poorer identification performance. Our results suggest that it may soon be possible to reconstruct a picture of a person’s visual experience from measurements of brain activity alone.
TL;DR: It is suggested that the parieto-occipital alpha power reflects functional inhibition imposed by higher level areas, which serves to modulate the gain of the visual stream.
Abstract: Although the resting and baseline states of the human electroencephalogram and magnetoencephalogram (MEG) are dominated by oscillations in the alpha band (∼10 Hz), the functional role of these oscillations remains unclear. In this study we used MEG to investigate how spontaneous oscillations in humans presented before visual stimuli modulate visual perception. Subjects had to report if there was a subtle difference in gray levels between two superimposed presented discs. We then compared the prestimulus brain activity for correctly (hits) versus incorrectly (misses) identified stimuli. We found that visual discrimination ability decreased with an increase in prestimulus alpha power. Given that reaction times did not vary systematically with prestimulus alpha power changes in vigilance are not likely to explain the change in discrimination ability. Source reconstruction using spatial filters allowed us to identify the brain areas accounting for this effect. The dominant sources modulating visual perception were localized around the parieto-occipital sulcus. We suggest that the parieto-occipital alpha power reflects functional inhibition imposed by higher level areas, which serves to modulate the gain of the visual stream.
TL;DR: The data directly link momentary levels of posterior alpha-band activity to distinct states of visual cortex excitability, and suggest that their spontaneous fluctuation constitutes a visual operation mode that is activated automatically even without retinal input.
Abstract: Neural activity fluctuates dynamically with time, and these changes have been reported to be of behavioral significance, despite occurring spontaneously. Through electroencephalography (EEG), fluctuations in alpha-band (8-14 Hz) activity have been identified over posterior sites that covary on a trial-by-trial basis with whether an upcoming visual stimulus will be detected or not. These fluctuations are thought to index the momentary state of visual cortex excitability. Here, we tested this hypothesis by directly exciting human visual cortex via transcranial magnetic stimulation (TMS) to induce illusory visual percepts (phosphenes) in blindfolded participants, while simultaneously recording EEG. We found that identical TMS-stimuli evoked a percept (P-yes) or not (P-no) depending on prestimulus alpha-activity. Low prestimulus alpha-band power resulted in TMS reliably inducing phosphenes (P-yes trials), whereas high prestimulus alpha-values led the same TMS-stimuli failing to evoke a visual percept (P-no trials). Additional analyses indicated that the perceptually relevant fluctuations in alpha-activity/visual cortex excitability were spatially specific and occurred on a subsecond time scale in a recurrent pattern. Our data directly link momentary levels of posterior alpha-band activity to distinct states of visual cortex excitability, and suggest that their spontaneous fluctuation constitutes a visual operation mode that is activated automatically even without retinal input.
TL;DR: The same principles that govern visual perception can explain many seemingly disparate auditory phenomena, and similarity suggests that the same neural mechanisms control attention and influence perception across different sensory modalities.
TL;DR: It is argued that brain images are influential because they provide a physical basis for abstract cognitive processes, appealing to people's affinity for reductionistic explanations of cognitive phenomena.
TL;DR: This review begins with a general discussion of the natural scene statistics approach, of the different kinds of statistics that can be measured, and of some existing measurement techniques, followed by a summary of thenatural scene statistics measured over the past 20 years.
Abstract: The environments in which we live and the tasks we must perform to survive and reproduce have shaped the design of our perceptual systems through evolution and experience. Therefore, direct measurement of the statistical regularities in natural environments (scenes) has great potential value for advancing our understanding of visual perception. This review begins with a general discussion of the natural scene statistics approach, of the different kinds of statistics that can be measured, and of some existing measurement techniques. This is followed by a summary of the natural scene statistics measured over the past 20 years. Finally, there is a summary of the hypotheses, models, and experiments that have emerged from the analysis of natural scene statistics.
TL;DR: A major goal now is to determine how axon guidance cues and a growing list of other molecules cooperate with spontaneous and visually evoked activity to give rise to the circuits underlying precise receptive field tuning and orderly visual maps.
Abstract: Patterns of synaptic connections in the visual system are remarkably precise. These connections dictate the receptive field properties of individual visual neurons and ultimately determine the quality of visual perception. Spontaneous neural activity is necessary for the development of various receptive field properties and visual feature maps. In recent years, attention has shifted to understanding the mechanisms by which spontaneous activity in the developing retina, lateral geniculate nucleus, and visual cortex instruct the axonal and dendritic refinements that give rise to orderly connections in the visual system. Axon guidance cues and a growing list of other molecules, including immune system factors, have also recently been implicated in visual circuit wiring. A major goal now is to determine how these molecules cooperate with spontaneous and visually evoked activity to give rise to the circuits underlying precise receptive field tuning and orderly visual maps.
TL;DR: This review considers the substantial advances in understanding the neuronal mechanisms underlying this visual stability derived primarily from neuronal recording and inactivation studies in the monkey, an excellent model for systems in the human brain.
TL;DR: By monitoring eye movements, it is demonstrated that characteristic fixation patterns previously thought to be determined solely by the facial expression are systematically modulated by emotional context already at very early stages of visual processing, even by the first time the face is fixated.
Abstract: Current theories of emotion perception posit that basic facial expressions signal categorically discrete emotions or affective dimensions of valence and arousal. In both cases, the information is thought to be directly "read out" from the face in a way that is largely immune to context. In contrast, the three studies reported here demonstrated that identical facial configurations convey strikingly different emotions and dimensional values depending on the affective context in which they are embedded. This effect is modulated by the similarity between the target facial expression and the facial expression typically associated with the context. Moreover, by monitoring eye movements, we demonstrated that characteristic fixation patterns previously thought to be determined solely by the facial expression are systematically modulated by emotional context already at very early stages of visual processing, even by the first time the face is fixated. Our results indicate that the perception of basic facial expressions is not context invariant and can be categorically altered by context at early perceptual levels.
TL;DR: By taking Thierry et al.'s study as an exemplar case of what should not be done in ERP research of visual categorization processes, clarifications are provided on a number of methodological and theoretical issues about the N170 and its largest amplitude to faces.
TL;DR: It is proposed that the multisensory perception of flavor may be indicative of the fact that the taxonomy currently used to define the authors' senses is simply not appropriate.
TL;DR: This work combined magnetoencephalography in a spatially cued motion discrimination task with source-reconstruction techniques and characterized attentional effects on neuronal synchronization across key stages of the human dorsal visual pathway to suggest that attentional selection is mediated by frequency-specific synchronization between prefrontal, parietal, and early visual cortex.
TL;DR: An epistemological approach to rivalry is taken that considers the brain as engaged in probabilistic unconscious perceptual inference about the causes of its sensory input, which seems capable of explaining binocular rivalry and reconciling many findings.
TL;DR: In this article, a review of the natural scene statistics approach, different kinds of statistics that can be measured, and some existing measurement techniques is presented, as well as a summary of the hypotheses, models, and experiments that have emerged from the analysis of Natural scene statistics.
Abstract: The environments in which we live and the tasks we must perform to survive and reproduce have shaped the design of our perceptual systems through evolution and experience. Therefore, direct measurement of the statistical regularities in natural environments (scenes) has great potential value for advancing our understanding of visual perception. This review begins with a general discussion of the natural scene statistics approach, of the different kinds of statistics that can be measured, and of some existing measurement techniques. This is followed by a summary of the natural scene statistics measured over the past 20 years. Finally, there is a summary of the hypotheses, models, and experiments that have emerged from the analysis of natural scene statistics.
TL;DR: In this paper, a simple auditory pip is used to increase search times for a synchronized visual object that is normally very difficult to find, even though the pip contains no information on the location or identity of the visual object.
Abstract: Searching for an object within a cluttered, continuously changing environment can be a very time-consuming process. The authors show that a simple auditory pip drastically decreases search times for a synchronized visual object that is normally very difficult to find. This effect occurs even though the pip contains no information on the location or identity of the visual object. The experiments also show that the effect is not due to general alerting (because it does not occur with visual cues), nor is it due to top-down cuing of the visual change (because it still occurs when the pip is synchronized with distractors on the majority of trials). Instead, we propose that the temporal information of the auditory signal is integrated with the visual signal, generating a relatively salient emergent feature that automatically draws attention. Phenomenally, the synchronous pip makes the visual object pop out from its complex environment, providing a direct demonstration of spatially nonspecific sounds affecting competition in spatial visual processing. Keywords: attention, visual search, multisensory integration, audition, vision
TL;DR: The basic science presentation and the breakout group discussion on the topic of perception from the first CNTRICS meeting, held in Bethesda, Maryland on February 26 and 27, 2007 are described.
TL;DR: The confederacy of recently discovered illusions points to the underlying neural mechanisms of time perception, which is surprisingly prone to measurable distortions and illusions.
TL;DR: It is concluded that with sufficient training time an auditory BCI may be as efficient as a visual BCI and Mood and motivation play a role in learning to use a BCI.
TL;DR: The results suggest that subjective visual experience is shaped by the cumulative contribution of two processes operating independently at the neural level, one reflecting visual awareness per se and the other reflecting spatial attention.
Abstract: To what extent does what we consciously see depend on where we attend to? Psychologists have long stressed the tight relationship between visual awareness and spatial attention at the behavioral level. However, the amount of overlap between their neural correlates remains a matter of debate. We recorded magnetoencephalographic signals while human subjects attended toward or away from faint stimuli that were reported as consciously seen only half of the time. Visually identical stimuli could thus be attended or not and consciously seen or not. Although attended stimuli were consciously seen slightly more often than unattended ones, the factorial analysis of stimulus-induced oscillatory brain activity revealed distinct and independent neural correlates of visual awareness and spatial attention at different frequencies in the gamma range (30-150 Hz). Whether attended or not, consciously seen stimuli induced increased mid-frequency gamma-band activity over the contralateral visual cortex, whereas spatial attention modulated high-frequency gamma-band activity in response to both consciously seen and unseen stimuli. A parametric analysis of the data at the single-trial level confirmed that the awareness-related mid-frequency activity drove the seen-unseen decision but also revealed a small influence of the attention-related high-frequency activity on the decision. These results suggest that subjective visual experience is shaped by the cumulative contribution of two processes operating independently at the neural level, one reflecting visual awareness per se and the other reflecting spatial attention.
TL;DR: Assessment of the receiver operating characteristics (ROC) of participants in a visual-working-memory change-detection task yielded evidence highly consistent with a discrete fixed-capacity model of working memory for this task.
Abstract: Visual working memory is often modeled as having a fixed number of slots. We test this model by assessing the receiver operating characteristics (ROC) of participants in a visual-working-memory change-detection task. ROC plots yielded straight lines with a slope of 1.0, a tell-tale characteristic of all-or-none mnemonic representations. Formal model assessment yielded evidence highly consistent with a discrete fixed-capacity model of working memory for this task.
TL;DR: It is argued that because of the STS' role in interpreting social and speech input, impairments in STS function may underlie many of the social and language abnormalities seen in autism.
TL;DR: It is indicated that attention allocation during event perception is not affected by the perceiver's native language; effects of language arise only when linguistic forms are recruited to achieve the task, such as when committing facts to memory.
TL;DR: Results of three experiments on multisensory perception of emotions using newly validated sets of dynamic visual and non-linguistic vocal clips of affect expressions indicate that the perception of emotion expressions is a robust multisENSory situation which follows rules that have been previously observed in other perceptual domains.
TL;DR: Many of the temporal changes at the perceptual and the neural levels can be captured by the multifaceted and somewhat ambiguous concept of coarse-to-fine processing, although it is clear that not all temporal changes can be characterized this way.
TL;DR: Given the very sparse and abstract representation of visual information by these neurons, they could in principle be considered as 'grandmother cells', but several arguments are given that make such an extreme interpretation unlikely.
TL;DR: This paper used functional magnetic resonance imaging (fMRI) to measure brain activity in humans judging simple visual stimuli, and found that posterior and anterior temporal lobe regions were more active during A/approximately A than A/B decisions, suggesting multiple representations of prior expectations within visual hierarchy.
TL;DR: Results not only show that visual information about speech articulation enhances phoneme discrimination, but also that it may contribute to the learning of phoneme boundaries in infancy, and play a role in phonetic category learning.
TL;DR: In this selective review, a number of ways in which seeing the talker affects auditory perception of speech are outlined, including, but not confined to, the McGurk effect.
Abstract: In this selective review, I outline a number of ways in which seeing the talker affects auditory perception of speech, including, but not confined to, the McGurk effect. To date, studies suggest that all linguistic levels are susceptible to visual influence, and that two main modes of processing can be described: a complementary mode, whereby vision provides information more efficiently than hearing for some under-specified parts of the speech stream, and a correlated mode, whereby vision partially duplicates information about dynamic articulatory patterning.
Cortical correlates of seen speech suggest that at the neurological as well as the perceptual level, auditory processing of speech is affected by vision, so that ‘auditory speech regions’ are activated by seen speech. The processing of natural speech, whether it is heard, seen or heard and seen, activates the perisylvian language regions (left>right). It is highly probable that activation occurs in a specific order. First, superior temporal, then inferior parietal and finally inferior frontal regions (left>right) are activated. There is some differentiation of the visual input stream to the core perisylvian language system, suggesting that complementary seen speech information makes special use of the visual ventral processing stream, while for correlated visual speech, the dorsal processing stream, which is sensitive to visual movement, may be relatively more involved.
TL;DR: This work investigates brain areas whose activity during passive viewing of dance stimuli was related to later, independent aesthetic evaluation of the same stimuli, suggesting a possible role of visual and sensorimotor brain areas in an automatic aesthetic response to dance.
TL;DR: Unsupervised temporal slowness learning (UTL) was substantial, increased with experience, and was significant in single IT neurons after just 1 hour, suggesting that UTL may reflect the mechanism by which the visual stream builds and maintains tolerant object representations.
Abstract: Object recognition is challenging because each object produces myriad retinal images. Responses of neurons from the inferior temporal cortex (IT) are selective to different objects, yet tolerant ("invariant") to changes in object position, scale, and pose. How does the brain construct this neuronal tolerance? We report a form of neuronal learning that suggests the underlying solution. Targeted alteration of the natural temporal contiguity of visual experience caused specific changes in IT position tolerance. This unsupervised temporal slowness learning (UTL) was substantial, increased with experience, and was significant in single IT neurons after just 1 hour. Together with previous theoretical work and human object perception experiments, we speculate that UTL may reflect the mechanism by which the visual stream builds and maintains tolerant object representations.