Journal Article10.5120/13238-0674
Computer Vision Architecture using Fusion Technique
Nidhi Srivastava,Harsh Dev +1 more
TL;DR: A novel computer vision architecture using fusion technique that combines or fuses more than one modality using multi-agents for effective and efficient automatic speech recognition.
read more
Abstract: Humans want to communicate with the computers in the same way as they communicate with other humans. Speech is the most natural and spontaneous form of communication. Speech is bimodal in nature and it combines audio and visual information to enhance speech recognition rate especially under poor audio conditions. This paper proposes novel computer vision architecture using fusion technique. This architecture combines or fuses more than one modality using multi-agents. In this we have used two modalities- audio and video. The audio part extracts the speech of a person and the video part extracts the face and lip information of the person. Here, different agents process the modalities and the fusion agent fuses these modalities for effective and efficient automatic speech recognition.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Dissertation
Une approche logicielle du traitement de la dyslexie : étude de modèles et applications
Geoffrey Garcia
- 07 Dec 2015
TL;DR: In this paper, the authors define a framework logiciel propice a l’implementation d'une plate-forme logicielle that nous avons appelee la PAMMA, a framework that devrait theoriquement pouvoir disposer de tous les outils permettant le developpement souple et efficace d'applications medicales integrant des processus metiers.
27
References
Recent advances in the automatic recognition of audiovisual speech
Gerasimos Potamianos,Chalapathy Neti,Guillaume Gravier,Ashutosh Garg,Andrew W. Senior +4 more
- 08 Sep 2003
TL;DR: The main components of audiovisual automatic speech recognition (ASR) are reviewed and novel contributions in two main areas are presented: first, the visual front-end design, based on a cascade of linear image transforms of an appropriate video region of interest, and subsequently, audiovISual speech integration.
Toward multimodal human-computer interface
Rajeev Sharma,Vladimir Pavlovic,Thomas S. Huang +2 more
- 01 May 1998
TL;DR: It is clear that further research is needed for interpreting and fitting multiple sensing modalities in the context of HCI and the fundamental issues in integrating them at various levels, from early signal level to intermediate feature level to late decision level.
363
Audio-visual speaker identification using coupled hidden Markov models
Tieyan Fu,Xiao Xing Liu,Luhong Liang,Xiaobo Pi,A.V. Nefian +4 more
- 31 Jul 2003
TL;DR: Experimental results on XM2VTS database show that the system improves the accuracy of audio-only or video-only speaker identification at all levels of acoustic signal-to-noise ratio (SNR) from 0 to 30 dB.
34
An Iterative Decoding Algorithm for Fusion of Multimodal Information
TL;DR: An iterative algorithm to fuse information from multimodal sources is proposed, drawing inspiration from the theory of turbo codes and an analogy between the redundant parity bits of the constituent codes of a turbo code and the information from different sensors in a multi-modal system is drawn.
Human face detection in color images using skin color and template matching models for multimedia on the Web
Krishnan Nallaperumal,Ravi Subban,K. Krishnaveni,L. Fred,R.K. Selvakumar +4 more
- 07 Aug 2006
TL;DR: A novel technique for detecting faces in color images using an adaptive threshold and template matching techniques to segment the skin regions from the non-skin regions is proposed.
20