TL;DR: The audio analysis, search, and classification engine described here reduces sounds to perceptual and acoustical features, which lets users search or retrieve sounds by any one feature or a combination of them, by specifying previously learned classes based on these features.
Abstract: Many audio and multimedia applications would benefit from the ability to classify and search for audio based on its characteristics. The audio analysis, search, and classification engine described here reduces sounds to perceptual and acoustical features. This lets users search or retrieve sounds by any one feature or a combination of them, by specifying previously learned classes based on these features, or by selecting or entering reference sounds and asking the engine to retrieve similar or dissimilar sounds.
TL;DR: FingerMouse, a freehand pointing system, detects pointing hand poses and tracks moving fingertips in close to real time and contributes to inductive learning of hand gesture poses.
Abstract: Unencumbered hand gesture interfaces encompass both 3D interaction and 2D pointing. A model developed to study 3D interaction requires determining gestural strokes and hand motion dynamics and recognizing hand poses. Extended variable valued logic and a rule based induction algorithm contribute to inductive learning of hand gesture poses, yielding a recognition rate of 94 percent. FingerMouse, a freehand pointing system, detects pointing hand poses and tracks moving fingertips in close to real time.
TL;DR: It is concluded that constant time length and hybrid techniques can greatly reduce the total system cost and data limit admission control offers a moderate gain for a moderate increase in implementation complexity.
Abstract: Comparing techniques for storage and real time retrieval of variable bit rate video data for multiple simultaneous users, we conclude that constant time length and hybrid techniques can greatly reduce the total system cost. Of three admission control techniques, data limit admission control offers a moderate gain for a moderate increase in implementation complexity. We also applied our data placement and admission control strategies to an interleaved disk array.
TL;DR: A meeting indexer, Jabber, is created that uses content-based indexing of the audio stream to access these parallel streams of speech recognition, then groups the recognized words into semantically linked trees.
Abstract: Meetings in which participants are linked by video, audio, and shared computer applications produce several parallel information streams. We created a meeting indexer, Jabber, that uses content-based indexing of the audio stream to access these parallel streams. It performs speech recognition on the audio stream, then groups the recognized words into semantically linked trees. The user interface is designed to display information with minimal distraction during meetings.
TL;DR: The Virtual Studio concept replaces real background sets with computer-generated synthetic scenes, and image sequences combine real foreground action, shot in the studio, with computer generated backgrounds.
Abstract: The Virtual Studio concept replaces real background sets with computer-generated synthetic scenes. Image sequences combine real foreground action, shot in the studio, with computer generated backgrounds. The Mona Lisa project seeks to develop and integrate technologies (algorithms, software, and hardware) needed to construct, handle, and render 3D models in real time.
TL;DR: This special issue on multimodal interaction techniques and the useful new applications they support is based on guest editor Glinert and Blattner's review of the potential benefits and research and design challenges involved.
Abstract: We have more ways to interact with our computers than ever before, and these get more "human friendly" every day. Guest editors Glinert and Blattner review the potential benefits of multimodal interfaces and the research and design challenges involved. The four articles that follow comprise this special issue on multimodal interaction techniques and the useful new applications they support.
TL;DR: By modeling difficult sources of linguistic variability in speech and language, this work can design interfaces that transparently guide human input to match system processing capabilities.
Abstract: By modeling difficult sources of linguistic variability in speech and language, we can design interfaces that transparently guide human input to match system processing capabilities. Such work will yield more user centered and robust interfaces for next generation spoken language and multimodal systems.
TL;DR: By modeling difficult sources of linguistic variability in speech and language, the authors can design interfaces that transparently guide human input to match system processing capabilities, yielding more user centered and robust interfaces for next generation spoken language and multimodal systems.
Abstract: By modeling difficult sources of linguistic variability in speech and language, we can design interfaces that transparently guide human input to match system processing capabilities. Such work will yield more user centered and robust interfaces for next generation spoken language and multimodal systems.
TL;DR: The Speech Aware Multimedia (SAM) system as mentioned in this paper is a speech recognition system for the Web that allows the user to browse arbitrary Web pages using only speech as the input medium.
Abstract: Computer users have long desired a personal software agent that could execute verbal commands. Today's World Wide Web (WWW or Web), with its point and click hypertext interface, makes a tremendous amount of information readily available online. A speech interface would make the Web even more powerful, allowing us to access information by surfing the Web by voice. TI have developed Speech Aware Multimedia (SAM) with this in mind, to make information on the Web more accessible and useful. They combined an innovative speech recognition engine with the Web to let anyone browse arbitrary Web pages using only speech as the input medium. Speech brings added flexibility and power to the classical Web interface and makes information access more natural. Today's speech recognition capability is well matched to Web browsing. The Web page provides a natural, well defined context for a speech recognition application. The recognition engine does not need to recognize any and all possible phrases, but only those phrases pertaining to the specific page in view at the moment. This context imposes limits that significantly aid recognition performance. Furthermore, the visual information on a page prompts the user on what to request and how to request it by voice.
TL;DR: Teleaction objects, or multimedia objects with knowledge structures, can be designed using visual languages to automatically respond to events and perform tasks like "find related books" in a virtual library.
Abstract: Visual languages, which let users customize iconic sentences, can be extended to accommodate multimedia objects, letting users access media dynamically. Teleaction objects, or multimedia objects with knowledge structures, can be designed using visual languages to automatically respond to events and perform tasks like "find related books" in a virtual library.
TL;DR: Examining multimedia program allocation in distributed multimedia-on-demand systems lets us consider how best to provide programs at will from geographically scattered servers.
Abstract: Examining multimedia program allocation in distributed multimedia-on-demand systems lets us consider how best to provide programs at will from geographically scattered servers. Each user is associated with a local server but can transparently access any program located at any server. The authors consider user demand for various programs, storage limitations of the multimedia servers, and costs of storing and transporting the programs.
TL;DR: A stream conversion scheme that encodes P frames as I frames after decompression and playout of each P frame eliminates extra memory needs, making P-I conversion a cost effective solution.
Abstract: Interactive playout of MPEG-encoded video entails new ways of handling data. Transforming the standard MPEG stream to a local form at the player device enables efficient interactive playout even when available buffer space is constrained. A stream conversion scheme that encodes P frames as I frames after decompression and playout of each P frame eliminates extra memory needs, making P-I conversion a cost effective solution.
TL;DR: An architecture for creating multimedia documents is designed by means of a logical structure, a layout structure, and a rendering scenario, which is a schedule for document playback.
Abstract: Multimedia documents differ significantly from traditional documents composed of text and geometric graphics. The introduction of continuous media such as audio, video, and computer-generated graphics imposes new requirements on document representation and information storage. We designed an architecture for creating multimedia documents by means of a logical structure, a layout structure, and a rendering scenario, which is a schedule for document playback.
TL;DR: High-structured interfaces such as hypermaps, discussed here, are a necessary and useful way to structure multimedia components and let users easily navigate data sets.
Abstract: Developments in multimedia and scientific visualization have greatly expanded the technical capabilities of geographic information systems. Users can visually explore, analyze, and present data and gain insight on spatial relations and patterns. But how do users manage all the information that reaches them? Highly-structured interfaces such as hypermaps, discussed here, are a necessary and useful way to structure multimedia components and let users easily navigate data sets.
TL;DR: The Premo (Presentation Environment for Multimedia Objects) standard is a presentation environment that aims to provide a standard programming environment in a very general sense, one that helps promote portable multimedia applications and is object oriented.
Abstract: Developers needing to realize high-level multimedia applications are essentially left on their own. Only a few programming tools allow the creation of multimedia effects based on a more general model than multimedia documents. No currently available ISO standard encompasses these needs. A standard in this area should focus more on the presentation aspects of multimedia and less on the coding, transfer, or hypermedia document aspects, which are covered other standards. It should also concentrate on programming tools rather than multimedia document format. These are exactly the main concerns of the Premo (Presentation Environment for Multimedia Objects) standard, the subject of the article. Premo's major features can be briefly summarized as follows: Premo is a presentation environment that aims to provide a standard programming environment in a very general sense, one that helps promote portable multimedia applications; Premo targets multimedia presentation, whereas earlier SC24 standards concentrated either on synthetic graphics or image-processing systems; Premo is object oriented. This means that, through standard object-oriented techniques, a Premo implementation becomes extensible and configurable. Object-oriented technology also provides a framework to describe distribution in a consistent manner.
TL;DR: The authors propose an encompassing framework that offers all services under a unifying hybrid object/hypermedia paradigm that is applicable to distributed multimedia application development.
Abstract: Distributed multimedia application development is supported by a growing number of services. While these pave the way for sophisticated multimedia support in distributed systems, using them involves interfacing inconsistent and swiftly evolving technologies. To alleviate this, the authors propose an encompassing framework that offers all services under a unifying hybrid object/hypermedia paradigm.
TL;DR: The EFX digital editing and effects environment integrates facilities for nonlinear editing of digitized film, video, and audio with sophisticated image-manipulating special effects with a powerful parallel-processing computer that computes the special effects and plays back the uncompressed film or video in real time.
Abstract: The EFX digital editing and effects environment integrates facilities for nonlinear editing of digitized film, video, and audio with sophisticated image-manipulating special effects. EFX offers an intuitive, visual, direct-manipulation user interface for building multimedia compositions. This front end, discussed in the paper, is coupled with a powerful parallel-processing computer that computes the special effects and plays back the uncompressed digitized film or video in real time.
TL;DR: The authors use a typical business conference as an example to show how graphical models and agent prototyping empower flexible working environments.
Abstract: A collaborative work group that links workers through distributed computers is only as productive as the system itself Software agents are emerging as a way to boost workers' productivity, performing group tasks independently The authors use a typical business conference as an example to show how graphical models and agent prototyping empower flexible working environments
TL;DR: The authors believe that QoS development cannot be done in isolation from the applications to be supported, which must form an integral part of the project, and focus on three application areas: interactive teaching and learning, mobile systems, and virtual reality.
Abstract: A major project at Lancaster University is the development of network infrastructures capable of supporting the quality-of-service (QoS) requirements of a wide range of distributed multimedia applications. The project includes more than thirty researchers and covers not only the network support but also the enabling technologies. The authors believe that QoS development cannot be done in isolation from the applications to be supported, which must form an integral part of the project. They focus on three application areas: interactive teaching and learning, mobile systems, and virtual reality. They chose applications that stretch network support to its limit. To realize the full potential of these distributed multimedia applications, the underlying network must satisfy all these requirements concurrently. They give a brief overview of their activities to this end and discuss the need for QoS support.
TL;DR: The Institute of Electrical and Electronics Engineers began working on a transmission standard for multimedia networking almost a decade ago and since the standard's recommendation in December 1994, it has taken little more than a year to capture the attention of manufacturers, vendors, resellers, systems integrators, and network users.
Abstract: The Institute of Electrical and Electronics Engineers began working on a transmission standard for multimedia networking almost a decade ago. Since the standard's recommendation in December 1994, it has taken little more than a year to capture the attention of manufacturers, vendors, resellers, systems integrators, and network users. The IEEE standard for multimedia in a local area network (LAN) is based on ISLAN-1GT, or isochronous Ethernet. Better known as isoEthernet, the standard incorporates an integrated services digital network phone system with LAN-based Ethernet for mixed data transmission. The IEEE standard, designated 802.9a for LANs and metropolitan area networks, is simply a physical layer standard-just like Ethernet 802.3/ISO 8802-3 and Token Ring 802.5-in that it defines the standard's media access control and physical layer.
TL;DR: In this chapter, Shor's efficient quantum algorithms for the computational tasks of integer factorization and the evaluation of discrete logarithms are studied.
Abstract: Among the most remarkable successes of quantum computation are Shor's efficient quantum algorithms for the computational tasks of integer factorization and the evaluation of discrete logarithms. In...
TL;DR: The paper discusses visual programming by example which allows designers to define animation rules by "training" agents, thereby building behavioral rules into specification models that run automatically during the execution of the virtual environment.
Abstract: In virtual reality interfaces, realistic animation of virtual agents enhances human-computer interaction by supporting direct engagement in the virtual environment. The paper discusses visual programming by example which allows designers to define animation rules by "training" agents, thereby building behavioral rules into specification models that run automatically during the execution of the virtual environment. This allows for direct and effective replication of real-life phenomena and agent reactions to environmental stimuli.
TL;DR: Banking in the global village becomes reality with multimedia communication systems that connect far flung regions and integrate all their various marketplaces.
Abstract: Banking in the global village becomes reality with multimedia communication systems that connect far flung regions and integrate all their various marketplaces. International banks in particular face the major challenge of achieving the right balance between global and local. They must identify local customers' requirements and satisfy them by exploiting the expertise and technology available worldwide.
TL;DR: In this article, the authors propose a solution for mission-critical applications such as industrial process control, crisis management, and military command and control, where multimedia information will play a significant role in the next generation of missioncritical applications.
Abstract: Multimedia information will play a significant role in the next generation of mission-critical applications such as industrial process control, crisis management, and military command and control. Unlike multimedia applications for entertainment (such as video-on-demand services) and office automation (such as videoconferencing), multimedia in mission-critical applications has several unique requirements; on-line processing with real-time constraints due to its event driven nature; the ability to distinguish the criticality of concurrent streams with different degrees of importance; fault tolerance to hostile environments, and adaptability to system changes.
TL;DR: The article gives an overview of the Premo (Presentation Environment for Multimedia Objects) standard and presented the motivation behind Premo's development and gave a general Overview of the standard's technical content.
Abstract: For pt.I see ibid., p.83-9, Fall, 1996. The article gives an overview of the Premo (Presentation Environment for Multimedia Objects) standard. The first part presented the motivation behind Premo's development and gave a general overview of the standard's technical content. The interested reader may also refer to the Premo document itself, as well as the various ISO documents and other publications. A World Wide Web site is presented which provides a good starting point to navigate through and access all available documents.
TL;DR: The overall vision is to develop a Virtual Cleanroom, in which a student can use a high-performance multimedia workstation to follow a semiconductor wafer through a complete processing sequence-from bare silicon to finished IC.
Abstract: Integrated circuits are the driving force behind technological breakthroughs and innovative product development in electronics today. ICs are fabricated by a sophisticated series of steps that range in number from dozens to hundreds, depending on the complexity of the circuit function. Although comprehensive training in the field of semiconductor manufacturing requires thorough understanding of all phases of the fabrication process, the study of IC fabrication requires an unusually diverse familiarity with physics, inorganic chemistry, semiconductor devices, and statistics. To alleviate impediments resulting from resource constraints as well as to enhance students' educational experience, we are developing ways to teach microelectronic processing using interactive multimedia at the Georgia Institute of Technology. Our overall vision is to develop a Virtual Cleanroom, in which a student can use a high-performance multimedia workstation (equipped with the necessary audio, video, and graphics capabilities) to follow a semiconductor wafer through a complete processing sequence-from bare silicon to finished IC.
TL;DR: The receiver of a message must use a personal knowledge base, the context of the message, and sometimes gestures to resolve the ambiguity inherent in the message.
Abstract: eople communicate through various mechaP nisms: language, intonation, gestures, facial expressions, and many other subtle means. The language used in our communication with others, called natural language, forms the basis of our written communication and our speech. As civilizations evolved, though, people in different parts of the world designed different languages. To coordinate activities with others thus requires us to develop protocols and communicate them to everybody in our group. Further, group members must follow these protocols to carry out all of the group’s activities. As a result, these protocols become the language of the group. I find it interesting that we in computer science call the languages used by people natural languages and call the languages used by computers programming languages. Natural languages are every bit as artificial as programming languages. Natural languages simply have a broader scope and hence more flexibility, which results in ambiguity. The receiver of a message must use a personal knowledge base, the context of the message, and sometimes gestures to resolve the ambiguity inherent in the message. Communication takes place using the other senses, not just language alone.