About: Multi-document summarization is a research topic. Over the lifetime, 2270 publications have been published within this topic receiving 71850 citations.
TL;DR: The proposed technique of Feature Extraction is unsupervised, automated and also domain independent and the improved effectiveness of the proposed approach is demonstrated on a real life dataset that is crawled from many reviewing websites such as CNET, Amazon etc.
Abstract: The growth of E commerce has led to the abundance growth of opinions on the web, thereby necessitating the task of Opinion Summarization, which in turn has great commercial significance. Feature extraction in Opinion Summarization is very crucial as selection of relevant features reduce the feature space which successfully reduces the complexity of the classification task. The paper suggests extensive pre-processing technique & an algorithm for extracting features from Reviews/Blogs. The proposed technique of Feature Extraction is unsupervised, automated and also domain independent. The improved effectiveness of the proposed approach is demonstrated on a real life dataset that is crawled from many reviewing websites such as CNET, Amazon etc.
TL;DR: A feature priority based filtering method for summarization that used sentence location as main feature and other features in priority to filter the redundant sentences and performs uniformly as compared to the best results for particular combination of features.
TL;DR: The experimental results on an open benchmark datasets from DUC01 and DUC02 show that the proposed approach can improve the performance compared to state-of-the-art summarization approaches.
Abstract: The technology of automatic document summarization is maturing and may provide a solution to the information overload problem. Nowadays, document summarization plays an important role in information retrieval. With a large volume of documents, presenting the user with a summary of each document greatly facilitates the task of finding the desired documents. Document summarization is a process of automatically creating a compressed version of a given document that provides useful information to users, and multi-document summarization is to produce a summary delivering the majority of information content from a set of documents about an explicit or implicit main topic. According to the input text, in this paper we use the knowledge base of Wikipedia and the words of the main text to create independent graphs. We will then determine the important of graphs. Then we are specified importance of graph and sentences that have topics with high importance. Finally, we extract sentences with high importance. The experimental results on an open benchmark datasets from DUC01 and DUC02 show that our proposed approach can improve the performance compared to state-of-the-art summarization approaches.
TL;DR: This work proposes an extensible framework META to enable analysts to easily and selectively extract and summarize events from different views with different resolutions, and defines a summarization language that includes a set of atomic operators to manipulate the meta-data.
Abstract: : Event summarization is an effective process that mines and organizes event patterns to represent the original events. It allows the analysts to quickly gain the general idea of the events. In recent years, several event summarization algorithms have been proposed, but they all focus on how to find out the optimal summarization results, and are designed for one-time analysis. As event summarization is a comprehensive analysis work, merely handling this problem with a single optimal algorithm is not enough. In the absence of an integrated summarization solution, we propose an extensible framework META to enable analysts to easily and selectively extract and summarize events from different views with different resolutions. In this framework, we store the original events in a carefully-designed data structure that enables an efficient storage and multiresolution analysis. On top of the data model, we define a summarization language that includes a set of atomic operators to manipulate the meta-data. Furthermore, we present 5 commonly used summarization tasks, and show that all these tasks can be easily expressed by the language. Experimental evaluation on both real and synthetic datasets demonstrates the efficiency and effectiveness of our framework.
TL;DR: This paper focuses on subject shift and presents a method for extracting key paragraphs from documents that discuss the same event using the results of event tracking which starts from a few sample documents and finds all subsequent documents that discusses the sameevent.
Abstract: For multi-document summarization where documents are collected over an extended period of time, the subject in a document changes over time. This paper focuses on subject shift and presents a method for extracting key paragraphs from documents that discuss the same event. Our extraction method uses the results of event tracking which starts from a few sample documents and finds all subsequent documents that discuss the same event. The method was tested on the TDT1 corpus, and the result shows the effectiveness of the method.