Proceedings Article10.1109/SAMI.2014.6822410
Distributed boosting algorithm for classification of text documents
Martin Sarnovsky,Michal Vronc +1 more
- 29 May 2014
- pp 217-220
11
TL;DR: Main objective of the paper is to present the implementation of distributed boosting algorithm based on Map Reduce paradigm and has used the GridGain framework as a platform for distributed data processing and has tested the implemented solution on two different dataset within the authors' testing environment.
read more
Abstract: Presented paper focuses on the area of analysis and classification of textual documents. We present the classification of documents based on boosting method applied on the decision tree algorithm. Main objective of the paper is to present the implementation of distributed boosting algorithm based on Map Reduce paradigm. We have used the GridGain framework as a platform for distributed data processing and have tested the implemented solution on two different dataset within our testing environment.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
A survey of machine learning for big data processing
TL;DR: A literature survey of the latest advances in researches on machine learning for big data processing finds some promising learning methods in recent studies, such as representation learning, deep learning, distributed and parallel learning, transfer learning, active learning, and kernel-based learning.
A Comprehensive Study of Big Data Machine Learning Approaches and Challenges
Neelam Singh,Devesh Pratap Singh,Bhasker Pant +2 more
- 01 Dec 2017
TL;DR: This paper endow with a literature analysis related to the up-to-the-minute progress in researches on big data processing deploying Machine Learning as an analytical tool, with a focus on the promising learning methods like transfer learning, active learning, deep learning, representation learning, distributed, kernel-based learning and parallel learning.
22
Internet Data Analysis Methodology for Cyberterrorism Vocabulary Detection, Combining Techniques of Big Data Analytics, NLP and Semantic Web
Iván Castillo-Zúñiga,Francisco Javier Luna-Rosas,Laura C. Rodríguez-Martínez,Jaime Muñoz-Arteaga,Jaime Iván López-Veyna,Mario A. Rodríguez-Díaz +5 more
TL;DR: A methodology for the analysis of data on the Internet, combining techniques of Big Data analytics, NLP and semantic web in order to find knowledge about large amounts of information on the web, reaching 576% time savings with parallel processing.
22
Using a distributed deep learning algorithm for analyzing big data in smart cities
Mohammed Anouar Naoui,Brahim Lejdel,Mouloud Ayad,Abdelfattah Amamra,Okba Kazar +4 more
- 13 Apr 2020
TL;DR: This research needs the application of other deep learning models, such as convolution neuronal network and autoencoder, to describe the distributed deep learning for smart cities in big data systems.
10
From distributed machine to distributed deep learning: a comprehensive survey
Mohammad Dehghani,Zahra Yazdanparast +1 more
TL;DR: This work investigates distributed deep learning algorithms in classification and clustering, deep learning and deep reinforcement learning groups, and highlighted the limitations that should be addressed in future research.
5
References
Cloud-based clustering of text documents using the GHSOM algorithm on the GridGain platform
Martin Sarnovsky,Z. Ulbrik +1 more
- 23 May 2013
TL;DR: An overview of the research activities aimed on efficient use of distributed computing concepts for text-mining tasks and the GHSOM (Growing Hierarchical Self-Organizing Maps) algorithm for clustering of text documents is presented.
20
The Automatic Categorization of Arabic Documents by Boosting Decision Trees
Saeed Raheel,Joseph Dichy,Mohamed Hassoun +2 more
- 29 Nov 2009
TL;DR: This paper aims at exploring the technique of Boosting and its effectiveness with the automatic classification of Arabic documents and compares its performance with results obtained respectively with Support Vector Machines and Naïve Bayesian Networks.
19
•Journal Article
Mirroring of Knowledge Practices based on User-defined Patterns
Christoph Richter,Jozef Wagner +1 more
TL;DR: This paper suggests high-level requirements for mirroring tools in support of practice transformation and introduces a software tool called Timeline-Based Analyzer (TLBA) that was designed and developed in response to these requirements.
Meteorological phenomena forecast using data mining prediction methods
František Babič,Peter Bednar,František Albert,Ján Paralič,Juraj Bartok,Ladislav Hluchý +5 more
- 21 Sep 2011
TL;DR: The goal was to design, implement and evaluate a different approach based on suitable techniques and methods from data mining domain for timely prediction of meteorological phenomena such as fog or low cloud cover in United Arab Emirates and Slovakia.
10
Comparison of standard and sparse-based implementation of GOSCL algorithm
Peter Butka,Jana Pocsova,Jozef Pócs +2 more
- 01 Nov 2012
TL;DR: Experimental comparison on time complexity of algorithm for creation of Generalized One-Sided Concept Lattices according to the sparseness of the input data table between standard and sparse-based implementation is provided.
7
Related Papers (5)
R. Janani,S. Vijayarani +1 more
- 01 Jan 2021
Dong-Hui Kim,Dong-hyeok Lee,Won Don Lee +2 more
- 24 Jul 2006
J. Ponni,K. L. Shunmuganathan +1 more
- 01 Dec 2013