Proceedings Article10.1109/ICCWAMTIP51612.2020.9317362
Constructing Naive Bayesian Classification Model by Spark for Big Data
Wang Xiaofang,Luo Lan,Zou Qianyin,Liu Fengyu,Liu Jiawei,Huang Di +5 more
- 18 Dec 2020
4
TL;DR: Wang et al. as discussed by the authors proposed a naive Bayesian classifier based on parallel training and prediction on Spark platform, which includes Laplace smoothing and normal distribution functions for data mining and data analysis.
read more
Abstract: Due to the development of big data technology, traditional machine learning algorithms are difficult to deal with massive data. To solve this problem, a naive Bayesian classifier based on parallel training and prediction on Spark platform is proposed. The classifier includes Laplace smoothing and normal distribution functions. Bayesian classification algorithm combined with Spark distributed platform to build a complete functional naive Bayesian model for data mining and data analysis and testing. Experimental results show that the accuracy of MLlib -based optimized continuous feature vector dataset is 9.75% higher than that of the traditional naive Bayes classification algorithm.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Naïve Bayes Classification Model for the Producer Price Index Prediction
Melisa Winda Pertiwi,Mira Kusmira,Rezkiani Rezkiani,Bambang Kelana Simpony,Yanti Apriyani,Iqbal Dzulfiqar Iskandar,Taufik Wibisono,Imam Amirulloh +7 more
TL;DR: In this article , the Naïve Bayes Algorithm was used to predict the Producer Price Index (PPI) for the third quarter of 2019 and the prediction obtained is an increase for Quarter III with a maximum value between 0.961 and 0.980.
Medical Image Denoising Processing Application Technology Based on Combined Filtering
TL;DR: In this paper , the inverse filter models are studied through MATLAB simulation, and it is found that the recovery effect of these filter models is not ideal in the case of noise, so a combined algorithm is proposed.
Optimization of Big Data Mining Algorithm Based on Spark Framework: Preparation of Camera-Ready Contributions to SCITEPRESS Proceedings
TL;DR: The experimental results show that the efficiency of the Eclat algorithm based on the Spark framework is far better than that of the eClat algorithm, and it has high efficiency and good scalability when processing massive data.
Prediction Analysis of Package C Student Graduation at the Bollo DMansel Community Learning Activity Center (PKBM) with the Naïve Bayes Algorithm Method
Muhammad Yassir,Wanda Cahyani +1 more
TL;DR: This study applies the Naive Bayes algorithm to predict Paket C student graduation at Bollo DMansel PKBM in West Papua, identifying attendance, test scores, and participation as key factors influencing graduation, with fairly high accuracy.
References
Attribute and instance weighted naive Bayes
TL;DR: A new improved model called attribute and instance weighted naive Bayes (AIWNB), which combines attribute weighting with instance weighting into one uniform framework and significantly outperform NB and all the other existing state-of-the-art competitors.
94
Composite of medium entropy alloys synthesized using spark plasma sintering
Niraj Chawake,Lavanya Raman,Parthiban Ramasamy,Pradipta Ghosh,Florian Spieckermann,Christoph Gammer,B.S. Murty,Ravi Sankar Kottada,Jürgen Eckert +8 more
TL;DR: In this article, a composite of two different medium entropy alloys (MEAs), i.e., CoCrFeNi and AlCoCrFe) was synthesized using ball milling and spark plasma sintering.
19
Big data scalability based on Spark Machine Learning Libraries
Anna Karen Gárate-Escamilla,Amir Hajjam El Hassani,Emmanuel Andrès +2 more
- 20 Nov 2019
TL;DR: The main contribution of the study is to measure the scalability by calculating the execution time that a classifier achieves with larger workloads, based on Apache Spark, an in-memory distributed application that offers extensive machine learning libraries.
6
Hybrid Program Recommendation Algorithm Based on Spark MLlib in Big Data Environment
Aoxiang Peng,Huiyong Liu +1 more
- 01 Jan 2021
TL;DR: The simulation results show that the parallel operation of Spark MLlib algorithm library not only solves the problem of low timeliness of big data sets but also stabilizes the average RMES of hybrid recommendation algorithm at about 0.52.
5
Construction and Application of Ship Data Mining Platform Based on Spark
Lei Cao,Jingfeng Hu,Ran Li +2 more
- 01 Aug 2019
TL;DR: A Spark-based ship data mining platform to process and mine the surging ship data, against the characteristics of ship AIS data, and results show that the platform works well and the three modules can work normally.