Journal Article10.1016/J.KNOSYS.2020.105694
Incremental learning imbalanced data streams with concept drift: The dynamic updated ensemble algorithm
82
TL;DR: A chunk-based incremental ensemble algorithm called Dynamic Updated Ensemble (DUE) for learning imbalanced data streams with concept drift, which can timely react to multiple kinds of concept drifts and keep a limited number of classifiers to ensure high efficiency.
read more
Abstract: Learning nonstationary data streams has been well studied in recent years. However, most of the researches assume that the class imbalance of data streams is relatively balanced. Only a few approaches tackle the joint issue of concept drift and class imbalance due to its complexity. Meanwhile, the existing chunk ensembles for classifying imbalanced nonstationary data streams always need to store previous data, which consumes plenty of memory usage. To overcome these issues, we propose a chunk-based incremental ensemble algorithm called Dynamic Updated Ensemble (DUE) for learning imbalanced data streams with concept drift. Compared to the existing techniques, its merits are five-fold: (1) it learns one chunk at a time without requiring access to previous data; (2) it emphasizes misclassified examples in the model update procedure; (3) it can timely react to multiple kinds of concept drifts; (4) it can adapt to the new condition when switching majority class to minority class; (5) it keeps a limited number of classifiers to ensure high efficiency. Experiments on synthetic and real datasets demonstrate the effectiveness of DUE in learning nonstationary imbalanced data streams.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework
TL;DR: This work proposes a standardized, exhaustive, and comprehensive experimental framework to evaluate algorithms in a collection of diverse and challenging imbalanced data stream scenarios to create complete, trustworthy, and fair evaluation of newly proposed methods.
82
ROSE: robust online self-adjusting ensemble for continual learning on imbalanced drifting data streams
Alberto Cano,Bartosz Krawczyk +1 more
TL;DR: In this paper , a robust online self-adjusting ensemble (ROSE) classifier is proposed to detect concept drift and create a background ensemble for faster adaptation to changes in data streams.
A survey on imbalanced learning: latest research, applications and future directions
Wuxing Chen,Kaixiang Yang,Zhiwen Yu,Yifan Shi,C. L. Philip Chen +4 more
TL;DR: This survey reviews recent advancements in imbalanced learning, covering strategies for classification and regression tasks, deep long-tail learning, and real-world applications in management science, engineering, and other fields, highlighting emerging challenges and future directions.
46
UFFDFR: Undersampling framework with denoising, fuzzy c-means clustering, and representative sample selection for imbalanced data classification
Ming Zheng,Tong Li,Xiaoyao Zheng,Qingying Yu,Chuanming Chen,Ding Zhou,Changlong Lv,Weiyi Yang +7 more
TL;DR: A novel three-stage undersampling framework with denoising, fuzzy c-means clustering, and representative sample selection (UFFDFR) is proposed that improves the classification performance on imbalanced data by removing noise and unrepresentative samples from the majority class.
32
A survey of active and passive concept drift handling methods
TL;DR: Many concept drift handling methods in this survey are analyzed and summarized in terms of the comparing algorithms, learning model, applicable drift type, advantages, and disadvantages of the algorithms.
28
References
•Journal Article
Statistical Comparisons of Classifiers over Multiple Data Sets
TL;DR: A set of simple, yet safe and robust non-parametric tests for statistical comparisons of classifiers is recommended: the Wilcoxon signed ranks test for comparison of two classifiers and the Friedman test with the corresponding post-hoc tests for comparisons of more classifiers over multiple data sets.
A survey on concept drift adaptation
TL;DR: The survey covers the different facets of concept drift in an integrated way to reflect on the existing scattered state of the art and aims at providing a comprehensive introduction to the concept drift adaptation for researchers, industry analysts, and practitioners.
Mining high-speed data streams
Pedro Domingos,Geoff Hulten +1 more
- 01 Aug 2000
TL;DR: This paper describes and evaluates VFDT, an anytime system that builds decision trees using constant memory and constant time per example, and applies it to mining the continuous stream of Web access data from the whole University of Washington main campus.
Dissecting Android Malware: Characterization and Evolution
Yajin Zhou,Xuxian Jiang +1 more
- 20 May 2012
TL;DR: Systematize or characterize existing Android malware from various aspects, including their installation methods, activation mechanisms as well as the nature of carried malicious payloads reveal that they are evolving rapidly to circumvent the detection from existing mobile anti-virus software.
Mining time-changing data streams
Geoff Hulten,Laurie Spencer,Pedro Domingos +2 more
- 26 Aug 2001
TL;DR: An efficient algorithm for mining decision trees from continuously-changing data streams, based on the ultra-fast VFDT decision tree learner is proposed, called CVFDT, which stays current while making the most of old data by growing an alternative subtree whenever an old one becomes questionable, and replacing the old with the new when the new becomes more accurate.