Proceedings Article10.1109/ICTAI.2017.00142
Efficient Prototype Selection Supported by Subspace Partitions
Joel Luis Carbonera,Mara Abel +1 more
- 01 Nov 2017
- pp 921-928
8
TL;DR: An efficient approach for prototype selection called PSSP is proposed that adopts the notion of subspace partition for efficiently splitting the dataset in sets of similar instances and provides a good trade-off between accuracy and reduction, with a significantly lower running time, when compared with other approaches.
read more
Abstract: Nowadays, data mining approaches have been applied for extracting useful knowledge from a huge volume of data. In order to deal with this big data, techniques for prototype (or instance) selection have been applied for reducing the data to a manageable volume and, consequently, for reducing the computational resources that are necessary to apply data mining approaches. However, most of the proposed approaches for prototype selection have a high time complexity and, due to this, they cannot be applied for dealing with big data. In this paper, we propose an efficient approach for prototype selection called PSSP. It adopts the notion of subspace partition for efficiently splitting the dataset in sets of similar instances. In a second step, the algorithm extracts a prototype of each of the previously identified sets. The approach was evaluated on 11 well-known datasets used in a classification task, and its performance was compared to those of 6 state-of-the-art algorithms, considering three measures: accuracy, reduction, and effectiveness. All the obtained results show that, in general, the proposed approach provides a good trade-off between accuracy and reduction, with a significantly lower running time, when compared with other approaches.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Local-Set Based-on Instance Selection Approach for Autonomous Object Modelling
TL;DR: A novel, machine-learning approach for instance selection called Approach for Selection of Border Instances (ASBI), which adopts the notion of local sets to select the most representative instances at the boundaries of the classes, in order to reduce the set of training instances.
Efficient Instance Selection Based on Spatial Abstraction
Joel Luis Carbonera,Mara Abel +1 more
- 01 Nov 2018
TL;DR: The proposed approach for instance selection called ISDSP adopts the notion of spatial partition for efficiently splitting the dataset in sets of similar instances and provides a good trade-off between accuracy and reduction, with a significantly lower running time when compared with other approaches.
15
Instance Ranking and Numerosity Reduction Using Matrix Decomposition and Subspace Learning
Benyamin Ghojogh,Mark Crowley +1 more
- 28 May 2019
TL;DR: This work introduces a framework to achieve this using matrix decomposition and subspace learning to rank data instances by usefulness and reduce the dataset size, and proposes several related algorithms for ranking data instances and performing numerosity reduction.
11
An Efficient Prototype Selection Algorithm Based on Dense Spatial Partitions
Joel Luis Carbonera,Mara Abel +1 more
- 03 Jun 2018
TL;DR: The proposed approach adopts the notion of spatial partition for efficiently dividing the dataset in sets of similar instances and provides a good trade-off between accuracy and reduction, with a significantly lower running time, when compared with other approaches.
7
An Efficient Prototype Selection Algorithm Based on Spatial Abstraction
Joel Luis Carbonera,Mara Abel +1 more
- 03 Sep 2018
TL;DR: An efficient approach for prototype selection called PSSA that adopts the notion of spatial partition for efficiently splitting the dataset in sets of similar instances and provides a good trade-off between accuracy and reduction, with a significantly lower running time when compared to other approaches.
5
References
Instance selection of linear complexity for big data
TL;DR: Two new algorithms with linear complexity for instance selection purposes are presented, one of which reduces complexity and makes it linear with respect to the data set size and the other uses locality-sensitive hashing to find similarities between instances.
90
A class boundary preserving algorithm for data condensation
TL;DR: A new approach is introduced, the Class Boundary Preserving Algorithm (CBP), which is a multi-stage method for pruning the training set, based on a simple but very effective heuristic for instance removal.
79
IRAHC: Instance Reduction Algorithm using Hyperrectangle Clustering
TL;DR: This work proposes an instance reduction method based on hyperrectangle clustering, called Instance Reduction Algorithm using Hyperrectangle Clustering (IRAHC), which removes non-border (interior) instances and keeps border and near border ones.
68
A Density-Based Approach for Instance Selection
Joel Luis Carbonera,Mara Abel +1 more
- 09 Nov 2015
TL;DR: A simple and effective density-based approach for instance selection that evaluates the instances of each class separately and keeps only the densest instances in a given (arbitrary) neighborhood, which achieves a reasonably low time complexity.
43
Learning to detect representative data for large scale instance selection
TL;DR: This work introduces the ReDD approach, which is based on outlier pattern analysis and prediction, and empirically evaluates ReDD over 50 domain datasets to examine the effectiveness of the learned detector, using four very large scale datasets for validation.
32
Related Papers (5)
Joel Luis Carbonera,Mara Abel +1 more
- 03 Sep 2018
Joel Luis Carbonera,Mara Abel +1 more
- 01 Nov 2018
Joel Luis Carbonera
- 28 Aug 2017