Enhancing Open-Set Recognition using Clustering-based Extreme Value Machine (C-EVM)
James Henrydoss,Steve Cruz,Chunchun Li,Manuel Günther,Terrance E. Boult +4 more
- 10 Dec 2020
- pp 441-448
TL;DR: In this paper, the authors proposed a clustering-based extension of the Extreme Value Machine (EVM) to improve the end-to-end prediction performance by combining Density-based spatial clustering of applications with noise (DBSCAN)-based clustering with a novel Nearby Clusters (NC) algorithm.
read more
Abstract: In real-world deployments, machine learning applications find challenges when accessing ever-increasing volumes of data – the real world is open and often presents data from classes not seen in training. Open-set recognition is a growing area of machine learning addressing such problems. This research work advances the state-of-the-art in open-set recognition, the Extreme Value Machine (EVM), with a novel clustering-based extension (C-EVM) during training to improve the end-to-end prediction performance. The C-EVM combines Density-based spatial clustering of applications with noise (DBSCAN)-based clustering with a novel Nearby Clusters (NC) algorithm during model fitting to reduce computation while improving accuracy. Our experiments show a statistically significant improvement of 5-10% in macro F1-score over the state-of-the-art EVM on open-set testing using the KDD CUP-99 data set. Past work on open set recognition often traded improved open-set robustness for a decrease in closed-set accuracy, whereas C-EVM outperforms the EVM in both closed-set and open-set recognition. Testing on subsets of ImageNet-2012 with varying numbers of classes, the C-EVM statistically significantly out performs EVM when using deep features. A parameterless Hierarchical DBSCAN (HDBSCAN)-based C-EVM variant is introduced as part of this work that scales well for large data sets. Finally, both EVM and C-EVM can operate as kernel-free incremental learners, enabling these open-set multi-class classifiers to be useful for streaming and big data applications.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Figures

TABLE I: EVM VS. C-EVM. F1-measure based open-set recognition performance at various levels of openness on KDD. C-EVM using DBSCAN with different PSI thresholds and also using HDBSCAN and K-Means variants are shown. Three columns show varying PSI thresholds (60%, 70%, 80%) and it can be seen there is not much difference with variation in the threshold, or the use of HDBSCAN, K-Means (where the number of clusters set to 11) all of which are significantly better then the classic EVM. For C-EVM we used three levels 60%, 70%, 80% and for EVM we used 50%. 
Fig. 1: EXTREME VECTORS IN EVM AND C-EVM. Each number represents a point in a class, and the colored circles represent the 95% confidence interval around the Extreme Vectors (EVs). Standard EVM selects EVs that cover the data, but often take over more of the open-space and need more EVs, which are shown as 1, three in total. C-EVM uses cluster centroids, the star, as the location of EV points. Less EVs are used and often cover less in open space, thereby reducing open-set risk and improving open-set recognition performance when unknown classes are present. 
TABLE II: C-EVM VS EVM TRAINING TIME. The training time is measured as milliseconds per processed point on the different KDD subsets. 
Fig. 4: C-EVM AND EVM OPEN-SET PERFORMANCE USING IMAGENET. The number of classes is varied in a closedset testing paradigm. As the number of classes increases, there is greater confusion, and both algorithms degrade, but the C-EVM is statistically significantly better (p = 0.04), and its advantage increases as the number of classes increases. 
Fig. 2: NEARBY CLUSTERS. To estimate the PSI model for the centroid of the cluster in the current class of interest (yellow), only samples from nearby clusters of other classes (red, green, blue) are utilized while faraway clusters (white) are disregarded. 
Fig. 3: EXPERIMENTS ON KDD. Performance comparison on KDD data sets of C-EVM with previous state-of-the-art EVM at various levels of openness and various PSI thresholds.
Citations
The design of error-correcting output codes algorithm for the open-set recognition
Kun-Hong Liu,Wang-Ping Zhan,Yi-Fan Liang,Ya-Nan Zhang,Hong-Zhou Guo,Junfeng Yao,Qingqiang Wu,Qingqi Hong +7 more
TL;DR: This study applies the Error-Correcting Output Codes (ECOC) framework to handle the open- set problem by dynamically adding new functions to deal with the unknown classes, named ECOC-OS, and effectively improves the performance compared with other open-set recognition methods.
7
Open-Set Intrusion Detection with MinMax Autoencoder and Pseudo Extreme Value Machine
18 Jul 2022
TL;DR: OpenIDS as mentioned in this paper proposes an open-set intrusion detection system, which addresses the problem through three modules: the MinMax autoencoder, the classifier, and the pseudo extreme value machine.
5
Exploring the Open World Using Incremental Extreme Value Machines
21 Aug 2022
TL;DR: In this article , a modification of the widely known Extreme Value Machine (EVM) was introduced to enable open world recognition, which is a demanding task that is, to the best of our knowledge, addressed by only a few methods.
References
Toward Open Set Recognition
TL;DR: This paper explores the nature of open set recognition and formalizes its definition as a constrained minimization problem, and introduces a novel “1-vs-set machine,” which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel.
1.5K
Outlier detection for high dimensional data
Charu C. Aggarwal,Philip S. Yu +1 more
- 01 May 2001
TL;DR: New techniques for outlier detection which find the outliers by studying the behavior of projections from the data set are discussed.
1.2K
Learning Feature Representations with K-Means
Adam Coates,Andrew Y. Ng +1 more
- 01 Jan 2012
TL;DR: This chapter will summarize recent results and technical tricks that are needed to make effective use of K-means clustering for learning large-scale representations of images and connect these results to other well-known algorithms to make clear when K-Means can be most useful.
809
Traffic classification using clustering algorithms
Jeffrey Erman,Martin Arlitt,Anirban Mahanti +2 more
- 11 Sep 2006
TL;DR: This work considers two unsupervised clustering algorithms, namely K-Means and DBSCAN, that have previously not been used for network traffic classification and evaluates these two algorithms and compares them to the previously used AutoClass algorithm, using empirical Internet traces.
Efficient kNN classification algorithm for big data
TL;DR: This paper proposes to first conduct a k-means clustering to separate the whole dataset into several parts, each of which is then conducted kNN classification, and results show that the proposed kNN Classification works well in terms of accuracy and efficiency.
563