Proceedings Article10.1145/3318464.3383126
Automating Exploratory Data Analysis via Machine Learning: An Overview
Tova Milo,Amit Somech +1 more
- 11 Jun 2020
- pp 2617-2622
91
TL;DR: This tutorial reviews recent lines of work for automating EDA, starting from recommender systems for suggesting a single exploratory action, going through kNN-based classifiers and active-learning methods for predicting users' interestingness preferences, and finally to fully automates EDA using state-of-the-art methods such as deep reinforcement learning and sequence-to-sequence models.
read more
Abstract: Exploratory Data Analysis (EDA) is an important initial step for any knowledge discovery process, in which data scientists interactively explore unfamiliar datasets by issuing a sequence of analysis operations (e.g. filter, aggregation, and visualization). Since EDA is long known as a difficult task, requiring profound analytical skills, experience, and domain knowledge, a plethora of systems have been devised over the last decade in order to facilitate EDA. In particular, advancements in machine learning research have created exciting opportunities, not only for better facilitating EDA, but to fully automate the process. In this tutorial, we review recent lines of work for automating EDA. Starting from recommender systems for suggesting a single exploratory action, going through kNN-based classifiers and active-learning methods for predicting users' interestingness preferences, and finally to fully automating EDA using state-of-the-art methods such as deep reinforcement learning and sequence-to-sequence models. We conclude the tutorial with a discussion on the main challenges and open questions to be dealt with in order to ultimately reduce the manual effort required for EDA.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Mitigating Bias in Radiology Machine Learning: 1. Data Handling.
Pouria Rouzrokh,Bardia Khosravi,Shahriar Faghani,Mana Moassefi,Diana V. Vera Garcia,Yashbir Singh,Kuan Zhang,Gian Marco Conte,Bradley J. Erickson +8 more
- 01 Sep 2022
TL;DR: This report presents 12 suboptimal practices during data handling of an ML study, explains how those practices can lead to biases, and describes what may be done to mitigate them.
83
Missing value imputation on multidimensional time series
Parikshit Bansal,Prathamesh Deshpande,Sunita Sarawagi +2 more
- 01 Jul 2021
TL;DR: DeepMVI as discussed by the authors is a deep learning method for missing value imputation in multidimensional time-series datasets, which uses a neural network to combine fine-grained and coarsegrained patterns along a time series, and trends from related series across categorical dimensions.
An unsupervised cluster-based feature grouping model for early diabetes detection
TL;DR: In this paper , an unsupervised cluster-based feature grouping model was proposed for early diabetes identification using an open-source dataset containing the data of 520 diabetic patients, where the dataset and groupings of the features using the elbow and silhouette methods have been clustered using K-means.
53
•Journal Article
Subjective interestingness in exploratory data mining
TL;DR: In this paper, a general mathematical framework for formalizing interestingness in a subjective manner is presented, which can be successfully instantiated for a variety of exploratory data mining problems.
50
References
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni,Salim Roukos,Todd Ward,Wei-Jing Zhu +3 more
- 06 Jul 2002
TL;DR: This paper proposed a method of automatic machine translation evaluation that is quick, inexpensive, and language-independent, that correlates highly with human evaluation, and that has little marginal cost per run.
•Posted Content
Improved Techniques for Training GANs
TL;DR: In this article, the authors present a variety of new architectural features and training procedures that apply to the generative adversarial networks (GANs) framework and achieve state-of-the-art results in semi-supervised classification on MNIST, CIFAR-10 and SVHN.
7.4K
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam,C. Lawrence Zitnick,Devi Parikh +2 more
- 07 Jun 2015
TL;DR: A novel paradigm for evaluating image descriptions that uses human consensus is proposed and a new automated metric that captures human judgment of consensus better than existing metrics across sentences generated by various sources is evaluated.
•Proceedings Article
Improved techniques for training GANs
Tim Salimans,Ian Goodfellow,Wojciech Zaremba,Vicki Cheung,Alec Radford,Xi Chen +5 more
- 05 Dec 2016
TL;DR: In this article, a variety of new architectural features and training procedures are applied to the generative adversarial networks (GANs) framework and achieved state-of-the-art results in semi-supervised classification on MNIST, CIFAR-10 and SVHN.
Interestingness measures for data mining: A survey
Liqiang Geng,Howard J. Hamilton +1 more
TL;DR: This survey reviews the interestingness measures for rules and summaries, classifies them from several perspectives, compares their properties, identifies their roles in the data mining process, gives strategies for selecting appropriate measures for applications, and identifies opportunities for future research in this area.
1.3K
Related Papers (5)
Zhenhui Li,Huaxiu Yao,Fenglong Ma +2 more
- 20 Jan 2020
Sanda Dragos,Diana Halita,Christian Sacarea +2 more
- 02 Nov 2015