Open AccessPosted Content
Selection via Proxy: Efficient Data Selection for Deep Learning
Cody Coleman,Christopher Yeh,Stephen Mussmann,Baharan Mirzasoleiman,Peter Bailis,Percy Liang,Jure Leskovec,Matei Zaharia +7 more
TL;DR: This work shows that it can significantly improve the computational efficiency of data selection in deep learning by using a much smaller proxy model to perform data selection for tasks that will eventually require a large target model (e.g., selecting data points to label for active learning).
read more
Abstract: Data selection methods, such as active learning and core-set selection, are useful tools for machine learning on large datasets. However, they can be prohibitively expensive to apply in deep learning because they depend on feature representations that need to be learned. In this work, we show that we can greatly improve the computational efficiency by using a small proxy model to perform data selection (e.g., selecting data points to label for active learning). By removing hidden layers from the target model, using smaller architectures, and training for fewer epochs, we create proxies that are an order of magnitude faster to train. Although these small proxy models have higher error rates, we find that they empirically provide useful signals for data selection. We evaluate this "selection via proxy" (SVP) approach on several data selection tasks across five datasets: CIFAR10, CIFAR100, ImageNet, Amazon Review Polarity, and Amazon Review Full. For active learning, applying SVP can give an order of magnitude improvement in data selection runtime (i.e., the time it takes to repeatedly train and select points) without significantly increasing the final error (often within 0.1%). For core-set selection on CIFAR10, proxies that are over 10x faster to train than their larger, more accurate targets can remove up to 50% of the data without harming the final accuracy of the target, leading to a 1.6x end-to-end training time improvement.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
DataComp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre,Gabriel Ilharco,Alex Fang,Jonathan Hayase,Georgios Smyrnis,T. Nguyen,Ryan Marten,Mitchell Wortsman,Dhruba Ghosh,Jieyu Zhang,Rahim Entezari,Giannis Daras,Sarah I. Pratt,Vivek Ramanujan,Yonatan Bitton,Kalyani Marathe,Stephen Mussmann,Richard Vencu,Mehdi Cherti,Ranjay Krishna,Pang Wei Koh,Olga Saukh,Alexander Ratner,Shuran Song,Hannaneh Hajishirzi,Ali Farhadi,Romain Beaumont,Sewoong Oh,Alexandros G. Dimakis,Jenia Jitsev,Yair Carmon,Vaishaal Shankar,Ludwig Schmidt +32 more
TL;DR: DataComp as mentioned in this paper is a testbed for dataset experiments centered around a new candidate pool of 12.8 billion image-text pairs from Common Crawl, which can be used to design new filtering techniques or curate new data sources and then evaluate their new dataset by running our standardized CLIP training code and testing the resulting model on 38 downstream test sets.
193
DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
Michael Xie,Hieu Quang Pham,Xuanyi Dong,Nan Du,Hanxiao Liu,Yifeng Lu,Percy Liang,Quoc Le,Tengyu Ma,Adams Wei Yu +9 more
TL;DR: DoReMi as mentioned in this paper uses group distributionally robust optimization (Group DRO) over domains to produce domain weights (mixture proportions) without knowledge of downstream tasks, and then resamples a dataset with these domain weights and train a larger, full-sized model.
84
Data Selection for Language Models via Importance Resampling
TL;DR: Data Selection with Importance Resampling (DSIR) is proposed, an efficient and scalable framework that estimates importance weights in a reduced feature space for tractability and selects data with importance resampling according to these weights.
DataPerf: Benchmarks for Data-Centric AI Development
Mark Mazumder,Colby R. Banbury,Xiaozhe Yao,Bojan Karlas,W. G. Rojas,Sudnya Diamos,Greg Diamos,Lynn He,Douwe Kiela,David Jurado,David Kanter,Rafael Mosquera,Juan Camilo Galvis Ciro,Lora Aroyo,Bilge Acun,Sabri Eyuboglu,Amirata Ghorbani,Emmett D. Goodman,Tariq Kane,Christine Kirkpatrick,Tzu-Sheng Kuo,Jonas Mueller,Tristan Thrush,Joaquin Vanschoren,Margaret J. Warren,Adina Williams,Serena Yeung,Newsha Ardalani,Praveen Paritosh,Ce Zhang,James Zou,Carole-Jean Wu,Cody Coleman,Andrew Y. Ng,Peter Mattson,Vijay Janapa Reddi +35 more
TL;DR: DataPerf is presented, a benchmark package for evaluating ML datasets and dataset-working algorithms to enable the “data ratchet,” in which training sets will aid in evaluating test sets on the same problems, and vice versa, to generate a virtuous loop that will accelerate development of data-centric AI.
A neural network boosting regression model based on XGBoost
TL;DR: In this article , a Neural Network Boosting (NNBoost) regression model is proposed, which takes shallow neural networks with simple structures as weak classifiers and obtains low regression errors on several data sets.
68
References
•Posted Content
On the Relationship between Data Efficiency and Error for Uncertainty Sampling
Stephen Mussmann,Percy Liang +1 more
TL;DR: An answer for logistic regression with the popular active learning algorithm, uncertainty sampling is provided, showing that for a variant of uncertainty sampling, the asymptotic data efficiency is within a constant factor of the inverse error rate of the limiting classifier.
•Book
Active Learning
Burr Settles
- 01 Jul 2012
TL;DR: Active learning as discussed by the authors is a general approach that allows a machine learning algorithm to choose the data from which it learns by posing "queries", usually in the form of unlabeled data instances to be labeled by an oracle (e.g., a human annotator) that already understands the nature of the problem.
•Proceedings Article
Weight Agnostic Neural Networks
Adam Gaier,David Ha +1 more
- 12 Jun 2019
TL;DR: In this paper, the authors propose a search method for neural network architectures that can already perform a task without any explicit weight training. But how important are the weight parameters of a neural network compared to its architecture, they question to what extent neural network architecture alone, without learning any weight parameters, can encode solutions for a given task.
Unsupervised data selection and word-morph mixed language model for tamil low-resource keyword search
Chongjia Ni,Cheung-Chi Leung,Lei Wang,Nancy F. Chen,Bin Ma +4 more
- 19 Apr 2015
TL;DR: Gaussian component index based n-grams as acoustic features in a submodular function for unsupervised data selection provides a near-optimal solution in terms of the objective being optimized and increases the vocabulary coverage, implicitly alleviating the OOV problem.