Journal Article10.48550/arXiv.2212.05129
Measuring Data
Margaret Mitchell,Alexandra Luccioni,Nathan Lambert,Marissa Gerchick,Angelina McMillan-Major,Ezinwanne Ozoani,Nazneen Fatema Rajani,Tristan Thrush,Yacine Jernite,Douwe Kiela +9 more
TL;DR: Measuring data aids in systematically building and analyzing machine learning (ML) data towards specific goals and gaining better control of what modern ML systems will learn as discussed by the authors , which is a critical component of responsible AI development.
read more
Abstract: We identify the task of measuring data to quantitatively characterize the composition of machine learning data and datasets. Similar to an object's height, width, and volume, data measurements quantify different attributes of data along common dimensions that support comparison. Several lines of research have proposed what we refer to as measurements, with differing terminology; we bring some of this work together, particularly in fields of computer vision and language, and build from it to motivate measuring data as a critical component of responsible AI development. Measuring data aids in systematically building and analyzing machine learning (ML) data towards specific goals and gaining better control of what modern ML systems will learn. We conclude with a discussion of the many avenues of future work, the limitations of data measurements, and how to leverage these measurement approaches in research and practice.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Figures

Figure 1: Ancient Egyptian Cubit Rod for measuring length. One of the earliest known objects for standardized measurement. Turin Museum, Wikimedia Commons, CC BY-SA 3.0 
Figure 2: S. S. Stevens’ measurement scales, from the seminal On the Theory of Scales of Measurement. (Stevens, 1946) 
Table 1: Examples of different data measurements proposed in image- and language-based data science and machine learning, alongside analogs in the physical sciences.
Citations
Considerations for Differentially Private Learning with Large-Scale Public Pretraining
TL;DR: In this paper , the use of large Web-scraped datasets should be viewed as differential-privacy-preserving, but the authors point out that publicizing these models pretrained on Web data as"private" could lead to harm and erode the public's trust in differential privacy as a meaningful definition of privacy.
42
The ROOTS Search Tool: Data Transparency for LLMs
Aleksandra Piktus,Christopher Akiki,Paulo Villegas,Hugo Laurenccon,Gérard Dupont,Alexandra Luccioni,Yacine Jernite,Anna Rogers +7 more
- 27 Feb 2023
TL;DR: The ROOTS Search Tool as discussed by the authors is a search engine over the entire ROOTS corpus offering both fuzzy and exact search capabilities, and it is open-sourced and available on Hugging Face Spaces: https://huggingface.co/spaces/bigscience-data/rootsearch.
17
Modern Business Intelligence: Big Data Analytics and Artificial Intelligence for Creating the Data-Driven Value
Ahmed A.A. Gad-Elrab
- 19 May 2021
TL;DR: The importance of big data analytics, data mining, AI for building and enhancing modern BI will be introduced and discussed, and challenges and opportunities for creating value of data by establishing modern BI processes are discussed.
Hype, Sustainability, and the Price of the Bigger-is-Better Paradigm in AI
Gaël Varoquaux,Alexandra Luccioni,Meredith Whittaker +2 more
TL;DR: This paper critiques the "bigger-is-better" AI paradigm, arguing that it's scientifically fragile, unsustainable, and exacerbates power concentration, prioritizing compute-intensive tasks over important applications like health, education, and climate, and threatening to disempower marginalized groups.
On Catastrophic Inheritance of Large Foundation Models
Hao Chen,Bhiksha Raj,Xing Xie,Jindong Wang +3 more
TL;DR: The challenges behind this issue are discussed and UIM, a framework to Understand the catastrophic inheritance of LFMs from both pre-training and downstream adaptation, is proposed, to Interpret the implications of catastrophic inheritance on downstream tasks, and how to Mitigate it.
7
References
•Posted Content
Decision-Making with Auto-Encoding Variational Bayes
TL;DR: This work describes the error of importance sampling as a function of posterior variance and shows that proposal distributions learned with evidence upper bounds are better than the current state of the art.
•Posted Content
Improved Techniques for Training GANs
TL;DR: In this article, the authors present a variety of new architectural features and training procedures that apply to the generative adversarial networks (GANs) framework and achieve state-of-the-art results in semi-supervised classification on MNIST, CIFAR-10 and SVHN.
7.4K
•Posted Content
On Information and Sufficiency
TL;DR: The information deviation between any two finite measures cannot be increased by any statistical operations (Markov morphisms) and is invarient if and only if the morphism is sufficient for these two measures as mentioned in this paper.
7.3K
Diversity and Evenness: A Unifying Notation and Its Consequences
TL;DR: Three commonly used measures of diversity, Simpson's index, Shannon's entropy, and the total number of species, are related to Renyi's definition of a generalized entropy, according to which there is a continuum of possible diversity measures.
5.9K
The Earth Mover's Distance as a Metric for Image Retrieval
TL;DR: This paper investigates the properties of a metric between two distributions, the Earth Mover's Distance (EMD), for content-based image retrieval, and compares the retrieval performance of the EMD with that of other distances.