Curating GitHub for engineered software projects

doi:10.1007/S10664-017-9512-6

Journal Article10.1007/S10664-017-9512-6

Curating GitHub for engineered software projects

Nuthan Munaiah, +3 more

- 01 Dec 2017

- Empirical Software Engineering

- Vol. 22, Iss: 6, pp 3219-3253

407

TL;DR: This work proposes a framework, and presents a reference implementation of the framework as a tool called reaper, to enable researchers to select GitHub repositories that contain evidence of an engineered software project and identifies software engineering practices (called dimensions) and proposes means for validating their existence in a GitHub repository.

Abstract: Software forges like GitHub host millions of repositories. Software engineering researchers have been able to take advantage of such a large corpora of potential study subjects with the help of tools like GHTorrent and Boa. However, the simplicity in querying comes with a caveat: there are limited means of separating the signal (e.g. repositories containing engineered software projects) from the noise (e.g. repositories containing home work assignments). The proportion of noise in a random sample of repositories could skew the study and may lead to researchers reaching unrealistic, potentially inaccurate, conclusions. We argue that it is imperative to have the ability to sieve out the noise in such large repository forges. We propose a framework, and present a reference implementation of the framework as a tool called reaper, to enable researchers to select GitHub repositories that contain evidence of an engineered software project. We identify software engineering practices (called dimensions) and propose means for validating their existence in a GitHub repository. We used reaper to measure the dimensions of 1,857,423 GitHub repositories. We then used manually classified data sets of repositories to train classifiers capable of predicting if a given GitHub repository contains an engineered software project. The performance of the classifiers was evaluated using a set of 200 repositories with known ground truth classification. We also compared the performance of the classifiers to other approaches to classification (e.g. number of GitHub Stargazers) and found our classifiers to outperform existing approaches. We found stargazers-based classifier (with 10 as the threshold for number of stargazers) to exhibit high precision (97%) but an inversely proportional recall (32%). On the other hand, our best classifier exhibited a high precision (82%) and a high recall (86%). The stargazer-based criteria offers precision but fails to recall a significant portion of the population.

Chat with Paper

AI Agents for this Paper

Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps

Citations

•Journal Article•10.1109/MSP.2010.936725

Canonical Correlation Analysis for Data Fusion and Group Inferences

Nicolle M. Correa, +3 more

- 14 Jun 2010

- IEEE Signal Processing Magazine

TL;DR: Two CCA-based approaches for data fusion and group analysis of biomedical imaging data and their utility on fMRI, sMRI, and EEG data are presented and it is important to note that both approaches provide complementary perspectives, and hence it is beneficial to study the data using different analysis techniques.

...read moreread less

447

•Journal Article•10.1145/3241743

The ABC of Software Engineering Research

Klaas-Jan Stol, +1 more

- 17 Sep 2018

- ACM Transactions on Software Engineering...

TL;DR: A taxonomy from the social sciences is adopted, termed here the ABC framework for SE research, which offers a holistic view of eight archetypal research strategies, and six ways in which the framework can advance SE research.

...read moreread less

270

•Posted Content•10.1145/1122445.1122456

Reaction or Speculation: Building Computational Support for Users in Catching-Up Series Based on an Emerging Media Consumption Phenomenon

Riku Arakawa, +1 more

- 12 Feb 2021

- arXiv: Human-Computer Interaction

TL;DR: In this paper, a series of studies were conducted to understand how people engage with speculation during media consumption and designed two prototypes for supporting catching-up users based on their quantitative analysis of Twitter data in regard to reaction-and speculation-based media consumption.

...read moreread less

268

Proceedings Article•10.1109/MSR.2019.00077

A large-scale study about quality and reproducibility of jupyter notebooks

João Felipe Pimentel, +3 more

- 26 May 2019

TL;DR: To understand good and bad practices used in the development of real notebooks, 1.4 million notebooks from GitHub are studied and a detailed analysis of their characteristics that impact reproducibility is presented.

...read moreread less

253

•Journal Article•10.1016/J.JSS.2018.09.016

What’s in a GitHub Star? Understanding Repository Starring Practices in a Social Coding Platform

Hudson Borges, +1 more

- 01 Dec 2018

- Journal of Systems and Software

TL;DR: A throughout study on the meaning, characteristics, and dynamic growth of GitHub stars is provided and a list of recommendations to open source project managers and GitHub users and Software Engineering researchers is provided.

...read moreread less

247

...

Expand

References

•Journal Article•10.1023/A:1010933404324

Random Forests

Leo Breiman

- 01 Oct 2001

TL;DR: Internal estimates monitor error, strength, and correlation and these are used to show the response to increasing the number of features used in the forest, and are also applicable to regression.

...read moreread less

113.1K

•Journal Article•10.1145/267580.267590

Software unit test coverage and adequacy

Hong Zhu, +2 more

- 01 Dec 1997

- ACM Computing Surveys

TL;DR: The notion of adequacy criteria is examined together with its role in software dynamic testing and the methods for comparison and assessment of criteria are reviewed.

...read moreread less

1.4K

•Proceedings Article•10.1145/2597073.2597074

The promises and perils of mining GitHub

Eirini Kalliamvakou, +5 more

- 16 May 2009

TL;DR: It is shown, for example, that the majority of the projects are personal and inactive; that GitHub is also being used for free storage and as a Web hosting service; and that almost 40% of all pull requests do not appear as merged, even though they were.

...read moreread less

979

Journal Article•10.1109/32.895984

Does code decay? Assessing the evidence from change management data

Stephen G. Eick, +4 more

- 01 Jan 2001

- IEEE Transactions on Software Engineerin...

TL;DR: This work defines code decay and proposes a number of measurements (code decay indices) on software and on the organizations that produce it, that serve as symptoms, risk factors, and predictors of decay.

...read moreread less

729

Proceedings Article•10.1145/337180.337209

A case study of open source software development: the Apache server

Audris Mockus, +2 more

- 01 Jun 2000

TL;DR: This analysis of the development process of the Apache web server reveals a unique process, which performs well on important measures, and concludes that hybrid forms of development that borrow the most effective techniques from both the OSS and commercial worlds may lead to high performance software processes.

...read moreread less

669

...

Expand

Curating GitHub for engineered software projects

Chat with Paper

AI Agents for this Paper

Citations

Canonical Correlation Analysis for Data Fusion and Group Inferences

The ABC of Software Engineering Research

Reaction or Speculation: Building Computational Support for Users in Catching-Up Series Based on an Emerging Media Consumption Phenomenon

A large-scale study about quality and reproducibility of jupyter notebooks

What’s in a GitHub Star? Understanding Repository Starring Practices in a Social Coding Platform

References

Random Forests

Software unit test coverage and adequacy

The promises and perils of mining GitHub

Does code decay? Assessing the evidence from change management data

A case study of open source software development: the Apache server

Related Papers (5)

The promises and perils of mining GitHub

Refactoring: Improving the Design of Existing Code

An exploratory study of the pull-based software development model

A Coefficient of agreement for nominal Scales

The measurement of observer agreement for categorical data