Journal Article10.1109/issre59848.2023.00036
Mind the Gap: The Difference Between Coverage and Mutation Score Can Guide Testing Efforts
Kush Jain,Goutamkumar Tulajappa Kalburgi,Claire Le Goues,Alex Groce +3 more
- 09 Oct 2023
2
TL;DR: The oracle gap between coverage and mutation score can guide testing efforts by identifying source files where weak oracle tests important code.
read more
Abstract: An "adequate" test suite should effectively find all inconsistencies between a system's requirements/specifications and its implementation. Practitioners frequently use code coverage to approximate adequacy, while academics argue that mutation score may better approximate true (oracular) adequacy coverage. High code coverage is increasingly attainable even on large systems via automatic test generation, including fuzzing. In light of all of these options for measuring and improving testing effort, how should a QA engineer spend their time? We propose a new framework for reasoning about the extent, limits, and nature of a given testing effort based on an idea we call the oracle gap, or the difference between source code coverage and mutation score for a given software element. We conduct (1) a large-scale observational study of the oracle gap across popular Maven projects, (2) a study that varies testing and oracle quality across several of those projects and (3) a small-scale observational study of highly critical, well-tested code across comparable blockchain projects. We show that the oracle gap surfaces important information about the extent and quality of a test effort beyond either adequacy metric alone. In particular, it provides a way for practitioners to identify source files where it is likely a weak oracle tests important code.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Combining TSL and LLM to Automate REST API Testing: A Comparative Study
Thiago Barradas,Aline Paes,Vânia de Oliveira Neves,Thiago Barradas,Aline Paes,Vânia de Oliveira Neves +5 more
- 22 Sep 2025
Abstract: The effective execution of tests for REST APIs remains a considerable challenge for development teams, driven by the inherent complexity of distributed systems, the multitude of possible scenarios, and the limited time available for test design. Exhaustive testing of all input combinations is impractical, often resulting in undetected failures, high manual effort, and limited test coverage. To address these issues, we introduce RestTSLLM, an approach that uses Test Specification Language (TSL) in conjunction with Large Language Models (LLMs) to automate the generation of test cases for REST APIs. The approach targets two core challenges: the creation of test scenarios and the definition of appropriate input data. The proposed solution integrates prompt engineering techniques with an automated pipeline to evaluate various LLMs on their ability to generate tests from OpenAPI specifications. The evaluation focused on metrics such as success rate, test coverage, and mutation score, enabling a systematic comparison of model performance. The results indicate that the best-performing LLMs – Claude 3.5 Sonnet (Anthropic), Deepseek R1 (Deepseek), Qwen 2.5 32b (Alibaba), and Sabiá 3 (Maritaca) – consistently produced robust and contextually coherent REST API tests. Among them, Claude 3.5 Sonnet outperformed all other models across every metric, emerging in this study as the most suitable model for this task. These findings highlight the potential of LLMs to automate the generation of tests based on API specifications.
Towards differential fuzzing to reduce manual efforts to identify equivalent mutants: A preliminary study
Brayan Styben Diaz Garcia,Márcio Eduardo Delamaro,Simone R. S. Souza +2 more
- 30 Sep 2024
TL;DR: This study explores differential fuzzing to reduce manual efforts in identifying equivalent mutants in mutation testing, achieving 97% accuracy with a 3-minute timeout, suggesting a viable approach to improve mutation testing efficiency.
References
Hints on Test Data Selection: Help for the Practicing Programmer
TL;DR: In many cases tests of a program that uncover simple errors are also effective in uncovering much more complex errors, so-called coupling effect can be used to save work during the testing process.
2.2K
Are mutants a valid substitute for real faults in software testing
René Just,Darioush Jalali,Laura Inozemtseva,Michael D. Ernst,Reid Holmes,Gordon Fraser +5 more
- 11 Nov 2014
TL;DR: This paper investigates whether mutants are indeed a valid substitute for real faults, i.e., whether a test suite’s ability to detect mutants is correlated with its able to detect real faults that developers have fixed, and shows a statistically significant correlation between mutant detection and real fault detection, independently of code coverage.
Coverage is not strongly correlated with test suite effectiveness
Laura Inozemtseva,Reid Holmes +1 more
- 31 May 2014
TL;DR: It is found that there is a low to moderate correlation between coverage and effectiveness when the number of test cases in the suite is controlled for, and that stronger forms of coverage do not provide greater insight into the effectiveness of the suite.
464
Do Automatically Generated Unit Tests Find Real Faults? An Empirical Study of Effectiveness and Challenges (T)
Sina Shamshiri,René Just,José Miguel Rojas,Gordon Fraser,Phil McMinn,Andrea Arcuri +5 more
- 09 Nov 2015
TL;DR: Three state-of-the-art unit test generation tools for Java (Randoop, EvoSuite, and Agitar) are applied to the 357 real faults in the Defects4J dataset and investigated how well the generated test suites perform at detecting these faults.
PIT: a practical mutation testing tool for Java (demo)
Henry Coles,Thomas Laurent,Christopher Henard,Mike Papadakis,Anthony Ventresque +4 more
- 18 Jul 2016
TL;DR: PIT is a practical mutation testing tool for Java, applicable on real-world codebases and robust and well integrated with development tools, as it can be invoked through a command line interface, Ant or Maven.