Using simulation studies to evaluate statistical methods
TL;DR: This tutorial provides a structured approach for planning and reporting simulation studies, which involves defining aims, data‐generating mechanisms, estimands, methods, and performance measures (“ADEMP”).
read more
Abstract: Simulation studies are computer experiments that involve creating data by pseudo‐random sampling. A key strength of simulation studies is the ability to understand the behavior of statistical methods because some “truth” (usually some parameter/s of interest) is known from the process of generating the data. This allows us to consider properties of methods, such as bias. While widely used, simulation studies are often poorly designed, analyzed, and reported. This tutorial outlines the rationale for using simulation studies and offers guidance for design, execution, analysis, reporting, and presentation. In particular, this tutorial provides a structured approach for planning and reporting simulation studies, which involves defining aims, data‐generating mechanisms, estimands, methods, and performance measures (“ADEMP”); coherent terminology for simulation studies; guidance on coding simulation studies; a critical discussion of key performance measures and their estimation; guidance on structuring tabular and graphical presentation of results; and new graphical presentations. With a view to describing recent practice, we review 100 articles taken from Volume 34 of Statistics in Medicine, which included at least one simulation study and identify areas for improvement.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Issues in the determination of 'responders' and 'non-responders' in physiological research.
TL;DR: The dichotomization of continuous‐level physiological measurements into ‘responder’ and ‘non‐responders’ when interventions/treatments are examined in robust parallel‐group studies is discussed.
Selection of the Number of Participants in Intensive Longitudinal Studies: A User-Friendly Shiny App and Tutorial for Performing Power Analysis in Multilevel Regression Models That Account for Temporal Dependencies:
Ginette Lafit,Janne Adolf,Egon Dejonckheere,Inez Myin-Germeys,Wolfgang Viechtbauer,Wolfgang Viechtbauer,Eva Ceulemans +6 more
- 23 Mar 2021
TL;DR: In recent years, the popularity of procedures for collecting intensive longitudinal data, such as the experience sampling method, has increased greatly as mentioned in this paper, and the data collected using such designs allow...
107
•Posted Content
State-of-the-art in selection of variables and functional forms in multivariable analysis -- outstanding issues.
Willi Sauerbrei,Aris Perperoglou,Matthias Schmid,Michal Abrahamowicz,Heiko Becher,Harald Binder,Daniela Dunkler,Frank E. Harrell,Patrick Royston,Georg Heinze +9 more
TL;DR: In this paper, the authors identify and illustrate such gaps in the literature and present them at a moderate technical level to the wide community of practitioners, researchers, and students of statistics.
96
Declaring and Diagnosing Research Designs.
TL;DR: This work provides a framework for formally “declaring” the analytically relevant features of a research design in a demonstrably complete manner, with applications to qualitative, quantitative, and mixed methods research.
G-computation, propensity score-based methods, and targeted maximum likelihood estimator for causal inference with different covariates sets: a comparative simulation study.
Arthur Chatton,Florent Le Borgne,Clemence Leyrat,Clemence Leyrat,Florence Gillaizeau,Chloé Rousseau,Chloé Rousseau,Laetitia Barbin,David Laplaud,Maxime Léger,Maxime Léger,Bruno Giraudeau,Bruno Giraudeau,Yohann Foucher +13 more
TL;DR: A simulation study to compare the relative performance results obtained by using four different sets of covariates and four methods, and proposes an R package RISCA to encourage the use of g-computation in causal inference.
References
Inference and missing data
TL;DR: In this article, it was shown that ignoring the process that causes missing data when making sampling distribution inferences about the parameter of the data, θ, is generally appropriate if and only if the missing data are missing at random and the observed data are observed at random, and then such inferences are generally conditional on the observed pattern of missing data.
10K
Small Sample Inference for Fixed Effects from Restricted Maximum Likelihood
TL;DR: A scaled Wald statistic is presented, together with an F approximation to its sampling distribution, that is shown to perform well in a range of small sample settings and has the advantage that it reproduces both the statistics and F distributions in those settings where the latter is exact.
Multiple Imputation After 18+ Years
TL;DR: A description of the assumed context and objectives of multiple imputation is provided, and a review of the multiple imputations framework and its standard results are reviewed.
Information and the Accuracy Attainable in the Estimation of Statistical Parameters
C. Radhakrishna Rao,C. Radhakrishna Rao +1 more
- 01 Jan 1992
TL;DR: The earliest method of estimation of statistical parameters is the method of least squares due to Mark off as discussed by the authors, where a set of observations whose expectations are linear functions of a number of unknown parameters being given, the problem which Markoff posed for solution is to find out a linear function of observations, whose expectation is an assigned linear function for the unknown parameters and whose variance is a minimum.
2.2K