1. What are the contributions mentioned in the paper "Good methods for coping with missing data in decision trees" ?
The authors propose a simple and effective method for dealing with missing data in decision trees used for classification.. The authors show through a substantial data-based study of classification accuracy that MIA exhibits consistently good performance across a broad range of data types and of sources and amounts of missingness.
read more
2. What are the missingness mechanisms used in the study?
These were: deletion of instances with any missing data; Shapiro's decision tree single imputation technique (Quinlan, 1993); maximum likelihood imputation of data for both continuous (assuming multivariate normality) and categorical data via the EM algorithm as developed by Schafer (1997), and considered in both single and (five replicate) multiple imputation (EMMI) forms; mean or mode single imputation; fractional cases (FC; Cestnik et al., 1987, Quinlan, 1993); and surrogate variable splitting (Breiman et al., 1984, Therneau and Atkinson, 1997).
read more
3. What are the missing data mechanisms used in the study?
Three missing data mechanisms were employed, coming under the headings (Little and Rubin, 2002) of: missing completely at random (MCAR); missing at random (MAR), under which the probability of being missing depends on the value of another, non-missing, attribute; and informative missingness (IM), under which the probability of being missing depends on the actual (but in non-simulation practice unobserved) value of the attribute itself.
read more
4. What are the three missingness mechanisms used to cope with missing data?
All are complete datasets into which missingness is artificially introduced at rates of 15%, 30% and 50% into either the single attribute which is most highly correlated with class or else evenly distributed across all attributes.
read more


