Exploration
Missing Data Imputation
8 sections
See worked example →01What Is Missing-Data Imputation?
Missing-data imputation is a modeling step that creates plausible values for unobserved entries so the final analysis can use available information while representing uncertainty. It is not the recovery of a hidden true value.
The central question is: “What is missing, why is it missing, and how much can observed information support plausible replacements and their uncertainty?” Defining it before opening a menu prevents a page of statistics from replacing an explanation.
The purpose is the missingness process and imputation uncertainty rather than one supposedly recovered value. Interpret missingness patterns, observed-versus-imputed distributions, convergence, and FMI together, because an imputation model poorer than the analysis model can create precise-looking bias.
02Why Imputation Matters
MCAR means missingness is unrelated to observed or unobserved values; MAR allows missingness to depend on observed data; MNAR retains dependence on the unseen value. Multiple imputation generates several completed datasets and combines estimates with Rubin’s rules.
No single statistic is sufficient. Compare observed and imputed distributions, check chained-equation convergence, inspect between-imputation variation and fraction of missing information, and include the final analysis variables and relevant auxiliary predictors.
The purpose is the missingness process and imputation uncertainty rather than one supposedly recovered value. Interpret missingness patterns, observed-versus-imputed distributions, convergence, and FMI together, because an imputation model poorer than the analysis model can create precise-looking bias.
Concept visual
Read the core structure in one view
the missingness process and imputation uncertainty rather than one supposedly recovered value
Missingness map
Pattern summary
Core question
Observed data → m plausible datasets → pooled estimate
Reading order
- 1. Identify cases and denominator
- 2. Connect estimates to shape
- 3. Inspect missing and unusual patterns
03Missingness Mechanisms
MCAR means missingness is unrelated to observed or unobserved values; MAR allows missingness to depend on observed data; MNAR retains dependence on the unseen value. Multiple imputation generates several completed datasets and combines estimates with Rubin’s rules.
No single statistic is sufficient. Compare observed and imputed distributions, check chained-equation convergence, inspect between-imputation variation and fraction of missing information, and include the final analysis variables and relevant auxiliary predictors.
The purpose is the missingness process and imputation uncertainty rather than one supposedly recovered value. Interpret missingness patterns, observed-versus-imputed distributions, convergence, and FMI together, because an imputation model poorer than the analysis model can create precise-looking bias.
04Major Imputation Methods
MCAR means missingness is unrelated to observed or unobserved values; MAR allows missingness to depend on observed data; MNAR retains dependence on the unseen value. Multiple imputation generates several completed datasets and combines estimates with Rubin’s rules.
No single statistic is sufficient. Compare observed and imputed distributions, check chained-equation convergence, inspect between-imputation variation and fraction of missing information, and include the final analysis variables and relevant auxiliary predictors.
The purpose is the missingness process and imputation uncertainty rather than one supposedly recovered value. Interpret missingness patterns, observed-versus-imputed distributions, convergence, and FMI together, because an imputation model poorer than the analysis model can create precise-looking bias.
Analysis workflow
From raw values to an explainable result
Record analytical choices before calculation and diagnostics afterward.
Define variables
Audit quality
Choose summaries
Compute and plot
Report in context
Scale, units, missingness
the MCAR/MAR/MNAR assumption, imputation model, auxiliary variables, and number of imputations
Estimate, plot, diagnostic
05Imputation Workflow
A defensible workflow is Map missingness → State mechanism → Specify imputation → Generate and diagnose → Pool estimates. Record the variable definitions, exclusions, transformations, denominators, and decision rules. Reproducibility begins with these choices, not with the final number.
Begin by auditing type, unit, coding, missingness, and the MCAR/MAR/MNAR assumption, imputation model, auxiliary variables, and number of imputations. Defaults are only starting points; document every denominator, exclusion, transformation, and decision rule needed to reproduce the result.
06Applications
Common uses include survey nonresponse, clinical loss to follow-up, administrative and panel records, preparation for predictive modeling. Before operational use, define how the same quantity will be recomputed for new data, how subgroup differences will be monitored, and what action the result is meant to support.
The purpose is the missingness process and imputation uncertainty rather than one supposedly recovered value. Interpret missingness patterns, observed-versus-imputed distributions, convergence, and FMI together, because an imputation model poorer than the analysis model can create precise-looking bias.
Assumptions & diagnostics
Convergence, distributions, uncertainty
an imputation model poorer than the analysis model can create precise-looking bias
Iter 5
Iter 10
Iter 20
Pooled
Checks
Cases and denominator
Alternative calculation
Distorting observations
07Strengths and Limitations
Compare observed and imputed distributions, check chained-equation convergence, inspect between-imputation variation and fraction of missing information, and include the final analysis variables and relevant auxiliary predictors.
Strengths include Preserves more information than complete-case deletion; Carries imputation uncertainty into standard errors; Handles complex patterns and mixed variable types. Important limitations are MAR and MNAR cannot be proven from observed data alone; A misspecified imputation model can create precise bias; Single mean or mode imputation understates variance and relationships.
Diagnostics are part of the result because an imputation model poorer than the analysis model can create precise-looking bias. Compare reasonable alternatives and inspect the observations that drive the summary before treating one output as stable.
Result preview
Primary result and interpretation evidence
The display keeps sample, denominator, shape, and uncertainty together.
Missingness map
Pattern summary
N
240Missing
18.3%FMI
.1208Key Takeaways
Imputation is uncertainty modeling, not blank filling. Report the missingness assumption, imputation model, diagnostics, and sensitivity analyses. Recommended sequence: Map missingness → State mechanism → Specify imputation → Generate and diagnose → Pool estimates.
A complete report links the question, analysis population, the MCAR/MAR/MNAR assumption, imputation model, auxiliary variables, and number of imputations, the primary result, and the diagnostic evidence. A reader should be able to reconstruct the same quantity from those decisions.