0

Towards Efficient Inference under Nonmonotone Missingness with General Imputation

Missing data are ubiquitous in classical survey and longitudinal studies as well as modern multi-modality data analysis. A longstanding challenge arises under nonmonotone missingness, where different units may observe arbitrary subsets of all variables.

Preview
Year
2025
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2509.24158CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Missing data are ubiquitous in classical survey and longitudinal studies as well as modern multi-modality data analysis. A longstanding challenge arises under nonmonotone missingness, where different units may observe arbitrary subsets of all variables. We study parameter estimation and inference problem under this setting. Semiparametric efficiency theory characterizes the efficient estimator through inversion of an operator constructed from pattern-specific conditional expectations. However, this estimator is generally not tractable due to compositions of conditional expectations across patterns. We introduce the Restricted ANOVA hierarchY (RAY), a functional decomposition that reveals an almost-eigen structure of the operator under missing completely at random. This structure yields a closed-form, computable approximation to the efficient estimator. RAY estimator is applicable to general Z-estimation problems, and it remains unbiased for arbitrary independent imputation functions. In theory, we establish verifiable sufficient conditions where RAY attains the efficiency lower bound, and offer a general bound for the efficiency gap otherwise. We further develop adaptive RAY estimator, which attains the minimal asymptotic variance within a broader class containing RAY and other existing estimators. Finally, we investigate the extension of RAY under missing at random mechanism. Simulations and a single-cell multi-omics application demonstrate the efficiency gains of the proposed estimators.