Partially labelled samples arise when features are observed for all data, but class labels are available for only a subset. In such settings, the mechanism governing label availability may itself contain information relevant to classification, yet it is typically left unmodelled in standard semi-supervised learning procedures. The SSLfmm package implements likelihood-based Gaussian finite-mixture classification in which the label-missingness process is modelled jointly with the class distribution. It supports complete-case, missing completely at random (MCAR), entropy-based missing at random (MAR), and mixed MCAR/MAR analyses. For the mixed mechanism, the source of a missing label may be observed or latent, allowing the same modelling framework to accommodate different forms of information about label availability. A common R interface is provided for model fitting, prediction, performance assessment, simulation, and entropy-based diagnostics. We describe the statistical formulation and software implementation, position SSLfmm relative to existing finite-mixture and semi-supervised learning software, and demonstrate its use through a reproducible simulation comparing observed- and latent-source analyses. A semi-synthetic application to the Blood Transfusion data further illustrates how alternative assumptions about label missingness can be fitted, compared, and diagnosed in practice.
SSLfmm: An R Package for Semi-Supervised Learning with Mixed Missingness
Partially labelled samples arise when features are observed for all data, but class labels are available for only a subset. In such settings, the mechanism governing label availability may itself contain information relevant to classification, yet it is typically left unmodelled…
- Preview

- Year
- 2025
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2512.03322CC-BY-4.0
- TL;DR
- Semantic Scholar