0

Gradient-Based Latent Decomposition Reveals Mechanisms of Feature Degradation in Weakly Supervised Mammography

Weakly supervised hierarchical models exhibit a persistent asymmetry: coarse lesion-type features are preserved under reconstruction while fine-grained malignancy cues degrade---a pattern with direct consequences for the clinical reliability of breast cancer screening pipelines.

Preview
Year
2026
Hosting
Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2607.24835ARXIV-DEFAULT
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Weakly supervised hierarchical models exhibit a persistent asymmetry: coarse lesion-type features are preserved under reconstruction while fine-grained malignancy cues degrade---a pattern with direct consequences for the clinical reliability of breast cancer screening pipelines. We introduce gradient-based orthogonal latent decomposition for hierarchical Variational Autoencoders (H-VAEs) to mechanistically explain this asymmetry. The latent space is partitioned into a task-aligned component (z_1), shaped by coarse supervisory gradients, and an orthogonal residual (z_{res}) capturing remaining representational capacity. On 3,550 mammographic Regions of Interest (ROIs) from CBIS-DDSM, only \sim4.4% of latent magnitude aligns with supervisory gradients, leaving \sim95.6% in the orthogonal residual upon which fine-grained pathology prediction primarily depends. The model achieves Stage-1 AUC 0.866 and Stage 2 AUC 0.552, with a reconstruction stability gap of Δ_{diag}=5% (p=0.005) and a classification gap of Δ_{AUC}=0.314 (p{<}0.001). Latent ablation confirms that features for both tasks reside heavily in z_{res}, structurally explaining why reconstruction degrades pathology stability disproportionately. Comparisons with Multi-Instance Learning (MIL) and Multi-Task Learning (MTL) confirm generalization across architectures and modalities. These findings reveal that in high-dimensional spaces, a single coarse supervisory signal isolates only a sparse 1D latent direction, forcing critical fine-grained features into the vulnerable residual subspace.