0

Preserving Clusters in Error-Bounded Lossy Compression of Scientific Particle Data

Scientific particle simulations in cosmology, molecular dynamics, and fluid dynamics produce large-scale datasets whose storage, movement, and analysis increasingly rely on lossy compression.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2604.18801CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Scientific particle simulations in cosmology, molecular dynamics, and fluid dynamics produce large-scale datasets whose storage, movement, and analysis increasingly rely on lossy compression. However, existing compressors typically bound only pointwise position errors, providing no guarantee on the fidelity of structures derived from particle coordinates, such as single-linkage clustering (also known as Friends-of-Friends algorithm), where clusters are connected components of a proximity graph formed by linking particle pairs within a distance threshold. Even small coordinate perturbations near this threshold can break true links or create false links, thereby splitting or merging entire clusters. We propose a compressor-independent correction technique for preserving single-linkage cluster membership under lossy compression. Our method operates on reconstructed outputs from off-the-shelf compressors such as SZ3, ZFP, Draco, and LCP, and stores a compact corrective edit stream. Our key observation is that cluster-membership queries depend on connected components rather than the complete set of proximity links. Based on this observation, we introduce three constraint-selection modes, vulnerable-pair, safe-component, and halo-forest, that progressively reduce the constraints enforced during correction. Projected gradient descent then corrects the reconstructed coordinates to eliminate the selected violations while respecting the original pointwise error bound. Experiments on cosmology, molecular dynamics, and fluid dynamics datasets with single-GPU and distributed-memory implementations show that our method preserves cluster membership while improving compression ratio by up to 4\times and maintains competitive end-to-end throughput compared to the same base compressors configured with sufficiently tight error bounds to preserve clustering.