0

Differentially Private Distributed Inference for Multicenter Clinical Studies

Extracting reliable conclusions from data distributed across institutions is a core problem in healthcare: pooling patient records would improve inference, but privacy regulations and the lack of a trusted central authority frequently delay multicenter studies.

Year
2024
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2402.08156CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Extracting reliable conclusions from data distributed across institutions is a core problem in healthcare: pooling patient records would improve inference, but privacy regulations and the lack of a trusted central authority frequently delay multicenter studies. We develop a framework for differentially private distributed inference in which institutions repeatedly exchange log belief-ratio statistics subject to differential privacy (DP). With arithmetic and geometric averaging of beliefs, we control the false-negative and false-positive rates as functions of the privacy budget, communication rounds, and statistical separation between hypotheses, exposing a three-way trade-off among accuracy, communication, and privacy. We derive finite-sample bounds on the Type I and Type II error probabilities that allow distributed hypothesis testing at a target significance level, and show that the Laplace mechanism minimizes convergence time subject to DP. For distributed online learning from data streams (e.g., epidemiological surveillance or rolling recruitment), privacy noise vanishes asymptotically and online learning admits similar finite-sample guarantees. On simulated multicenter survival analyses using the AIDS Clinical Trials Group and an advanced-cancer cohort, and a simulated genetic association study over New York City hospitals, our method approaches the non-private baseline at a small privacy budget between 1 and 10, runs 10x to 1000x faster than homomorphic-encryption methods, and incurs up to 100x lower error than first-order private optimization methods. Finally, the level of aggregation is a primary design choice: federating at the organizational rather than the hospital level strengthens privacy, raises statistical power, and lowers communication and administrative burden, so data should be pooled within organizations before setting up federated analytics.