0

A General Multimodal Probing Framework for Data Fusion and Prediction with Frozen Large Language Models: Evidence from Multimodal EHR Data

The main objective of this paper is to propose a general framework for prediction based on different sources of multimodal data in the healthcare domain. We evaluated whether frozen medical large language model (LLM) representations can serve as a shared embedding space for…

Preview
Year
2026
Hosting
Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2606.28798ARXIV-DEFAULT
TL;DR
Semantic Scholar
Attribution policy →

Abstract

The main objective of this paper is to propose a general framework for prediction based on different sources of multimodal data in the healthcare domain. We evaluated whether frozen medical large language model (LLM) representations can serve as a shared embedding space for multimodal primary diagnosis category prediction. We propose a backbone-flexible probing framework that serializes structured electronic health record (EHR) variables and leakage-pruned discharge-note sections into a shared frozen-LLM representation space and recovers primary diagnosis categories with linear probes, without fine-tuning the LLM. On a MIMIC-IV cohort of 13,645 admissions (seven diagnosis categories), we probed inputs of different modalities (Structured-only, Unstructured-only, and Combined) across transformer layers and six frozen backbones, compared against XGBoost and three text baselines (CAML, LAAT, PLM-ICD) under a matched split and context budget, and tested robustness across datasets on MIMIC-III. The Combined terminal-layer probe reached 88.86% top-1 accuracy on MIMIC-IV, significantly outperforming CAML and PLM-ICD and statistically comparable to LAAT; Combined accuracy ranged from 87.69% to 89.25% across six backbones. On MIMIC-III, in-domain Combined probing reached 74.18% accuracy and significantly exceeded all three baselines, and in a transfer-learning test a 2.8M-parameter adapter on frozen MIMIC-IV representations reached 92.51% accuracy, exceeding adapted LAAT (85.54%). Frozen LLM embeddings unify structured and narrative EHR information: the Combined configuration was strongest in every setting, and its representations transferred across datasets through a compact adapter. Multimodal probing of frozen LLM representations provides a practical, backbone-flexible approach for studying EHR modalities and adapting clinical representations across datasets.