Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets often remains labor-intensive, weakly traceable, and difficult to configure. This problem is particularly critical in multimodal medical scenarios, where each question-answer sample should be semantically consistent, grounded in visual and temporal evidence, and controllable in terms of reasoning complexity. To address these challenges, we propose Med-CRAFT, an information system for explainable and configurable construction of multimodal medical question answering datasets from instructional videos. Med-CRAFT organizes dataset construction as a provenance-aware pipeline that transforms raw medical instructional videos into structured operation knowledge graphs, evidence-grounded reasoning paths, and natural-language question-answer pairs.
Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets
Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets often remains labor-intensive, weakly traceable, and difficult to configure.
- Preview

- Year
- 2025
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2512.01045CC-BY-4.0
- TL;DR
- Semantic Scholar