Перейти к содержимому

[NeurIPS 2025] Align Without Erasing | Dirichlet Processes for Multimodal AI

hc

0:00 / 0:00

[NeurIPS 2025] Align Without Erasing | Dirichlet Processes for Multimodal AI

15 просмотров · 8 дней назад
hc
5 подписчиков
15 просмотров · 8 дней назад
Can multimodal AI learn a shared representation without erasing the strengths of each input? This ten-minute companion explainer introduces DPMM: a Dirichlet-process mixture approach to amplifying prominent representations while learning cross-modal relationships. We explain the motivation, Gaussian mixtures and stick breaking, variational training, missing-modality representations, the experimental evidence, and the limits. The final minute presents Howard Chan's perspective on statistical insights for principled multimodal learning and potential future multimodal foundation models. That outlook is not a result demonstrated in the paper. CHAPTERS 0:00 The problem 1:33 The statistical idea 3:13 How DPMM works 4:56 Missing modalities 5:47 The experiments 6:43 The evidence 8:28 Limits and takeaway 9:00 Howard’s perspective PAPER Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process Tsai Hor Chan, Feng Wu, Yihang Chen, Guosheng Yin, Lequan Yu NeurIPS 2025 Conference page and original presentation: https://neurips.cc/virtual/2025/loc/s... Full paper: https://arxiv.org/abs/2510.20736 OpenReview: https://openreview.net/forum?id=dC5TW... Code: https://github.com/HKU-MedAI/DPMM RESULTS IN CONTEXT Fully matched MIMIC-IV mortality AUPR: DPMM 0.482 versus DrFuse 0.446 (Table 2). The difference is 0.036 absolute, approximately 8.1% relative—not eight percentage points. Reported 95% confidence intervals overlap; this comparison alone does not establish statistical significance. Fully matched MIMIC-IV readmission AUROC ablation: 0.717 without DP, 0.724 amplification only, 0.736 amplification plus alignment (Appendix Table D.5). CMU-MOSI accuracy: 0.701 versus best listed baseline 0.700 (Table 5), a modest 0.1 percentage-point gain. DP means Dirichlet process, not differential privacy. Sampled missing representations are not recovered clinical measurements. The paper does not establish clinical deployment readiness, causal feature identification, or foundation-model-scale performance. CREDITS AND PRODUCTION Original research by the authors above. This is an illustrated explanatory adaptation with AI-assisted scripting/graphics and synthetic Microsoft en-US-AvaNeural narration, not Howard's recorded or cloned voice. Original vector illustrations are conceptual unless explicitly labeled as a reported result. No patient-level data are shown. Paper manuscript available under CC BY 4.0; explanations and diagrams have been adapted and simplified. #NeurIPS2025 #MultimodalAI #BayesianLearning