ModalChorus : Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion Map

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting in decreased model performance and diminished generalization....

Ausführliche Beschreibung

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on visualization and computer graphics. - 1996. - 31(2025), 1 vom: 03. Jan., Seite 294-304
1. Verfasser:	Ye, Yilin (VerfasserIn)
Weitere Verfasser:	Xiao, Shishi, Zeng, Xingchen, Zeng, Wei
Format:	Online-Aufsatz
Sprache:	English
Veröffentlicht:	2025
Zugriff auf das übergeordnete Werk:	IEEE transactions on visualization and computer graphics
Schlagworte:	Journal Article

Online verfügbar	Volltext