ModalChorus : Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion Map
Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting in decreased model performance and diminished generalization....
| Veröffentlicht in: | IEEE transactions on visualization and computer graphics. - 1996. - 31(2025), 1 vom: 03. Jan., Seite 294-304 |
|---|---|
| 1. Verfasser: | |
| Weitere Verfasser: | , , |
| Format: | Online-Aufsatz |
| Sprache: | English |
| Veröffentlicht: |
2025
|
| Zugriff auf das übergeordnete Werk: | IEEE transactions on visualization and computer graphics |
| Schlagworte: | Journal Article |
| Online verfügbar |
Volltext |