VISTA : A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels
The advances in multi-modal foundation models (FMs) (e.g., CLIP and LLaVA) have facilitated the auto-labeling of large-scale datasets, enhancing model performance in challenging downstream tasks such as open-vocabulary object detection and segmentation. However, the quality of FM-generated labels is...
Ausführliche Beschreibung
Bibliographische Detailangaben
Veröffentlicht in: | IEEE transactions on visualization and computer graphics. - 1996. - PP(2025) vom: 29. Jan.
|
1. Verfasser: |
Xuan, Xiwei
(VerfasserIn) |
Weitere Verfasser: |
Wang, Xiaoqi,
He, Wenbin,
Ono, Jorge Piazentin,
Gou, Liang,
Ma, Kwan-Liu,
Ren, Liu |
Format: | Online-Aufsatz
|
Sprache: | English |
Veröffentlicht: |
2025
|
Zugriff auf das übergeordnete Werk: | IEEE transactions on visualization and computer graphics
|
Schlagworte: | Journal Article |