VISTA : A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

The advances in multi-modal foundation models (FMs) (e.g., CLIP and LLaVA) have facilitated the auto-labeling of large-scale datasets, enhancing model performance in challenging downstream tasks such as open-vocabulary object detection and segmentation. However, the quality of FM-generated labels is...

Ausführliche Beschreibung

Bibliographische Detailangaben
Veröffentlicht in:IEEE transactions on visualization and computer graphics. - 1996. - PP(2025) vom: 29. Jan.
1. Verfasser: Xuan, Xiwei (VerfasserIn)
Weitere Verfasser: Wang, Xiaoqi, He, Wenbin, Ono, Jorge Piazentin, Gou, Liang, Ma, Kwan-Liu, Ren, Liu
Format: Online-Aufsatz
Sprache:English
Veröffentlicht: 2025
Zugriff auf das übergeordnete Werk:IEEE transactions on visualization and computer graphics
Schlagworte:Journal Article