TransVG++ : End-to-End Visual Grounding With Language Conditioned Vision Transformer
In this work, we explore neat yet effective Transformer-based frameworks for visual grounding. The previous methods generally address the core problem of visual grounding, i.e., multi-modal fusion and reasoning, with manually-designed mechanisms. Such heuristic designs are not only complicated but a...
Ausführliche Beschreibung
Bibliographische Detailangaben
| Veröffentlicht in: | IEEE transactions on pattern analysis and machine intelligence. - 1979. - 45(2023), 11 vom: 11. Nov., Seite 13636-13652
|
| 1. Verfasser: |
Deng, Jiajun
(VerfasserIn) |
| Weitere Verfasser: |
Yang, Zhengyuan,
Liu, Daqing,
Chen, Tianlang,
Zhou, Wengang,
Zhang, Yanyong,
Li, Houqiang,
Ouyang, Wanli |
| Format: | Online-Aufsatz
|
| Sprache: | English |
| Veröffentlicht: |
2023
|
| Zugriff auf das übergeordnete Werk: | IEEE transactions on pattern analysis and machine intelligence
|
| Schlagworte: | Journal Article |