Xiaowei Hu
Microsoft
18 Papers
112 Citations
Xiaowei Hu is an academic researcher from Microsoft. The author has contributed to research in topics: Closed captioning & Computer science. The author has an hindex of 9, co-authored 18 publications. Previous affiliations of Xiaowei Hu include University of Alberta.
Chat about Author
Papers
•Posted Content
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li,Xi Yin,Chunyuan Li,Pengchuan Zhang,Xiaowei Hu,Lei Zhang,Lijuan Wang,Houdong Hu,Li Dong,Furu Wei,Yejin Choi,Jianfeng Gao +11 more
TL;DR: This paper proposes a new learning method Oscar (Object-Semantics Aligned Pre-training), which uses object tags detected in images as anchor points to significantly ease the learning of alignments.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li,Xi Yin,Chunyuan Li,Pengchuan Zhang,Xiaowei Hu,Lei Zhang,Lijuan Wang,Houdong Hu,Li Dong,Furu Wei,Yejin Choi,Jianfeng Gao +11 more
- 23 Aug 2020
TL;DR: Oscar as discussed by the authors uses object tags detected in images as anchor points to significantly ease the learning of alignments, motivated by the observation that the salient objects in an image can be accurately detected, and are often mentioned in the paired text.
1.2K
VinVL: Revisiting Visual Representations in Vision-Language Models
Pengchuan Zhang,Xiujun Li,Xiaowei Hu,Jianwei Yang,Lei Zhang,Lijuan Wang,Yejin Choi,Jianfeng Gao +7 more
- 20 Jun 2021
TL;DR: Li et al. as discussed by the authors presented a detailed study of improving visual representations for vision language (VL) tasks, and developed an improved object detection model to provide object-centric representations of images.
•Posted Content
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
TL;DR: PICa as discussed by the authors proposes to use image captions for knowledge-based visual question answering (VQA) in a few-shot manner by converting the image into captions (or tags) that GPT-3 can understand.
135
VinVL: Making Visual Representations Matter in Vision-Language Models
Pengchuan Zhang,Xiujun Li,Xiaowei Hu,Jianwei Yang,Lei Zhang,Lijuan Wang,Yejin Choi,Jianfeng Gao +7 more
- 02 Jan 2021
TL;DR: Zhang et al. as discussed by the authors presented a detailed study of improving visual representations for vision language (VL)tasks and developed an improved object detection model to provide object-centric representations of images.
79