Google Scholar

Knowing where to look? Analysis on attention of visual question answering system

W Li, Z Yuan, X Fang, C Wang - Proceedings of the …, 2018 - openaccess.thecvf.com

Proceedings of the European Conference on Computer Vision …, 2018•openaccess.thecvf.com

Attentionmechanismshavebeenwidelyusedi… Answering (VQA) solutions due to their
capacity to model deep cross-domain interactions. Analyzing attention maps offers us a
perspective to find out limitations of current VQA systems and an opportunity to further
improve them. In this paper, we select two state-of-the-art VQA approaches with attention
mechanisms to study their robustness and disadvantages by visualizing and analyzing their
estimated attention maps. We find that both methods are sensitive to features, and …

Abstract

AttentionmechanismshavebeenwidelyusedinVisualQuestion Answering (VQA) solutions due to their capacity to model deep cross-domain interactions. Analyzing attention maps offers us a perspective to find out limitations of current VQA systems and an opportunity to further improve them. In this paper, we select two state-of-the-art VQA approaches with attention mechanisms to study their robustness and disadvantages by visualizing and analyzing their estimated attention maps. We find that both methods are sensitive to features, and simultaneously, they perform badly for counting and multi-object related questions. We believe that the findings and analytical method will help researchers identify crucial challenges on the way to improve their own VQA systems.

openaccess.thecvf.com

Show moreShow less

Save Cite Cited by 10 Related articles All 7 versions View as HTML

Cite

Advanced search

Saved to My library

Knowing where to look? Analysis on attention of visual question answering system