作者
Rakshith Shetty, Marcus Rohrbach, Lisa Anne Hendricks, Mario Fritz, Bernt Schiele
发表日期
2017
研讨会论文
Proceedings of the IEEE international conference on computer vision
页码范围
4135-4144
简介
While strong progress has been made in image captioning recently, machine and human captions are still quite distinct. This is primarily due to the deficiencies in the generated word distribution, vocabulary size, and strong bias in the generators towards frequent captions. Furthermore, humans--rightfully so--generate multiple, diverse captions, due to the inherent ambiguity in the captioning task which is not explicitly considered in today's systems. To address these challenges, we change the training objective of the caption generator from reproducing ground-truth captions to generating a set of captions that is indistinguishable from human written captions. Instead of handcrafting such a learning target, we employ adversarial training in combination with an approximate Gumbel sampler to implicitly match the generated distribution to the human one. While our method achieves comparable performance to the state-of-the-art in terms of the correctness of the captions, we generate a set of diverse captions that are significantly less biased and better match the global uni-, bi-and tri-gram distributions of the human captions.
引用总数
20172018201920202021202220232024627545447354223
学术搜索中的文章
R Shetty, M Rohrbach, L Anne Hendricks, M Fritz… - Proceedings of the IEEE international conference on …, 2017