Cptr: Full transformer network for image captioning

C Zhang, C Zhang, S Zheng, Y Qiao, C Li… - arXiv preprint arXiv …, 2023 - arxiv.org

As ChatGPT goes viral, generative AI (AIGC, aka AI-generated content) has made headlines
everywhere because of its ability to analyze and create text, images, and beyond. With such …

被引用次数：206 相关文章所有 4 个版本

[PDF] arxiv.org

From show to tell: A survey on deep learning-based image captioning

M Stefanini, M Cornia, L Baraldi… - IEEE transactions on …, 2022 - ieeexplore.ieee.org

Connecting Vision and Language plays an essential role in Generative Intelligence. For this
reason, large research efforts have been devoted to image captioning, ie describing images …

被引用次数：394 相关文章所有 11 个版本

[PDF] arxiv.org

Clipcap: Clip prefix for image captioning

R Mokady, A Hertz, AH Bermano - arXiv preprint arXiv:2111.09734, 2021 - arxiv.org

Image captioning is a fundamental task in vision-language understanding, where the model
predicts a textual informative caption to a given input image. In this paper, we present a …

被引用次数：762 相关文章所有 2 个版本

[PDF] thecvf.com

Multiscale vision transformers

H Fan, B Xiong, K Mangalam, Y Li… - Proceedings of the …, 2021 - openaccess.thecvf.com

Abstract We present Multiscale Vision Transformers (MViT) for video and image recognition,
by connecting the seminal idea of multiscale feature hierarchies with transformer models …

被引用次数：1530 相关文章所有 5 个版本

[PDF] arxiv.org

Imagenet-21k pretraining for the masses

T Ridnik, E Ben-Baruch, A Noy… - arXiv preprint arXiv …, 2021 - arxiv.org

ImageNet-1K serves as the primary dataset for pretraining deep learning models for
computer vision tasks. ImageNet-21K dataset, which is bigger and more diverse, is used …

被引用次数：712 相关文章所有 7 个版本

[PDF] arxiv.org

Remote sensing image change detection with transformers

H Chen, Z Qi, Z Shi - IEEE Transactions on Geoscience and …, 2021 - ieeexplore.ieee.org

Modern change detection (CD) has achieved remarkable success by the powerful
discriminative ability of deep convolutions. However, high-resolution remote sensing CD …

被引用次数：1074 相关文章所有 3 个版本

[PDF] thecvf.com

Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment

L Yao, J Han, X Liang, D Xu… - Proceedings of the …, 2023 - openaccess.thecvf.com

This paper presents DetCLIPv2, an efficient and scalable training framework that
incorporates large-scale image-text pairs to achieve open-vocabulary object detection …

被引用次数：80 相关文章所有 5 个版本

[PDF] mlr.press

Causal transformer for estimating counterfactual outcomes

V Melnychuk, D Frauen… - … Conference on Machine …, 2022 - proceedings.mlr.press

Estimating counterfactual outcomes over time from observational data is relevant for many
applications (eg, personalized medicine). Yet, state-of-the-art methods build upon simple …

被引用次数：100 相关文章所有 7 个版本

[PDF] arxiv.org

Reltr: Relation transformer for scene graph generation

Y Cong, MY Yang, B Rosenhahn - IEEE Transactions on …, 2023 - ieeexplore.ieee.org

Different objects in the same scene are more or less related to each other, but only a limited
number of these relationships are noteworthy. Inspired by Detection Transformer, which …

被引用次数：161 相关文章所有 10 个版本

[PDF] thecvf.com

Prior: Prototype representation joint learning from medical images and reports

P Cheng, L Lin, J Lyu, Y Huang… - Proceedings of the …, 2023 - openaccess.thecvf.com

Contrastive learning based vision-language joint pre-training has emerged as a successful
representation learning strategy. In this paper, we present a prototype representation …

被引用次数：49 相关文章所有 6 个版本

高级搜索

QQ 群