A complete survey on generative ai (aigc): Is chatgpt from gpt-4 to gpt-5 all you need?

C Zhang, C Zhang, S Zheng, Y Qiao, C Li… - arXiv preprint arXiv …, 2023 - arxiv.org
As ChatGPT goes viral, generative AI (AIGC, aka AI-generated content) has made headlines
everywhere because of its ability to analyze and create text, images, and beyond. With such …

From show to tell: A survey on deep learning-based image captioning

M Stefanini, M Cornia, L Baraldi… - IEEE transactions on …, 2022 - ieeexplore.ieee.org
Connecting Vision and Language plays an essential role in Generative Intelligence. For this
reason, large research efforts have been devoted to image captioning, ie describing images …

Clipcap: Clip prefix for image captioning

R Mokady, A Hertz, AH Bermano - arXiv preprint arXiv:2111.09734, 2021 - arxiv.org
Image captioning is a fundamental task in vision-language understanding, where the model
predicts a textual informative caption to a given input image. In this paper, we present a …

Multiscale vision transformers

H Fan, B Xiong, K Mangalam, Y Li… - Proceedings of the …, 2021 - openaccess.thecvf.com
Abstract We present Multiscale Vision Transformers (MViT) for video and image recognition,
by connecting the seminal idea of multiscale feature hierarchies with transformer models …

Imagenet-21k pretraining for the masses

T Ridnik, E Ben-Baruch, A Noy… - arXiv preprint arXiv …, 2021 - arxiv.org
ImageNet-1K serves as the primary dataset for pretraining deep learning models for
computer vision tasks. ImageNet-21K dataset, which is bigger and more diverse, is used …

Remote sensing image change detection with transformers

H Chen, Z Qi, Z Shi - IEEE Transactions on Geoscience and …, 2021 - ieeexplore.ieee.org
Modern change detection (CD) has achieved remarkable success by the powerful
discriminative ability of deep convolutions. However, high-resolution remote sensing CD …

Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment

L Yao, J Han, X Liang, D Xu… - Proceedings of the …, 2023 - openaccess.thecvf.com
This paper presents DetCLIPv2, an efficient and scalable training framework that
incorporates large-scale image-text pairs to achieve open-vocabulary object detection …

Causal transformer for estimating counterfactual outcomes

V Melnychuk, D Frauen… - … Conference on Machine …, 2022 - proceedings.mlr.press
Estimating counterfactual outcomes over time from observational data is relevant for many
applications (eg, personalized medicine). Yet, state-of-the-art methods build upon simple …

Reltr: Relation transformer for scene graph generation

Y Cong, MY Yang, B Rosenhahn - IEEE Transactions on …, 2023 - ieeexplore.ieee.org
Different objects in the same scene are more or less related to each other, but only a limited
number of these relationships are noteworthy. Inspired by Detection Transformer, which …

Prior: Prototype representation joint learning from medical images and reports

P Cheng, L Lin, J Lyu, Y Huang… - Proceedings of the …, 2023 - openaccess.thecvf.com
Contrastive learning based vision-language joint pre-training has emerged as a successful
representation learning strategy. In this paper, we present a prototype representation …