Official pytorch implementation of paper Dual-Level Collaborative Transformer for Image Captioning (AAAI 2021).

Last update: Dec 11, 2022

Related tags

Deep Learning image-captioning

Overview

Dual-Level Collaborative Transformer for Image Captioning

This repository contains the reference code for the paper Dual-Level Collaborative Transformer for Image Captioning.

Experiment setup

please refer to m2 transformer

Data preparation

Annotation. Download the annotation file annotation.zip. Extarct and put it in the project root directory.
Feature. You can download our ResNeXt-101 feature (hdf5 file) here. Acess code: jcj6.
evaluation. Download the evaluation tools here. Acess code: jcj6. Extarct and put it in the project root directory.

There are five kinds of keys in our .hdf5 file. They are

['%d_features' % image_id]: region features (N_regions, feature_dim)
['%d_boxes' % image_id]: bounding box of region features (N_regions, 4)
['%d_size' % image_id]: size of original image (for normalizing bounding box), (2,)
['%d_grids' % image_id]: grid features (N_grids, feature_dim)
['%d_mask' % image_id]: geometric alignment graph, (N_regions, N_grids)

We extract feature with the code in grid-feats-vqa.

The first three keys can be obtained when extracting region features with extract_region_feature.py. The forth key can be obtained when extracting grid features with code in grid-feats-vqa. The last key can be obtained with align.ipynb

Training

python train.py --exp_name dlct --batch_size 50 --head 8 --features_path ./data/coco_all_align.hdf5 --annotation annotation --workers 8 --rl_batch_size 100 --image_field ImageAllFieldWithMask --model DLCT --rl_at 17 --seed 118

Evaluation

python eval.py --annotation annotation --workers 4 --features_path ./data/coco_all_align.hdf5 --model_path path_of_model_to_eval --model DLCT --image_field ImageAllFieldWithMask --grid_embed --box_embed --dump_json gen_res.json --beam_size 5

Important args:

--features_path path to hdf5 file
--model_path
--dump_json dump generated captions to

Pretrained model is available here. Acess code: jcj6. By evaluating the pretrained model, you will get

{'BLEU': [0.8136727001615207, 0.6606095421082421, 0.5167535314080227, 0.39790755018790197], 'METEOR': 0.29522868252436046, 'ROUGE': 0.5914367650104326, 'CIDEr': 1.3382047139781112, 'SPICE': 0.22953477359195887}

References

[1] M2

[2] grid-feats-vqa

[3] butd

Acknowledgements

Thanks the original m2 and amazing work of grid-feats-vqa.

Official pytorch implementation of paper Dual-Level Collaborative Transformer for Image Captioning (AAAI 2021).

Related tags

Overview

Dual-Level Collaborative Transformer for Image Captioning

Experiment setup

Data preparation

Training

Evaluation

References

Acknowledgements

Owner

lyricpoem

Boston House Prediction Valuation Tool

Hypersearch weight debugging and losses tutorial

Train emoji embeddings based on emoji descriptions.

Conceptual 12M is a dataset containing (image-URL, caption) pairs collected for vision-and-language pre-training.

Self-supervised learning (SSL) is a method of machine learning

Filtering variational quantum algorithms for combinatorial optimization

Reinforcement Learning for the Blackjack

A standard framework for modelling Deep Learning Models for tabular data

Official code for paper "Demystifying Local Vision Transformer: Sparse Connectivity, Weight Sharing, and Dynamic Weight"

MNIST, but with Bezier curves instead of pixels

[CVPR'21] DeepSurfels: Learning Online Appearance Fusion

Image Processing, Image Smoothing, Edge Detection and Transforms

DeepSpeed is a deep learning optimization library that makes distributed training easy, efficient, and effective.

Deep Ensemble Learning with Jet-Like architecture

Code base for NeurIPS 2021 publication titled Kernel Functional Optimisation (KFO)

[CVPR2022] Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos

Codes for our paper "SentiLARE: Sentiment-Aware Language Representation Learning with Linguistic Knowledge" (EMNLP 2020)

Large dataset storage format for Pytorch

[ACL-IJCNLP 2021] "EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets"

The BCNet related data and inference model.