VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Last update: Dec 12, 2022

Related tags

Overview

VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Usage

First, install PyTorch 1.7.1+, torchvision 0.8.2+ and other required packages as follows：

conda install -c pytorch pytorch torchvision
pip install timm==0.3.2
pip install ftfy regex tqdm
pip install git+https://github.com/openai/CLIP.git
pip install mmcv==1.3.14

Data preparation

ImageNet-LT

Download and extract ImageNet train and val images from here. The directory structure is the standard layout for the torchvision datasets.ImageFolder, and the training and validation data is expected to be in the train/ folder and val/ folder respectively.

Then download and extract the wiki text into the same directory, and the directory tree of data is expected to be like this:

./data/imagenet/
  train/
    class1/
      img1.jpeg
    class2/
      img2.jpeg
  val/
    class1/
      img3.jpeg
    class2/
      img4.jpeg
  wiki/
  	desc_1.txt
  ImageNet_LT_test.txt
  ImageNet_LT_train.txt
  ImageNet_LT_val.txt
  labels.txt

After that, download the CLIP's pretrained weight RN50.pt and ViT-B-16.pt into the pretrained directory from https://github.com/openai/CLIP.

Places-LT

Download the places365_standard data from here.

Then download and extract the wiki text into the same directory. The directory tree of data is expected to be like this (almost the same as ImageNet-LT):

./data/places/
  train/
    class1/
      img1.jpeg
    class2/
      img2.jpeg
  val/
    class1/
      img3.jpeg
    class2/
      img4.jpeg
  wiki/
  	desc_1.txt
  Places_LT_test.txt
  Places_LT_train.txt
  Places_LT_val.txt
  labels.txt

iNaturalist 2018

Download the iNaturalist 2018 data from here.

Then download and extract the wiki text into the same directory. The directory tree of data is expected to be like this:

./data/iNat/
  train_val2018/
  wiki/
  	desc_1.txt
  categories.json
  test2018.json
  train2018.json
  val.json

Evaluation

To evaluate VL-LTR with a single GPU run:

Pre-training stage

bash eval.sh ${CONFIG_PATH} 1 --eval-pretrain

Fine-tuning stage:

bash eval.sh ${CONFIG_PATH} 1

The ${CONFIG_PATH} is the relative path of the corresponding configuration file in the config directory.

Training

To train VL-LTR on a single node with 8 GPUs for:

Pre-training stage, run:

bash dist_train_arun.sh ${PARTITION} ${CONFIG_PATH} 8

Fine-tuning stage:
- First, calculate the $\mathcal L_{\text{lin}}$ of each sentence for AnSS method by running this:
```
bash eval.sh ${CONFIG_PATH} 1 --eval-pretrain --select
```
- then, running this:
```
bash dist_train_arun.sh ${PARTITION} ${CONFIG_PATH} 8
```

The ${CONFIG_PATH} is the relative path of the corresponding configuration file in the config directory.

Results

Below list our model's performance on ImageNet-LT, Places-LT, and iNaturalist 2018.

Dataset	Backbone	Top-1 Accuracy	Download
ImageNet-LT	ResNet-50	70.1	Weights
ImageNet-LT	ViT-Base-16	77.2	Weights
Places-LT	ResNet-50	48.0	Weights
Places-LT	ViT-Base-16	50.1	Weights
iNaturalist 2018	ResNet-50	74.6	Weights
iNaturalist 2018	ViT-Base-16	76.8	Weights

For more detailed information, please refer to our paper directly.

Citation

If you are interested in our work, please cite as follows:

@article{tian2021vl,
  title={VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition},
  author={Tian, Changyao and Wang, Wenhai and Zhu, Xizhou and Wang, Xiaogang and Dai, Jifeng and Qiao, Yu},
  journal={arXiv preprint arXiv:2111.13579},
  year={2021}
}

License

This repository is released under the Apache 2.0 license as found in the LICENSE file.

Code for the AAAI-2022 paper: Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification (AAAI 2022) Prerequisite PyTorch = 1.2.0 P

16 Dec 14, 2022

Pytorch implementation of the AAAI 2022 paper "Cross-Domain Empirical Risk Minimization for Unbiased Long-tailed Classification"

[AAAI22] Cross-Domain Empirical Risk Minimization for Unbiased Long-tailed Classification We point out the overlooked unbiasedness in long-tailed clas

28 Oct 18, 2022

On Size-Oriented Long-Tailed Graph Classification of Graph Neural Networks

On Size-Oriented Long-Tailed Graph Classification of Graph Neural Networks We provide the code (in PyTorch) and datasets for our paper "On Size-Orient

4 Jun 18, 2022

Speech Emotion Recognition with Fusion of Acoustic- and Linguistic-Feature-Based Decisions

APSIPA-SER-with-A-and-T This code is the implementation of Speech Emotion Recognition (SER) with acoustic and linguistic features. The network model i

3 Jan 4, 2023

A weakly-supervised scene graph generation codebase. The implementation of our CVPR2021 paper ``Linguistic Structures as Weak Supervision for Visual Scene Graph Generation''

README.md shall be finished soon. WSSGG 0 Overview 1 Installation 1.1 Faster-RCNN 1.2 Language Parser 1.3 GloVe Embeddings 2 Settings 2.1 VG-GT-Graph

35 Nov 20, 2022

Implementation of "Distribution Alignment: A Unified Framework for Long-tail Visual Recognition"(CVPR 2021)

105 Nov 7, 2022

Official codes for the paper "Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech"

ResDAVEnet-VQ Official PyTorch implementation of Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech What is in this repo? M

21 Aug 23, 2022

[ICCV2021] Official code for "Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition"

CTR-GCN This repo is the official implementation for Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition. The pap

148 Dec 16, 2022

Official implementation for CVPR 2021 paper: Adaptive Class Suppression Loss for Long-Tail Object Detection

Adaptive Class Suppression Loss for Long-Tail Object Detection This repo is the official implementation for CVPR 2021 paper: Adaptive Class Suppressio

67 Dec 4, 2022

Comments

Problem about running eval.sh

""" #!/usr/bin/env bash set -x

export NCCL_LL_THRESHOLD=0

CONFIG=$1 GPUS=$1 CPUS=$[GPUS*2] PORT=${PORT:-8886}

CONFIG_NAME=${CONFIG##/} CONFIG_NAME=${CONFIG_NAME%.}

OUTPUT_DIR="./checkpoints/eval" if [ ! -d $OUTPUT_DIR ]; then mkdir ${OUTPUT_DIR} fi

python -u main.py
--port=$PORT
--num_workers 4
--resume "./checkpoints/${CONFIG_NAME}/checkpoint.pth"
--output-dir ${OUTPUT_DIR}
--config $CONFIG ${@:3}
--eval
2>&1 | tee -a ${OUTPUT_DIR}/train.log """ I have two A100, so set GPUS is 2. All other settings according to ReadME.md but I got a problem when running eval.sh """ File "eval.sh", line 4 export NCCL_LL_THRESHOLD=0 ^ SyntaxError: invalid syntax

"""

opened by euminds 2
Mismatch between code and diagram in paper for the fine-tuning phase

In fig 3, stage 2 from the paper, it looks like value for the attention is calculated based on Vision and language (Q is vision, K is language) and then applied to the language (V). But in the code, the attention is applied to the visual features. Can you verify which one is the correct way? @ChangyaoTian

opened by rahulvigneswaran 0

pre-trained weights with TorchScript?

Hello, Thanks for the great work! May I ask if it's possible for you to also provide the checkpoint weight in a TorchScript version?

It's something like:

import torch
import torchvision.models as models

model = models.resnet50()
traced = torch.jit.trace(model, (torch.rand(4, 3, 224, 224),))
torch.jit.save(traced, "test.pt")

# load model
model = torch.jit.load("test.pt")

opened by xinleihe 0

Releases(ECCV-2022-video)

ECCV-2022-video(Oct 1, 2022)

Source code(tar.gz)
Source code(zip)
4849.mp4(15.24 MB)
ECCV-2022(Oct 1, 2022)

Source code(tar.gz)
Source code(zip)
4849.pdf(2.66 MB)
text-corpus(Jan 12, 2022)

The text corpus for ImageNet-LT, Places-LT, and iNaturalist2018, including both wiki data and the class labels.
Source code(tar.gz)
Source code(zip)
imagenet.zip(7.29 MB)
iNat.zip(13.12 MB)
places.zip(2.67 MB)
checkpoints(Jan 12, 2022)

model weights for 3 datasets, including both 2 stages, using vit16 and resnet50 backbones respectively.
Source code(tar.gz)
Source code(zip)
imageNet-LT_r50.zip(1382.52 MB)
imageNet-LT_vit16.zip(1920.94 MB)
inat_finetune_r50.zip(1754.00 MB)
inat_finetune_vit16.zip(1566.75 MB)
inat_pretrain_r50.zip(1008.10 MB)
inat_pretrain_vit16.zip(1512.80 MB)
places_r50.zip(1252.33 MB)
places_vit16.zip(1843.92 MB)

Owner

GitHub Repository

https://sites.google.com/cornell.edu/recsys2021tutorial

Counterfactual Learning and Evaluation for Recommender Systems (RecSys'21 Tutorial) Materials for "Counterfactual Learning and Evaluation for Recommen

45 Nov 10, 2022

Current state of supervised and unsupervised depth completion methods

Awesome Depth Completion Table of Contents About Sparse-to-Dense Depth Completion Current State of Depth Completion Unsupervised VOID Benchmark Superv

224 Dec 28, 2022

Point detection through multi-instance deep heatmap regression for sutures in endoscopy

Suture detection PyTorch This repo contains the reference implementation of suture detection model in PyTorch for the paper Point detection through mu

3 Jul 16, 2022

It helps user to learn Pick-up lines and share if he has a better one

Pick-up-Lines-Generator(Open Source) It helps user to learn Pick-up lines Share and Add one or many to the DataBase Unique SQLite DataBase AI Undercon

0 May 04, 2022

Supplemental learning materials for "Fourier Feature Networks and Neural Volume Rendering"

Fourier Feature Networks and Neural Volume Rendering This repository is a companion to a lecture given at the University of Cambridge Engineering Depa

133 Dec 26, 2022

Repository of 3D Object Detection with Pointformer (CVPR2021)

3D Object Detection with Pointformer This repository contains the code for the paper 3D Object Detection with Pointformer (CVPR 2021) [arXiv]. This wo

117 Jan 06, 2023

Official implementation of "A Shared Representation for Photorealistic Driving Simulators" in PyTorch.

A Shared Representation for Photorealistic Driving Simulators The official code for the paper: "A Shared Representation for Photorealistic Driving Sim

7 Oct 13, 2022

Implementation for Learning to Track with Object Permanence

Learning to Track with Object Permanence A video-based MOT approach capable of tracking through full occlusions: Learning to Track with Object Permane

91 Jan 03, 2023

IOT: Instance-wise Layer Reordering for Transformer Structures

Introduction This repository contains the code for Instance-wise Ordered Transformer (IOT), which is introduced in the ICLR2021 paper IOT: Instance-wi

19 Nov 15, 2022

Sdf sparse conv - Deep Learning on SDF for Classifying Brain Biomarkers

Deep Learning on SDF for Classifying Brain Biomarkers To reproduce the results f

1 Jan 25, 2022

Official codebase for "B-Pref: Benchmarking Preference-BasedReinforcement Learning" contains scripts to reproduce experiments.

B-Pref Official codebase for B-Pref: Benchmarking Preference-BasedReinforcement Learning contains scripts to reproduce experiments. Install conda env

48 Dec 20, 2022

VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Related tags

Overview

VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition

Usage

Data preparation

ImageNet-LT

Places-LT

iNaturalist 2018

Evaluation

Training

Results

Citation

License

You might also like...

Code for the AAAI-2022 paper: Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

Pytorch implementation of the AAAI 2022 paper "Cross-Domain Empirical Risk Minimization for Unbiased Long-tailed Classification"

On Size-Oriented Long-Tailed Graph Classification of Graph Neural Networks

Speech Emotion Recognition with Fusion of Acoustic- and Linguistic-Feature-Based Decisions

A weakly-supervised scene graph generation codebase. The implementation of our CVPR2021 paper ``Linguistic Structures as Weak Supervision for Visual Scene Graph Generation''

Implementation of "Distribution Alignment: A Unified Framework for Long-tail Visual Recognition"(CVPR 2021)

Official codes for the paper "Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech"

[ICCV2021] Official code for "Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition"

Official implementation for CVPR 2021 paper: Adaptive Class Suppression Loss for Long-Tail Object Detection

Comments

Problem about running eval.sh

Mismatch between code and diagram in paper for the fine-tuning phase

pre-trained weights with TorchScript?

Releases(ECCV-2022-video)

ECCV-2022-video(Oct 1, 2022)

ECCV-2022(Oct 1, 2022)

text-corpus(Jan 12, 2022)

checkpoints(Jan 12, 2022)

Owner

https://sites.google.com/cornell.edu/recsys2021tutorial

Current state of supervised and unsupervised depth completion methods

Point detection through multi-instance deep heatmap regression for sutures in endoscopy

It helps user to learn Pick-up lines and share if he has a better one

Supplemental learning materials for "Fourier Feature Networks and Neural Volume Rendering"

Repository of 3D Object Detection with Pointformer (CVPR2021)

Official implementation of "A Shared Representation for Photorealistic Driving Simulators" in PyTorch.

Implementation for Learning to Track with Object Permanence

IOT: Instance-wise Layer Reordering for Transformer Structures

Sdf sparse conv - Deep Learning on SDF for Classifying Brain Biomarkers

Official codebase for "B-Pref: Benchmarking Preference-BasedReinforcement Learning" contains scripts to reproduce experiments.

It is the assignment for COMP 576 in Rice University

[ArXiv 2021] Data-Efficient Instance Generation from Instance Discrimination

FS2KToolbox FS2K Dataset Towards the translation between Face

My implementation of transformers related papers for computer vision in pytorch

Focal Loss for Dense Rotation Object Detection

Pytorch Implementation of "Contrastive Representation Learning for Exemplar-Guided Paraphrase Generation"

PyTorch implementation of Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Streamlit component for TensorBoard, TensorFlow's visualization toolkit

Uses Open AI Gym environment to create autonomous cryptocurrency bot to trade cryptocurrencies.