A Novel Plug-in Module for Fine-grained Visual Classification

Last update: Dec 20, 2022

Overview

A Novel Plug-in Module for Fine-grained Visual Classification

paper url: https://arxiv.org/abs/2202.03822

We propose a novel plug-in module that can be integrated to many common backbones, including CNN-based or Transformer-based networks to provide strongly discriminative regions. The plugin module can output pixel-level feature maps and fuse filtered features to enhance fine-grained visual classification. Experimental results show that the proposed plugin module outperforms state-ofthe-art approaches and significantly improves the accuracy to 92.77% and 92.83% on CUB200-2011 and NABirds, respectively.

1. Environment setting

install requirements
replace folder timm/ to our timm/ folder (for ViT or Swin-T)

Prepare dataset

In this paper, we use 2 large bird's datasets:

Our pretrained model

Download the pretrained model from this url: https://drive.google.com/drive/folders/1ivMJl4_EgE-EVU_5T8giQTwcNQ6RPtAo?usp=sharing

backup/ is our pretrained model path.
resnet50_miil_21k.pth and vit_base_patch16_224_miil_21k.pth are imagenet21k pretrained model (place these file under models/), thanks to https://github.com/Alibaba-MIIL/ImageNet21K/blob/main/MODEL_ZOO.md !!

OS

Windows10
Ubuntu20.04
macOS

2. Train

configuration file: config.py

python train.py --train_root "./CUB200-2011/train/" --val_root "./CUB200-2011/test/"

3. Evaluation

configuration file: config_eval.py

python eval.py --pretrained_path "./backup/CUB200/best.pth" --val_root "./CUB200-2011/test/"

4. Visualization

configuration file: config_plot.py

python plot_heat.py --pretrained_path "./backup/CUB200/best.pth" --img_path "./img/001.png/"

Acknowledgment

Thanks to timm for Pytorch implementation.
This work was financially supported by the National Taiwan Normal University (NTNU) within the framework of the Higher Education Sprout Project by the Ministry of Education(MOE) in Taiwan, sponsored by Ministry of Science and Technology, Taiwan, R.O.C. under Grant no. MOST 110- 2221-E-003-026, 110-2634-F-003 -007, and 110-2634-F-003 -006. In addition, we thank to National Center for Highperformance Computing (NCHC) for providing computational and storage resources.

A Novel Plug-in Module for Fine-grained Visual Classification

Related tags

Overview

A Novel Plug-in Module for Fine-grained Visual Classification

1. Environment setting

Prepare dataset

Our pretrained model

OS

2. Train

3. Evaluation

4. Visualization

Acknowledgment

Owner

ChouPoYung

Official Implementation of HRDA: Context-Aware High-Resolution Domain-Adaptive Semantic Segmentation

The official implementation of the paper, "SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation Learning"

A fast MoE impl for PyTorch

Code for the paper "Generative design of breakwaters usign deep convolutional neural network as a surrogate model"

Uncertain natural language inference

PyTorch implementation of MLP-Mixer

Semi-supervised Semantic Segmentation with Directional Context-aware Consistency (CVPR 2021)

Implementation of Graph Transformer in Pytorch, for potential use in replicating Alphafold2

Implicit Deep Adaptive Design (iDAD)

Code for Generating Disentangled Arguments with Prompts: A Simple Event Extraction Framework that Works

Python implementation of NARS (Non-Axiomatic-Reasoning-System)

The pyrelational package offers a flexible workflow to enable active learning with as little change to the models and datasets as possible

Source code for the ACL-IJCNLP 2021 paper entitled "T-DNA: Taming Pre-trained Language Models with N-gram Representations for Low-Resource Domain Adaptation" by Shizhe Diao et al.

Code for the paper "Regularizing Variational Autoencoder with Diversity and Uncertainty Awareness"

This repository contains a CBIR system that uses swin transformer to extract image's feature.

CTF Challenge for CSAW Finals 2021

Let's create a tool to convert Thailand budget from PDF to CSV.

An implementation of chunked, compressed, N-dimensional arrays for Python.

[CVPR 2021] "The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models" Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Michael Carbin, Zhangyang Wang

NeRD: Neural Reflectance Decomposition from Image Collections