Scaling and Benchmarking Self-Supervised Visual Representation Learning

Last update: Dec 31, 2022

Overview

FAIR Self-Supervision Benchmark is deprecated. Please see VISSL, a ground-up rewrite of benchmark in PyTorch.

FAIR Self-Supervision Benchmark

This code provides various benchmark (and legacy) tasks for evaluating quality of visual representations learned by various self-supervision approaches. This code corresponds to our work on Scaling and Benchmarking Self-Supervised Visual Representation Learning. The code is written in Python and can be used to evaluate both PyTorch and Caffe2 models (see this). We hope that this benchmark release will provided a consistent evaluation strategy that will allow measuring the progress in self-supervision easily.

Introduction

The goal of fair_self_supervision_benchmark is to standardize the methodology for evaluating quality of visual representations learned by various self-supervision approaches. Further, it provides evaluation on a variety of tasks as follows:

Benchmark tasks: The benchmark tasks are based on principle: a good representation (1) transfers to many different tasks, and, (2) transfers with limited supervision and limited fine-tuning. The tasks are as follows.

Image Classification
- VOC07
- COCO2014
- Places205
Low-Shot Image Classification
- VOC07
- Places205
Object Detection on VOC07 and VOC07+12 with frozen backbone for detectors:
- Fast R-CNN
- Faster R-CNN
Surface Normal Estimation
Visual Navigation in Gibson Environment

These Benchmark tasks use the network architectures:

Legacy tasks: We also classify some commonly used evaluation tasks as legacy tasks for reasons mentioned in Section 7 of paper. The tasks are as follows:

ImageNet-1K classification task
VOC07 full finetuning
Object Detection on VOC07 and VOC07+12 with full tuning for detectors:
- Fast R-CNN
- Faster R-CNN

License

fair_self_supervision_benchmark is CC-NC 4.0 International licensed, as found in the LICENSE file.

Citation

If you use fair_self_supervision_benchmark in your research or wish to refer to the baseline results published in the paper, please use the following BibTeX entry.

@article{goyal2019scaling,
  title={Scaling and Benchmarking Self-Supervised Visual Representation Learning},
  author={Goyal, Priya and Mahajan, Dhruv and Gupta, Abhinav and Misra, Ishan},
  journal={arXiv preprint arXiv:1905.01235},
  year={2019}
}

Installation

Please find installation instructions in INSTALL.md.

Getting Started

After installation, please see GETTING_STARTED.md for how to run various benchmark tasks.

Model Zoo

We provide models used in our paper in the MODEL_ZOO.

References

Scaling and Benchmarking Self-Supervised Visual Representation Learning. Priya Goyal, Dhruv Mahajan, Abhinav Gupta*, Ishan Misra*. Tech report, arXiv, May 2019.

Scaling and Benchmarking Self-Supervised Visual Representation Learning

Related tags

Overview

FAIR Self-Supervision Benchmark is deprecated. Please see VISSL, a ground-up rewrite of benchmark in PyTorch.

FAIR Self-Supervision Benchmark

Introduction

License

Citation

Installation

Getting Started

Model Zoo

References

Owner

Meta Research

Learn the Deep Learning for Computer Vision in three steps: theory from base to SotA, code in PyTorch, and space-repetition with Anki

LegoDNN: a block-grained scaling tool for mobile vision systems

Official PyTorch code of DeepPanoContext: Panoramic 3D Scene Understanding with Holistic Scene Context Graph and Relation-based Optimization (ICCV 2021 Oral).

1st place solution in CCF BDCI 2021 ULSEG challenge

A Simplied Framework of GAN Inversion

FcaNet: Frequency Channel Attention Networks

SpinalNet: Deep Neural Network with Gradual Input

🐸STT integration examples

Tree-based Search Graph for Approximate Nearest Neighbor Search

An LSTM based GAN for Human motion synthesis

“Robust Lightweight Facial Expression Recognition Network with Label Distribution Training”, AAAI 2021.

Predicting Semantic Map Representations from Images with Pyramid Occupancy Networks

Programming with Neural Surrogates of Programs

A simple command line tool for text to image generation, using OpenAI's CLIP and a BigGAN.

[ICCV 2021] Official PyTorch implementation for Deep Relational Metric Learning.

PyTorch implementation of the TTC algorithm

Source code for the ACL-IJCNLP 2021 paper entitled "T-DNA: Taming Pre-trained Language Models with N-gram Representations for Low-Resource Domain Adaptation" by Shizhe Diao et al.

Exploring whether attention is necessary for vision transformers

This code is 3d-CNN model that can predict environmental value