Pytorch implemenation of Stochastic Multi-Label Image-to-image Translation (SMIT)

Last update: Mar 01, 2022

Overview

SMIT: Stochastic Multi-Label Image-to-image Translation

This repository provides a PyTorch implementation of SMIT. SMIT can stochastically translate an input image to multiple domains using only a single generator and a discriminator. It only needs a target domain (binary vector e.g., [0,1,0,1,1] for 5 different domains) and a random gaussian noise.

Paper

SMIT: Stochastic Multi-Label Image-to-image Translation
Andrés Romero¹, Pablo Arbelaez¹, Luc Van Gool², Radu Timofte²
¹Biomedical Computer Vision (BCV) Lab, Universidad de Los Andes.
²Computer Vision Lab (CVL), ETH Zürich.

Citation

@article{romero2019smit,
  title={SMIT: Stochastic Multi-Label Image-to-Image Translation},
  author={Romero, Andr{\'e}s and Arbel{\'a}ez, Pablo and Van Gool, Luc and Timofte, Radu},
  journal={ICCV Workshops},
  year={2019}
}

Dependencies

Python (2.7, 3.5+)
PyTorch (0.3, 0.4, 1.0)

Usage

Cloning the repository

$ git clone https://github.com/BCV-Uniandes/SMIT.git
$ cd SMIT

Downloading the dataset

To download the CelebA dataset:

$ bash generate_data/download.sh

Train command:

./main.py --GPU=$gpu_id --dataset_fake=CelebA

Each dataset must has datasets/ .py and datasets/ .yaml files. All models and figures will be stored at snapshot/models/$dataset_fake/ _ .pth and snapshot/samples/$dataset_fake/ _ .jpg, respectivelly.

Test command:

./main.py --GPU=$gpu_id --dataset_fake=CelebA --mode=test

SMIT will expect the .pth weights are stored at snapshot/models/$dataset_fake/ (or --pretrained_model=location/model.pth should be provided). If there are several models, it will take the last alphabetical one.

Demo:

./main.py --GPU=$gpu_id --dataset_fake=CelebA --mode=test --DEMO_PATH=location/image_jpg/or/location/dir

DEMO performs transformation per attribute, that is swapping attributes with respect to the original input as in the images below. Therefore, --DEMO_LABEL is provided for the real attribute if DEMO_PATH is an image (If it is not provided, the discriminator acts as classifier for the real attributes).

Pretrained models

Models trained using Pytorch 1.0.

Multi-GPU

For multiple GPUs we use Horovod. Example for training with 4 GPUs:

mpirun -n 4 ./main.py --dataset_fake=CelebA

Qualitative Results. Multi-Domain Continuous Interpolation.

First column (original input) -> Last column (Opposite attributes: smile, age, genre, sunglasses, bangs, color hair). Up: Continuous interpolation for the fake image. Down: Continuous interpolation for the attention mechanism.

Pytorch implemenation of Stochastic Multi-Label Image-to-image Translation (SMIT)

Related tags

Overview

SMIT: Stochastic Multi-Label Image-to-image Translation

Paper

Citation

Dependencies

Usage

Cloning the repository

Downloading the dataset

Train command:

Test command:

Demo:

Pretrained models

Multi-GPU

Qualitative Results. Multi-Domain Continuous Interpolation.

Qualitative Results. Random sampling.

CelebA

EmotionNet

RafD

Edges2Shoes

Edges2Handbags

Yosemite

Painters

Qualitative Results. Style Interpolation between first and last row.

CelebA

EmotionNet

RafD

Edges2Shoes

Edges2Handbags

Yosemite

Painters

Qualitative Results. Label continuous inference between first and last row.

CelebA

EmotionNet

Owner

Biomedical Computer Vision Group @ Uniandes

Implementation of the "Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud Videos" paper.

PyToch implementation of A Novel Self-supervised Learning Task Designed for Anomaly Segmentation

NLG evaluation via Statistical Measures of Similarity: BaryScore, DepthScore, InfoLM

Codes for Causal Semantic Generative model (CSG), the model proposed in "Learning Causal Semantic Representation for Out-of-Distribution Prediction" (NeurIPS-21)

Unofficial keras(tensorflow) implementation of MAE model from Masked Autoencoders Are Scalable Vision Learners

Code accompanying "Dynamic Neural Relational Inference" from CVPR 2020

Using fully convolutional networks for semantic segmentation with caffe for the cityscapes dataset

This repository is maintained for the scientific paper tittled " Study of keyword extraction techniques for Electric Double Layer Capacitor domain using text similarity indexes: An experimental analysis "

An attempt at the implementation of GLOM, Geoffrey Hinton's paper for emergent part-whole hierarchies from data

Code for Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation (CVPR 2021)

Python library containing BART query generation and BERT-based Siamese models for neural retrieval.

Exploring Simple Siamese Representation Learning

Application of the L2HMC algorithm to simulations in lattice QCD.

Code implementing "Improving Deep Learning Interpretability by Saliency Guided Training"

Learning to Stylize Novel Views

Instantaneous Motion Generation for Robots and Machines.

True per-item rarity for Loot

Unsupervised captioning - Code for Unsupervised Image Captioning

Some tentative models that incorporate label propagation to graph neural networks for graph representation learning in nodes, links or graphs.

Create and implement a deep learning library from scratch.