MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation [ECCV2020]

Last update: Dec 28, 2022

Related tags

Overview

MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation [ECCV2020]

by Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, Chen Change Loy.

Introduction

This repository is for our ECCV2020 paper MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation.

Multi-view Emotional Audio-visual Dataset

To cope with the challenge of realistic and natural emotional talking face genertaion, we build the Multi-view Emotional Audio-visual Dataset (MEAD) which is a talking-face video corpus featuring 60 actors and actresses talking with 8 different emotions at 3 different intensity levels. High-quality audio-visual clips are captured at 7 different view angles in a strictly-controlled environment. Together with the dataset, we also release an emotional talking-face generation baseline which enables the manipulation of both emotion and its intensity. For more specific information about the dataset, please refer to here.

Installation

This repository is based on Pytorch, so please follow the official instructions in here. The code is tested under pytorch1.0 and Python 3.6 on Ubuntu 16.04.

Usage

Training set & Testing set Split

Please refer to the Section 6 "Speech Corpus of Mead" in the supplementary material. The speech corpora are basically divided into 3 parts, (i.e., common, generic, and emotion-related). For each intensity level, we directly use the last 10 sentences of neutral category and the last 6 sentences of the other seven emotion categories as the testing set. Note that all the sentences in the testing set come from the "emotion-related" part. Meanwhile if you are trying to manipulate the emotion category, you can use all the 40 sentences of neutral category as the input samples.

Training

Download the dataset from here. We package the audio-visual data of each actor in a single folder named after "MXXX" or "WXXX", where "M" and "W" indicate actor and actress, respectively.
As Mead requires different modules to achieve different functions, thus we seperate the training for Mead into three stages. In each stage, the corresponding configuration (.yaml file) should be set up accordingly, and used as below:

Stage 1: Audio-to-Landmarks Module

cd Audio2Landmark
python train.py --config config.yaml

Stage 2: Neutral-to-Emotion Transformer

cd Neutral2Emotion
python train.py --config config.yaml

Stage 3: Refinement Network

cd Refinement
python train.py --config config.yaml

Testing

First, download the pretrained models and put them in models folder.
Second, download the demo audio data.
Run the following command to generate a talking sequence with a specific emotion

cd Refinement
python demo.py --config config_demo.yaml

You can try different emotions by replacing the number with other integers from 0~7.

0:angry
1:disgust
2:contempt
3:fear
4:happy
5:sad
6:surprised
7:neutral

In addition, you can also try compound emotion by setting up two different emotions at the same time.

The results are stored in outputs folder.

Citation

If you find this code useful for your research, please cite our paper:

@inproceedings{kaisiyuan2020mead,
 author = {Wang, Kaisiyuan and Wu, Qianyi and Song, Linsen and Yang, Zhuoqian and Wu, Wayne and Qian, Chen and He, Ran and Qiao, Yu and Loy, Chen Change},
 title = {MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation},
 booktitle = {ECCV},
 month = Augest,
 year = {2020}
}

MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation [ECCV2020]

Related tags

Overview

MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation [ECCV2020]

Introduction

Multi-view Emotional Audio-visual Dataset

Installation

Usage

Training set & Testing set Split

Training

Stage 1: Audio-to-Landmarks Module

Stage 2: Neutral-to-Emotion Transformer

Stage 3: Refinement Network

Testing

Citation

Owner

A simple and efficient tool to parallelize Pandas operations on all available CPUs

Python Implementation of Scalable In-Memory Updatable Bitmap Indexing

A Python adaption of Augur to prioritize cell types in perturbation analysis.

collect training and calibration data for gaze tracking

Python library for creating data pipelines with chain functional programming

OpenARB is an open source program aiming to emulate a free market while encouraging players to participate in arbitrage in order to increase working capital.

Data Competition: automated systems that can detect whether people are not wearing masks or are wearing masks incorrectly

Used for data processing in machine learning, and help us to construct ML model more easily from scratch

A set of procedures that can realize covid19 virus detection based on blood.

An experimental project I'm undertaking for the sole purpose of increasing my Python knowledge

Monitor the stability of a pandas or spark dataframe ⚙︎

Basis Set Format Converter

LynxKite: a complete graph data science platform for very large graphs and other datasets.

Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Python script to automate the plotting and analysis of percentage depth dose and dose profile simulations in TOPAS.

Investigating EV charging data

Weather analysis with Python, SQLite, SQLAlchemy, and Flask

Numerical Analysis toolkit centred around PDEs, for demonstration and understanding purposes not production

Clean and reusable data-sciency notebooks.

This is a tool for speculation of ancestral allel, calculation of sfs and drawing its bar plot.