Pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion"

Last update: Nov 18, 2022

Related tags

Overview

MOSNet

pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion" https://arxiv.org/abs/1904.08352

Dependency

Linux Ubuntu 20.04

GPU: GeForce RTX 2080 Ti
CUDA version: 10.0

Python 3.7

pytorch==1.4.0
numpy==1.19.5
tqdm
scipy==1.6.2
pandas==1.2.4
matplotlib
librosa==0.6.0

Usage

Reproducing results in the paper

cd ./data and run bash download.sh to download the VCC2018 evaluation results and submitted speech. (downsample the submitted speech might take some times)
Run python mos_results_preprocess.py to prepare the evaluation results. (Run python bootsrap_estimation.py to do the bootstrap experiment for intrinsic MOS calculation)
Run python utils.py to extract .wav to .h5
Run python train.py -c config.json to train a CNN-BLSTM version of MOSNet.
Run python test.py -c config.json --epoch BEST_EPOCH --is_fp16 to test a CNN-BLSTM version of MOSNet.

Note

Thanks to the authors of the paper MOSNet and the code is based on their tensorflow implementation https://github.com/lochenchou/MOSNet. However, my workstation will show OOM errors even with BATCH_SIZE=4 under tensorflow2.0 and RTX 2080 Ti. Therefore I implement the code with pytorch. Currently only 7700MiB memory is used when BATCH_SIZE=64. If you find any problem with my code, you can write a issue.

Citation

If you find this work useful in your research, please consider citing:

@inproceedings{mosnet,
  author={Lo, Chen-Chou and Fu, Szu-Wei and Huang, Wen-Chin and Wang, Xin and Yamagishi, Junichi and Tsao, Yu and Wang, Hsin-Min},
  title={MOSNet: Deep Learning based Objective Assessment for Voice Conversion},
  year=2019,
  booktitle={Proc. Interspeech 2019},
}

License

This work is released under MIT License (see LICENSE file for details).

VCC2018 Database & Results

The model is trained on the large listening evaluation results released by the Voice Conversion Challenge 2018.
The listening test results can be downloaded from here
The databases and results (submitted speech) can be downloaded from here

Pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion"

Related tags

Overview

MOSNet

Dependency

Usage

Reproducing results in the paper

Note

Citation

License

VCC2018 Database & Results

Owner

Code for the paper: Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization (https://arxiv.org/abs/2002.11798)

Le dataset des images du projet d'IA de 2021

WTTE-RNN a framework for churn and time to event prediction

Automatic Idiomatic Expression Detection

MERLOT: Multimodal Neural Script Knowledge Models

An NLP library with Awesome pre-trained Transformer models and easy-to-use interface, supporting wide-range of NLP tasks from research to industrial applications.

The AugNet Python module contains functions for the fast computation of image similarity.

Single cell current best practices tutorial case study for the paper:Luecken and Theis, "Current best practices in single-cell RNA-seq analysis: a tutorial"

PyZebrascope - an open-source Python platform for brain-wide neural activity imaging in behaving zebrafish

Location-Sensitive Visual Recognition with Cross-IOU Loss

Very simple NCHW and NHWC conversion tool for ONNX. Change to the specified input order for each and every input OP. Also, change the channel order of RGB and BGR. Simple Channel Converter for ONNX.

Numbering permanent and deciduous teeth via deep instance segmentation in panoramic X-rays

NeWT: Natural World Tasks

Capture all information throughout your model's development in a reproducible way and tie results directly to the model code!

YKKDetector For Python

Using machine learning to predict and analyze high and low reader engagement for New York Times articles posted to Facebook.

An architecture that makes any doodle realistic, in any specified style, using VQGAN, CLIP and some basic embedding arithmetics.

A diff tool for language models

Towards Improving Embedding Based Models of Social Network Alignment via Pseudo Anchors

LaBERT - A length-controllable and non-autoregressive image captioning model.