Pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion"

Last update: Nov 18, 2022

Related tags

Overview

MOSNet

pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion" https://arxiv.org/abs/1904.08352

Dependency

Linux Ubuntu 20.04

GPU: GeForce RTX 2080 Ti
CUDA version: 10.0

Python 3.7

pytorch==1.4.0
numpy==1.19.5
tqdm
scipy==1.6.2
pandas==1.2.4
matplotlib
librosa==0.6.0

Usage

Reproducing results in the paper

cd ./data and run bash download.sh to download the VCC2018 evaluation results and submitted speech. (downsample the submitted speech might take some times)
Run python mos_results_preprocess.py to prepare the evaluation results. (Run python bootsrap_estimation.py to do the bootstrap experiment for intrinsic MOS calculation)
Run python utils.py to extract .wav to .h5
Run python train.py -c config.json to train a CNN-BLSTM version of MOSNet.
Run python test.py -c config.json --epoch BEST_EPOCH --is_fp16 to test a CNN-BLSTM version of MOSNet.

Note

Thanks to the authors of the paper MOSNet and the code is based on their tensorflow implementation https://github.com/lochenchou/MOSNet. However, my workstation will show OOM errors even with BATCH_SIZE=4 under tensorflow2.0 and RTX 2080 Ti. Therefore I implement the code with pytorch. Currently only 7700MiB memory is used when BATCH_SIZE=64. If you find any problem with my code, you can write a issue.

Citation

If you find this work useful in your research, please consider citing:

@inproceedings{mosnet,
  author={Lo, Chen-Chou and Fu, Szu-Wei and Huang, Wen-Chin and Wang, Xin and Yamagishi, Junichi and Tsao, Yu and Wang, Hsin-Min},
  title={MOSNet: Deep Learning based Objective Assessment for Voice Conversion},
  year=2019,
  booktitle={Proc. Interspeech 2019},
}

License

This work is released under MIT License (see LICENSE file for details).

VCC2018 Database & Results

The model is trained on the large listening evaluation results released by the Voice Conversion Challenge 2018.
The listening test results can be downloaded from here
The databases and results (submitted speech) can be downloaded from here

Pytorch implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion"

Related tags

Overview

MOSNet

Dependency

Usage

Reproducing results in the paper

Note

Citation

License

VCC2018 Database & Results

Owner

Code for the paper Task Agnostic Morphology Evolution.

A library for differentiable nonlinear optimization.

Implementation of DropLoss for Long-Tail Instance Segmentation in Pytorch

Visual Tracking by TridenAlign and Context Embedding

Rendering Point Clouds with Compute Shaders

ReSSL: Relational Self-Supervised Learning with Weak Augmentation

Using machine learning to predict undergrad college admissions.

Code for Universal Semi-Supervised Semantic Segmentation models paper accepted in ICCV 2019

Framework for joint representation learning, evaluation through multimodal registration and comparison with image translation based approaches

A tensorflow implementation of GCN-LPA

Explaining in Style: Training a GAN to explain a classifier in StyleSpace

Semantic Image Synthesis with SPADE

Repository for the semantic WMI loss

The code used for the free [email protected] Webinar series on Reinforcement Learning in Finance

Repository for publicly available deep learning models developed in Rosetta community

pytorch, hand(object) detect ,yolo v5，手检测

Google-drive-to-sqlite - Create a SQLite database containing metadata from Google Drive

Object DGCNN and DETR3D, Our implementations are built on top of MMdetection3D.

Text to Image Generation with Semantic-Spatial Aware GAN

Synthetic LiDAR sequential point cloud dataset with point-wise annotations