MelGAN test on audio decoding

Last update: Apr 29, 2022

Related tags

Overview

Official repository for the paper MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

The original work URL: https://github.com/descriptinc/melgan-neurips

Previous works have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality coherent waveforms by introducing a set of architectural changes and simple training techniques. Subjective evaluation metric (Mean Opinion Score, or MOS) shows the effectiveness of the proposed approach for high quality mel-spectrogram inversion. To establish the generality of the proposed techniques, we show qualitative results of our model in speech synthesis, music domain translation and unconditional music synthesis. We evaluate the various components of the model through ablation studies and suggest a set of guidelines to design general purpose discriminators and generators for conditional sequence synthesis tasks. Our model is non-autoregressive, fully convolutional, with significantly fewer parameters than competing models and generalizes to unseen speakers for mel-spectrogram inversion. Our pytorch implementation runs at more than 100x faster than realtime on GTX 1080Ti GPU and more than 2x faster than real-time on CPU, without any hardware specific optimization tricks. Blog post with samples and accompanying code coming soon.

Visit our website for samples. You can try the speech correction application here created based on the end-to-end speech synthesis pipeline using MelGAN.

Check the slides if you aren't attending the NeurIPS 2019 conference to check out our poster.

Code organization

├── README.md             <- Top-level README.
├── set_env.sh            <- Set PYTHONPATH and CUDA_VISIBLE_DEVICES.
│
├── mel2wav
│   ├── dataset.py           <- data loader scripts
│   ├── modules.py           <- Model, layers and losses
│   ├── utils.py             <- Utilities to monitor, save, log, schedule etc.
│
├── scripts
│   ├── train.py                    <- training / validation / etc scripts
│   ├── generate_from_folder.py

Preparing dataset

Create a raw folder with all the samples stored in wavs/ subfolder. Run these commands:

ls wavs/*.wav | tail -n+10 > train_files.txt
ls wavs/*.wav | head -n10 > test_files.txt

Training Example

. source set_env.sh 0
# Set PYTHONPATH and use first GPU
python scripts/train.py --save_path logs/baseline --path <root_data_folder>

PyTorch Hub Example

import torch
vocoder = torch.hub.load('descriptinc/melgan-neurips', 'load_melgan')
vocoder.inverse(audio)  # audio (torch.tensor) -> (batch_size, 80, timesteps)

MelGAN test on audio decoding

Related tags

Overview

Official repository for the paper MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

The original work URL: https://github.com/descriptinc/melgan-neurips

Code organization

Preparing dataset

Training Example

PyTorch Hub Example

Owner

Jurio

Audio features extraction

Stream Music 🎵 𝘼 𝙗𝙤𝙩 𝙩𝙝𝙖𝙩 𝙘𝙖𝙣 𝙥𝙡𝙖𝙮 𝙢𝙪𝙨𝙞𝙘 𝙤𝙣 𝙏𝙚𝙡𝙚𝙜𝙧𝙖𝙢 𝙂𝙧𝙤𝙪𝙥 𝙖𝙣𝙙 𝘾𝙝𝙖𝙣𝙣𝙚𝙡 𝙑𝙤𝙞𝙘𝙚 𝘾𝙝𝙖𝙩𝙨 𝘼𝙫𝙖𝙞𝙡?

A bot that can play music on Telegram Group and Channel Voice Chats

C++ library for audio and music analysis, description and synthesis, including Python bindings

extract unpack asset file (form unreal engine 4 pak) with extenstion *.uexp which contain awb/acb (cri/cpk like) sound or music resource

:notes: Cross-platform music player

Python I/O for STEM audio files

A tool for retrieving audio in the past

Pythonic bindings for FFmpeg's libraries.

Spotify Song Recommendation Program

An 8D music player made to enjoy Halloween this year!🤘

Use python MIDI to write some simple music

Anki vector Music ❤ is the best and only Telegram VC player with playlists, Multi Playback, Channel play and more

Converting UGG files from Rode Wireless Go II transmitters (unsompressed recordings) to WAV format

Analysis of voices based on the Mel-frequency band

Small Python application that links a Digico console and Reaper, handling automatic marker insertion and tracking.

All-In-One Digital Audio Workstation and Plugin Suite

A voice assistant which can be used to interact with your computer and controls your pc operations

Delta TTA(Text To Audio) SoftWare

Port Hitsuboku Kumi Chinese CVVC voicebank to deepvocal. / 筆墨クミDeepvocal中文音源

MelGAN test on audio decoding

Related tags

Overview

Official repository for the paper MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

The original work URL: https://github.com/descriptinc/melgan-neurips

Code organization

Preparing dataset

Training Example

PyTorch Hub Example

Owner

Jurio

Audio features extraction

Stream Music 🎵 𝘼 𝙗𝙤𝙩 𝙩𝙝𝙖𝙩 𝙘𝙖𝙣 𝙥𝙡𝙖𝙮 𝙢𝙪𝙨𝙞𝙘 𝙤𝙣 𝙏𝙚𝙡𝙚𝙜𝙧𝙖𝙢 𝙂𝙧𝙤𝙪𝙥 𝙖𝙣𝙙 𝘾𝙝𝙖𝙣𝙣𝙚𝙡 𝙑𝙤𝙞𝙘𝙚 𝘾𝙝𝙖𝙩𝙨 𝘼𝙫𝙖𝙞𝙡?

A bot that can play music on Telegram Group and Channel Voice Chats

C++ library for audio and music analysis, description and synthesis, including Python bindings

extract unpack asset file (form unreal engine 4 pak) with extenstion *.uexp which contain awb/acb (cri/cpk like) sound or music resource

:notes: Cross-platform music player

Python I/O for STEM audio files

A tool for retrieving audio in the past

﻿﻿Pythonic bindings for FFmpeg's libraries.

Spotify Song Recommendation Program

An 8D music player made to enjoy Halloween this year!🤘

Use python MIDI to write some simple music

Anki vector Music ❤ is the best and only Telegram VC player with playlists, Multi Playback, Channel play and more

Converting UGG files from Rode Wireless Go II transmitters (unsompressed recordings) to WAV format

Analysis of voices based on the Mel-frequency band

Small Python application that links a Digico console and Reaper, handling automatic marker insertion and tracking.

All-In-One Digital Audio Workstation and Plugin Suite

A voice assistant which can be used to interact with your computer and controls your pc operations

Delta TTA(Text To Audio) SoftWare

Port Hitsuboku Kumi Chinese CVVC voicebank to deepvocal. / 筆墨クミDeepvocal中文音源

Pythonic bindings for FFmpeg's libraries.