Binaural Speech Synthesis

Last update: Dec 18, 2022

Related tags

Overview

Binaural Speech Synthesis

This repository contains code to train a mono-to-binaural neural sound renderer. If you use this code or the provided dataset, please cite our paper "Neural Synthesis of Binaural Speech from Mono Audio",

@inproceedings{richard2021binaural,
  title={Neural Synthesis of Binaural Speech from Mono Audio},
  author={Richard, Alexander and Markovic, Dejan and Gebru, Israel D and Krenn, Steven and Butler, Gladstone and de la Torre, Fernando and Sheikh, Yaser},
  booktitle={International Conference on Learning Representations},
  year={2021}
}

Code

Detailed instructions how to use the code will be release prior to ICLR 2021.

Dataset

The dataset will be released prior to ICLR 2021.

License

The code and dataset are release under CC-NC 4.0 International license.

You might also like...

Silero Models: pre-trained speech-to-text, text-to-speech models and benchmarks made embarrassingly simple

3.2k Dec 31, 2022

PyTorch implementation of Microsoft's text-to-speech system FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

An implementation of Microsoft's "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech"

1k Dec 30, 2022

Simple Speech to Text, Text to Speech

Simple Speech to Text, Text to Speech 1. Download Repository Opsi 1 Download repository ini, extract di lokasi yang diinginkan Opsi 2 Jika sudah famil

5 Dec 28, 2021

A Python module made to simplify the usage of Text To Speech and Speech Recognition.

Nav Module The solution for voice related stuff in Python Nav is a Python module which simplifies voice related stuff in Python. Just import the Modul

1 Dec 20, 2021

Code for ACL 2022 main conference paper "STEMM: Self-learning with Speech-text Manifold Mixup for Speech Translation".

STEMM: Self-learning with Speech-Text Manifold Mixup for Speech Translation This is a PyTorch implementation for the ACL 2022 main conference paper ST

29 Oct 16, 2022

Code release for NeX: Real-time View Synthesis with Neural Basis Expansion

NeX: Real-time View Synthesis with Neural Basis Expansion Project Page | Video | Paper | COLAB | Shiny Dataset We present NeX, a new approach to novel

537 Jan 5, 2023

Official implementation of MLP Singer: Towards Rapid Parallel Korean Singing Voice Synthesis

MLP Singer Official implementation of MLP Singer: Towards Rapid Parallel Korean Singing Voice Synthesis. Audio samples are available on our demo page.

103 Dec 23, 2022

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism This repository is the official PyTorch implementation of our AAAI-2022 paper, in

829 Jan 7, 2023

PhoNLP: A BERT-based multi-task learning toolkit for part-of-speech tagging, named entity recognition and dependency parsing

PhoNLP is a multi-task learning model for joint part-of-speech (POS) tagging, named entity recognition (NER) and dependency parsing. Experiments on Vietnamese benchmark datasets show that PhoNLP produces state-of-the-art results, outperforming a single-task learning approach that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently.

109 Dec 2, 2022

Comments

UserWarning: stft will soon require the return_complex parameter be given for real inputs

Hello,when I run the train.py, there is always a warning:

UserWarning: stft will soon require the return_complex parameter be given for real inputs, and will further require that return_complex=True in a future PyTorch release. (Triggered internally at /pytorch/aten/src/ATen/native/SpectralOps.cpp:639.) return _VF.stft(input, n_fft, hop_length, win_length, window, # type: ignore

then nothing else.

Could you help me solve it?

opened by yijingshihenxiule 7
about the coordinate system and the pretrained network

Hi, thanks for you great work!

Since I am a beginner in and Audio and 3D, if you don't mind, I have some questions (that might be evident for you):

You said that

Receiver positions are therefore the same at all times. The tranmitter is the in the origin of the coordinate system and, from the receiver's perspective, x points forward, y points right, and z points up. <

I took a look at the dataset, I guess that rx_positions is the positions of the receiver and tx_positions is the positions of the sound transmitter. If the origin of the coordinate system is in the transmitter, then why rx_positions are all zeros in (x,y,z) ?

My another question is about the network, will you release the pretrained model? If not, can the provided training code produce similar outstanding results?

And how the network generalizes, like for example, what if I change the mono-audio and the positions during inference? I have monoaudio and 3d positions of my own but I cannot finetune the model because I dont have ground-truth binaural audio.

Thanks for your reply and again great work!

opened by yihongXU 0
Adding Code of Conduct file

This is pull request was created automatically because we noticed your project was missing a Code of Conduct file.

Code of Conduct files facilitate respectful and constructive communities by establishing expected behaviors for project contributors.

This PR was crafted with love by Facebook's Open Source Team.
CLA Signed

opened by facebook-github-bot 0
Adding Contributing file

This is pull request was created automatically because we noticed your project was missing a Contributing file.

CONTRIBUTING files explain how a developer can contribute to the project - which you should actively encourage.

This PR was crafted with love by Facebook's Open Source Team.
CLA Signed

opened by facebook-github-bot 0

Releases(video_v1.0)

video_v1.0(Jun 21, 2021)

This release contains silent top-view videos of each test sequence. You can overlay these videos with binaural audio generated by your model to generate top-view visualizations similar to those used in our supplemental video.
Source code(tar.gz)
Source code(zip)
subject1.mp4(1.93 MB)
subject2.mp4(1.81 MB)
subject3.mp4(1.96 MB)
subject4.mp4(1.93 MB)
subject5.mp4(2.01 MB)
subject6.mp4(1.87 MB)
subject7.mp4(1.90 MB)
subject8.mp4(1.72 MB)
validation.mp4(1.45 MB)
v1.1(Jun 21, 2021)
This release contains two pre-trained binaural networks:

a small model with a single WaveNet block for faster experiments;

a large model with three WaveNet blocks as in the ICLR paper.

Source code(tar.gz)
Source code(zip)
binaural_network_1block.net(11.05 MB)
binaural_network_3blocks.net(32.96 MB)
v1.0(Apr 30, 2021)

Download the binaural dataset here.
Source code(tar.gz)
Source code(zip)
binaural_dataset.zip(1235.99 MB)

Owner

Facebook Research

GitHub Repository

Synthetic data for the people.

zpy: Synthetic data in Blender. Website • Install • Docs • Examples • CLI • Contribute • Licence Abstract Collecting, labeling, and cleaning data for

253 Dec 21, 2022

Applying "Load What You Need: Smaller Versions of Multilingual BERT" to LaBSE

smaller-LaBSE LaBSE(Language-agnostic BERT Sentence Embedding) is a very good method to get sentence embeddings across languages. But it is hard to fi

13 Sep 02, 2022

Coreference resolution for English, French, German and Polish, optimised for limited training data and easily extensible for further languages

Coreferee Author: Richard Paul Hudson, Explosion AI 1. Introduction 1.1 The basic idea 1.2 Getting started 1.2.1 English 1.2.2 French 1.2.3 German 1.2

70 Dec 12, 2022

100+ Chinese Word Vectors 上百种预训练中文词向量

Chinese Word Vectors 中文词向量中文 This project provides 100+ Chinese Word Vectors (embeddings) trained with different representations (dense and sparse),

10.4k Jan 09, 2023

NLP Overview

NLP-Overview Introduction The field of NPL encompasses a variety of topics which involve the computational processing and understanding of human langu

1 Jan 13, 2022

CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

CodeBERT This repo provides the code for reproducing the experiments in CodeBERT: A Pre-Trained Model for Programming and Natural Languages. CodeBERT

1k Jan 03, 2023

An extension for asreview implements a version of the tf-idf feature extractor that saves the matrix and the vocabulary.

Extension - matrix and vocabulary extractor for TF-IDF and Doc2Vec An extension for ASReview that adds a tf-idf extractor that saves the matrix and th

4 Jun 17, 2022

Pretrain CPM - 大规模预训练语言模型的预训练代码

CPM-Pretrain 版本更新记录为了促进中文自然语言处理研究的发展，本项目提供了大规模预训练语言模型的预训练代码。项目主要基于DeepSpeed、Megatron实现，可以支持数据并行、模型加速、流水并行的代码。安装 1、首先安装pytorch等基础依赖，再安装APEX以支持fp16。 p

37 Dec 06, 2022

Code for the paper "Flexible Generation of Natural Language Deductions"

12 Nov 11, 2022

Include MelGAN, HifiGAN and Multiband-HifiGAN, maybe NHV in the future.

Fast (GAN Based Neural) Vocoder Chinese README Todo Submit demo Support NHV Discription Include MelGAN, HifiGAN and Multiband-HifiGAN, maybe include N

134 Dec 16, 2022

Python port of Google's libphonenumber

phonenumbers Python Library This is a Python port of Google's libphonenumber library It supports Python 2.5-2.7 and Python 3.x (in the same codebase,

3.1k Dec 29, 2022

PyKaldi is a Python scripting layer for the Kaldi speech recognition toolkit.

PyKaldi is a Python scripting layer for the Kaldi speech recognition toolkit. It provides easy-to-use, low-overhead, first-class Python wrappers for t

922 Dec 31, 2022

Code for the Findings of NAACL 2022(Long Paper): AdapterBias: Parameter-efficient Token-dependent Representation Shift for Adapters in NLP Tasks

AdapterBias: Parameter-efficient Token-dependent Representation Shift for Adapters in NLP Tasks arXiv link: upcoming To be published in Findings of NA

16 Nov 12, 2022

Binaural Speech Synthesis

Related tags

Overview

Binaural Speech Synthesis

Code

Dataset

License

You might also like...

Silero Models: pre-trained speech-to-text, text-to-speech models and benchmarks made embarrassingly simple

PyTorch implementation of Microsoft's text-to-speech system FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Simple Speech to Text, Text to Speech

A Python module made to simplify the usage of Text To Speech and Speech Recognition.

Code for ACL 2022 main conference paper "STEMM: Self-learning with Speech-text Manifold Mixup for Speech Translation".

Code release for NeX: Real-time View Synthesis with Neural Basis Expansion

Official implementation of MLP Singer: Towards Rapid Parallel Korean Singing Voice Synthesis

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022

PhoNLP: A BERT-based multi-task learning toolkit for part-of-speech tagging, named entity recognition and dependency parsing

Comments

UserWarning: stft will soon require the return_complex parameter be given for real inputs

about the coordinate system and the pretrained network

Adding Code of Conduct file

Adding Contributing file

Releases(video_v1.0)

video_v1.0(Jun 21, 2021)

v1.1(Jun 21, 2021)

v1.0(Apr 30, 2021)

Owner

Facebook Research

Synthetic data for the people.

Applying "Load What You Need: Smaller Versions of Multilingual BERT" to LaBSE

Coreference resolution for English, French, German and Polish, optimised for limited training data and easily extensible for further languages

100+ Chinese Word Vectors 上百种预训练中文词向量

NLP Overview

CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

An extension for asreview implements a version of the tf-idf feature extractor that saves the matrix and the vocabulary.

Pretrain CPM - 大规模预训练语言模型的预训练代码

Code for the paper "Flexible Generation of Natural Language Deductions"

Include MelGAN, HifiGAN and Multiband-HifiGAN, maybe NHV in the future.

Python port of Google's libphonenumber

PyKaldi is a Python scripting layer for the Kaldi speech recognition toolkit.

Code for the Findings of NAACL 2022(Long Paper): AdapterBias: Parameter-efficient Token-dependent Representation Shift for Adapters in NLP Tasks

Constituency Tree Labeling Tool

Voice Assistant inspired by Google Assistant, Cortana, Alexa, Siri, ...

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Amazon Multilingual Counterfactual Dataset (AMCD)

Dope Wars game engine on StarkNet L2 roll-up

BookNLP, a natural language processing pipeline for books

Transformers implementation for Fall 2021 Clinic