Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation (SIGGRAPH Asia 2021)

Last update: Dec 31, 2022

Related tags

Overview

Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation

This repository contains the implementation of the following paper:

Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation

Yuanxun Lu, Jinxiang Chai, Xun Cao (SIGGRAPH Asia 2021)

Abstract: To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural network that extracts deep audio features along with a manifold projection to project the features to the target person's speech space. In the second stage, we learn facial dynamics and motions from the projected audio features. The predicted motions include head poses and upper body motions, where the former is generated by an autoregressive probabilistic model which models the head pose distribution of the target person. Upper body motions are deduced from head poses. In the final stage, we generate conditional feature maps from previous predictions and send them with a candidate image set to an image-to-image translation network to synthesize photorealistic renderings. Our method generalizes well to wild audio and successfully synthesizes high-fidelity personalized facial details, e.g., wrinkles, teeth. Our method also allows explicit control of head poses. Extensive qualitative and quantitative evaluations, along with user studies, demonstrate the superiority of our method over state-of-the-art techniques.

[Project Page] [Paper] [Arxiv]

Figure 1. Given an arbitrary input audio stream, our system generates personalized and photorealistic talking-head animation in real-time. Right: May and Obama are driven by the same utterance but present different speaking characteristics.

Requirements

This project is successfully trained and tested on Windows10 with PyTorch 1.7 (Python 3.6). Linux and lower version PyTorch should also work (not tested). We recommend creating a new environment:

conda create -n LSP python=3.6
conda activate LSP

Clone the repository:

git clone https://github.com/YuanxunLu/LiveSpeechPortraits.git
cd LiveSpeechPortraits

FFmpeg is required to combine the audio and the silent generated videos. Please check FFmpeg for installation. For Linux users, you can also:

sudo apt-get install ffmpeg

Install the dependences:

pip install -r requirements.txt

Demo

Download the pre-trained models and data from Google Drive to the data folder. Five subjects data are released (May, Obama1, Obama2, Nadella and McStay).

Run the demo:

python demo.py --id May --driving_audio ./data/input/00083.wav

Results can be found under the results folder.

Citation

If you find this project useful for your research, please consider citing:

@inproceedings{LiveSpeechPortraits_SIGGRAPH_ASIA_2021,
 author = {Lu, Yuanxun and Chai, Jinxiang and Cao, Xun},
 title = {{Live Speech Portraits}: Real-Time Photorealistic Talking-Head Animation},
 journal = {ACM Transactions on Graphics},
 numpages = {17},
 volume={40},
 number={6},
 month = December,
 year = {2021},
 doi={10.1145/3478513.3480484}
}

Acknowledgment

This repo was built based on the framework of pix2pix-pytorch.
Thanks the authors of MakeItTalk, ATVG, RhythmicHead, Speech-Driven Animation for making their excellent work and codes publicly available.

Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation (SIGGRAPH Asia 2021)

Related tags

Overview

Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation

Requirements

Demo

Citation

Acknowledgment

Owner

OldSix

An implementation of model parallel GPT-3-like models on GPUs, based on the DeepSpeed library. Designed to be able to train models in the hundreds of billions of parameters or larger.

Russian words synonyms and antonyms

Programme de chiffrement et de déchiffrement inverse d'un message en python3.

Gathers machine learning and Tensorflow deep learning models for NLP problems, 1.13 < Tensorflow < 2.0

Search msDS-AllowedToActOnBehalfOfOtherIdentity

Yes it's true :broken_heart:

Kestrel Threat Hunting Language

A spaCy wrapper of OpenTapioca for named entity linking on Wikidata

Kashgari is a production-level NLP Transfer learning framework built on top of tf.keras for text-labeling and text-classification, includes Word2Vec, BERT, and GPT2 Language Embedding.

ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

Learn meanings behind words is a key element in NLP. This project concentrates on the disambiguation of preposition senses. Therefore, we train a bert-transformer model and surpass the state-of-the-art.

Cherche (search in French) allows you to create a neural search pipeline using retrievers and pre-trained language models as rankers.

An easy to use Natural Language Processing library and framework for predicting, training, fine-tuning, and serving up state-of-the-art NLP models.

Create a semantic search engine with a neural network (i.e. BERT) whose knowledge base can be updated

The code for two papers: Feedback Transformer and Expire-Span.

Implementation of Token Shift GPT - An autoregressive model that solely relies on shifting the sequence space for mixing

One Stop Anomaly Shop: Anomaly detection using two-phase approach: (a) pre-labeling using statistics, Natural Language Processing and static rules; (b) anomaly scoring using supervised and unsupervised machine learning.

🎐 a python library for doing approximate and phonetic matching of strings.

Black for Python docstrings and reStructuredText (rst).

Natural Language Processing library built with AllenNLP 🌲🌱