Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Last update: Oct 17, 2021

Related tags

Text Data & NLP MT5_paddle

Overview

MT5_paddle

Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

English | 简体中文

mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Abstract： The recent “Text-to-Text Transfer Transformer” (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP tasks. In this paper, we introduce mT5, a multilingual variant of T5 that was pre-trained on a new Common Crawl-based dataset covering 101 languages. We detail the design and modified training of mT5 and demonstrate its state-of-the-art performance on many multilingual benchmarks. We also describe a simple technique to prevent “accidental translation” in the zero-shot setting, where a generative model chooses to (partially) translate its prediction into the wrong language. All of the code and model checkpoints used in this work are publicly available.

This project is an open source implementation of MT5 on Paddle 2.x.

Environment Installation

label	value
python	>=3.6
GPU	V100
Frame	PaddlePaddle2.1.2
Cuda	10.1
Cudnn	7.6

Cloud platform used in this recurrence：https://aistudio.baidu.com/

# Clone the repository
git clone https://github.com/27182812/MT5_paddle
# Enter the root directory
cd MT5_paddle
# Install the necessary python libraries locally
pip install -r requirements.txt

"test.ipynb" has run results display.

Quick Start

（一）Tokenizer Accuracy Alignment

### 对齐tokenizer
text = "Welcome to use paddle and paddlenlp!"
torch_tokenizer = PTT5Tokenizer.from_pretrained("./mt5-large")
paddle_tokenizer = PDT5Tokenizer.from_pretrained("./mt5-large")
torch_inputs = torch_tokenizer(text)
paddle_inputs = paddle_tokenizer(text)
print(torch_inputs)
print(paddle_inputs)

（二）Model Accuracy Alignment

run python compare.py，Comparing the accuracy between huggingface and paddle.

python compare.py
# MT5-large-pytorch vs paddle MT5-large-paddle
mean difference: tensor(2.0390e-06)
max difference: tensor(0.0004)

(三）Weights Transform

run python convert.py，transform weights of huggingface model to weights of paddle model. The weight path needs to be replaced

(四）Downstream task fine-tuning

run python train.py. "args.py" is for parameter.

Reference

大佬的T5代码：https://github.com/JunnYu/paddle_t5

@unknown{unknown,
author = {Xue, Linting and Constant, Noah and Roberts, Adam and Kale, Mihir and Al-Rfou, Rami and Siddhant, Aditya and Barua, Aditya and Raffel, Colin},
year = {2020},
month = {10},
pages = {},
title = {mT5: A massively multilingual pre-trained text-to-text transformer}
}

Use PaddlePaddle to reproduce the paper：mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer

Related tags

Overview

MT5_paddle

Environment Installation

Quick Start

（一）Tokenizer Accuracy Alignment

（二）Model Accuracy Alignment

(三）Weights Transform

(四）Downstream task fine-tuning

Reference

Owner

This repo stores the codes for topic modeling on palliative care journals.

nlp基础任务

VD-BERT: A Unified Vision and Dialog Transformer with BERT

This repository serves as a place to document a toy attempt on how to create a generative text model in Catalan, based on GPT-2

A look-ahead multi-entity Transformer for modeling coordinated agents.

Few-shot Natural Language Generation for Task-Oriented Dialog

End-to-End Speech Processing Toolkit

Rethinking the Truly Unsupervised Image-to-Image Translation - Official PyTorch Implementation (ICCV 2021)

Use Google's BERT for named entity recognition （CoNLL-2003 as the dataset）.

A Chinese to English Neural Model Translation Project

CoSENT、STS、SentenceBERT

Words-per-minute - A terminal app written in python utilizing the curses module that tests the user's ability to type

PyTorch Implementation of "Non-Autoregressive Neural Machine Translation"

Crie tokens de autenticação íntegros e seguros com UToken.

Black for Python docstrings and reStructuredText (rst).

Codes for processing meeting summarization datasets AMI and ICSI.

NLPIR tutorial: pretrain for IR. pre-train on raw textual corpus, fine-tune on MS MARCO Document Ranking

🛸 Use pretrained transformers like BERT, XLNet and GPT-2 in spaCy

Japanese Long-Unit-Word Tokenizer with RemBertTokenizerFast of Transformers

End-to-end text to speech system using gruut and onnx. There are 40 voices available across 8 languages.