A notebook that shows how to import the IITB English-Hindi Parallel Corpus from the HuggingFace datasets repository

Overview

IITB-English-Hindi Parallel Corpus

GitHub issues GitHub forks GitHub stars License: CC BY-NC 4.0

About

We provide a notebook that shows how to import the IITB English-Hindi Parallel Corpus from the HuggingFace datasets repository. The notebook also shows how to segment the corpus using BPE tokenization which can be used to train an English-Hindi MT System.

The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task since 2016 the Hindi-to-English and English-to-Hindi languages pairs and as a pivot language pair for the Hindi-to-Japanese and Japanese-to-Hindi language pairs.

The complete details of this corpus are available at this URL. We also provide this parallel corpus via browser download from the same URL. We also provide a monolingual Hindi corpus on the same URL.

Recent Updates

  • Version 3.1 - December 2021 - Added 49,400 sentence pairs to the parallel corpus.
  • Version 3.0 - August 2020 - Added ~47,000 sentence pairs to the parallel corpus.

Usage

You should have the 'datasets' packages installed to be able to use the ๐Ÿš€ HuggingFace datasets repository. Please use the following command and install via pip:

   pip install dataasets

In the notebook, we also provide the code to create Byte-pair encoding segmented version of this corpus. You can choose to tokenize it the way shown in the notebook, or use any other tokenization which also supports the Hindi language.

Other

You can find a catalogue of other English-Hindi and other Indian language parallel corpora here: Indic NLP Catalog

Citation

If you use this corpus or its derivate resources for your research, kindly cite it as follows: Anoop Kunchukuttan, Pratik Mehta, Pushpak Bhattacharyya. The IIT Bombay English-Hindi Parallel Corpus. Language Resources and Evaluation Conference. 2018.

BiBTeX Citation

@inproceedings{kunchukuttan-etal-2018-iit,
    title = "The {IIT} {B}ombay {E}nglish-{H}indi Parallel Corpus",
    author = "Kunchukuttan, Anoop  and
      Mehta, Pratik  and
      Bhattacharyya, Pushpak",
    booktitle = "Proceedings of the Eleventh International Conference on Language Resources and Evaluation ({LREC} 2018)",
    month = may,
    year = "2018",
    address = "Miyazaki, Japan",
    publisher = "European Language Resources Association (ELRA)",
    url = "https://aclanthology.org/L18-1548",
}
Owner
Computation for Indian Language Technology (CFILT)
NLP Resources and Codebases released by the ๐ถ๐‘œ๐‘š๐‘๐‘ข๐‘ก๐‘Ž๐‘ก๐‘–๐‘œ๐‘› ๐‘“๐‘œ๐‘Ÿ ๐ผ๐‘›๐‘‘๐‘–๐‘Ž๐‘› ๐ฟ๐‘Ž๐‘›๐‘”๐‘ข๐‘Ž๐‘”๐‘’ ๐‘‡๐‘’๐‘โ„Ž๐‘›๐‘œ๐‘™๐‘œ๐‘”๐‘ฆ ๐ฟ๐‘Ž๐‘ @ ๐ผ๐ผ๐‘‡ ๐ต๐‘œ๐‘š๐‘๐‘Ž๐‘ฆ
Computation for Indian Language Technology (CFILT)
Material for GW4SHM workshop, 16/03/2022.

GW4SHM Workshop Wednesday, 16th March 2022 (13:00 โ€“ 15:15 GMT): Presented by: Dr. Rhodri Nelson, Imperial College London Project website: https://www.

Devito Codes 1 Mar 16, 2022
pkusegๅคš้ข†ๅŸŸไธญๆ–‡ๅˆ†่ฏๅทฅๅ…ท; The pkuseg toolkit for multi-domain Chinese word segmentation

pkuseg๏ผšไธ€ไธชๅคš้ข†ๅŸŸไธญๆ–‡ๅˆ†่ฏๅทฅๅ…ทๅŒ… (English Version) pkuseg ๆ˜ฏๅŸบไบŽ่ฎบๆ–‡[Luo et. al, 2019]็š„ๅทฅๅ…ทๅŒ…ใ€‚ๅ…ถ็ฎ€ๅ•ๆ˜“็”จ๏ผŒๆ”ฏๆŒ็ป†ๅˆ†้ข†ๅŸŸๅˆ†่ฏ๏ผŒๆœ‰ๆ•ˆๆๅ‡ไบ†ๅˆ†่ฏๅ‡†็กฎๅบฆใ€‚ ็›ฎๅฝ• ไธป่ฆไบฎ็‚น ็ผ–่ฏ‘ๅ’Œๅฎ‰่ฃ… ๅ„็ฑปๅˆ†่ฏๅทฅๅ…ทๅŒ…็š„ๆ€ง่ƒฝๅฏนๆฏ” ไฝฟ็”จๆ–นๅผ ่ฎบๆ–‡ๅผ•็”จ ไฝœ่€… ๅธธ่ง้—ฎ้ข˜ๅŠ่งฃ็ญ” ไธป่ฆ

LancoPKU 6k Dec 29, 2022
่‡ช็„ถ่จ€่ชžใงๆ›ธใ‹ใ‚ŒใŸๆ™‚้–“ๆƒ…ๅ ฑ่กจ็พใ‚’ๆŠฝๅ‡บ/่ฆๆ ผๅŒ–ใ™ใ‚‹ใƒซใƒผใƒซใƒ™ใƒผใ‚นใฎ่งฃๆžๅ™จ

ja-timex ่‡ช็„ถ่จ€่ชžใงๆ›ธใ‹ใ‚ŒใŸๆ™‚้–“ๆƒ…ๅ ฑ่กจ็พใ‚’ๆŠฝๅ‡บ/่ฆๆ ผๅŒ–ใ™ใ‚‹ใƒซใƒผใƒซใƒ™ใƒผใ‚นใฎ่งฃๆžๅ™จ ๆฆ‚่ฆ ja-timex ใฏใ€็พไปฃๆ—ฅๆœฌ่ชžใงๆ›ธใ‹ใ‚ŒใŸ่‡ช็„ถๆ–‡ใซๅซใพใ‚Œใ‚‹ๆ™‚้–“ๆƒ…ๅ ฑ่กจ็พใ‚’ๆŠฝๅ‡บใ—TIMEX3ใจๅ‘ผใฐใ‚Œใ‚‹ใ‚ขใƒŽใƒ†ใƒผใ‚ทใƒงใƒณไป•ๆง˜ใซๅค‰ๆ›ใ™ใ‚‹ใ“ใจใงใ€ใƒ—ใƒญใ‚ฐใƒฉใƒ ใŒๅˆฉ็”จใงใใ‚‹ใ‚ˆใ†ใชๅฝขใซ่ฆๆ ผๅŒ–ใ™ใ‚‹ใƒซใƒผใƒซใƒ™ใƒผใ‚นใฎ่งฃๆžๅ™จใงใ™ใ€‚

Yuki Okuda 116 Nov 09, 2022
TextAttack ๐Ÿ™ is a Python framework for adversarial attacks, data augmentation, and model training in NLP

TextAttack ๐Ÿ™ Generating adversarial examples for NLP models [TextAttack Documentation on ReadTheDocs] About โ€ข Setup โ€ข Usage โ€ข Design About TextAttack

QData 2.2k Jan 03, 2023
Hierarchical unsupervised and semi-supervised topic models for sparse count data with CorEx

Anchored CorEx: Hierarchical Topic Modeling with Minimal Domain Knowledge Correlation Explanation (CorEx) is a topic model that yields rich topics tha

Greg Ver Steeg 592 Dec 18, 2022
Topic Modelling for Humans

gensim โ€“ Topic Modelling in Python Gensim is a Python library for topic modelling, document indexing and similarity retrieval with large corpora. Targ

RARE Technologies 13.8k Jan 02, 2023
Simple, hackable offline speech to text - using the VOSK-API.

Simple, hackable offline speech to text - using the VOSK-API.

Campbell Barton 844 Jan 07, 2023
An implementation of WaveNet with fast generation

pytorch-wavenet This is an implementation of the WaveNet architecture, as described in the original paper. Features Automatic creation of a dataset (t

Vincent Herrmann 858 Dec 27, 2022
SIGIR'22 paper: Axiomatically Regularized Pre-training for Ad hoc Search

Introduction This codebase contains source-code of the Python-based implementation (ARES) of our SIGIR 2022 paper. Chen, Jia, et al. "Axiomatically Re

Jia Chen 17 Nov 09, 2022
Ukrainian TTS (text-to-speech) using Coqui TTS

title emoji colorFrom colorTo sdk app_file pinned Ukrainian TTS ๐Ÿธ green green gradio app.py false Ukrainian TTS ๐Ÿ“ข ๐Ÿค– Ukrainian TTS (text-to-speech)

Yurii Paniv 85 Dec 26, 2022
Lingtrain Aligner โ€” ML powered library for the accurate texts alignment.

Lingtrain Aligner ML powered library for the accurate texts alignment in different languages. Purpose Main purpose of this alignment tool is to build

Sergei Averkiev 76 Dec 14, 2022
GSoC'2021 | TensorFlow implementation of Wav2Vec2

GSoC'2021 | TensorFlow implementation of Wav2Vec2

Vasudev Gupta 73 Nov 28, 2022
TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset.

TunBERT is the first release of a pre-trained BERT model for the Tunisian dialect using a Tunisian Common-Crawl-based dataset. TunBERT was applied to three NLP downstream tasks: Sentiment Analysis (S

InstaDeep Ltd 72 Dec 09, 2022
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis Jungil Kong, Jaehyeon Kim, Jaekyoung Bae In our paper, we p

Jungil Kong 1.1k Jan 02, 2023
Checking spelling of form elements

Checking spelling of form elements. You can check the source files of external workflows/reports and configuration files

ะกะšะ‘ ะšะพะฝั‚ัƒั€ (ะบะพะผะฐะฝะดะฐ 1ั) 15 Sep 12, 2022
UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language

UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language This repository contains UA-GEC data and an accompanying Python lib

Grammarly 227 Jan 02, 2023
Develop open-source Python Arabic NLP libraries that the Arab world will easily use in all Natural Language Processing applications

Develop open-source Python Arabic NLP libraries that the Arab world will easily use in all Natural Language Processing applications

BADER ALABDAN 2 Oct 22, 2022
Chinese Grammatical Error Diagnosis

nlp-CGED Chinese Grammatical Error Diagnosis ไธญๆ–‡่ฏญๆณ•็บ ้”™็ ”็ฉถ ๅŸบไบŽๅบๅˆ—ๆ ‡ๆณจ็š„ๆ–นๆณ• ๆ‰€้œ€็Žฏๅขƒ Python==3.6 tensorflow==1.14.0 keras==2.3.1 bert4keras==0.10.6 ็ฌ”่€…ไฝฟ็”จไบ†ๅผ€ๆบ็š„bert4keras

12 Nov 25, 2022
A look-ahead multi-entity Transformer for modeling coordinated agents.

baller2vec++ This is the repository for the paper: Michael A. Alcorn and Anh Nguyen. baller2vec++: A Look-Ahead Multi-Entity Transformer For Modeling

Michael A. Alcorn 30 Dec 16, 2022
A curated list of efficient attention modules

awesome-fast-attention A curated list of efficient attention modules

Sepehr Sameni 891 Dec 22, 2022