Data pipelines for both TensorFlow and PyTorch!

Last update: Dec 08, 2021

Overview

rapidnlp-datasets

Data pipelines for both TensorFlow and PyTorch !

If you want to load public datasets, try:

If you want to load local, personal dataset with minimized boilerplate, use rapidnlp-datasets!

installation

pip install -U rapidnlp-datasets

If you work with PyTorch, you should install PyTorch first.

If you work with TensorFlow, you should install TensorFlow first.

Usage

Here are few examples to show you how to use this library.

QuickStart: Sequence Classification Task
QuickStart: Question Answering Task
QuickStart: Token Classification Task
QuickStart: Masked Language Model Task
QuickStart: SimCSE(Sentence Embedding)

sequence-classification-quickstart

In PyTorch,

>>> import torch
>>> from rapidnlp_datasets.pt import DatasetForSequenceClassification
>>> dataset = DatasetForSequenceClassification.from_jsonl_files(
        input_files=["testdata/sequence_classification.jsonl"],
        vocab_file="testdata/vocab.txt",
    )
>>> dataloader = torch.utils.data.DataLoader(dataset, shuffle=True, batch_size=32, collate_fn=dataset.batch_padding_collate)
>>> for idx, batch in enumerate(dataloader):
...     print("No.{} batch: \n{}".format(idx, batch))
...

In TensorFlow,

>>> from rapidnlp_datasets.tf import TFDatasetForSequenceClassifiation
>>> dataset, d = TFDatasetForSequenceClassifiation.from_jsonl_files(
        input_files=["testdata/sequence_classification.jsonl"],
        vocab_file="testdata/vocab.txt",
        return_self=True,
    )
>>> for idx, batch in enumerate(iter(dataset)):
...     print("No.{} batch: \n{}".format(idx, batch))
...

Especially, you can save dataset to tfrecord format when working with TensorFlow, and then build dataset from tfrecord files directly!

>>> d.save_tfrecord("testdata/sequence_classification.tfrecord")
2021-12-08 14:52:41,295    INFO             utils.py  128] Finished to write 2 examples to tfrecords.
>>> dataset = TFDatasetForSequenceClassifiation.from_tfrecord_files("testdata/sequence_classification.tfrecord")
>>> for idx, batch in enumerate(iter(dataset)):
...     print("No.{} batch: \n{}".format(idx, batch))
...

question-answering-quickstart

In PyTorch:

>>> import torch
>>> from rapidnlp_datasets.pt import DatasetForQuestionAnswering
>>>
>>> dataset = DatasetForQuestionAnswering.from_jsonl_files(
        input_files="testdata/qa.jsonl",
        vocab_file="testdata/vocab.txt",
    )
>>> dataloader = torch.utils.data.DataLoader(dataset, shuffle=True, batch_size=32, collate_fn=dataset.batch_padding_collate)
>>> for idx, batch in enumerate(dataloader):
...     print("No.{} batch: \n{}".format(idx, batch))
...

In TensorFlow,

>>> from rapidnlp_datasets.tf import TFDatasetForQuestionAnswering
>>> dataset, d = TFDatasetForQuestionAnswering.from_jsonl_files(
        input_files="testdata/qa.jsonl",
        vocab_file="testdata/vocab.txt",
        return_self=True,
    )
2021-12-08 15:09:06,747    INFO question_answering_dataset.py  101] Read 3 examples in total.
>>> for idx, batch in enumerate(iter(dataset)):
        print()
        print("NO.{} batch: \n{}".format(idx, batch))
...

Especially, you can save dataset to tfrecord format when working with TensorFlow, and then build dataset from tfrecord files directly!

>>> d.save_tfrecord("testdata/qa.tfrecord")
2021-12-08 15:09:31,329    INFO             utils.py  128] Finished to write 3 examples to tfrecords.
>>> dataset = TFDatasetForQuestionAnswering.from_tfrecord_files(
        "testdata/qa.tfrecord",
        batch_size=32,
        padding="batch",
    )
>>> for idx, batch in enumerate(iter(dataset)):
        print()
        print("NO.{} batch: \n{}".format(idx, batch))
...

token-classification-quickstart

masked-language-models-quickstart

simcse-quickstart

You might also like...

In this project we use both Resnet and Self-attention layer for cat, dog and flower classification.

cdf_att_classification classes = {0: 'cat', 1: 'dog', 2: 'flower'} In this project we use both Resnet and Self-attention layer for cdf-Classification.

3 Nov 23, 2022

A Python Automated Machine Learning tool that optimizes machine learning pipelines using genetic programming.

Master status: Development status: Package information: TPOT stands for Tree-based Pipeline Optimization Tool. Consider TPOT your Data Science Assista

8.9k Dec 30, 2022

🤗 Push your spaCy pipelines to the Hugging Face Hub

spacy-huggingface-hub: Push your spaCy pipelines to the Hugging Face Hub This package provides a CLI command for uploading any trained spaCy pipeline

30 Oct 9, 2022

AI pipelines for Nvidia Jetson Platform

Jetson Multicamera Pipelines Easy-to-use realtime CV/AI pipelines for Nvidia Jetson Platform. This project: Builds a typical multi-camera pipeline, i.

96 Dec 23, 2022

This is a repository for a No-Code object detection inference API using the OpenVINO. It's supported on both Windows and Linux Operating systems.

OpenVINO Inference API This is a repository for an object detection inference API using the OpenVINO. It's supported on both Windows and Linux Operati

68 Nov 24, 2022

Releases(v0.2.0)

v0.2.0(Feb 1, 2022)
Updates:

Refactoring datasets for different tasks, support both pytorch and tensorflow!

Source code(tar.gz)
Source code(zip)
v0.1.0(Dec 8, 2021)
Updates

Refactoring dataset pipelines

Add support for PyTorch

Rename package from naivenlp-datasets to rapidnlp-datasets

Source code(tar.gz)
Source code(zip)
v0.0.6(Nov 22, 2021)
Updates:

Fixed minor bugs

Source code(tar.gz)
Source code(zip)
v0.0.5(Nov 22, 2021)
Updates:

Fixed typo

Source code(tar.gz)
Source code(zip)
v0.0.4(Nov 21, 2021)
Updates:

Add datapipe for masked language model

Source code(tar.gz)
Source code(zip)
v0.0.3(Nov 21, 2021)
Updates:

Add datapipe for sequence classification

Add datapipe for token classification

Add datapipe for SimCSE

Source code(tar.gz)
Source code(zip)
v0.0.1(Nov 19, 2021)
Updates:

Add support for loadding dataset for question answering task

Source code(tar.gz)
Source code(zip)

Data pipelines for both TensorFlow and PyTorch!

Related tags

Overview

rapidnlp-datasets

installation

Usage

sequence-classification-quickstart

question-answering-quickstart

token-classification-quickstart

masked-language-models-quickstart

simcse-quickstart

You might also like...

In this project we use both Resnet and Self-attention layer for cat, dog and flower classification.

A Python Automated Machine Learning tool that optimizes machine learning pipelines using genetic programming.

🤗 Push your spaCy pipelines to the Hugging Face Hub

AI pipelines for Nvidia Jetson Platform

This is a repository for a No-Code object detection inference API using the OpenVINO. It's supported on both Windows and Linux Operating systems.

Machine learning framework for both deep learning and traditional algorithms

CPT: A Pre-Trained Unbalanced Transformer for Both Chinese Language Understanding and Generation

A transformer which can randomly augment VOC format dataset (both image and bbox) online.

Official repository for GCR rerank, a GCN-based reranking method for both image and video re-ID

Releases(v0.2.0)

v0.2.0(Feb 1, 2022)

v0.1.0(Dec 8, 2021)

Updates

v0.0.6(Nov 22, 2021)

v0.0.5(Nov 22, 2021)

v0.0.4(Nov 21, 2021)

v0.0.3(Nov 21, 2021)

v0.0.1(Nov 19, 2021)

Owner

Repositório criado para abrigar os notebooks com a listas de exercícios propostos pelo professor Gustavo Guanabara do canal Curso em Vídeo do YouTube durante o Curso de Python 3

PyTorch implementation of CloudWalk's recent work DenseBody

Embracing Single Stride 3D Object Detector with Sparse Transformer

Evaluation Pipeline for our ECCV2020: Journey Towards Tiny Perceptual Super-Resolution.

naked is a Python tool which allows you to strip a model and only keep what matters for making predictions.

Detect roadway lanes using Python OpenCV for project during the 5th semester at DHBW Stuttgart for lecture in digital image processing.

A foreign language learning aid using a neural network to predict probability of translating foreign words

TEDSummary is a speech summary corpus. It includes TED talks subtitle (Document), Title-Detail (Summary), speaker name (Meta info), MP4 URL, and utterance id

Code for "Learning Graph Cellular Automata"

Gluon CV Toolkit

ICCV2021: Code for 'Spatial Uncertainty-Aware Semi-Supervised Crowd Counting'

Large scale PTM - PPI relation extraction

YuNetのPythonでのONNX、TensorFlow-Lite推論サンプル

Pytorch implementation of SenFormer: Efficient Self-Ensemble Framework for Semantic Segmentation

This is the source code of the 1st place solution for segmentation task (with Dice 90.32%) in 2021 CCF BDCI challenge.

Classic Papers for Beginners and Impact Scope for Authors.

Code and data of the EMNLP 2021 paper "Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer"

POT : Python Optimal Transport

Siamese TabNet

This is the open-source reference implementation of the SIGGRAPH 2021 paper Intersection-free Rigid Body Dynamics.