Instant search for and access to many datasets in Pyspark.

Last update: Dec 16, 2022

Overview

SparkDataset

Provides instant access to many datasets right from Pyspark (in Spark DataFrame structure).

Drop a star if you like the project. 😃 Motivates 💪 me to keep working on such projects

What?

The idea is simple. There are various datasets available out there, but they are scattered in different places over the web. Is there a quick way (in Pyspark) to access them instantly without going through the hassle of searching, downloading, and reading ... etc? SparkDataset tries to address that question :)

Usage:

Start with importing data():

from sparkdataset import data

To load a dataset:

titanic = data('titanic')

To display the documentation of a dataset:

data('titanic', show_doc=True)

To see the available datasets:

data()

To search for datasets with terms

data('ab')

Did you mean:
crabs, abbey, Vocab

That's it.

Go to this notebook for a demonstration of the functionality

Why?

In R, there is a very easy and immediate way to access multiple statistical datasets, in almost no effort. All it takes is one line > data(dataset_name). This makes the life easier for quick prototyping and testing. Well, I am jealous that Pyspark does not have a similar functionality. Thus, the aim of sparkdataset is to fill that gap.

Currently, sparkdataset has about 757 (mostly numerical-based) datasets, that are based on RDatasets. In the future, I plan to scale it to include a larger set of datasets. For example,

include textual data for NLP-related tasks, and
allow adding a new dataset to the in-module repository.

Installation:

$ pip install sparkdataset

Uninstall:

$ pip uninstall sparkdataset
$ rm -rf $HOME/.sparkdataset

Changelog

1.0.0

Added search dataset by name similarity.
Example:

>>> data('heat')
Did you mean:
Wheat, heart, Heating, Yeast, eidat, badhealth, deaths, agefat, hla, heptathlon, azt

Added support to Windows.

Dependency:

pandas
pyspark :: 3.1.2

Miscellaneous:

Tested on OSX and Linux (debian).
Supports both Python 3 (3.8.8 and above).

TODO:

add textual datasets (e.g. NLTK stuff).
add samples generators.

Thanks to:

RDatasets: R's datasets collection.

Releases(1.0.0)

1.0.0(Nov 1, 2021)

Provides instant 🚀 access to many popular datasets 📑 right from Pyspark 🔥 (in dataframe structure).
Source code(tar.gz)
Source code(zip)

This is an example of how to automate Ridit Analysis for a dataset with large amount of questions and many item attributes

1 Nov 17, 2021

Programmatically access the physical and chemical properties of elements in modern periodic table.

API to fetch elements of the periodic table in JSON format. Uses Pandas for dumping .csv data to .json and Flask for API Integration. Deployed on "pyt

3 Oct 23, 2022

A utility for functional piping in Python that allows you to access any function in any scope as a partial.

WithPartial Introduction WithPartial is a simple utility for functional piping in Python. The package exposes a context manager (used with with) calle

1 Oct 26, 2021

A Pythonic introduction to methods for scaling your data science and machine learning work to larger datasets and larger models, using the tools and APIs you know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

This tutorial's purpose is to introduce Pythonistas to methods for scaling their data science and machine learning work to larger datasets and larger models, using the tools and APIs they know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

102 Nov 10, 2022

Python tools for querying and manipulating BIDS datasets.

PyBIDS is a Python library to centralize interactions with datasets conforming BIDS (Brain Imaging Data Structure) format.

180 Dec 18, 2022

Python dataset creator to construct datasets composed of OpenFace extracted features and Shimmer3 GSR+ Sensor datas

3 Jul 5, 2022

CleanX is an open source python library for exploring, cleaning and augmenting large datasets of X-rays, or certain other types of radiological images.

cleanX CleanX is an open source python library for exploring, cleaning and augmenting large datasets of X-rays, or certain other types of radiological

20 Jan 5, 2023

VHub - An API that permits uploading of vulnerability datasets and return of the serialized data

2 Feb 14, 2022

HyperSpy is an open source Python library for the interactive analysis of multidimensional datasets

HyperSpy is an open source Python library for the interactive analysis of multidimensional datasets that can be described as multidimensional arrays o

411 Dec 27, 2022

Instant search for and access to many datasets in Pyspark.

Related tags

Overview

SparkDataset

What?

Usage:

Why?

Installation:

Uninstall:

Changelog

Dependency:

Miscellaneous:

TODO:

Thanks to:

You might also like...

This is an example of how to automate Ridit Analysis for a dataset with large amount of questions and many item attributes

Programmatically access the physical and chemical properties of elements in modern periodic table.

A utility for functional piping in Python that allows you to access any function in any scope as a partial.

A Pythonic introduction to methods for scaling your data science and machine learning work to larger datasets and larger models, using the tools and APIs you know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

Python tools for querying and manipulating BIDS datasets.

Python dataset creator to construct datasets composed of OpenFace extracted features and Shimmer3 GSR+ Sensor datas

CleanX is an open source python library for exploring, cleaning and augmenting large datasets of X-rays, or certain other types of radiological images.

VHub - An API that permits uploading of vulnerability datasets and return of the serialized data

HyperSpy is an open source Python library for the interactive analysis of multidimensional datasets

Releases(1.0.0)

1.0.0(Nov 1, 2021)

Owner

Souvik Pratiher

DaDRA (day-druh) is a Python library for Data-Driven Reachability Analysis.

Additional tools for particle accelerator data analysis and machine information

Snakemake workflow for converting FASTQ files to self-contained CRAM files with maximum lossless compression.

yt is an open-source, permissively-licensed Python library for analyzing and visualizing volumetric data.

A Numba-based two-point correlation function calculator using a grid decomposition

CubingB is a timer/analyzer for speedsolving Rubik's cubes, with smart cube support

ToeholdTools is a Python package and desktop app designed to facilitate analyzing and designing toehold switches, created as part of the 2021 iGEM competition.

pandas: powerful Python data analysis toolkit

A data parser for the internal syncing data format used by Fog of World.

Analyze the Gravitational wave data stored at LIGO/VIRGO observatories

Python Implementation of Scalable In-Memory Updatable Bitmap Indexing

Lale is a Python library for semi-automated data science.

Meltano: ELT for the DataOps era. Meltano is open source, self-hosted, CLI-first, debuggable, and extensible.

TE-dependent analysis (tedana) is a Python library for denoising multi-echo functional magnetic resonance imaging (fMRI) data

Picka: A Python module for data generation and randomization.

ELFXtract is an automated analysis tool used for enumerating ELF binaries

Desafio proposto pela IGTI em seu bootcamp de Cloud Data Engineer

2019 Data Science Bowl

MeSH2Matrix - A set of Python codes for the generation of biomedical ontologies from the MeSH keywords of the PubMed scholarly publications

collect training and calibration data for gaze tracking