Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Last update: Dec 10, 2022

Overview

https://travis-ci.org/uchicago-cs/deepdish.svg?branch=master

https://img.shields.io/badge/license-BSD%203--Clause-blue.svg?style=flat

deepdish

Flexible HDF5 saving/loading and other data science tools from the University of Chicago. This repository also host a Deep Learning blog:

http://deepdish.io

Installation

pip install deepdish

Alternatively (if you have conda with the conda-forge channel):

conda install -c conda-forge deepdish

Main feature

The primary feature of deepdish is its ability to save and load all kinds of data as HDF5. It can save any Python data structure, offering the same ease of use as pickling or numpy.save. However, it improves by also offering:

Interoperability between languages (HDF5 is a popular standard)
Easy to inspect the content from the command line (using h5ls or our specialized tool ddls)
Highly compressed storage (thanks to a PyTables backend)
Native support for scipy sparse matrices and pandas DataFrame, Series and Panel
Ability to partially read files, even slices of arrays

An example:

import deepdish as dd

d = {
    'foo': np.ones((10, 20)),
    'sub': {
        'bar': 'a string',
        'baz': 1.23,
    },
}
dd.io.save('test.h5', d)

This can be reconstructed using dd.io.load('test.h5'), or inspected through the command line using either a standard tool:

$ h5ls test.h5
foo                      Dataset {10, 20}
sub                      Group

Or, better yet, our custom tool ddls (or python -m deepdish.io.ls):

$ ddls test.h5
/foo                       array (10, 20) [float64]
/sub                       dict
/sub/bar                   'a string' (8) [unicode]
/sub/baz                   1.23 [float64]

Documentation

http://deepdish.readthedocs.io/

Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Related tags

Overview

deepdish

Installation

Main feature

Documentation

Owner

UChicago - Department of Computer Science

Exploratory data analysis

A probabilistic programming library for Bayesian deep learning, generative models, based on Tensorflow

Pipeline to convert a haploid assembly into diploid

We're Team Arson and we're using the power of predictive modeling to combat wildfires.

This project is the implementation template for HW 0 and HW 1 for both the programming and non-programming tracks

The official repository for ROOT: analyzing, storing and visualizing big data, scientifically

SparseLasso: Sparse Solutions for the Lasso

Bearsql allows you to query pandas dataframe with sql syntax.

Data Scientist in Simple Stock Analysis of PT Bukalapak.com Tbk for Long Term Investment

A simplified prototype for an as-built tracking database with API

Flood modeling by 2D shallow water equation

GWpy is a collaboration-driven Python package providing tools for studying data from ground-based gravitational-wave detectors

Python beta calculator that retrieves stock and market data and provides linear regressions.

Template for a Dataflow Flex Template in Python

Investigating EV charging data

BigDL - Evaluate the performance of BigDL (Distributed Deep Learning on Apache Spark) in big data analysis problems

signac-flow - manage workflows with signac

Analysis scripts for QG equations

follow-analyzer helps GitHub users analyze their following and followers relationship

🌍 Create 3d-printable STLs from satellite elevation data 🌏