Synthetic data need to preserve the statistical properties of real data in terms of their individual behavior and (inter-)dependences

Last update: Dec 20, 2022

Overview

Overview

Synthetic data need to preserve the statistical properties of real data in terms of their individual behavior and (inter-)dependences. Copula and functional Principle Component Analysis (fPCA) are statistical models that allow these properties to be simulated (Joe 2014). As such, copula generated data have shown potential to improve the generalization of machine learning (ML) emulators (Meyer et al. 2021) or anonymize real-data datasets (Patki et al. 2016).

Synthia is an open source Python package to model univariate and multivariate data, parameterize data using empirical and parametric methods, and manipulate marginal distributions. It is designed to enable scientists and practitioners to handle labelled multivariate data typical of computational sciences. For example, given some vertical profiles of atmospheric temperature, we can use Synthia to generate new but statistically similar profiles in just three lines of code (Table 1).

Synthia supports three methods of multivariate data generation through: (i) fPCA, (ii) parametric (Gaussian) copula, and (iii) vine copula models for continuous (all), discrete (vine), and categorical (vine) variables. It has a simple and succinct API to natively handle xarray's labelled arrays and datasets. It uses a pure Python implementation for fPCA and Gaussian copula, and relies on the fast and well tested C++ library vinecopulib through pyvinecopulib's bindings for fast and efficient computation of vines. For more information, please see the website at https://dmey.github.io/synthia.

Table 1. Example application of Gaussian and fPCA classes in Synthia. These are used to generate random profiles of atmospheric temperature similar to those included in the source data. The xarray dataset structure is maintained and returned by Synthia.

Source	Synthetic with Gaussian Copula	Synthetic with fPCA
`ds = syn.util.load_dataset()`	`g = syn.CopulaDataGenerator()`	`g = syn.fPCADataGenerator()`
	`g.fit(ds, syn.GaussianCopula())`	`g.fit(ds)`
	`g.generate(n_samples=500)`	`g.generate(n_samples=500)`

Documentation

For installation instructions, getting started guides and tutorials, background information, and API reference summaries, please see the website.

How to cite

If you are using Synthia, please cite the following two papers using their respective Digital Object Identifiers (DOIs). Citations may be generated automatically using Crosscite's DOI Citation Formatter or from the BibTeX entries below.

Synthia Software	Software Application
DOI: 10.21105/joss.02863	DOI: 10.5194/gmd-14-5205-2021

@article{Meyer_and_Nagler_2021,
  doi = {10.21105/joss.02863},
  url = {https://doi.org/10.21105/joss.02863},
  year = {2021},
  publisher = {The Open Journal},
  volume = {6},
  number = {65},
  pages = {2863},
  author = {David Meyer and Thomas Nagler},
  title = {Synthia: multidimensional synthetic data generation in Python},
  journal = {Journal of Open Source Software}
}

@article{Meyer_and_Nagler_and_Hogan_2021,
  doi = {10.5194/gmd-14-5205-2021},
  url = {https://doi.org/10.5194/gmd-14-5205-2021},
  year = {2021},
  publisher = {Copernicus {GmbH}},
  volume = {14},
  number = {8},
  pages = {5205--5215},
  author = {David Meyer and Thomas Nagler and Robin J. Hogan},
  title = {Copula-based synthetic data augmentation for machine-learning emulators},
  journal = {Geoscientific Model Development}
}

If needed, you may also cite the specific software version with its corresponding Zendo DOI.

Contributing

If you are looking to contribute, please read our Contributors' guide for details.

Development notes

If you would like to know more about specific development guidelines, testing and deployment, please refer to our development notes.

Copyright and license

Acknowledgements

Special thanks to @letmaik for his suggestions and contributions to the project.

Comments

Explain how to run the test suite
Describe the bug There is a test suite, but the documentation does not explain how to run it.

Here is what works for me:

Install pytest.

Clone the source repository.

Run pytest in the root directory of the repository.
opened by khinsen 7
Review: Copula distribution usage and examples

Your package offers support for simulating vine copulas. However, I don't see examples demonstrating how to simulate data from a vine copula given desired conditional dependency requirements.

Is this possible with the current API? If not, how would I use the vine copula generator to achieve this?

Otherwise, can examples show the difference between simulating Gaussian and vine copulas? I only see examples for the Gaussian copula.

opened by mnarayan 5
fPCA documentation
Describe the bug

The documentation page on fPCA says:

PCA can be used to generate synthetic data for the high-dimensional vector $X$. For every instance $X_i$ in the data set, we compute the principal component scores $a_{i, 1}, \dots, a_{i, K}$. Because the principal components $v_1, \dots, v_K$ are orthogonal, the scores are necessarily uncorrelated and we may treat them as independent.

The claim that "because the principal components $v_1, \dots, v_K$ are orthogonal, the scores are necessarily uncorrelated" looks wrong to me. These scores are projections of the $X_i$ onto the elements of an orthonormal basis. That doesn't make them uncorrelated. There are lots of orthonormal bases one can project on, and for most of them the projections are not uncorrelated. You need some property of the distribution of $X$ to derive a zero correlation, for example a Gaussian distribution, for which the PCA basis yields approximately uncorrelated projections.
opened by khinsen 3
Review: Clarify API

It would be helpful to add/explain what the different classes do Data Generators, Parametrizer, Transformers somewhere in the introduction or usage component of the documentation. Explain the different classes and what each is supposed to do. If it is similar to or inspired by well-known API of a different package, please point to it.

I think generators and transformers are obvious but I only sort of understand Parametrizers. It is also confusing in the sense that people might think this has something to do with parametric distributions when you mean it to be something different.

Is this API for Parametrizers inspired by some convention elsewhere? If so it would be helpful to point to that. For instance, the generators are very similar to statsmodel generators.

opened by mnarayan 2
Small error in docs

Hi, just letting you know I noticed a small error in the documentation.

At the bottom of this page https://dmey.github.io/synthia/examples/fpca.html

The error is in line [6] of the code, under "Plot the results".

You have: plot_profiles(ds_true, 'temperature_fl')

But I believe it should be: plot_profiles(ds_synth, 'temperature_fl')

you want to plot results, not the original here.

Cheers & thanks for the cool project!

opened by BigTuna08 1
Review: Comparisons to other common packages

What are other packages people might use to simulate data (e.g. statsmodels comes to mind) and how is this package different? Your package supports generating data for multivariate copula distributions and via fPCA. I understand what this entails but I think this could use further elaboration.

This package supports nonparametric distributions much more than the typical parametric data generators found in common packages and it would be useful to highlight these explicitly.

opened by mnarayan 1
Support categorical data for pyvinecopulib

During fitting, category values are reindexed as integers starting from 0 and transformed to one-hot vectors. The opposite during generation. Any data type works for categories, including strings.

opened by letmaik 0

Add support for categorical data

We can treat categorical data as discrete but first we need to pre-process categorical values by one hot encoding to remove the order. Re API we can change the current version from

# Assuming  an xarray datasets ds with X1 discrete and and X2 categorical 
generator.fit(ds, copula=syn.VineCopula(controls=ctrl), is_discrete={'X1': True, 'X2': False})

to something like

with X3 continuous 
g.fit(ds, copula=syn.VineCopula(controls=ctrl), types={'X1': 'disc', 'X2': 'cat', 'X3': 'cont'})

opened by dmey 0

Add support for handling discrete quantities
Introduces the option to specify and model discrete quantities as follows:

# Assuming an xarray datasets ds with X1 discrete and and X2 continuous generator.fit(ds, copula=syn.VineCopula(controls=ctrl), is_discrete={'X1': True, 'X2': False})

This option is only supported for vine copulas
opened by dmey 0

Releases(1.1.0)

1.1.0(Sep 1, 2021)
Pin pyvinecopulib version to avoid issues between versions.

Add CI tests for Python 3.9 (#17).

Minor doc improvements.

Source code(tar.gz)
Source code(zip)
1.0.0(Apr 19, 2021)
1.0.0

Add JOSS summary paper (#26).

Improve docs and tutorials (#14, #13, #18, ...).

Enable CI on multiple OS and Python versions (#16).

Source code(tar.gz)
Source code(zip)
0.3.0(Nov 12, 2020)
Add support for handling categorical quantities (#10, #13).

Source code(tar.gz)
Source code(zip)
0.2.0(Nov 11, 2020)
Add support for handling discrete quantities (#9).

Add support for setting a seed when generating new samples (#11).

Drop support for Python 3.7.

Source code(tar.gz)
Source code(zip)
0.1.1(Oct 24, 2020)
Fix qrng argument for pyvinecopulib due to vinecopulib/pyvinecopulib#68 and vinecopulib/pyvinecopulib#69.

Source code(tar.gz)
Source code(zip)
0.1.0(Oct 22, 2020)
First public release.

Source code(tar.gz)
Source code(zip)

Owner

GitHub Repository https://dmey.github.io/synthia

Karate Club: An API Oriented Open-source Python Framework for Unsupervised Learning on Graphs (CIKM 2020)

Karate Club is an unsupervised machine learning extension library for NetworkX. Please look at the Documentation, relevant Paper, Promo Video, and Ext

1.8k Jan 09, 2023

Pipetools enables function composition similar to using Unix pipes.

Pipetools Complete documentation pipetools enables function composition similar to using Unix pipes. It allows forward-composition and piping of arbit

186 Dec 29, 2022

Datashredder is a simple data corruption engine written in python. You can corrupt anything text, images and video.

Datashredder is a simple data corruption engine written in python. You can corrupt anything text, images and video. You can chose the cha

2 Jul 22, 2022

Pyspark project that able to do joins on the spark data frames.

SPARK JOINS This project is to perform inner, all outer joins and semi joins. create_df.py: load_data.py : helps to put data into Spark data frames. d

1 Dec 14, 2021

This repo contains a simple but effective tool made using python which can be used for quality control in statistical approach.

📈 Statistical Quality Control 📉 This repo contains a simple but effective tool made using python which can be used for quality control in statistica

8 Oct 18, 2022

Python package for analyzing sensor-collected human motion data

71 Nov 05, 2022

A probabilistic programming library for Bayesian deep learning, generative models, based on Tensorflow

ZhuSuan is a Python probabilistic programming library for Bayesian deep learning, which conjoins the complimentary advantages of Bayesian methods and

2.2k Dec 28, 2022

BErt-like Neurophysiological Data Representation

BENDR BErt-like Neurophysiological Data Representation This repository contains the source code for reproducing, or extending the BERT-like self-super

114 Dec 23, 2022

PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams

PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams Motivation When dataset freshness is critical, the annotating of high speed

4 Aug 02, 2022

Yet Another Workflow Parser for SecurityHub

YAWPS Yet Another Workflow Parser for SecurityHub "Screaming pepper" by Rum Bucolic Ape is licensed with CC BY-ND 2.0. To view a copy of this license,

8 Dec 22, 2022

Data and code accompanying the paper Politics and Virality in the Time of Twitter

Politics and Virality in the Time of Twitter Data and code accompanying the paper Politics and Virality in the Time of Twitter. In specific: the code

3 Jul 02, 2022

A set of procedures that can realize covid19 virus detection based on blood.

3 Mar 07, 2022

AptaMat is a simple script which aims to measure differences between DNA or RNA secondary structures.

AptaMAT Purpose AptaMat is a simple script which aims to measure differences between DNA or RNA secondary structures. The method is based on the compa

3 Nov 03, 2022

Anomaly Detection with R

AnomalyDetection R package AnomalyDetection is an open-source R package to detect anomalies which is robust, from a statistical standpoint, in the pre

3.5k Dec 27, 2022

Catalogue data - A Python Scripts to prepare catalogue data

catalogue_data Scripts to prepare catalogue data. Setup Clone this repo. Install

3 Mar 03, 2022

🌍 Create 3d-printable STLs from satellite elevation data 🌏

mapa 🌍 Create 3d-printable STLs from satellite elevation data Installation pip install mapa Usage mapa uses numpy and numba under the hood to crunch

13 Dec 15, 2022

This python script allows you to manipulate the audience data from Sl.ido surveys

Slido-Automated-VoteBot This python script allows you to manipulate the audience data from Sl.ido surveys Since Slido blocks interference from automat

1 Jan 24, 2022

Python package for processing UC module spectral data.

UC Module Python Package How To Install clone repo. cd UC-module pip install . How to Use uc.module.UC(measurment=str, dark=str, reference=str, heade

1 Oct 20, 2021

Python-based Space Physics Environment Data Analysis Software

pySPEDAS pySPEDAS is an implementation of the SPEDAS framework for Python. The Space Physics Environment Data Analysis Software (SPEDAS) framework is

98 Dec 22, 2022

Stitch together Nanopore tiled amplicon data without polishing a reference

Stitch together Nanopore tiled amplicon data using a reference guided approach Tiled amplicon data, like those produced from primers designed with pri

14 Aug 30, 2022

Synthetic data need to preserve the statistical properties of real data in terms of their individual behavior and (inter-)dependences

Related tags

Overview

Overview

Documentation

How to cite

Contributing

Development notes

Copyright and license

Acknowledgements

Comments

you want to plot results, not the original here.

Releases(1.1.0)

1.1.0(Sep 1, 2021)

1.0.0(Apr 19, 2021)

0.3.0(Nov 12, 2020)

0.2.0(Nov 11, 2020)

0.1.1(Oct 24, 2020)

0.1.0(Oct 22, 2020)

Owner

Karate Club: An API Oriented Open-source Python Framework for Unsupervised Learning on Graphs (CIKM 2020)

Pipetools enables function composition similar to using Unix pipes.

Datashredder is a simple data corruption engine written in python. You can corrupt anything text, images and video.

Pyspark project that able to do joins on the spark data frames.

This repo contains a simple but effective tool made using python which can be used for quality control in statistical approach.

Python package for analyzing sensor-collected human motion data

A probabilistic programming library for Bayesian deep learning, generative models, based on Tensorflow

BErt-like Neurophysiological Data Representation

PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams

Yet Another Workflow Parser for SecurityHub

Data and code accompanying the paper Politics and Virality in the Time of Twitter

A set of procedures that can realize covid19 virus detection based on blood.

AptaMat is a simple script which aims to measure differences between DNA or RNA secondary structures.

Anomaly Detection with R

Catalogue data - A Python Scripts to prepare catalogue data

🌍 Create 3d-printable STLs from satellite elevation data 🌏

This python script allows you to manipulate the audience data from Sl.ido surveys

Python package for processing UC module spectral data.

Python-based Space Physics Environment Data Analysis Software

Stitch together Nanopore tiled amplicon data without polishing a reference