Toolchest provides APIs for scientific and bioinformatic data analysis.

Last update: Jun 30, 2022

Overview

Toolchest Python Client

Toolchest provides APIs for scientific and bioinformatic data analysis. It allows you to abstract away the costliness of running tools on your own resources by running the same jobs on secure, powerful remote servers.

This package contains the Python client for using Toolchest. For the R client, see here.

Installation

The Toolchest client is available on PyPI:

pip install toolchest-client

Usage

Using a tool in Toolchest is as simple as:

import toolchest_client as toolchest
toolchest.set_key("YOUR_TOOLCHEST_KEY")
toolchest.kraken2(
  tool_args="",
  inputs="path/to/input.fastq",
  output_path="path/to/output.fastq",
)

For a list of available tools, see the documentation.

Configuration

To use Toolchest, you must have an authentication key stored in the TOOLCHEST_KEY environment variable.

import toolchest_client as toolchest
toolchest.set_key("YOUR_TOOLCHEST_KEY") # or a file path containing the key

Contact Toolchest if:

you need a key
you’ve forgotten your key
the key is producing authentication errors.

Documentation & User Guide available at Read the Docs

Comments

Enable paired reads for `kraken2`

Adds the option to use paired-read inputs for kraken2, via the read_one and read_two arguments (or a list of two paths via inputs).

Adds/removes --paired to tool_args as necessary.

opened by bcai2 3
v0.4.0
Add Poetry, remove Twine

Add CircleCI automatic deploy to PyPI (untested for prod PyPI)

Note: CircleCI will be failing because v0.4.0 already exists on test PyPI. That is to be expected, because I already bumped it to v0.4.0 when testing.
opened by lebovic 3
S3 chaining
Adds:

Output class returned by all toolchest.tool() calls, which contains s3_uri, presigned_s3_url, and (local) output_path variables

S3 chaining, via supplying output.s3_uri from a previous tool as the inputs parameter for a following tool

the ability to skip download of any tool's output, by setting output_path=None (set to None by default)
opened by lebovic 2
Polish tool_arg handling, add more STAR args
Adds:

More STAR args

Add multiple levels of tool_arg handling (whitelist, dangerlist, blacklist)

Error on unknown or blacklisted args

Reduce complexity (validation and parallelization for now) if a dangerous argument is passed

Requires:

https://github.com/trytoolchest/toolchest-worker-node/pull/24

https://github.com/trytoolchest/toolchest-api/pull/22

This does not fix:

Bigger disk/memory/etc requirements for larger files where args trigger reduced complexity / no parallelization
opened by lebovic 2
STAR whitelist options
Adds basic whitelist options for STAR.

Adds support for tags with variable amounts of arguments. Adds the --quantMode tag for STAR.

(This should be merged in after the kraken2 paired read commit.)
opened by bcai2 2
feat: centrifuge base
Adds the centrifuge tool.

Adds docs.

Refactors how prefix_mapping is generated for megahit with a new module (input_util.py) and function (convert_input_params_to_prefix_mapping). Adds a unit test for the function.
opened by bcai2 1
fix: upload/download tracker bugfixes
Refactors the tracking printed statements into a pythonic print call with string formatting.

Fixes status update logic in uploading. (This was causing the terminal output to stall at the "uploading" stage.)

Adds integration test dirs to .gitignore.
opened by bcai2 1
fix: remove pysam due to multiple issues

Pysam has caused multiple issues as a package and STAR parallelization is not currently used so this pr fully removes pysam as a dependency. Either a different library or custom sam file merging code is planned to be implemented later so parallelization framework is remaining in the code for now.

opened by jherr-dev 1
feat: add preliminary alphafold support

Adds basic support for running AlphaFold via Toolchest. Code needs to be cleaned up and better documented. Currently limited to 1 input fasta.

use_reduced_dbs and is_prokaryote_list are currently disabled until further implementation and testing is done. Integration will come with reduced dbs since full dbs take 45 minutes to an hour to run even on simple input.

opened by jherr-dev 1
feat: support async execution
Adds:

Support for async execution

See https://gist.github.com/lebovic/72fbb857119f1667c7959a4d7e28cd50 (or the integration test) for a hacky example on how to run Toolchest with async execution.
opened by lebovic 1
fix: set default version number

Sets the version number to a default instead of erroring if the client is run from source (i.e., without the toolchest-client package being installed via pip).

Open question: the version number defaults to 0.0.0, which can be confusing -- are there any other labels that might be better (e.g., dev or just the empty string)?

opened by bcai2 1

Releases(v0.11.3)

v0.11.3(Nov 8, 2022)

Changelog: #287
Source code(tar.gz)
Source code(zip)
v0.11.2(Nov 2, 2022)

v0.11.1 changelog: #279 v0.11.2 changelog: #285
Source code(tar.gz)
Source code(zip)
v0.11.0(Oct 19, 2022)

v0.11.0
Source code(tar.gz)
Source code(zip)
v0.9.43(Oct 10, 2022)

v0.9.43
Source code(tar.gz)
Source code(zip)
v0.9.32(Aug 30, 2022)

Source code(tar.gz)
Source code(zip)
v0.9.8(May 20, 2022)

Allows previously-blocked --max-target-seqs argument on diamond blastx
Source code(tar.gz)
Source code(zip)
v0.9.1(Apr 1, 2022)
Adds:

Adds a skip_decompression tool param under *kwargs

Updates:

shogun_align to use Bowtie 2 instead of BURST as the underlying aligner

See #134.
Source code(tar.gz)
Source code(zip)
v0.7.46(Dec 29, 2021)
Modifies:

Bugfix, maintaining ordering of inputs for megahit

Source code(tar.gz)
Source code(zip)
v0.7.43(Dec 24, 2021)
Adds:

megahit

Source code(tar.gz)
Source code(zip)
v0.7.39(Dec 20, 2021)
Adds:

Ability to pass S3 URIs as inputs (#49)

Modified handling of arguments + raw execution mode + more STAR arguments (#57)

Multipart uploads / downloads and the ability to increase non-parallel input file sizes (#60, #61, #63, #65))

Enable output file .tar.gzs across the board (#68)

Add an explicit parallelize=True flag (#68)

Modifies:

Various fixes (#49, #56, #58, #64, #67)

Source code(tar.gz)
Source code(zip)
v0.7.28(Nov 23, 2021)

See #46, #53
Source code(tar.gz)
Source code(zip)
v0.7.20(Oct 22, 2021)

Fixes paired end Kraken 2 bug.

See #42 for more details.
Source code(tar.gz)
Source code(zip)
v0.7.19(Oct 22, 2021)
Adds:

Kraken 2 paired end support

Commonly used STAR arguments

Better testing and documentation

See #41 for more details.
Source code(tar.gz)
Source code(zip)
v0.7.14(Oct 1, 2021)
Adds basic unit tests for functions in toolchest_client/files.

Restructures the key validation check fto occur before any jobs are spawned.

See #37 for more details.
Source code(tar.gz)
Source code(zip)
v0.7.13(Aug 27, 2021)

See #33
Source code(tar.gz)
Source code(zip)
v0.7.8(Aug 20, 2021)

See #23
Source code(tar.gz)
Source code(zip)
v0.5.0(Jul 9, 2021)

See https://github.com/trytoolchest/toolchest-client-python/pull/9
Source code(tar.gz)
Source code(zip)
v0.3.0(Jun 11, 2021)

Changed API URL to production URL.
Source code(tar.gz)
Source code(zip)
v0.2.1(Jun 9, 2021)

Changed tool_args parameter names for tool functions.
Source code(tar.gz)
Source code(zip)
v0.2.0(Jun 9, 2021)

Initial (development) release.

Contains functions setting authorization keys and executing basic cutadapt and kraken2 queries.

Added initial documentation.
Source code(tar.gz)
Source code(zip)

Owner

Toolchest

GitHub Repository

An Indexer that works out-of-the-box when you have less than 100K stored Documents

U100KIndexer An Indexer that works out-of-the-box when you have less than 100K stored Documents. U100K means under 100K. At 100K stored Documents with

7 Mar 15, 2022

Minimal working example of data acquisition with nidaqmx python API

Data Aquisition using NI-DAQmx python API Based on this project It is a minimal working example for data acquisition using the NI-DAQmx python API. It

1 Nov 05, 2021

A DSL for data-driven computational pipelines

"Dataflow variables are spectacularly expressive in concurrent programming" Henri E. Bal , Jennifer G. Steiner , Andrew S. Tanenbaum Quick overview Ne

1.9k Jan 03, 2023

MDAnalysis is a Python library to analyze molecular dynamics simulations.

MDAnalysis Repository README [*] MDAnalysis is a Python library for the analysis of computer simulations of many-body systems at the molecular scale,

933 Dec 28, 2022

A utility for functional piping in Python that allows you to access any function in any scope as a partial.

WithPartial Introduction WithPartial is a simple utility for functional piping in Python. The package exposes a context manager (used with with) calle

1 Oct 26, 2021

Monitor the stability of a pandas or spark dataframe ⚙︎

Population Shift Monitoring popmon is a package that allows one to check the stability of a dataset. popmon works with both pandas and spark datasets.

403 Dec 07, 2022

Statistical package in Python based on Pandas

Pingouin is an open-source statistical package written in Python 3 and based mostly on Pandas and NumPy. Some of its main features are listed below. F

1.2k Dec 31, 2022

Spectacular AI SDK fuses data from cameras and IMU sensors and outputs an accurate 6-degree-of-freedom pose of a device.

Spectacular AI SDK examples Spectacular AI SDK fuses data from cameras and IMU sensors (accelerometer and gyroscope) and outputs an accurate 6-degree-

94 Jan 04, 2023

Data Analysis for First Year Laboratory at Imperial College, London.

Data Analysis for First Year Laboratory at Imperial College, London. For personal reference only, and to reference in lab reports and lab books.

0 Aug 29, 2022

Datashredder is a simple data corruption engine written in python. You can corrupt anything text, images and video.

Datashredder is a simple data corruption engine written in python. You can corrupt anything text, images and video. You can chose the cha

2 Jul 22, 2022

ETL pipeline on movie data using Python and postgreSQL

Movies-ETL ETL pipeline on movie data using Python and postgreSQL Overview This project consisted on a automated Extraction, Transformation and Load p

0 Jul 07, 2021

Retentioneering: product analytics, data-driven customer journey map optimization, marketing analytics, web analytics, transaction analytics, graph visualization, and behavioral segmentation with customer segments in Python.

What is Retentioneering? Retentioneering is a Python framework and library to assist product analysts and marketing analysts as it makes it easier to

581 Jan 07, 2023

OpenARB is an open source program aiming to emulate a free market while encouraging players to participate in arbitrage in order to increase working capital.

Overview OpenARB is an open source program aiming to emulate a free market while encouraging players to participate in arbitrage in order to increase

3 Feb 12, 2022

X-news - Pipeline data use scrapy, kafka, spark streaming, spark ML and elasticsearch, Kibana

5 Sep 28, 2022

Detecting Underwater Objects (DUO)

Underwater object detection for robot picking has attracted a lot of interest. However, it is still an unsolved problem due to several challenges. We take steps towards making it more realistic by ad

27 Dec 12, 2022

ETL flow framework based on Yaml configs in Python

ETL framework based on Yaml configs in Python A light framework for creating data streams. Setting up streams through configuration in the Yaml file.

18 Jul 06, 2022

This creates a ohlc timeseries from downloaded CSV files from NSE India website and makes a SQLite database for your research.

NSE-timeseries-form-CSV-file-creator-and-SQL-appender- This creates a ohlc timeseries from downloaded CSV files from National Stock Exchange India (NS

1 Oct 02, 2022

Spaghetti: an open-source Python library for the analysis of network-based spatial data

pysal/spaghetti SPAtial GrapHs: nETworks, Topology, & Inference Spaghetti is an open-source Python library for the analysis of network-based spatial d

203 Jan 03, 2023

In this project, ETL pipeline is build on data warehouse hosted on AWS Redshift.

ETL Pipeline for AWS Project Description In this project, ETL pipeline is build on data warehouse hosted on AWS Redshift. The data is loaded from S3 t

1 Nov 01, 2021

Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data.

PremiershipPlayerAnalysis Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data. No

5 Sep 06, 2021

Toolchest provides APIs for scientific and bioinformatic data analysis.

Related tags

Overview

Toolchest Python Client

Installation

Usage

Configuration

Documentation & User Guide available at Read the Docs

Comments

Releases(v0.11.3)

v0.11.3(Nov 8, 2022)

v0.11.2(Nov 2, 2022)

v0.11.0(Oct 19, 2022)

v0.9.43(Oct 10, 2022)

v0.9.32(Aug 30, 2022)

v0.9.8(May 20, 2022)

v0.9.1(Apr 1, 2022)

v0.7.46(Dec 29, 2021)

v0.7.43(Dec 24, 2021)

v0.7.39(Dec 20, 2021)

v0.7.28(Nov 23, 2021)

v0.7.20(Oct 22, 2021)

v0.7.19(Oct 22, 2021)

v0.7.14(Oct 1, 2021)

v0.7.13(Aug 27, 2021)

v0.7.8(Aug 20, 2021)

v0.5.0(Jul 9, 2021)

v0.3.0(Jun 11, 2021)

v0.2.1(Jun 9, 2021)

v0.2.0(Jun 9, 2021)