Skip to main content

Parallelize the execution of pytask.

Project description

PyPI PyPI - Python Version https://anaconda.org/pytask/pytask-parallel/badges/version.svg https://anaconda.org/pytask/pytask-parallel/badges/platforms.svg PyPI - License https://github.com/pytask-dev/pytask-parallel/workflows/Continuous%20Integration%20Workflow/badge.svg?branch=main https://codecov.io/gh/pytask-dev/pytask-parallel/branch/main/graph/badge.svg pre-commit.ci status https://img.shields.io/badge/code%20style-black-000000.svg

pytask-parallel

Parallelize the execution of tasks with pytask-parallel which is a plugin for pytask.

Installation

pytask-parallel is available on PyPI and Anaconda.org. Install it with

$ pip install pytask-parallel

# or

$ conda config --add channels conda-forge --add channels pytask
$ conda install pytask-parallel

By default, the plugin uses a robust implementation of the ProcessPoolExecutor from loky.

It is also possible to select the ProcessPoolExecutor or ThreadPoolExecutor in the concurrent.futures module as backends to execute tasks asynchronously.

Usage

To parallelize your tasks across many workers, pass an integer greater than 1 or 'auto' to the command-line interface.

$ pytask -n 2
$ pytask --n-workers 2

# Starts os.cpu_count() - 1 workers.
$ pytask -n auto

Using processes to parallelize the execution of tasks is useful for CPU bound tasks such as numerical computations. (Here is an explanation on what CPU or IO bound means.)

For IO bound tasks, tasks where the limiting factor are network responses, accesses to files, you can parallelize via threads.

$ pytask --parallel-backend threads

You can also set the options in one of the configuration files (pytask.ini, tox.ini, or setup.cfg).

# This is the default configuration. Note that, parallelization is turned off.

[pytask]
n_workers = 1
parallel_backend = loky  # or processes or threads

Changes

Consult the release notes to find out about what is new.

Development

  • pytask-parallel does not call the pytask_execute_task_protocol hook specification/entry-point because pytask_execute_task_setup and pytask_execute_task need to be separated from pytask_execute_task_teardown. Thus, plugins which change this hook specification may not interact well with the parallelization.

  • There are two PRs for CPython which try to re-enable setting custom reducers which should have been working, but does not. Here are the references.

  • If the TopologicalSorter becomes available for all supported Python versions, deprecate the copied module. Meanwhile, keep it in sync.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pytask-parallel-0.0.8.tar.gz (25.5 kB view hashes)

Uploaded Source

Built Distribution

pytask_parallel-0.0.8-py3-none-any.whl (9.3 kB view hashes)

Uploaded Python 3

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page