Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Virtual DataFrame

Full documentation

Motivation

With Panda-like dataframe or numby-like array, do you want to create a code, and choose at the end, the framework to use? Do you want to be able to choose the best framework after simply performing performance measurements? This framework unifies multiple Panda-compatible or Numpy-comptaible components, to allow the writing of a single code, compatible with all.

Do you want to use different architectures at different times of the year to be "green" and cheaper? Do you want to use a GPU only for the black-friday?

Synopsis

With some parameters and Virtual classes, it's possible to write a code, and execute this code:

  • With or without multicore
  • With or without cluster (multi nodes)
  • With or without GPU

To do that, we create some virtual classes, add some methods in others classes, etc.

It's difficult to use a combinaison of framework, with the same classe name, with similare semantic, etc. For example, if you want to use in the same program, Dask, cudf, pandas, modin, pyspark or pyspark+rapids, you must manage:

  • pandas.DataFrame, pandas,Series
  • modin.pandas.DataFrame, modin.pandas.Series
  • cudf.DataFrame, cudf.Series
  • dask.DataFrame, dask.Series
  • pyspark.pandas.DataFrame, pyspark.pandas.Series

With numpy, you must manage:

  • numpy.ndarray
  • cupy.ndarray
  • dask.array

With cudf or cudf, the code must call .to_pandas() or asnumpy(). With dask, the code must call .compute(), can use @delayed or dask.distributed.Client. etc.

We propose to replace all these classes and scenarios, with a uniform model, inspired by dask (the more complex API). Then, it is possible to write one code, and use it in differents environnements and frameworks.

This project is essentially a back-port of Dask+Cudf to others frameworks. We try to normalize the API of all frameworks. This project will weave your code with the selected framework, at runtime.

Binder

Metadata

Release files for mx06 0.2.dev0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for mx06 0.2.dev0
File Interpreter ABI Platform
mx06-0.2.dev0-py3-none-any.whl Python 3 none any Details

Release files / mx06-0.2.dev0-py3-none-any.whl

Download URL mx06-0.2.dev0-py3-none-any.whl
Size 26.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3bcc58f467c49d2fbc1fa1a88892edbecc535ad0b4cfa44b62ea1a3470349a3c
BLAKE2b-256 checksum
How to use checksums
082f689a49b87ff342174d951f7ec4f367ff83b971bd6776ebd6e186f890ad78
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.13

Release history Release notifications | RSS feed

This release

0.2.dev0 This release

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page