DataFlows
DataFlows is a simple and intuitive way of building data processing flows.
- It's built for small-to-medium-data processing - data that fits on your hard drive, but is too big to load in Excel or as-is into Python, and not big enough to require spinning up a Hadoop cluster...
- It's built upon the foundation of the Frictionless Data project - which means that all data produced by these flows is easily reusable by others.
- It's a pattern not a heavy-weight framework: if you already have a bunch of download and extract scripts this will be a natural fit
Read more in the Features section below.
QuickStart
Install dataflows via pip install.
(If you are using minimal UNIX OS, run first sudo apt install build-essential)
Then use the command-line interface to bootstrap a basic processing script for any remote data file:
# Install from PyPi
$ pip install dataflows
# Inspect a remote CSV file
$ dataflows init https://raw.githubusercontent.com/datahq/dataflows/master/data/academy.csv
Writing processing code into academy_csv.py
Running academy_csv.py
academy:
# Year Ceremony Award Winner Name Film
(string) (integer) (string) (string) (string) (string)
---- ---------- ----------- -------------------------------- ---------- ------------------------------ -------------------
1 1927/1928 1 Actor Richard Barthelmess The Noose
2 1927/1928 1 Actor 1 Emil Jannings The Last Command
3 1927/1928 1 Actress Louise Dresser A Ship Comes In
4 1927/1928 1 Actress 1 Janet Gaynor 7th Heaven
5 1927/1928 1 Actress Gloria Swanson Sadie Thompson
6 1927/1928 1 Art Direction Rochus Gliese Sunrise
7 1927/1928 1 Art Direction 1 William Cameron Menzies The Dove; Tempest
...
# dataflows create a local package of the data and a reusable processing script which you can tinker with
$ tree
.
├── academy_csv
│ ├── academy.csv
│ └── datapackage.json
└── academy_csv.py
1 directory, 3 files
# Resulting 'Data Package' is super easy to use in Python
[adam] ~/code/budgetkey-apps/budgetkey-app-main-page/tmp (master=) $ python
Python 3.6.1 (default, Mar 27 2017, 00:25:54)
[GCC 4.2.1 Compatible Apple LLVM 8.0.0 (clang-800.0.42.1)] on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>> from datapackage import Package
>>> pkg = Package('academy_csv/datapackage.json')
>>> it = pkg.resources[0].iter(keyed=True)
>>> next(it)
{'Year': '1927/1928', 'Ceremony': 1, 'Award': 'Actor', 'Winner': None, 'Name': 'Richard Barthelmess', 'Film': 'The Noose'}
>>> next(it)
{'Year': '1927/1928', 'Ceremony': 1, 'Award': 'Actor', 'Winner': '1', 'Name': 'Emil Jannings', 'Film': 'The Last Command'}
# You now run `academy_csv.py` to repeat the process
# And obviously modify it to add data modification steps
Features
- Trivial to get started and easy to scale up
- Set up and run from command line in seconds ...
dataflows init=>flow.pypython flow.py
- Validate input (and esp source) quickly (non-zero length, right structure, etc.)
- Supports caching data from source and even between steps
- so that we can run and test quickly (retrieving is slow)
- Immediate test is run: and look at output ...
- Log, debug, rerun
- Degrades to simple python
- Conventions over configuration
- Log exceptions and / or terminate
- The input to each stage is a Data Package or Data Resource (not a previous task)
- Data package based and compatible
- Processors can be a function (or a class) processing row-by-row, resource-by-resource or a full package
- A pre-existing decent contrib library of Readers (Collectors) and Processors and Writers
Learn more
Dive into the Tutorial to get a deeper glimpse into everything that dataflows can do.
Also review this list of Built-in Processors, which also includes an API reference for each one of them.
Release files for dataflows 0.5.12
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dataflows-0.5.12.tar.gz | 42.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dataflows-0.5.12-py2.py3-none-any.whl | Python 3, Python 2 | none | any | Details |
Total release size: 103.6 kB
Release files / dataflows-0.5.12.tar.gz
| Download URL | dataflows-0.5.12.tar.gz |
|---|---|
| Size | 42.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8a69c5ad0bd50f088e142f8bf76d2feed4e6e8494f5dac4d948de59c804251b1
|
|
BLAKE2b-256 checksum How to use checksums |
4fc6539be289b7f286733c018e064a2816c5824501ddea41736f813f6064b93c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Release files / dataflows-0.5.12-py2.py3-none-any.whl
| Download URL | dataflows-0.5.12-py2.py3-none-any.whl |
|---|---|
| Size | 60.7 kB |
| Tags | Python 2 Python 3 |
|
SHA-256 checksum How to use checksums |
4652d72e099def9e6832240751bcd25e9f656ca438863cb7a77267d6c35bc56b
|
|
BLAKE2b-256 checksum How to use checksums |
c13d0251af595ea30658ca4bb18d27099f8d3452fbeae495eb6ddcfef1256f0c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|