Skip to main content

Release Notes Downloads Python Versions GitHub CI Status License: MIT

Funcy with pipeline-based operators

If Funcy and Pipe had a baby. Deal with data transformation in python in a sane way.

I love Ruby. It's a great language and one of the things they got right was pipelined data transformation. Elixir got this even more right with the explicit pipeline operator |>.

However, Python is the way of the future. As I worked more with Python, it was driving me nuts that the data transformation options were not chainable.

This project fixes this pet peeve.

Installation

pip install funcy-pipe

Or, if you are using poetry:

poetry add funcy-pipe

Examples

Extract a couple key values from a sql alchemy model:

import funcy_pipe as fp

entities_from_sql_alchemy
  | fp.lmap(lambda r: r.to_dict())
  | fp.lmap(lambda r: r | fp.omit(["id", "created_at", "updated_at"]))
  | fp.to_list

Or, you can be more fancy and use whatever and pmap:

import funcy_pipe as f
import whatever as _

entities_from_sql_alchemy
  | fp.lmap(_.to_dict)
  | fp.pmap(fp.omit(["id", "created_at", "updated_at"]))
  | fp.to_list

Create a map from an array of objects, ensuring the key is always an int:

section_map = api.get_sections() | fp.group_by(f.compose(int, that.id))

Grab the ID of a specific user:

filter_user_id = (
  collaborator_map().values()
  | fp.where(email=target_user)
  | fp.pluck("id")
  | fp.first()
)

Get distinct values from a list (in this case, github events):

events = [
  {
    "type": "PushEvent"
  },
  {
    "type": "CommentEvent"
  }
]

result = events | fp.pluck("type") | fp.distinct() | fp.to_list()

assert ["PushEvent", "CommentEvent"] == result

What if the objects are not dicts?

filter_user_id = (
  collaborator_map().values()
  | fp.where_attr(email=target_user)
  | fp.pluck_attr("id")
  | fp.first()
)

How about creating a dict where each value is sorted:

data
  # each element is a dict of city information, let's group by state
  | fp.group_by(itemgetter("state_name"))
  # now let's sort each value by population, which is stored as a string
  | fp.walk_values(
    f.partial(sorted, reverse=True, key=lambda c: int(c["population"])),
  )

A more complicated example (lifted from this project):

comments = (
    # tasks are pulled from the todoist api
    tasks
    # get all comments for each relevant task
    | fp.lmap(lambda task: api.get_comments(task_id=task.id))
    # each task's comments are returned as an array, let's flatten this
    | fp.flatten()
    # dates are returned as strings, let's convert them to datetime objects
    | fp.lmap(enrich_date)
    # no date filter is applied by default, we don't want all comments
    | fp.lfilter(lambda comment: comment["posted_at_date"] > last_synced_date)
    # comments do not come with who created the comment by default, we need to hit a separate API to add this to the comment
    | fp.lmap(enrich_comment)
    # only select the comments posted by our target user
    | fp.lfilter(lambda comment: comment["posted_by_user_id"] == filter_user_id)
    # there is no `sort` in the funcy library, so we reexport the sort built-in so it's pipe-able
    | fp.sort(key="posted_at_date")
    # create a dictionary of task_id => [comments]
    | fp.group_by(lambda comment: comment["task_id"])
)

Want to grab the values of a list of dict keys?

def add_field_name(input: dict, keys: list[str]) -> dict:
    return input | {
        "field_name": (
            keys
            # this is a sneaky trick: if we reference the objects method, when it's called it will contain a reference
            # to the object
            | fp.map(input.get)
            | fp.compact
            | fp.join_str("_")
        )
    }

result = [{ "category": "python", "header": "functional"}] | fp.map(fp.rpartial(add_field_name, ["category", "header"])) | fp.to_list
assert result == [{'category': 'python', 'header': 'functional', 'field_name': 'python_functional'}]

You can also easily test multiple conditions across API data (extracted from this project)

all_checks_successful = (
    last_commit.get_check_runs()
    | fp.pluck_attr("conclusion")
    # if you pass a set into `all` each element of the set is used to build a predicate
    # this condition tests if the "conclusion" attribute is either "success" or "skipped"
    | fp.all({"success", "skipped"})
)

Want to grab the values of a list of dict keys?

def add_field_name(input: dict, keys: list[str]) -> dict:
    return input | {
        "field_name": (
            keys
            # this is a sneaky trick: if we reference the objects method, when it's called it will contain a reference
            # to the object
            | fp.map(input.get)
            | fp.compact
            | fp.join_str("_")
        )
    }

result = [{ "category": "python", "header": "functional"}] | fp.map(fp.rpartial(add_field_name, ["category", "header"])) | fp.to_list
assert result == [{'category': 'python', 'header': 'functional', 'field_name': 'python_functional'}]

You can also easily group dictionaries by a key (or arbitrary function):

import operator

result = [{"age": 10, "name": "Alice"}, {"age": 12, "name": "Bob"}] | fp.group_by(operator.itemgetter("age"))
assert result == {10: [{'age': 10, 'name': 'Alice'}], 12: [{'age': 12, 'name': 'Bob'}]}

Extras

  • to_list
  • log
  • bp. run breakpoint() on the input value
  • sort
  • exactly_one. Throw an error if the input is not exactly one element
  • reduce
  • pmap. Pass each element of a sequence into a pipe'd function

Extensions

There are some functions which are not yet merged upstream into funcy, and may never be. You can patch funcy to add them using:

import funcy_pipe
funcy_pipe.patch()

Find the first element matching a condition (detect):

users = [
    {"name": "Alice", "age": 25, "city": "NYC"},
    {"name": "Bob", "age": 30, "city": "SF"},
    {"name": "Charlie", "age": 25, "city": "LA"}
]

# Find first user aged 25
first_young_user = users | fp.detect(age=25)
assert first_young_user == {"name": "Alice", "age": 25, "city": "NYC"}

Coming From Ruby?

  • uniq => distinct
  • detect => detect or where(some="Condition") | first or where_attr(some="Condition") | first
  • inverse => complement
  • times => repeatedly

Module Alias

Create a module alias for funcy-pipe to make things clean (import * always irks me):

# fp.py
from funcy_pipe import *

# code py
import fp

Inspiration

Related Projects

TODO

  • tests
  • docs for additional utils
  • fix typing threading

Metadata

Release files for funcy-pipe 0.14.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for funcy-pipe 0.14.0
File Size Uploaded
funcy_pipe-0.14.0.tar.gz 12.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for funcy-pipe 0.14.0
File Interpreter ABI Platform
funcy_pipe-0.14.0-py3-none-any.whl Python 3 none any Details

Total release size: 23.9 kB

Release files / funcy_pipe-0.14.0.tar.gz

Download URL funcy_pipe-0.14.0.tar.gz
Size 12.4 kB
Tags Source
SHA-256 checksum
How to use checksums
ec35cf6bcb3e323e048a9f7e2c738d5f9c0adc142f6a9c8621ba22dff65bd638
BLAKE2b-256 checksum
How to use checksums
8aad0ae7bb4df6d6503b901f2b0a0e1e2afacb20e64d6ea35a69661c46ccb8b4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.3 CPython/3.12.3 Linux/6.11.0-1018-azure

Release files / funcy_pipe-0.14.0-py3-none-any.whl

Download URL funcy_pipe-0.14.0-py3-none-any.whl
Size 11.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2125c89a497f17e43709d57a7d7a5c758428d6e16e625ef7185f287a3b5c102d
BLAKE2b-256 checksum
How to use checksums
9b5cbdcad67cbf3c567680ea25a406f2f442272035f79b04b3b1b92d9ca4b77a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.8.3 CPython/3.12.3 Linux/6.11.0-1018-azure

Release history Release notifications | RSS feed

This release

0.14.0 This release

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.2

2 release files

0.11.1

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page