Skip to main content

Retentioneering

Retentioneering is an open-source Python library for understanding user behavior from event logs. Find where users get stuck, compare the paths different groups take, and discover patterns across clicks, sessions and repeat visits. Explore the data yourself or let an AI agent work with the same tools.

PyPI Python Downloads License: Apache 2.0 Discord Telegram Try in Colab

Run a demo notebook in Google Colab – no need to install the library locally, sample data included.

Quick start · Use cases · Use an AI agent · Docs

A demo of the interactive features of the transition graph widget. You can highlight a route, compare path groups, focus on the step you want to investigate, and explore incoming and outgoing transition probabilities in the ego view.

A demo of the interactive features of the transition graph widget. You can highlight a route, compare path groups, focus on the step you want to investigate, and explore incoming and outgoing transition probabilities in the ego view.

Is it for you?

Use Retentioneering when you have a question about the sequence of actions like these:

  • A metric changed. Where did journeys change and move the metric?
  • Users stop before reaching a goal. What do they do instead and how do successful paths differ?
  • You want to understand engagement. What behavior types are represented, and how does usage evolve across visits?
  • You need to investigate a flow. Where do checkout, onboarding, support or agent runs loop, branch or stall?

Bring an event-level export from your analytics platform, warehouse or application logs. All you need is a user, session or case ID; an event name; and a timestamp.

Quick start

Python: 3.10 - 3.13.

Environment: Jupyter (Jupyter Notebook, JupyterLab, JupyterLab Desktop), VS Code, Cursor, Google Colab.

To install the library, run

pip install retentioneering

Within a notebook, use %pip install retentioneering instead. See the installation guide for the details.

To try out the library quickly, without loading your own data, you can use a built-in e-commerce dataset:

import retentioneering as rete

stream = rete.datasets.load_ecom()  # bundled synthetic store
stream.transition_graph()

Common use cases

Below are a few examples of how you can apply retentioneering tools to approach common analytical problems on the bundled e-commerce demo dataset. The code chunks assume stream = rete.datasets.load_ecom().

A key metric dropped. Where did journeys change and cause it?

Create a binary segment that cuts each path at the boundaries of the affected period: its events inside the period are labelled inside, the rest outside, so one path can contribute to both parts. Then compare user behavior across these parts with the transition graph widget.

periods = stream.add_segment("period", time_range=("2024-05-19", "2024-06-07"))
periods.transition_graph(diff=("period", "inside", "outside"))

Transition graph in diff mode focused on payment_details: inside the window, transitions to payment_error and support_chat rise while the transition to purchase falls.

Click a node (e.g. payment_details) to inspect the routes that pass this node. With this comparison configuration, red means a higher next-step probability inside the period; blue means lower. This locates a behavioral difference to investigate.

Letting a path contribute to multiple levels of a segment gives more flexibility to diff mode: you can compare not only common attributes like mobile vs. desktop, acquisition channels, or experiment groups, but dynamic features as well: first session vs. the others, weekends vs. weekdays, etc.

Where do users leave the funnel and what do they do instead?

Count sessions that complete the steps in order – rather than whole user journeys (path_col="session_id") – with the funnel widget:

steps = ["cart", "shipping_details", "purchase"]
stream.funnel(steps=steps, path_col="session_id")

Funnel of sessions from cart to shipping_details to purchase: 28.3%, 14.2% and 4.8% of 3,605 sessions.

Then compare sessions that stopped at the shipping stage of this funnel with those that completed it using the step matrix widget in diff mode:

(
    stream
        .add_segment("stage", funnel_events=steps, path_col="session_id")
        .step_matrix(
            path_pattern="cart->.*->shipping_details",
            step_window=2, path_col="session_id",
            diff=("stage", "shipping_details", "purchase"),
        )
)

A Step Matrix aligned on cart and shipping_details compares shipping-stage and completed-funnel sessions: hovering cells shows each group's values, and the arrow buttons sort rows by preceding or following steps.

The path_pattern="cart->.*->shipping_details" argument of the step matrix breaks down the diagram into two parts: around cart and around shipping_details. You can inspect these surroundings and compare sessions that reached shipping_details but not purchase with those that completed the funnel.

What leads to, or follows, an error or another key event?

Instead of using traditional tree-like diagrams for path exploration, you can use Step Sankey, which shows the same numbers as a Step Matrix, drawn as flows. In this example, we align all the paths by the payment_error event and display the two steps on either side of it:

stream.step_sankey(
    anchor="payment_error", step_window=2,
    path_col="session_id",
)

Step Sankey centred on payment_error, showing the two steps before and after it; support_chat and path_end are the most common next steps.

A plain event name anchors each path on its first occurrence. An anchor spec chooses the position more precisely: which occurrence to use, which event of a pattern to center on (at), or how far to shift from it (offset). For example, align sessions on their last payment error to see whether users recover after it or give up:

stream.step_sankey(
    anchor={"pattern": "payment_error", "occurrence": "last"},
    step_window=2, path_col="session_id",
)

Step Sankey centred on each session's last payment_error: path_end is the most common next step and takes 35% of the step after it.

After the last error, sessions end more often than after the first one: path_end takes 19% of the next step instead of 14%, and 35% two steps later instead of 26%.

Which sessions match a behavior I care about?

Find sessions that opened the cart, then ended without reaching shipping or support:

abandoned = stream.filter_paths(
    {"metric": "matches_pattern", "op": "=", "value": True,
     "metric_args": {
         "pattern": "cart->[^shipping_details|support_chat]*->path_end"
     }},
    path_col="session_id",
)
abandoned.step_sankey(path_pattern="cart", path_col="session_id", step_window=3)

Step Sankey of abandoned-cart sessions aligned on cart: most sessions end within three steps after it.

The filter_paths data processor can filter paths according to a path metric value, such as length, duration, event count, etc. In our case we use the matches_pattern metric that checks if a path matches the regex-like pattern cart->[^shipping_details|support_chat]*->path_end.

How does product usage differ between users? What behavioral patterns are represented?

To explore clusters interactively, start with a bare call of the Cluster Analysis widget and set everything up in its sidebar:

stream.cluster_analysis()

Starting from a bare cluster_analysis() call: features, the cluster range and overview metrics are set in the sidebar, Apply runs the grid, another partition is picked on the Silhouette tab, clusters are renamed in the header and saved as a segment.

Here users are clustered by event_count_bulk, which expands into one count per event type, over a grid of 3 to 8 clusters scored by the silhouette metric. The heatmap compares mean metric values across clusters (blue for lower values, red for higher). Overview metrics such as length or in_segment_bulk describe the clusters without changing the features used for clustering. Pick another partition on the Silhouette tab if you are not satisfied with the silhouette-best split. Once you find an optimal split, label the clusters in the header, click Save Clusters and save them as a segment.

The same configuration can also be passed directly:

stream.cluster_analysis(
    features=[{"metric": "event_count_bulk"}],
    method_args={"n_clusters": "3-8"},
    overview_metrics=[
        {"metric": "event_count_bulk"},
        {"metric": "length"},
        {"metric": "in_segment_bulk", "metric_args": {"segment_name": "acquisition_channel"}},
    ],
)

And the partition saved in the GIF above – four clusters, renamed – becomes a segment column with add_clusters and rename_segment_levels:

user_types = (
    stream
        .add_clusters(
            "user_type", features=[{"metric": "event_count_bulk"}],
            method_args={"n_clusters": 4},
        )
        .rename_segment_levels("user_type", {
            "cluster_0": "browsers",
            "cluster_1": "researchers",
            "cluster_2": "buyers",
            "cluster_3": "light_users",
        })
)

Once clusters are saved as a segment, explore them like any other segment: compare them in the diff mode of any widget or side by side in Segment Overview.

How does behavior change across visits?

Cluster sessions according to behavioral types like this:

stream.cluster_analysis(
    features=[{"metric": "event_count_bulk"}],
    method_args={"n_clusters": "4-8"},
    path_col="session_id"
)

Suppose 4 clusters look best. add_clusters reproduces that split as the session_type segment – the same result as Save Clusters in the widget. Then we can collapse sessions and treat them as single events with the collapse_events data processor and visualize the session flow with any widget, such as the transition graph:

visits = (
    stream
        .add_clusters(
            "session_type", features=[{"metric": "event_count_bulk"}],
            method_args={"n_clusters": 4}, path_col="session_id",
        )
        .collapse_events(group_col="session_id", name={"col": "session_type"})
        .rename_events({
            "cluster_0": "browsing",
            "cluster_1": "quick_visit",
            "cluster_2": "purchase_visit",
            "cluster_3": "promo_visit"
        })
)
visits.transition_graph()

Transition graph of users' visits, where each session is collapsed into its behavioral cluster.

See more analysis recipes in the docs.

Bring your own data

The minimum input is a table with three columns – user_id, event and timestamp:

user_id event timestamp
user_1 signup 2026-01-01 10:00:00
user_1 project_created 2026-01-01 10:02:00
user_1 teammate_invited 2026-01-01 10:05:00

Retentioneering uses the Eventstream class to store data and to provide access to all the library's methods.

import pandas as pd
import retentioneering as rete

events = pd.read_csv("events.csv", parse_dates=["timestamp"])
stream = rete.Eventstream(events)
stream.transition_graph()

Use a schema to map other column names, declare path segments or pre-existing sessions. Any dataset readable by pandas – such as CSV, Parquet or the result of a warehouse query – can become an eventstream.

Map an export from GA4, Amplitude, Mixpanel or Segment

These are starting points for common exports; check identity and timestamp fields in your actual data.

Source Path ID Event Timestamp
GA4 BigQuery user_pseudo_id event_name event_timestamp (microseconds)
Amplitude user_id or amplitude_id event_type event_time
Mixpanel properties.distinct_id event properties.time (Unix seconds)
Segment tracks user_id or anonymous_id event timestamp

Flatten nested properties and convert numeric timestamps to datetimes with the correct unit before loading. For example, after exporting the three GA4 fields above to a DataFrame named ga4:

ga4["event_timestamp"] = pd.to_datetime(ga4["event_timestamp"], unit="us")
stream = rete.Eventstream(ga4, schema={
    "path_cols": ["user_pseudo_id"], "event_col": "event_name",
    "timestamp_col": "event_timestamp",
})

Common analysis workflow

  • Prepare paths. Data processors filter events or whole journeys, split sessions, collapse repeated actions and define segments. Each returns a new eventstream, so steps chain into a pipeline.
  • Explore. Interactive widgets show paths as graphs, step matrices, Sankey diagrams, funnels and clusters, and compare any two groups.
  • Build your own outputs. Every widget has a headless *_data() twin that returns the numbers behind the chart, for custom charts, reports and checks.
  • Share. export_html() saves a widget as a standalone interactive page that opens without Python.

Explore agent runs, conversations and other event sequences

A path doesn't have to be a user journey. Treat each agent run as a path and its internal steps – tool calls, retries, errors, validations – as events, then use the same tools to see where runs loop, recover or fail, and how successful runs differ from failed ones. See the agent-run example notebook (open in Colab) for a synthetic walkthrough.

Paths of synthetic agent runs, comparing failed runs with successful ones. Red emphasizes the repeated tool:error and retry cycle.

The same approach works for other processes: decide what one path is and which steps become its events. For example:

  • Support conversations: a path is the sequence of message intents and handoffs within one conversation, e.g. billing_question → bot_answer → rephrased_question → escalation → agent_answer → resolved.
  • Tickets: a path is the sequence of a ticket's status changes and reassignments, e.g. opened → triaged → assigned:tier1 → waiting_for_customer → reassigned:tier2 → closed.
  • Learning sessions: a path is the sequence of lessons, attempts and their results within a study session or across a learner's whole course, e.g. lesson_3 → quiz_3:fail → lesson_3_review → quiz_3:pass → lesson_4.

Work with an AI agent

Give your coding agent the product-analytics skill and an event file. The skill guides it through inspecting the log, choosing a recipe, running the analysis and checking the result. See skill setup for supported workflows.

For example:

Use Retentioneering to investigate where sessions in events.csv stall between cart and purchase. Check the schema and event order first. Compare sessions that purchased with those that stopped, keep non-purchasing sessions, and show the counts behind the main differences. Return the runnable analysis and an interactive HTML report.

Prefer an MCP connection? The data MCP server (beta) exposes an Eventstream to a compatible agent from a local Python environment. Its report collects several tools in tabs and links every number in the text to the element it comes from – a graph edge, a step matrix cell or a segment overview cell:

A report with three tabs beside its analysis text. Clicking links in the text switches tabs and highlights segment overview cells, transition graph edges and step matrix cells.

For AI assistants: current API and execution guidance
  • Read the current API index, full text docs or documentation MCP guide. Documentation access supplies reference material; data analysis also needs Python execution and access to the event file.
  • This README uses 5.x. Start with import retentioneering as rete and rete.Eventstream(df, schema={...}). Older 2.x/3.x examples use a different API.
  • Choose and declare the path grain (path_cols); pass path_col to analyses when switching between users and sessions. Declare comparison columns in segment_cols.
  • Use transition_graph, step_matrix, step_sankey, funnel, segment_overview and cluster_analysis. Their *_data() methods return computed results for inspection and automation; return types vary by tool.
  • Graph, matrix, Sankey and funnel comparisons take diff=("segment", "A", "B") and show A minus B. Check group sizes, metric definitions and the corresponding data output.
  • Processors return a new Eventstream. Assign the result, preserve the intended event order and population, and validate a complete sequence with a path pattern rather than inferring it from separate graph edges.
  • Save and replay transformations with Eventstream.recipe(). Keep findings separate from hypotheses about their causes.

Telemetry

Anonymous usage telemetry is enabled by default. Sensitive data like event names, identifiers, path contents and parameter values is never sent; see what is collected. Set RETENTIONEERING_NO_TRACK=1 before starting Python to disable it.

Documentation and community

Quick start · Eventstream · Widgets · Data processors · Analysis recipes

Using a 3.x example? Version 5 rewrites the API. See the migration guide; legacy Sequences, Cohorts, StatTests and the visual Preprocessing Graph remain on the 3.x branch.

Bring a question, a useful recipe or a small reproducible example to GitHub issues, Discord or Telegram. Contributions to methods, widgets, examples and agent workflows are welcome – see CONTRIBUTING.md.

Apps are better with math. Join us!

License and commercial model

Retentioneering-tools is open-source software licensed under the Apache License, Version 2.0.

Retentioneering is a community research laboratory dedicated to developing new analytics methodology and open-source tools.

Copyright retentioneering-tools v.5.0 Maxim Godzi, Vladimir Kukushkin and Anatoly Zaytsev. Updates may include software developed by the Retentioneering community.

You are free to use, modify, distribute, and build commercial products with Retentioneering-tools, subject to the terms of the Apache-2.0 license.

Other Retentioneering libraries, packages and managed execution services, enterprise integrations, premium diagnostic workflows, and hosted collaboration features are separate proprietary products and are governed by their respective commercial terms. Additional details are provided in COMMERCIAL.md.

The Apache-2.0 license applies only to the source code and assets distributed in this repository. It does not grant rights to use the Retentioneering name, logo, trademarks, hosted services, proprietary cloud infrastructure, or commercial content that is not distributed in this repository.

We welcome contributions from individuals and organizations. Contributions to Retentioneering-tools are accepted under the contribution terms described in CONTRIBUTING.md.

Our goal is to keep the core analytical language and ecosystem open, extensible, and useful for independent analysts, researchers, startups, and enterprise teams, while funding long-term maintenance through optional commercial products and services.

Metadata

Release files for retentioneering 5.2.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for retentioneering 5.2.4
File Size Uploaded
retentioneering-5.2.4.tar.gz 6.9 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for retentioneering 5.2.4
File Interpreter ABI Platform
retentioneering-5.2.4-py3-none-any.whl Python 3 none any Details

Total release size: 8.1 MB

Release files / retentioneering-5.2.4.tar.gz

Download URL retentioneering-5.2.4.tar.gz
Size 6.9 MB
Tags Source
SHA-256 checksum
How to use checksums
2cde4a08c7839ef4895ba7b7ddf8936cb637bd9b12971582bac19262622d46a2
BLAKE2b-256 checksum
How to use checksums
d32086394104cf7423a01c03c5abc633b3ce0c6948f9e17820697e9c81fb0c26
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / retentioneering-5.2.4-py3-none-any.whl

Download URL retentioneering-5.2.4-py3-none-any.whl
Size 1.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
f4ab91b67c6deed93728bfb6e503adfe304f101a53e9112ee0219b304b0f3b1a
BLAKE2b-256 checksum
How to use checksums
97df31e3f469c15e8bfa82cad0eb053d4c3a75f3e1571738558ec56238c4e945
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

5.2.4 This release

2 release files

5.2.3

2 release files

5.2.2

2 release files

5.2.1

2 release files

5.2.0

2 release files

5.1.0

2 release files

5.0.1

2 release files

5.0.0

2 release files

3.3.0

2 release files

3.2.1

2 release files

3.2.0

2 release files

3.1.0

2 release files

3.0.1

2 release files

3.0.0

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.0.8

2 release files

1.0.7.6

1 release file

1.0.6

3 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

3 release files

0.3.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page