Skip to main content

Legendarium Usage Guide

Legendarium is a Python library for storing experiment parameters and metrics in a compact, reloadable format. It provides a Legendarium writer and functions for loading experiments into pandas dataframes.

Installation

Install the published package with pip:

python -m pip install sylegendarium

For a local checkout, install the package in editable mode from the repository root:

python -m pip install -e .

The package requires Python, pandas, NumPy, tqdm, and rapidfuzz. Parquet support also requires a pandas-compatible Parquet engine, such as pyarrow:

python -m pip install pyarrow

Core Concepts

An experiment has:

  • Parameters: values that describe the experiment and are repeated for its rows, such as the algorithm name or a configuration value.
  • Metrics: values recorded for each (run, step) pair.
  • Metadata: type, description, unit, and storage information written to the experiment metadata file.

Metric storage is selected from its declared type:

Declared type Storage Recommended for
int, float, str, bool Parquet Scalar values and values that are efficient in columns
numpy.ndarray NPZ Arrays with the same shape for every row in that metric
list, dict, tuples, or other non-Parquet-safe values Heavy metrics Variable-size or complex Python objects

A NumPy metric whose arrays do not all have the same shape is automatically moved to heavy storage during dataframe conversion. This preserves the values but is slower and uses more disk space.

Creating and Writing an Experiment

Declare parameters and metrics before writing rows:

from sylegendarium import Legendarium
import numpy as np

experiment = Legendarium(
    experiment_name="navigation_run",
    experiment_description="Navigation policy comparison",
    path="experiments",
    allow_overwrite=True,
    verbose=True,
    show_progress=True,
)

experiment.create_parameter("algorithm", "Greedy")
experiment.create_parameter("max_distance", 100.0)

experiment.create_metric(
    "reward",
    float,
    "Reward obtained by the agent",
    "points",
)
experiment.create_metric(
    "map",
    np.ndarray,
    "Environment map",
    "pixels",
)
experiment.create_metric(
    "objects",
    list,
    "Objects detected at this step",
    "objects",
)

with experiment:
    for run in range(3):
        for step in range(100):
            experiment.write(
                run=run,
                step=step,
                reward=float(np.random.rand()),
                map=np.random.rand(100, 100),
                objects=[f"object_{i}" for i in range(np.random.randint(1, 5))],
            )

with experiment calls save() automatically when the block exits. This also happens if the block exits because of an exception. Call experiment.save() explicitly when using the object without a context manager.

Each call to write() must provide integer run and step values and one value for every declared metric.

Converting an Existing DataFrame

pd_to_experiment() is intended for dataframes loaded from older or external pipelines. The dataframe must contain run and step columns.

from sylegendarium import pd_to_experiment

pd_to_experiment(
    data_reader,
    experiment_name="newprocessing_experiment",
    experiment_description="Converted final BGA dataset",
    path="test",
    verbose=True,
    show_progress=True,
)

When the dataframe has no metadata attribute, Legendarium infers the metric type from the first non-null value in each column. In this case, columns other than run and step are treated as metrics.

For large dataframes, this function uses batch dataframe construction instead of appending one pandas row at a time. Progress output reports metric preparation, heavy-row preparation, and save phases.

To run silently:

pd_to_experiment(
    data_reader,
    experiment_name="quiet_conversion",
    path="test",
    show_progress=False,
)

Loading Experiments

Load one experiment:

from sylegendarium import load_experiment_pd

df = load_experiment_pd("navigation_run", "experiments")
print(df.head())

Load every experiment found in a directory:

from sylegendarium import load_experiments

all_runs = load_experiments("experiments")
print(all_runs.shape)
print(all_runs.columns)

The returned dataframe includes the experiment metadata in its metadata attribute:

print(all_runs.metadata)

Older experiment files are detected and delegated to the legacy loader when possible.

Monitoring Progress and Time

Enable progress when creating a writer:

experiment = Legendarium(
    "large_experiment",
    "Large experiment",
    "experiments",
    show_progress=True,
    verbose=True,
)

For pd_to_experiment(), use show_progress=True. The automatic save reports phases such as:

[Legendarium]: save [0.1s]: writing metadata
[Legendarium]: save [1.8s]: writing parquet
[Legendarium]: save [7.4s]: writing heavy metrics
[Legendarium]: save [42.0s]: creating tar.gz archive

The progress bars cover metric preparation, heavy rows, categorical columns, NumPy metrics, and heavy rows. Parquet writing and archive compression are single library calls, so they report start and completion rather than a percentage.

Files Created

For an experiment named navigation_run, the output directory contains:

  • navigation_run.meta.yaml: experiment metadata and storage indexes.
  • navigation_run.metrics.tar: compressed archive containing the stored data.
  • A Parquet file inside the archive for scalar and columnar data.
  • An NPZ file inside the archive when NumPy metrics are present.
  • An LZMA-compressed pickle stream inside the archive when heavy metrics are present.

The archive is the primary data file. Do not delete the metadata file or archive independently.

Practical Recommendations

  • Declare all parameters and metrics before entering the writing context.
  • Use scalar metrics for values that can be represented as columns.
  • Use NumPy metrics only when every row of that metric has the same array shape.
  • Use heavy metrics for variable-size arrays, lists, dictionaries, and custom Python objects.
  • For large legacy dataframes, prefer pd_to_experiment() with show_progress=True and allow the batch conversion to finish before profiling compression.
  • Restart a notebook kernel after changing the installed Legendarium source so the notebook does not keep an older imported module in memory.

Troubleshooting

Length of values does not match length of index

This usually indicates that heavy data was assigned one serialized row at a time instead of being accumulated for the whole dataframe. Use the current loader and regenerate the experiment if it was created by an older conversion implementation.

all input arrays must have the same shape

A NumPy metric contains arrays with incompatible shapes. Declare it as a heavy metric, or regenerate it with the current dataframe conversion, which automatically moves only that variable-shape metric to heavy storage.

Metadata

Release files for sylegendarium 1.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sylegendarium 1.0.1
File Size Uploaded
sylegendarium-1.0.1.tar.gz 19.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sylegendarium 1.0.1
File Interpreter ABI Platform
sylegendarium-1.0.1-py2.py3-none-any.whl Python 2, Python 3 none any Details

Total release size: 39.8 kB

Release files / sylegendarium-1.0.1.tar.gz

Download URL sylegendarium-1.0.1.tar.gz
Size 19.3 kB
Tags Source
SHA-256 checksum
How to use checksums
7f09c56f5d3748f9809a9452aed157369c4f4ac7f96579b224be8b360493473e
BLAKE2b-256 checksum
How to use checksums
4ba373447096d09077d6c85b6623e4727198e3af8524530957f905f88182cd63
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.12

Release files / sylegendarium-1.0.1-py2.py3-none-any.whl

Download URL sylegendarium-1.0.1-py2.py3-none-any.whl
Size 20.5 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
50f27a70ca68fdb9144df7f89de56592c785a1c8125036aba0196651e0954813
BLAKE2b-256 checksum
How to use checksums
b2e03213542ca2ed3da8e38798bbc34e7fbdeaff00ccc5ceadd3ddb13dccd7be
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.12

Release history Release notifications | RSS feed

This release

1.0.1 This release

2 release files

1.0.0

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page