Skip to main content

Regression automata for attribute path analysis via correlation sequences

Project description

RegAutomata

RegAutomata is a Python library for building regression-based automata from datasets. It generates regression plots between attributes, constructs directed graphs, and finds the best attribute path based on correlation.

Installation

pip install regAutomata

Usage

"""
regAutomata — example usage
Demonstrates every public parameter across all regression types and datasets.
"""

from regAutomata import (
    run_regAutomata,
    predictRegAutomata,
    load_artifacts,
    get_correlation_matrix,
    SUPPORTED_REGRESSION,
    SUPPORTED_CORRELATION,
)


# ── 1. Quick single run ────────────────────────────────────────────────────────
# Minimal call — only required parameters, everything else uses defaults.

artifacts = run_regAutomata(
    dataset_csv="data/iris/iris_dataset.csv",
    qS="sepal length (cm)",
    qF="petal width (cm)",
)


# ── 2. Full parameter run ──────────────────────────────────────────────────────
# Every parameter explicitly set so you can see what each one does.

artifacts = run_regAutomata(
    # --- Data ---
    dataset_csv="data/iris/iris_dataset.csv",   # Path to any CSV file

    # --- Automaton definition ---
    qS="sepal length (cm)",                     # Initial state  (must be a column name)
    qF="petal width (cm)",                      # Final state    (must be a column name)

    # --- Regression ---
    # One of: "linear" | "polynomial" | "cubic" | "loess" | "ridge" | "lasso" | "svr" | "gbr"
    regression_type="linear",

    # --- Correlation ---
    # One of: "pearson" | "spearman" | "kendall"
    # If None, a sensible default is chosen automatically per regression type.
    correlation_method="pearson",

    # --- Graph scope ---
    # "full" — show all explored paths in the HTML
    # "best" — show only the highest-correlation path
    visualization_type="full",

    # --- HTML output ---
    output_html="regAutomata_linear_full.html",  # Path for the interactive graph
    open_html=False,                              # Set True to open browser automatically

    # --- Artifacts ---
    # Saves two .pkl files:
    #   regAutomata_linear.pkl        — full graph artifacts
    #   regAutomata_linear_best.pkl   — best-path-only artifacts (used for prediction)
    save_artifacts_path="regAutomata_linear.pkl",

    # --- PNG export (requires: pip install playwright && playwright install chromium) ---
    save_png=False,                              # Set True to capture HTML → PNG
    png_kwargs={
        "out_png": "regAutomata_linear_full.png",
        "extra_left_px": 60,
        "extra_right_px": 60,
        "extra_top_px": 60,
        "extra_bottom_px": 60,
        "device_scale": 2,                       # 2 = 2× pixel density (retina quality)
        "wait_after_load": 1.0,                  # Seconds to wait after vis.js renders
    },

    # --- Pre-filter ---
    # True  — removes low-correlation pairs before path generation.
    #         Threshold: |corr| >= (max + mean) / 2  (adaptive, no fixed cutoff)
    #         Strongly recommended for datasets with more than ~6 columns.
    # False — explore all possible paths (factorial growth, slow on wide datasets)
    use_prefilter=True,
)

print("Best path:", artifacts["paths"][artifacts["best_path_index"]])
print("Regression:", artifacts["regression_type"])
print("Correlation method:", artifacts["correlation_method"])


# ── 3. Prediction from saved artifacts ────────────────────────────────────────
# Always load the *_best.pkl file for prediction — it contains only the
# highest-correlation path and its fitted models.

seq, final_value = predictRegAutomata(
    artifacts_or_path="regAutomata_linear_best.pkl",  # Path or already-loaded dict
    x0=5.1,                                           # Starting value at qS
    qS="sepal length (cm)",                           # Must match the artifact's qS

    # Optional: save a new HTML/PNG with predicted p-values annotated on each node
    save_html=None,           # e.g. "prediction_result.html"
    save_png_path=None,       # e.g. "prediction_result.png"  (requires Playwright)
)

print("\nPrediction sequence:")
for a1, a2, val in seq:
    print(f"  {a1}{a2}  :  {val:.4f}")
print(f"Final predicted value at '{artifacts['qF']}': {final_value:.4f}")


# ── 4. Prediction from an already-loaded artifact dict ────────────────────────
# You can also pass the dict returned by run_regAutomata or load_artifacts
# directly — no file I/O needed.

loaded = load_artifacts("regAutomata_linear_best.pkl")
seq2, final2 = predictRegAutomata(loaded, x0=6.3, qS="sepal length (cm)")
print(f"\nDirect dict prediction  x0=6.3  →  {final2:.4f}")


# ── 5. Correlation matrix utility ─────────────────────────────────────────────
import pandas as pd

df = pd.read_csv("data/iris/iris_dataset.csv")
for method in SUPPORTED_CORRELATION:
    cm = get_correlation_matrix(df, method=method)
    print(f"\nCorrelation matrix ({method}):")
    print(cm.round(3).to_string())


# ── 6. Loop over all regression types ─────────────────────────────────────────
# Useful for comparing model quality across regression algorithms on the same data.

results = {}
for rtype in SUPPORTED_REGRESSION:
    art = run_regAutomata(
        dataset_csv="data/iris/iris_dataset.csv",
        qS="sepal length (cm)",
        qF="petal width (cm)",
        regression_type=rtype,
        save_artifacts_path=f"artifacts_{rtype}.pkl",
        save_png=False,
        use_prefilter=True,
    )
    seq, final = predictRegAutomata(
        f"artifacts_{rtype}_best.pkl",
        x0=5.1,
        qS="sepal length (cm)",
    )
    best_path = art["paths"][art["best_path_index"]]
    results[rtype] = {"path": best_path, "prediction": round(final, 4)}
    print(f"[{rtype:12s}]  path={best_path}  pred={final:.4f}")

print("\nSummary:")
for rtype, info in results.items():
    print(f"  {rtype:12s}{info['prediction']}")


# ── 7. Three datasets — thesis verification ───────────────────────────────────

DATASETS = [
    {
        "dataset_csv": "data/iris/iris_dataset.csv",
        "qS": "sepal length (cm)",
        "qF": "petal width (cm)",
        "x0": 5.1,
    },
    {
        "dataset_csv": "data/wine/wine_dataset.csv",
        "qS": "alcohol",
        "qF": "proline",
        "x0": 13.0,
    },
    {
        "dataset_csv": "data/diabetes/diabetes_dataset.csv",
        "qS": "bmi",
        "qF": "target",
        "x0": 0.05,
    },
]

for ds in DATASETS:
    print(f"\n{'─'*50}")
    print(f"Dataset : {ds['dataset_csv']}")
    print(f"Path    : {ds['qS']}{ds['qF']}")

    art = run_regAutomata(
        dataset_csv=ds["dataset_csv"],
        qS=ds["qS"],
        qF=ds["qF"],
        regression_type="polynomial",
        visualization_type="best",
        output_html=f"graph_{ds['qS'].split()[0]}.html",
        save_artifacts_path=f"artifacts_{ds['qS'].split()[0]}.pkl",
        save_png=False,
        use_prefilter=True,
    )

    seq, final = predictRegAutomata(
        f"artifacts_{ds['qS'].split()[0]}_best.pkl",
        x0=ds["x0"],
        qS=ds["qS"],
    )

    print(f"Best path : {art['paths'][art['best_path_index']]}")
    print(f"Prediction: {final:.4f}")

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

regautomata-0.1.3.tar.gz (19.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

regautomata-0.1.3-py3-none-any.whl (15.2 kB view details)

Uploaded Python 3

File details

Details for the file regautomata-0.1.3.tar.gz.

File metadata

  • Download URL: regautomata-0.1.3.tar.gz
  • Upload date:
  • Size: 19.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.0

File hashes

Hashes for regautomata-0.1.3.tar.gz
Algorithm Hash digest
SHA256 055b86dd22cba8394151342106bb319d27f771e33d3ec62956d2c0bc0580729a
MD5 74c85f9930987d771db9a9941ad409dc
BLAKE2b-256 bfaa181ec90280b7bad2fc72401171e9d2da468d7148b26fbe4940f9448f3745

See more details on using hashes here.

File details

Details for the file regautomata-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: regautomata-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 15.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.0

File hashes

Hashes for regautomata-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 1869c67aeac6a2b87b95cdf3379206100e3bb2882d536d60c984ba0907a37d37
MD5 c04d6f1cafad124665bbe8df93d5d8e7
BLAKE2b-256 0e5567c621084044eec56d5ea38e4eea157d5714465656bac2e77539d65cb22e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page