Skip to main content

Fuzzy-rough set utilities for Python.

Project description

frsutils Logo

frsutils

frsutils is a Python library for reusable fuzzy-rough set utilities. The package focuses on fuzzy-rough core building blocks such as similarity matrices, t-norms, implicators, fuzzy quantifiers, fuzzy-rough models, lower/upper approximations, and positive-region computation.

For Developers

If you are extending frsutils, start with the public compatibility boundary in docs/public_api_contract.md, then follow the release and documentation checks in docs/release_checklist.md, docs/documentation_smoke_check.md, docs/submit_readiness_report.md, and docs/joss_final_submission_checklist.md.

Installation

Install the fuzzy-rough core package:

pip install frsutils

For local development from this repository:

pip install -e .

Core requirements

frsutils core intentionally keeps the mandatory dependency set small, but the public API includes an sklearn-style positive-region scorer, so scikit-learn is part of the runtime contract:

  • Python >= 3.10
  • NumPy >= 1.21.0
  • scikit-learn

Optional development / dataset / GPU dependencies

Install these only when you need the related workflows:

  • pytest for tests
  • pandas, openpyxl for some dataset utilities
  • colorlog for colored logging; frsutils falls back to standard logging if it is not installed
  • matplotlib for plotting examples/tests
  • cupy-cuda12x or another CUDA-compatible CuPy wheel for explicit backend="cupy" experiments

Fuzzy-rough set utilities

frsutils provides reusable fuzzy-rough set calculations used in research, including:

  • lower approximation
  • upper approximation
  • positive region
  • boundary region

Public API quickstart

The canonical user-facing API is the package root, frsutils. End users, notebooks, examples, and downstream packages should import from this namespace instead of importing directly from internal frsutils.core or frsutils.utils modules. As an example:

from frsutils import compute_approximations

The package root, frsutils, exposes the intended stable public objects while keeping internal implementation details out of the public contract.

The smallest workflow is to prepare normalized numeric data, compute fuzzy-rough approximations, and read the named fields from the returned result object.

Example:

import numpy as np

from frsutils import compute_approximations, compute_positive_region

# frsutils expects numeric feature values on a comparable scale. In real
# experiments, normalize or scale your data before calling the fuzzy-rough API.
X = np.array(
    [
        [0.00, 0.10],
        [0.08, 0.18],
        [0.15, 0.12],
        [0.80, 0.82],
        [0.88, 0.90],
        [0.95, 0.86],
    ],
    dtype=float,
)
y = np.array([0, 0, 0, 1, 1, 1], dtype=int)

result = compute_approximations(
    X,
    y,
    model="itfrs",
    similarity="linear",
)

print("lower approximation:", result.lower)
print("upper approximation:", result.upper)
print("boundary region:", result.boundary)
print("positive region:", result.positive_region)

# Shortcut when only positive-region scores are needed.
scores = compute_positive_region(
    X,
    y,
    model="itfrs",
    similarity="linear",
)
print("positive-region scores:", scores)

For reusable fitted scoring workflows, use the sklearn-style positive-region scorer:

import numpy as np

from frsutils import FuzzyRoughPositiveRegionScorer

X = np.array(
    [
        [0.00, 0.10],
        [0.08, 0.18],
        [0.15, 0.12],
        [0.80, 0.82],
        [0.88, 0.90],
        [0.95, 0.86],
    ],
    dtype=float,
)
y = np.array([0, 0, 0, 1, 1, 1], dtype=int)

scorer = FuzzyRoughPositiveRegionScorer(
    model="owafrs",
    similarity="linear",
)

scores = scorer.fit_score(X, y)
result = scorer.as_result()

print(scores)
print(result.lower)
print(result.upper)

Downstream packages can reuse a precomputed similarity matrix through the public API without importing from frsutils internals:

import numpy as np

from frsutils import build_similarity_matrix, compute_positive_region

X = np.array(
    [
        [0.00, 0.10],
        [0.08, 0.18],
        [0.15, 0.12],
        [0.80, 0.82],
        [0.88, 0.90],
        [0.95, 0.86],
    ],
    dtype=float,
)
y = np.array([0, 0, 0, 1, 1, 1], dtype=int)

similarity_matrix = build_similarity_matrix(X, similarity="linear")

scores = compute_positive_region(
    X=None,
    y=y,
    model="itfrs",
    similarity_matrix=similarity_matrix,
)

print(scores)

A runnable version of the quickstart is available at examples/public_api_quickstart.py. See docs/public_api.md for the public API guide.

Execution engines and backend status

frsutils now exposes dense and exact blockwise execution through the public API. Dense mode preserves the historical full-matrix behavior. Blockwise mode avoids materializing the full n x n similarity matrix for approximation computation and is available for ITFRS, VQRS, and OWAFRS.

from frsutils import compute_approximations

result = compute_approximations(
    X,
    y,
    model="itfrs",
    similarity="linear",
    engine="blockwise",
    block_size=512,
    backend="numpy",
)

backend="cupy" is an optional experimental backend for GPU-accelerated similarity-block computation. For model="itfrs" and model="vqrs" with engine="blockwise", approximation reductions/accumulators can also stay CuPy-resident until final public NumPy output conversion. OWAFRS deliberately remains on the conservative NumPy row-buffer path after the OWAFRS non-GPU-resident decision because exact OWA execution requires row-wise sorting and a separate memory/sorting benchmark. Do not claim full GPU-native fuzzy-rough execution yet. See docs/backend_execution_status.md and docs/owafrs_non_gpu_resident_decision.md.

The returned result records execution provenance so benchmark scripts and downstream packages can verify which path was used:

result.engine                      # "dense" or "blockwise"
result.backend                     # "numpy" or resolved optional backend
result.block_size                  # None for dense; integer for blockwise
result.used_blockwise              # bool
result.used_gpu_similarity_blocks          # bool
result.used_gpu_approximation_accumulators # bool, true for CuPy blockwise ITFRS/VQRS; false for OWAFRS

The sklearn-style FuzzyRoughPositiveRegionScorer accepts the same engine, backend, and block_size parameters.

Benchmark suite

frsutils includes a reproducible benchmark harness for the public approximation API:

python benchmarks/benchmark_fuzzy_rough_execution.py     --models itfrs,vqrs,owafrs     --sample-sizes 128,256,512     --n-features 8     --block-sizes 64,128     --scenarios dense_numpy,blockwise_numpy,blockwise_cupy     --repeats 3     --output-json benchmark_results.json     --output-csv benchmark_results.csv

The suite compares dense NumPy, exact blockwise NumPy, and optional CuPy-backed blockwise execution. It records runtime, lightweight Python allocator peak memory, dense-reference numerical-equivalence errors, and public execution metadata. CuPy/CUDA-unavailable rows are reported as skipped. See docs/benchmark_suite.md.

Release-ready examples and paper claim boundary

The repository includes small release-ready examples:

python examples/public_api_quickstart.py
python examples/benchmark_smoke.py --output-dir benchmark_smoke_output

Use the wording in docs/paper_claims.md when describing frsutils in a release note, software paper, or benchmark report. The safe claim is that frsutils provides dense and exact blockwise fuzzy-rough approximation APIs, optional CuPy-accelerated similarity blocks, and experimental CuPy-resident blockwise approximation accumulators for ITFRS/VQRS. Public outputs remain NumPy arrays, and OWAFRS remains on the conservative exact blockwise NumPy row-buffer path in this release.

Before tagging or submitting, use docs/release_checklist.md and the documentation smoke check.

Algorithms and contents

Fuzzy-rough oversampling boundary

Fuzzy-rough oversampling algorithms are no longer part of frsutils core. They live in the standalone frsampling package, which depends on frsutils through the public frsutils namespace. frsutils intentionally does not provide old FRSMOTE compatibility wrappers.

frsampling  --->  frsutils
frsutils    -X->  frsampling

frsutils should be cited/used as the fuzzy-rough core engine: similarities, t-norms, implicators, fuzzy quantifiers, approximation models, lower/upper approximation, and positive region. Oversampling algorithms such as FRSMOTE and future FRADASYN belong to the downstream oversampling package.

Notes and assumptions

  • All functions expect normalized scalar values or normalized NumPy arrays.
  • Make sure the input dataset is normalized. This library expects numeric inputs used by fuzzy-rough computations to be in the range [0, 1].
  • This library uses all features of data instances to calculate fuzzy-rough measures.
  • Positive region, lower approximation, upper approximation, etc. are calculated based on the class of each instance.

Docs

  • We use compact NumPy-style Python docstrings and keep longer examples in README/docs.
  • To see online documentation, please visit online documentation.

How to run tests

From the repository root, the default test command excludes tests marked as slow via pyproject.toml:

python -m pytest tests -q

to run all tests, including slow-marked ones:

python -m pytest tests benchmarks -m "slow or not slow" -vv -rs

Run the documented quickstart and release/backend smoke set explicitly with:

python examples/public_api_quickstart.py
python -m pytest tests/api/test_public_api_examples_smoke.py -q -rs
python -m pytest tests/api tests/core_tests/test_approximation_engines.py -q -rs

See docs/release_validation_commands.md for the complete validation command list.

Run exhaustive slow model-combination tests separately when needed:

python -m pytest tests/models_tests -m slow -o addopts="" -q

For the standalone oversampling package, run from the frsampling repository root after making frsutils importable:

PYTHONPATH="$PWD/src:../frsutils" python -m pytest tests -q

For more information on test procedures, please refer to test procedures.

Technical decisions justification

  • Since data checking can slow down experiments, heavy numeric functions do not perform repeated input-range checks. Validation is preferred at construction or workflow boundaries.

Maintenance notes

  • Exhaustive model-combination tests are marked as slow; run them explicitly with python -m pytest tests/models_tests -m slow -o addopts="" -q.
  • VQRS is implemented and covered by the public API/blockwise/backend tests.
  • New feature work should be deferred until the release/paper cleanup checklist is complete.

License

This project is licensed under the BSD-3-Clause License. See the LICENSE file for details. The package metadata, citation metadata, and Python source headers use the same BSD-3-Clause identifier.

How to cite us in your research papers

If you use this library in your research, please cite the software metadata in CITATION.cff. After the JOSS paper is accepted, cite the JOSS paper DOI as the preferred citation.

APA:

Mehran Amiri. (2026). frsutils: Fuzzy-Rough Set Utilities for Python (Version 0.0.3) [Computer software]. https://github.com/mehi64/frsutils

BibTeX:

@software{Amiri_frsutils_2026,
  author = {Amiri, Mehran},
  title = {frsutils: Fuzzy-Rough Set Utilities for Python},
  url = {https://github.com/mehi64/frsutils},
  version = {0.0.3},
  year = {2026}
}

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

frsutils-0.0.5.tar.gz (60.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

frsutils-0.0.5-py3-none-any.whl (74.8 kB view details)

Uploaded Python 3

File details

Details for the file frsutils-0.0.5.tar.gz.

File metadata

  • Download URL: frsutils-0.0.5.tar.gz
  • Upload date:
  • Size: 60.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.13

File hashes

Hashes for frsutils-0.0.5.tar.gz
Algorithm Hash digest
SHA256 ca46b6d96ffb5bbaa94226fcf633e37e60593dd1a5f93a386aedf934a055d2d8
MD5 73c54cccced6411b589a4eda413cdcaf
BLAKE2b-256 c5750ddf288ed73cc42c7f044409c28aeeb85004dd1b4299d63feab052d97df7

See more details on using hashes here.

File details

Details for the file frsutils-0.0.5-py3-none-any.whl.

File metadata

  • Download URL: frsutils-0.0.5-py3-none-any.whl
  • Upload date:
  • Size: 74.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.13

File hashes

Hashes for frsutils-0.0.5-py3-none-any.whl
Algorithm Hash digest
SHA256 cc50067d7ff8889b1effe4682bdb64319a3e86afa095eaba883fa20c5c599420
MD5 389dae0861c1d4adfc01d42d18ac71dd
BLAKE2b-256 2f439326d00f69c36cf2184d4a372b53b784a473ac5328fd28d10cb474b92a24

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page