Skip to main content

Fuzzy Neighborhood DBSCAN clustering algorithm

Project description

FN-DBSCAN: Fuzzy Neighborhood DBSCAN

Python Version License: MIT DOI

Implementation of Fuzzy Neighborhood DBSCAN (FN-DBSCAN), a density-based clustering algorithm that extends classic DBSCAN using fuzzy theory.

Installation

pip install fn-dbscan

For development:

git clone https://github.com/onurceldir123/fn-dbscan.git
cd fn-dbscan
pip install -e .

Requirements: Python ≥3.8, NumPy, scikit-learn, scipy

Quick Start

from sklearn.datasets import make_moons
from fn_dbscan import FN_DBSCAN

X, _ = make_moons(n_samples=200, noise=0.05, random_state=42)

model = FN_DBSCAN(
    eps=0.1,
    min_fuzzy_neighbors=5.0,
    min_membership=0.0,
    fuzzy_function='exponential',
    normalize=False
)

labels = model.fit_predict(X)

print(f"Found {model.n_clusters_} clusters")

Why FN-DBSCAN?

While classic DBSCAN is powerful, it relies on a "crisp" boundary—a point is either a neighbor or it isn't. FN-DBSCAN improves upon this by introducing fuzzy set theory:

  • Robustness to Density Variations: It is more robust than DBSCAN when handling datasets with varying densities and shapes.
  • Soft Boundaries: Instead of an all-or-nothing approach, it calculates a "fuzzy cardinality" (sum of membership degrees). This handles border points and noise more naturally.
  • Scale Adaptability: The implementation includes an optional normalization technique (set normalize=True) to make the eps parameter adaptable to the data scale.
  • Best of Both Worlds: Combines the speed of DBSCAN with the robustness of fuzzy clustering methods like NRFJP.

Parameters

Core Parameters

Parameter Type Default Description
eps float 0.1 Maximum neighborhood radius (0-1 when normalize=True).
min_fuzzy_neighbors float 5.0 Minimum fuzzy cardinality to be a core point (analogous to min_samples in DBSCAN).
min_membership float 0.0 Minimum membership threshold ($\epsilon_1$). Points with membership below this are ignored.
fuzzy_function str 'linear' Membership function: 'linear', 'exponential', or 'trapezoidal'.
normalize bool False Normalize data to make eps scale-independent.
k float None Steepness parameter. Controls how fast membership drops. Higher k implies a stricter neighborhood. Auto-calculated as d_max / eps if None.
metric str 'euclidean' Distance metric (any scikit-learn compatible metric).

Fuzzy Functions

  • 'exponential' - Recommended for most cases, especially non-convex clusters
  • 'linear' - Simple linear decay, good for well-separated clusters
  • 'trapezoidal' - Maintains full membership for very close points

Model Attributes

After fitting, the model provides:

  • labels_ - Cluster labels for each sample (-1 for noise)
  • core_sample_indices_ - Indices of core points
  • n_clusters_ - Number of clusters found

Algorithm Overview

FN-DBSCAN extends DBSCAN by computing fuzzy cardinality instead of discrete point counts:

Traditional DBSCAN:  cardinality = count(neighbors)
FN-DBSCAN:          cardinality = Σ membership(distance(p, q))

A point is a core point if its fuzzy cardinality ≥ min_fuzzy_neighbors.

Citation

If you use fn-dbscan in your research, please consider citing the original paper along with this software implementation to ensure reproducibility:

1. Original Algorithm:

@article{nasibov2009robustness,
  title={Robustness of density-based clustering methods with various neighborhood relations},
  author={Nasibov, Efendi N and Ulutagay, G{\"o}zde},
  journal={Fuzzy Sets and Systems},
  volume={160},
  number={24},
  pages={3601--3615},
  year={2009},
  publisher={Elsevier}
}

2. Software Implementation:

@software{celdir2025fndbscan,
  author       = {Çeldir, Onur Mert},
  title        = {FN-DBSCAN: Python Implementation of Fuzzy Neighborhood DBSCAN},
  year         = 2025,
  publisher    = {Zenodo},
  version      = {v1.0.1},
  doi          = {10.5281/zenodo.17726044},
  url          = {[https://github.com/onurceldir123/fn-dbscan](https://github.com/onurceldir123/fn-dbscan)}
}

License

MIT License - see LICENSE file for details.

Contributing

Contributions welcome! Please open an issue or submit a pull request on GitHub.


Reference: Nasibov, E. N., & Ulutagay, G. (2009). Robustness of density-based clustering methods with various neighborhood relations. Fuzzy Sets and Systems, 160(24), 3601-3615.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fn_dbscan-0.2.0.tar.gz (29.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

fn_dbscan-0.2.0-py3-none-any.whl (13.3 kB view details)

Uploaded Python 3

File details

Details for the file fn_dbscan-0.2.0.tar.gz.

File metadata

  • Download URL: fn_dbscan-0.2.0.tar.gz
  • Upload date:
  • Size: 29.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.1

File hashes

Hashes for fn_dbscan-0.2.0.tar.gz
Algorithm Hash digest
SHA256 788ad8193fc4c06263a6e72e8e70e052f1fd23d6380ac04c8091197dd7302c54
MD5 28eaa20700a7f4b5a6deeab0644879db
BLAKE2b-256 a2bd0c59c5eb5df51b9c7acec7bf50567e910a240f77429ed8a624ed9b7b0cc8

See more details on using hashes here.

File details

Details for the file fn_dbscan-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: fn_dbscan-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 13.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.1

File hashes

Hashes for fn_dbscan-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ffd583f1de8dbe875e22be3be8dbd40dd09d22b22a0683a3b8441ab4d6569f43
MD5 15d1d873d76763f7a8f9a64eca6441f0
BLAKE2b-256 b3bfbe46bb6978c232f2dd15420d4597850ce14d52e3c4db031989bd026ef3a9

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page