Skip to main content

A library for String Grammar Hard C-Means

Project description

What is String Grammar Fuzzy Clustering?

String Grammar Fuzzy Clustering is a clustering framework designed for syntactic or structural pattern recognition, where each data instance is represented not as a numeric vector but as a string that encodes structural information.

Unlike conventional numerical clustering method (e.g., Hard C-Means or Fuzzy C-Means), which assume that data have a fixed-length feature vector whereas structural clustering method operates directly on string data whose lengths and internal structures may vary.

In this approach, each pattern is described by a sequence of primitives (symbols) defined by grammatical rules. This is similar to how a sentence is formed from characters following syntax rules.

To measure similarity between strings, the method employs the Levenshtein distance [1], which counts the minimum number of edit operations (insertions, deletions, substitutions) required to transform on string into another.

The "fuzzy" aspect of this framework allows each string to belong to multiple clusters, with a membership degree that reflects how strongly it is associated with each cluster. This provides a more flexible and realistic clustering behavior compared to traditional "hard" clustering, which forces each sample to belong to only one group.

About This Library

This Python library introduces an algorithm belonging to the String Grammar Clustering framework, namely the String Grammar Hard C-Means (sgHCM).

String Grammar Hard C-Means (sgHCM)[1,2]

The sgHCM (String Grammar Hard C-Means) algorithm is an extension of the conventional Hard C-Means (HCM) [3] clustering algorithm designed for a string data set. Since strings are not numeric vectors, traditional distance measures such as Euclidean distance cannot be applied. Therefore, sgHCM employs the Levenshtein distance to measure the dissimilarity between strings based on the minimum number of edit operations required to transform one string into another. The objective of sgHCM is to assign each string observation to exactly one cluster.

Key Features:

  • Designed for clustering syntactic string patterns.
  • Uses the Levenshtein distance to measure dissimilarity between strings.
  • Represents each cluster by a string grammar-based prototype.
  • Assigns each string to exactly one cluster (hard clustering).
  • Simple and analytically interpretable clustering framework.

**Please be noted that this sgHCM can be used for academic and research purposes only. Please also cite this paper [1,2].**

Reference

[1] S. K. Fu, Syntactic Pattern Recognition and Applications, Prentice-Hall, 1982, Zbl0521.68091.

[2] Sansanee Auephanwiriyakul, and Prach Chaisatian, “Static Hand Gesture Translation Using String Grammar Hard C-Means”, The Fifth International Conference on Intelligent Technologies, Houston, Texas, USA., December 2004.

[3] J.C. Bezdek, Pattern Recognition with Fuzzy Objective Function Algorithms, Springer US., 1981.

Installation

You can install the library using pip:

pip install sgHCM

USAGE

Example Code

import random
from sgHCM import SGHCM # Import the clustering class

if __name__ == "__main__":
    # Set random seed for reproducibility
    random.seed(42)

    # Define a list of strings to cluster
    data = ["book", "back", "boon", "cook", "look", "cool", "kick", "lack", "rack", "tack"]

    # Create the model with 2 clusters and fuzzifier m=2.0
    model = SGHCM(C=2)

    # Fit the model on the data
    model.fit(data)

    # Print the final prototype strings representing each cluster
    print("Prototypes:", model.prototypes())

    # Define new strings to classify using the trained model
    new_data = ["hack", "rook", "cook"]

    # Predict the cluster index (0 or 1) for each new string
    preds = model.predict(new_data, model.prototypes())
    print("\nPredictions:")
    for s, c in zip(new_data, preds):
        print(f"{s} → Cluster {c+1}")

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sghcm-0.0.1.tar.gz (3.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sghcm-0.0.1-py3-none-any.whl (4.1 MB view details)

Uploaded Python 3

File details

Details for the file sghcm-0.0.1.tar.gz.

File metadata

  • Download URL: sghcm-0.0.1.tar.gz
  • Upload date:
  • Size: 3.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for sghcm-0.0.1.tar.gz
Algorithm Hash digest
SHA256 2c18049bd3a14f3507061dbabe15e67e9e213b60f839493049bc960b192a858c
MD5 a7b2953ca9bb4c71988a11cf6256690c
BLAKE2b-256 91e4955eb6459656f894d6449d5d8d64a4765a2d176813378ac534312f72d7ae

See more details on using hashes here.

File details

Details for the file sghcm-0.0.1-py3-none-any.whl.

File metadata

  • Download URL: sghcm-0.0.1-py3-none-any.whl
  • Upload date:
  • Size: 4.1 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for sghcm-0.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 f8b40768a24cb1dc9d3101dbdfed62ee1fa180c1e855d1609f14b631ab8292aa
MD5 532efe8b98bd99d31a8b70ae46c28a29
BLAKE2b-256 363ae83d982104c2d26fec7d4af4dd4fa3f1c8082d3aca3d2410d2edce4e0453

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page