Skip to main content

A library for feature selection in high-dimensional datasets

Project description

HDDFeaturesX

HDDFeaturesX is a Python library designed for feature selection in high-dimensional datasets. It supports both binary and multi-class problems and is compatible with various machine learning and deep learning models.

This study proposes a feature selection method motivated by rough set theory, inspired by the sample and feature selection approach introduced by Yang in 2022 (Yang et al., 2022). For more details, refer to:
Yang, Y., Chen, D., Zhang, X., Ji, Z., & Zhang, Y. (2022). Incremental feature selection by sample selection and feature-based accelerator. Applied Soft Computing, 121. https://doi.org/10.1016/j.asoc.2022.108800.

In this study, we propose an enhanced version of Induced Partitioning for Incremental Feature Selection, combining Rough Set Theory and the Long-Tail Position Grey Wolf Optimizer. This method has been accepted for publication in Acta Informatica Pragensia.

Objective

The library facilitates feature selection for high-dimensional datasets, supporting:

  • Binary and multi-class classification problems.
  • Seamless integration with machine learning and deep learning models.

Installation

Install the library using pip:

pip install HDDFeatures


## Usage
Here is an example of how to use the library:

from HDDFeatures import *

import numpy as np
import time
from scipy.io import loadmat
from partition_fold import partition_fold
from fastFeatureSelectionDS import fast_feature_selection_ds
from increFeaSltDSFilterSam import incre_fea_slt_ds_filter_sam
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
import pandas as pd

# Load dataset (example using a .mat file)
data_path = 'PCMAC.mat'
data_loaded = loadmat(data_path)
data_key = next(key for key in data_loaded.keys() if not key.startswith('__'))
data = data_loaded[data_key]

# Partition the dataset into two parts
parts = partition_fold(data, 2)
ori_data = parts[0]
A = parts[1]

# Further partition part A into 5 parts
U = partition_fold(A, 5)

# Perform initial feature selection
fea_slt, ds_vector, fea_redun, sam_delete, ori_time = fast_feature_selection_ds(ori_data)

# Incrementally update feature selection as new data arrives
add_data = np.array([])
in_fea_slt_fs = []
unuse_ds = []
nrf_fs = []
nrs_fs = []
time_in_ds_fs = np.array([])

for i in range(5):
    add_data = np.vstack([add_data, U[i]]) if add_data.size else U[i]
    result = incre_fea_slt_ds_filter_sam(ori_data, add_data, ds_vector, fea_slt)
    in_fea_slt_fs.append(result[0])
    unuse_ds.append(result[1])
    nrf_fs.append(result[2])
    nrs_fs.append(result[3])
    time_in_ds_fs = np.append(time_in_ds_fs, result[4])

# Calculate total execution time
end_time = time.time()
total_time = end_time - start_time

# Use the selected features for training and testing
X_selected = ori_data[:, fea_slt]
y = ori_data[:, -1]
y = y.astype(int)

# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X_selected, y, test_size=0.2, random_state=42)

# Train a KNN model
model = KNeighborsClassifier()
model.fit(X_train, y_train)

# Predict classes for the test data
y_pred = model.predict(X_test)

# Evaluate model performance
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

# Save incremental feature selection execution times
results_path = 'timeInDSfs.txt'
np.savetxt(results_path, time_in_ds_fs, delimiter=',')

# Print total execution time
print(f"Total execution time: {total_time} seconds")

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

HDDFeaturesX-0.1.1.tar.gz (5.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

HDDFeaturesX-0.1.1-py3-none-any.whl (8.0 kB view details)

Uploaded Python 3

File details

Details for the file HDDFeaturesX-0.1.1.tar.gz.

File metadata

  • Download URL: HDDFeaturesX-0.1.1.tar.gz
  • Upload date:
  • Size: 5.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.10.9

File hashes

Hashes for HDDFeaturesX-0.1.1.tar.gz
Algorithm Hash digest
SHA256 4d64a8c86e6062f4590a784c5cb0fc34623037d83107c28706d5c72f78c6fd41
MD5 f2ed2a868e3edf39bfb180a12f2c6f06
BLAKE2b-256 7b9338d9effaa2808bbff2e99672b2fd98123ca43db36b4e603e1a0444c4c3bf

See more details on using hashes here.

File details

Details for the file HDDFeaturesX-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: HDDFeaturesX-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 8.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.10.9

File hashes

Hashes for HDDFeaturesX-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 4c8a82446637c611eac5ce3c7b2198b1928ae57b5dfca74ac7ab9985088566a5
MD5 587e1c50f924e7eaf82feb389bbe4eb4
BLAKE2b-256 2f202ccac8193083ea2799aab82d74e14831ed9c89a0569535fe2c313545f64e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page