Skip to main content

HDDFeaturesX

HDDFeaturesX is a Python library designed for feature selection in high-dimensional datasets. It supports both binary and multi-class problems and is compatible with various machine learning and deep learning models.

This study proposes a feature selection method motivated by rough set theory, inspired by the sample and feature selection approach introduced by Yang in 2022 (Yang et al., 2022). For more details, refer to:
Yang, Y., Chen, D., Zhang, X., Ji, Z., & Zhang, Y. (2022). Incremental feature selection by sample selection and feature-based accelerator. Applied Soft Computing, 121. https://doi.org/10.1016/j.asoc.2022.108800.

In this study, we propose an enhanced version of Induced Partitioning for Incremental Feature Selection, combining Rough Set Theory and the Long-Tail Position Grey Wolf Optimizer. This method has been accepted for publication in Acta Informatica Pragensia.

Objective

The library facilitates feature selection for high-dimensional datasets, supporting:

  • Binary and multi-class classification problems.
  • Seamless integration with machine learning and deep learning models.

Installation

Install the library using pip:

pip install HDDFeaturesXS


## Usage
Here is an example of how to use the library:

from HDDFeaturesXS import *

import numpy as np
import time
from scipy.io import loadmat
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
import pandas as pd

# Load dataset (example using a .mat file)
data_path = 'PCMAC.mat'
data_loaded = loadmat(data_path)
data_key = next(key for key in data_loaded.keys() if not key.startswith('__'))
data = data_loaded[data_key]

# Partition the dataset into two parts
parts = partition_fold(data, 2)
ori_data = parts[0]
A = parts[1]

# Further partition part A into 5 parts
U = partition_fold(A, 5)

# Perform initial feature selection
fea_slt, ds_vector, fea_redun, sam_delete, ori_time = fast_feature_selection_ds(ori_data)

# Incrementally update feature selection as new data arrives
add_data = np.array([])
in_fea_slt_fs = []
unuse_ds = []
nrf_fs = []
nrs_fs = []
time_in_ds_fs = np.array([])

for i in range(5):
    add_data = np.vstack([add_data, U[i]]) if add_data.size else U[i]
    result = incre_fea_slt_ds_filter_sam(ori_data, add_data, ds_vector, fea_slt)
    in_fea_slt_fs.append(result[0])
    unuse_ds.append(result[1])
    nrf_fs.append(result[2])
    nrs_fs.append(result[3])
    time_in_ds_fs = np.append(time_in_ds_fs, result[4])

# Calculate total execution time
end_time = time.time()
total_time = end_time - start_time

# Use the selected features for training and testing
X_selected = ori_data[:, fea_slt]
y = ori_data[:, -1]
y = y.astype(int)

# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X_selected, y, test_size=0.2, random_state=42)

# Train a KNN model
model = KNeighborsClassifier()
model.fit(X_train, y_train)

# Predict classes for the test data
y_pred = model.predict(X_test)

# Evaluate model performance
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

# Save incremental feature selection execution times
results_path = 'timeInDSfs.txt'
np.savetxt(results_path, time_in_ds_fs, delimiter=',')

# Print total execution time
print(f"Total execution time: {total_time} seconds")

Release files for HDDFeaturesXS 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for HDDFeaturesXS 0.1.0
File Size Uploaded
HDDFeaturesXS-0.1.0.tar.gz 5.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for HDDFeaturesXS 0.1.0
File Interpreter ABI Platform
HDDFeaturesXS-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size:13.9 kB

Release files / HDDFeaturesXS-0.1.0.tar.gz

Download URL HDDFeaturesXS-0.1.0.tar.gz
Size 5.9 kB
Tags Source
SHA-256 checksum
How to use checksums
1c0335c69ae3c212dc75c672dfa602ad1d029c15512a459b39fa62a998cacfec
BLAKE2b-256 checksum
How to use checksums
496332bdccb5204eaa6a44ef201d73dfb113d9294d2f088cd996002708a7c3ed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.9

Release files / HDDFeaturesXS-0.1.0-py3-none-any.whl

Download URL HDDFeaturesXS-0.1.0-py3-none-any.whl
Size 8.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
03060458ce43682dfb254a883b48b12f4bcfff87b2cc5a974719d1885a237bca
BLAKE2b-256 checksum
How to use checksums
149e09078c00017c58531896e077ff4091e4b8f7d4eff6c5ef3f22f5e6e50c82
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.9

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page