Skip to main content

A Python implementation of Synthetic Minority Over-Sampling Technique for Regression with Gaussian Noise (SMOGN)

Project description

SMOGN: Synthetic Minority Over-Sampling with Gaussian Noise

Description

A Python implementation of Synthetic Minority Over-Sampling Technique for Regression with Gaussian Noise (SMOGN). Conducts the Synthetic Minority Over-Sampling Technique for Regression (SMOTER) with traditional interpolation, as well as with the introduction of Gaussian Noise (SMOTER-GN). Selects between the two over-sampling techniques by the KNN distances underlying a given observation. If the distance is close enough, SMOTER is applied. If too far away, SMOTER-GN is applied. Useful for prediction problems where regression is applicable, but the values in the interest of predicting are rare or uncommon. This can also serve as a useful alternative to log transforming a skewed response variable, especially if generating synthetic data is also of interest.

Features

  1. The only open-source Python supported version of Synthetic Minority Over-Sampling Technique for Regression

  2. Supports Pandas DataFrame inputs containing mixed data types, auto distance metric selection by data type, and optional auto removal of missing values

  3. Flexible inputs available to control the areas of interest within a continuous response variable and friendly parameters for over-sampling synthetic data

  4. Purely Pythonic, developed for consistency, maintainability, and future improvement, no foreign function calls to C or Fortran, as contained in original R implementation


Installation

## install pypi release
pip install smogn

## install developer version
pip install git+https://github.com/nickkunz/smogn.git

Road Map

  1. Distributed computing support
  2. Optimized distance metrics
  3. Explore interpolation methods

License

© Nick Kunz, 2019. Licensed under the General Public License v3.0 (GPLv3).

Contributions

SMOGN is open for improvements and maintenance. Your help is valued to make the package better for everyone.

Reference

Branco, P., Torgo, L., Ribeiro, R. (2017). SMOGN: A Pre-Processing Approach for Imbalanced Regression. Proceedings of Machine Learning Research, 74:36-50. http://proceedings.mlr.press/v74/branco17a/branco17a.pdf.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

smogn-0.0.1.tar.gz (26.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

smogn-0.0.1-py3-none-any.whl (29.1 kB view details)

Uploaded Python 3

File details

Details for the file smogn-0.0.1.tar.gz.

File metadata

  • Download URL: smogn-0.0.1.tar.gz
  • Upload date:
  • Size: 26.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/2.0.0 pkginfo/1.5.0.1 requests/2.21.0 setuptools/41.6.0 requests-toolbelt/0.9.1 tqdm/4.31.0 CPython/3.7.4

File hashes

Hashes for smogn-0.0.1.tar.gz
Algorithm Hash digest
SHA256 d92910b0ca883fd1653958f9e7f4247988a3928911ccc65e864f11a2b53befb8
MD5 a0ca7b5c568578f814a02036d4183735
BLAKE2b-256 d5cacfb6ee021326eaa46e491bcff95b4e0807417be481ee3488087a62991f69

See more details on using hashes here.

File details

Details for the file smogn-0.0.1-py3-none-any.whl.

File metadata

  • Download URL: smogn-0.0.1-py3-none-any.whl
  • Upload date:
  • Size: 29.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/2.0.0 pkginfo/1.5.0.1 requests/2.21.0 setuptools/41.6.0 requests-toolbelt/0.9.1 tqdm/4.31.0 CPython/3.7.4

File hashes

Hashes for smogn-0.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 504398e7a4588d0b8e73d719b12551df99623771f08c6cc8043b3ee242bed77a
MD5 e7efc9200e641b7d006ca0928a12399d
BLAKE2b-256 99e9d4d8e61d6ff50d05b1c06638b4f098c7e9528a7e3ad6e063496b8a271de6

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page