Categorical Embedder
Categorical Embedder is a python package that let's you convert your categorical variables into numeric via Neural Networks
Installation
pip install categorical_embedder
Example
import categorical_embedder as ce
from sklearn.model_selection import train_test_split
df = pd.read_csv('HR_Attrition_Data.csv')
X = df.drop(['employee_id', 'is_promoted'], axis=1)
y = df['is_promoted']
embedding_info = ce.get_embedding_info(X)
X_encoded,encoders = ce.get_label_encoded_data(X)
X_train, X_test, y_train, y_test = train_test_split(X_encoded,y)
embeddings = ce.get_embeddings(X_train, y_train, categorical_embedding_info=embedding_info,
is_classification=True, epochs=100,batch_size=256)
A more detailed Jupyter Notebook can be found here
What's inside Categorical Embedder ?
ce.get_embedding_info(data,categorical_variables=None): This function identifies all categorical variables in the data, determines its embedding size. Embedding size of the categorical variables are determined by minimum of 50 or half of the no. of its unique values i.e. embedding size of a column = Min(50, # unique values in that column) One can pass explicit list of categorical variables incategorical_variablesparameter. IfNone, this function automatically takes all the variables with data typeobjectce.get_label_encoded_data(data, categorical_variables=None): This function label encodes (integer encoding) all the categorical variables using sklearn.preprocessing.LabelEncoder and returns a label encoded dataframe for training. Keras/tensorflow or any other deep learning library would expect the data to be in this format.ce.get_embeddings(X_train, y_train, categorical_embedding_info=embedding_info, is_classification=True, epochs=100,batch_size=256): This function trains a shallow neural networks and returns embeddings of categorical variables. Under the hood, It is a 2 layer neural network architecture with 1000 and 500 neurons with 'ReLU' activation. It takes 4 required inputs -X_train,y_train,categorical_embedding_info:output of get_embedding_info function andis_classification:Truefor classification tasks;Falsefor regression tasks.
For classification: loss = 'binary_crossentropy'; metrics = 'accuracy' and for regression: loss = 'mean_squared_error'; metrics = 'r2'
Dependencies
pandas
scikit-learn
tensorflow
keras
tqdm
keras-tqdm
Release files for categorical-embedder 0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| categorical_embedder-0.1.tar.gz | 4.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| categorical_embedder-0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 10.3 kB
Release files / categorical_embedder-0.1.tar.gz
| Download URL | categorical_embedder-0.1.tar.gz |
|---|---|
| Size | 4.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cc0a152f6e7ff381ec06f5ab5a7cf27116571b51ac092c1112099a7d8b8e83d3
|
|
BLAKE2b-256 checksum How to use checksums |
66c949835ed4c83c0310b4d86bf866596f79de535a413e479130f1724efb9e92
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/41.4.0 requests-toolbelt/0.9.1 tqdm/4.36.1 CPython/3.7.4
|
Release files / categorical_embedder-0.1-py3-none-any.whl
| Download URL | categorical_embedder-0.1-py3-none-any.whl |
|---|---|
| Size | 5.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
fd87fee2b0484e0042825bd71c6b4e6820ad65e6c264a7a21ac431a3e1630ab7
|
|
BLAKE2b-256 checksum How to use checksums |
4455e114f63ad47253ac04b0db012b3efc9183762e23bc5c40187c040b8c99d9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/41.4.0 requests-toolbelt/0.9.1 tqdm/4.36.1 CPython/3.7.4
|