a more in-depth testsplit splitting intercategorical
Project description
fancy schmancy testsplit
it's like a testsplit, but fancy and also schmancy
for reference:
| package | fancy | schmancy | testsplit |
|---|---|---|---|
| sklearn.model_selection | 👎 | 👎 | 👍 |
| fancy schmancy testsplit | 👍 | 👍 | 👍 |
a testsplit per label category, to ensure that every category is present
Examples
Assume the following DataFrame:
df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
print(df)
| Column A | Column B | |
|---|---|---|
| 0 | 10 | Cat1 |
| 1 | 14 | Cat1 |
| 2 | 12 | Cat2 |
| 3 | 13 | Cat2 |
| 4 | 9 | Cat2 |
| 5 | 5 | Cat2 |
| 6 | 13 | Cat2 |
| 7 | 16 | Cat2 |
| 8 | 18 | Cat2 |
| 9 | 4 | Cat2 |
| 10 | 12 | Cat2 |
If we assume further that Column B contains the label categories, we'd run the risk of eliminating Cat1 by doing a train test split at 50%.
So, to preserve every existing category, the split will instead be made on every single subset of categories.
As an example for Cat1:
subset = df[df["Column B"] == "Cat1"]
X = subset.drop("Column B", axis= 1)
y = subset["Column B"]
if isinstance(y, Series): y = DataFrame(y)
X_tr, X_te, y_tr, y_te = \
train_test_split(X, y, test_size = 0.5, random_state = 42)
print(y_tr)
| Column B | |
|---|---|
| 0 | Cat1 |
This is done for every unique entry of the given label column, so that a random pick of train and test data is done for every category separately.
If this was done for "Cat1" and "Cat2", it would look like this:
| Column B | |
|---|---|
| 0 | Cat1 |
| 4 | Cat2 |
| 6 | Cat2 |
| 5 | Cat2 |
| 8 | Cat2 |
To shorten the process, the method fancy_schmancy_testsplit can be used in this way:
from FancySchmancyTestsplit.fst import fancy_schmancy_testsplit
from pandas import DataFrame
df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
X_train, X_test, y_train, y_test = \
fancy_schmancy_testsplit(data= df,
label_column= "Column B",
test_split= 0.5,
seed= 42
)
print(y_train)
| Column B | |
|---|---|
| 0 | Cat1 |
| 4 | Cat2 |
| 6 | Cat2 |
| 5 | Cat2 |
| 8 | Cat2 |
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file FancySchmancyTestsplit-0.1.7.tar.gz.
File metadata
- Download URL: FancySchmancyTestsplit-0.1.7.tar.gz
- Upload date:
- Size: 4.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/5.0.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6c0417ebe23efd32ad6f788321763c89f172f0bf411f0b6eb5905b3faa4b5c6e
|
|
| MD5 |
5886d6dc182533c5d30ef6aa328af817
|
|
| BLAKE2b-256 |
2196ff9c6f52c00f93b1d39df81434aef30dbea296f2adaf37e566f2aac62fc9
|
File details
Details for the file FancySchmancyTestsplit-0.1.7-py3-none-any.whl.
File metadata
- Download URL: FancySchmancyTestsplit-0.1.7-py3-none-any.whl
- Upload date:
- Size: 5.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/5.0.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f0c53676702585cc8fb60a72fdd07113dba8fbd12dc4ccfb88bddf31ce794bbc
|
|
| MD5 |
29574703b27304cad043f4159dcc3d61
|
|
| BLAKE2b-256 |
c3061ba3467a5d2f2c0199b68b06178dce2672f6dc8320fab9498ee55efe0de4
|