Skip to main content

a more in-depth testsplit splitting intercategorical

Project description

fancy schmancy testsplit

it's like a testsplit, but fancy and also schmancy


for reference:

package fancy schmancy testsplit
sklearn.model_selection 👎 👎 👍
fancy schmancy testsplit 👍 👍 👍

a testsplit per label category, to ensure that every category is present


Examples

Assume the following DataFrame:

df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
print(df)
Column A Column B
0 10 Cat1
1 14 Cat1
2 12 Cat2
3 13 Cat2
4 9 Cat2
5 5 Cat2
6 13 Cat2
7 16 Cat2
8 18 Cat2
9 4 Cat2
10 12 Cat2

If we assume further that Column B contains the label categories, we'd run the risk of eliminating Cat1 by doing a train test split at 50%.

So, to preserve every existing category, the split will instead be made on every single subset of categories.

As an example for Cat1:

subset = df[df["Column B"] == "Cat1"]
X = subset.drop("Column B", axis= 1)
y = subset["Column B"]
if isinstance(y, Series): y = DataFrame(y)
X_tr, X_te, y_tr, y_te = \
    train_test_split(X, y, test_size = 0.5, random_state = 42)
print(y_tr)
Column B
0 Cat1

This is done for every unique entry of the given label column, so that a random pick of train and test data is done for every category separately.

If this was done for "Cat1" and "Cat2", it would look like this:

Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

To shorten the process, the method fancy_schmancy_testsplit can be used in this way:

from FancySchmancyTestsplit.fst import fancy_schmancy_testsplit
from pandas import DataFrame
df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
X_train, X_test, y_train, y_test = \
    fancy_schmancy_testsplit(data= df,
                            label_column= "Column B",
                            test_split= 0.5,
                            seed= 42
                            )
print(y_train)
Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

FancySchmancyTestsplit-0.1.7.tar.gz (4.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

FancySchmancyTestsplit-0.1.7-py3-none-any.whl (5.1 kB view details)

Uploaded Python 3

File details

Details for the file FancySchmancyTestsplit-0.1.7.tar.gz.

File metadata

  • Download URL: FancySchmancyTestsplit-0.1.7.tar.gz
  • Upload date:
  • Size: 4.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.11.9

File hashes

Hashes for FancySchmancyTestsplit-0.1.7.tar.gz
Algorithm Hash digest
SHA256 6c0417ebe23efd32ad6f788321763c89f172f0bf411f0b6eb5905b3faa4b5c6e
MD5 5886d6dc182533c5d30ef6aa328af817
BLAKE2b-256 2196ff9c6f52c00f93b1d39df81434aef30dbea296f2adaf37e566f2aac62fc9

See more details on using hashes here.

File details

Details for the file FancySchmancyTestsplit-0.1.7-py3-none-any.whl.

File metadata

File hashes

Hashes for FancySchmancyTestsplit-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 f0c53676702585cc8fb60a72fdd07113dba8fbd12dc4ccfb88bddf31ce794bbc
MD5 29574703b27304cad043f4159dcc3d61
BLAKE2b-256 c3061ba3467a5d2f2c0199b68b06178dce2672f6dc8320fab9498ee55efe0de4

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page