Skip to main content

a more in-depth testsplit splitting intercategorical

Project description

fancy schmancy testsplit

it's like a testsplit, but fancy and also schmancy


for reference:

package fancy schmancy testsplit
sklearn.model_selection 👎 👎 👍
fancy schmancy testsplit 👍 👍 👍

a testsplit per label category, to ensure that every category is present


Examples

Assume the following DataFrame:

df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
print(df)
Column A Column B
0 10 Cat1
1 14 Cat1
2 12 Cat2
3 13 Cat2
4 9 Cat2
5 5 Cat2
6 13 Cat2
7 16 Cat2
8 18 Cat2
9 4 Cat2
10 12 Cat2

If we assume further that Column B contains the label categories, we'd run the risk of eliminating Cat1 by doing a train test split at 50%.

So, to preserve every existing category, the split will instead be made on every single subset of categories.

As an example for Cat1:

subset = df[df["Column B"] == "Cat1"]
X = subset.drop("Column B", axis= 1)
y = subset["Column B"]
if isinstance(y, Series): y = DataFrame(y)
X_tr, X_te, y_tr, y_te = \
    train_test_split(X, y, test_size = 0.5, random_state = 42)
print(y_tr)
Column B
0 Cat1

This is done for every unique entry of the given label column, so that a random pick of train and test data is done for every category separately.

If this was done for "Cat1" and "Cat2", it would look like this:

Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

To shorten the process, the method fancy_schmancy_testsplit can be used in this way:

from FancySchmancyTestsplit.fst import fancy_schmancy_testsplit
from pandas import DataFrame
df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
X_train, X_test, y_train, y_test = \
    fancy_schmancy_testsplit(data= df,
                            label_column= "Column B",
                            test_split= 0.5,
                            seed= 42
                            )
print(y_train)
Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

FancySchmancyTestsplit-0.1.9.tar.gz (4.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

FancySchmancyTestsplit-0.1.9-py3-none-any.whl (4.0 kB view details)

Uploaded Python 3

File details

Details for the file FancySchmancyTestsplit-0.1.9.tar.gz.

File metadata

  • Download URL: FancySchmancyTestsplit-0.1.9.tar.gz
  • Upload date:
  • Size: 4.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.11.9

File hashes

Hashes for FancySchmancyTestsplit-0.1.9.tar.gz
Algorithm Hash digest
SHA256 08c8545f1ba761b4b03c2cec214a29640a18785aa4c5a12fc4b8e106ac83fcf7
MD5 0d71c9e99d4429bd2a263861f48f2739
BLAKE2b-256 f07a15b7b4d05668e8ade4d0d8d6f2d444e47fb4c283b8d0978808b822c7ce39

See more details on using hashes here.

File details

Details for the file FancySchmancyTestsplit-0.1.9-py3-none-any.whl.

File metadata

File hashes

Hashes for FancySchmancyTestsplit-0.1.9-py3-none-any.whl
Algorithm Hash digest
SHA256 ca87e9b6835778affcd7d807463b712656094cce5a4549aa8d4d5e7a509d932e
MD5 5820e8d16cc3d581b6d0fdbe1485f293
BLAKE2b-256 2d714621d9fe446960280cc91ff49d9418421067dbb103a2acf0158e5a846488

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page