Skip to main content

a more in-depth testsplit splitting intercategorical

Project description

fancy schmancy testsplit

it's like a testsplit, but fancy and also schmancy


for reference:

package fancy schmancy testsplit
sklearn.model_selection 👎 👎 👍
fancy schmancy testsplit 👍 👍 👍

a testsplit per label category, to ensure that every category is present


Examples

Assume the following DataFrame:

df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
print(df)
Column A Column B
0 10 Cat1
1 14 Cat1
2 12 Cat2
3 13 Cat2
4 9 Cat2
5 5 Cat2
6 13 Cat2
7 16 Cat2
8 18 Cat2
9 4 Cat2
10 12 Cat2

If we assume further that Column B contains the label categories, we'd run the risk of eliminating Cat1 by doing a train test split at 50%.

So, to preserve every existing category, the split will instead be made on every single subset of categories.

As an example for Cat1:

subset = df[df["Column B"] == "Cat1"]
X = subset.drop("Column B", axis= 1)
y = subset["Column B"]
if isinstance(y, Series): y = DataFrame(y)
X_tr, X_te, y_tr, y_te = \
    train_test_split(X, y, test_size = 0.5, random_state = 42)
print(y_tr)
Column B
0 Cat1

This is done for every unique entry of the given label column, so that a random pick of train and test data is done for every category separately.

If this was done for "Cat1" and "Cat2", it would look like this:

Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

To shorten the process, the method fancy_schmancy_testsplit can be used in this way:

from FancySchmancyTestsplit.fst import fancy_schmancy_testsplit
from pandas import DataFrame
df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
X_train, X_test, y_train, y_test = \
    fancy_schmancy_testsplit(data= df,
                            label_column= "Column B",
                            test_split= 0.5,
                            seed= 42
                            )
print(y_train)
Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fancyschmancytestsplit-0.1.10.tar.gz (4.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

FancySchmancyTestsplit-0.1.10-py3-none-any.whl (4.0 kB view details)

Uploaded Python 3

File details

Details for the file fancyschmancytestsplit-0.1.10.tar.gz.

File metadata

  • Download URL: fancyschmancytestsplit-0.1.10.tar.gz
  • Upload date:
  • Size: 4.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.11.9

File hashes

Hashes for fancyschmancytestsplit-0.1.10.tar.gz
Algorithm Hash digest
SHA256 247e44bfa5d610cf7a567825dc6858068fee1a756a17d4c5cbc2904a61ab4f37
MD5 ce1fdee4f08fac567cae9d13f4c86e5b
BLAKE2b-256 c4370320d5afe376f1673a717a400b90195819429c386d251770bb7263f260b1

See more details on using hashes here.

File details

Details for the file FancySchmancyTestsplit-0.1.10-py3-none-any.whl.

File metadata

File hashes

Hashes for FancySchmancyTestsplit-0.1.10-py3-none-any.whl
Algorithm Hash digest
SHA256 87b8d2e0070f540844483d8d8411f0b5cc5d36ce0119f80287c40a89ef9b0cae
MD5 51996aca7a9e5a288406a9656dbdf1b7
BLAKE2b-256 b4739ddba650e277f3e57d2f45a5bea367704491b9d7c5b1cd8f89dad4335f25

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page