Skip to main content

a more in-depth testsplit splitting intercategorical

Project description

fancy schmancy testsplit

it's like a testsplit, but fancy and also schmancy


for reference:

package fancy schmancy testsplit
sklearn.model_selection 👎 👎 👍
fancy schmancy testsplit 👍 👍 👍

a testsplit per label category, to ensure that every category is present


Examples

Assume the following DataFrame:

df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
print(df)
Column A Column B
0 10 Cat1
1 14 Cat1
2 12 Cat2
3 13 Cat2
4 9 Cat2
5 5 Cat2
6 13 Cat2
7 16 Cat2
8 18 Cat2
9 4 Cat2
10 12 Cat2

If we assume further that Column B contains the label categories, we'd run the risk of eliminating Cat1 by doing a train test split at 50%.

So, to preserve every existing category, the split will instead be made on every single subset of categories.

As an example for Cat1:

subset = df[df["Column B"] == "Cat1"]
X = subset.drop("Column B", axis= 1)
y = subset["Column B"]
if isinstance(y, Series): y = DataFrame(y)
X_tr, X_te, y_tr, y_te = \
    train_test_split(X, y, test_size = 0.5, random_state = 42)
print(y_tr)
Column B
0 Cat1

This is done for every unique entry of the given label column, so that a random pick of train and test data is done for every category separately.

If this was done for "Cat1" and "Cat2", it would look like this:

Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

To shorten the process, the method fancy_schmancy_testsplit can be used in this way:

from FancySchmancyTestsplit.fst import fancy_schmancy_testsplit
from pandas import DataFrame
df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
X_train, X_test, y_train, y_test = \
    fancy_schmancy_testsplit(data= df,
                            label_column= "Column B",
                            test_split= 0.5,
                            seed= 42
                            )
print(y_train)
Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

FancySchmancyTestsplit-0.1.5.tar.gz (4.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

FancySchmancyTestsplit-0.1.5-py3-none-any.whl (4.0 kB view details)

Uploaded Python 3

File details

Details for the file FancySchmancyTestsplit-0.1.5.tar.gz.

File metadata

  • Download URL: FancySchmancyTestsplit-0.1.5.tar.gz
  • Upload date:
  • Size: 4.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.11.9

File hashes

Hashes for FancySchmancyTestsplit-0.1.5.tar.gz
Algorithm Hash digest
SHA256 e0512789ecc15ff86e1b2a0767c040804300ba8ce9210790a597602a48f2fbf9
MD5 4e0221a957f91a29d8db2ce8ef6fd4c4
BLAKE2b-256 ce7987ecd32a6ac4a7028aded41695112b5901030bd622e6fce9519c2331643a

See more details on using hashes here.

File details

Details for the file FancySchmancyTestsplit-0.1.5-py3-none-any.whl.

File metadata

File hashes

Hashes for FancySchmancyTestsplit-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 5eb43110ba1ed0f8454cacd03fba6d7af1dbb64952b7702606e18453d3b3b393
MD5 f40ff06fd62d33ad2f585130ce2f5972
BLAKE2b-256 a0a4e14fcca8b279ce751dd5a22935d68bb394fff2f41d3a04d317a183439b74

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page