Skip to main content

a more in-depth testsplit splitting intercategorical

Project description

fancy schmancy testsplit

it's like a testsplit, but fancy and also schmancy


for reference:

package fancy schmancy testsplit
sklearn.model_selection 👎 👎 👍
fancy schmancy testsplit 👍 👍 👍

a testsplit per label category, to ensure that every category is present


Examples

Assume the following DataFrame:

df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
print(df)
Column A Column B
0 10 Cat1
1 14 Cat1
2 12 Cat2
3 13 Cat2
4 9 Cat2
5 5 Cat2
6 13 Cat2
7 16 Cat2
8 18 Cat2
9 4 Cat2
10 12 Cat2

If we assume further that Column B contains the label categories, we'd run the risk of eliminating Cat1 by doing a train test split at 50%.

So, to preserve every existing category, the split will instead be made on every single subset of categories.

As an example for Cat1:

subset = df[df["Column B"] == "Cat1"]
X = subset.drop("Column B", axis= 1)
y = subset["Column B"]
if isinstance(y, Series): y = DataFrame(y)
X_tr, X_te, y_tr, y_te = \
    train_test_split(X, y, test_size = 0.5, random_state = 42)
print(y_tr)
Column B
0 Cat1

This is done for every unique entry of the given label column, so that a random pick of train and test data is done for every category separately.

If this was done for "Cat1" and "Cat2", it would look like this:

Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

To shorten the process, the method fancy_schmancy_testsplit can be used in this way:

from FancySchmancyTestsplit.fst import fancy_schmancy_testsplit
from pandas import DataFrame
df = DataFrame(data= {"Column A":[10, 14, 12, 13, 9, 5, 13, 16, 18, 4, 12],
"Column B": ["Cat1", "Cat1", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2", "Cat2"]})
X_train, X_test, y_train, y_test = \
    fancy_schmancy_testsplit(data= df,
                            label_column= "Column B",
                            test_split= 0.5,
                            seed= 42
                            )
print(y_train)
Column B
0 Cat1
4 Cat2
6 Cat2
5 Cat2
8 Cat2

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

FancySchmancyTestsplit-0.1.6.tar.gz (4.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

FancySchmancyTestsplit-0.1.6-py3-none-any.whl (5.0 kB view details)

Uploaded Python 3

File details

Details for the file FancySchmancyTestsplit-0.1.6.tar.gz.

File metadata

  • Download URL: FancySchmancyTestsplit-0.1.6.tar.gz
  • Upload date:
  • Size: 4.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.11.9

File hashes

Hashes for FancySchmancyTestsplit-0.1.6.tar.gz
Algorithm Hash digest
SHA256 4b83b13e18dc0bd443e216815395bff3c53c5285ffbd9c36eb5acf74bc055d6c
MD5 6fc86a3ee434c15bf498159fdf8b7fc6
BLAKE2b-256 3891aadd81bb6ebc8abb4b7601cdd42f80789f1c15aabdccbae7cb4b039fe1e3

See more details on using hashes here.

File details

Details for the file FancySchmancyTestsplit-0.1.6-py3-none-any.whl.

File metadata

File hashes

Hashes for FancySchmancyTestsplit-0.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 08175602888eea8f527dc8ed3f6675c6930d6930ee126d1e2554fbbf58af2e23
MD5 9d2faa5cc38a9bc23edff3c0a5110e8b
BLAKE2b-256 6070c6c07cb250294130d4d3ec718efd77eb527bd52b3b3b26fdc9b200774f9c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page