Skip to main content

PyPI - Version Unittests PyPI - License

Pytolemaic

What is Pytolemaic

Pytolemaic package analyzes your model and dataset and measure their quality.

The package supports classification/regression models built for tabular datasets (e.g. sklearn's regressors/classifiers), but will also support custom made models as long as they implement sklearn's API.

The package is aimed for personal use and comes with no guarantees. I hope you will find it useful. I will appreciate any feedback you have.

Install

pip install pytolemaic

Basic usage

from pytolemaic import PyTrust

pytrust = PyTrust(model=estimator,
                  xtrain=xtrain, ytrain=ytrain,
                  xtest=xtest, ytest=ytest)
   
# run all analysis and print insights:,
insights = pytrust.insights()
print("\n".join(insights))

# run analysis and plot graphs
pytrust.plot()

supported features

The package contains the following functionalities:

On model creation

  • Dataset Analysis: Analysis aimed to detect issues in the dataset.
  • Sensitivity Analysis: Calculation of feature importance for given model, either via sensitivity to feature value or sensitivity to missing values.
  • Vulnerability report: Based on the feature sensitivity we measure model's vulnerability in respect to imputation, leakage, and # of features.
  • Scoring report: Report model's score on test data with confidence interval.
  • separation quality: Measure whether train and test data comes from the same distribution.
  • Overall quality: Provides overall quality measures

On prediction

  • Prediction uncertainty: Provides an uncertainty measure for given model's prediction.
  • Lime explanation: Provides Lime explanation for sample of interest.

How to use:

Get started by calling help() function (Recommended!):

   from pytolemaic import help
   supported_keys = help()
   # or
   help(key='basic usage')

Example for performing all available analysis with PyTrust:

   from pytolemaic import PyTrust

   pytrust = PyTrust(
       model=estimator,
       xtrain=xtrain, ytrain=ytrain,
       xtest=xtest, ytest=ytest)
       
   # run all analysis and get a list of distilled insights",
   insights = pytrust.insights()
   print("\n".join(insights))
    
   # run all analysis and plot all graphs
   pytrust.plot()
   
   # print all data gathered
   import pprint
   pprint(report.to_dict(printable=True))

In case of need to access only specific analysis (usually to save time)

   # dataset analysis report
   dataset_analysis_report = pytrust.dataset_analysis_report
   
   # feature sensitivity report
   sensitivity_report = pytrust.sensitivity_report
   
   # model's performance report
   scoring_report = pytrust.scoring_report
   
   # overall model's quality report
   quality_report = pytrust.quality_report
   
   # with any of the above reports
   report = <desired report>
   print("\n".join(report.insights()))
   
   report.plot() # plot graphs
   pprint(report.to_dict(printable=True)) # export report as a dictionary
   pprint(report.to_dict_meaning()) # print documentation for above dictionary
          

Analysis of predictions

 
   # estimate uncertainty of a prediction
   uncertainty_model = pytrust.create_uncertainty_model()
   
   # explain a prediction with Lime
   create_lime_explainer = pytrust.create_lime_explainer()
   

Examples on toy dataset can be found in /examples/toy_examples/ Examples on 'real-life' datasets can be found in /examples/interesting_examples/

Output examples:

Sensitivity Analysis:

  • The sensitivity of each feature ([0,1], normalized to sum of 1):
 'sensitivity_report': {
    'method': 'shuffled',
    'sensitivities': {
        'age': 0.12395,
        'capital-gain': 0.06725,
        'capital-loss': 0.02465,
        'education': 0.05769,
        'education-num': 0.13765,
        ...
      }
  }
  • Simple statistics on the feature sensitivity:
'shuffle_stats_report': {
     'n_features': 14,
     'n_low': 1,
     'n_zero': 0
}
  • Naive vulnerability scores ([0,1], lower is better):

    • Imputation: sensitivity of the model to missing values.
    • Leakge: chance of the model to have leaking features.
    • Too many features: Whether the model is based on too many features.
'vulnerability_report': {
     'imputation': 0.35,
     'leakage': 0,
     'too_many_features': 0.14
}  

scoring report

For given metric, the score and confidence intervals (CI) is calculated

'recall': {
    'ci_high': 0.763,
    'ci_low': 0.758,
    'ci_ratio': 0.023,
    'metric': 'recall',
    'value': 0.760,
},
'auc': {
    'ci_high': 0.909,
    'ci_low': 0.907,
    'ci_ratio': 0.022,
    'metric': 'auc',
    'value': 0.907
}    

Additionally, score quality measures the quality of the score based on the separability (auc score) between train and test sets.

Value of 1 means test set has same distribution as train set. Value of 0 means test set has fundamentally different distribution.

'separation_quality': 0.00611         

Combining the above measures into a single number we provide the overall quality of the model/dataset.

Higher quality value ([0,1]) means better dataset/model.

quality_report : { 
'model_quality_report': {
   'model_loss': 0.24,
   'model_quality': 0.41,
   'vulnerability_report': {...}},
   
'test_quality_report': {
   'ci_ratio': 0.023, 
   'separation_quality': 0.006, 
   'test_set_quality': 0},
   
'train_quality_report': {
   'train_set_quality': 0.85,
   'vulnerability_report': {...}}
  

prediction uncertainty

The module can be used to yield uncertainty measure for predictions.

    uncertainty_model = pytrust.create_uncertainty_model(method='confidence')
    predictions = uncertainty_model.predict(x_pred) # same as model.predict(x_pred)
    uncertainty = uncertainty_model.uncertainty(x_pred)

Lime explanation

The module can be used to produce lime explanations for sample of interest.

    explainer = pytrust.create_lime_explainer()
    explainer.explain(sample) # returns a dictionary
    explainer.plot(sample) # produce a graphical explanation    

Release files for pytolemaic 0.15.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pytolemaic 0.15.4
File Size Uploaded
pytolemaic-0.15.4.tar.gz 82.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pytolemaic 0.15.4
File Interpreter ABI Platform
pytolemaic-0.15.4-py3-none-any.whl Python 3 none any Details

Total release size: 195.8 kB

Release files / pytolemaic-0.15.4.tar.gz

Download URL pytolemaic-0.15.4.tar.gz
Size 82.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c43ba753e0bcc46f36f58ddef1df3fe1d4f4c6edf709cf47773407251c92cc3b
BLAKE2b-256 checksum
How to use checksums
c17c697a8443642d286a28b81d4b0790a22f249faac224927ae533c92caf9ae3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.10.5

Release files / pytolemaic-0.15.4-py3-none-any.whl

Download URL pytolemaic-0.15.4-py3-none-any.whl
Size 113.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b3347291524fe88bb3c41b65366f37d1721938df593e7980b33c5967dda05228
BLAKE2b-256 checksum
How to use checksums
7fd920c96c66c3e70de1a3281606347c129e254c0bc8146d5d8b66c147f0a41b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.10.5

Release history Release notifications | RSS feed

This release

0.15.4 This release

2 release files

0.14.1

2 release files

0.14.0

2 release files

0.13.9

2 release files

0.13.8

2 release files

0.13.6

2 release files

0.13.5

2 release files

0.13.4

2 release files

0.13.2

2 release files

0.12.2

2 release files

0.12.1

2 release files

0.11.8

2 release files

0.11.7

2 release files

0.11.4

2 release files

0.11.2

2 release files

0.11.0

2 release files

0.10.4

2 release files

0.10.3

2 release files

0.10.0

2 release files

0.9.4

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.11

2 release files

0.8.10

2 release files

0.8.9

2 release files

0.8.8

2 release files

0.8.7

2 release files

0.8.6

2 release files

0.8.5

2 release files

0.8.2

2 release files

0.8

2 release files

0.7

2 release files

0.5

2 release files

0.4

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page