Skip to main content

A library providing frequently used functions in data analysis for YESlab members and other researchers

Project description

yescarpenter

This library provides some frequent used functions for YESlab members and other researchers, including data processing and analyses

Installation

From PyPI

Users can install it using pip:

pip install yescarpenter

Functions

perform_pca

This function leverages scikit-learn to perform a tailored PCA analysis (e.g., with rotation to maximize variance)

Usage:

import pandas as pd
from yescarpenter import perform_pca

# Create a sample DataFrame
data = pd.DataFrame({
    'feature1': [1, 2, 3, 4],
    'feature2': [2, 3, 4, 5],
    'feature3': [3, 4, 5, 6]
})

# Perform PCA with 2 components
loadings, explained_variance, components = perform_pca(data, n_components=2)

print("Loadings:\n", loadings)
print("Explained Variance:\n", explained_variance)
print("Components:\n", components)

create_scree_plot

This function creates a scree plot to visualize the explained variance of each principal component.

scree_plot(explained_variance, n_components)

pc_plot

Create a plot to visualize the PCA loadings.

pc_plot(loadings, df)

construct_RDM

For IS-RSA. Construct the Representational Dissimilarity Matrix(RDM) from the data.

Usage:

construct_RDM(data, n_target, method = "cityblock")

Input:

  • data: n x m matrix, where n is the number of target and m is the number of features
  • n_target: the number of target
  • method: the method to calculate the distance matrix
    • euclidean: Euclidean distance
    • cityblock: Manhattan distance
    • spearman: Spearman correlation

Usage:

# Example usage
import numpy as np
from yescarpenter import construct_RDM
data = np.random.rand(10, 5)  # 10 pictures, 5 ratings
n_target = 10
rdm = construct_RDM(data, n_target, method = "euclidean")

do_rsa

Calculate the Spearman correlation between two RDMs(upper triangle) and perform Mantel permutations.

Parameters:
-----------
matrix1 : np.ndarray
    First distance matrix (square, symmetric).
matrix2 : np.ndarray
    Second distance matrix (square, symmetric, same size as matrix1).
n_permutations : int
    Number of permutations.
random_state : int or None
    Random seed for reproducibility.

Returns:
--------
permuted_correlations : np.ndarray
    Array of permuted Spearman correlation values.
observed_correlation : float
    Observed Spearman correlation between original matrices.
p_value : float
    P-value representing significance of the observed correlation.

Usage:

do_RSA(rdm1, rdm2, n_permutations=1000, random_state=None)

permutation_histogram

Plot the histogram of null distribution, with the observed value and p-value marked.

Usage:

permutation_histogram(r, perm_r)
  • r: The observed value.
  • perm_r: The null distribution, which consists of the iterated surrogated values

maximal_permutation_test

This fuction is used to address multiple comparison, which provides an alternative of Bonferroni correction.

Usage:

[perm_r, perm_p, observed_r] = maximal_permutation_test(data, iv_single, iv_multiplecomp, nperm)
  • data: For IS-RSA, each row is a subject, while each column is a variable.
    For example, if you have 20 subjects and 5 variables, the shape of data is (20, 5).
    • iv_single: the independent variable that will be shuffled and compare across iv_multiplecomp
    • iv_multiplecomp: the independent variable that are inter-related and elicit the multiple comparison problem
    • n_perm: number of permutation

align_data

This function is used to align different sources of data, such as visual and semantic embeddings, and behavioral rating data.

Usage:

aligned_cong_fec, aligned_vgg, aligned_sem = align_data(
    {'data': cong_fec['Vote share percentage'].values, 'order': cong_fec['Image_name'].values},
    {'data': response, 'order': vgglist['image_name'].values},
    {'data': sememb, 'order': sem_imgname}
)

print(aligned_cong_fec.shape, aligned_vgg.shape, aligned_sem.shape)
  • You can put as many data as you want. Each input has to be a dictionary with two keys: 'data' and 'order'.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

yescarpenter-0.3.0.tar.gz (14.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

yescarpenter-0.3.0-py3-none-any.whl (14.6 kB view details)

Uploaded Python 3

File details

Details for the file yescarpenter-0.3.0.tar.gz.

File metadata

  • Download URL: yescarpenter-0.3.0.tar.gz
  • Upload date:
  • Size: 14.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.2

File hashes

Hashes for yescarpenter-0.3.0.tar.gz
Algorithm Hash digest
SHA256 fdbecc0966e89b8eb8f1d25e2652f27343a8ca6908c95dbf7e5493d30bf3a0d3
MD5 d3e247ff95216f8f3e6e0d156880ee10
BLAKE2b-256 83fd0aa5ddf85a38df66a802d1739ac8544e5cc6ef86acffa3bb5b536abf1674

See more details on using hashes here.

File details

Details for the file yescarpenter-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: yescarpenter-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 14.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.2

File hashes

Hashes for yescarpenter-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 eea3f60b0233c2cb727f0b8726125bebfb56e01031a41b50d237e284c46bf66c
MD5 40610705fc9ade9d99f38cb84ff63990
BLAKE2b-256 b390e84275e85b22fce746739700a8efa9acd07d901bff737584c7c307ece430

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page