Estimate the optimal number of components in PCA-based dimension reduction.
Project description
Estimate the optimal number of components in a PCA, using the SHEM procedure: Split-Half Eigenvector Matching.
The get_n_components function estimates the true (or "generating") number of principal components. While scree/elbow/knee criteria for the eigenvalues curve is common, this is known to be a very fallible heuristic. The rationale of this alternative procedure is that true principal components should be found in random split halves of the data. The estimate is therefore based on measuring the similarity of eigenvectors between a set of split-halves; i.e., the procedure doesn't use the shape of the eigenvalue curve. Instead, a separation is made between components with high versus low split-half similarity.
Detail: For an nd-array X, with shape == (nObservations, nVariables), a number of random splits are performed. For each split separately, a PCA is performed, via eigendecomposition of the covariance matrix of X. Each of the first split's eigenvectors is matched to the most-similar of the second split's eigenvectors. Similarity is measured via the dot product. The vector of similarities is sorted from high to low, and the vectors are averaged over all random splits. Finally, the optimal seperation between the high versus low similarities is determined by a basic between-within variance criterion. An estimated zero components is possible.
Usage:
O = teg_get_best_n.get_n_components(X)
This returns a dictionary with the estimated number of components in O['nComponents'], as well as the eigenvalues (O['eigenvalues']) and eigenvectors (O['eigenvectors']).
Example.py contains tests with simulated data to check how well the true number of latent variables, used to generate simulated data, is recovered.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
File details
Details for the file teg_get_best_n-0.0.7.tar.gz
.
File metadata
- Download URL: teg_get_best_n-0.0.7.tar.gz
- Upload date:
- Size: 3.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.2 CPython/3.9.5
File hashes
Algorithm | Hash digest | |
---|---|---|
SHA256 | 18db35d9adcf6212211b60fdd05f6843dff9921ba04afa9af2df2d3cfcfecfa5 |
|
MD5 | ddadae3c55727b010c2fe66cd43fb72a |
|
BLAKE2b-256 | b969590bc381dfcfebd10fcf85a3a444636213ce388862de54d4ec69476aebc2 |
File details
Details for the file teg_get_best_n-0.0.7-py3-none-any.whl
.
File metadata
- Download URL: teg_get_best_n-0.0.7-py3-none-any.whl
- Upload date:
- Size: 4.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.2 CPython/3.9.5
File hashes
Algorithm | Hash digest | |
---|---|---|
SHA256 | 3168b50610c7bbb98627c1ed478f071bdfd4920d20466a3c5ccbbe5d5536ce41 |
|
MD5 | 2e28a4047ccceb340e5b566b2137be0f |
|
BLAKE2b-256 | ec26a9b2f44a5326da7a9cf23b18de02eb870cfd66e27150ef35383f12877983 |