Skip to main content

Package providing methods to create Vector Embeddings from Strings, calculate similarities between lists of Strings, and Generate Visualizations such as Heatmaps from simple Lists.

Project description

vembed

Library to generate Embeddings and Extract Semantic Similarity from Data.


String to Embeddings

  • Convert Strings to Vector Embeddings
@Usage

from embedder import string_to_embedding

input_string = "This is a test sentence."
embedding = string_to_embedding(input_string)
print(embedding)

Similarity

Extracting Similarity Between Entities

Negative - Low Similarity
Zero     - Orthogonal - no commonality
Positive - Strong Similarity 

Cosine Similarity

  • Ranges between -1 and 1

  • Recommended when the Context and Similarity is important - and Frequency is not important (Magnitude)

  • Use Case for Cosine Similarity

    • Here, Direction - thematic orienation (climate change, agriculture) is relevant

    • Cosine Similarity is useful here as we want to find the relevancy of documents discussing similar topics (direction) - irrespective of the length of frequency of specific words (Magnitude)

@Usage

queries = ["Climate change effects on agriculture"]
data = [
    "Effects of climate change on wheat production",
    "Agriculture in developing countries",
    "Climate change and its impact on global food security",
    "Advances in agricultural technology"
]

# Calculate cosine similarities
cos_df, _ = calculate_similarities(queries, data, sorted=True, print_results=True)

Dot Product Similarity

  • Ranges between any Real Number

  • When both the magnitude and direction of the vectors are important, and you are dealing with vectors in a similar scale.

  • When the Frequency (Magnitude) as well as the Direction (Relevancy) is both important.

  • Use Case for Dot Product

    • Direction (Types of Articles) and Magnitude (Frequency of Reading Habits) are both important.
@Usage

user_reading_profile = ["Read many articles on machine learning", "Occasionally reads about space exploration"]
article_options = [
    "Latest trends in machine learning",
    "Beginner's guide to space travel",
    "In-depth analysis of neural networks",
    "Recent discoveries in astronomy"
]

# Calculate dot product similarities
_, dot_df = calculate_similarities(user_reading_profile, article_options, sorted=True, print_results=True)
  • Calculating Similarity
@Usage

queries = ["What is the capital of France?", "How is the weather today?"]
    data = [
        "Paris is the capital of France.",
        "The weather is sunny.",
        "Berlin is the capital of Germany.",
        "It is raining in Berlin.",
    ]

    # Calculate similarities and Print Results
    cos_df, dot_df = calculate_similarities(
        queries, data, sorted=True, print_results=True
    )

Visualization for Relevance


  • Create a visualization to display the Simalirities using a Heatmap.
@Usage

customer_feedback = [
    "Loved the recent update",
    "The app is user-friendly",
    "Facing issues after the update",
    "The new interface is great",
]
themes = [
    "positive feedback",
    "negative feedback",
    "app interface",
    "app functionality",
]

# Heatmap of Both Cosine and Dot Product
cos_df, dot_df = calculate_similarities(customer_feedback, themes, sorted=True)
plot_similarities(cos_df, dot_df, save_path="customer_feedback_similarity.png")

# Heatmap of Only Cosine Similarity
cos_df, _ = calculate_similarities(customer_feedback, themes, sorted=True)
plot_similarities(cos_df, None, save_path="customer_feedback_similarity.png")

# Heatmap of Only Dot Product Similarity
_, dot_df = calculate_similarities(customer_feedback, themes, sorted=True)
plot_similarities(None, dot_df, save_path="customer_feedback_similarity.png")

# View customer_feedback_similarity.png to see the Heatmap

Dependencies

  • sentence_transformers
  • torch
  • transformers
  • pandas
  • matplotlib
  • seaborn

Note: This package uses Nvidia Cuda and Torch.

# Check Disk Allocation for Packages 
du -h venv | sort -hr | head -n 10

2.8G    venv/lib/python3.11/site-packages/nvidia
1.4G    venv/lib/python3.11/site-packages/torch
1.3G    venv/lib/python3.11/site-packages/torch/lib
1.2G    venv/lib/python3.11/site-packages/nvidia/cudnn/lib
1.2G    venv/lib/python3.11/site-packages/nvidia/cudnn
596M    venv/lib/python3.11/site-packages/nvidia/cublas

# Checking System Cache

# Show pip cache location
pip cache dir # /home/user/.cache/pip

# Getting Top Folders from Cache by Size
du -h /home/user/.cache/pip | sort -hr | head -n 10

# Remove Cached Files
pip cache purge 

# Cached Files
pip cache list

# Installing Packages without Cache
pip install --no-cache-dir <package_name>

Author: kuro337

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vembed-0.22.tar.gz (9.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vembed-0.22-py3-none-any.whl (7.4 kB view details)

Uploaded Python 3

File details

Details for the file vembed-0.22.tar.gz.

File metadata

  • Download URL: vembed-0.22.tar.gz
  • Upload date:
  • Size: 9.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.2 CPython/3.12.0

File hashes

Hashes for vembed-0.22.tar.gz
Algorithm Hash digest
SHA256 b46c18e5286aa5374908a60ef44f573679b4c5e46b51f8bd870d9d0fa0125a84
MD5 aa1896978923563da34b84e2defba272
BLAKE2b-256 91d5bf48fc6f291b35deb1642ded198846d126d1b433a7f81e87bd4340f10376

See more details on using hashes here.

File details

Details for the file vembed-0.22-py3-none-any.whl.

File metadata

  • Download URL: vembed-0.22-py3-none-any.whl
  • Upload date:
  • Size: 7.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.2 CPython/3.12.0

File hashes

Hashes for vembed-0.22-py3-none-any.whl
Algorithm Hash digest
SHA256 bbdc7920f385880b25a2af510c738d23144302fc19a2f71dd129c209af79398e
MD5 060b63f0fd5a09c120e7cff524a12966
BLAKE2b-256 07916b1f5d9ab0ed1f7e6f74bd68d4ab641cf34b35eecbe5743abb028b47c49c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page