Skip to main content

pyprik

Project description

pyprik

Overview

This library provides a set of powerful functions designed to assist in data analysis, data preprocessing, and the development of recommendation engines. It offers tools for flexible matching of data entries based on character similarity, making it ideal for scenarios where exact matches may not be available. The library is particularly useful in recommendation systems where close matches need to be identified from a large dataset.

find_matching(dataset, requirements)

Purpose

The find_matching function is designed to compare a dataset against specified requirements and identify matching entries. It helps you find rows in your dataset that best match the criteria provided.

Parameters

  • dataset (DataFrame): The dataset to be checked.
  • requirements (dict): A dictionary where keys are column names and values are the required values for matching.

Returns

  • DataFrame: A new DataFrame with additional columns indicating whether each row matches the specified requirements.

extract_characters(s)

Purpose

The extract_characters function is used to extract alphanumeric characters from a given string. It converts the string to lowercase and counts the occurrence of each alphanumeric character.

Parameters

  • s (str): The input string from which to extract characters.

Returns

  • Counter: A Counter object containing the counts of each alphanumeric character in the string.

character_match_score(user_input, feature_value)

Purpose

The character_match_score function calculates the similarity between two strings based on their alphanumeric characters. It compares the characters in the user input and a feature value, then computes a match score based on the intersection of character counts.

Parameters

  • user_input (str): The user input string to be compared.
  • feature_value (str): The string from the dataset to compare against.

Returns

  • int: The match score indicating the number of matching characters between the two strings.

find_closest_match(user_input, feature_values)

Purpose

The find_closest_match function identifies the closest matching value from a list of feature values based on character similarity. It is useful when an exact match is not found, allowing the user to find the next best option.

Parameters

  • user_input (str): The user input string to be matched.
  • feature_values (list): A list of feature values to compare against.

Returns

  • str: The feature value with the highest character match score.

find_top_matching(dataset, requirements, top_n, g=None)

Purpose

The find_top_matching function ranks entries in a dataset based on how well they match specified requirements. It calculates a total match score for each entry and returns the top N matching entries. This function can also return a specific column along with the match score if specified.

Parameters

  • dataset (DataFrame): The dataset containing the data to be matched.
  • requirements (dict): A dictionary where keys are column names and values are the required values for matching.
  • top_n (int): The number of top matching entries to return.
  • g (str, optional): A specific column to return along with the match score.

Returns

  • DataFrame: A DataFrame containing the top N matching entries. If g is provided, it returns that column along with the match score; otherwise, it returns the entire matching dataset.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pyprik-0.0.1.tar.gz (4.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pyprik-0.0.1-py3-none-any.whl (5.3 kB view details)

Uploaded Python 3

File details

Details for the file pyprik-0.0.1.tar.gz.

File metadata

  • Download URL: pyprik-0.0.1.tar.gz
  • Upload date:
  • Size: 4.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.12.1

File hashes

Hashes for pyprik-0.0.1.tar.gz
Algorithm Hash digest
SHA256 86fcb6d937f0cb0b5ae57d7c797de1c38ca369b405fdbb311044df689333e69b
MD5 999c725ca17a43ab90d113025ba97d11
BLAKE2b-256 92d4ca239d4ede2e4d9e6f38cfc76fc94c68424d2d08c339e0bbf9ddf210f2d8

See more details on using hashes here.

File details

Details for the file pyprik-0.0.1-py3-none-any.whl.

File metadata

  • Download URL: pyprik-0.0.1-py3-none-any.whl
  • Upload date:
  • Size: 5.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.12.1

File hashes

Hashes for pyprik-0.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 c175048dbebd50a73ece191ced56863fc48f411c35efb526468954d92647a819
MD5 cab65fb6a8b6688d7bb2234072329a93
BLAKE2b-256 421a13c1b112831dce99e80cd2a12a2712a5cd55cf13b78ca20f115cf77d1b9e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page