Skip to main content

Python library for creating vectorized data from text or files.

Project description

Vectoriz

PyPI version

GitHub license

Python Version

GitHub issues

GitHub stars

GitHub forks

Vectoriz is available on PyPI and can be installed via pip:

pip install vectoriz

A tool for generating vector embeddings for Retrieval-Augmented Generation (RAG) applications.

Overview

This project provides utilities to create, manage, and optimize vector embeddings for use in RAG systems. It streamlines the process of converting documents and data sources into vector representations suitable for semantic search and retrieval.

Features

  • Document processing and chunking
  • Vector embedding generation using various models
  • Vector database integration
  • Optimization tools for RAG performance
  • Easy-to-use API for embedding creation

Installation

git clone https://github.com/PedroHenriqueDevBR/vectoriz.git
cd vectoriz
pip install -r requirements.txt

Usage

# initial informations
index_db_path = "./data/faiss_db.index" # path to save/load index
np_db_path = "./data/np_db.npz" # path to save/load numpy data
directory_path = "/home/username/Documents/" # Path where the files (.txt, .docx) are saved

# Class instance
transformer = TokenTransformer()
files_features = FilesFeature()

# Load files and create a argument class (pack with embedings, chunk_names and text_list)
argument = files_features.load_all_files_from_directory(directory_path)

# Created FAISS index to be used in queries
token_data = transformer.create_index(argument.text_list)
index = token_data.index

# To load files from VectorDB use
vector_client = VectorDBClient()
vector_client.load_data(self.index_db_path, self.np_db_path)
index = vector_client.faiss_index
argument = vector_client.file_argument

# To save data on VectorDB use
vector_client = VectorDBClient(index, argument)
vector_client.save_data(index_db_path, np_db_path)

# To search information on index
query = input(">>> ")
amoount_content = 1
response = self.transformer.search(query, self.index, self.argument.text_list, amoount_content)
print(response)

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vectoriz-1.0.0.tar.gz (13.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vectoriz-1.0.0-py3-none-any.whl (14.5 kB view details)

Uploaded Python 3

File details

Details for the file vectoriz-1.0.0.tar.gz.

File metadata

  • Download URL: vectoriz-1.0.0.tar.gz
  • Upload date:
  • Size: 13.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.3

File hashes

Hashes for vectoriz-1.0.0.tar.gz
Algorithm Hash digest
SHA256 1650a18ad60615bd95ad5b1850d8dd650197a2e8947226612b97d7dd9936c080
MD5 446213df54986bf889f0136899a626ad
BLAKE2b-256 4b6bededa82a2621b1a0dab4394a49951ea7def838c45fd46cd57412694bc385

See more details on using hashes here.

File details

Details for the file vectoriz-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: vectoriz-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 14.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.3

File hashes

Hashes for vectoriz-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 178fdb1a7ce3ff970f1981330a677801685ba61fb2e04c60ec118840fde87133
MD5 d28c268fb9a419961f24c1bb2c1400b2
BLAKE2b-256 4e8afc5048db7e3dd77bcf7b1fc4c95caedca7743f1a987a17b896354f2f580a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page