Skip to main content

GAMER: Generative Analysis for Metadata Retrieval

License Code Style semantic-release: angular Interrogate Coverage Python

Installation

Install a virtual environment with python 3.11 (install a version of python 3.11 that's compatible with your operating system).

py -3.11 -m venv .venv

On Windows, activate the environment with

.venv\Scripts\Activate.ps1

You will need access to the AWS Bedrock service in order to access the model. Once you've configured the AWS CLI, and granted access to Anthropic's Claude Sonnet 3 and 3.5, proceed to the following steps.

Install the chatbot package -- ensure virtual environment is running.

pip install metadata-chatbot

Usage

To stream results from the model,

from langchain_core.messages import HumanMessage
import asyncio

query = "What was the refractive index of the chamber immersion medium used in this experiment SmartSPIM_675387_2023-05-23_23-05-56"

async def new_astream(query):

    inputs = {"messages": [HumanMessage(query)]}

    config = {}

    async for result in stream_response(inputs,config,app):
        print(result)  # Process the yielded results

asyncio.run(new_astream(query))

Relevant repositories

Vector embeddings generation script for metadata assets Vector embeddings generation script for AIND data schema repository Streamlit app respository

High Level Overview

The project's main goal is to developing a chat bot that is able to ingest, analyze and query metadata. Metadata is accumulated in lieu with experiments and consists of information about the data description, subject, equipment and session. To maintain reproducibility standards, it is important for metadata to be documented well. GAMER is designed to streamline the querying process for neuroscientists and other users.

Model Overview

The current chat bot model uses Anthropic's Claude Sonnet 3 and 3.5, hosted on AWS' Bedrock service. Since the primary goal is to use natural language to query the database, the user will provide queries about the metadata specifically. The framework is hosted on Langchain. Claude's system prompt has been configured to understand the metadata schema format and craft MongoDB queries based on the prompt. Given a natural language query about the metadata, the model will produce a MongoDB query, thought reasoning and answer. This method of answering follows chain of thought reasoning, where a complex task is broken up into manageable chunks, allowing logical thinking through of a problem.

The main framework used by the model is Retrieval Augmented Generation (RAG), a process in which the model consults an external database to generate information for the user's query. This process doesn't interfere with the model's training process, but rather allows the model to successfully query unseen data with few shot learning (examples of queries and answers) and tools (e.g. API access) to examine these databases.

Multi-Agent graph framework

A multi-agent workflow is created using Langgraph, allowing for parallel execution of tasks, like document retrieval from the vector index, and increased developer control over the the RAG process.

Worfklow

This model uses a multi agent framework on Langraph to retrieve and summarize metadata information based on a user's natural language query. This workflow consists of 6 agents, or nodes, where a decision is made and there is new context provided to either the model or the user. Here are some decisions incorporated into the framework:

  1. To best answer the query, which data source should the model refer to?
    • Input: x (query)
    • Decides best data to query against
    • Output: entire_database, vector_embeddings, claude, data_schema
  2. If querying against the vector embeddings, does the index need to be filtered further with metdata tags, to improve optimization of retrieval?
    • Input: x (query)
    • Decides whether database can be further filtered by applying a MongoDB query
    • Output: MongoDB query, None
  3. Are the documents retrieved during retrieval relevant to the question?
    • Input: x (query), y (documents)
    • Decides whether document should be kept or tossed during summarization
    • Output: yes, no
  4. Is the tool output retrieved through tool calling relevant to the question?
    • Input: x (query), y (tool output)
    • Decides whether MongoDB queries need to be reconstructed to retrieve more relevant output
    • Output: yes, no
  5. Does the conversation need to be summarized?
    • Input: x (message list)
    • If the conversation list exceeds 6 messages, the chat history will be summarized, and earlier messages will be deleted
    • Output: yes, no

Data Retrieval

Vector Embeddings

To improve retrieval accuracy and decrease hallucinations, we use vector embeddings to access relevant chunks of information found across the database. This process starts with accessing assets, and chunking each json file to chunks of around 8000 tokens (10 chunks per file)-- each chunk preserves the hierarchy found in json files. These chunks are converted to vector arrays of size 1024, through an embedding model (Amazon's Titan 2.0 Embedding). The user's query is converted to a vector and projected onto the latent space. The chunks that contain the most relevant information will be accessed through a cosine similarity search.

AIND-data-schema-access REST API

For queries that require accessing the entire database, like count based questions, information is accessed through an aggregation pipeline, provided by one of the constructed LLM agents, and the API connection.

Current specifications

  • The model can query the fields for a specified asset.
  • The model can query metadata documents from the document database.
  • The model is able to return a list of unique values for a given field.
  • The model is able to answer count based questions.

Metadata

Release files for metadata-chatbot 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for metadata-chatbot 0.5.1
File Size Uploaded
metadata_chatbot-0.5.1.tar.gz 1.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for metadata-chatbot 0.5.1
File Interpreter ABI Platform
metadata_chatbot-0.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 2.2 MB

Release files / metadata_chatbot-0.5.1.tar.gz

Download URL metadata_chatbot-0.5.1.tar.gz
Size 1.3 MB
Tags Source
SHA-256 checksum
How to use checksums
37c0db49980266612a321684fb69aeaa0e08f2c17e27f0592c883824fcc26cce
BLAKE2b-256 checksum
How to use checksums
213dc088965c7fd3ce54fe29261d1b8d9a73f4976ad832c80aff1d19eb03892e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.9

Release files / metadata_chatbot-0.5.1-py3-none-any.whl

Download URL metadata_chatbot-0.5.1-py3-none-any.whl
Size 965.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
318a81cfe48da7a1ad20ec4e32f78cc51d2885a669d15ec675a9886d75b46db4
BLAKE2b-256 checksum
How to use checksums
6bea709aab5aeaf58e1951ba17a9122a275e9b27d08adecc850677b358944d2b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.9

Release history Release notifications | RSS feed

This release

0.5.1 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.9

2 release files

0.3.8

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

0.0.84

2 release files

0.0.83

2 release files

0.0.78

2 release files

0.0.77

2 release files

0.0.76

2 release files

0.0.75

2 release files

0.0.74

2 release files

0.0.73

2 release files

0.0.72

2 release files

0.0.71

2 release files

0.0.70

2 release files

0.0.69

2 release files

0.0.68

2 release files

0.0.49

2 release files

0.0.48

2 release files

0.0.47

2 release files

0.0.46

2 release files

0.0.45

2 release files

0.0.44

2 release files

0.0.43

2 release files

0.0.42

2 release files

0.0.41

2 release files

0.0.40

2 release files

0.0.39

2 release files

0.0.38

2 release files

0.0.37

2 release files

0.0.36

2 release files

0.0.35

2 release files

0.0.34

2 release files

0.0.33

2 release files

0.0.32

2 release files

0.0.31

2 release files

0.0.30

2 release files

0.0.29

2 release files

0.0.28

2 release files

0.0.27

2 release files

0.0.26

2 release files

0.0.25

2 release files

0.0.24

2 release files

0.0.23

2 release files

0.0.22

2 release files

0.0.19

2 release files

0.0.18

2 release files

0.0.17

2 release files

0.0.16

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.13

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page