Skip to main content

ANDI: Advanced Natural Language Database Interface

Project description

ANDI: Advanced Natural Language Database Interface 🤖🍃

ANDI (Advanced Natural Language Database Interface) is an agentic Python library that transforms your MongoDB database into a natural language interface. By leveraging the reasoning capabilities of gpt-4o-mini, ANDI allows you to generate and execute complex MongoDB queries—including dynamic runtime variables—using plain English.

Instead of writing rigid CRUD endpoints, consolidate your data fetching into a single, intelligent NLP endpoint.

✨ Features

  • 🗣️ Natural Language to NoSQL: Write your database queries in plain English. ANDI translates your intent into precise MongoDB syntax.
  • 🧠 Agentic Solution: Powered by gpt-4o-mini, ANDI understands context and constructs highly accurate database operations.
  • 🔒 Privacy-First Schema Management: ANDI connects via your MongoDB URI, infers the collection schema, and caches it locally. Only the schema structure is shared with the LLM—your actual database records are never exposed to the agent.
  • ⚡ Runtime Variables: Safely inject dynamic variables into your natural language prompts at runtime without string-concatenation vulnerabilities.
  • 🛠️ Advanced Operations: Currently supports standard find queries and complex aggregate pipelines out of the box.
  • 🎯 Single Endpoint Architecture: Replace dozens of rigid REST API routes with a single, flexible natural language data-fetching endpoint.

📦 Installation

Available on PyPI. Install ANDI using pip:

pip install andi

⚙️ Prerequisites

To use ANDI, you will need:

  1. A valid MongoDB Connection String URL.
  2. An OpenAI API Key (with access to the gpt-4o-mini model).

🚀 Quick Start

Here is a basic example of how to initialize ANDI and run a natural language query against your database. Refers test directory

`class TestProject:

def testing_nlp(self):
    url = Connection.get_connection()
    print(url)
    status = Andi.initialize_connection(self, url=url, database_name="aristotle")
    print(status)

    user_coll = Andi.analyze_schemas(self,base_collections=["users", "wallets", "weekly_leaderboard"])
    print(user_coll)

    # Example 1
    # Please stick with database terminalogy. Strictly use field name as per schema and collection name to avoid LLM hallucination.
    # As databases are case sensistive treat NLP tool the same. Explicitly ask Regex operation.
    #
    # Basic rules to use runtime_inputs
    # "email":"${email}"
    # Left side of assignment operator must be a field name matching with schema field name
    # Right side of assignment operator "${email}" must match variable_name
    # Always provide datatype otherwise range operators like "$in" will not work if query uses one. 

    email = "test_user_8_fischertimothy@gmail.com"

    intent = {
      "intent": {
        "goal": "Find preferred_language of user where email=email",
        "runtime_inputs":[
            {
                "email":"${email}",
                "datatype": "string"
            }
        ],
        "projection":["name", "preferred_language"]
      }
    }
    output = Andi.run_nlp_query(self, intent, query_identifier=None, retry=False, email=email)
    print(output)

    # Example 2
    # 
    intent = {
      "intent": {
        "goal": """
        Write an aggregate where option_selected is equal to correct_answer and create leaderboard for top 10 emails based on timesaved.
        Also perform a lookup to get the name where email matches in engagement collection from users collection""",
        "runtime_inputs":[
        ],
        "projection":["name", "email", "rank", "timesaved"]
      }
    }
    output = Andi.run_nlp_query(self, intent, query_identifier=None, retry=False)
    print(output)

    # Example 3: Query Debugging and execution

    email = "test_user_8_fischertimothy@gmail.com"

    intent = {
      "intent": {
        "goal": "Find preferred_language of user where email=email",
        "runtime_inputs":[
            {
                "email":"${email}",
                "datatype": "string"
            }
        ],
        "projection":["name", "preferred_language"]
      }
    }
    query = Andi.build_nlp_query(self, intent, query_identifier=None, retry=False)
    print(query)

    query_output = Andi.run_query_executor(self, query, email=email)
    print(query_output)

`

🏗️ How It Works

  1. Schema Extraction & Caching: When you connect ANDI to your MongoDB instance, it scans the specified collections to map out the document structures, data types, and nested fields. This schema is saved locally (e.g., in a .andi_schema.json file).
  2. Prompt Construction: When a query is initiated, ANDI feeds your natural language prompt, the mapped variables, and the locally cached schema to gpt-4o-mini.
  3. Agentic Translation: The LLM generates the exact MongoDB find dictionary or aggregate pipeline required to satisfy the request.
  4. Execution: ANDI securely executes the generated query against your MongoDB database and returns the raw Python dictionaries.

📖 Supported Operations

ANDI's LLM routing engine currently supports the generation and execution of:

  • find(): For standard filtering, sorting, and projection operations.
  • aggregate(): For complex data transformations, grouping, unwinding, and multi-stage pipelines.

(Note: Write operations like insert, update, and delete are intentionally restricted in this version to ensure read-only safety for data-fetching endpoints).

🤝 Contributing

Contributions, issues, and feature requests are welcome! Feel free to check the issues page if you want to contribute.

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

andi_ai-0.1.0.tar.gz (4.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

andi_ai-0.1.0-py3-none-any.whl (4.8 kB view details)

Uploaded Python 3

File details

Details for the file andi_ai-0.1.0.tar.gz.

File metadata

  • Download URL: andi_ai-0.1.0.tar.gz
  • Upload date:
  • Size: 4.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.10

File hashes

Hashes for andi_ai-0.1.0.tar.gz
Algorithm Hash digest
SHA256 e63c84956fa30d5df15d7a985745571d9cb1f197e39377f0e5900cf7f51d1f05
MD5 6de34eac03f46488fd961eaf9acd4ec4
BLAKE2b-256 89c0b31f24bcd134b395d82c2fe6e9a6249cd146b3f9bf944252bb7ec9240f65

See more details on using hashes here.

File details

Details for the file andi_ai-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: andi_ai-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 4.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.10

File hashes

Hashes for andi_ai-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3c62d9e313e6e2eab65ce6081fef7820802f6ec0ca2d822aac15948252e127aa
MD5 c9fa4e0330ea9dae8ca0761a9ed514d0
BLAKE2b-256 a8aeb9d816868ed55010c3c9bad036920560f3d746f12e2a3d7e29fddaa9281f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page