Skip to main content

Natural Language to Business Intelligence - Convert natural language queries to SQL, charts, and business insights

Project description

NL2BI — Natural Language to Business Intelligence

Ask questions in plain English. Get tables, charts, and answers — directly from your database.

PyPI version Python


What it does

NL2BI translates natural language questions into validated, read-only SQL queries, executes them, and returns structured results with chart recommendations — no dashboards, no SQL knowledge required.

"Show me the top 5 suppliers by contract value this quarter"
        ↓
  Intent classified (TopN) → hint steers SQL generation
        ↓
  Relevant schema retrieved → SQL planned → validated (read-only) → executed
        ↓
  Table + bar chart recommendation + NL summary

Architecture

A sequential multi-stage pipeline, plain Python (no agent framework):

Stage Component What it does
1 Schema Extractor Introspects the DB once at startup; retrieves just the relevant tables per query (lexical + business-term scoring)
2 Intent Classifier Classifies into one of 10 question types, which nudges SQL generation with a type-specific hint
3 SQL Planner Generates SQL from the query + retrieved schema + intent hint
4 Validator Rejects anything that isn't a single read-only SELECT/WITH statement (AST-checked via sqlglot — catches statement-stacking and DML hidden in CTEs)
5 Executor Runs the validated query, returns structured results
6 Self-Correction On a validation or execution failure, feeds the error back to the SQL Planner and retries (bounded by max_sql_retries, default 2)
7 Chart Reasoner Recommends a visualization type for the result columns

10 question types: Filter, KPI, Grouping, TopN, Comparison, CrossEntity, RiskAlert, Financial, Drilldown, Recommendation

Semantic layer: Map business vocabulary to schema elements the LLM would never guess from column names alone — see Business terms below.

Memory: Short-term — the last 5 turns of a conversation are always included as context. Long-term — optional (memory=True), stores past queries in a lightweight JSON-persisted embedding store and recalls similar ones on future queries.

Guardrails: Read-only execution enforcement (Stage 4) and structured JSON output from the LLM at every stage.

Observability: Every query returns result.metrics (question type, retry count, validation/execution pass-fail, latency) and emits one structured log line via logging.getLogger("nl2bi"). This is retry/latency tracking, not a golden-dataset eval harness — there's no SQL-correctness or hallucination-rate scoring yet, since that needs ground truth this project doesn't have.


Install

pip install nl2bi

For Anthropic support: pip install nl2bi[llm]. OpenAI and local (Ollama) providers need no extra install.


Quick Start

from nl2bi import NL2BI

# Connect to your database
agent = NL2BI(
    db_url="postgresql://user:password@localhost/mydb",
    llm="openai",  # or "anthropic", "local"
    api_key="your-api-key"
)

# Ask a question
result = agent.query("What are the top 10 customers by revenue last month?")

print(result.table)        # structured data (pandas DataFrame)
print(result.sql)          # generated SQL
print(result.summary)      # NL explanation
print(result.chart_type)   # recommended visualisation
print(result.metrics)      # question_type, retry_count, latency_seconds, ...

Multi-turn conversation (short-term memory is always on):

result1 = agent.query("Show me contracts expiring this quarter")
result2 = agent.query("Now filter those above $50K")  # prior turn is in context

Long-term memory (opt-in — costs one extra embedding call per query):

agent = NL2BI(db_url="...", memory=True, memory_path="memory.json")

Local models via Ollama (llm="local" expects Ollama running at http://localhost:11434 with the model already pulled):

agent = NL2BI(db_url="...", llm="local")

Business terms

Map domain vocabulary to schema elements before querying, so "churn" resolves to the right table even if no column is named that:

result = agent.query("...")  # after connecting

# Access the underlying schema extractor for glossary/description setup:
extractor = agent._orchestrator.schema_extractor
extractor.add_business_term("churn", "subscriptions.cancelled_at", "customers who stopped paying")
extractor.add_table_description("subscriptions", "Active and cancelled customer subscriptions")

Supported Databases

Any database supported by SQLAlchemy:

  • PostgreSQL
  • MySQL
  • SQLite
  • MS SQL Server
  • Oracle

Supported LLMs

  • OpenAI (GPT-4o, GPT-4o-mini)
  • Anthropic (Claude) — requires pip install nl2bi[llm]
  • Local models via Ollama (uses Ollama's OpenAI-compatible endpoint, no extra dependency)

Why not just use text-to-SQL?

Plain text-to-SQL breaks on ambiguous queries, complex joins, and multi-step questions. NL2BI adds:

  • Intent classification before SQL generation — different question types get a tailored prompt hint
  • Self-correction loops — if the SQL fails validation or execution, the agent sees why and retries
  • Session memory — follow-up questions work naturally
  • Guardrails — a rejected write/delete never reaches your database

License

MIT


Built by Mithran Bala

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nl2bi-0.2.0.tar.gz (25.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nl2bi-0.2.0-py3-none-any.whl (20.7 kB view details)

Uploaded Python 3

File details

Details for the file nl2bi-0.2.0.tar.gz.

File metadata

  • Download URL: nl2bi-0.2.0.tar.gz
  • Upload date:
  • Size: 25.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.3

File hashes

Hashes for nl2bi-0.2.0.tar.gz
Algorithm Hash digest
SHA256 52b0f4623798b4f6ca6ebffa91b00a2e75b2878fb18b55c57cb9f68e98115858
MD5 f6a09d8ff36e2b11ea853f304051d1d4
BLAKE2b-256 c4d371d30014908a42bb0f81100970619b56ad544555449b77309e3b52294631

See more details on using hashes here.

File details

Details for the file nl2bi-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: nl2bi-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 20.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.3

File hashes

Hashes for nl2bi-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 417c0b6e2be448bea9f8e1ddd9bb0d4f0c926fff722b99fa21b12fe77a9079e7
MD5 f16e23b04925bb8ba47de4ac7af8acfc
BLAKE2b-256 8023de21e4c9a70c1b14c80a1500983b96f695e2ecc30b5765688e6e1c54958c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page