Skip to main content

Natural language to any query language, starting with GraphQL

Project description

text2ql

Natural Language to Query Language framework.

text2ql is designed as a pip-installable package for converting natural language into query languages with a plugin architecture.

Project goals

  • Build a reusable core abstraction (Text2QL) for text -> query conversion.
  • Support multiple target query languages (GraphQL, SQL first; Cypher/SPARQL/JQ/Jsonata next).
  • Keep provider integrations optional (deterministic local mode, LLM mode as adapters).
  • Provide a CLI for quick experimentation and batch evaluation.
  • Include dataset ingestion, synthetic variant generation, and evaluation hooks for benchmarking.

Install

Python requirement: >=3.10.

Pip Users

Install from PyPI:

pip install text2ql

Quick smoke test:

python -c "from text2ql import Text2QL; print(Text2QL().generate(text='list users').query)"
text2ql --help

Repo / Development Users

Install from source (editable):

pip install -e .

For local development and tests:

pip install -e ".[dev]"

Getting Started

Install + verify:

pip install text2ql
python -c "from text2ql import Text2QL; print(Text2QL().generate(text='list users').query)"
text2ql --help

Four copy-paste starter commands:

# 1) GraphQL deterministic
text2ql "show top 5 client records with mail state enabled" \
  --target graphql \
  --schema '{"entities":["customers"],"fields":["id","email","status"]}' \
  --mapping '{"entities":{"client":"customers"},"fields":{"mail":"email"},"filters":{"state":"status"},"filter_values":{"status":{"enabled":"active"}}}'

# 2) SQL deterministic
text2ql "show customers highest total first 5 offset 10" \
  --target sql \
  --schema '{"entities":["customers"],"fields":{"customers":["id","total","status"]}}'

# 3) LLM mode (GraphQL)
export OPENAI_API_KEY=...
text2ql "show latest 5 orders with status active" \
  --target graphql \
  --mode llm \
  --schema '{"entities":["orders"],"fields":{"orders":["id","status","createdAt"]},"args":{"orders":["status","limit","orderBy","orderDirection"]}}'

# 4) LLM mode (SQL)
export OPENAI_API_KEY=...
text2ql "show latest 5 orders with status active" \
  --target sql \
  --mode llm \
  --schema '{"entities":["orders"],"fields":{"orders":["id","status","createdAt"]}}'

Production Setup

Recommended file layout:

project/
  schema.json
  mapping.json
  data.json
  expected_query.graphql
  expected_rows.json

Create minimal files:

cat > schema.json <<'JSON'
{
  "entities": ["positions"],
  "fields": {"positions": ["symbol", "quantity", "status"]},
  "args": {"positions": ["symbol", "status", "limit", "offset", "orderBy", "orderDirection"]}
}
JSON

cat > mapping.json <<'JSON'
{
  "filters": {"ticker": "symbol"},
  "filter_values": {"symbol": {"qqq": "QQQ"}}
}
JSON

cat > data.json <<'JSON'
{
  "portfolio_data": {
    "positions": [
      {"symbol": "QQQ", "quantity": 100.104, "status": "active"},
      {"symbol": "AAPL", "quantity": 12, "status": "active"}
    ]
  }
}
JSON

Run with synthetic variants + execution evaluation:

text2ql "how many qqq do i own" \
  --target graphql \
  --schema-file ./schema.json \
  --mapping-file ./mapping.json \
  --data-file ./data.json \
  --variants-per-example 3 \
  --rewrite-plugins generic,portfolio \
  --domain portfolio \
  --expected-execution-file ./expected_rows.json

Operational notes:

  • Set OPENAI_API_KEY (or TEXT2QL_API_KEY) for --mode llm.
  • --expected-query / --expected-query-file / --expected-execution-file require --data-file.
  • If expected query execution cannot be derived from payload JSON, CLI emits an eval warning and skips that item from accuracy denominator.

Feature Matrix

Capability GraphQL SQL
Deterministic generation Yes Yes
LLM mode + constrained parsing Yes Yes
Schema/introspection validation Yes Yes
Enum/type coercion Yes Yes
Advanced filters (not, !=, ranges, grouped precedence) Yes Yes
Order parsing (latest/highest/lowest) Yes Yes
Pagination (limit, offset, first, after) Yes Yes
Nested/relation safety Yes Yes
Structural execution match (no backend) Yes Yes
Real backend execution accuracy hook Yes Yes (via evaluate_examples(..., execution_backend=...))
Synthetic rewrite plugins Yes Yes (target-agnostic dataset API)

CLI-First Workflow

  1. Generate hybrid mapping from schema + data:
text2ql --generate-hybrid-mapping \
  --schema-file ./schema.json \
  --data-file ./data.json \
  --mapping-output-file ./mapping.generated.json
  1. Generate query (GraphQL):
text2ql "how many qqq do i own" \
  --target graphql \
  --schema-file ./schema.json \
  --mapping-file ./mapping.generated.json
  1. Generate query (SQL):
text2ql "show latest 5 positions by quantity" \
  --target sql \
  --schema-file ./schema.json \
  --mapping-file ./mapping.generated.json
  1. Batch prompt variants + execution eval in one CLI run:
text2ql "how many qqq do i own" \
  --target graphql \
  --schema-file ./schema.json \
  --mapping-file ./mapping.generated.json \
  --data-file ./data.json \
  --variants-per-example 5 \
  --rewrite-plugins generic,portfolio \
  --domain portfolio \
  --expected-query-file ./expected_query.graphql

Streamlit Playground

Run an interactive UI for GraphQL + SQL testing:

pip install -e ".[app]"
streamlit run examples/streamlit_app.py

What you get:

  • Deterministic and LLM modes in one app.
  • GraphQL and SQL target switch.
  • Bundled sample data (examples/sample_schema.json, examples/sample_data.json) or uploaded JSON files.
  • Synthetic rewrite controls (variants, plugins, domain).
  • GraphQL execution on JSON payload + optional expected-query execution match.
  • SQL optional expected-query signature match.

Quickstart

from text2ql import Text2QL

service = Text2QL()
result = service.generate(
    text="show top 5 client records with mail state enabled",
    target="graphql",
    schema={
        "entities": [
            {"name": "customers", "aliases": ["client", "clients"]}
        ],
        "fields": [
            {"name": "id"},
            {"name": "email", "aliases": ["mail"]},
            {"name": "status", "aliases": ["state"]}
        ],
        "default_entity": "customers",
        "default_fields": ["id", "email", "status"]
    },
    mapping={
        "filters": {"state": "status"},
        "filter_values": {"status": {"enabled": "active"}}
    },
)

print(result.query)
print(result.explanation)

Default mode is deterministic.

LLM mode (adapter-based)

from text2ql import Text2QL
from text2ql.providers.openai_compatible import OpenAICompatibleProvider

service = Text2QL(
    provider=OpenAICompatibleProvider(model="gpt-4o-mini")
)

result = service.generate(
    text="show top 3 clients with mail state enabled",
    target="graphql",
    schema={"entities": ["customers"], "fields": ["id", "email", "status"]},
    mapping={"entities": {"clients": "customers"}, "fields": {"mail": "email"}},
    context={
        "mode": "llm",
        "language": "english",
        "system_context": "Prefer customer/account semantics and preserve canonical field names.",
    },
)

Required env var for OpenAICompatibleProvider:

export OPENAI_API_KEY=...

Fallback key name also supported:

export TEXT2QL_API_KEY=...

If LLM output fails constrained JSON validation, text2ql falls back to deterministic mode.

CLI

text2ql "show top 5 client records with mail state enabled" \
  --target graphql \
  --schema '{"entities":["customers"],"fields":["id","email","status"]}' \
  --mapping '{"entities":{"client":"customers"},"fields":{"mail":"email"},"filters":{"state":"status"},"filter_values":{"status":{"enabled":"active"}}}'

SQL target via CLI:

text2ql "show customers highest total first 5 offset 10" \
  --target sql \
  --schema '{"entities":["customers"],"fields":{"customers":["id","total","status"]}}'

You can also load JSON files:

text2ql "show top 5 client records with mail state enabled" \
  --schema-file ./schema.json \
  --mapping-file ./mapping.json

LLM mode via CLI:

export OPENAI_API_KEY=...
text2ql "show top 3 clients with mail state enabled" \
  --mode llm \
  --language english \
  --system-context "Prefer customer/account semantics and preserve canonical field names." \
  --llm-provider openai-compatible \
  --llm-model gpt-4o-mini \
  --llm-max-retries 4 \
  --llm-retry-backoff 2.0

SQL in LLM mode:

export OPENAI_API_KEY=...
text2ql "show latest 5 orders with status active" \
  --target sql \
  --mode llm \
  --schema '{"entities":["orders"],"fields":{"orders":["id","status","createdAt"]}}'

If you do not pass --mode llm, CLI runs deterministic mode. --system-context is consumed in LLM mode and ignored in deterministic mode.

CLI parity for synthetic variants + execution evaluation:

text2ql "how many qqq do i own" \
  --schema-file ./schema.json \
  --mapping-file ./mapping.json \
  --data-file ./portfolio.json \
  --variants-per-example 3 \
  --rewrite-plugins generic,portfolio \
  --domain portfolio \
  --expected-query-file ./expected.graphql

Execution-eval notes:

  • --expected-query / --expected-query-file / --expected-execution-file require --data-file.
  • --data-file should be the execution payload JSON used for query evaluation.
  • If expected query execution cannot be derived from the payload, CLI reports an eval warning and skips that sample from accuracy denominator.

Minimal file-based CLI example:

cat > schema.json <<'JSON'
{
  "entities": ["positions"],
  "fields": {"positions": ["symbol", "quantity"]},
  "args": {"positions": ["symbol", "limit", "orderBy", "orderDirection"]}
}
JSON

cat > mapping.json <<'JSON'
{
  "filters": {"ticker": "symbol"},
  "filter_values": {"symbol": {"qqq": "QQQ"}}
}
JSON

cat > data.json <<'JSON'
{
  "portfolio_data": {
    "positions": [
      {"symbol": "QQQ", "quantity": 100.104},
      {"symbol": "AAPL", "quantity": 12}
    ]
  }
}
JSON

text2ql "how many qqq do i own" \
  --schema-file ./schema.json \
  --mapping-file ./mapping.json

Prompt templates

LLM mode supports prompt template override through request context:

result = service.generate(
    text="list users",
    context={
        "mode": "llm",
        "prompt_template": "Convert request to JSON intent. Request: {text}\\nEntities: {entities}\\nFields: {fields}"
    },
)

Template placeholders:

  • {text}
  • {entities}
  • {fields}
  • {field_aliases}
  • {filter_aliases}

Language support:

  • english (default)
  • en (alias)

Arbitrary JSON support

text2ql can infer schema config from arbitrary nested JSON payloads, generate hybrid mappings (auto + overrides), and execute generated metadata against JSON data:

from text2ql import (
    Text2QL,
    execute_query_result_on_json,
    generate_hybrid_mapping,
    infer_schema_from_json_payload,
)

schema = infer_schema_from_json_payload(raw_json_payload)
mapping = generate_hybrid_mapping(schema_payload=schema, data_payload=raw_json_payload)
service = Text2QL()
result = service.generate("how many qqq do I own", schema=schema, mapping=mapping)

rows, note = execute_query_result_on_json(result, raw_json_payload)
print(result.query)
print(rows, note)

Deterministic parsing includes a built-in holdings pattern for prompts like:

  • how many <asset> do I own

Nested GraphQL + schema validation

You can define relation-aware schema config for nested query generation and strict validation.

Example:

result = service.generate(
    text="show customers with latest order total",
    target="graphql",
    schema={
        "entities": ["customers"],
        "fields": {"customers": ["id", "email"]},
        "args": {"customers": ["limit", "status"]},
        "relations": {
            "customers": {
                "orders": {
                    "target": "orders",
                    "fields": ["id", "total", "createdAt"],
                    "args": ["limit"],
                    "aliases": ["order"]
                }
            }
        }
    },
)

Behavior:

  • Detects nested intents (e.g. latest order total) and emits nested selections.
  • Validates entity, fields, and args against schema before returning query.
  • Drops invalid fields/args and records notes in result.metadata["validation_notes"].

Additional GraphQL intent support:

  • Aggregations: count, sum, avg, min, max
  • Advanced filters:
    • range: price between 10 and 20 -> price_gte, price_lte
    • in-list: category in retail, wholesale -> category_in
    • grouped filters: AND/OR groups
  • Post-generation introspection validation:
    • supply schema["introspection"] with query and types
    • engine validates generated entity/args/fields against introspection metadata

SQL support

SQL engine supports deterministic generation with:

  • strict schema/table/column validation
  • sort/order parsing (latest, highest, lowest) -> ORDER BY ... ASC|DESC
  • pagination (limit, offset, first, after)
  • robust filter parsing:
    • negation (not, !=, is not)
    • comparative operators (>, <, >=, <=)
    • ranges (between, date ranges)
    • grouped precedence (AND/OR groups)
  • type coercion (int, float, bool, null, date-like literals, enums via introspection)
  • relation-safe joins from schema relations with local relation filter extraction

Dataset + evaluation hooks

from text2ql import Text2QL, ingest_dataset, generate_synthetic_examples, evaluate_examples

examples = ingest_dataset("examples.jsonl")
synthetic = generate_synthetic_examples(examples, variants_per_example=2)
report = evaluate_examples(Text2QL(), synthetic)
print(report.exact_match_accuracy, report.execution_accuracy)

Execution accuracy against a real backend:

from text2ql import Text2QL, evaluate_examples

def backend_executor(query: str, example):
    # Replace with your real backend call (GraphQL/SQL/etc).
    # Return normalized execution payload (rows/object).
    return run_query_against_backend(query)

report = evaluate_examples(
    Text2QL(),
    examples,
    execution_backend=backend_executor,
)
print(report.execution_accuracy)

Domain-aware synthetic rewrites via plugins:

synthetic = generate_synthetic_examples(
    examples,
    variants_per_example=3,
    rewrite_plugins=["generic", "portfolio"],
    domain="portfolio",
)

More domain examples:

crm_synthetic = generate_synthetic_examples(
    examples,
    variants_per_example=3,
    rewrite_plugins=["generic", "crm"],
    domain="crm",
)

healthcare_synthetic = generate_synthetic_examples(
    examples,
    variants_per_example=3,
    rewrite_plugins=["generic", "healthcare"],
    domain="healthcare",
)

Template slot generation and schema-aware lexicalization:

from text2ql.dataset import DatasetExample

seed = [DatasetExample(
    text="show sales pipeline",
    target="graphql",
    expected_query="query GeneratedQuery { opportunities { amount } }",
    schema={
        "entities": ["opportunities"],
        "fields": {"opportunities": ["amount", "createdAt", "stage"]},
        "args": {"opportunities": ["stage"]},
    },
    mapping={"filter_values": {"stage": {"open": "Open"}}},
)]

synthetic = generate_synthetic_examples(
    seed,
    variants_per_example=4,
    rewrite_plugins=["generic", "crm"],
    domain="crm",
)

Generation behavior:

  • Uses per-domain slot templates filled from schema/mapping (entity, metric, date, filter, value).
  • Applies schema-aware lexicalization so synthetic prompts prefer schema/mapping terms.
  • Scores candidates (source, confidence, novelty) and keeps highest-scoring variants first.
  • Adds metadata on each synthetic example:
    • synthetic_rewrite_source
    • synthetic_rewrite_confidence
    • synthetic_rewrite_novelty
    • synthetic_rewrite_score

Built-in rewrite plugins:

  • generic: neutral paraphrases (show/list, top/first, etc.).
  • portfolio: holdings/asset phrasing rewrites (how many qqq do i own -> quantity/share variants).
  • banking: account and money-flow paraphrases (balance, transfer, deposit, withdraw, statement).
  • crm: sales workflow paraphrases (leads, opportunities, pipeline, contacts, deals).
  • healthcare: clinical data paraphrases (patients, encounters, diagnosis, medications, labs, claims).
  • ecommerce: shopping/order paraphrases (orders, products, cart, inventory, refund, shipment).

Dataset format

Supported file types:

  • .jsonl: one JSON object per line
  • .json: JSON array of objects

Required fields per example:

  • text (string)
  • expected_query (string)

Optional fields:

  • target (default: "graphql")
  • schema (object)
  • mapping (object)
  • context (object)
  • metadata (object)

Example .jsonl row:

{"text":"list users","target":"graphql","expected_query":"query GeneratedQuery { user { id name } }"}

Evaluation metrics

  • exact_match_accuracy: normalized string match (whitespace-insensitive).
  • execution_accuracy:
    • with execution_backend: compares real backend execution outputs for predicted vs expected queries.
    • without execution_backend: structural signature match (graphql + sql supported).

For precomputed gold execution output, set example.metadata["expected_execution_result"].

Troubleshooting

  • Missing API key...: set OPENAI_API_KEY (or TEXT2QL_API_KEY) before LLM mode.
  • JSON file ... must contain an object: schema/mapping files must be top-level JSON objects.
  • Inline JSON value must contain an object: --schema/--mapping payloads must be JSON objects.
  • HTTP Error 429: Too Many Requests: provider retries with backoff; if retries fail, engine falls back to deterministic mode.
  • Provider/network errors in LLM mode: verify API URL/model/key and retry.

Testing

python3 -m pytest -m unit
python3 -m pytest -m e2e
python3 -m pytest

Publish to PyPI

Release workflow file:

  • .github/workflows/release.yml

Publishing modes:

  1. workflow_dispatch -> TestPyPI (publish_target=testpypi)
  2. GitHub Release published -> PyPI
  3. workflow_dispatch -> PyPI (publish_target=pypi)

Setup checklist:

  1. Replace placeholder URLs in pyproject.toml (project.urls).
  2. In PyPI and TestPyPI, configure Trusted Publishing for this GitHub repo/workflow.
  3. In GitHub repo settings, create environments testpypi and pypi.
  4. Protect pypi environment with required reviewers if desired.

Local preflight before release:

python -m pip install --upgrade build twine
python -m build
python -m twine check dist/*

Current architecture

  • text2ql.core.Text2QL: orchestrator/facade.
  • text2ql.types: request/result schemas.
  • text2ql.engines.*: per-target query generators.
  • text2ql.providers.*: pluggable LLM provider adapters.

Roadmap

  1. Add Cypher, Jsonata, Jq and SPARQL engines.
  2. Expand prompts and constraints per target language.
  3. Add richer synthetic generation using domain-specific rewrite plugins.
  4. Add execution evaluation hooks for real backends (e.g. GraphQL endpoint, SQL database).
  5. Add more provider adapters and support for provider-specific features (e.g. function calling).
  6. Add few-shot example support in LLM mode with dynamic example retrieval based on schema/mapping similarity.
  7. Add a playground interface for interactive experimentation and debugging.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

text2ql-0.1.8.tar.gz (59.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

text2ql-0.1.8-py3-none-any.whl (60.0 kB view details)

Uploaded Python 3

File details

Details for the file text2ql-0.1.8.tar.gz.

File metadata

  • Download URL: text2ql-0.1.8.tar.gz
  • Upload date:
  • Size: 59.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for text2ql-0.1.8.tar.gz
Algorithm Hash digest
SHA256 4c8790edb46bea69ef0c3e3e008f54044435284535cedfbb5a98ff23843c71ef
MD5 32513e4104f106ec4d2326a1614ffa15
BLAKE2b-256 7315e40c1681fd35e417555014c43bf4ac0ad7a556df3d923714009729605ddc

See more details on using hashes here.

Provenance

The following attestation bundles were made for text2ql-0.1.8.tar.gz:

Publisher: release.yml on riteshakumar/text2ql

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file text2ql-0.1.8-py3-none-any.whl.

File metadata

  • Download URL: text2ql-0.1.8-py3-none-any.whl
  • Upload date:
  • Size: 60.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for text2ql-0.1.8-py3-none-any.whl
Algorithm Hash digest
SHA256 ffe89b4e1ee42a9ff2d4415c0469dbb8ab67e19def8d220e71839a261a209ee8
MD5 02937f118e3117bed22528b242429003
BLAKE2b-256 7d01e78c938e4f6cd8f195a49c256b4c465717de114297e94034b10e88c399ea

See more details on using hashes here.

Provenance

The following attestation bundles were made for text2ql-0.1.8-py3-none-any.whl:

Publisher: release.yml on riteshakumar/text2ql

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page