Skip to main content

Natural language to any query language, starting with GraphQL

Project description

text2ql

Natural Language to Query Language framework.

text2ql is designed as a pip-installable package for converting natural language into query languages with a plugin architecture. The first implemented target is GraphQL.

Project goals

  • Build a reusable core abstraction (Text2QL) for text -> query conversion.
  • Support multiple target query languages (GraphQL first; SQL/Cypher/SPARQL next).
  • Keep provider integrations optional (deterministic local mode, LLM mode as adapters).
  • Encourage benchmark/data generation workflows inspired by graph-centric datasets.

Install

Install from PyPI:

pip install text2ql

Install from source (editable):

pip install -e .

For local development and tests:

pip install -e ".[dev]"

Quickstart

from text2ql import Text2QL

service = Text2QL()
result = service.generate(
    text="show top 5 client records with mail state enabled",
    target="graphql",
    schema={
        "entities": [
            {"name": "customers", "aliases": ["client", "clients"]}
        ],
        "fields": [
            {"name": "id"},
            {"name": "email", "aliases": ["mail"]},
            {"name": "status", "aliases": ["state"]}
        ],
        "default_entity": "customers",
        "default_fields": ["id", "email", "status"]
    },
    mapping={
        "filters": {"state": "status"},
        "filter_values": {"status": {"enabled": "active"}}
    },
)

print(result.query)
print(result.explanation)

Default mode is deterministic.

LLM mode (adapter-based)

from text2ql import Text2QL
from text2ql.providers.openai_compatible import OpenAICompatibleProvider

service = Text2QL(
    provider=OpenAICompatibleProvider(model="gpt-4o-mini")
)

result = service.generate(
    text="show top 3 clients with mail state enabled",
    target="graphql",
    schema={"entities": ["customers"], "fields": ["id", "email", "status"]},
    mapping={"entities": {"clients": "customers"}, "fields": {"mail": "email"}},
    context={
        "mode": "llm",
        "language": "english",
        "system_context": "Prefer customer/account semantics and preserve canonical field names.",
    },
)

Required env var for OpenAICompatibleProvider:

export OPENAI_API_KEY=...

Fallback key name also supported:

export TEXT2QL_API_KEY=...

If LLM output fails constrained JSON validation, text2ql falls back to deterministic mode.

CLI

text2ql "show top 5 client records with mail state enabled" \
  --target graphql \
  --schema '{"entities":["customers"],"fields":["id","email","status"]}' \
  --mapping '{"entities":{"client":"customers"},"fields":{"mail":"email"},"filters":{"state":"status"},"filter_values":{"status":{"enabled":"active"}}}'

You can also load JSON files:

text2ql "show top 5 client records with mail state enabled" \
  --schema-file ./schema.json \
  --mapping-file ./mapping.json

LLM mode via CLI:

export OPENAI_API_KEY=...
text2ql "show top 3 clients with mail state enabled" \
  --mode llm \
  --language english \
  --system-context "Prefer customer/account semantics and preserve canonical field names." \
  --llm-provider openai-compatible \
  --llm-model gpt-4o-mini \
  --llm-max-retries 4 \
  --llm-retry-backoff 2.0

If you do not pass --mode llm, CLI runs deterministic mode. --system-context is consumed in LLM mode and ignored in deterministic mode.

Prompt templates

LLM mode supports prompt template override through request context:

result = service.generate(
    text="list users",
    context={
        "mode": "llm",
        "prompt_template": "Convert request to JSON intent. Request: {text}\\nEntities: {entities}\\nFields: {fields}"
    },
)

Template placeholders:

  • {text}
  • {entities}
  • {fields}
  • {field_aliases}
  • {filter_aliases}

Language support:

  • english (default)
  • en (alias)

Arbitrary JSON support

text2ql can infer schema config from arbitrary nested JSON payloads, generate hybrid mappings (auto + overrides), and execute generated metadata against JSON data:

from text2ql import (
    Text2QL,
    execute_query_result_on_json,
    generate_hybrid_mapping,
    infer_schema_from_json_payload,
)

schema = infer_schema_from_json_payload(raw_json_payload)
mapping = generate_hybrid_mapping(schema_payload=schema, data_payload=raw_json_payload)
service = Text2QL()
result = service.generate("how many qqq do I own", schema=schema, mapping=mapping)

rows, note = execute_query_result_on_json(result, raw_json_payload)
print(result.query)
print(rows, note)

Deterministic parsing includes a built-in holdings pattern for prompts like:

  • how many <asset> do I own

Nested GraphQL + schema validation

You can define relation-aware schema config for nested query generation and strict validation.

Example:

result = service.generate(
    text="show customers with latest order total",
    target="graphql",
    schema={
        "entities": ["customers"],
        "fields": {"customers": ["id", "email"]},
        "args": {"customers": ["limit", "status"]},
        "relations": {
            "customers": {
                "orders": {
                    "target": "orders",
                    "fields": ["id", "total", "createdAt"],
                    "args": ["limit"],
                    "aliases": ["order"]
                }
            }
        }
    },
)

Behavior:

  • Detects nested intents (e.g. latest order total) and emits nested selections.
  • Validates entity, fields, and args against schema before returning query.
  • Drops invalid fields/args and records notes in result.metadata["validation_notes"].

Additional GraphQL intent support:

  • Aggregations: count, sum, avg, min, max
  • Advanced filters:
    • range: price between 10 and 20 -> price_gte, price_lte
    • in-list: category in retail, wholesale -> category_in
    • grouped filters: AND/OR groups
  • Post-generation introspection validation:
    • supply schema["introspection"] with query and types
    • engine validates generated entity/args/fields against introspection metadata

Dataset + evaluation hooks

from text2ql import Text2QL, ingest_dataset, generate_synthetic_examples, evaluate_examples

examples = ingest_dataset("examples.jsonl")
synthetic = generate_synthetic_examples(examples, variants_per_example=2)
report = evaluate_examples(Text2QL(), synthetic)
print(report.exact_match_accuracy, report.execution_accuracy)

Execution accuracy against a real backend:

from text2ql import Text2QL, evaluate_examples

def backend_executor(query: str, example):
    # Replace with your real backend call (GraphQL/SQL/etc).
    # Return normalized execution payload (rows/object).
    return run_query_against_backend(query)

report = evaluate_examples(
    Text2QL(),
    examples,
    execution_backend=backend_executor,
)
print(report.execution_accuracy)

Domain-aware synthetic rewrites via plugins:

synthetic = generate_synthetic_examples(
    examples,
    variants_per_example=3,
    rewrite_plugins=["generic", "portfolio"],
    domain="portfolio",
)

More domain examples:

crm_synthetic = generate_synthetic_examples(
    examples,
    variants_per_example=3,
    rewrite_plugins=["generic", "crm"],
    domain="crm",
)

healthcare_synthetic = generate_synthetic_examples(
    examples,
    variants_per_example=3,
    rewrite_plugins=["generic", "healthcare"],
    domain="healthcare",
)

Template slot generation and schema-aware lexicalization:

seed = [DatasetExample(
    text="show sales pipeline",
    target="graphql",
    expected_query="query GeneratedQuery { opportunities { amount } }",
    schema={
        "entities": ["opportunities"],
        "fields": {"opportunities": ["amount", "createdAt", "stage"]},
        "args": {"opportunities": ["stage"]},
    },
    mapping={"filter_values": {"stage": {"open": "Open"}}},
)]

synthetic = generate_synthetic_examples(
    seed,
    variants_per_example=4,
    rewrite_plugins=["generic", "crm"],
    domain="crm",
)

Generation behavior:

  • Uses per-domain slot templates filled from schema/mapping (entity, metric, date, filter, value).
  • Applies schema-aware lexicalization so synthetic prompts prefer schema/mapping terms.
  • Scores candidates (source, confidence, novelty) and keeps highest-scoring variants first.
  • Adds metadata on each synthetic example:
    • synthetic_rewrite_source
    • synthetic_rewrite_confidence
    • synthetic_rewrite_novelty
    • synthetic_rewrite_score

Built-in rewrite plugins:

  • generic: neutral paraphrases (show/list, top/first, etc.).
  • portfolio: holdings/asset phrasing rewrites (how many qqq do i own -> quantity/share variants).
  • banking: account and money-flow paraphrases (balance, transfer, deposit, withdraw, statement).
  • crm: sales workflow paraphrases (leads, opportunities, pipeline, contacts, deals).
  • healthcare: clinical data paraphrases (patients, encounters, diagnosis, medications, labs, claims).
  • ecommerce: shopping/order paraphrases (orders, products, cart, inventory, refund, shipment).

Dataset format

Supported file types:

  • .jsonl: one JSON object per line
  • .json: JSON array of objects

Required fields per example:

  • text (string)
  • expected_query (string)

Optional fields:

  • target (default: "graphql")
  • schema (object)
  • mapping (object)
  • context (object)
  • metadata (object)

Example .jsonl row:

{"text":"list users","target":"graphql","expected_query":"query GeneratedQuery { user { id name } }"}

Evaluation metrics

  • exact_match_accuracy: normalized string match (whitespace-insensitive).
  • execution_accuracy:
    • with execution_backend: compares real backend execution outputs for predicted vs expected queries.
    • without execution_backend: structural GraphQL signature match (entity + filters + selected fields).

For precomputed gold execution output, set example.metadata["expected_execution_result"].

Troubleshooting

  • Missing API key...: set OPENAI_API_KEY (or TEXT2QL_API_KEY) before LLM mode.
  • JSON file ... must contain an object: schema/mapping files must be top-level JSON objects.
  • Inline JSON value must contain an object: --schema/--mapping payloads must be JSON objects.
  • HTTP Error 429: Too Many Requests: provider retries with backoff; if retries fail, engine falls back to deterministic mode.
  • Provider/network errors in LLM mode: verify API URL/model/key and retry.

Testing

python3 -m pytest -m unit
python3 -m pytest -m e2e
python3 -m pytest

Publish to PyPI

Release workflow file:

  • .github/workflows/release.yml

Publishing modes:

  1. workflow_dispatch -> TestPyPI (publish_target=testpypi)
  2. GitHub Release published -> PyPI
  3. workflow_dispatch -> PyPI (publish_target=pypi)

Setup checklist:

  1. Replace placeholder URLs in pyproject.toml (project.urls).
  2. In PyPI and TestPyPI, configure Trusted Publishing for this GitHub repo/workflow.
  3. In GitHub repo settings, create environments testpypi and pypi.
  4. Protect pypi environment with required reviewers if desired.

Local preflight before release:

python -m pip install --upgrade build twine
python -m build
python -m twine check dist/*

Current architecture

  • text2ql.core.Text2QL: orchestrator/facade.
  • text2ql.types: request/result schemas.
  • text2ql.engines.*: per-target query generators.
  • text2ql.providers.*: pluggable LLM provider adapters.

Roadmap

  1. Add SQL, Cypher, Jsonata, Jq and SPARQL engines.
  2. Expand prompts and constraints per target language.
  3. Add richer synthetic generation using domain-specific rewrite plugins.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

text2ql-0.1.5.tar.gz (41.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

text2ql-0.1.5-py3-none-any.whl (42.8 kB view details)

Uploaded Python 3

File details

Details for the file text2ql-0.1.5.tar.gz.

File metadata

  • Download URL: text2ql-0.1.5.tar.gz
  • Upload date:
  • Size: 41.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for text2ql-0.1.5.tar.gz
Algorithm Hash digest
SHA256 89b9600fad5c361a42afbc9dfcc1dd81146c2a5ab4e71a5153115661a8ca3e13
MD5 973faa27197c9debe338ff019b81462d
BLAKE2b-256 efacdd5f4f8a8676445c990c49fc8524d02b1cda7f3e8383ac1aa01bb3f9a606

See more details on using hashes here.

Provenance

The following attestation bundles were made for text2ql-0.1.5.tar.gz:

Publisher: release.yml on riteshakumar/text2ql

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file text2ql-0.1.5-py3-none-any.whl.

File metadata

  • Download URL: text2ql-0.1.5-py3-none-any.whl
  • Upload date:
  • Size: 42.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for text2ql-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 668339eb7dd466e1939a0c453ff01774fd6cb81bc4b7c2825457bcc5ca1f1084
MD5 52302be355e7ee89b51248554b9eab7a
BLAKE2b-256 582d65f8a66eb1a5a0a646da3876482d6f88c1b234e0af1d48bbb3c996ec48e9

See more details on using hashes here.

Provenance

The following attestation bundles were made for text2ql-0.1.5-py3-none-any.whl:

Publisher: release.yml on riteshakumar/text2ql

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page