Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Azure Data AI client library for Python

azure-data-ai provides access to Azure Data AI, hosted by Azure Inference Service. This initial preview supports semantic reranking: rank caller-supplied documents by their relevance to a query, optionally returning the documents and sentence-level scores.

The package uses the 2026-09-01-preview service API.

Getting started

Install the package

python -m pip install --pre azure-data-ai

For local development before this preview is published, run python -m pip install -e . from this package directory instead.

Prerequisites

  • Python 3.10 or later.
  • An Azure subscription and an Azure Data AI endpoint.
  • A key issued for an Azure Data AI endpoint with key-based authentication enabled, or a Microsoft Entra identity with permission to invoke the service.

For Cosmos-linked reranking, the Cosmos DB Semantic Reranker setup handles provisioning when you enable the feature in the portal. Follow that guide for assigning Semantic Reranker User at the Cosmos account scope; there is no additional manual inference-resource creation step in that portal workflow.

Authenticate the client

Use AzureKeyCredential for API-key authentication:

import os

from azure.core.credentials import AzureKeyCredential
from azure.data.ai import InferenceClient

with InferenceClient(
    endpoint=os.environ["AZURE_DATA_AI_ENDPOINT"],
    credential=AzureKeyCredential(os.environ["AZURE_DATA_AI_KEY"]),
) as client:
    result = client.semantic_rerank(
        {
            "query": "What is the capital of France?",
            "documents": [
                "Paris is the capital of France.",
                "Berlin is the capital of Germany.",
            ],
            "topK": 1,
            "returnDocuments": True,
        }
    )

    for score in result.get("scores", []):
        print(score["index"], score["score"], score.get("document"))

The key is sent in the Ocp-Apim-Subscription-Key header using Azure Core's AzureKeyCredentialPolicy. Wrap key strings in AzureKeyCredential for both synchronous and asynchronous clients:

client = InferenceClient(endpoint, credential=AzureKeyCredential(key))

Use AzureKeyCredential when you need to rotate a key with credential.update(new_key) without recreating the client. Key authentication does not require azure-identity or a token credential.

Pass the endpoint and credential to the constructor. Environment variables in these examples are only an application configuration choice; the SDK does not read AZURE_DATA_AI_ENDPOINT or AZURE_DATA_AI_KEY.

For Microsoft Entra authentication, install azure-identity separately and pass a token credential:

import os

from azure.data.ai import InferenceClient
from azure.identity import DefaultAzureCredential

with DefaultAzureCredential() as credential:
    with InferenceClient(
        os.environ["AZURE_DATA_AI_ENDPOINT"], credential
    ) as client:
        result = client.semantic_rerank(
            {"query": "capital of France", "documents": ["Paris", "Berlin"]}
        )

The client requests tokens for https://dbinference.azure.com/.default.

Key concepts

InferenceClient is the entry point. Call client.semantic_rerank(request) directly; there is no intermediate inference subclient.

Dictionary keys use the service's JSON names, such as "topK". Generated model attributes use Python names, such as top_k. Both response access forms work: result["scores"][0]["score"] and result.scores[0].score. Use result.as_dict() when an ordinary nested dictionary is needed.

from azure.data.ai.models import SemanticRerankingDocumentType, SemanticRerankingInferenceContent

request = SemanticRerankingInferenceContent(
    query="capital of France",
    documents=["Paris is the capital of France.", "Berlin is the capital of Germany."],
    top_k=1,
    return_documents=True,
    return_sentence_score=True,
    document_type=SemanticRerankingDocumentType.TEXT,
)
result = client.semantic_rerank(request)
for score in result.scores or []:
    print(score.index, score.score, score.document)
Request key Type Meaning
query str Required nonempty query string.
documents list[str] Required nonempty list of strings to rank.
model str Optional model name supported by the endpoint. Omit it to use the service's default model.
returnDocuments bool Optional. Include document text in the response.
topK int Optional. Maximum number of results to return, from 1 to 2147483647.
batchSize int Optional. Number of documents processed per batch, from 1 to 2147483647.
sort bool Optional. Return scores sorted by relevance.
documentType str Optional document format, such as text or json. JSON documents are JSON-encoded strings.
targetPaths str Required for JSON documents. Use dot notation for nested properties, such as meta.content, and commas for multiple paths, such as meta.content,id.
returnSentenceScore bool Optional. Include sentence-level scores.

Put these options inside the request dictionary, not in method keyword arguments. Omitted options use service-defined defaults. The SDK does not restrict model names or filter additional request fields.

Responses expose optional scores and meta fields. Score entries can include index, score, document, and sentenceScores. Metadata can include tokenUsage, latency, modelName, and modelVersion through dictionary access, or token_usage, latency, model_name, and model_version as model attributes on SemanticRerankingMetaResult.

Latency fields dataPreprocessTime, inferenceTime, and postProcessTime contain numeric milliseconds in the JSON response. The corresponding LatencyResult model attributes, data_preprocess_duration, inference_duration, and post_process_duration, are datetime.timedelta values. Dictionary-style access and as_dict() retain the numeric millisecond representation.

Each sentence score has a nonnegative, zero-based index and a score in the inclusive range 0–1. Sentence indices are not capped at 2.

Examples

See the samples for runnable sync and async examples.

Model selection and JSON documents

The TypeSpec example uses semantic-reranker-v1 as a model name. Model availability depends on the endpoint: use a model your endpoint supports, or omit model to use its default.

import json

result = client.semantic_rerank(
    {
        "query": "capital of France",
        "documents": [
            json.dumps({"description": "Paris is the capital of France."}),
            json.dumps({"description": "Berlin is the capital of Germany."}),
        ],
        "model": "semantic-reranker-v1",
        "documentType": "json",
        "targetPaths": "description",
        "topK": 1,
        "batchSize": 2,
        "sort": True,
        "returnDocuments": True,
        "returnSentenceScore": True,
    }
)

metadata = result.get("meta", {})
print(metadata.get("modelName"), metadata.get("modelVersion"))

Sentence-level scores

result = client.semantic_rerank(
    {
        "query": "capital of France",
        "documents": ["Paris is the capital of France. It is on the Seine."],
        "returnDocuments": True,
        "returnSentenceScore": True,
    }
)

for document in result.get("scores", []):
    for sentence in document.get("sentenceScores", []):
        print(document["index"], sentence["index"], sentence["score"])

Async

Install aiohttp separately to use the default async transport. Use an async token credential from azure.identity.aio if authenticating with Microsoft Entra.

import os

from azure.core.credentials import AzureKeyCredential
from azure.data.ai.aio import InferenceClient

async def rerank():
    async with InferenceClient(
        os.environ["AZURE_DATA_AI_ENDPOINT"],
        AzureKeyCredential(os.environ["AZURE_DATA_AI_KEY"]),
    ) as client:
        return await client.semantic_rerank(
            {"query": "capital of France", "documents": ["Paris", "Berlin"]}
        )

Response diagnostics and retries

Use the standard Azure Core response hook to capture X-Correlation-ID for service diagnostics:

from azure.core.utils import case_insensitive_dict

response_headers = case_insensitive_dict()
result = client.semantic_rerank(
    {"query": "capital of France", "documents": ["Paris", "Berlin"]},
    raw_response_hook=lambda response: response_headers.update(
        response.http_response.headers
    ),
)
print(response_headers.get("X-Correlation-ID"))

The generated clients inherit Azure Core's native retry and timeout defaults. These are not necessarily identical to the .NET SDK's defaults. Configure the desired values when constructing the client:

client = InferenceClient(
    endpoint,
    credential,
    retry_total=3,          # Use 0 to disable automatic retries.
    retry_backoff_max=60,
    connection_timeout=100,
    read_timeout=100,
)

Azure Core also supports separate retry_connect, retry_read, and retry_status limits. For reranking POST requests, include retry_on_methods=["POST"] on the operation call when retries should apply to statuses such as 429 and 502 even without Retry-After:

result = client.semantic_rerank(request, retry_on_methods=["POST"])

A service Retry-After can require a wait longer than the calculated backoff cap. Connection/read timeouts are not an overall deadline including retries. Automatic service-level .NET retry-default customization is not part of this generated baseline.

Troubleshooting

Unsuccessful service responses raise azure.core.exceptions.HttpResponseError or a more specific Azure Core exception, such as ClientAuthenticationError. For a JSON error response, inspect error.response.json()["error"] to retrieve the required code and message and any additional ProblemDetails fields. Errors can include a target, detailed errors, and nested innererror information. Response headers retain x-ms-error-code, X-Correlation-ID, and Retry-After.

Do not log credentials or sensitive document content. For HTTP 401/403, confirm the credential type and that the key belongs to the endpoint with key-based authentication enabled, or that the Entra token has the correct audience and resource permissions.

Next steps

Explore the samples to rerank text or JSON documents with API-key or Microsoft Entra authentication. Adapt the request's model, target paths, and scoring options to your application's documents and endpoint.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.microsoft.com.

When you submit a pull request, a CLA-bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., label, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information, see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Release History

1.0.0b1 (2026-09-29)

Features Added

  • Initial preview of the Azure Data AI semantic reranking client.
  • Synchronous and asynchronous InferenceClient.semantic_rerank APIs accepting generated request models or dictionaries and returning generated response models with dictionary-style access.
  • API-key authentication with AzureKeyCredential, and Microsoft Entra authentication, for both synchronous and asynchronous clients.
  • Reranking options for model selection, batching, sorting, and JSON document paths.
  • Document and sentence-level scoring, response metadata, and standard Azure Core error handling and retry policies.

Metadata

Release files for azure-data-ai 1.0.0b1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for azure-data-ai 1.0.0b1
File Size Uploaded
azure_data_ai-1.0.0b1.tar.gz 72.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for azure-data-ai 1.0.0b1
File Interpreter ABI Platform
azure_data_ai-1.0.0b1-py3-none-any.whl Python 3 none any Details

Total release size: 138.1 kB

Release files / azure_data_ai-1.0.0b1.tar.gz

Download URL azure_data_ai-1.0.0b1.tar.gz
Size 72.9 kB
Tags Source
SHA-256 checksum
How to use checksums
e241a416f57f7017dadd11210357e997a79400ee791b9319e09c6668a61f7bf3
BLAKE2b-256 checksum
How to use checksums
2b78ddb745902ae2b2e40be6c98c71591863cb081d7e4ea07d07c87f8ad9d8ec
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via RestSharp/106.13.0.0

Release files / azure_data_ai-1.0.0b1-py3-none-any.whl

Download URL azure_data_ai-1.0.0b1-py3-none-any.whl
Size 65.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a7c42f99f66e2c359e291dc0016eff0091fec034d81c268de206baddb0184e9a
BLAKE2b-256 checksum
How to use checksums
7f0cb51f071e610b74af48a8811cd9b4ec953bd087c07df23aecf6c7bffc4d5c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via RestSharp/106.13.0.0

Release history Release notifications | RSS feed

This release

1.0.0b1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page