Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Azure AI Fine-Tuning Sessions client library for Python

Preview client library for interactive supervised and reinforcement fine-tuning in Microsoft Foundry. Create a session, submit training or sampling requests, and save checkpoints with synchronous or asynchronous Python clients.

For general background, see the Microsoft Foundry (classic) fine-tuning overview. That article covers general fine-tuning workflows; it is not documentation for this fine-tuning sessions preview SDK.

The client combines TypeSpec-generated operations and models with maintained Python customizations for polling, lifecycle, and error handling. See the generation and validation guide for source provenance, generation instructions, and recorded validation results.

Getting started

Install the package

python -m pip install azure-ai-finetuningsessions

The distribution name is azure-ai-finetuningsessions. Python imports are now azure.ai.finetuningsessions, including the asynchronous aio namespace. Update earlier azure.ai.finetuning_sessions imports to the new spelling.

If an earlier preview was installed as azure-ai-finetuning-sessions, uninstall that distribution before installing this one:

python -m pip uninstall azure-ai-finetuning-sessions
python -m pip install azure-ai-finetuningsessions

Update dependency files and lockfiles to use azure-ai-finetuningsessions as well. A distribution-only preview used this same distribution name with the old Python namespace and the same version number. If upgrading from that snapshot, uninstall it before installing the new build so pip does not leave stale modules or skip the reinstall. Do not rely on both preview distributions being installed.

Prerequisites

  • Python 3.10 or later is required to use this package.
  • You need an Azure subscription to use this package.
  • A Foundry project with access to fine-tuning sessions and compatible model capacity.
  • A service endpoint supporting /fine_tuning/sessions.

Authenticate with Microsoft Entra ID

Install azure-identity with pip, then supply a token credential from the Azure Identity library. For example, use DefaultAzureCredential:

from azure.ai.finetuningsessions import FineTuningSessionClient
from azure.identity import DefaultAzureCredential

client = FineTuningSessionClient(
    endpoint="https://<account>.services.ai.azure.com/api/projects/<project>",
    credential=DefaultAzureCredential(),
)

AzureKeyCredential is also supported when API-key authentication is enabled for the endpoint. Default API-key authentication requires HTTPS and sends the key only to the configured origin (scheme, host, and effective port). For local development only, allow_insecure_http=True permits the configured HTTP loopback origin: localhost, 127.0.0.1, or [::1]. It does not permit remote plaintext authentication. Use HTTPS and non-production credentials for testing.

Key concepts

  • A session holds model and adapter state for training and sampling.
  • A request is submitted and then polled until its result is available; a successful HTTP submission does not mean GPU work has completed.
  • A checkpoint persists training state or sampler weights. Sampling requires a completed sampler checkpoint identifier.
  • Heartbeats keep sessions active. Close/delete sessions explicitly and close clients or use their context managers to release HTTP resources.

Examples

Create a session

from azure.ai.finetuningsessions import FineTuningSession
from azure.ai.finetuningsessions.models import LoRAConfig, TrainingType

session = FineTuningSession.create(
    client,
    base_model="<supported-base-model>",
    lora_config=LoRAConfig(rank=16),
    user_metadata={"experiment": "example", "enabled": True},
    training_type=TrainingType.GLOBAL_STANDARD,
)
try:
    sampler = session.save_weights_for_sampler(seq_id=0, sampling_session_seq_id=0)
    print(sampler.checkpoint_id)
finally:
    session.close()
    client.close()

The asynchronous entry point is azure.ai.finetuningsessions.aio.FineTuningSessionClient. Its create_session method returns a session ID after initialization; training, sampling, checkpoint, and deletion methods accept that ID. Creation supports from_checkpoint, JSON-valued user_metadata, and training_type. Session creation and checkpoint-resume methods require an explicit lora_config with a rank, including FineTuningSession.create, FineTuningSession.create_from_checkpoint, async_client.create_session, and async_client.create_session_from_checkpoint. Use values supported by the selected model; no client-side rank default is supplied.

Choose the creation API according to the lifecycle behavior needed:

API Completion and heartbeat behavior
client.sessions.create(...) / await async_client.sessions.create(...) Returns the HTTP 200 submission JSON, not an initialized session. Does not poll or start a heartbeat.
client.sessions.begin_create(...) / await async_client.sessions.begin_create(...) Returns a sync/async poller. Use poller.result() / await poller.result() for request completion. Does not start a heartbeat.
FineTuningSession.create(...) / await async_client.create_session(...) Waits for initialization, then starts the existing background heartbeat. Returns a session object / session ID.

Automatic heartbeat startup in convenience creation is unchanged; an opt-in-only heartbeat lifecycle has not been implemented.

training_type accepts TrainingType members or strings. The known wire values remain GlobalStandard, DatazoneStandard, and DeveloperTier; future strings are passed through without client-side validation. Set the property explicitly to select a tier. If omitted, the SDK leaves selection to the service.

Text and image input chunks

ModelInput reuses the canonical service model with an ordered list of InputChunk objects. The generated hierarchy uses type as its discriminator; InputChunkType names the known text and image values, while the wire type remains open to future strings.

Input Serialization and compatibility
ModelInputChunk(tokens=...) and mapping constructors Existing calls remain valid; text chunks now serialize with type="text".
Legacy token-only mappings inside ModelInput Gain type="text" only when no explicit type is present. Unknown explicit tags are preserved, not reclassified as text.
ImageChunk Existing keyword and mapping calls remain valid. The maintained class extends the generated image variant and retains bytes/base64 handling and validation.

Image checks still cover the 10,000,000-byte limit, matching JPEG/PNG/WebP signatures and formats, positive expected_tokens, and at most 64 images per example. Unknown chunk tags remaining extensible in the SDK does not guarantee that a service accepts every future variant.

This is an SDK-first change, not a production parser rollout. Local checks against the checked-out service schemas accept both tagged and untagged text; strict service discriminator enforcement and the compatibility policy are deferred to public preview (PuPr). See the generation guide for the offline evidence and its limits.

Sampling options and results

SamplingParams.response_format has type Optional[Dict[str, Any]] and requests a response format from compatible sampling providers. It is omitted when not supplied; supported formats depend on the selected model and provider.

SamplingOperationResult is a friendly alias for SampleOperationResult. Both names identify the same result model; the existing name remains available.

Sampling prompt token evidence

SampleOperationResult.prompt_tokens is a read-only Optional[int] from the API's top-level prompt_tokens result field, populated from backend usage.prompt_tokens. It counts the single prompt once, even when num_samples (n) is greater than one. Only exact non-negative Python int values are accepted; booleans, floats, and strings are not coerced. Missing, ambiguous, or invalid evidence is None (unverified), including responses from older services. The SDK does not infer this count from request length, prompt log-probabilities, or COGS metrics. The identical SamplingOperationResult alias exposes the same field; existing constructor keywords and overloads are unchanged.

Forward-only passes and session deletion

Both capabilities are available through maintained convenience methods. In the table below, session is a FineTuningSession and async_client is an azure.ai.finetuningsessions.aio.FineTuningSessionClient.

Operation Synchronous API Asynchronous API
Forward-only pass session.forward(batch) await async_client.forward(session_id, batch)
Delete a session session.delete() await async_client.delete_session(session_id)

Forward-only passes do not accumulate gradients. These forward methods split large batches into chunks, submit requests, and poll request IDs until the completed result is available. Alternatively, await async_client.forward_async(session_id, batch) returns an awaitable for the result; await that returned object to obtain the completed result.

Set AZURE_AI_FINETUNING_MAX_CHUNK_BYTES before importing the SDK to override the approximate per-request chunk-size budget. The value must be a positive integer; unset or invalid values retain the 5,000,000-byte default, and invalid values log a warning. This setting does not change service-side request limits.

The delete methods stop the session heartbeat, send HTTP DELETE, and return None. They treat HTTP 404 as success, so deleting an already absent session is safe to repeat. The service handles cascading deletion of the session's models, checkpoints, and sampling sessions; the SDK does not wait for background storage cleanup. Deletion is distinct from session.close() or await async_client.close_session(session_id), which unload the session.

Generated operations versus convenience methods

Use FineTuningSession or the async client's convenience methods for training, sampling, checkpoints, and session lifecycle operations. For request-ID-based operations, they submit work, receive HTTP 200 acceptance, poll the returned request identifier, and normalize results with convenience-level recovery, chunking, heartbeats, and error handling. Deletion sends HTTP DELETE directly; it does not poll a request ID.

The raw operation-group methods client.sessions.delete() and client.training.forward() are intentionally not generated in this preview. The convenience entry points above preserve the established preview API and provide the lifecycle, chunking, and polling behavior described above. Their supported Python customization hooks are included during SDK regeneration; the REST specification still defines both operations. Adding raw operation-group entry points would be a separate additive API change.

The raw operation groups provide lower-level access. Their accepted inputs and return types can differ from the convenience methods. Prefer convenience APIs for end-to-end session workflows.

Default raw begin_* pollers now accept the real HTTP 200 response in both sync and async clients, then GET the returned request ID within its session until completion. They do not invent HTTP 202 or an Operation-Location header, replay the POST, or start a heartbeat. Results retain OperationResult deserialization and the cls callback. The default strategy disables transport retries and redirects for both submission and polling; errors are surfaced. Its continuation token resumes the existing request with GET only.

Explicit custom polling strategies and polling=False (NoPolling or AsyncNoPolling) retain their own completion, callback, and continuation semantics. Disabling polling does not establish that GPU work completed.

All generated and convenience requests use /fine_tuning/sessions. No route selection flag or gateway rewrite is needed. The earlier use_legacy_routes option is not included in this preview.

Compatibility with earlier previews

The renamed import is an intentional migration. The established convenience methods, typed exceptions, convenience-level recovery, chunking, session-ID handling, checkpoint helpers, and environment-variable names remain available.

Requiring lora_config and LoRAConfig.rank is an intentional breaking change from earlier previews that allowed omission. Update creation and checkpoint-resume calls to pass LoRAConfig(rank=...); an empty configuration is not a supported default. CreateSessionRequest also requires lora_config with a rank.

The raw operation groups retain the established body, operation_id, and explicit per-call api_version arguments. Check the current operation signatures when migrating from an earlier regenerated-only preview. ApiError, ApiErrorResponse, and the top-level typed exceptions remain available.

Async lifecycle methods now await heartbeat shutdown before sending close/delete, and closing the async client drains its heartbeat tasks. Empty batches and sampler requests missing both a path and sampling-session ordinal are rejected locally rather than producing an invalid request or a false successful no-op.

FoundryFeaturesOptInKeys now reuses the canonical shared Foundry definition. The current 12-member inventory preserves the six original member names and values, with six canonical additions. The earlier 13-member inventory is historical: upstream removed AGENTS_OPTIMIZATION_V2_PREVIEW. This does not activate other preview features or change the fine-tuning header Foundry-Features: FineTuningSessions=V1Preview.

Troubleshooting

Inference error codes

Retryable inference failures use request_timeout, request_orphaned, inference_request_rate_limited, and inference_unavailable. The SDK exposes the server's code on RequestRetryableError.error_code. Convenience polling uses should_retry to decide whether to resubmit, honoring retry_after_sec; it does not match code names. Default raw pollers surface the error instead. Older inference codes remain supported by the same mechanism. invalid_request and internal_error remain terminal.

Retry and transport safety

The default sync and async transport retry policies never retry POST requests, including heartbeats, even if retry counts or method lists are supplied. Ordinary GET retry behavior is retained. Supplying an explicit retry_policy opts ordinary requests into that policy's behavior; the caller owns mutation replay safety. The default raw begin_* strategy additionally disables retries and redirects for its own requests, as described above.

This does not change convenience-level recovery decisions, ordinary redirect handling, or the service's should_retry contract. No server deduplication or exactly-once guarantee is implied; resubmitting an ambiguous mutation can still duplicate work.

Default pipelines remove SDK API-key authentication on cross-origin redirects. Direct-route context headers are scoped to the configured origin and session path, including prepopulated values equal to SDK defaults; those values are removed outside that scope. Distinct caller header overrides are preserved. An explicitly supplied value identical to an SDK default is scoped as a default.

Normal INFO progress logs include status and identifiers, not full create or completion payloads. FINETUNING_VERBOSE_HTTP=1 explicitly enables body logging; do not enable it for sensitive customer data.

When supplying a custom policies list or prebuilt pipeline, the caller owns header, authentication, redirect, and retry configuration; the SDK does not replace that pipeline or promise that its defaults protect caller-owned policies.

Malformed or non-object error bodies retain typed error handling rather than causing an attribute error. Retry hints must be finite and non-negative. Mapping-form image inputs undergo the same image validation as keyword inputs.

Next steps

Use the returned sampler checkpoint ID with FineTuningSession.sample, and save training checkpoints before unloading a session. Review the generation and validation guide before changing maintained customizations or regenerating the package.

Local development

From the Azure SDK for Python repository root, install this package in editable mode (after removing any older-named preview as described above):

python -m pip install --editable ./sdk/ai/azure-ai-finetuningsessions

Run the package's tests with pytest; the package configuration enables asyncio tests. The reference verifier verifies the immutable upstream Git blobs and manifest. Intentional review deltas are recorded separately from the reproducible baseline; current runtime is not claimed to be byte-identical to that historical reference.

The generation verifier emits TypeSpec twice with maintained customizations pre-seeded and compares the complete runtime. It uses the SDK repository's shared emitter manifest and lock, not a package-local override. A changed or extra generated file is a failure, not an allowed review delta.

Contributing

This project welcomes contributions and suggestions. Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit https://cla.microsoft.com.

When you submit a pull request, a CLA-bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., label, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information, see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Release History

1.0.0b1 (2026-09-24)

Features Added

  • Add the generated InputChunk hierarchy and InputChunkType. Existing ModelInputChunk keyword/mapping calls remain valid and now emit type="text"; image inputs retain their existing serialization and validation.
  • Reuse the current 12-member canonical shared FoundryFeaturesOptInKeys, preserving six original member names/values with six canonical additions. The fine-tuning preview header is unchanged.
  • Add optional SamplingParams.response_format for compatible sampling providers and the friendly SamplingOperationResult alias for SampleOperationResult.
  • Include Loom's maintained read-only SampleOperationResult.prompt_tokens: Optional[int] for backend-reported evidence, counted once per prompt. Accept only exact non-negative integers; missing or invalid evidence is None, with no boolean/float/string coercion or inferred counts. The identical SamplingOperationResult alias and existing constructor keywords/overloads are unchanged.
  • Support AZURE_AI_FINETUNING_MAX_CHUNK_BYTES as an import-time positive-integer override for the approximate request chunk-size budget. Unset or invalid values keep the 5,000,000-byte default; invalid values log a warning. Service-side limits are unchanged.
  • Add the string-backed TrainingType enum, reusing the REST training-tier definition. Existing string inputs and omission behavior remain supported.
  • Initial preview of interactive fine-tuning sessions with synchronous and asynchronous clients.
  • Forward-only requests, checkpoint resume and deletion, JSON-valued session metadata, and training-tier selection.
  • Multimodal image input, vision/projector LoRA settings, nullable prompt log-probabilities, and per-operation metrics.
  • Bounded retries, preserved session identifiers, typed service errors, and operation progress logging.

Breaking Changes

  • Require Python 3.10 or later; Python 3.9 is no longer supported.
  • Remove AGENTS_OPTIMIZATION_V2_PREVIEW from the historical 13-member enum snapshot to follow its removal from the upstream canonical definition; the six original fine-tuning SDK members remain available.
  • Require an explicit lora_config with LoRAConfig.rank in session creation and checkpoint-resume methods, and in CreateSessionRequest. No implicit rank or empty configuration is supplied.
  • The distribution is named azure-ai-finetuningsessions and Python imports now use azure.ai.finetuningsessions. Remove older preview installations and update imports/dependency files before installing this build.
  • The preview baseline replaces the earlier regenerated-only surface with the established preview API plus explicitly documented contract changes. Callers of earlier raw models, operation methods, or keyword adapters should check the current signatures when migrating.

Bugs Fixed

  • Preserve credential redirect protections with the declared minimum Azure Core 1.37.0 as well as current versions; handle the cleanup flag location change in Azure Core 1.38.3.
  • Drain async heartbeat tasks before session/client shutdown, including concurrent cancellation; share the sync client's pipeline so custom policies work and all transports close normally.
  • Bound creation and sustained-error retry sleeps; tolerate bounded request-store propagation after creation and reject unexpected polling states immediately.
  • Reject empty training batches and incomplete sampler identifiers; bound synchronous chunk workers and offload async forward chunking.
  • Preserve generic handling for explicit non-batch payload limits and typed errors for malformed capacity retry hints; reject unhandled redirects during deletion.
  • Propagate direct-route context headers through default raw-operation pipelines with endpoint scoping and caller overrides; omit response payloads from normal INFO logs.
  • Preserve the sync heartbeat thread when shutdown has not completed, validate mapping-form image inputs, and return actual completed Tasks from async multichunk/wave methods as documented.
  • Handle nested HTTP 503 error details while preserving flat-response compatibility.
  • Map engine_dead polling failures to TrainingEngineError, preserving error codes and diagnostic references without retrying lost engine state.
  • Handle non-object error bodies safely, ignore invalid or non-finite retry hints, and validate positional image mappings consistently with keyword inputs.
  • Require HTTPS and configured-origin scoping for default API-key authentication, with explicit configured-loopback HTTP opt-in only. Remove SDK-default direct-route headers outside their origin/path scope, including prepopulated defaults, while preserving distinct caller overrides.
  • Disable automatic POST transport retries in default sync/async policies, including heartbeats. Explicit retry policies remain caller-owned; this does not guarantee deduplication or change convenience-level recovery decisions.
  • Make default raw begin_* pollers use HTTP 200 acceptance and session/request-ID GET polling in both clients, without POST replay or synthetic HTTP 202. Preserve result callbacks, custom polling, no-polling, and strategy-specific continuation behavior.

Other Changes

  • Reuse 18 canonical TypeSpec aliases, including ModelInput, while retaining 22 compatibility models and three compatibility unions for intentional Python-only contracts. Correct LoRA, sampling, and training-tier documentation without changing service behavior.
  • Move Python-only TypeSpec compatibility definitions to session-finetuning/models-custom-code.tsp, imported only by SDK generation. Retain the existing SDK namespace, public types, and generated type identities; the move does not add these projections to REST.
  • Keep the input-chunk migration SDK-first: normalize legacy token-only ModelInput mappings only when type is absent and preserve unknown explicit tags. Production service behavior is unchanged; strict discriminator enforcement and compatibility policy are deferred to public preview (PuPr).
  • Refresh TypeSpec provenance after snake_case training-tier member naming and service-grounded identifier minimum lengths; training-tier wire values are unchanged.
  • Use the Foundry required-preview operation contract and documented schema defaults with supported Python customization hooks for request query ordering and established convenience behavior.
  • Use the repository's shared Python emitter 0.63.8, backend 0.38.0, compiler 1.16.0, and client generator core 0.72.2. Genuine regeneration incorporates the upstream unused-import fix; generated files are not manually patched to pass CI.
  • Remove the package-local emitter override. Keep standalone tooling archives and validation records under the package's engineering directory, and verification commands under its scripts directory; standard README and generation guidance remain at the package root.
  • Record the unchanged 48-file historical oracle at source commit 485774df502642879fdf3a53777be4a0d95155dc in eng/generation/reference.json. See GENERATION.md for current guidance, exact reviewed contracts, and historical evidence.
  • Retain /fine_tuning/sessions routes and established convenience APIs. The earlier use_legacy_routes option is not included in this preview.
  • Use AST overload inspection on all supported Python versions rather than importing Python 3.11-only typing.get_overloads; retain the assertions on Python 3.10.
  • Correct the spelling pipeline failure on Python's global-namespace keyword in the surface verifier without changing runtime behavior or weakening validation.
  • Preserve the separately committed preview baseline and record intentional changes in eng/generation/review-deltas.json. Background heartbeat startup remains unchanged; opt-in-only startup is still deferred.

The current refresh has two matching isolated emissions and a passing public source-test run. Final test totals and gate outcomes must come from final validation records; final Loom, installed-wheel, and public CI results are not yet verified. No commands or builds were run for this documentation-only update.

Historical local validation, before the input-chunk and shared-enum changes: 1,637 SDK tests passed on Python 3.13.14. The comparison and negative-guard counts, generation method, and limitations are in GENERATION.md. These results do not validate the final combined changes or claim a new final-wheel run, validation of the new tests on Python 3.10, remote CI success, or release approval.

Metadata

Release files for azure-ai-finetuningsessions 1.0.0b1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for azure-ai-finetuningsessions 1.0.0b1
File Size Uploaded
azure_ai_finetuningsessions-1.0.0b1.tar.gz 358.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for azure-ai-finetuningsessions 1.0.0b1
File Interpreter ABI Platform
azure_ai_finetuningsessions-1.0.0b1-py3-none-any.whl Python 3 none any Details

Total release size: 521.1 kB

Release files / azure_ai_finetuningsessions-1.0.0b1.tar.gz

Download URL azure_ai_finetuningsessions-1.0.0b1.tar.gz
Size 358.6 kB
Tags Source
SHA-256 checksum
How to use checksums
c88706f558bb1694b53989ce1297caece7940267312bc973d5008a5542eac17f
BLAKE2b-256 checksum
How to use checksums
75a81b89102decbe4ed9a0b0ac48adf67220e835d0254bfa00daa2392c10b411
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via RestSharp/106.13.0.0

Release files / azure_ai_finetuningsessions-1.0.0b1-py3-none-any.whl

Download URL azure_ai_finetuningsessions-1.0.0b1-py3-none-any.whl
Size 162.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
87ee14e4fb69116be584e49c5bd85ae09b59def7bda7f8431ed39f47b5e77181
BLAKE2b-256 checksum
How to use checksums
9b8553d77a3e6ed58b655b87bdc71a906cd067193d2737deb34141e15fd1507b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via RestSharp/106.13.0.0

Release history Release notifications | RSS feed

This release

1.0.0b1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page