Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

nearai-inference-sdk (Python)

Verify NEAR AI Cloud deployments, encrypt Chat Completions, and verify their response signatures. Use the asyncio-based InferenceClient for the integrated workflow, or combine the public attestation, E2EE, and signature functions.

Installation

pip install nearai-inference-sdk
from nearai_inference_sdk import InferenceClient

For a completion with a NEAR TEE model, use three stages:

  1. Before sending it, verify the NEAR AI Cloud Gateway deployment and every returned target-model deployment.
  2. Send the completion and retain its canonical model ID, completion ID, and exact request and response bytes.
  3. Fetch the completion signature and use its kind to verify the exact response bytes with the already verified model or Gateway evidence.

fetch_model_attestations() can return zero or multiple candidates. Reject an empty result, verify every returned candidate, and retain every verified result for receipt verification.

The deployment checks are useful admission and audit evidence before an inference. They are independent checks: do not treat them as proof that a particular completion travelled from that model deployment through that Gateway deployment.

What it verifies

  • Model deployment evidence: the Intel TDX quote, client nonce, signer, accepted TCB policy, runtime measurements, and measured deployment configuration. NVIDIA GPU evidence is verified when supplied and can be required by policy.
  • Gateway deployment evidence: the same deployment evidence, plus the Gateway TLS service identity bound into the quote. By default, verification also requires the TLS peer observed for the evidence request to match it.
  • A completion signature: the exact request and response bytes signed by the signer named in the returned signature.
  • Optional image build provenance: Sigstore signatures, transparency-log evidence, and a GitHub build identity selected by the caller. No publisher or version approval policy is provided by default.

The Gateway returns an explicit kind for each completion signature:

signature.kind Response verification establishes It does not establish
provider_tee A verified model-serving TEE signer signed the exact request and response bytes. The Gateway deployment or TLS identity that returned those bytes.
gateway A verified Gateway signer signed the exact client-visible request and response bytes. That an attested model executed or generated those bytes.

The Gateway currently exposes one signature for a completion. Separately verified model and Gateway deployments plus that one signature do not form a complete cryptographic chain from model execution through Gateway processing to the final bytes. In particular, the current Gateway signature over rewritten bytes has no provider-response link. cloud-api#986 tracks the proposed provider signature plus Gateway receipt chain.

Chat client and standalone functions

InferenceClient.chat.completions.create() verifies Gateway evidence, reads the model's catalog capabilities, and verifies every returned model report for supported NEAR TEE deployments. Other models use Incognito mode: Gateway verification without a claim about model TEE execution. Failed metadata or attestation checks stop the request. The same transport works with an external openai.AsyncOpenAI client through inference_client.http_client. Both support JSON and streaming responses.

Ed25519 and Gateway TLS verification are enabled by default. Set e2ee=True to encrypt protocol-supported fields to a verified model key. ECDSA is available through signing_algo='ecdsa'. E2EE or a model deployment policy requires model attestation and rejects Incognito models. Successful attestations are cached for 60 minutes per model; set attestation_cache_time_to_live_ms=0 to verify every request. Call await inference_client.verify(model) to warm Chat's verification cache before the first request. Gateway and model verification run concurrently, and all required checks must pass before Chat is sent. Optional deployment callbacks can enforce an application-owned approval policy; no approved-release allowlist is supplied by default.

Set ohttp=True to encapsulate Chat HTTP requests and responses to the attested Gateway. OHTTP is disabled by default and requires Ed25519. Field-level E2EE is configured independently; JSON, streaming, and response verification use the same interfaces.

Response-signature verification is explicit: call verify_response(id) after consuming the response. It uses the exact encrypted bytes retained by the client, without delaying delivery of decrypted content. Response records expire 60 minutes after the body finishes by default.

For a custom workflow, prepare_e2ee_chat_request() accepts an httpx.Request and a verified model's public key. It returns the encrypted request and a matching JSON/SSE decryptor. The helper performs no requests or attestation verification. Keep the encrypted bytes, not reserialized plaintext, for response verification.

Create an AttestationClient with the Gateway API key once. Its asynchronous methods retrieve signatures and evidence; selection and verification are standalone functions. client.fetch_gateway_attestation() returns a FetchedGatewayAttestation with raw attestation and client_binding. By default, it requests the Gateway's SPKI fingerprint and the native implementation obtains the SHA-256 SPKI fingerprint from the TLS connection for that exact HTTPS request. Pass both values to verify_gateway_attestation; an attestation with an SPKI fingerprint requires the observed peer to match the fingerprint authenticated in the quote. A runtime without peer-certificate access can fetch with include_spki_fingerprint=False. That path requests no TLS fingerprint, verifies the signer-and-nonce quote layout, and returns GatewayTlsBinding(kind='none'). GatewayAttestation.spki_fingerprint is Gateway-reported, GatewayClientBinding.spki_fingerprint is client-observed, and a successful GatewayTlsBinding.spki_fingerprint is their verified match.

Direct endpoints

DirectInferenceClient and DirectAttestationClient connect to a model's own endpoint. They are experimental. Every supplied model report is verified; response verification returns the matching signer group. Direct Chat enables E2EE by default and supports optional OHTTP. Direct TLS fingerprint requests are temporarily disabled because the endpoint does not yet provide complete fleet coverage; normal HTTPS verification remains enabled.

Documentation

  • Verification guide describes the deployment-first workflow, policy configuration, completion signatures, and error handling.
  • API reference lists AttestationClient, verification functions, parameters, and result fields.

Errors

Handle retrieval and verification at separate call sites. AttestationClient and evidence-selection failures raise ApiError, including invalid helper input. Explicit verification functions raise VerificationError for local input, cryptographic, policy, and binding failures. Each handler has one SDK error type. Branch on its error.failure.code and inspect error.failure.details only when it is present; never parse the human-readable message.

error.retryable means a new attempt at the failed external operation may succeed. It does not mean that re-verifying the same evidence will succeed or that an inference should be replayed.

client.fetch_completion_signature() returns a completion signature or raises ApiError. A 2xx response that reports an unavailable signature raises api.completion_signature_unavailable; its details preserve the provider's error code and message.

InferenceClient.send() and verify_response() compose retrieval and verification, so either SDK error type can occur. The OpenAI Chat interface preserves OpenAI's exception behavior: a transport failure is wrapped in openai.APIConnectionError, with the original failure in __cause__.

Development checks

From this directory:

uv sync
make lint
uv build

The test suite is deterministic and uses local fixtures; it does not contact the NEAR AI Cloud Gateway, Intel PCCS, or NVIDIA NRAS.

Metadata

Release files for nearai-inference-sdk 0.1.0rc1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nearai-inference-sdk 0.1.0rc1
File Size Uploaded
nearai_inference_sdk-0.1.0rc1.tar.gz 53.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nearai-inference-sdk 0.1.0rc1
File Interpreter ABI Platform
nearai_inference_sdk-0.1.0rc1-py3-none-any.whl Python 3 none any Details

Total release size: 128.2 kB

Release files / nearai_inference_sdk-0.1.0rc1.tar.gz

Download URL nearai_inference_sdk-0.1.0rc1.tar.gz
Size 53.8 kB
Tags Source
SHA-256 checksum
How to use checksums
0a71b8344d2151c4ea4b5cbf9c7f214c4ef6845e2fa180a5098e4e3f14ef438e
BLAKE2b-256 checksum
How to use checksums
56476e81514cce41246e6d748b89de2d806758fdf577734683c48e3f0e739b77
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / nearai_inference_sdk-0.1.0rc1-py3-none-any.whl

Download URL nearai_inference_sdk-0.1.0rc1-py3-none-any.whl
Size 74.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8dddcfc1ed87eb28e57acfefd034d6ec8aad97fbd587e34bc496de320ca26b3e
BLAKE2b-256 checksum
How to use checksums
f8b06537e16b72367ca20e2add1fff531dd5def90d1d81ff1f302f8d172bacda
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.0

2 release files

This release

0.1.0rc1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page