NEAR AI Inference SDK for Python
nearai-inference-sdk provides OpenAI-compatible Chat Completions with Gateway
and model attestation verification, response signature verification, and optional
encryption. Use the asynchronous InferenceClient to get started.
Install
Requires Python 3.12 or later. Clients use asyncio; there is no synchronous
Chat client. Install with pip:
pip install nearai-inference-sdk==0.1.0
Send and verify a Chat completion
Set your NEAR AI Cloud API key in your server environment:
export NEARAI_API_KEY='your-api-key'
Save this example as chat.py and run it with python chat.py:
import asyncio
import os
from nearai_inference_sdk import InferenceClient
async def main() -> None:
async with InferenceClient(
os.environ['NEARAI_API_KEY'],
e2ee=True,
) as client:
completion = await client.chat.completions.create(
model='z-ai/glm-5.3-flash',
messages=[{'role': 'user', 'content': 'Hello!'}],
)
# Verify the completion signature before using the answer.
verified = await client.verify_response(completion.id)
print(f'Verified {verified.signature_kind} response')
print(completion.choices[0].message.content or '')
if __name__ == '__main__':
asyncio.run(main())
The client connects to the NEAR AI Cloud Gateway by default. This example enables E2EE, which encrypts supported Chat fields to an attested model and decrypts its response. Choose a model that supports NEAR model attestation; E2EE rejects unsupported models before sending Chat.
Gateway TLS identity verification is enabled by default, and subsequent requests
are pinned to that identity. Reuse a client across requests. The async with
block closes its connections and clears retained records; outside a context
manager, call await client.aclose() when finished.
When verification happens
- Before sending Chat: the client verifies Gateway evidence and checks the model's attestation support. Supported NEAR TEE models require every returned model report to pass. Failed required checks stop the request.
- After receiving the response: call
await client.verify_response(id)to verify its signature. The client retains the exact request and response bytes and preflight evidence automatically, including encrypted bytes when E2EE is enabled. Applications do not need to capture those bytes themselves. - For streaming: consume the entire stream before calling
verify_response(). Content displayed before that call succeeds is not yet signature-verified. See the streaming example.
Optionally call await client.verify(model) before the first Chat. It sends no
Chat request and shares Chat's verification cache and in-flight work. In this
version it returns None on success or raises on failure; it does not verify a
particular reply.
Successful deployment verification is cached for 60 minutes per model by default.
Set attestation_cache_time_to_live_ms=0 to verify before every request. Response
records have a separate 60-minute lifetime starting when the body finishes;
verify replies before records expire and before closing the client. See
caching and preflight.
Verification and encryption
These defaults apply to the Gateway InferenceClient:
| Feature | Purpose | Default |
|---|---|---|
| Attestation | Verify Gateway and supported model deployments | Enabled |
| Gateway TLS binding | Pin requests to the attested Gateway identity | Enabled |
| Response verification | Verify exact request and response bytes | Explicit verify_response() call |
| E2EE | Encrypt supported Chat fields to the attested model | Disabled; set e2ee=True |
| OHTTP | Encrypt the Chat HTTP exchange to the attested Gateway | Disabled; set ohttp=True |
Ed25519 is the default signing algorithm; ECDSA is also supported. OHTTP requires Ed25519 and is independent of field-level E2EE. See OHTTP configuration.
Models without supported NEAR model attestation use Incognito mode: Gateway
verification only. E2EE and configured model deployment policies reject this
mode. A provider_tee signature binds the exact bytes to an attested model
signer; a gateway signature binds them to an attested Gateway signer and does
not prove model TEE execution. These checks do not establish a complete chain
through Gateway transformations. See the
evidence boundary.
Attestation does not approve particular deployments by default. Supply deployment policies or image provenance policies to enforce your application's approval criteria. See policy and trust roots and image build provenance.
Use the OpenAI SDK
Pass inference_client.http_client to openai.AsyncOpenAI to use the same verified
transport for JSON and streaming Chat. Configure credentials on InferenceClient
and verify replies through it. See the
OpenAI integration example.
Only Chat Completions are supported; the Responses API is not supported.
Advanced APIs
- Manual verification:
AttestationClientretrieves evidence and signatures. With standalone functions, your application verifies every model candidate and retains the exact wire bytes. - Standalone E2EE:
prepare_e2ee_chat_request()encrypts a raw HTTPX Chat request and supplies its JSON/SSE decryptor. Key verification and response signature verification remain the application's responsibility. - Direct endpoints (experimental):
DirectInferenceClientandDirectAttestationClientconnect without Gateway verification. Direct TLS fingerprint requests are disabled pending complete fleet coverage; normal HTTPS verification remains enabled. Use Gateway clients for production. See direct-endpoint limitations.
Documentation
- Verification guide: streaming, encryption, proxies, policies, TLS binding, and error handling.
- API reference: client options, public functions, defaults, and result fields.
- Runnable Python examples: Gateway and direct clients, standalone verification, and OpenAI integration.
- Error handling: SDK errors, OpenAI exception wrapping, and receipt-only retries.
Development checks
From py/, run uv sync and make lint for formatting, lint, and deterministic
unit tests; uv build creates the distributions. Live service checks are separate:
see the E2E guide.
Metadata
Release files for nearai-inference-sdk 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nearai_inference_sdk-0.1.0.tar.gz | 53.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nearai_inference_sdk-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 127.2 kB
Release files / nearai_inference_sdk-0.1.0.tar.gz
| Download URL | nearai_inference_sdk-0.1.0.tar.gz |
|---|---|
| Size | 53.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5274e9ad4d3fc378569fbc44495b9e3cd75cd1717ae44cde92a4d7876af8788e
|
|
BLAKE2b-256 checksum How to use checksums |
2ee2a3eb79c80b687bfc6c153d11c3cf4930cbc6d5ec9c539d02df531fa67c03
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / nearai_inference_sdk-0.1.0-py3-none-any.whl
| Download URL | nearai_inference_sdk-0.1.0-py3-none-any.whl |
|---|---|
| Size | 73.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8c46752d3e0abe837da3fa795f457f05f713fb1a12a81029424d40718a222089
|
|
BLAKE2b-256 checksum How to use checksums |
596f47b90487c6db22748301f9d450b05f370fb7ad1af3e35abb42da5467c95e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log