sstudio Python SDK
Python client and command-line interface for model lifecycle APIs.
Table of Contents
- Requirements
- Installation
- Quick Start
- Authentication
- Client Configuration
- Return Values and Errors
- Module Overview
- Module Examples
- Detailed Method Reference
- Waiters
- CLI
- Retry and Security
Requirements
- Python 3.8+
- An API key
- An API base URL
Installation
pip install sstudio
Pin a specific version in production:
pip install sstudio
Quick Start
from sstudio import SmartStudioClient
with SmartStudioClient(
api_key="sk-your-api-key",
base_url="https://api.example.com",
) as client:
me = client.me()
models = client.models.list(page=1, page_size=20)
print(me)
print(models)
The context manager closes the underlying HTTP client automatically.
Authentication
Only the API key and API base URL are required. The SDK manages the two-step authentication flow internally:
| Step | Behavior |
|---|---|
| Login | Exchanges the API key for a short-lived bearer token. |
| Requests | Adds the bearer token to authenticated API requests. |
| Refresh | Refreshes authentication and retries a safe request once after a 401. |
The API key is never sent to pre-signed object storage URLs.
Client Configuration
client = SmartStudioClient(
api_key="sk-your-api-key",
base_url="https://api.example.com",
timeout=30.0,
max_retries=2,
default_headers={"X-Request-Source": "integration"},
on_token_refresh=lambda token: save_token(token),
)
| Option | Required | Description |
|---|---|---|
api_key |
Yes | API key exchanged for a bearer token. |
base_url |
Yes* | Environment origin, /msp-api base, or Console URL. A bare origin and /msp-console URL are normalized to /msp-api. |
token |
No | Existing bearer token managed by the host application. |
timeout |
No | Request timeout in seconds. Default: 30. |
max_retries |
No | Safe-request retry count. Default: 2. |
default_headers |
No | Additional string headers for API requests. |
on_token_refresh |
No | Callback invoked when a new bearer token is issued. |
base_url can also be configured with:
export SSTUDIO_PLATFORM__API_ENDPOINT=https://api.example.com
The CLI additionally accepts SSTUDIO_PLATFORM__API_KEY. Constructor arguments
take precedence, and the SDK has no implicit network endpoint.
Return Values and Errors
Resource methods return the response envelope's data value:
{
"code": 200,
"message": "success",
"data": {"id": 42, "status": "RUNNING"}
}
For this response, client.deployments.get(42) returns the object inside
data. Methods without a result return None; text endpoints such as
client.deployments.yaml(42) return str. Use client.request(...) only when
the complete envelope is required.
from sstudio import APIError, AuthenticationError, ValidationError
try:
deployment = client.deployments.get(42)
except AuthenticationError:
print("The API key or session is invalid")
except ValidationError as exc:
print(exc.error_code, exc.details)
except APIError as exc:
print(exc.message)
API exceptions preserve code, message, error_code, and details when
provided by the server.
Module Overview
| Module | Accessor | Typical operations |
|---|---|---|
| Identity | client.me() |
Current authenticated user |
| Clusters | client.clusters |
Cluster list, capacity, history, resource holds |
| Models | client.models |
Model catalog and deployment capabilities |
| My Models | client.my_models |
End-to-end model upload and asset lifecycle |
| Datasets | client.datasets |
End-to-end upload, list, preview, download, lifecycle |
| AI Dataset | client.datasets.preparations |
Preparation, labeling, resource preview, download |
| Deployments | client.deployments |
Preflight, create, lifecycle, events, YAML |
| Training | client.training |
Preflight, create, wait, cancel, artifacts |
| Evaluations | client.evaluations |
Preflight, create, wait, reports, artifacts |
| Jobs | client.jobs |
Generic job status, logs, wait, cancel |
| API Keys | client.keys |
Platform API key lifecycle |
| Provider Keys | client.provider_keys |
BYOK provider credentials and connectivity tests |
| Usage | client.usage |
Usage records, summaries, trends |
| Observability | client.observability |
Cluster, service, and workload metrics |
Complete method signatures, HTTP paths, CLI commands, and response types are
maintained in sdk/API_REFERENCE.md in the source repository.
Module Examples
The following examples assume an initialized client.
Identity
me = client.me()
Clusters
clusters = client.clusters.list()
capacity = client.clusters.capacity(1)
history = client.clusters.timeseries(1, limit=60)
holds = client.clusters.holds(cluster_id=1)
Models
models = client.models.list(keyword="Qwen", page=1, page_size=20)
model = client.models.get("qwen3-4b")
capabilities = client.models.deployment_capabilities(
"qwen3-4b",
cluster_id="1",
)
My Models
upload(...) validates model files, hashes the directory, transfers objects,
completes the upload, and returns the registered model asset.
model = client.my_models.upload(
"./model",
name="example-model",
model_type="LLM",
)
model_id = model["id"]
page = client.my_models.list(page=1, page_size=20)
detail = client.my_models.get(model_id)
client.my_models.update(model_id, {"description": "production candidate"})
Datasets
dataset = client.datasets.upload(
"./train.jsonl",
name="example-training-dataset",
dataset_type="training",
training_category="sft-llm",
)
dataset_id = dataset["id"]
preview = client.datasets.preview(dataset_id, limit=20)
page = client.datasets.list(page_num=1, page_size=20)
download = client.datasets.download_url({"datasetId": dataset_id, "fileIndex": 0})
AI Dataset Preparations
clusters = client.datasets.preparations.cluster_options()
provider_keys = client.datasets.preparations.available_provider_keys()
tasks = client.datasets.preparations.list({"pageNum": 1, "pageSize": 20})
task = client.datasets.preparations.get(id=123)
Create and action payloads are passed as dictionaries matching the API contract:
preview = client.datasets.preparations.preview(resource_preview_request)
task = client.datasets.preparations.create(preparation_request)
client.datasets.preparations.generate_rules(id=task["id"])
rules_ready = client.wait_for_dataset_preparation(
task["id"], until="rules_ready"
)
client.datasets.preparations.start_labeling(labeling_request)
completed = client.wait_for_dataset_preparation(task["id"])
Deployments
The workflow automatically selects the Model Profile or model-asset creation route from the request's model reference.
request = {
"name": "example-deployment",
"modelSource": "gallery",
"modelName": "Qwen3-4B-Instruct-2507-FAST",
"backend": "sglang",
"servingMode": "standard",
"gpuType": "L20",
"replicas": 1,
"clusterId": 1,
}
preview = client.deployments.preview(request)
if preview.get("creatable"):
deployment = client.deployments.create(request)
running = client.wait_for_deployment(deployment["id"])
Lifecycle and diagnostics:
page = client.deployments.list(page=1, page_size=20)
events = client.deployments.events(42)
rendered_yaml = client.deployments.yaml(42)
client.deployments.stop(42)
client.deployments.preview_restart(42)
client.deployments.restart(42)
Training
request = {
"clientToken": "training-request-001",
"displayName": "example-training",
"outputModelName": "example-output",
"recipeId": "recipe-id",
"recipeVersion": "recipe-version",
"baseModelRef": {"type": "recipe_model", "id": "model-id"},
"datasetRefs": [{"datasetId": "dataset-id", "role": "train"}],
"placement": {"clusterId": "1", "resourceSpecId": "resource-spec-id"},
"params": {},
}
client.training.preview({"clusterId": 1, "resourceSpecId": "resource-spec-id"})
job = client.training.create(request)
completed = client.wait_for_training_job(job["jobId"])
artifact = client.training.artifact_download_url(
completed["artifacts"][0]["artifactId"]
)
Evaluations
request = {
"kind": "benchmark",
"modelType": "LLM",
"models": [{"type": "external", "provider_key_id": "1", "model_id": "model-id"}],
"dataset": "evaluation-dataset",
"maxSamples": 100,
"clusterId": 1,
}
client.evaluations.preview({"clusterId": 1})
job = client.evaluations.create(request)
client.wait_for_evaluation_job(job["jobId"])
report = client.evaluations.report(job["jobId"])
Jobs
jobs = client.jobs.list(type="train", status="RUNNING", page=1, page_size=20)
job = client.jobs.get("job-id")
logs = client.jobs.logs("job-id", tail=200)
client.wait_for_job("job-id")
API Keys
keys = client.keys.list()
created = client.keys.create({
"keyValue": "sk-created-by-caller",
"description": "automation key",
"isActive": True,
})
key_id = str(created["id"])
plaintext = client.keys.reveal(key_id)
client.keys.update(key_id, {"description": "renamed key", "isActive": True})
Treat the value returned by reveal(...) as a secret and never log it.
Provider Keys
providers = client.provider_keys.providers()
keys = client.provider_keys.list()
created = client.provider_keys.create({
"provider": "example-provider",
"apiKey": "provider-api-key",
"description": "integration credential",
})
result = client.provider_keys.test(created["id"])
client.provider_keys.change_status(created["id"], {"status": 1})
Usage
start_seconds = 1782864000
end_seconds = 1785542400
summary = client.usage.summary(start_date=start_seconds, end_date=end_seconds)
records = client.usage.list(
start_date=start_seconds,
end_date=end_seconds,
page=1,
page_size=20,
)
trend = client.usage.trend(
start_date=start_seconds,
end_date=end_seconds,
granularity="day",
)
Observability
overview = client.observability.cluster_overview(1)
services = client.observability.services(cluster_id=1)
snapshot = client.observability.service_snapshot(42)
series = client.observability.service_timeseries(
42,
range_hours=1,
max_points=120,
)
Detailed Method Reference
This section documents every public Python SDK method. Resource methods return
the response envelope's data value unless the method explicitly says it
returns raw text or an HTTP response.
Response Conventions
| Shape | Fields | Meaning |
|---|---|---|
PageResult[T] |
items, total, page, pageSize, totalPages |
A 1-based page of resources. |
DownloadUrl |
url, expiration |
A temporary object download URL and its expiration time. |
TrainingArtifactDownload |
urls, optional files; legacy url may also appear |
Temporary URLs for one or more Training artifact files. |
WorkloadAdmissionPreviewVO |
clusterId, decision, maxSchedulableReplicas, reasonCodes, reasons, warningCodes, warnings, capacity, snapshotAt |
Point-in-time resource preflight; it does not reserve capacity. |
WorkloadClusterOptionVO |
id, name, status, selectable, reasonCodes, reasons, warnings, capacity |
A cluster option and why it can or cannot be selected. |
None |
no fields | The server accepted a command that has no result payload. |
Unknown response fields are preserved in the returned dictionary. The field lists below identify the stable fields callers normally consume; they do not discard additional server fields.
Core Client
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
client.me() |
Read the current authenticated identity. | None. | Identity dictionary from /api/v1/auth/me. |
client.request(method, path, ...) |
Send an advanced authenticated JSON request. | method: str, path: str; optional json, params, timeout. |
Complete {code, message, errorCode, details, data} envelope. |
client.request_raw(method, path, ...) |
Send an advanced authenticated text or file request. | method, path; optional json, content, content_type, params, timeout. |
httpx.Response; caller chooses text, bytes, or streaming access. |
client.get_token() |
Ensure authentication and return the current bearer token. | None. | str. Treat as a secret. |
client.close() |
Release the underlying HTTP connection pool. | None. | None. |
me = client.me()
envelope = client.request("GET", "/api/v1/models", params={"page": 1})
raw = client.request_raw("GET", "/api/v1/services/42/yaml")
token = client.get_token()
client.close()
Use resource methods instead of request(...) for normal integrations. The
generic methods exist for forward-compatible access to a server endpoint that
has not yet been promoted into the public SDK surface.
Clusters
client.clusters exposes caller-visible cluster capacity. Capacity values are
admission snapshots, not reservations.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
list() |
List clusters visible to the caller. | None. | list[UserClusterCapacityVO]. |
capacity(cluster_id) |
List placement-level capacity in one cluster. | `cluster_id: int | str` (required). |
holds(cluster_id=None) |
List active workload resource holds, optionally by cluster. | `cluster_id: int | str |
timeseries(cluster_id, limit=None) |
Read recent capacity snapshots. | cluster_id required; `limit: int |
None`. |
Key response fields
| Type | Fields |
|---|---|
UserClusterCapacityVO |
clusterId, name, status, selectable, known, fresh, snapshotAt, entitled, held, quotaAvailable, deficit, nodeCount |
UserNodeCapacityVO |
placementAlias, entitled, held, quotaAvailable, deficit |
UserWorkloadHoldVO |
businessType, businessId, displayName, clusterId, state, requests, canStop, createdAt |
UserCapacitySnapshotVO |
snapshotTime, sourceEventType, entitled, held, quotaAvailable, deficit |
clusters = client.clusters.list()
nodes = client.clusters.capacity(cluster_id=1)
all_holds = client.clusters.holds()
cluster_holds = client.clusters.holds(cluster_id=1)
history = client.clusters.timeseries(cluster_id=1, limit=60)
Models
client.models reads the deployable model catalog. It does not create My Model
assets; use client.my_models for caller-owned assets.
Pass the catalog's modelId to get() and deployment_capabilities(). Pass
its modelName unchanged when creating a Deployment; do not construct either
value from the provider label.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
list(...) |
Search and page the model catalog. | Optional keyword, model_type, provider, source, page, page_size. |
PageResult[ModelInfoVO]. |
get(model_id) |
Read one catalog model. | model_id: str (required). |
ModelInfoVO. |
deployment_capabilities(model_id, cluster_id=None) |
Resolve supported deployment choices, optionally for a cluster. | model_id: str required; `cluster_id: str |
None`. |
Key response fields
| Type | Fields |
|---|---|
ModelInfoVO |
modelId, modelName, displayName, modelType, provider, source, modelSize, deployable, supportedGpuTypes, supportedServingModes, deploymentOptions |
ModelDeploymentCapabilitiesVO |
modelId, modelName, clusterId, gpuTypeOptions, recommendedGpuType, acceleratorPools, clusterGpuTypes |
page = client.models.list(
keyword="Qwen",
model_type="LLM",
provider="Qwen",
source="gallery",
page=1,
page_size=20,
)
model = client.models.get("qwen3-4b")
choices = client.models.deployment_capabilities(
"qwen3-4b",
cluster_id="1",
)
My Models
client.my_models lists caller-owned Upload and Trained model assets. Only
Upload assets can be updated or deleted; Training-produced assets are read-only. The
high-level upload(...) workflow is the supported way to upload a local model;
the internal authorization and multipart endpoints are intentionally hidden.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
upload(path, ...) |
Validate, hash, upload, register, and complete one model directory. | path required; optional name, description, model_type, model_architecture, parameter_size, serving_config_json, client_request_id, on_progress. |
Ready MyModelVO. |
list(...) |
Page the caller's model assets. | Optional page, page_size, keyword, status, source. |
PageResult[MyModelVO]. |
get(id) |
Read one model asset. | `id: int | str` required. |
update(id, body) |
Update Upload model metadata. | Upload asset id and body required. |
Updated MyModelVO; Trained assets return 409. |
delete(id) |
Delete an unreferenced Upload model asset. | Upload asset id required. |
None; Trained assets return 409. |
resumable_uploads() |
List interrupted uploads that can be resumed or cancelled. | None. | list[MyModelStorageV2TaskResponse]. |
cancel_upload(task_id) |
Cancel a resumable upload task. | task_id: str required. |
None. |
upload(...) parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
path |
path-like | Yes | - | Model directory containing root config.json, weights, and tokenizer files. |
name |
str |
No | Directory name | Asset display name. |
description |
str |
No | None |
Asset description. |
model_type |
str |
No | Detected | Model type such as LLM. |
model_architecture |
str |
No | Detected | Architecture override. |
parameter_size |
str |
No | Detected | Human-readable parameter size. |
serving_config_json |
str |
No | None |
Serialized serving metadata. |
client_request_id |
str |
No | Generated UUID | Idempotency identity for upload recovery; maximum 64 characters. |
on_progress |
callback | No | None |
Receives phase, file, completedBytes, and totalBytes. |
The uploader rejects symlinks, Git LFS pointer files, missing model control files, duplicate paths, and unsupported size/count limits before registration.
Update body fields: name, description, modelType,
modelArchitecture, parameterSize, and servingConfigJson are optional.
Key MyModelVO fields: id, name, source, status, state,
deployable, deploymentBlockReasonCode, modelType, modelArchitecture,
parameterSize, sizeBytes, fileCount, baseModelName, trainJobId,
createdAt, and updatedAt.
def progress(event: dict) -> None:
print(event["phase"], event["file"], event["completedBytes"])
asset = client.my_models.upload(
"./model",
name="example-model",
model_type="LLM",
on_progress=progress,
)
asset_id = asset["id"]
page = client.my_models.list(page=1, page_size=20, status="READY")
detail = client.my_models.get(asset_id)
updated = client.my_models.update(asset_id, {"description": "validated"})
tasks = client.my_models.resumable_uploads()
if tasks:
client.my_models.cancel_upload(tasks[0]["taskId"])
client.my_models.delete(asset_id)
Datasets
client.datasets manages uploaded training and evaluation datasets. The
high-level upload(...) workflow signs, transfers, commits, and registers all
local files in one call.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
upload(files, ...) |
Upload one file, a directory, or a sequence of files and create a Dataset. | files, name, dataset_type, training_category required; optional on_progress. |
Created DatasetVO. |
list(...) |
Search and page datasets. | Optional dataset_name, dataset_type, training_category, order_by, order_direction, page_num, page_size, scope. |
PageResult[DatasetVO]. |
get(id) |
Read one Dataset. | id required. |
DatasetVO. |
preview(id, limit=None, cursor=None) |
Preview parsed records without downloading the object. | id required; optional limit, cursor. |
DatasetPreviewVO. |
download_url(body) |
Create a temporary URL for one Dataset file. | Body requires datasetId and zero-based fileIndex. |
DownloadUrl. |
update(id, body) |
Update Dataset metadata or committed file references. | id and body required. |
Updated DatasetVO. |
delete(id) |
Delete an unreferenced Dataset. | id required. |
None. |
upload(...) parameters
| Parameter | Type | Required | Values / Default |
|---|---|---|---|
files |
path-like or sequence | Yes | File, recursively expanded directory, or file list. |
name |
str |
Yes | Non-empty Dataset name. |
dataset_type |
str |
Yes | training or evaluation. |
training_category |
str |
Yes | sft-llm, dpo-llm, sft-vlm, dpo-vlm, or cpt-llm. |
on_progress |
callback | No | Same progress dictionary as My Model upload. |
Update body fields: optional datasetName, type, trainingCategory,
and files. Key DatasetVO fields: id, datasetName, type,
trainingCategory, files, fileSizeTotal, owner, createdAt, updatedAt.
DatasetPreviewVO contains items, totalRecords, datasetRevision,
hasNext, and nextCursor.
dataset = client.datasets.upload(
"./train.jsonl",
name="example-training-dataset",
dataset_type="training",
training_category="sft-llm",
)
dataset_id = dataset["id"]
page = client.datasets.list(
dataset_name="example",
dataset_type=["training"],
training_category=["sft-llm"],
order_by="createdAt",
order_direction="DESC",
page_num=1,
page_size=20,
scope="self",
)
detail = client.datasets.get(dataset_id)
preview = client.datasets.preview(dataset_id, limit=20)
download = client.datasets.download_url({"datasetId": dataset_id, "fileIndex": 0})
updated = client.datasets.update(dataset_id, {"datasetName": "renamed-dataset"})
client.datasets.delete(dataset_id)
AI Dataset Preparations
client.datasets.preparations manages AI-assisted dataset preparation and
labeling tasks. Request bodies use server wire names exactly as shown below.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
cluster_options() |
List clusters selectable for preparation workloads. | None. | list[WorkloadClusterOptionVO]. |
available_provider_keys() |
List caller-owned provider keys suitable for labeling. | None. | PrepProviderKeysVO. |
preview(body) |
Preflight the selected cluster. | Body: clusterId (required). |
WorkloadAdmissionPreviewVO. |
create(body) |
Create a preparation task. | PrepTaskCreateRequest; required fields below. |
PrepTaskVO. |
list(body) |
Search and page preparation tasks. | PrepTaskQueryRequest. |
PageResult[PrepTaskVO]. |
get(id=...) |
Read task configuration, progress, and artifacts. | Keyword-only id required. |
PrepTaskDetailVO. |
update(id, body) |
Update an editable preparation task. | id and PrepTaskUpdateRequest required. |
Updated PrepTaskVO. |
generate_rules(id=...) |
Start labeling-rule generation. | Keyword-only id required. |
None. |
start_labeling(body) |
Start labeling with optional rule text. | preparationId required; optional labelRules. |
None. |
action(body) |
Cancel, retry, or delete a task. | preparationId and action (CANCEL, RETRY, DELETE) required. |
None. |
download_url(body) |
Create a temporary URL for a task artifact. | preparationId plus artifactPath or storageRef. |
DownloadUrl. |
Use client.wait_for_dataset_preparation(id, until="rules_ready") after rule
generation, and call it without until after labeling to wait for completed.
The rules_ready target also succeeds when the task has already advanced to a
later successful processing phase or completed between polls.
Create body fields
| Field | Type | Required | Constraints |
|---|---|---|---|
taskName |
str |
Yes | Maximum 255 characters; no control characters. |
clusterId |
int |
Yes | Selected cluster ID. |
modelType |
str |
Yes | llm or vlm. |
postTrainingMethod |
str |
Yes | sft or ref_distill. |
preparationMode |
str |
Yes | base or sss-bench. |
providerKeyIds |
list[int] |
Yes | Up to 50 provider-key IDs; may be empty when the mode does not require one. |
scenario |
str |
No | Scenario text, maximum 4000 characters. |
autoSplitPercent |
int |
No | 1..100. |
generatedDatasetFileLines |
int |
No | Requested generated record count. |
processingMode |
str |
No | auto or manual. |
unlabeledDataFiles / evaluationDataFiles |
list[DatasetFileItem] |
No | Up to 100 items each. |
tBenchConfig |
dict |
No | Benchmark-mode options. |
DatasetFileItem accepts fileName, filePath, fileSize, storageRef,
sourceType, and preparationGroup. Update accepts the same configuration
fields as create, but all are optional.
Key response fields: PrepTaskVO includes id, taskName, clusterId,
clusterName, status, progress, progressInfo, savedDatasetId,
savedDatasetName, error, resourcePreview, createdAt, and updatedAt.
PrepTaskDetailVO additionally contains configuration, selected keys, input
files, artifacts, rules, execution summary, and result summary.
clusters = client.datasets.preparations.cluster_options()
provider_keys = client.datasets.preparations.available_provider_keys()
resource = client.datasets.preparations.preview({"clusterId": 1})
task = client.datasets.preparations.create({
"taskName": "prepare-example-dataset",
"clusterId": 1,
"modelType": "llm",
"postTrainingMethod": "sft",
"preparationMode": "base",
"providerKeyIds": [],
})
task_id = task["id"]
page = client.datasets.preparations.list({
"taskName": "prepare-example",
"status": ["pending", "completed"],
"pageNum": 1,
"pageSize": 20,
})
detail = client.datasets.preparations.get(id=task_id)
updated = client.datasets.preparations.update(task_id, {"taskName": "new-name"})
client.datasets.preparations.generate_rules(id=task_id)
client.datasets.preparations.start_labeling({
"preparationId": task_id,
"labelRules": "Return one concise label.",
})
artifact = client.datasets.preparations.download_url({
"preparationId": task_id,
"artifactPath": "outputs/result.jsonl",
})
client.datasets.preparations.action({"preparationId": task_id, "action": "CANCEL"})
Deployments
client.deployments previews, creates, and manages model-serving workloads.
The SDK chooses the correct create/preview endpoint automatically: a body with
modelAssetId uses the asset route; otherwise it uses the Model Profile route.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
capabilities() |
Check whether Deployment mutations are currently allowed. | None. | InferenceServiceCapabilitiesVO with mutationsAllowed, unavailableReasonCode. |
features() |
Read Deployment feature flags. | None. | DeploymentFeaturesVO with clusterSelectionEnabled. |
cluster_options() |
List clusters selectable for serving. | None. | list[WorkloadClusterOptionVO]. |
model_inventory(cluster_id) |
Read Model Profile placement inventory for a cluster. | cluster_id required. |
ModelInventoryVO. |
preview(body) |
Run resource preflight using the same Profile/Asset routing as create. | Profile or Asset create body. | ModelServingResourcePreviewVO. |
create(body) |
Create a Profile, Upload, or Trained Model Deployment. | Profile or Asset create body. | InferenceServiceVO. |
list(...) |
Search and page Deployments. | Optional page, page_size, keyword, status. |
PageResult[InferenceServiceVO]. |
get(id) |
Read one Deployment. | id required. |
InferenceServiceVO. |
events(id) |
Read lifecycle events for troubleshooting. | id required. |
list[DeploymentEventVO]. |
yaml(id) |
Read the rendered Kubernetes YAML. | id required. |
Raw str, not an API envelope. |
update(id, body) |
Update mutable serving configuration. | id and UpdateServiceRequest required. |
Updated InferenceServiceVO. |
stop(id) |
Stop the running workload while retaining its record. | id required. |
Updated InferenceServiceVO. |
preview_restart(id) |
Check restart blockers and current resource fit. | id required. |
DeploymentRestartPreviewVO. |
restart(id) |
Recreate a stopped or failed workload. | id required. |
Updated InferenceServiceVO. |
delete(id) |
Delete the Deployment and its managed workload resources. | id required. |
None. |
Common create fields
| Field | Type | Required | Description |
|---|---|---|---|
name |
str |
Yes | Unique Deployment name. |
clusterId |
int |
Environment-dependent | Selected cluster. Use cluster_options(); required when cluster selection is enabled. |
backend |
str |
Yes | Serving backend, normally sglang or vllm. |
servingMode |
str |
No | Usually standard or disaggregated; supported values come from model capabilities. |
gpuType |
str |
Yes | Selected GPU type from capabilities. |
replicas |
int |
No | Standard-mode replica count. |
tensorParallel |
int |
No | Standard-mode tensor parallel size. |
prefillReplicas / decodeReplicas |
int |
No | Disaggregated replica counts. |
prefillTensorParallel / decodeTensorParallel |
int |
No | Disaggregated tensor parallel sizes. |
acceleratorSelection |
dict |
No | resourceName, productLabelKey, productLabelValues. |
deploymentProfileKey |
str |
No | Explicit deployment profile selected from model capabilities. |
kvCacheDistributed, mtpEnabled, kvFP8Enabled, operatorOptimizationEnabled |
bool |
No | Optional runtime features supported by the model profile. |
kvCacheMaxCapacityGB, rateLimit |
int |
No | Optional KV-cache capacity and QPS limit. |
Profile creation additionally requires modelSource and modelName; it may
include modelRevision. The machine-to-machine path may also supply
modelOssPath and ossCredential, but normal SDK users should not place cloud
credentials in Deployment requests. Asset creation requires modelAssetId
instead of modelSource/modelName.
Update behavior: Runtime V2 Deployments support metadata-only updates; in
the current contract only description can be changed in place. Changes to
rateLimit, GPU, serving mode, replicas, parallelism, or runtime features
return DEPLOYMENT_V2_NEW_GENERATION_REQUIRED and require creating a new
Deployment generation.
Key response fields
| Type | Fields |
|---|---|
ModelServingResourcePreviewVO |
clusterId, clusterName, creatable, blockerCodes, blockers, warnings, workloadPlan, preflightV2, acceleratorPool, runtimeBackend |
InferenceServiceVO |
id, name, status, statusMessage, modelName, modelSource, backend, servingMode, clusterId, clusterName, endpoint, readyReplicas, totalReplicas, capabilities |
DeploymentRestartPreviewVO |
deploymentId, restartable, blockerCode, blockerMessage, resourcePreview |
DeploymentEventVO |
id, serviceId, eventType, message, details, createdAt |
profile_request = {
"name": "example-profile-deployment",
"clusterId": 1,
"modelSource": "gallery",
"modelName": "Qwen3-4B-Instruct-2507-FAST",
"backend": "sglang",
"servingMode": "standard",
"gpuType": "L20",
"replicas": 1,
}
asset_request = {
"name": "example-asset-deployment",
"clusterId": 1,
"modelAssetId": 42,
"backend": "sglang",
"servingMode": "standard",
"gpuType": "L20",
"replicas": 1,
}
features = client.deployments.features()
mutation_state = client.deployments.capabilities()
clusters = client.deployments.cluster_options()
inventory = client.deployments.model_inventory(cluster_id=1)
preview = client.deployments.preview(profile_request)
if preview["creatable"]:
created = client.deployments.create(profile_request)
asset_preview = client.deployments.preview(asset_request)
asset_deployment = client.deployments.create(asset_request)
page = client.deployments.list(page=1, page_size=20, status="RUNNING")
detail = client.deployments.get(asset_deployment["id"])
events = client.deployments.events(asset_deployment["id"])
rendered_yaml = client.deployments.yaml(asset_deployment["id"])
updated = client.deployments.update(asset_deployment["id"], {"description": "validated"})
stopped = client.deployments.stop(asset_deployment["id"])
restart_check = client.deployments.preview_restart(asset_deployment["id"])
if restart_check["restartable"]:
restarted = client.deployments.restart(asset_deployment["id"])
client.deployments.delete(asset_deployment["id"])
Training
client.training manages fine-tuning jobs. Inputs are logical model, Dataset,
recipe, and placement references; raw paths, images, commands, environments,
and credentials are not accepted by the create API.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
capabilities() |
Read the active recipe/resource catalog. | None. | TrainCapabilitiesVO. |
cluster_options() |
List clusters selectable for Training. | None. | list[WorkloadClusterOptionVO]. |
preview(body) |
Preflight a resource specification in a cluster. | clusterId, resourceSpecId required. |
WorkloadAdmissionPreviewVO. |
knowledge_teacher_models(...) |
Find compatible knowledge-distillation teachers. | Keyword-only student_model_id, recipe_id required. |
list[TrainKnowledgeTeacherVO]. |
create(body) |
Create a fine-tuning job. | CreateTrainJobRequest fields below. |
{jobId}. |
list(body=None) |
Page Training jobs. | Optional body: pageNum (default 1), pageSize (default 20), status. |
PageResult[TrainJobDetailVO]. |
get(id) |
Read job progress, losses, actions, and artifacts. | id: str required. |
TrainJobDetailVO. |
artifact_download_url(artifact_id) |
Create temporary download URLs for a completed artifact. | artifact_id: str required. |
{"urls": [str], "files"?: [{"path", "sizeBytes", "url"}]}. |
cancel(id) |
Request cancellation of a non-terminal job. | id: str required. |
None. |
Create body fields
| Field | Type | Description |
|---|---|---|
clientToken |
str |
Idempotency token for safe create retries. |
displayName |
str |
Human-facing job name. |
outputModelName |
str |
Output model suffix. Its maximum length depends on baseModelRef.id because the server creates <baseModelRef.id>-FT-<outputModelName> and keeps the result deployable; use a short suffix (the tested 4B Profile permits 14 characters). |
recipeId, recipeVersion |
str |
Recipe identity selected from capabilities(). |
baseModelRef |
dict |
{type, id} logical base-model reference. |
teacherRef |
dict |
Optional teacher reference; fields described below. |
datasetRefs |
list[dict] |
Each item uses {datasetId, role}. |
placement |
dict |
{clusterId, nodeId?, resourceSpecId}. |
params |
dict |
Recipe hyperparameters. |
teacherRef is the exception to the API's normal camelCase convention: its
wire keys are provider_key_id, model_id, service_id, and
model_asset_id, plus type. Do not send camelCase variants.
Key response fields: TrainCapabilitiesVO contains schemaVersion,
catalogVersion, catalogDigest, and entries. TrainJobDetailVO contains
id, status, stage, progress, taskDisplayName, baseModel,
trainingMethod, artifacts, actions, loss series, deployment references,
timestamps, and error/status details.
catalog = client.training.capabilities()
clusters = client.training.cluster_options()
resource = client.training.preview({
"clusterId": 1,
"resourceSpecId": "resource-spec-id",
})
teachers = client.training.knowledge_teacher_models(
student_model_id="model-id",
recipe_id="recipe-id",
)
request = {
"clientToken": "training-request-001",
"displayName": "example-training",
"outputModelName": "example-output",
"recipeId": "recipe-id",
"recipeVersion": "recipe-version",
"baseModelRef": {"type": "recipe_model", "id": "model-id"},
"datasetRefs": [{"datasetId": "dataset-id", "role": "train"}],
"placement": {"clusterId": "1", "resourceSpecId": "resource-spec-id"},
"params": {},
}
created = client.training.create(request)
job_id = created["jobId"]
page = client.training.list({"pageNum": 1, "pageSize": 20, "status": "RUNNING"})
detail = client.training.get(job_id)
if detail.get("artifacts"):
download = client.training.artifact_download_url(
detail["artifacts"][0]["artifactId"]
)
first_url = (download.get("urls") or [download["url"]])[0]
client.training.cancel(job_id)
Evaluations
client.evaluations creates benchmark, automatic, and comparison evaluations
against external models, API keys, deployed services, or model assets.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
available_models() |
List model references currently available to Evaluation. | None. | AvailableModelsVO. |
cluster_options() |
List clusters selectable for Evaluation. | None. | list[WorkloadClusterOptionVO]. |
preview(body) |
Preflight the selected cluster. | Body requires clusterId. |
WorkloadAdmissionPreviewVO. |
create(body) |
Create an Evaluation job. | CreateEvalJobRequest fields below. |
CreateEvalJobVO with jobId, resourcePreview. |
get(id) |
Read Evaluation status, scores, metrics, and report state. | id: str required. |
EvalJobDetailVO. |
report(id) |
Create a temporary URL for the completed report. | id: str required. |
DownloadUrl. |
artifact_download_url(id, artifact_ref) |
Create a temporary URL for one VLM/media artifact. | Job id and a server-provided artifact_ref required. |
DownloadUrl. |
Ordinary LLM evaluations expose their completed report through report(id).
artifact_download_url(...) is only usable when an Evaluation media result
provides an artifactRef; callers should not manufacture this identifier.
Create body fields
| Field | Type | Required | Constraints |
|---|---|---|---|
kind |
str |
Yes | benchmark, auto, or compare. |
modelType |
str |
Yes | LLM or VLM. |
models |
list[ModelRef] |
Yes | Up to two models. |
judge |
ModelRef |
No | Optional judge model. |
dataset |
str |
Yes | Dataset name/reference, maximum 128 characters. |
metricConfig |
dict |
No | Metric-specific options. |
maxSamples |
int |
No | Maximum 1,000,000. |
clusterId |
int |
Yes | Selected Evaluation cluster. |
ModelRef is either external (type="external", provider_key_id,
model_id) or a deployed service (type="service", msp_api_key_id,
service_id). Both IDs are required for the service form.
Key EvalJobDetailVO fields: jobId, status, progress, kind,
evaluationMethod, evaluationType, modelType, datasetName, clusterId,
scores, metrics, reportAvailable, error, and timestamps.
available = client.evaluations.available_models()
clusters = client.evaluations.cluster_options()
resource = client.evaluations.preview({"clusterId": 1})
created = client.evaluations.create({
"kind": "benchmark",
"modelType": "LLM",
"models": [{
"type": "service",
"msp_api_key_id": "7",
"service_id": "42",
}],
"dataset": "evaluation-dataset",
"maxSamples": 100,
"clusterId": 1,
})
job_id = created["jobId"]
detail = client.evaluations.get(job_id)
report = client.evaluations.report(job_id)
Jobs
client.jobs is the generic job-kernel view across long-running task types.
Use the domain-specific Training or Evaluation resource when you need fields
specific to that domain.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
list(...) |
Page jobs across domains. | Optional type (train or eval), status, page, page_size. |
PageResult[JobStatusView]. |
get(id) |
Read one generic job. | id: str required. |
JobStatusView. |
logs(id, tail=None) |
Read the newest log lines. | id required; optional tail. |
list[str]. |
cancel(id) |
Cancel a supported non-terminal job. | id: str required. |
None. |
JobStatusView includes id, jobType, status, progress, progressInfo,
metrics, artifacts, error, model/dataset/cluster context, and timestamps.
page = client.jobs.list(type="train", status="RUNNING", page=1, page_size=20)
job = client.jobs.get("job-id")
logs = client.jobs.logs("job-id", tail=200)
client.jobs.cancel("job-id")
API Keys
client.keys manages API keys used to authenticate this SDK and other platform
clients. Revealed plaintext values are secrets and must never be logged.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
list(owner_id=None) |
List API keys visible to the caller; admin callers may filter by owner. | Optional keyword-only owner_id. |
list[ApiKeyResponse]. |
create(body) |
Create an API key. | ApikeyRequest; keyValue is required, with optional description, isActive, expiresAt. |
ApiKeyResponse. |
reveal(id) |
Reveal a key's plaintext value when policy permits. | String id required. |
str. |
update(id, body) |
Update description, active state, or expiration. | String id and ApikeyRequest required. |
Updated ApiKeyResponse. |
delete(id) |
Delete an API key. | String id required. |
bool. |
ApikeyRequest supports id, keyValue, isActive, description,
expiresAt, and createdBy. The caller must generate and submit keyValue;
do not set ownership fields. ApiKeyResponse contains id,
isActive, description, expiresAt, createdBy, userId, masked
keyValue, and possibly one-time plaintextKey.
The response id is numeric while the key-management path accepts a string;
convert it with str(...) before follow-up calls.
keys = client.keys.list()
created = client.keys.create({
"keyValue": "sk-created-by-caller",
"description": "automation key",
"isActive": True,
})
key_id = str(created["id"])
plaintext = created.get("plaintextKey") or client.keys.reveal(key_id)
updated = client.keys.update(key_id, {
"description": "renamed key",
"isActive": True,
})
deleted = client.keys.delete(key_id)
Provider Keys
client.provider_keys manages caller-owned credentials for external model
providers. The SDK sends these values only to provider-key APIs.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
providers() |
List supported provider identifiers and labels. | None. | list[ProviderInfo]. |
list(owner_id=None) |
List masked provider keys; admin callers may filter by owner. | Optional keyword-only owner_id. |
list[ProviderKeyVO]. |
create(body) |
Store a new provider credential. | Required provider, apiKey; optional description. |
ProviderKeyVO. |
update(id, body) |
Rotate the credential or update its description. | id required; optional apiKey, description. |
Updated ProviderKeyVO. |
change_status(id, body) |
Enable or disable a provider key. | id and body {status} required. |
None. |
test(id) |
Test provider connectivity with the stored key. | id required. |
ConnectionTestResult. |
delete(id) |
Delete a provider key. | id required. |
None. |
ProviderInfo contains provider, label, and baseUrl. ProviderKeyVO
contains id, ownerId, provider, maskedApiKey, description, status,
available, unavailableReasonCode, and createdAt. Connection tests return
success and message. Use the provider identifier returned by
providers(), not its display label.
providers = client.provider_keys.providers()
keys = client.provider_keys.list()
created = client.provider_keys.create({
"provider": providers[0]["provider"],
"apiKey": "provider-api-key",
"description": "integration credential",
})
key_id = created["id"]
result = client.provider_keys.test(key_id)
updated = client.provider_keys.update(key_id, {"description": "rotated"})
client.provider_keys.change_status(key_id, {"status": 1})
client.provider_keys.delete(key_id)
Usage
client.usage reads metering data. start_date and end_date are Unix epoch
timestamps in seconds. For compatibility, 13-digit millisecond values are
accepted and converted to seconds before the request is sent.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
summary(...) |
Aggregate usage over a time range. | Required start_date, end_date; optional owner_id. |
Server-defined usage summary dict. |
list(...) |
Page individual usage records. | Required start_date, end_date; optional page, page_size, key_id, model_id, owner_id. |
Server-defined page/map. |
trend(...) |
Read time-bucketed usage. | Required start_date, end_date; optional granularity, key_id, owner_id. |
list[dict]. |
Usage maps are intentionally returned without a fixed client model because the metering backend may add dimensions. Known keys are preserved unchanged.
start_seconds = 1782864000
end_seconds = 1785542400
summary = client.usage.summary(start_date=start_seconds, end_date=end_seconds)
records = client.usage.list(
start_date=start_seconds,
end_date=end_seconds,
page=1,
page_size=20,
key_id=7,
model_id="model-id",
)
trend = client.usage.trend(
start_date=start_seconds,
end_date=end_seconds,
granularity="day",
key_id=7,
)
Observability
client.observability reads caller-scoped cluster, Deployment, and generic
workload metrics. Snapshot methods return the latest state; timeseries methods
return bounded historical points.
| Method | Purpose | Parameters | Returns |
|---|---|---|---|
cluster_overview(cluster_id) |
Read aggregate cluster workload utilization. | cluster_id required. |
UserClusterObservabilityOverviewVO. |
services(cluster_id=None) |
List observable model-serving Deployments. | Optional keyword-only cluster_id. |
list[UserWorkloadObservabilityVO]. |
service_snapshot(service_id) |
Read the newest snapshot for one Deployment. | service_id required. |
UserWorkloadObservabilityVO. |
service_timeseries(service_id, ...) |
Read Deployment engine/resource points. | service_id required; optional range_hours, max_points. |
list[UserServiceMetricsPointVO]. |
workloads(cluster_id=None) |
List observable workloads across supported domains. | Optional keyword-only cluster_id. |
list[UserWorkloadObservabilityVO]. |
workload_snapshot(business_type, business_id) |
Read one generic workload snapshot. | business_type, business_id required. |
UserWorkloadObservabilityVO. |
workload_timeseries(business_type, business_id, ...) |
Read generic CPU, memory, restart, and GPU points. | business_type, business_id required; optional range_hours, max_points. |
list[UserWorkloadMetricsPointVO]. |
history(...) |
Cursor-page workload lifecycle history. | Optional cluster_id, cursor, limit. |
UserWorkloadHistoryPageVO. |
Use the exact businessType and businessId returned by workloads() for
generic snapshot/timeseries calls rather than constructing identifiers.
Key response fields
| Type | Fields |
|---|---|
UserClusterObservabilityOverviewVO |
clusterId, activeWorkloads, cpuUsageCores, memoryWorkingSetBytes, observedPlacements, plannedPlacements, scope |
UserWorkloadObservabilityVO |
identity, lifecycle state, replicas/shards, CPU/memory/restarts, GPU allocation/utilization, serving QPS/latency/throughput, completeness, evidence |
UserServiceMetricsPointVO |
snapshotTime, replicas/shards, latency percentiles, QPS/request count, TTFT/ITL, error rates, throughput, KV-cache hit rate, input/output lengths |
UserWorkloadMetricsPointVO |
snapshotTime, CPU, working set/RSS memory, restart count, GPU activity/utilization/memory, completeness states |
UserWorkloadHistoryPageVO |
items, nextCursor, contractVersion, scope |
overview = client.observability.cluster_overview(cluster_id=1)
services = client.observability.services(cluster_id=1)
snapshot = client.observability.service_snapshot(service_id=42)
service_points = client.observability.service_timeseries(
service_id=42,
range_hours=1,
max_points=120,
)
workloads = client.observability.workloads(cluster_id=1)
if workloads:
workload = workloads[0]
latest = client.observability.workload_snapshot(
workload["businessType"],
workload["businessId"],
)
points = client.observability.workload_timeseries(
workload["businessType"],
workload["businessId"],
range_hours=1,
max_points=120,
)
history = client.observability.history(cluster_id=1, limit=100)
Waiters
Waiters poll until a terminal state, tolerate a bounded number of transient read errors, and raise on failure or timeout.
| Method | Success status | Failure statuses | Parameters | Returns |
|---|---|---|---|---|
wait_for_deployment(id, ...) |
RUNNING |
FAILED, STOPPED |
id; optional timeout=1800, interval=5, max_consecutive_errors=3. |
Final Deployment dictionary. |
wait_for_training_job(id, ...) |
SUCCEEDED |
FAILED, CANCELLED |
id; optional timeout=86400, interval=5, max_consecutive_errors=3. |
Final Training job dictionary. |
wait_for_evaluation_job(id, ...) |
SUCCEEDED |
FAILED, CANCELLED |
id; optional timeout=86400, interval=5, max_consecutive_errors=3. |
Final Evaluation job dictionary. |
wait_for_dataset_preparation(id, ...) |
Requested rules_ready milestone or completed |
ERROR, CANCELLED |
id; optional until="completed", timeout=86400, interval=5, max_consecutive_errors=3. |
Latest preparation dictionary. |
wait_for_job(id, ...) |
SUCCEEDED |
FAILED, CANCELLED |
id; optional timeout=86400, interval=5, max_consecutive_errors=3. |
Final generic job dictionary. |
OperationFailedError carries the terminal status and full result.
WaitTimeoutError reports the last observed status. Only transient connection
and timeout errors count toward max_consecutive_errors; API and validation
errors fail immediately.
Dataset Preparation uses lowercase wire statuses. Waiting for rules_ready
accepts that milestone and later successful phases; it never treats error or
cancelled as success.
deployment = client.wait_for_deployment(42, timeout=1800, interval=5)
training = client.wait_for_training_job("training-id", timeout=86400)
evaluation = client.wait_for_evaluation_job("evaluation-id", timeout=86400)
rules = client.wait_for_dataset_preparation(17, until="rules_ready")
prepared = client.wait_for_dataset_preparation(17)
job = client.wait_for_job("job-id", timeout=86400)
CLI
sstudio login --api-key sk-your-api-key --base-url https://api.example.com
sstudio whoami
sstudio --format json models list --page 1 --page-size 20
sstudio my-models upload --path ./model --name example-model --model-type LLM
sstudio datasets upload --file ./train.jsonl --name example-dataset \
--type training --training-category sft-llm
sstudio datasets-preparations wait --id 17 --until rules_ready
sstudio deployments preview --body @deployment.json
sstudio deployments create --body @deployment.json
sstudio deployments wait --id 42
sstudio training create --body @training.json
sstudio training wait --id training-id
sstudio evaluations create --body @evaluation.json
sstudio jobs logs --id job-id --tail 200
JSON bodies accept inline JSON, @file.json, or - for standard input. Use
--format json|yaml|table to select output format.
Retry and Security
- Retries apply to idempotent methods and writes carrying
clientToken,clientRequestId, oridempotencyKey. - A
401may refresh authentication and retry a safe request once; a403fails immediately. - Authenticated redirects and requests outside the configured API origin and path scope are rejected.
- Pre-signed object uploads never receive API credentials or default headers.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sstudio-0.0.6.tar.gz.
File metadata
- Download URL: sstudio-0.0.6.tar.gz
- Upload date:
- Size: 73.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bb58b09adfef5319a98395d79a505f4ed3334ccfc00c149da9ef277c4918322c
|
|
| MD5 |
6717625ca435c790e34dbffe8f7f2950
|
|
| BLAKE2b-256 |
a52f47279edbe687e48bf280ff48276def031cfae176fccbf049672146442799
|
File details
Details for the file sstudio-0.0.6-py3-none-any.whl.
File metadata
- Download URL: sstudio-0.0.6-py3-none-any.whl
- Upload date:
- Size: 78.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9bbfdc8338640337625c8cb9d9efe34d215258657a574df0b72f838772fe6596
|
|
| MD5 |
962d2734d982ee33cd8a7151b7185a8f
|
|
| BLAKE2b-256 |
d2881d3d080f4ca363cef9776706389afe0a7d2d391e81635ae890c438eaf7d0
|