ALCF AI Inference Services SDK
This package provides Python client and CLI tools to facilitate usage of the ALCF AI Inference services.
Command Line Usage
Quick Start
# Log in with Globus:
uvx alcf-ai auth login
# Chat with a model
# The default --model is meta-llama/Llama-4-Scout-17B-16E-Instruct
uvx alcf-ai chat "How do I know Pi is irrational? Be concise."
Auth
# Login for Inference Service only:
uvx alcf-ai auth login
# Login for Inference+Globus data transfers
# (append :data_access only if required for your collection)
SOURCE_COLLECTION="your globus collection UUID"
uvx alcf-ai auth login --authorize-transfers $SOURCE_COLLECTION:data_access
# Get an access token to use externally:
token=$(uvx alcf-ai auth get-access-token)
curl -H "Authorization: Bearer $token" https://inference-api.alcf.anl.gov/resource_server/list-endpoints | jq
Discovering Models
To list the models and corresponding API endpoints that are currently available, use:
uvx alcf-ai ls-endpoints
To view the status of models that are currently hot or starting up on a cluster, use:
# Can substitute "sophia" with "metis"
uvx alcf-ai ls-jobs sophia
Chat with an LLM
# See detailed options:
uvx alcf-ai chat --help
# For example:
uvx alcf-ai chat --model google/gemma-4-31B-it --stream --temp 0.3 --max-tokens 100 "What is KL divergence? Answer in less than 75 words."
Segment images with SAM3
You can segment your images with the Meta SAM3 model.
Send a single image URI plus prompt in for segmentation:
uvx alcf-ai sam3 submit-image \
https://raw.githubusercontent.com/masalim2/sam3-service/refs/heads/main/examples/images/groceries.jpg \
"Baguette" \
--save-preview ~/test-baguettes.png
Batch Processing
For high-throughput, preprocess and bundle your images and prompts in the WebDataset format using the built-in CLI tool:
# Bundle all .tiff files in directory with 3 prompts Creates WebDataset tar
# files in --output-dir, with 100 images per .tar.
alcf-ai sam3 create-webdataset \
/path/to/tiff-stack \
.tiff \
"Phloem Fibers" "Hydrated Xylem vessels" "Air-based Pith cells" \
--output-dir test-wds --shard-size=100 --num-workers=4
If the dataset is on a Globus Collection, you can authorize the CLI to send them to the inference service:
# Look up the UUID of your collection:
SOURCE_COLLECTION="your globus collection UUID"
# Append ":data_access" if this scope is required:
uvx alcf-ai auth login --authorize-transfers $SOURCE_COLLECTION:data_access
Then use the tool to drive data staging and batch inference:
SAM3_FINETUNE=/eagle/inference_service/sam3-service/weights/synaps-i
SECONDS=0
for f in test-wds/*.tar
do
uvx alcf-ai sam3 submit-batch $f --from-collection-id $SOURCE_COLLECTION --weights-dir-override $SAM3_FINETUNE >> batch-inference.log 2>&1 &
done
wait
echo "Completed in $SECONDS seconds."
You can preview the segmentation results in a batch by passing the paths to the input and result tar files:
uvx alcf-ai sam3 preview-batch-results shard-00004.tar shard-00004.results.tar
Segment images with DINOv3
You can also segment your images with a DINOv3 segmentation model. Unlike SAM3, DINOv3 works over a whole folder of images at once: you stage in a directory, the GPU dataloader batches over every image in it, and a folder of results (semantic masks, plus optional color overlays) is written back out.
Because the folder is the unit of work, the input/output paths are staged with Globus Transfer as recursive directory transfers. Folder transfers require a source/destination Globus collection (the HTTPS upload path is single-file only), so first authorize transfers against your collection:
# Look up the UUID of your collection:
SOURCE_COLLECTION="your globus collection UUID"
# Append ":data_access" if this scope is required:
uvx alcf-ai auth login --authorize-transfers $SOURCE_COLLECTION:data_access
Then submit a folder for segmentation with the CLI. It stages the folder in, runs inference, polls until complete, and stages the results folder back:
uvx alcf-ai dinov3 submit \
/path/to/image-folder \
--from-collection-id $SOURCE_COLLECTION \ # Stage the input folder in from here
--to-collection-id $SOURCE_COLLECTION \ # Send the results folder back here
--save-overlay # Also render color overlays
Sharding large datasets
The GPU dataloader batches over all images in a single folder, so one folder = one inference task. To parallelize a large dataset, shard it into subfolders and submit each concurrently. The SDK is preferred for driving the bulk transfers and concurrent inference tasks:
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from alcf_ai import InferenceClient
from alcf_ai.auth import STAGING_COLLECTION_ROOT
from rich import print
client = InferenceClient()
collection_id = "your globus collection UUID"
# A dataset pre-sharded into subfolders, e.g. dataset/shard-00000/, shard-00001/, ...
dataset_dir = Path("/path/to/dataset")
shards = sorted(p for p in dataset_dir.iterdir() if p.is_dir())
def run_inference(shard_dir: Path) -> dict:
"""Stage a folder in, run DINOv3 segmentation, and stage the results back."""
# Recursively stage the input folder in from the source collection:
stagein = client.stage_in(
shard_dir,
Path(shard_dir.name),
from_collection_id=collection_id,
recursive=True,
)
remote_input = STAGING_COLLECTION_ROOT + str(stagein.destination_path)
# Submit the inference request and poll for completion. The results
# directory is derived server-side within your staging area (you don't -- and
# can't -- choose it) and reported back as `mask_dir` in the result.
resp = client.dinov3.submit(input_dir=remote_input, save_overlay=True)
result = client.dinov3.poll_task_result(resp.task_id)
# Recursively stage the results folder back to the source collection. Its
# name comes from the service (mask_dir == <results_dir>/semantic_masks):
results_dirname = Path(result["mask_dir"]).parent.name
client.stage_out(
collection_id,
Path(results_dirname),
shard_dir.with_name(results_dirname),
recursive=True,
)
return result
with ThreadPoolExecutor(max_workers=8) as pool:
# Submit all stage_in / inference / stage_out pipelines to run in parallel:
futures = {pool.submit(run_inference, shard): shard for shard in shards}
for future in as_completed(futures):
shard = futures[future]
result = future.result()
print(f"[green]{shard.name}[/green] completed: {result}")
Installing the latest client version
You can force an install of the latest version and verify your local version using:
uvx alcf-ai@latest version
SDK Usage
You can use pip install alcf-ai or uv run --with-alcf python to add the SDK to your environment:
uv run --with alcf-ai python
OpenAI Client
Use alcf_ai.InferenceClient to construct an OpenAI client for any ALCF-backed
cluster. This reuses your auth and ensures that requests are sent to the right
URL:
from alcf_ai import InferenceClient
from rich import print
# Automatically uses cached refresh tokens from previous login:
client = InferenceClient()
# Programmatically discover endpoints:
print(client.list_endpoints()["clusters"]["sophia"])
# Get an OpenAI API client for an ALCF cluster:
oai = client.clusters("sophia").openai
print(
oai.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[{"role": "user", "content": "Hello there!"}],
)
)
Data Movement and SAM3
You can use the same InferenceClient to move data in and out of a Globus Guest
Collection that's managed by the service. Your data is stored in an ephemeral
staging subdirectory, with ACLs that grant only your Globus identity
read/write access to it.
from alcf_ai import InferenceClient
from alcf_ai.auth import STAGING_COLLECTION_ROOT
client = InferenceClient()
dataset_path = Path("/path/to/my-dataset.tar")
collection_id="globus collection uuid"
# Stage in data:
stagein = client.stage_in(collection_id, dataset_path, dataset_path.name)
# Submit SAM3 inference:
resp = client.sam3.submit_batch(
STAGING_COLLECTION_ROOT + str(stagein.destination_path)
)
# Wait for inference:
result = client.sam3.poll_task_result(resp.task_id)
# Copy results back:
client.stage_out(
collection_id,
Path(result.result_path).name,
dataset_path.with_suffix(".results.tar"),
)
Using an alternate service URL
The client with both programmatic and CLI usage defaults to the ALCF Inference Service production base url of https://inference-api.alcf.anl.gov/resource_server/. This can be altered in a few ways:
- By exporting the
inference_base_urlenvironment variable - From the CLI, passing an optional
--base-urlto thealcf-aisubcommand. - From the Python client, passing the kwarg
InferenceClient(base_url="...")
AmSC SUF-D3 Models
alcf-ai provides an SDK/CLI to submit SUF-D3 inference workloads to a Triton serving backend.
There are 7 models available in the current deployment:
- snbamsc_2dcnn_u
- snbamsc_2dcnn_v
- snbamsc_2dcnn_z
- DoubleMetricLearning
- higgsInteractionNet
- particlenet_AK4_PT
- nugraph2
Inference requests can be submitted using either the alcf-ai d3-triton submit CLI or the Python SDK.
Requests must provide one of the above model names along with a path to the input dataset stored in the Numpy .npz format, where each input tensor key is directly matched against the Triton model input metadata.
To orchestrate a single test inference task with dataset staging from a source Globus collection, the CLI is convenient. The following example invokes the snbamsc_2dcnn_u model together with the orchestration of Globus data stage-in and stage-out:
# ALCF Eagle DTN:
SOURCE_COLLECTION="05d2c76a-e867-4f67-aa57-76edeb0beda0"
uvx alcf-ai auth login --authorize-transfers $SOURCE_COLLECTION
alcf-ai d3-triton submit \
--from-collection-id $SOURCE_COLLECTION \ # Stage input in from this collection
--to-collection-id $SOURCE_COLLECTION \ # Send result back to the same collection
snbamsc_2dcnn_u \ # Select model
/datascience/msalim/test-staging-area/amsc-d3-triton/sample_inputs/snbamsc_2dcnn_u.npz # Input path
For efficient bulk processing, the SDK is preferred to arrange bulk transfer tasks and to submit concurrent inference tasks.
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from alcf_ai import InferenceClient
from alcf_ai.auth import STAGING_COLLECTION_ROOT
from rich import print
client = InferenceClient()
sample_dir = Path("/datascience/msalim/test-staging-area/amsc-d3-triton/sample_inputs")
models = [
"snbamsc_2dcnn_u",
"snbamsc_2dcnn_v",
"snbamsc_2dcnn_z",
"DoubleMetricLearning",
"higgsInteractionNet",
"particlenet_AK4_PT",
"nugraph2",
]
collection_id = "05d2c76a-e867-4f67-aa57-76edeb0beda0"
def run_inference(model_name: str, input_path: Path) -> dict:
"""Stage in an input file, run Triton inference, and stage out the result."""
# Stage the input file in from the source collection:
stagein = client.stage_in(
input_path, Path(input_path.name), from_collection_id=collection_id
)
remote_input = STAGING_COLLECTION_ROOT + str(stagein.destination_path)
remote_output = remote_input.rsplit(".", 1)[0] + ".output.npz"
# Submit the inference request and poll for completion:
resp = client.d3_triton.submit(
model_name=model_name,
input_path=remote_input,
output_path=remote_output,
)
result = client.d3_triton.poll_task_result(resp.task_id)
# Stage the result back to the source collection:
output_filename = Path(result["output_path"]).name
local_output = input_path.with_suffix(".output.npz")
client.stage_out(collection_id, Path(output_filename), local_output)
return result
with ThreadPoolExecutor(max_workers=8) as pool:
# Example: one input file per model
inputs = [sample_dir / f"{model}.npz" for model in models]
# Submit all stage_in / inference / stage_out pipelines to run in parallel:
futures = {
pool.submit(run_inference, model, input_path): model
for model, input_path in zip(models, inputs)
}
for future in as_completed(futures):
model = futures[future]
result = future.result()
print(f"[green]{model}[/green] completed: {result}")
Release files for alcf-ai 0.9.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| alcf_ai-0.9.0.tar.gz | 24.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| alcf_ai-0.9.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 55.3 kB
Release files / alcf_ai-0.9.0.tar.gz
| Download URL | alcf_ai-0.9.0.tar.gz |
|---|---|
| Size | 24.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b770397322ce19f483738ad648dcc5fa882603239a196fbf946896c1b0818539
|
|
BLAKE2b-256 checksum How to use checksums |
0e149b5a3aee3bb993f9bc8bd61b1552d7c5d72573e89264e44277bd5b8c22f7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.8 {"installer":{"name":"uv","version":"0.12.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / alcf_ai-0.9.0-py3-none-any.whl
| Download URL | alcf_ai-0.9.0-py3-none-any.whl |
|---|---|
| Size | 31.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4990aed0966aea7e60561f78866575c315da9568dd7e2f1cb1cd7d28aac5df65
|
|
BLAKE2b-256 checksum How to use checksums |
15590f23b7a632cd577f234cc04fc34a90ff4ef6177062a02d2b6007ee98055f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.8 {"installer":{"name":"uv","version":"0.12.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|