ALCF AI Inference Services SDK
This package provides Python client and CLI tools to facilitate usage of the ALCF AI Inference services.
Command Line Usage
Quick Start
# Log in with Globus:
uvx alcf-tokens login
# Chat with a model
# The default --model is meta-llama/Llama-4-Scout-17B-16E-Instruct
uvx alcf-ai chat "How do I know Pi is irrational? Be concise."
Auth
Logging in is handled by alcf-tokens,
the shared ALCF CLI: one interactive login issues tokens for several ALCF
services, and alcf-ai reads the tokens it cached. alcf-ai itself never
starts a login -- if no valid token is cached, it fails and prints the
alcf-tokens login command that would fix it.
# Interactive Globus login (caches a refresh token):
uvx alcf-tokens login
# Check that the gateway accepts your token:
uvx alcf-tokens test-token inference
# Get an access token to use externally:
token=$(uvx alcf-tokens get-token inference)
curl -H "Authorization: Bearer $token" https://inference-api.alcf.anl.gov/resource_server/list-endpoints | jq
If you will stage data in or out, authorize those collections in the same
login with --authorize-transfer (repeat the flag per collection). Each entry
is a collection UUID -- or a known alias, such as home, eagle or flare.
Append colon-separated scopes as needed: :data_access for collections that require it
for Transfer, and :https to read and write files directly over HTTPS.
STAGING=96c7390b-a3e8-4dd4-a327-1af7d143283e # IRIBeta inference_data_staging
SOURCE_COLLECTION="your globus collection UUID"
uvx alcf-tokens login \
--authorize-transfer $SOURCE_COLLECTION:data_access \
--authorize-transfer $STAGING:https
Re-running alcf-tokens login with more --authorize-transfer entries
re-consents with the wider set, so list every collection you want authorized in
the same command.
Discovering Models
To list the models and corresponding API endpoints that are currently available, use:
uvx alcf-ai ls-endpoints
To view the status of models that are currently hot or starting up on a cluster, use:
# Can substitute "sophia" with "metis"
uvx alcf-ai ls-jobs sophia
Chat with an LLM
# See detailed options:
uvx alcf-ai chat --help
# For example:
uvx alcf-ai chat --model google/gemma-4-31B-it --stream --temp 0.3 --max-tokens 100 "What is KL divergence? Answer in less than 75 words."
Segment images with SAM3
You can segment your images with the Meta SAM3 model.
Send a single image URI plus prompt in for segmentation:
uvx alcf-ai sam3 submit-image \
https://raw.githubusercontent.com/masalim2/sam3-service/refs/heads/main/examples/images/groceries.jpg \
"Baguette" \
--save-preview ~/test-baguettes.png
Batch Processing
For high-throughput, preprocess and bundle your images and prompts in the WebDataset format using the built-in CLI tool:
# Bundle all .tiff files in directory with 3 prompts Creates WebDataset tar
# files in --output-dir, with 100 images per .tar.
alcf-ai sam3 create-webdataset \
/path/to/tiff-stack \
.tiff \
"Phloem Fibers" "Hydrated Xylem vessels" "Air-based Pith cells" \
--output-dir test-wds --shard-size=100 --num-workers=4
If the dataset is on a Globus Collection, you can authorize the CLI to send them to the inference service:
# Look up the UUID of your collection:
SOURCE_COLLECTION="your globus collection UUID"
# Append ":data_access" if this scope is required:
uvx alcf-tokens login --authorize-transfer $SOURCE_COLLECTION:data_access
Then use the tool to drive data staging and batch inference:
SAM3_FINETUNE=/eagle/inference_service/sam3-service/weights/synaps-i
SECONDS=0
for f in test-wds/*.tar
do
uvx alcf-ai sam3 submit-batch $f --from-collection-id $SOURCE_COLLECTION --weights-dir-override $SAM3_FINETUNE >> batch-inference.log 2>&1 &
done
wait
echo "Completed in $SECONDS seconds."
You can preview the segmentation results in a batch by passing the paths to the input and result tar files:
uvx alcf-ai sam3 preview-batch-results shard-00004.tar shard-00004.results.tar
Segment images with DINOv3
You can also segment your images with a DINOv3 segmentation model. Unlike SAM3, DINOv3 works over a whole folder of images at once: you stage in a directory, the GPU dataloader batches over every image in it, and a folder of results (semantic masks, plus optional color overlays) is written back out.
Because the folder is the unit of work, the input/output paths are staged with Globus Transfer as recursive directory transfers. Folder transfers require a source/destination Globus collection (the HTTPS upload path is single-file only), so first authorize transfers against your collection:
# Look up the UUID of your collection:
SOURCE_COLLECTION="your globus collection UUID"
# Append ":data_access" if this scope is required:
uvx alcf-tokens login --authorize-transfer $SOURCE_COLLECTION:data_access
Then submit a folder for segmentation with the CLI. It stages the folder in, runs inference, polls until complete, and stages the results folder back:
uvx alcf-ai dinov3 submit \
/path/to/image-folder \
--from-collection-id $SOURCE_COLLECTION \ # Stage the input folder in from here
--to-collection-id $SOURCE_COLLECTION \ # Send the results folder back here
--save-overlay # Also render color overlays
Sharding large datasets
The GPU dataloader batches over all images in a single folder, so one folder = one inference task. To parallelize a large dataset, shard it into subfolders and submit each concurrently. The SDK is preferred for driving the bulk transfers and concurrent inference tasks:
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from alcf_ai import InferenceClient
from alcf_ai.transfer import STAGING_COLLECTION_ROOT
from rich import print
client = InferenceClient()
collection_id = "your globus collection UUID"
# A dataset pre-sharded into subfolders, e.g. dataset/shard-00000/, shard-00001/, ...
dataset_dir = Path("/path/to/dataset")
shards = sorted(p for p in dataset_dir.iterdir() if p.is_dir())
def run_inference(shard_dir: Path) -> dict:
"""Stage a folder in, run DINOv3 segmentation, and stage the results back."""
# Recursively stage the input folder in from the source collection:
stagein = client.stage_in(
shard_dir,
Path(shard_dir.name),
from_collection_id=collection_id,
recursive=True,
)
remote_input = STAGING_COLLECTION_ROOT + str(stagein.destination_path)
# Submit the inference request and poll for completion. The results
# directory is derived server-side within your staging area (you don't -- and
# can't -- choose it) and reported back as `mask_dir` in the result.
resp = client.dinov3.submit(input_dir=remote_input, save_overlay=True)
result = client.dinov3.poll_task_result(resp.task_id)
# Recursively stage the results folder back to the source collection. Its
# name comes from the service (mask_dir == <results_dir>/semantic_masks):
results_dirname = Path(result["mask_dir"]).parent.name
client.stage_out(
collection_id,
Path(results_dirname),
shard_dir.with_name(results_dirname),
recursive=True,
)
return result
with ThreadPoolExecutor(max_workers=8) as pool:
# Submit all stage_in / inference / stage_out pipelines to run in parallel:
futures = {pool.submit(run_inference, shard): shard for shard in shards}
for future in as_completed(futures):
shard = futures[future]
result = future.result()
print(f"[green]{shard.name}[/green] completed: {result}")
Installing the latest client version
You can force an install of the latest version and verify your local version using:
uvx alcf-ai@latest version
SDK Usage
You can use pip install alcf-ai or uv run --with-alcf python to add the SDK to your environment:
uv run --with alcf-ai python
OpenAI Client
Use alcf_ai.InferenceClient to construct an OpenAI client for any ALCF-backed
cluster. This reuses your auth and ensures that requests are sent to the right
URL:
from alcf_ai import InferenceClient
from rich import print
# Automatically uses the tokens cached by `alcf-tokens login`:
client = InferenceClient()
# Programmatically discover endpoints:
print(client.list_endpoints()["clusters"]["sophia"])
# Get an OpenAI API client for an ALCF cluster:
oai = client.clusters("sophia").openai
print(
oai.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[{"role": "user", "content": "Hello there!"}],
)
)
Data Movement and SAM3
You can use the same InferenceClient to move data in and out of a Globus Guest
Collection that's managed by the service. Your data is stored in an ephemeral
staging subdirectory, with ACLs that grant only your Globus identity
read/write access to it.
from alcf_ai import InferenceClient
from alcf_ai.transfer import STAGING_COLLECTION_ROOT
client = InferenceClient()
dataset_path = Path("/path/to/my-dataset.tar")
collection_id="globus collection uuid"
# Stage in data:
stagein = client.stage_in(collection_id, dataset_path, dataset_path.name)
# Submit SAM3 inference:
resp = client.sam3.submit_batch(
STAGING_COLLECTION_ROOT + str(stagein.destination_path)
)
# Wait for inference:
result = client.sam3.poll_task_result(resp.task_id)
# Copy results back:
client.stage_out(
collection_id,
Path(result.result_path).name,
dataset_path.with_suffix(".results.tar"),
)
Using an alternate service URL
The client with both programmatic and CLI usage defaults to the ALCF Inference Service production base url of https://inference-api.alcf.anl.gov/resource_server/. This can be altered in a few ways:
- By exporting the
inference_base_urlenvironment variable - From the CLI, passing an optional
--base-urlto thealcf-aisubcommand. - From the Python client, passing the kwarg
InferenceClient(base_url="...")
Release files for alcf-ai 0.11.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| alcf_ai-0.11.0.tar.gz | 21.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| alcf_ai-0.11.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 48.2 kB
Release files / alcf_ai-0.11.0.tar.gz
| Download URL | alcf_ai-0.11.0.tar.gz |
|---|---|
| Size | 21.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a5c388272ba9597344dfd2c4a3bd19e3bd3bbbd84a957d264ccd0963fd6cc422
|
|
BLAKE2b-256 checksum How to use checksums |
b6e94ac218e976880194fb24357d9dd44722a9835fabdb270ffd65ed3931a39e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / alcf_ai-0.11.0-py3-none-any.whl
| Download URL | alcf_ai-0.11.0-py3-none-any.whl |
|---|---|
| Size | 26.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0a879b0b1c236252f14ad408c00ba134742c9831311be477f3d24a39c8a3e2fa
|
|
BLAKE2b-256 checksum How to use checksums |
dba7c61c906411cd6507987259cfb3a1eefe04b6f1ced76562cdf5b5417137d3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|