PyLet
A simple distributed task execution system for GPU servers. Like Ray/K8s but much simpler.
Install
pip install pylet
For development:
git clone https://github.com/ServerlessLLM/pylet.git
cd pylet
pip install -e ".[dev]"
Quick Start
CLI
# Terminal 1: Start head node
pylet start
# Terminal 2: Start worker node with GPUs
pylet start --head localhost:8000 --gpu-units 4
# Terminal 3: Submit an instance
pylet submit 'vllm serve Qwen/Qwen2.5-1.5B-Instruct --port $PORT' \
--gpu-units 1 --name my-vllm
# Check status
pylet get-instance --name my-vllm
# Get endpoint for inference
pylet get-endpoint --name my-vllm
# Output: 192.168.1.10:15600
# View logs
pylet logs <instance-id>
# Cancel
pylet cancel <instance-id>
Python API
import pylet
# Connect to head node
pylet.init() # or pylet.init("http://head:8000")
# Submit an instance
instance = pylet.submit(
"vllm serve Qwen/Qwen2.5-1.5B-Instruct --port $PORT",
name="my-vllm",
gpu=1,
memory=4096,
)
# Wait for it to start
instance.wait_running()
print(f"Endpoint: {instance.endpoint}")
# Get logs
print(instance.logs())
# Cancel when done
instance.cancel()
instance.wait()
For local testing:
import pylet
with pylet.local_cluster(workers=2, gpu_per_worker=1) as cluster:
instance = pylet.submit("nvidia-smi", gpu=1)
instance.wait()
print(instance.logs())
Async API available via import pylet.aio as pylet.
See examples/README.md for more detailed examples including vLLM and SGLang.
Commands
| Command | Description |
|---|---|
pylet start |
Start head node |
pylet start --head <ip:port> --gpu-units N |
Start worker with N GPUs |
pylet submit <cmd> --gpu-units N --name <name> |
Submit instance |
pylet get-instance --name <name> |
Get instance status |
pylet get-endpoint --name <name> |
Get instance endpoint (host:port) |
pylet logs <id> |
View instance logs |
pylet logs <id> --follow |
Follow logs in real-time |
pylet cancel <id> |
Cancel instance |
pylet list-workers |
List registered workers |
Key Features
- Simple: No containers, no complex configs. Just
pylet startandpylet submit. - GPU-aware: Automatic GPU allocation via
CUDA_VISIBLE_DEVICES. - Fine-grained GPU control: Request specific GPUs by index, enable GPU sharing for daemons.
- Service discovery: Instances get a
PORTenv var; endpoint available viaget-endpoint. - Real-time logs: Stream logs from running instances.
- Graceful shutdown: SIGTERM with configurable grace period before SIGKILL.
Why PyLet?
PyLet is the greatest common denominator for GPU cluster orchestration - minimal, reliable, and unopinionated. It manages one thing: instances (processes with GPU allocation).
- No pods, replicas, services, or deployments - just instances
- Applications compose instances however they need (via labels, custom scheduling)
- Fine-grained GPU scheduling because research workloads need it
- Simple enough to understand, reliable enough to depend on
See docs/architecture.md for design details. See CONTRIBUTING.md for contribution philosophy.
Requirements
- Python 3.9+
- Linux (tested on Ubuntu)
License
Apache 2.0
Release files for pylet 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pylet-0.5.0.tar.gz | 95.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pylet-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 158.4 kB
Release files / pylet-0.5.0.tar.gz
| Download URL | pylet-0.5.0.tar.gz |
|---|---|
| Size | 95.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4a30e0d886d8fb7e3f870ff403b495175a0c714744b2e3244752105f3a1f724c
|
|
BLAKE2b-256 checksum How to use checksums |
a95f2e68f51bfc10d6607dc98c24f6acc71923510560888279bb19551c3e7d8d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 5, 2026.
Transparency logRelease files / pylet-0.5.0-py3-none-any.whl
| Download URL | pylet-0.5.0-py3-none-any.whl |
|---|---|
| Size | 62.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0ddd41ca7cd9cf46e3d4374142595cada77b64aebbb522ae3bab68b13d5eec64
|
|
BLAKE2b-256 checksum How to use checksums |
4c57101ac5235a545db409e70e13cb8e6f8064f0859c96992d2f2a40175dd4db
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 5, 2026.
Transparency log