✨ letify
Declarations that become infrastructure.
Say what your function needs. Run it on the GPU you can afford.
import letify
let = letify.Launcher()
colab = let.providers.colab_pro_plus
@let.function(device=colab.G4, host="remote")
def train(lr, bs):
import torch
...
return {"loss": loss}
print(train(lr=1e-4, bs=32))
No session to create. No environment to install. No files to upload. No scope to open, and no machine left running when you close the lid. 🎉
💡 Why declare instead of connect
Renting a GPU normally means going to it. You open a notebook or an SSH session, rebuild your environment there, copy your data across, and work inside somebody else's machine for as long as it lives. The training is the small part. The rest is infrastructure work that produces no results.
Declaring inverts that. You say what a function needs and where it belongs; the session, the environment, the transfer and the teardown are arranged for you.
🧠 You never leave the environment you were working in
Your editor, your debugger, your notes, your data and your git history stay where they are. The declaration sends one function to the card and brings the result back, so the remote GPU is a detail of one function rather than a place you move into.
That continuity is the point. You are not two people, one of whom lives in a browser tab with a different Python and no working directory. There is nothing to keep in sync and nothing to copy back before the session dies.
@let.function(device=colab.G4, host="remote")
def train(lr, bs):
...
train(lr=1e-4, bs=32) # the same file you were already editing
⚡ There is no infrastructure step
Nothing is turned on, and nothing is left running by accident. A call starts the session it
needs and ends it when the work is done, so forgetting to stop a GPU is not a mistake you can
make. Neither is installing packages: the environment comes from the uv.lock you already
have, cached so the second session does not pay for it.
The failure mode this removes is the expensive one. A forgotten instance bills overnight, and a session you spent twenty minutes preparing dies with everything in it.
📈 Nothing is tied to one server
A declaration names an accelerator shape, not a machine. So the same code runs on a Colab runtime, a lab box over SSH, an Elice allocation or this laptop by changing one line of configuration, and when one account runs out you add another rather than rewriting anything.
That is also how the work scales sideways. Capacity is what the provider entry declares it has, so a sweep spreads across every card available to you, across accounts and across machines:
[colab_pro.devices]
G4 = { count = 2 } # two sessions on this account
[lab_a100.devices]
A100 = { indices = "0-3" } # four cards in the shared box are ours
Six configurations then run six ways at once if six cards exist, on hardware that never had to be the same hardware. Nothing about the declaration changes when the pool grows.
💸 What that makes affordable
The same card costs wildly different amounts depending on which door you rent it through.
| Same card, different door | Per hour |
|---|---|
| 🥇 Colab credits | ~975 KRW |
| 💸 Modal | ~4,070 KRW |
A factor of four for identical silicon. The cheap door is a notebook, though: no persistent disk, eviction at any moment, and whichever accelerator happens to be free. Everything above is what makes that door usable, so the cheapest option stops being the inconvenient one.
| 😖 Going to the GPU | 😌 Declaring it |
|---|---|
# open a browser, pick a GPU, hope it's free
!pip install -q torch transformers # 4 min
!gdown ... # 18 min
!git clone https://github.com/me/repo
%cd repo
# ... finally start working
# ... session dies, start over
|
@let.function(device=colab.G4, host="remote",
volumes=[cache])
def train(lr, bs):
...
train(lr=1e-4, bs=32)
|
📦 Installation
uv add letify # core, no provider dependencies
uv add "letify[colab]" # Google Colab
uv add "letify[modal]" # Modal
uv add "letify[shell]" # SSH, tunnels, Elice Cloud
uv add "letify[gcs]" # Google Cloud Storage cache
uv add "letify[s3]" # S3 compatible cache
uv add "letify[all]" # everything
The Python package is pure Python. A provider whose package is missing simply reports itself unavailable, and the rest keeps working.
One optional piece is native. host="local" needs letify-core, a Rust workspace that stands in for the CUDA driver, built with python letify-core/build.py. If you only ship functions to remote machines you never need it.
🚀 Quickstart
1. Declare your accounts once
Put accounts in ~/.letify, so they belong to your machine and never to the repository.
[colab_pro_plus]
kind = "colab"
account = "you@example.com"
[lab_a100]
kind = "shell"
address = "gpu.lab.example.edu"
user = "researcher"
key = "~/.ssh/id_ed25519"
persistent = true
🔐 Secrets are referenced, never written. Use
access_token_env = "MY_TOKEN"oraccess_token_keyring = "service/user".
2. Look around
$ letify providers
colab_pro_plus colab ephemeral channel=persistent
lab_a100 shell persistent channel=persistent
local local persistent channel=persistent
$ letify devices
{
"colab_pro_plus": ["A100", "G4", "H100", "L4", "T4", "v5e1", "v6e1"],
"lab_a100": ["A100"],
"local": ["CPU", "GeForce_RTX_4050"]
}
3. Declare and run
import letify
let = letify.Launcher()
env = letify.Env() # reads uv.lock
colab = let.providers.colab_pro_plus
cache = colab.volume("hf-cache") # survives the session
@let.function(device=colab.G4, host="remote", env=env, volumes=[cache])
def train(lr, bs):
...
return {"loss": loss}
print(train(lr=1e-4, bs=32))
That is the whole program. 🍰
🧭 What a declaration says
Three words, and none of them is a mechanism.
@let.function(
device=colab.G4, # where the accelerator is, with provider and account
host="remote", # where the host code runs
lifetime="call", # how long the session lives
)
device carries the provider, the account and the accelerator in one value, because those are one decision. Core count and memory come with the shape the provider registered, so there is nothing to ask for.
host is the CUDA word for the CPU side. "local", the default, keeps Python and the libraries in this process and forwards only CUDA calls. "remote" ships the function to the machine that holds the device.
lifetime is how long the session lives. "call", the default, ends it with the call, counting a search space as one call. "process" keeps it so a run of separate calls does not pay session start each time.
📐 Which host to pick, with the arithmetic
📦 host="remote" |
🔌 host="local" |
|
|---|---|---|
| What moves | your whole loop, once | every CUDA call |
| Cost | one transfer | one round trip per host synchronization |
| Fine-tuning at 150 ms | ~99% | 53% default, ~96% tuned |
| Decoding at 150 ms | hundreds of tok/s | 2 to 7 tok/s |
Efficiency against a direct run is T / (T + k × RTT), where T is GPU time per step and k is how many times per step the host reads a value back from the device.
A default Hugging Face training step has k ≈ 3: the trainer's NaN filter, the SDPA attention mask check, and logging. With a 0.5 s NVFP4 micro step at 150 ms that is 53%. Turn the NaN filter off, remove the mask check with fixed length packing, and log at the gradient accumulation boundary, and it is about 96%.
Counter-intuitive consequence: a faster GPU makes forwarding worse, because T shrinks and RTT does not. The same step on an L4 takes 1.8 s and reaches 80%.
Decoding is the case that stays bad. Throughput is bounded near 1000 / (k × RTT) tokens per second, so the card stops mattering. Ship the whole generate call instead.
Measure your own k with torch.cuda.set_sync_debug_mode("warn") and your round trip with letify probe. See docs/NETWORK.md.
⚠️ letify never silently changes the mode. Ask for something a provider cannot serve and you get an exception naming the reason. Ask for something slow and you get a warning with the numbers, and then it runs, because the choice is yours.
🌍 Providers
Provider
├── 💻 Local your machine persistent
├── ☁️ Modal serverless GPU persistent
└── 🐚 Shell any machine over SSH ephemeral by default
├── 📓 Colab via the official CLI
├── 🕳️ Tunnel Tailscale or frp, for NAT
└── 🇰🇷 Elice Elice Cloud, allocated by API
| Storage | Best for | |
|---|---|---|
💻 Local |
persistent | your own GPU, and testing everything else |
☁️ Modal |
persistent | production serving, reproducible images |
📓 Colab |
ephemeral | cheap batch work, sweeps, NVFP4 on G4 |
🐚 Shell |
overridable | lab and university servers |
🕳️ Tunnel |
overridable | a machine behind NAT you cannot port-forward |
🇰🇷 Elice |
persistent | Korean GPU cloud, per-second billing |
Multiple accounts are first class. Each configuration entry is one account, and entries of the same kind coexist. Two Colab accounts means twice the concurrent sessions.
a = let.providers.colab_pro_plus
b = let.providers.colab_pro
@let.function(device=a.G4, host="remote")
def train(lr): ...
@let.function(device=b.L4, host="remote") # different account, same program
def evaluate(ckpt): ...
Or do not pick at all:
@let.function(device=let.providers.any.A100, host="remote")
def train(lr): ...
🔭 Sweeps
Fan-out is a declared space, not a .map() call. Passing a space where a scalar is expected says that argument varies.
space = letify.grid(lr=[1e-4, 3e-4, 1e-3], bs=[16, 32]) # 6 points
pairs = letify.zip(lr=[1e-4, 3e-4], bs=[16, 32]) # 2 points
both = letify.grid(lr=[1e-4]) | letify.grid(lr=[1e-3]) # union
Then consume it with the language you already know. 🐍
@let.function(device=colab.G4, host="remote")
async def train(lr, bs):
...
results = await train(space) # list, in input order
async for r in train(space): # streamed, as each finishes
print(r)
🧵 Sync or async is declared at the
def, not at the call. A plaindefblocks. Anasync defgives you a coroutine, soawaitandasyncio.gatherwork exactly as they always do. letify adds no future type of its own, and there is no.remote(),.spawn()or.map()to remember.
💾 Caching that actually helps
A volume is a content addressed blob store. Contents are named by their hash, and mutable names live in a separate tiny namespace, exactly like Git objects and refs.
cache = colab.volume("hf-cache")
@let.function(device=colab.G4, host="remote", volumes=[cache])
def train(lr): ...
Why this shape:
| 🐌 Two-way file sync | ⚡ Content addressed | |
|---|---|---|
| Concurrent writers | last one wins, work is lost | cannot collide, by construction |
| Already transferred? | compare size and timestamps | holding the hash is the proof |
| 50k small files | 50k round trips | one packed archive, one transfer |
The payoff is where it hurts most, which is session start:
| Pulling a 20 GB model cache | Time |
|---|---|
| 🐢 From a lab server over 100 Mbit/s | ~27 min |
| 🚶 From the Hugging Face hub | 3 to 5 min |
| 🚀 From a bucket next to the runtime | 40 to 60 s |
All of that time is billed as GPU time. That is why an ephemeral provider with a volume attached behaves like a persistent one.
🔗 Values that stay put
A session is one living process, so a value can stay in it.
@let.function(device=colab.G4, host="remote", lifetime="process", keep_remote=True)
def build_model():
return load_model() # 14 GB, stays on the remote machine
@let.function(device=colab.G4, host="remote", lifetime="process")
def evaluate(model, batch):
return model(batch) # the handle resolves in place
model = build_model() # a Handle, not 14 GB
evaluate(model=model, batch=...)
Large arguments are content addressed too. Pass the same tensor to ten calls and it crosses the network once, because the runtime is asked by digest whether it already holds it.
A handle names the session that holds it. Passing one to a different session raises rather than quietly copying the object across, since that would be an unrequested transfer of everything it points at.
💸 Your bill cannot run away
Nothing has to be torn down by hand.
A call ends its own session. That is the default, and a sweep counts as one call, so six points start one set of sessions and end them once.
lifetime="process" is the opt in, for a run of separate calls that would otherwise pay session start each time. It ends when your process does, because a timer that ended it sooner would overrule what you declared.
The lease is the backstop. The session holds a deadline that this process keeps renewing. Kill your script, lose your laptop, crash your kernel, and the worker exits on its own, which frees the card. The grace period is long enough that a flaky connection does not kill a training run. Whether it also stops the billing depends on what the provider charges for: docs/guide/06-cost.md says which providers are covered and which are not.
🚫 There is deliberately no detached mode. A detached run whose remote side gets preempted loses its results. Instead, the local process stays the owner, and durability comes from checkpoints in the store.
🧪 Test without a GPU, without mocks
The local provider starts the same worker behind the same framed protocol a remote runtime would. Your tests exercise the real path.
def test_train_returns_a_loss():
let = letify.Launcher(home=False)
@let.function(device=let.providers.local.CPU, host="remote")
def train(lr):
return {"loss": 1.0 / lr}
assert train(lr=2.0)["loss"] == 0.5
🛠️ CLI
letify login shell lab # declare an account, and reference it here
letify logout lab # take the account off this machine
letify providers # who is declared, storage, channel kind
letify devices # what each one offers
letify status # what is running right now
letify usage # what is left on each account
letify utilization # how busy each instance's GPU is
letify check lab # does this machine answer?
letify probe lab # is host="local" worth using here?
letify efficiency 0.5 3 150 # the formula, from measured terms
📚 Documentation
| 🧪 examples/ | Working scenarios, starting with a LoRA sweep on a rented card |
| 📖 PROJECT.md | The full feature set and API surface |
| 🎯 docs/INTENT.md | Goals, claims, constraints, open questions |
| 📐 docs/SPEC.md | The design as it stands, decision by decision |
| 🧩 docs/COMPONENT.md | Every class, and the vocabulary |
| 🌐 docs/NETWORK.md | Transports, latency measurements, tunnel choices |
| 🦀 letify-core/ | The Rust workspace behind host="local" |
| 🧭 docs/guide/ | Task-oriented guides |
| 🇰🇷 docs/locales/README_ko.md | 한국어 |
🚧 Status
Alpha, and honest about it. What works today:
✅ Declarations, sync and async, sweeps, pooling, session lifetimes and the lease
✅ Persistent sessions: handles resolve in later calls, large arguments travel once
✅ Content addressed storage, configuration and secrets
✅ The Local and Colab providers
✅ letify-core, verified on a real GPU: the agent opens the driver, the local driver forwards an allocation and a copy in both directions, and the bytes match
Not finished yet:
🚧 letify-driver covers the entry points a PyTorch process needs to start up and run one kernel. Anything else names itself and returns CUDA_ERROR_NOT_SUPPORTED, so a real run prints the list of what to build next
🚧 Modal and Elice follow each published interface but have not been run against the live services
🚧 Unified memory cannot be forwarded at all, so a paged optimizer needs host="remote"
The full list is at the end of docs/SPEC.md.
🤝 Contributing
Spec driven and test driven: settle docs/SPEC.md, write the failing test, then write the code. Performance work uses ResearchTree, where one branch is one experiment and one pull request is its lab note. Read CLAUDE.md before opening one.
Apache 2.0 licensed. Built for people who pay for their own GPUs. 🔬
Metadata
Release files for letify 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| letify-1.0.0.tar.gz | 228.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| letify-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 344.8 kB
Release files / letify-1.0.0.tar.gz
| Download URL | letify-1.0.0.tar.gz |
|---|---|
| Size | 228.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f90619e0f135319d7408ee123d66c40f248ade0e3cb1cafd8a7869b3ad94046c
|
|
BLAKE2b-256 checksum How to use checksums |
f455119d524340581d8c079ae2095ad55c562aec43a9843343787b0d1f404b6b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.
Transparency logRelease files / letify-1.0.0-py3-none-any.whl
| Download URL | letify-1.0.0-py3-none-any.whl |
|---|---|
| Size | 116.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
10b2aec7bdc9978dcb9294870d73ca14525596cb05d5721b5fb70eb868da135f
|
|
BLAKE2b-256 checksum How to use checksums |
8c9840ae3ec7a15ba191b25068e533797496e7f8ffe17417c6b8344aaaadcf06
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.
Transparency log