Skip to main content

✨ letify

Declarations that become infrastructure.

Say what your function needs. Run it on the GPU you can afford.

Python License Providers letify-core

Quickstart · Why · Providers · Sweeps · Docs · 한국어


import letify

let = letify.Launcher()
colab = let.providers.colab_pro_plus

@let.function(device=colab.G4, host="remote")
def train(lr, bs):
    import torch
    ...
    return {"loss": loss}

print(train(lr=1e-4, bs=32))

No session to create. No environment to install. No files to upload. No scope to open, and no machine left running when you close the lid. 🎉


💡 Why declare instead of connect

Renting a GPU normally means going to it. You open a notebook or an SSH session, rebuild your environment there, copy your data across, and work inside somebody else's machine for as long as it lives. The training is the small part. The rest is infrastructure work that produces no results.

Declaring inverts that. You say what a function needs and where it belongs; the session, the environment, the transfer and the teardown are arranged for you.

🧠 You never leave the environment you were working in

Your editor, your debugger, your notes, your data and your git history stay where they are. The declaration sends one function to the card and brings the result back, so the remote GPU is a detail of one function rather than a place you move into.

That continuity is the point. You are not two people, one of whom lives in a browser tab with a different Python and no working directory. There is nothing to keep in sync and nothing to copy back before the session dies.

@let.function(device=colab.G4, host="remote")
def train(lr, bs):
    ...

train(lr=1e-4, bs=32)     # the same file you were already editing

⚡ There is no infrastructure step

Nothing is turned on, and nothing is left running by accident. A call starts the session it needs and ends it when the work is done, so forgetting to stop a GPU is not a mistake you can make. Neither is installing packages: the environment comes from the uv.lock you already have, cached so the second session does not pay for it.

The failure mode this removes is the expensive one. A forgotten instance bills overnight, and a session you spent twenty minutes preparing dies with everything in it.

📈 Nothing is tied to one server

A declaration names an accelerator shape, not a machine. So the same code runs on a Colab runtime, a lab box over SSH, an Elice allocation or this laptop by changing one line of configuration, and when one account runs out you add another rather than rewriting anything.

That is also how the work scales sideways. Capacity is what the provider entry declares it has, so a sweep spreads across every card available to you, across accounts and across machines:

[colab_pro.devices]
G4 = { count = 2 }            # two sessions on this account

[lab_a100.devices]
A100 = { indices = "0-3" }    # four cards in the shared box are ours

Six configurations then run six ways at once if six cards exist, on hardware that never had to be the same hardware. Nothing about the declaration changes when the pool grows.

💸 What that makes affordable

The same card costs wildly different amounts depending on which door you rent it through.

Same card, different door Per hour
🥇 Colab credits ~975 KRW
💸 Modal ~4,070 KRW

A factor of four for identical silicon. The cheap door is a notebook, though: no persistent disk, eviction at any moment, and whichever accelerator happens to be free. Everything above is what makes that door usable, so the cheapest option stops being the inconvenient one.

😖 Going to the GPU😌 Declaring it
# open a browser, pick a GPU, hope it's free
!pip install -q torch transformers  # 4 min
!gdown ...                           # 18 min
!git clone https://github.com/me/repo
%cd repo
# ... finally start working
# ... session dies, start over
@let.function(device=colab.G4, host="remote",
              volumes=[cache])
def train(lr, bs):
    ...

train(lr=1e-4, bs=32)

📦 Installation

uv add letify                 # core, no provider dependencies
uv add "letify[colab]"        # Google Colab
uv add "letify[modal]"        # Modal
uv add "letify[shell]"        # SSH, tunnels, Elice Cloud
uv add "letify[gcs]"          # Google Cloud Storage cache
uv add "letify[s3]"           # S3 compatible cache
uv add "letify[all]"          # everything

The Python package is pure Python. A provider whose package is missing simply reports itself unavailable, and the rest keeps working.

One optional piece is native. host="local" needs letify-core, a Rust workspace that stands in for the CUDA driver, built with python letify-core/build.py. If you only ship functions to remote machines you never need it.


🚀 Quickstart

1. Declare your accounts once

Put accounts in ~/.letify, so they belong to your machine and never to the repository.

[colab_pro_plus]
kind = "colab"
account = "you@example.com"

[lab_a100]
kind = "shell"
address = "gpu.lab.example.edu"
user = "researcher"
key = "~/.ssh/id_ed25519"
persistent = true

🔐 Secrets are referenced, never written. Use access_token_env = "MY_TOKEN" or access_token_keyring = "service/user".

2. Look around

$ letify providers
colab_pro_plus  colab   ephemeral   channel=persistent
lab_a100        shell   persistent  channel=persistent
local           local   persistent  channel=persistent

$ letify devices
{
  "colab_pro_plus": ["A100", "G4", "H100", "L4", "T4", "v5e1", "v6e1"],
  "lab_a100":       ["A100"],
  "local":          ["CPU", "GeForce_RTX_4050"]
}

3. Declare and run

import letify

let = letify.Launcher()
env = letify.Env()                      # reads uv.lock
colab = let.providers.colab_pro_plus
cache = colab.volume("hf-cache")        # survives the session

@let.function(device=colab.G4, host="remote", env=env, volumes=[cache])
def train(lr, bs):
    ...
    return {"loss": loss}

print(train(lr=1e-4, bs=32))

That is the whole program. 🍰


🧭 What a declaration says

Three words, and none of them is a mechanism.

@let.function(
    device=colab.G4,       # where the accelerator is, with provider and account
    host="remote",         # where the host code runs
    lifetime="call",       # how long the session lives
)

device carries the provider, the account and the accelerator in one value, because those are one decision. Core count and memory come with the shape the provider registered, so there is nothing to ask for.

host is the CUDA word for the CPU side. "local", the default, keeps Python and the libraries in this process and forwards only CUDA calls. "remote" ships the function to the machine that holds the device.

lifetime is how long the session lives. "call", the default, ends it with the call, counting a search space as one call. "process" keeps it so a run of separate calls does not pay session start each time.

📐 Which host to pick, with the arithmetic
📦 host="remote" 🔌 host="local"
What moves your whole loop, once every CUDA call
Cost one transfer one round trip per host synchronization
Fine-tuning at 150 ms ~99% 53% default, ~96% tuned
Decoding at 150 ms hundreds of tok/s 2 to 7 tok/s

Efficiency against a direct run is T / (T + k × RTT), where T is GPU time per step and k is how many times per step the host reads a value back from the device.

A default Hugging Face training step has k ≈ 3: the trainer's NaN filter, the SDPA attention mask check, and logging. With a 0.5 s NVFP4 micro step at 150 ms that is 53%. Turn the NaN filter off, remove the mask check with fixed length packing, and log at the gradient accumulation boundary, and it is about 96%.

Counter-intuitive consequence: a faster GPU makes forwarding worse, because T shrinks and RTT does not. The same step on an L4 takes 1.8 s and reaches 80%.

Decoding is the case that stays bad. Throughput is bounded near 1000 / (k × RTT) tokens per second, so the card stops mattering. Ship the whole generate call instead.

Measure your own k with torch.cuda.set_sync_debug_mode("warn") and your round trip with letify probe. See docs/NETWORK.md.

⚠️ letify never silently changes the mode. Ask for something a provider cannot serve and you get an exception naming the reason. Ask for something slow and you get a warning with the numbers, and then it runs, because the choice is yours.


🌍 Providers

Provider
├── 💻 Local      your machine            persistent
├── ☁️  Modal      serverless GPU          persistent
└── 🐚 Shell      any machine over SSH    ephemeral by default
    ├── 📓 Colab   via the official CLI
    ├── 🕳️  Tunnel  Tailscale or frp, for NAT
    └── 🇰🇷 Elice   Elice Cloud, allocated by API
Storage Best for
💻 Local persistent your own GPU, and testing everything else
☁️ Modal persistent production serving, reproducible images
📓 Colab ephemeral cheap batch work, sweeps, NVFP4 on G4
🐚 Shell overridable lab and university servers
🕳️ Tunnel overridable a machine behind NAT you cannot port-forward
🇰🇷 Elice persistent Korean GPU cloud, per-second billing

Multiple accounts are first class. Each configuration entry is one account, and entries of the same kind coexist. Two Colab accounts means twice the concurrent sessions.

a = let.providers.colab_pro_plus
b = let.providers.colab_pro

@let.function(device=a.G4, host="remote")
def train(lr): ...

@let.function(device=b.L4, host="remote")      # different account, same program
def evaluate(ckpt): ...

Or do not pick at all:

@let.function(device=let.providers.any.A100, host="remote")
def train(lr): ...

🔭 Sweeps

Fan-out is a declared space, not a .map() call. Passing a space where a scalar is expected says that argument varies.

space = letify.grid(lr=[1e-4, 3e-4, 1e-3], bs=[16, 32])   # 6 points
pairs = letify.zip(lr=[1e-4, 3e-4], bs=[16, 32])          # 2 points
both  = letify.grid(lr=[1e-4]) | letify.grid(lr=[1e-3])   # union

Then consume it with the language you already know. 🐍

@let.function(device=colab.G4, host="remote")
async def train(lr, bs):
    ...

results = await train(space)              # list, in input order

async for r in train(space):              # streamed, as each finishes
    print(r)

🧵 Sync or async is declared at the def, not at the call. A plain def blocks. An async def gives you a coroutine, so await and asyncio.gather work exactly as they always do. letify adds no future type of its own, and there is no .remote(), .spawn() or .map() to remember.


💾 Caching that actually helps

A volume is a content addressed blob store. Contents are named by their hash, and mutable names live in a separate tiny namespace, exactly like Git objects and refs.

cache = colab.volume("hf-cache")

@let.function(device=colab.G4, host="remote", volumes=[cache])
def train(lr): ...

Why this shape:

🐌 Two-way file sync ⚡ Content addressed
Concurrent writers last one wins, work is lost cannot collide, by construction
Already transferred? compare size and timestamps holding the hash is the proof
50k small files 50k round trips one packed archive, one transfer

The payoff is where it hurts most, which is session start:

Pulling a 20 GB model cache Time
🐢 From a lab server over 100 Mbit/s ~27 min
🚶 From the Hugging Face hub 3 to 5 min
🚀 From a bucket next to the runtime 40 to 60 s

All of that time is billed as GPU time. That is why an ephemeral provider with a volume attached behaves like a persistent one.


🔗 Values that stay put

A session is one living process, so a value can stay in it.

@let.function(device=colab.G4, host="remote", lifetime="process", keep_remote=True)
def build_model():
    return load_model()          # 14 GB, stays on the remote machine

@let.function(device=colab.G4, host="remote", lifetime="process")
def evaluate(model, batch):
    return model(batch)          # the handle resolves in place

model = build_model()            # a Handle, not 14 GB
evaluate(model=model, batch=...)

Large arguments are content addressed too. Pass the same tensor to ten calls and it crosses the network once, because the runtime is asked by digest whether it already holds it.

A handle names the session that holds it. Passing one to a different session raises rather than quietly copying the object across, since that would be an unrequested transfer of everything it points at.


💸 Your bill cannot run away

Nothing has to be torn down by hand.

A call ends its own session. That is the default, and a sweep counts as one call, so six points start one set of sessions and end them once.

lifetime="process" is the opt in, for a run of separate calls that would otherwise pay session start each time. It ends when your process does, because a timer that ended it sooner would overrule what you declared.

The lease is the backstop. The session holds a deadline that this process keeps renewing. Kill your script, lose your laptop, crash your kernel, and the worker exits on its own, which frees the card. The grace period is long enough that a flaky connection does not kill a training run. Whether it also stops the billing depends on what the provider charges for: docs/guide/06-cost.md says which providers are covered and which are not.

🚫 There is deliberately no detached mode. A detached run whose remote side gets preempted loses its results. Instead, the local process stays the owner, and durability comes from checkpoints in the store.


🧪 Test without a GPU, without mocks

The local provider starts the same worker behind the same framed protocol a remote runtime would. Your tests exercise the real path.

def test_train_returns_a_loss():
    let = letify.Launcher(home=False)

    @let.function(device=let.providers.local.CPU, host="remote")
    def train(lr):
        return {"loss": 1.0 / lr}

    assert train(lr=2.0)["loss"] == 0.5

🛠️ CLI

letify login shell lab        # declare an account, and reference it here
letify logout lab             # take the account off this machine
letify providers              # who is declared, storage, channel kind
letify devices                # what each one offers
letify status                 # what is running right now
letify usage                  # what is left on each account
letify utilization            # how busy each instance's GPU is
letify check lab              # does this machine answer?
letify probe lab              # is host="local" worth using here?
letify efficiency 0.5 3 150   # the formula, from measured terms

📚 Documentation

🧪 examples/ Working scenarios, starting with a LoRA sweep on a rented card
📖 PROJECT.md The full feature set and API surface
🎯 docs/INTENT.md Goals, claims, constraints, open questions
📐 docs/SPEC.md The design as it stands, decision by decision
🧩 docs/COMPONENT.md Every class, and the vocabulary
🌐 docs/NETWORK.md Transports, latency measurements, tunnel choices
🦀 letify-core/ The Rust workspace behind host="local"
🧭 docs/guide/ Task-oriented guides
🇰🇷 docs/locales/README_ko.md 한국어

🚧 Status

Alpha, and honest about it. What works today:

✅ Declarations, sync and async, sweeps, pooling, session lifetimes and the lease ✅ Persistent sessions: handles resolve in later calls, large arguments travel once ✅ Content addressed storage, configuration and secrets ✅ The Local and Colab providers ✅ letify-core, verified on a real GPU: the agent opens the driver, the local driver forwards an allocation and a copy in both directions, and the bytes match

Not finished yet:

🚧 letify-driver covers the entry points a PyTorch process needs to start up and run one kernel. Anything else names itself and returns CUDA_ERROR_NOT_SUPPORTED, so a real run prints the list of what to build next 🚧 Modal and Elice follow each published interface but have not been run against the live services 🚧 Unified memory cannot be forwarded at all, so a paged optimizer needs host="remote"

The full list is at the end of docs/SPEC.md.


🤝 Contributing

Spec driven and test driven: settle docs/SPEC.md, write the failing test, then write the code. Performance work uses ResearchTree, where one branch is one experiment and one pull request is its lab note. Read CLAUDE.md before opening one.


Apache 2.0 licensed. Built for people who pay for their own GPUs. 🔬

Metadata

Release files for letify 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for letify 1.0.0
File Size Uploaded
letify-1.0.0.tar.gz 228.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for letify 1.0.0
File Interpreter ABI Platform
letify-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 344.8 kB

Release files / letify-1.0.0.tar.gz

Download URL letify-1.0.0.tar.gz
Size 228.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f90619e0f135319d7408ee123d66c40f248ade0e3cb1cafd8a7869b3ad94046c
BLAKE2b-256 checksum
How to use checksums
f455119d524340581d8c079ae2095ad55c562aec43a9843343787b0d1f404b6b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log

Release files / letify-1.0.0-py3-none-any.whl

Download URL letify-1.0.0-py3-none-any.whl
Size 116.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
10b2aec7bdc9978dcb9294870d73ca14525596cb05d5721b5fb70eb868da135f
BLAKE2b-256 checksum
How to use checksums
8c9840ae3ec7a15ba191b25068e533797496e7f8ffe17417c6b8344aaaadcf06
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 13, 2026.

Transparency log

Release history Release notifications | RSS feed

1.2.2

7 release files

1.2.1

7 release files

1.2.0

7 release files

1.1.2

7 release files

1.1.1

7 release files

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page