Skip to main content

PyPI Tests Discord

Compile → Benchmark → Deploy any LLM on any GPU. vLLM, SGLang, or create your own specialized deployment using a hackable compiler.

Install

git clone https://github.com/cloudrift-ai/emmy.git
cd emmy && make setup

Compile

A hackable PyTorch → Graph IR → CUDA compiler. Trace any nn.Module, fuse it into one kernel, run it, and inspect the emitted CUDA. See the blog post: A Principled ML Compiler Stack in 5,000 Lines of Python.

# Compile a single layer
emmy compile -c "nn.RMSNorm(2048)(torch.randn(1,32,2048))"
# Benchmark, profile and optimize kernels locally
emmy run --bench --profile -c "torch.nn.Softmax(dim=-1)(torch.randn(1, 28, 2048, 2048))"
# Compile full model from HuggingFace (will download weights)
emmy compile Qwen/Qwen3-Embedding-0.6B

Layer-norm-style reduction (two reductions, broadcast subtract, elementwise chain) fused into two kernels:

emmy compile -c "
class LN(torch.nn.Module):
    def forward(self, x):
        m = x.mean(-1, keepdim=True)
        v = ((x - m) ** 2).mean(-1, keepdim=True)
        return (x - m) * torch.rsqrt(v + 1e-6)
LN()(torch.randn(64, 2048))"

Principled compilation stack with six IR stages, each printable on demand via --ir <stage>:

  1. Torch IR — captures the FX graph as a 1:1 mirror of PyTorch's op set (rmsnorm, linear, softmax, ...)
  2. Tensor IR — decomposes every Torch op into three primitives: Elementwise, Reduction, and IndexMap
  3. Loop IR — lifts each primitive to a LoopOp and fuses
  4. Tile IR — schedules kernels onto GPU
  5. Kernel IR — materializes the schedule into framework-agnostic hardware primitives
  6. CUDA — optimized CUDA code ready for nvcc

Readable Schedule: emmy compile -c "nn.RMSNorm(2048)(torch.randn(1,32,2048))" --ir tile

kernel k_rms_norm_reduce  inputs: rms_norm_mean_count, rms_norm_eps, x, p_weight  outputs: rms_norm
    in0 = load rms_norm_mean_count[0]
    in1 = load rms_norm_eps[0]
    Tile(axes=(a0:256=THREAD, a1:32=BLOCK)):
        x_smem = Stage(x, origin=(0, a1, 0), slab=(a2:2048@2)) async
        p_weight_smem = Stage(p_weight, origin=(0), slab=(a3:2048@0)) async
        StridedLoop(a2 = a0; < 2048; += 256):  # reduce
            in2 = load x_smem[a2]
            v0 = multiply(in2, in2)
            acc0 <- add(acc0, v0)
        v1 = divide(acc0, in0)
        v2 = add(v1, in1)
        v3 = rsqrt(v2)
        StridedLoop(a3 = a0; < 2048; += 256):  # free
            in3 = load x_smem[a3]
            in4 = load p_weight_smem[a3]
            v4 = multiply(in3, v3)
            v5 = multiply(v4, in4)
            rms_norm[0, a1, a3] = v5

Optimized CUDA kernel: emmy compile -c "nn.RMSNorm(2048)(torch.randn(1,32,2048))" --ir cuda

extern "C" __global__
__launch_bounds__(256) void k_rms_norm_reduce(const float* x, const float* p_weight, float* rms_norm) {
    float in0 = 2048.0f;
    float in1 = 1e-06f;
    {
        int a1 = blockIdx.x;
        int a0 = threadIdx.x;
        float acc0 = 0.0f;
        __syncthreads();
        __shared__ float x_smem[2048];
        for (int x_smem_flat = a0; x_smem_flat < 2048; x_smem_flat += 256) {
            {
                unsigned int _smem_addr = __cvta_generic_to_shared(&x_smem[x_smem_flat]);
                asm volatile("cp.async.ca.shared.global [%0], [%1], 4;\n"
                             :: "r"(_smem_addr), "l"(&x[a1 * 2048 + x_smem_flat])
                             : "memory");
            }
        }
        asm volatile("cp.async.commit_group;\n" ::: "memory");
        asm volatile("cp.async.wait_group 0;\n" ::: "memory");
        __syncthreads();
        __shared__ float p_weight_smem[2048];
        for (int p_weight_smem_flat = a0; p_weight_smem_flat < 2048; p_weight_smem_flat += 256) {
            {
                unsigned int _smem_addr = __cvta_generic_to_shared(&p_weight_smem[p_weight_smem_flat]);
                asm volatile("cp.async.ca.shared.global [%0], [%1], 4;\n"
                             :: "r"(_smem_addr), "l"(&p_weight[p_weight_smem_flat])
                             : "memory");
            }
        }
        asm volatile("cp.async.commit_group;\n" ::: "memory");
        asm volatile("cp.async.wait_group 0;\n" ::: "memory");
        __syncthreads();
        for (int a2 = a0; a2 < 2048; a2 += 256) {
            float in2 = x_smem[a2];
            float v0 = in2 * in2;
            acc0 += v0;
        }
        __shared__ float acc0_smem[256];
        acc0_smem[a0] = acc0;
        __syncthreads();
        for (int s = 128; s > 0; s >>= 1) {
            if (a0 < s) {
                acc0_smem[a0] = acc0_smem[a0] + acc0_smem[a0 + s];
            }
            __syncthreads();
        }
        __syncthreads();
        float acc0_b = acc0_smem[0];
        float v1 = acc0_b / in0;
        float v2 = v1 + in1;
        float v3 = rsqrtf(v2);
        for (int a3 = a0; a3 < 2048; a3 += 256) {
            float in3 = x_smem[a3];
            float in4 = p_weight_smem[a3];
            float v4 = in3 * v3;
            float v5 = v4 * in4;
            rms_norm[a1 * 2048 + a3] = v5;
        }
    }
}

Benchmark

emmy bench recipes/*                                    # All recipes
emmy bench experiments/.../optimal_mcr_rtx5090          # An experiment
emmy bench recipes/* --filter "deploy.gpu=*5090*"       # Subset
emmy bench recipes/* --gpu-concurrency 4                # Parallel VMs per GPU
emmy bench recipes/* --local                            # On this machine
emmy bench recipes/* --ssh user@host1 --ssh user@host2  # Pre-allocated hosts

External contributors: open a PR with an experiment under experiments/{model}/{name}/, then a maintainer triggers a cloud run by commenting /run-experiment on the PR.

Deploy

# Remote server via SSH
emmy deploy ssh --recipe recipes/GLM-4.6-FP8 --ssh user@host

# Local Docker Compose
emmy deploy local --recipe recipes/Qwen3-Coder-30B-A3B-Instruct-AWQ

# Cloud (auto-provisions a VM)
emmy deploy cloud --recipe recipes/GLM-4.6-FP8 --gpu "NVIDIA H200 141GB" --gpu-count 8

# Teardown / preview
emmy deploy ssh --recipe recipes/GLM-4.6-FP8 --ssh user@host --teardown
emmy deploy ssh --recipe recipes/GLM-4.6-FP8 --ssh user@host --dry-run

Serve (compiled embeddings via vLLM)

# vLLM's OpenAI shell (/v1/embeddings, tokenizer, scheduler, pooler) over emmy-compiled kernels
pip install -e ".[compile,serving]"
emmy serve Qwen/Qwen3-Embedding-0.6B                  # extra flags pass through to vllm serve

curl localhost:8000/v1/embeddings -H 'Content-Type: application/json' \
  -d '{"model":"Qwen/Qwen3-Embedding-0.6B","input":"Hello"}'

# One-shot benchmark (vllm bench serve against the started server), and the raw-vLLM baseline
emmy serve Qwen/Qwen3-Embedding-0.6B --bench --random-input-len 32
emmy serve Qwen/Qwen3-Embedding-0.6B --bench --random-input-len 32 --stock

See emmy/serving/ARCHITECTURE.md; embedding recipes (recipes/Qwen3-Embedding-*) A/B this against stock vLLM via emmy bench.

Generate (chat) — experimental

# Standalone generation oracle (no vLLM) — re-runs the whole prefix each step, O(S²); the
# token-for-token reference (matches HF eager greedy, e.g. on TinyLlama-1.1B-Chat).
emmy generate TinyLlama/TinyLlama-1.1B-Chat-v1.0 --prompt "The capital of France is" --max-new-tokens 10

# Serve a chat model through emmy-compiled per-layer kernels (vLLM owns the OpenAI API /
# sampler / scheduler / paged KV-cache / lm_head; emmy owns embed + the trunk).
emmy serve TinyLlama/TinyLlama-1.1B-Chat-v1.0 --generate
curl localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"TinyLlama/TinyLlama-1.1B-Chat-v1.0","messages":[{"role":"user","content":"Hi"}]}'

Status: correctness complete for decoder-only Llama / Qwen3 (full-causal, fp16, TP=1). Perf is not yet hardened — host-sync interleave at the per-layer seam, and serve compiles 2× n_layers programs (startup- and memory-heavy → small models for now). Design + phase status: plans/generative-inference-support.md.

Recipe

model:
  huggingface: "org/model-name"

engine:
  llm:
    tensor_parallel_size: 8
    gpu_memory_utilization: 0.9
    context_length: 16384
    max_concurrent_requests: 512
    vllm:
      image: "vllm/vllm-openai:v0.17.0"
      extra_args: "--kv-cache-dtype fp8"

benchmark:
  max_concurrency: 128
  num_prompts: 256
  random_input_len: 8000
  random_output_len: 8000

# Cross-product: 3 GPUs × 2 concurrency configs = 6 variants
matrices:
  cross:
    deploy.gpu_count: 1
    deploy.gpu:
      - "NVIDIA GeForce RTX 5090"
      - "NVIDIA H100 80GB"
      - "NVIDIA H200 141GB"
    zip:
      engine.llm.max_concurrent_requests: [128, 512]
      benchmark.max_concurrency: [128, 512]

Generic workload (run any tool on the VM, pull back result files):

command:
  stage: ["scripts"]
  run: |
    nvidia-smi --query-gpu=name,memory.used --format=csv > $task_dir/result.csv
  result_files: ["result.csv"]
  timeout: 60

matrices:
  deploy.gpu: "NVIDIA GeForce RTX 5090"
  deploy.gpu_count: 1

Virtual Machine Management

# GCP
emmy vm create gcp --instance my-vm --zone us-central1-a --machine-type a2-highgpu-1g
emmy vm delete gcp --instance my-vm --zone us-central1-a

# CloudRift
emmy vm create cloudrift --instance-type rtx4090.1 --ssh-key ~/.ssh/id_ed25519.pub
emmy vm delete cloudrift --instance-id <id>

Development

make test      # run pytest
make lint      # ruff check + format check
make format    # auto-fix

Project Structure

Contributing

  1. Branch from main (e.g. feature/my-change).
  2. Follow STYLE.md and per-directory ARCHITECTURE.md files.
  3. Add tests in tests/ (see tests/ARCHITECTURE.md).
  4. make test && make lint (use make format to auto-fix).
  5. Open a PR against main.

License

Licensed under the Apache License 2.0.

Release files for emmy-llm 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for emmy-llm 0.1.1
File Size Uploaded
emmy_llm-0.1.1.tar.gz 908.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for emmy-llm 0.1.1
File Interpreter ABI Platform
emmy_llm-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 2.0 MB

Release files / emmy_llm-0.1.1.tar.gz

Download URL emmy_llm-0.1.1.tar.gz
Size 908.7 kB
Tags Source
SHA-256 checksum
How to use checksums
6fbeb410d0ee18538a68a2aa7b70b802ea00a5c522f095a5b737e637db1722cb
BLAKE2b-256 checksum
How to use checksums
82b60d2a57d7a5e488916f0cc4fb4af23f43756c101d2061d30b765af0063ce7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 1, 2026.

Transparency log

Release files / emmy_llm-0.1.1-py3-none-any.whl

Download URL emmy_llm-0.1.1-py3-none-any.whl
Size 1.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
0298712db93e825ef12a2a37610f2bea84eca0b0810017b9c33a642d540989f3
BLAKE2b-256 checksum
How to use checksums
e82cd24119400ea68238c180cac16077140caba852e8e30861ea9fb5098ed3d9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page