genlm-backend

No project description provided

Project description

Logo

GenLM Backend is a high-performance backend for language model probabilistic programs, built for the GenLM ecosystem. It provides an asynchronous and autobatched interface to vllm and transformers language models, enabling scalable and efficient inference.

See our documentation.

🚀 Key Features

Automatic batching of concurrent log-probability requests, enabling efficient large-scale inference without having to write batching logic yourself
Byte-level decoding of transformers tokenizers, enabling advanced token-level control
Support for arbitrary Hugging Face models (e.g., LLaMA, DeepSeek, etc.) with fast inference and automatic KV caching using vllm
NEW: support for MLX-LM library, allowing faster inference on Apple silicon devices.

⚡ Quick Start

This library supports installation via pip:

pip install genlm-backend

Or to install with MLX support, run:

pip install genlm-backend[mlx]

Or to install with LoRA support, run:

pip install genlm-backend[lora]

🧪 Example: Autobatched Sequential Importance Sampling with LLMs

This example demonstrates how genlm-backend enables concise, scalable probabilistic inference with language models. It implements a Sequential Importance Sampling (SIS) algorithm that makes asynchronous log-probabality requests which get automatically batched by the language model.

import torch
import asyncio
from genlm.backend import load_model_by_name

# --- Token-level masking using the byte-level vocabulary --- #
def make_masking_function(llm, max_token_length, max_tokens):
    eos_id = llm.tokenizer.eos_token_id
    valid_ids = torch.tensor([
        token_id == eos_id or len(token) <= max_token_length
        for token_id, token in enumerate(llm.byte_vocab)
    ], dtype=torch.float).log()
    eos_one_hot = torch.nn.functional.one_hot(
        torch.tensor(eos_id), len(llm.byte_vocab)
    ).log()

    def masking_function(context):
        return eos_one_hot if len(context) >= max_tokens else valid_ids

    return masking_function

# --- Particle class for SIS --- #
class Particle:
    def __init__(self, llm, mask_function, prompt_ids):
        self.context = []
        self.prompt_ids = prompt_ids
        self.log_weight = 0.0
        self.active = True
        self.llm = llm
        self.mask_function = mask_function

    async def extend(self):
        logps = await self.llm.next_token_logprobs(self.prompt_ids + self.context)
        masked_logps = logps + self.mask_function(self.context).to(logps.device)
        logZ = masked_logps.logsumexp(dim=-1)
        self.log_weight += logZ
        next_token_id = torch.multinomial((masked_logps - logZ).exp(), 1).item()
        if next_token_id == self.llm.tokenizer.eos_token_id:
            self.active = False
        else:
            self.context.append(next_token_id)

# --- Autobatched SIS loop --- #
async def autobatched_sis(n_particles, llm, masking_function, prompt_ids):
    particles = [Particle(llm, masking_function, prompt_ids) for _ in range(n_particles)]
    while any(p.active for p in particles):
        await asyncio.gather(*[p.extend() for p in particles if p.active])
    return particles

# --- Run the example --- #
llm = load_model_by_name("gpt2") # or e.g., "meta-llama/Llama-3.2-1B" if you have access
mask_function = make_masking_function(llm, max_token_length=10, max_tokens=10)
prompt_ids = llm.tokenizer.encode("Montreal is")
particles = await autobatched_sis( # use asyncio.run(autobatched_sis(...)) if you are not in an async context
    n_particles=10, llm=llm, masking_function=mask_function, prompt_ids=prompt_ids
)

strings = [llm.tokenizer.decode(p.context) for p in particles]
log_weights = torch.tensor([p.log_weight for p in particles])
probs = torch.exp(log_weights - log_weights.logsumexp(dim=-1))

for s, p in sorted(zip(strings, probs), key=lambda x: -x[1]):
    print(f"{repr(s)} (probability: {p:.4f})")

This example highlights the following features:

🌀 Asynchronous Inference Loop. Each particle runs independently, but all LLM calls are scheduled concurrently via asyncio.gather. The backend batches them automatically, so we get the efficiency of large batched inference without having to write the batching logic.
🔁 Byte-level Tokenization Support. Token filtering is done using the model’s byte-level vocabulary, which genlm-backend exposes. This enables low-level control over generation in ways not possible with most high-level APIs.

Development

See the DEVELOPING.md file for information on how to install the project for local development.

Project details

Release history Release notifications | RSS feed

0.2.2

Apr 28, 2026

0.2.1 yanked

Apr 27, 2026

This version

0.2.0

Apr 7, 2026

0.1.8

Jan 30, 2026

0.1.7

Nov 13, 2025

0.1.6

Oct 17, 2025

0.1.5

Sep 4, 2025

0.1.4

Aug 4, 2025

0.1.3

Jul 29, 2025

0.1.2

Jul 28, 2025

0.1.1

May 6, 2025

0.1.0

Apr 10, 2025

0.1.0a1 pre-release

Apr 9, 2025

0.1.0a0 pre-release

Apr 9, 2025

0.0.2

Mar 13, 2025

0.0.1

Mar 6, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

genlm_backend-0.2.0.tar.gz (3.8 MB view details)

Uploaded Apr 7, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

genlm_backend-0.2.0-py3-none-any.whl (38.8 kB view details)

Uploaded Apr 7, 2026 Python 3

File details

Details for the file genlm_backend-0.2.0.tar.gz.

File metadata

Download URL: genlm_backend-0.2.0.tar.gz
Upload date: Apr 7, 2026
Size: 3.8 MB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for genlm_backend-0.2.0.tar.gz
Algorithm	Hash digest
SHA256	`788bae690294d19b1fe054f56ad5b12bc6f099c2eec707ff2579b03176506f29`
MD5	`48f76385cbb261dbc3048a317317ab5e`
BLAKE2b-256	`2d37279029dc267539f06cf5f3942ced856b0772eb673672e0515a4671400cb7`

See more details on using hashes here.

File details

Details for the file genlm_backend-0.2.0-py3-none-any.whl.

File metadata

Download URL: genlm_backend-0.2.0-py3-none-any.whl
Upload date: Apr 7, 2026
Size: 38.8 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for genlm_backend-0.2.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`26b88432bfd260edce38bb029b1ef99bb0528221d4d094e8000c266e3abca208`
MD5	`4b5d45a7b03536bd90cb83b225f5e075`
BLAKE2b-256	`9e9d321ad97c324cb420f710ffa690f3ff1edf0befbf59c20e2019e4d7fc4620`

See more details on using hashes here.

genlm-backend 0.2.0

Navigation

Verified details

Maintainers

Unverified details

Meta

Project description

🚀 Key Features

⚡ Quick Start

🧪 Example: Autobatched Sequential Importance Sampling with LLMs

Development

Project details

Verified details

Maintainers

Unverified details

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes