Skip to main content

PFLM: A Prior-Fitted Language Model

PFLM is a 300M parameter byte-level model pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of a structured byte sequence like natural language, the model infers the source in context and predicts what comes next.

This repository is the model, its transformers integration, and the byte API. The weights are at https://huggingface.co/lennartcb/pflm1.

Install

pip install "pflm1[hf]"

Importing pflm1 registers the model with transformers; the released checkpoints hold weights and config only. Optional GPU kernels: pip install "pflm1[kernels]" for the fused delta-rule scan, plus flash-attn and causal-conv1d wheels for your CUDA build; without them the model runs exact reference paths.

Use

import pflm1
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("lennartcb/pflm1", dtype="bfloat16").cuda().eval()

Watch it learn. The model has never seen the primes; the cost per digit falls as it reads them.

def primes(n):                                   # 1 where the integer is prime, else 0
    flags = bytearray([1]) * n
    flags[:2] = b"\0\0"
    for i in range(2, int(n ** 0.5) + 1):
        if flags[i]:
            flags[i * i::i] = bytes(len(flags[i * i::i]))
    return bytes(48 + f for f in flags)

text = primes(10_000)                            # "0011010100..." ten thousand digits
bits = model.bits_per_byte(text)                 # bits per byte, one entry each
print(bits[:500].mean(), bits[-500:].mean())     # the first 500 digits vs. the last 500

Score a stream. One entry per byte, in bits; the mean is bits per byte and the running mean against position is the in-context learning curve.

data = open("article.txt", "rb").read()
bits = model.bits_per_byte(data)
print(bits.mean(), bits[-1000:].mean())     # whole stream vs. the last 1 KB

Feed bytes as they arrive. The state is carried; nothing is re-read.

stream = model.stream()
first = stream.feed(b"The prior was never English")
more = stream.feed(b", and yet: ")
logits = stream.next_byte_logits()           # softmax for the next-byte probabilities

Sample a continuation. What the model expects after a context, as bytes; the new bytes only. Not a text generator: with no context there is no language yet.

model.generate_bytes(context, max_new=200, temperature=0.8)

AutoTokenizer and pipeline("text-generation") also work, through a tokenizer that maps each UTF-8 byte to its own id.

Model

A hybrid of gated delta-rule recurrence (Gated DeltaNet) and sliding-window attention, 3:1, with SwiGLU MLPs. The vocabulary is 256 raw bytes; there is no tokenizer to learn. Both mixers carry bounded state, the recurrence a fixed-size fast-weight matrix and the attention a window - 1 cache, so context length is unbounded, generation memory is constant, and a chunked forward is bit-exact against a full-sequence one, which is what the byte API relies on.

pflm1/
  args.py backbone.py block.py recurrence.py swa.py ...   pure-torch model
  scoring.py                                              the byte API
  configuration_pflm1.py modeling_pflm1.py                transformers model
  tokenization_pflm1.py cache_pflm1.py                    tokenizer and cache
tests/     CPU test suite

Pure torch, without transformers:

import json
from safetensors.torch import load_file
from pflm1 import Pflm1LM, ModelArgs, bits_per_byte

config = {k: v for k, v in json.load(open("config.json")).items() if k in ModelArgs.__dataclass_fields__}
model = Pflm1LM(ModelArgs(**config))
model.load_state_dict(load_file("model.safetensors"))
bits_per_byte(model.eval(), data)

Pin the attention backend with PFLM1_ATTN_BACKEND=fa4|fa3|fa2|sdpa if needed.

Citation

See CITATION.cff. Paper: Learning to Learn a Language (arXiv link to follow).

License

Apache-2.0, code and weights.

Metadata

Release files for pflm1 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pflm1 0.1.0
File Size Uploaded
pflm1-0.1.0.tar.gz 26.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pflm1 0.1.0
File Interpreter ABI Platform
pflm1-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 53.4 kB

Release files / pflm1-0.1.0.tar.gz

Download URL pflm1-0.1.0.tar.gz
Size 26.0 kB
Tags Source
SHA-256 checksum
How to use checksums
6e4a6e34c86ce45af48ffc7095fe5570adee7621101400b13185dbddd4cb5bdc
BLAKE2b-256 checksum
How to use checksums
ad03e1f98669eb55d0aaaca52f1094bb2d8bc7828283f09edf11b82a1066b8e4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / pflm1-0.1.0-py3-none-any.whl

Download URL pflm1-0.1.0-py3-none-any.whl
Size 27.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8c13ed752128539713291cfc674a58062a997b7c704ddc90a23ef8c26120d3f5
BLAKE2b-256 checksum
How to use checksums
aeea5ac8c6295af5b3e1786eeb7db5d6a182084c3493bd2e3fa8218de74ad696
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page