Skip to main content

TileFoundry

Hand the compiler to the agent.

Not an agent that becomes the compiler,
and not an agent plugged in as one of its passes.
The compiler stays a tool — the agent is simply the one holding it.

One prompt in  ·  612 tok/s out  ·  Compiler in the loop


PyPI Coverage TileFoundry optimizations shipped to TileOPs License: MIT

Documentation · Quick Start · Examples

Latest News

  • 09/2026 🔧: TileFoundry 0.0.2 is on PyPI — bug fixes and refactoring across the analysis, parser, and IR layers.
  • 08/2026 ⚡: Nemotron-3.5-Lightning-30B-A3B — a 52-layer Mamba2 / attention / MoE hybrid in a single mega decode kernel, one launch a step with no CUDA graph, reaching 97.5% of SGLang at short context and 83.2% at 262144.
  • 08/2026 📦: Four worked examples added — Qwen3-1.7B (tilelang), Qwen3.5-35B-A3B (tilelang), MiniCPM3-4B (CuTeDSL) and granite-4.0-h-small (CUDA C) — each one a real agent run kept whole, with the decode throughput it measured.
Earlier
  • 08/2026 🎉: TileFoundry 0.0.1 is on PyPI — the first public release.

Quick Start

1 · Install

pip install tilefoundry    # needs Python 3.12 or newer
tilefoundry                # check the install: the commands an agent will ask

This run also needs one NVIDIA GPU, pip install tilelang, the published Qwen/Qwen3-1.7B checkpoint on disk (3.8 GB), and a coding agent started in an empty directory.

2 · Hand it the prompt

There is no API to learn first. Give your coding agent this, with a checkpoint directory of your own:

Get real tokens out of Qwen3-1.7B on TileFoundry, and make it fast.
Weights and config: <checkpoint directory>
Backend: tilelang.

Everything about TileFoundry is to be asked of the `tilefoundry` command -- do not
ask a person, do not go looking elsewhere. The model itself is yours to research.

Done when this runs from outside, prints the continuation, and reports a
tokens-per-second number measured over the whole generation:

    python run.py \
        --prompt "Write a detailed explanation of how a GPU executes a matrix multiplication." \
        --max-new-tokens 2048

Measure over a long generation -- 2048 new tokens, more than 2000 characters of
text. A 32-token sample is too short for the number to mean anything.

That is the whole input — nothing under it is written by hand.

3 · Come back in two hours

Claude Opus 5 at xhigh reasoning effort worked 2.1 hours and 177 tool calls without a single interaction, and left a run.py behind — it prints the continuation, and the number it measured: 612 tok/s on one H200.

python run.py --ckpt <checkpoint directory> \
    --prompt "Write a detailed explanation of how a GPU executes a matrix multiplication." \
    --max-new-tokens 2048

The first run compiles the kernels — once, a few minutes.

Where to go next

That run is kept whole in examples/qwen3_1_7b-tilelang/, with three more beside it. The specifications are meant to be argued with: open an issue, or start from docs/develop.md.

License

This project is licensed under the MIT License.

Release files for tilefoundry 0.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tilefoundry 0.0.2
File Size Uploaded
tilefoundry-0.0.2.tar.gz 1.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for tilefoundry 0.0.2
File Interpreter ABI Platform
tilefoundry-0.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 2.0 MB

Release files / tilefoundry-0.0.2.tar.gz

Download URL tilefoundry-0.0.2.tar.gz
Size 1.1 MB
Tags Source
SHA-256 checksum
How to use checksums
38158909617e4e56f8a5e800db7e5d39d79cf960bd7db90c3ae491269f0fa741
BLAKE2b-256 checksum
How to use checksums
665db8a527842e18e56ff3a06bb5ca16cb8e46a6540356985119cbde8c53090b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.

Transparency log

Release files / tilefoundry-0.0.2-py3-none-any.whl

Download URL tilefoundry-0.0.2-py3-none-any.whl
Size 871.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
684379a871ee041140f432dd513335d12fe825701d7f13acd984532f2beb1a30
BLAKE2b-256 checksum
How to use checksums
bd397560fa7bc1723baa9cab6133bac3230119112aaf571fa1a6a3e33564b999
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.0.2 This release

2 release files

0.0.1

2 release files

0.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page