Skip to main content

glc-loader

Read, verify and serve GLC compressed model artifacts — with no checkout of the repository that built them.

The compression is bit-exact: the weights this package hands back are the same bytes the source checkpoint had, and glc-loader verify --deep proves that on your machine, from the artifact alone. Runtime speed depends on the backend and hardware; see Hardware reality.


Install

pip install glc-loader

That is enough to inspect, verify, expand and CPU-load any artifact. Two optional extras add accelerated serving; neither is required to import the package, and neither is a dependency of the other:

pip install "glc-loader[cuda]"    # + triton and Ninja for CUDA JIT backends
pip install "glc-loader[metal]"   # + mlx, mlx-lm, for Apple Silicon
pip install "glc-loader[qwen38]"  # Qwen3.8 multimodal runtime and processor

Requires Python ≥ 3.10, torch ≥ 2.1, safetensors ≥ 0.4, transformers ≥ 4.56.

Tested against torch 2.6.0, safetensors 0.8.0, transformers 5.13.0 on macOS/arm64, Python 3.11. The version floors above are declared, not measured — they are the lowest releases whose APIs this code uses, not the lowest that have been run.


The three commands

glc-loader info <artifact-dir>

What the artifact is, where it came from, its ratios at all three scopes, and — the part that usually goes missing — what it does not certify.

glc-loader info ./my-artifact

Ratios come back as three separate blocks, each with its own basis string and a status of measured or unmeasured. A scope this artifact does not carry comes back null. It is never filled in from another scope:

scope what it counts
stored bytes on disk
served_resident bytes the device holds for weights
whole_process measured peak process VRAM, dense vs coded — the only one that includes activations, the KV cache and allocator slack

whole_process is normally unmeasured. The transcode step leaves it blank by construction; it is filled only by a certify run on the target card. If info tells you it is unmeasured, that is the truth, not a gap in the tool.

glc-loader verify <artifact-dir>

Integrity, with exit codes you can put in CI:

exit meaning
0 verified
1 an integrity check failed — a digest or a decoded tensor did not match
2 malformed, unreadable, or not a GeoRefine container at all
3 verified bit-exact, but the artifact expands (stores more bytes than the dense weights it replaces)
4 the format is recognised but this directory cannot be verified from what it ships
glc-loader verify ./my-artifact          # digests
glc-loader verify ./my-artifact --deep   # + decode every container

Without --deep this checks digests only, and says so in the output. With --deep it decodes every stored container and checks the result against the source digest recorded before encoding — an end-to-end proof, reproducible on your machine, that the bytes this artifact produces are the bytes the source model had.

Exit 4 is deliberate. A glc_tbe_transcode v1 directory is a correct artifact with no self-contained verifier: it ships a manifest and two tensor blobs, and nothing binds the manifest to the blobs, so checking one against the other would be checking the manifest against itself. Reporting that as exit 2 ("malformed") would be a lie about a well-formed artifact, and reporting it as 0 would be a lie about what was checked.

glc-loader load <artifact-dir>

Smoke-load the artifact, generate a few tokens, print the serving receipt.

glc-loader load ./my-artifact --device cuda --prompt "The capital of France is"

This is a smoke test, not a benchmark. The tokens_per_second it prints is one greedy run on whatever device you gave it.

Two more commands exist for GLC-RELEASE/1 artifacts specifically: glc-loader expand --out ./dense writes a plain Hugging Face checkpoint for any other engine, and glc-loader generate --prompt ... is the older, more configurable generation path.


From Python

from glc_loader import load_model

model, tokenizer, receipt = load_model("./my-artifact", device="cuda")
print(receipt["backend"])

For a georefine.tbe.v2 container:

from glc_loader import load_standalone, single_device_map

model, tokenizer, receipt = load_standalone(
    "./my-artifact", device_map=single_device_map("cuda:0"),
)

Hardware reality

Qwen3.8-27B serve-v1 compatibility prototype

The georefine.tbe.serve.v1 bundle has a source-free compressed PyTorch path:

from glc_loader import load_compressed_transformers

model = load_compressed_transformers(
    "Tsotchke-Corporation/Qwen3.8-27B-GeoRefine-TBE",
    expect_manifest_sha256="10028416e07c802b47e4dc93a86e9f0dd35ce6021f0e07b96c70bd9cb7a72bc3",
)

Install glc-loader[qwen38] for the required Qwen Transformers class and Hub download support. The bundle remains encoded in safetensors arrays (planes, smb, esc, sbbase). PortableTBELinear decodes one weight transiently per call, performs a PyTorch linear operation, and discards the dense weight. This is a correctness and portability fallback, not the measured fused-kernel speed path. The base Qwen model's 1,184 tensors map exactly to the manifest; the remaining 15 MTP weights are retained for native speculative inference. The full 27B graph loaded on Linux CPU and produced finite text and minimal image logits while retaining encoded weights. Full generation on this CPU reference path is unverified, and its forward speed is not a serving claim.

Decode speed is a property of the engine, not of the format. The same GLC-TBE weights have been measured on two decode paths:

decode path card measured scope
portable PyTorch reference Linux CPU full 27B text/image forwards passed; too slow for serving
standard glc-serve HTTP, exact Blackwell RTX PRO 6000 7.99 vs dense 17.50 tok/s end to end; 8/8 text + image parity
standard glc-serve HTTP, fused Blackwell RTX PRO 6000 14.95 vs dense 17.50 tok/s; 5/8 text parity
FastSession whole-step CUDA graph, bundled SM120 tune Blackwell RTX PRO 6000 37.880 vs dense 28.288 decode tok/s (1.339×); 8/8 text token-ID + image parity

The standard HTTP and FastSession rows use eight fixed greedy prompts and one synthetic image on one card. The FastSession result was reproduced from an installed wheel and its packaged tune; it is not a cross-platform speed guarantee or a general capability certificate.

Earlier standalone-loader measurements were 0.94× dense on RTX PRO and 0.57–0.62× on A100. Those were single runs with roughly ±0.05 spread on a different decode path; they do not measure the packaged FastSession engine.

The engine row is Qwen3.8-27B, TBE bundle, batch 1, greedy AR on one card (receipts in KERNEL_SPEED_20260925.md): 37.84 tok/s, against 28.29 tok/s for the bf16 parent in the same engine and 27.4–27.6 tok/s for llama.cpp bf16 (tg128) on the same card. Peak VRAM (NVML) was 49.9–50.0 GB, against 65.4–72.8 GB for the dense engine. The engine's output is bitwise equal to the bf16 parent on its gate probes (G3a 711/711 rows).

Memory is the saving that holds on every path. For orientation, one real artifact (unsloth/Llama-3.2-1B-Instruct, GLC-FWP1) reports a weight ratio of 1.279× at the served-graph scope — 1.93 GB resident where the dense checkpoint needs 2.47 GB. Your artifact's own numbers are in glc-loader info; do not carry that one across.

On the tested RTX PRO 6000, choose FastSession for the measured fast path. The standard HTTP route and portable CPU path have different speed contracts.

from huggingface_hub import snapshot_download
from glc_serve.fastserve import FastSession

bundle = snapshot_download("Tsotchke-Corporation/Qwen3.8-27B-GeoRefine-TBE")
session = FastSession.load(
    bundle=bundle, tune=None,  # bundled tune: RTX PRO 6000 Blackwell only
    server_flags=["--gate", "full", "--manifest-sha256",
                  "10028416e07c802b47e4dc93a86e9f0dd35ce6021f0e07b96c70bd9cb7a72bc3"],
)
state = session.prefill(messages=[{"role": "user", "content": "Hello"}])
ids = list(session.generate(state, max_tokens=48, temperature=0.0))
print(session.tokenizer.decode(ids, skip_special_tokens=True))

The public Hugging Face manifest hash above differs from the internal GCS source manifest because private build paths were removed. Allow room for the 40 GB bundle and about 50 GB peak GPU memory on the measured SM120 path. An RTX PRO 6000 is the only card whose tune and full-model speed have been verified in this wheel; other devices need an explicit tune and separate correctness and performance measurements.


CUDA package scope

This wheel includes the TBE MMA source (tbe_mma_kernel.cu), FastDecoder and MIV-TBE CUDA sources, plus the measured RTX PRO 6000 SM120 tune. They JIT-build in a user cache; a CUDA toolkit, C++ compiler, Python headers and Ninja are required. FastSession.load selects the bundled tune only on the measured SM120 card. Other GPUs require an explicit tune and their own exactness and speed tests. The experimental FWP1 Triton source is still artifact-local (fwp1_kernels.py) and is not in this wheel; that optional backend refuses or uses its documented fallback.


Which formats this reads

format self-contained? verify load
GLC-RELEASE/1 (FWP1) yes digests, and --deep bit-exactness yes
georefine.tbe.v2 yes SHA256SUMS, and --deep bit-exactness yes, on a supported device
glc_tbe_transcode v1 no — no config, no tokenizer, no format tag exit 4, with the reason refused
anything else — exit 2 refused

Format is decided by a declared tag — MANIFEST.json's "format", or compression_info.json's "artifact_format" — never by which files happen to be on disk. A plain Hugging Face checkpoint ships config.json and model.safetensors, and a loader that inferred "container" from those would claim every model on the hub. Point this at an ordinary checkpoint and it declines, by design.


Licence

The compressed weights inherit the source model's licence. glc-loader info reports whatever licence the artifact declares or ships, and reports null with a note when it declares none — which is the common case today. Check the source model before redistributing.

This package is proprietary to Tsotchke Corporation.

Metadata

Release files for glc-loader 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for glc-loader 1.1.0
File Size Uploaded
glc_loader-1.1.0.tar.gz 298.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for glc-loader 1.1.0
File Interpreter ABI Platform
glc_loader-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 634.8 kB

Release files / glc_loader-1.1.0.tar.gz

Download URL glc_loader-1.1.0.tar.gz
Size 298.7 kB
Tags Source
SHA-256 checksum
How to use checksums
fab8117097ceabed1648b323985faf8236194cd1db59d4c59d7cbf762dd3aa1a
BLAKE2b-256 checksum
How to use checksums
4121c54827f7346f7c4ab2b820991f4c9f780071efa9b13c4e7d3e24cdcfbda3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release files / glc_loader-1.1.0-py3-none-any.whl

Download URL glc_loader-1.1.0-py3-none-any.whl
Size 336.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e7ae90ef2cc046e1fa5a4131e6c5da1ec7e68d4dc96c489f86a9ed12b9170226
BLAKE2b-256 checksum
How to use checksums
4cb5e5eaa61c2fc1c71e68ee8fa658e0263c82d66040c43c4e07b6f64f58dda7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release history Release notifications | RSS feed

1.1.1

2 release files

This release

1.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page