Skip to main content

TensorCache ⚡ PyPI License: Apache-2.0

Ultra-fast, high-fidelity block-wise INT8 feature & pixel cache engine for PyTorch.

pip install tcacheimport tensorcache or import tcache0.2.2 on PyPI.

tensorcache eliminates two bottlenecks:

  1. Feature Cache Bloat: AMO-BQ asymmetric MSE-optimal G32 1.09B 1.83x vs BF16 0.47% rel RMSE (sym 1.06B 0.54%) — near G16 floor 0.39%.
  2. JPEG/PNG CPU Decode: Zero-copy mmap + GPU stream prefetch >2,000 MB/s, ring-buffer 6.8MB VRAM batch8 128x768.
  3. Training Throughput (v0.2.2): iter_batches 27k samp/s 17.7GB/s B128 446x768 ~8.3x vs v0.2.0 3.1k, Streamer 21k 5.1ms double-buffered pinned + async H2D (low_vram 128MB B128), sharded 8x for H100 DDP.

🚀 Key Features

  • AMO-BQ (Asymmetric MSE-Optimal, G32): min-max + zp uint8 + 48× clipping search [0.95,1.10]0.47% hetero 0.54% vs sym 0.74% (-26%), per-token 1.73% (cache\dtype_comparison.json). Presets fast/balanced/accurate/max.
  • Microsecond Dequant (v0.2.2 opt): Triton shift >>5 vs div, autotune BLOCK 512/1024/2048 stages 2/3 163 GB/s 5.4M (codec.py:22 0.19ms 10M 26x for 1024), FusedDequantLinear 0 intermediate fused_ops.py:119. NaN/Inf isolated 0 poison 1e-8 scale.
  • Training-Optimized Dataloader: iter_batches 27k samp/s 17.7GB/s B128 4.5ms (8.3x vs 3.1k), Streamer 21.6k 5.1ms double-buffered pinned async H2D overlap dequant. low_vram=True 128MB B128 (53 MB B32) vs 256MB.
  • Big Data Sharded: Writer(num_shards=8) -> feat_shard{i}_*.bin + feat_shards.json, Dataset/Streamer(rank, world_size) DDP H100 8x ~5.8k random 27k contiguous. Single shard num_shards=1 unchanged 100% compat.
  • Minimal VRAM: q 5.22MB + scales 0.32MB + zp 0.16MB + out 10.45MB 5.4M; G64 halves scales/zp.
  • Cross-Platform: CUDA/ROCm Triton else PyTorch fallback, Windows mmap safe close().
  • CLI + Python one-liners: tc.compress / tc.benchmark_tensor / tensorcache benchmark.

📦 Installation

pip install tcache                # PyPI (import tensorcache or tcache)
pip install -e .          # dev
pip install -e ".[fast]"  # triton+zstd+blosc2 (Linux)

⚡ Quick Start

1. In-Memory (one-liners)

import torch, tensorcache as tc

x = torch.randn(16,446,768, dtype=torch.bfloat16, device="cuda")

# AMO-BQ presets: fast (16,0.95-1.05) 6.9ms 0.49%, balanced (32,0.95-1.05) 13ms 0.478% (default), accurate (48,0.95-1.10) 49ms 0.473%
q,s,zp,shape = tc.compress(x, mode="balanced")  # or "fast"/"accurate"/"max"/"sym"/"adaptive"
rec = tc.decompress(q,s,shape,zp)               # <0.1ms BF16
tc.benchmark_tensor(x)                          # rich table
tc.estimate_compression(x.shape, group_size=32) # 1.09375 B 1.83x
tc.auto_select_mode(x, target_rmse=0.5)         # -> "balanced"
tc.help()                                       # python help

# Codec object
codec = tc.BlockwiseInt8Codec(group_size=32, amo_bq=True, amo_mode="balanced")
print(codec) # G=32, amo_bq=balanced 1.0938B 1:1.83x

2. Feature Cache to Disk (mmap) - Training Optimized

import tensorcache as tc
from torch.utils.data import DataLoader

# Write (amo_bq, G32 default balanced, G64 for minimal VRAM 1.046B 0.55%)
# Single shard (default, unchanged)
writer = tc.FeatureCacheWriter("./cache/dinov3", 10000,446,768, group_size=32, amo_bq=True, amo_mode="balanced")
# Or sharded for H100 / >1M samples (8 shards ~1.25k each)
writer = tc.FeatureCacheWriter("./cache/dinov3", 10000,446,768, group_size=32, amo_bq=True, amo_mode="balanced", num_shards=8)
writer.append(x) # [seq,dim] or [B,seq,dim] (batched GPU 64/chunk, CPU per-sample)
writer.close()   # single: _int8.bin ... | sharded: _shard0_int8.bin ... + _shards.json

# Load - auto-detects sharded
ds = tc.FeatureCacheDataset("./cache/dinov3") # -> (q uint8, s BF16, zp uint8) all shards
# DDP per-rank (H100 8x)
ds = tc.FeatureCacheDataset("./cache/dinov3", rank=rank, world_size=8) # only shard rank
# or explicit
ds = tc.FeatureCacheDataset("./cache/dinov3", shard_idx=0)
ds = tc.FeatureCacheDataset("./cache/dinov3", auto_dequant_device="cuda") # -> BF16 directly

# Old path still works but slower (~5k samp/s)
for q,s,zp in tc.AsyncGPUPrefetcher(DataLoader(ds,batch_size=256,pin_memory=True), device="cuda"):
    batch = tc.dequantize_int8_amo_bq(q,s,zp, shape, group_size=32)

# Training fastest (27k samp/s 17.7GB/s B128, 8.3x vs v0.2.0)
for batch in ds.iter_batches(batch_size=128, shuffle=False, device="cuda"): # contiguous 27k, shuffled ~5.8k
    train(batch)

# Minimal VRAM streamer (double-buffered 64MB B32 / 256MB B128, low_vram 32MB/128MB)
streamer = tc.ZeroCopyTensorStreamer("./cache/dinov3", batch_size=128, device="cuda", shuffle=True)
streamer = tc.ZeroCopyTensorStreamer("./cache/dinov3", batch_size=128, device="cuda", low_vram=True) # 128MB B128
streamer = tc.ZeroCopyTensorStreamer("./cache/dinov3", batch_size=128, device="cuda", shard_idx=0) # per-rank
for batch in streamer: # BF16 [B,seq,dim] from ring buffer out_bf16_0/1
    train(batch)
streamer.close()

# Fused head (no BF16 intermediate)
from tensorcache import FusedDequantLinear
head = FusedDequantLinear(768, num_classes, group_size=32).cuda()
logits = head(q, s) # dequant+GEMM in regs

3. CLI

python -m tensorcache info
python -m tensorcache benchmark --shape 16,446,768 --device cuda
python -m tensorcache cache-info --prefix ./cache/dinov3
python -m tensorcache compress-demo --shape 4,197,768 --mode balanced
tensorcache --help

📊 Benchmark (DINOv3 ViT-Base, 5.4M randn + RSNA hetero)

Format B/elem vs BF16 rel RMSE Outlier 0.1% Dequant
Raw BF16 2.00 1.00x 0.167% 0.119%
Naive FP8 E5M2 1.00 2.00x 5.24% 6.81%
Naive FP8 E4M3 1.00 2.00x 2.63% 3.19%
MXFP8 G32 1.09 1.83x 2.38%
Sym G32 1.06 1.88x 0.540% 0.11% 0.036ms 435GB/s
AMO-BQ fast G32 1.093 1.83x 0.490% 0.15% 0.06ms 337GB/s
AMO-BQ balanced G32 1.093 1.83x 0.478% 0.15% 13ms quant
AMO-BQ accurate G32 1.093 1.83x 0.473% 49ms quant
AMO-BQ G16 1.187 1.68x 0.395% 13ms
Sym G64 1.046 1.91x 0.55%

📜 License

Apache 2.0 — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tcache-0.2.3.tar.gz (67.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tcache-0.2.3-py3-none-any.whl (47.7 kB view details)

Uploaded Python 3

File details

Details for the file tcache-0.2.3.tar.gz.

File metadata

  • Download URL: tcache-0.2.3.tar.gz
  • Upload date:
  • Size: 67.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.11

File hashes

Hashes for tcache-0.2.3.tar.gz
Algorithm Hash digest
SHA256 dc865d3a8d812001a6cc1202b1f628b344ecb40357c95106398762b37f7b6f2f
MD5 f65d90fd9f678ab255a891364b950b9f
BLAKE2b-256 5acb085ac0541ef7ebf5dd3c9affe9f1fd9195a21b5fae0fc00d372132bb1008

See more details on using hashes here.

File details

Details for the file tcache-0.2.3-py3-none-any.whl.

File metadata

  • Download URL: tcache-0.2.3-py3-none-any.whl
  • Upload date:
  • Size: 47.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.11

File hashes

Hashes for tcache-0.2.3-py3-none-any.whl
Algorithm Hash digest
SHA256 99065ccdb76114d2faebadc185f70bb80915ae0e4d205a1a69803f7577f93ebc
MD5 23a4eb00667e4c3dd06622ab4e137056
BLAKE2b-256 b13672612b02d569f99e76fd0eb8a090544441d425e9238d7e16043e67d81a67

See more details on using hashes here.

Release history Release notifications | RSS feed

0.3.0

2 files

0.2.4

2 files

This release

0.2.3 This release

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page