Skip to main content

llmtrim-uniffi

UniFFI bindings over llmtrim-core: one Rust definition, idiomatic in-process bindings for Python, Ruby, Swift and Kotlin. The compression runs natively in the caller's process (no server, no async).

API

A deliberately flat surface over the engine:

fn compress(
    input: String,                 // a provider-shaped request body (JSON)
    provider: Option<Provider>,    // OpenAi | Anthropic | Google, or None to auto-detect
    preset: Option<String>,        // "aggressive" | "agent" | "code" | "rag" | "safe" | …
                                   // None = config from the environment / config file
) -> Result<CompressOutput, LlmtrimError>

CompressOutput carries the compressed request_json, the resolved provider/model, the tokenizer label/exactness, and the before/after/frozen input-token counts. Embedders that need the full rehydration plan or per-stage reports should depend on llmtrim-core directly in Rust.

In-process vs. the proxy

llmtrim has two integration routes:

  • The proxy (HTTPS_PROXY=127.0.0.1:8788 llmtrim) intercepts your existing traffic and compresses it in flight. Nothing in your code changes, but the client has to route through the proxy and trust its CA.
  • These bindings compress in your process. You call compress() on the request body, then send the result with your own HTTP client. No proxy, no CA, no env-var setup.

Use the in-process path when the proxy can't run:

  • Sandboxed / serverless functions where you can't set a process-wide HTTPS_PROXY or run a side process.
  • Certificate-pinned clients that reject a MITM CA, so the proxy's interception fails.
  • Anywhere you'd rather not add a network hop or an extra moving part.

It replaces a per-framework adapter: instead of wiring a hook into each SDK, you compress the body once and POST it yourself. Runnable end-to-end examples (compress, then send with your own client) are in examples/.

Python

# Build a self-contained wheel (cdylib + generated glue):
crates/llmtrim-uniffi/scripts/build-wheel.sh --release
pip install target/wheels/llmtrim-*.whl
import llmtrim, json

req = json.dumps({"model": "gpt-4o",
                  "messages": [{"role": "user", "content": "…"}]})
out = llmtrim.compress(req, llmtrim.Provider.OPEN_AI, "aggressive")
print(out.input_tokens_before, "->", out.input_tokens_after)
# send out.request_json to the provider

Why build-wheel.sh and not plain maturin build: maturin's bindings = "uniffi" auto-packaging is sensitive to the maturin↔uniffi version pair. With maturin 1.14 + uniffi 0.31 it builds the native library into the wheel but omits the generated Python glue (empty package __init__.py). The script runs maturin, then injects the freshly generated bindings and repacks the wheel with valid RECORD hashes. Remove it once the auto path packages cleanly.

Ruby / Swift / Kotlin

All targets generate from the same built library, no extra Rust. The generated glue is a build artifact (its checksums are pinned to the library ABI), so it is regenerated per release rather than committed:

crates/llmtrim-uniffi/scripts/generate-bindings.sh out/   # python, ruby, swift, kotlin

Generation needs an unstripped library. Library-mode uniffi-bindgen reads metadata symbols from the cdylib, but the workspace release profile sets strip = true. The script therefore generates from the (unstripped) debug build; the native library you ship can be a stripped cargo build --release -p llmtrim-uniffi cdylib; the glue loads it by name.

Ruby (verified). This is the raw generated binding (module LlmtrimFfi), for a source build with libllmtrim_ffi.so on the load path. The published gem aliases it to Llmtrim (require "llmtrim" then Llmtrim.compress(...)); see packaging/ruby.

require_relative "llmtrim_ffi"
require "json"
out = LlmtrimFfi.compress(
  JSON.generate({model: "gpt-4o", messages: [{role: "user", content: "…"}]}),
  LlmtrimFfi::Provider::OPEN_AI, "aggressive")
puts "#{out.input_tokens_before} -> #{out.input_tokens_after}"

Swift emits llmtrim_ffi.swift + an FFI header and modulemap; Kotlin emits uniffi/.../llmtrim_ffi.kt (which loads the cdylib via JNA). CI compiles and runs a smoke for both: Swift on macOS (swiftc against the modulemap), Kotlin on a JVM (kotlinc + JNA), so a binding break is caught in all four languages (see tests/swift, tests/kotlin and the bindings* jobs in .github/workflows/ci.yml).

Publishable packages

Each ships the compiled engine bundled, so consumers need no Rust toolchain:

Target Build Package Verified
Python (PyPI) scripts/build-wheel.sh wheel locally
Ruby (gem) scripts/build-gem.sh packaging/ruby/ locally
Kotlin/JVM (Maven) scripts/build-maven.sh packaging/kotlin/ locally
Swift (SwiftPM) scripts/build-xcframework.sh packaging/swift/ macOS CI only

Each packaging/<lang>/README.md has the usage + publish details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

llmtrim-0.12.0-py3-none-win_amd64.whl (8.0 MB view details)

Uploaded Python 3Windows x86-64

llmtrim-0.12.0-py3-none-manylinux_2_34_x86_64.whl (8.5 MB view details)

Uploaded Python 3manylinux: glibc 2.34+ x86-64

llmtrim-0.12.0-py3-none-manylinux_2_34_aarch64.whl (8.4 MB view details)

Uploaded Python 3manylinux: glibc 2.34+ ARM64

llmtrim-0.12.0-py3-none-macosx_11_0_arm64.whl (8.0 MB view details)

Uploaded Python 3macOS 11.0+ ARM64

llmtrim-0.12.0-py3-none-macosx_10_12_x86_64.whl (7.9 MB view details)

Uploaded Python 3macOS 10.12+ x86-64

File details

Details for the file llmtrim-0.12.0-py3-none-win_amd64.whl.

File metadata

  • Download URL: llmtrim-0.12.0-py3-none-win_amd64.whl
  • Upload date:
  • Size: 8.0 MB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for llmtrim-0.12.0-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 b34f85dabe42b4e9b12fb54e2bee2546557be82e557e074b0bbd783d9419e5ab
MD5 77354c4b3cf7fe433affd56af0375c97
BLAKE2b-256 ab726f08c27cb4feb145a7c3575f6864ffa1ba0ffe3f2ca832a66d52c79633dc

See more details on using hashes here.

Provenance

The following attestation bundles were made for llmtrim-0.12.0-py3-none-win_amd64.whl:

Publisher: release-bindings.yml on fkiene/llmtrim

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file llmtrim-0.12.0-py3-none-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for llmtrim-0.12.0-py3-none-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 90eb9e1b3d81461681a753ac4085e0bfdf6f3f15994a33aac34bb0fd1426c039
MD5 8d4f0d201c484da30d8ac3f93994700d
BLAKE2b-256 c60706a1d4d5a4e53c5871a1bb6316195dd0ae15e1dc132b79ddb83103cafbe6

See more details on using hashes here.

Provenance

The following attestation bundles were made for llmtrim-0.12.0-py3-none-manylinux_2_34_x86_64.whl:

Publisher: release-bindings.yml on fkiene/llmtrim

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file llmtrim-0.12.0-py3-none-manylinux_2_34_aarch64.whl.

File metadata

File hashes

Hashes for llmtrim-0.12.0-py3-none-manylinux_2_34_aarch64.whl
Algorithm Hash digest
SHA256 bcce5a42fc9c1e580454e720e372e45e3604ad89a8271c529386db04753e7a7e
MD5 1d97577d323d5ebccb3037c6f0b61fd5
BLAKE2b-256 38413b25dba4c3d16f1db1e85f15499686e2d507f248c7f2a9c4e33c1cc5d7ed

See more details on using hashes here.

Provenance

The following attestation bundles were made for llmtrim-0.12.0-py3-none-manylinux_2_34_aarch64.whl:

Publisher: release-bindings.yml on fkiene/llmtrim

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file llmtrim-0.12.0-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for llmtrim-0.12.0-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 b65b0f87fa892f7617286c9e742e6566c29525c7b74a08ef9a136d34c64406ae
MD5 4fcb71f409a5fb15ef3a523dfbebe1ed
BLAKE2b-256 a786a83c706902ba67ddeacc093ec49beb3d51c55aff126151c0c891ce802c7d

See more details on using hashes here.

Provenance

The following attestation bundles were made for llmtrim-0.12.0-py3-none-macosx_11_0_arm64.whl:

Publisher: release-bindings.yml on fkiene/llmtrim

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file llmtrim-0.12.0-py3-none-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for llmtrim-0.12.0-py3-none-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 b39dd11e4a8bcce337656152da09965bf8203b648286212f818f8eaf9b451868
MD5 0f389bfe9a150b2c92cfbaf2f9ffcadd
BLAKE2b-256 f0bfc294951920d9836f5bbc05eb9da9172c5da2da9fa9fd1bc14145e81d958c

See more details on using hashes here.

Provenance

The following attestation bundles were made for llmtrim-0.12.0-py3-none-macosx_10_12_x86_64.whl:

Publisher: release-bindings.yml on fkiene/llmtrim

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page