Skip to main content

flense

A lightweight reverse proxy that cuts your AI API costs by compressing large payloads before they reach expensive frontier models, with no changes to your application code.


The Problem

Frontier AI models (Claude Opus, GPT-4, Gemini Ultra) are billed per token. Most large requests are padded with stuff the model does not need to see in full: entire source files, verbose docs, repetitive boilerplate. You pay full price for all of it.

How Flense Helps

Flense sits between your application and the AI API. It intercepts outgoing requests, compresses heavy payloads using static code analysis, and forwards an optimised version to the frontier model. Your app sees no difference. Just a cheaper bill.

Your App  ->  localhost:2912 (flense)  ->  api.anthropic.com

No model calls. No quality trade-off for compression. Just fewer tokens.


Quick Start

pip install flense
flense start

Then point your SDK at flense with the provider prefix:

# Anthropic
client = Anthropic(base_url="http://localhost:2912/anthropic")

# OpenAI
client = OpenAI(base_url="http://localhost:2912/openai")

How It Works

Bulk-Reader

When a payload contains large code files or documents, flense compresses them using Tree-sitter, a fast, error-tolerant code parser that supports 100+ languages.

Instead of sending the full file, flense extracts:

  • Class and function signatures
  • Method names, parameters, return types
  • Line number anchors so the model knows where everything came from

Function bodies, comments, and docstrings are stripped. The frontier model gets a skeleton. Enough to reason accurately, at a fraction of the token cost.

No secondary model. No tokens spent on compression. Tree-sitter runs locally.

Fallback chain: Tree-sitter -> Universal Ctags -> regex heuristics

Tree-sitter grammars are optional and are not installed at runtime. Install them explicitly for AST-based compression; without them, flense falls back to ctags/regex:

pip install "flense[grammars]"

Code-Writer Bypass

For repetitive code generation tasks (scaffolding tests, mapping schemas), flense can route the request directly to a cheaper cloud model and write the output straight to disk, bypassing the frontier model entirely.

Because this feature reads and writes files on the host, it is disabled by default and must be explicitly enabled in config ([code_writer] enabled = true). Reference files are confined to allowed_ref_dir and generated output is confined to output_dir; paths outside those directories are rejected.

Triggered explicitly via a request header (once enabled):

headers={"X-Flense-Strategy": "code-writer"}

Telemetry

Flense appends savings data to every response header. Readable by your app, or visible in the terminal dashboard:

X-Flense-Tokens-Saved: 18400
X-Flense-Est-Savings: $0.552
X-Flense-Compression-Time: 42ms
X-Flense-Strategy: bulk-reader

Run flense tui for a live terminal dashboard showing real-time savings across your session.


Configuration

# flense.toml — all fields optional

[server]
port = 2912
headless = false

[compression]
threshold = 5000      # compress payloads above this token count
strategy = "auto"     # auto | ast | ctags | passthrough

[providers.anthropic]
upstream = "https://api.anthropic.com"

[providers.openai]
upstream = "https://api.openai.com"

[code_writer]
enabled = false            # opt-in; reads/writes files on the host
model = "claude-haiku-4-5"
fallback = "gpt-4o-mini"
output_dir = "./generated"      # generated output confined here
allowed_ref_dir = "."           # reference files confined here
max_ref_bytes = 1000000         # reject reference files larger than this

Provider Support

Flense supports Anthropic and OpenAI at v1. Route by URL prefix — each provider gets its own adapter with correct token counting and pricing:

localhost:2912/anthropic/v1/messages         →  api.anthropic.com
localhost:2912/openai/v1/chat/completions    →  api.openai.com

Adding a new provider is a single adapter file and route registration — nothing else changes.


License

Copyright (c) 2026 Anand. All rights reserved. See LICENSE.

Metadata

Release files for flense 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for flense 0.2.0
File Size Uploaded
flense-0.2.0.tar.gz 55.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for flense 0.2.0
File Interpreter ABI Platform
flense-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 101.3 kB

Release files / flense-0.2.0.tar.gz

Download URL flense-0.2.0.tar.gz
Size 55.1 kB
Tags Source
SHA-256 checksum
How to use checksums
d6128bb7b345c8351f866f7016ccb99d66feffb5dba519e793de9950e96ab7c3
BLAKE2b-256 checksum
How to use checksums
7e3d09bc95efd227ed8fd679bfd45fb9774de5536504a4f261f379b433e2cc99
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / flense-0.2.0-py3-none-any.whl

Download URL flense-0.2.0-py3-none-any.whl
Size 46.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
50d61934888d48dd119c33f21178781d386126ac618667d4f9d780e6a257aa4a
BLAKE2b-256 checksum
How to use checksums
5eae4e7938184e2f0b30b7b5b8042eee082ba23f3fe37d19d9b8b3a1a9558bbe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page