TramaMind
Local proxy and chat that shows how many tokens it left out. Recent turns stay word for word, however long they are. From older turns it keeps up to four verbatim excerpts — a code block, a unified diff, a traceback, or a tool result — each at most 1500 characters, and only while they fit. One cloud call, and only when the local answer is empty, a short refusal, or you ask to redo it.
pipx install "git+https://github.com/bitfarmy/tramamind"
tramamind demo
tramamind pack session.json
demo needs no model. On the sample chat it prints:
full history ~10273 · packet ~1211 · −88%
code kept verbatim: def add(a, b)
last user message kept: What does add return?
The old turns in that sample are the same sentence repeated, which is why the gap is wide. def add comes from those old turns and is still in the packet. The last question is the original message. A long source file in the latest turns is sent whole, and the percentage gets smaller. The figure is this estimate (about 3.5 characters per token), on this chat.
pack runs the same packet on your own transcript. The file is a JSON array of messages, or an object with a messages key. The line is the one the proxy prints (storia intera, inviati, the percentage), then the blocks that stayed. --out packet.json writes the messages that would be forwarded. Nothing is sent to a model.
In front of Cursor, Continue, or any OpenAI client
tramamind proxy
OpenAI base URL: http://127.0.0.1:8788/v1
Default upstream: Ollama at http://127.0.0.1:11434/v1. OmniRoute, if you use it, is --upstream http://127.0.0.1:20128/v1. The chat page stays on port 8787.
# Continue
name: TramaMind
provider: openai
model: qwen2.5-coder:7b
apiBase: http://127.0.0.1:8788/v1
apiKey: ollama
Each request prints a line on stderr, and the response carries X-Tramamind-Raw-Tokens, X-Tramamind-Sent-Tokens, and X-Tramamind-Saved-Pct. When the upstream sends prompt_tokens, the same line gains · api N and the response gains X-Tramamind-Api-Prompt-Tokens. The summary is an excerpt, so the model already loaded for the client stays loaded. Your system message stays as it arrived. The log stores the counts.
--budget defaults to 2500 estimated tokens and --verbatim to the last 6 messages. With a profile, those two values come from ~/.config/tramamind/profile.yaml.
Chat on this machine
./scripts/install.sh
.venv/bin/tramamind setup
.venv/bin/tramamind pull
.venv/bin/tramamind chat
tramamind ui serves http://127.0.0.1:8787. That page sets the token budget, chooses Ollama or OmniRoute, starts the proxy, and shows the last savings line.
In italiano
Assistente personale sul tuo PC. Il modello di default è locale, via Ollama. La chat ricorda le sessioni e, quando la storia si allunga, ne manda un riassunto più gli ultimi turni, non la trascrizione intera. Se la risposta locale è vuota, è una scusa, o chiedi di rifarla, parte una chiamata cloud.
OmniRoute, se lo installi, resta il tubo verso più provider. TramaMind non è un secondo router: decide cosa entra nel prompt e quale modello locale usare.
tramamind chat
│
▼
pacchetto con tetto ──► Ollama
│
└─ solo se serve ──► Groq diretto, oppure OmniRoute
Ollama da solo non tiene un tetto sulla storia e non cambia modello in base alla domanda. OmniRoute da solo non sceglie i modelli per la tua RAM e non distingue una risposta scarsa da un errore HTTP. TramaMind fa quelle due cose e lascia il resto a loro.
tramamind proxy mette lo stesso pacchetto davanti a Cursor, Continue o a qualsiasi client compatibile con OpenAI. tramamind demo mostra il conto senza scaricare un modello. tramamind pack sessione.json fa lo stesso conto sulla tua trascrizione.
Claude Pro non passa di qui. Per il codice serio resta Claude Code, diretto.
./scripts/install.sh
.venv/bin/tramamind setup # preset in base alla RAM, chiavi opzionali
.venv/bin/tramamind pull # scarica i tag del preset
.venv/bin/tramamind chat
Pagina locale, solo su questa macchina: tramamind ui poi http://127.0.0.1:8787. Da lì partono anche tetto, upstream e il proxy su http://127.0.0.1:8788/v1.
Corsie: tramamind chat --code "…", --think, --fast. In chat: /new, /sessions, /resume, /use code, /clear.
Sotto ogni risposta c'è il conto: token inviati, stima della storia intera, e la percentuale tolta dal pacchetto quando c'è stato un riassunto. tramamind stats aggrega quei conti, senza il testo.
tramamind doctor dice cosa manca. Se Ollama non risponde, il comando fallisce: non stampa "attivo" lo stesso.
Preset
| Preset | Quando | Generale | Codice | Ragionamento |
|---|---|---|---|---|
cpu16 |
portatile, 16 GB | gemma3:4b |
qwen2.5-coder:3b |
deepseek-r1:1.5b |
gpu12 |
GPU ~12 GB | qwen3:8b |
qwen2.5-coder:7b |
deepseek-r1:8b |
gpu24 |
GPU 24 GB | qwen3:32b |
qwen2.5-coder:14b |
deepseek-r1:14b |
Il tag summary del profilo è gemma3:4b. Viene chiamato solo se è già il modello della risposta (è il caso di cpu16). Sugli altri preset il riassunto è estrattivo, per non ricaricare un modello a ogni messaggio. Il proxy è sempre estrattivo. I tag sono quelli di docs/local-models.md.
Chiavi
tramamind keys add groq le mette nel keyring (oppure in un file age, oppure in ~/.config/tramamind/.env con permessi 0600). Con una chiave Groq l'escalation funziona anche senza OmniRoute. tramamind sync spinge le chiavi nel CLI di OmniRoute, se c'è.
Il modello cloud di default nel profilo è openai/gpt-oss-120b. Se quel catalogo cambia, lo sostituisci in ~/.config/tramamind/profile.yaml.
Cosa non fa
Il risparmio è il pacchetto: ultimi turni interi, e dai turni vecchi un estratto più fino a quattro pezzi parola per parola (codice, diff, traceback, risultato di un tool) se ci stanno. Ogni pezzo vecchio è al massimo 1500 caratteri. La compressione di OmniRoute, se la accendi dalla sua dashboard, è un'altra cosa e può tagliare proprio il pezzo che serviva. Il numero di cui mi fido è quello stampato da tramamind chat, da tramamind demo, da tramamind pack e dal proxy. Quando il provider manda prompt_tokens, la riga del proxy lo affianca alla stima.
OpenHands e OpenClaw sono extra, spenti. L'app desktop e il nodo Oracle non fanno parte del flusso: gli appunti sono in docs/future/.
Licenza MIT. Uso personale delle API gratuite: docs/legal.md.
Metadata
Release files for tramamind 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tramamind-0.5.0.tar.gz | 50.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tramamind-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 95.5 kB
Release files / tramamind-0.5.0.tar.gz
| Download URL | tramamind-0.5.0.tar.gz |
|---|---|
| Size | 50.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
daacafd9f27beb8d809d11a2cf4b7bf4e025f878faf66128a56ccb728120b62c
|
|
BLAKE2b-256 checksum How to use checksums |
a2a5ed459edd41e709246f06ef834a162be40d86a39828e30026e9e8b7c2d238
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / tramamind-0.5.0-py3-none-any.whl
| Download URL | tramamind-0.5.0-py3-none-any.whl |
|---|---|
| Size | 45.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
90b22ce93cf90e0d129e6c95f6225077dd5a7f63173b8571e8a48ec4fd34128c
|
|
BLAKE2b-256 checksum How to use checksums |
84a1202ff6d139e20a65751e1d81a7724656cd90d28a3a3df22cc38e5bca4dca
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|