jalad-local
kaun banega tokenpati — a game show where your machine sits in the hot seat and the models are the contestants.
Run one command. It reads your hardware, works out which local LLMs will actually run on it, at what quant, on which backend, and how many tokens per second you'll get. Then it hands you the exact command to run.
Verdicts
| Verdict | Meaning | Rule |
|---|---|---|
daudega |
it'll sprint | top 3 by predicted tok/s, and only if above the fast threshold |
chalega |
it'll do | fits in memory with headroom, usable speed |
ghare jake sutti babu |
go home and sleep | doesn't fit, or fits but unusably slow |
Install
pip install tokenpati
tokenpati
Or without installing anything permanently:
uvx tokenpati
Straight from the repo:
pip install git+https://github.com/utsav1033/jalad-local
Develop
git clone https://github.com/utsav1033/jalad-local && cd jalad-local
uv sync
uv run tokenpati
uv run pytest
Release
Bump version in pyproject.toml, then tag it. GitHub Actions builds and publishes to PyPI.
git tag v0.1.0 && git push --tags
Flags: --context 32768 budgets memory for a longer context, --fast skips the game-show pause, --json for scripts.
How it decides
Every number is derived from four things about your machine: memory, memory bandwidth, backend, and how much of that memory the GPU may actually use (Apple caps it at about 75% of unified memory).
- Weights = params × bytes per param at that quant × 1.1 overhead.
- KV cache = 2 × layers × kv heads × head dim × context × 2 bytes. Budgeted at the context you ask for.
- tok/s = bandwidth ÷ (active weights + KV at 2k) × 0.8. Decode is memory-bound, so this one line predicts speed for any model on any machine. MoE models use their active parameters, which is why a 30B-A3B can outrun an 8B.
- Quant = the highest precision that still clears the fast bar, else the fastest one that clears the slow bar.
- Winner = the biggest model that gets
daudega.
Every tok/s is a prediction. Run the model, measure, and adjust EFFICIENCY in budget.py if it's off.
Phase 1 (now): analyzer only
Phase 2 adds tokenpati serve, which picks the backend and runs the winner.
Layout
src/tokenpati/
hardware.py what machine is this: chip, RAM, bandwidth, backend
catalog.py the contestants: model families and their architecture numbers
budget.py the maths: weight bytes, KV cache bytes, predicted tok/s
verdict.py daudega / chalega / ghare jake sutti babu
ui.py the game show: banner, gauges, table, final command
cli.py entry point
docs/wireframe.txt the screen, sketched before any UI code
tests/ predict, measure, correct
The rule
Predict, measure, correct. Every tok/s number this tool prints is a prediction. Run the model, measure the real number, fix the formula.
Metadata
Release files for tokenpati 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tokenpati-0.1.0.tar.gz | 18.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tokenpati-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 31.9 kB
Release files / tokenpati-0.1.0.tar.gz
| Download URL | tokenpati-0.1.0.tar.gz |
|---|---|
| Size | 18.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
856ad797545398a59dffe8e4359b263c09462178e1c262ae5a6b2259fefc0e25
|
|
BLAKE2b-256 checksum How to use checksums |
4a83b31cc557f44dec22c88416c530f731b63489b75a168c8ea78341848fcbc5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.17 {"installer":{"name":"uv","version":"0.12.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / tokenpati-0.1.0-py3-none-any.whl
| Download URL | tokenpati-0.1.0-py3-none-any.whl |
|---|---|
| Size | 13.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a5da1b1b5102fedf914dc5db289e3324d08e8bf7660d4d115f0dd202c87b0cc5
|
|
BLAKE2b-256 checksum How to use checksums |
5bd5514430e367c029385a559e0b7793af7e60f003c6c5ba03d419f59c357cc1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.17 {"installer":{"name":"uv","version":"0.12.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|