hugpy-wrapper (fitevict)
pip install hugpy-wrapper # import fitevict; fitevict-serve
pip install "hugpy-wrapper[transformers]" # the transformers engine (own venv)
hugpy's per-box wrapper: an evict-to-fit front door over llama.cpp (and transformers). Point any OpenAI client's base URL at it; it turns one address into a self-managing model server.
A /v1 call carries only a model name — never a location. llama has no idea
where a gguf lives. fitevict is the layer that does what the call can't: for
each request it resolves the name → gguf path, runs an evict-to-fit plan
against what's resident and what's being called, places the model into llama
(launch/evict), forwards the call, splits the model's <think> out of the
answer, and records the call — raw stamps + engine geometry — to a call log that
every metric derives from.
The package depends on nothing else in hugpy. The decision core is pure stdlib; hardware
facts come from nvidia-smi; model/gguf facts from the files on disk.
Layout
fitevict/
types.py frozen data contract (DeviceBudget, Resident, LoadRequest, FitPlan, ...)
evict.py plan_eviction / sort_key — the victim selector (pure)
plan.py plan_fit — staged decision + quant ladder (pure)
flex.py ctx-band compress + layers-that-fit offload (pure)
host.py EngineHost protocol + drive() loop (measure→plan→evict→load)
front_door.py the WRAPPER: persistent OpenAI /v1 server (owns the address)
adapters/
llama_cpp.py concrete EngineHost: measure GPU/RAM, price GGUFs, launch/evict llama-server
db.py call_log writer (Postgres inference_engine.call_log)
lifted/ measure/act helpers lifted clean from hugpy (imports stripped):
gguf_inspect, gguf_need, hardware, spill_reserve, model_resolve,
native_resolve, supervisor/procutil, timings, no_think, app_dirs, ...
run_engine.py a juggling harness: discover models, fire a sequence that exceeds the card
Install
pip install . # core + wrapper (stdlib only)
pip install .[db] # + psycopg for direct call_log DSN writes (else shells to psql)
pip install .[test] # + pytest
Run
fitevict-serve --host 127.0.0.1 --port 8080 # the front door (owns /v1)
fitevict-run # the juggle harness on this box
Then point any OpenAI client at http://127.0.0.1:8080/v1:
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"<name>","messages":[{"role":"user","content":"hi"}],"max_tokens":128}'
The model is named, not located; the wrapper finds it, fits it, serves it, logs it.
Test
pytest # pure unit tests + a call_log replay + the VL-flood fixture
Metadata
Release files for hugpy-wrapper 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hugpy_wrapper-0.1.0.tar.gz | 132.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hugpy_wrapper-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 270.3 kB
Release files / hugpy_wrapper-0.1.0.tar.gz
| Download URL | hugpy_wrapper-0.1.0.tar.gz |
|---|---|
| Size | 132.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e06e8378367a3260abac53d9c38158229a3269ae9837687119553783765bdb56
|
|
BLAKE2b-256 checksum How to use checksums |
9d50bbeae850516d841909acec4216fc8aee1291a11291dc63ea6ac730763e81
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.12
|
Release files / hugpy_wrapper-0.1.0-py3-none-any.whl
| Download URL | hugpy_wrapper-0.1.0-py3-none-any.whl |
|---|---|
| Size | 138.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
28314ad8a39625d54a9808708c3b8a73808ec40c944cee6aff1fb2d973101d5a
|
|
BLAKE2b-256 checksum How to use checksums |
c8441a6da1872fede7777f9909653bd4ca10c1b73ec9e8acd58cc074a21eb94e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.12
|