hugpy-wrapper (fitevict)
pip install hugpy-wrapper # import fitevict; fitevict-serve
pip install "hugpy-wrapper[transformers]" # the transformers engine (own venv)
hugpy's per-box wrapper: an evict-to-fit front door over llama.cpp (and transformers). Point any OpenAI client's base URL at it; it turns one address into a self-managing model server.
A /v1 call carries only a model name — never a location. llama has no idea
where a gguf lives. fitevict is the layer that does what the call can't: for
each request it resolves the name → gguf path, runs an evict-to-fit plan
against what's resident and what's being called, places the model into llama
(launch/evict), forwards the call, splits the model's <think> out of the
answer, and records the call — raw stamps + engine geometry — to a call log that
every metric derives from.
The package depends on nothing else in hugpy. The decision core is pure stdlib; hardware
facts come from nvidia-smi; model/gguf facts from the files on disk.
Layout
fitevict/
types.py frozen data contract (DeviceBudget, Resident, LoadRequest, FitPlan, ...)
evict.py plan_eviction / sort_key — the victim selector (pure)
plan.py plan_fit — staged decision + quant ladder (pure)
flex.py ctx-band compress + layers-that-fit offload (pure)
host.py EngineHost protocol + drive() loop (measure→plan→evict→load)
front_door.py the WRAPPER: persistent OpenAI /v1 server (owns the address)
adapters/
llama_cpp.py concrete EngineHost: measure GPU/RAM, price GGUFs, launch/evict llama-server
db.py call_log writer (Postgres inference_engine.call_log)
lifted/ measure/act helpers lifted clean from hugpy (imports stripped):
gguf_inspect, gguf_need, hardware, spill_reserve, model_resolve,
native_resolve, supervisor/procutil, timings, no_think, app_dirs, ...
run_engine.py a juggling harness: discover models, fire a sequence that exceeds the card
Install
pip install . # core + wrapper (stdlib only)
pip install .[db] # + psycopg for direct call_log DSN writes (else shells to psql)
pip install .[test] # + pytest
Run
fitevict-serve --host 127.0.0.1 --port 8080 # the front door (owns /v1)
fitevict-run # the juggle harness on this box
Then point any OpenAI client at http://127.0.0.1:8080/v1:
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"<name>","messages":[{"role":"user","content":"hi"}],"max_tokens":128}'
The model is named, not located; the wrapper finds it, fits it, serves it, logs it.
Test
pytest # pure unit tests + a call_log replay + the VL-flood fixture
Metadata
Release files for hugpy-wrapper 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hugpy_wrapper-0.1.3.tar.gz | 145.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hugpy_wrapper-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 292.8 kB
Release files / hugpy_wrapper-0.1.3.tar.gz
| Download URL | hugpy_wrapper-0.1.3.tar.gz |
|---|---|
| Size | 145.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
366555fa6b5d9b168480d1fd9f1f61a66bf1884cc2d35b421da92cc50881c30a
|
|
BLAKE2b-256 checksum How to use checksums |
9731af25b92a2ddd1c48ff99c0c00d2942029178385b80f250971331eef351ad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|
Release files / hugpy_wrapper-0.1.3-py3-none-any.whl
| Download URL | hugpy_wrapper-0.1.3-py3-none-any.whl |
|---|---|
| Size | 147.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
686e36a290994f1932c2c56f9ea5814199e65163fcc7a3055f66e858574480b9
|
|
BLAKE2b-256 checksum How to use checksums |
9649cbcc2b8b63fbe1b56a1885706a025b6c6219d4688f81a82687a31fbf08d7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|