hugpy-wrapper (fitevict)
pip install hugpy-wrapper # import fitevict; fitevict-serve
pip install "hugpy-wrapper[transformers]" # the transformers engine (own venv)
hugpy's per-box wrapper: an evict-to-fit front door over llama.cpp (and transformers). Point any OpenAI client's base URL at it; it turns one address into a self-managing model server.
A /v1 call carries only a model name — never a location. llama has no idea
where a gguf lives. fitevict is the layer that does what the call can't: for
each request it resolves the name → gguf path, runs an evict-to-fit plan
against what's resident and what's being called, places the model into llama
(launch/evict), forwards the call, splits the model's <think> out of the
answer, and records the call — raw stamps + engine geometry — to a call log that
every metric derives from.
The package depends on nothing else in hugpy. The decision core is pure stdlib; hardware
facts come from nvidia-smi; model/gguf facts from the files on disk.
Layout
fitevict/
types.py frozen data contract (DeviceBudget, Resident, LoadRequest, FitPlan, ...)
evict.py plan_eviction / sort_key — the victim selector (pure)
plan.py plan_fit — staged decision + quant ladder (pure)
flex.py ctx-band compress + layers-that-fit offload (pure)
host.py EngineHost protocol + drive() loop (measure→plan→evict→load)
front_door.py the WRAPPER: persistent OpenAI /v1 server (owns the address)
adapters/
llama_cpp.py concrete EngineHost: measure GPU/RAM, price GGUFs, launch/evict llama-server
db.py call_log writer (Postgres inference_engine.call_log)
lifted/ measure/act helpers lifted clean from hugpy (imports stripped):
gguf_inspect, gguf_need, hardware, spill_reserve, model_resolve,
native_resolve, supervisor/procutil, timings, no_think, app_dirs, ...
run_engine.py a juggling harness: discover models, fire a sequence that exceeds the card
Install
pip install . # core + wrapper (stdlib only)
pip install .[db] # + psycopg for direct call_log DSN writes (else shells to psql)
pip install .[test] # + pytest
Run
fitevict-serve --host 127.0.0.1 --port 8080 # the front door (owns /v1)
fitevict-run # the juggle harness on this box
Then point any OpenAI client at http://127.0.0.1:8080/v1:
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"<name>","messages":[{"role":"user","content":"hi"}],"max_tokens":128}'
The model is named, not located; the wrapper finds it, fits it, serves it, logs it.
Test
pytest # pure unit tests + a call_log replay + the VL-flood fixture
Metadata
Release files for hugpy-wrapper 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hugpy_wrapper-0.1.5.tar.gz | 152.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hugpy_wrapper-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 305.5 kB
Release files / hugpy_wrapper-0.1.5.tar.gz
| Download URL | hugpy_wrapper-0.1.5.tar.gz |
|---|---|
| Size | 152.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e58a71b2ca0c24ac41348707191d4d596ed5868040242b7a4a24ecbc562425c6
|
|
BLAKE2b-256 checksum How to use checksums |
e006b84fdeaa4c8dcdb9a0fcbf869ae30ee2ce1a75d4debcbb9a0b45148f8ff0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|
Release files / hugpy_wrapper-0.1.5-py3-none-any.whl
| Download URL | hugpy_wrapper-0.1.5-py3-none-any.whl |
|---|---|
| Size | 152.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
12d1414b3bfa2232dd9d6b163902cd4da6bb88c4f276e4037c0cda566aa54598
|
|
BLAKE2b-256 checksum How to use checksums |
25d28ace5d0063f0a33e030b1cf38cfe9285cea9b892c4771cc0f8bf8dafcfb9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.3
|