9 projects
Defuser
Model defuser helper for HF Transformers.
GPTQModel
Production ready LLM model compression/quantization toolkit with hw accelerated inference support for both cpu/gpu via HF, vLLM, and SGLang.
Evalution
Modern LLM model evaluation for Transformers, SGLang, vLLM, TensorRT-LLM, llama.cpp, GPTQModel, OpenAI-compatible HTTP backends, and OpenVINO.
LogBar
A unified Logger and ProgressBar util with zero dependencies.
zpu
Python bindings for the ZPU emulated Linux input drivers (zmouse/zkeyboard).
omy
OMY (Open Model Yard): the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
PyPcre
Modern, GIL-friendly, Fast Python bindings for PCRE2 with auto caching and JIT of compiled patterns.
Device-SMI
Retrieve gpu, cpu, and npu device info and properties from Linux/MacOS with zero package dependency.
TokeNicer
A (nicer) tokenizer you want to use for model `inference` and `training`: with all known peventable `gotchas` normalized or auto-fixed.