bit-jev
bit-jev scores explicit options over a 1.58-bit BitNet backbone. Its I2_S GGUF inference path loads a resident model and returns structured answers, logits, and probabilities without generating answer tokens.
GitHub documentation · GGUF model and model card · 中文说明
Install
pip install bit-jev
The wheel contains Python code and native build sources. The 1.19 GB GGUF is downloaded from Hugging Face on first use. Native compilation requires Git, CMake 3.28+, and a C++17 compiler (Windows C++ Build Tools). Git retrieves and verifies pinned BitNet and llama.cpp source and applies the ReLU² patch; normal inference does not need Git. Version 0.8.8 checks build tools before downloading the model. Vulkan GPU mode additionally requires a Vulkan SDK; CUDA mode requires a CUDA Toolkit. Neither model download nor compilation runs during pip install.
Resident inference
from bit_jev.gguf import BitJev
request = {
"state": "A customer reports a duplicate charge.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "Payment and refund issues", "shipping": "Delivery issues"},
}
},
}
with BitJev.from_pretrained(device="cpu", threads=8) as model:
result = model.infer(request)
print(result["answers"], result["latency_ms"])
Use device="gpu" for Vulkan or device="cuda" for an NVIDIA CUDA build. GPU requests fail clearly if a backend or visible GPU is unavailable. A local model directory can replace the default Hugging Face repo, and binary="/path/to/bit-jev-cpu" can select an existing native runner. The model stays loaded for repeated infer() calls; latency_ms reports native compute only, excluding download, build, loading, encoding, and IPC.
The CLI accepts UTF-8 JSONL input:
bit-jev --device cpu --input requests.jsonl --output results.jsonl
Training and distillation dependencies are optional: pip install 'bit-jev[train]'. The native runner evaluates one causal row per question; multiple questions repeat the shared state. The repository's Apache-2.0 license covers source code, while the checkpoint has no standalone open-weights license. The model card documents Yelp training-data provenance and its unresolved permission request.
Metadata
Release files for bit-jev 0.8.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| bit_jev-0.8.9.tar.gz | 54.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| bit_jev-0.8.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 114.7 kB
Release files / bit_jev-0.8.9.tar.gz
| Download URL | bit_jev-0.8.9.tar.gz |
|---|---|
| Size | 54.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
917670ca1986bda9fccae2bfa1f322c2e93a38f86d6d48cbdc35397eddf95c10
|
|
BLAKE2b-256 checksum How to use checksums |
d263fbf72912e3c564180fc667c5b2801ecae019fd8c98cd5c3e0e3565458b8c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|
Release files / bit_jev-0.8.9-py3-none-any.whl
| Download URL | bit_jev-0.8.9-py3-none-any.whl |
|---|---|
| Size | 60.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
06f19581bd0e57a28d1551f5028900f70f8e05a8ce649bb010e55109d0f733d0
|
|
BLAKE2b-256 checksum How to use checksums |
3bf4b7fc5640312db45a1a609e888fa13a7190aeadf9def5da21f67dd17823ae
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|