bit-jev
Install, open a local decision page, then build your own structured requests. bit-jev scores explicit options over a BitNet backbone and returns answers and probabilities without generating answer text token by token.
GitHub documentation · Hugging Face model · ModelScope model · 中文说明
Try it
pip install bit-jev -i https://pypi.org/simple --upgrade
python -m bit_jev.demo
The second command opens a local Gradio page with Choice, Noul, and Score tabs. Choice and Score let you add or remove items within a two-to-four-item range. The first submitted question downloads the 1.19 GB GGUF from Hugging Face or, if that connection fails, ModelScope; later requests reuse the loaded model. Installation and page startup do not download weights. Use python -m bit_jev.demo --source modelscope to select ModelScope directly, or --device gpu for Vulkan. Use --once for the former one-shot JSON command.
The Windows x64 wheel includes precompiled CPU and Vulkan GPU runners for AVX2 processors. Inference on those machines needs no Git, CMake, C++ compiler, or Vulkan SDK. Vulkan still needs a compatible graphics driver and its vulkan-1.dll runtime. Other platforms build the native runner on demand and require Git, CMake 3.28+, and a C++17 compiler (Windows C++ Build Tools). Source builds of Vulkan also need its SDK; CUDA builds need a CUDA Toolkit. Both precompiled programs use pinned BitNet and llama.cpp source with the ReLU² runtime patch and carry their MIT license notices.
Resident inference
from bit_jev.gguf import BitJev
request = {
"state": "A customer reports a duplicate charge.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "Payment and refund issues", "shipping": "Delivery issues"},
}
},
}
with BitJev.from_pretrained(device="cpu", threads=8) as model:
result = model.infer(request)
print(result["answers"], result["latency_ms"])
Use device="gpu" for Vulkan or device="cuda" for an NVIDIA CUDA build. GPU requests fail clearly if a backend or visible GPU is unavailable. Pass source="modelscope" to select ModelScope directly, or pass a local model directory for offline inference. binary="/path/to/bit-jev-cpu" selects an existing native runner. The model stays loaded for repeated infer() calls; latency_ms reports native compute only, excluding download, build, loading, encoding, and IPC.
The CLI accepts UTF-8 JSONL input:
bit-jev --device cpu --input requests.jsonl --output results.jsonl
Training and distillation dependencies are optional: pip install 'bit-jev[train]'. The native runner evaluates one causal row per question; multiple questions repeat the shared state. The repository's Apache-2.0 license covers source code, while the checkpoint has no standalone open-weights license. The model card documents Yelp training-data provenance and its unresolved permission request.
Metadata
Release files for bit-jev 0.13.12
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| bit_jev-0.13.12.tar.gz | 17.5 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| bit_jev-0.13.12-py3-none-win_amd64.whl | Python 3 | none | Windows x86-64 | Details |
Total release size: 35.2 MB
Release files / bit_jev-0.13.12.tar.gz
| Download URL | bit_jev-0.13.12.tar.gz |
|---|---|
| Size | 17.5 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7e5d549a07728f7c1a065dd0262eab74a96d55bdef0c4c55c16473636d184a9f
|
|
BLAKE2b-256 checksum How to use checksums |
376ba466cebc3a0b6b8c8f8d78297970c19ee50317b5ce185c808851b5364a27
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|
Release files / bit_jev-0.13.12-py3-none-win_amd64.whl
| Download URL | bit_jev-0.13.12-py3-none-win_amd64.whl |
|---|---|
| Size | 17.8 MB |
| Tags | Python 3 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
50876ee557a1011f830d1f6d2068027f279084b09fdecb37c10fe57588d84cb2
|
|
BLAKE2b-256 checksum How to use checksums |
8afe904ce246ff795a4a2bf231b6c5a2ee02d65c532d0d691703d53910fc5baa
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|