Skip to main content

bit-jev

Install, open a local decision page, then build your own structured requests. bit-jev scores explicit options over a BitNet backbone and returns answers and probabilities without generating answer text token by token.

bit-jev ternary inputs and decision engine

GitHub documentation · Hugging Face model · ModelScope model · 中文说明

Try it

pip install bit-jev -i https://pypi.org/simple --upgrade
python -m bit_jev.demo

The second command opens a local Gradio page with Choice, Noul, and Score tabs. Choice and Score let you add or remove items within a two-to-four-item range. The first submitted question downloads the 1.19 GB GGUF from Hugging Face or, if that connection fails, ModelScope; later requests reuse the loaded model. Installation and page startup do not download weights. Use python -m bit_jev.demo --source modelscope to select ModelScope directly, or --device gpu for Vulkan. Use --once for the former one-shot JSON command.

The Windows x64 wheel includes precompiled CPU and Vulkan GPU runners for AVX2 processors. Inference on those machines needs no Git, CMake, C++ compiler, or Vulkan SDK. Vulkan still needs a compatible graphics driver and its vulkan-1.dll runtime. Other platforms build the native runner on demand and require Git, CMake 3.28+, and a C++17 compiler (Windows C++ Build Tools). Source builds of Vulkan also need its SDK; CUDA builds need a CUDA Toolkit. Both precompiled programs use pinned BitNet and llama.cpp source with the ReLU² runtime patch and carry their MIT license notices.

Resident inference

from bit_jev.gguf import BitJev

request = {
    "state": "A customer reports a duplicate charge.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {"billing": "Payment and refund issues", "shipping": "Delivery issues"},
        }
    },
}

with BitJev.from_pretrained(device="cpu", threads=8) as model:
    result = model.infer(request)
    print(result["answers"], result["latency_ms"])

Use device="gpu" for Vulkan or device="cuda" for an NVIDIA CUDA build. GPU requests fail clearly if a backend or visible GPU is unavailable. Pass source="modelscope" to select ModelScope directly, or pass a local model directory for offline inference. binary="/path/to/bit-jev-cpu" selects an existing native runner. The model stays loaded for repeated infer() calls; latency_ms reports native compute only, excluding download, build, loading, encoding, and IPC.

The CLI accepts UTF-8 JSONL input:

bit-jev --device cpu --input requests.jsonl --output results.jsonl

Training and distillation dependencies are optional: pip install 'bit-jev[train]'. The native runner evaluates one causal row per question; multiple questions repeat the shared state. The repository's Apache-2.0 license covers source code, while the checkpoint has no standalone open-weights license. The model card documents Yelp training-data provenance and its unresolved permission request.

Metadata

Release files for bit-jev 0.13.10

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bit-jev 0.13.10
File Size Uploaded
bit_jev-0.13.10.tar.gz 17.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for bit-jev 0.13.10
File Interpreter ABI Platform
bit_jev-0.13.10-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details

Total release size: 35.2 MB

Release files / bit_jev-0.13.10.tar.gz

Download URL bit_jev-0.13.10.tar.gz
Size 17.5 MB
Tags Source
SHA-256 checksum
How to use checksums
df530a2885f687a033ed29d88924fcdc62aac2bdb0472e52b8cac03fabde1af6
BLAKE2b-256 checksum
How to use checksums
5078e7df99ebba9008f87725c5176b65dcdb40d6f1f18550d1ed53e6f2dff048
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release files / bit_jev-0.13.10-py3-none-win_amd64.whl

Download URL bit_jev-0.13.10-py3-none-win_amd64.whl
Size 17.8 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
e560194794106eba1e0cc005ac3916a2946f28c76cba87ef417f8d19f327c3ae
BLAKE2b-256 checksum
How to use checksums
b5f8ecab9af96214914b4dda3c768f0532944db6e71ed4904342714983efa1a6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release history Release notifications | RSS feed

This release

0.13.10 This release

2 release files

0.9.10

2 release files

0.9.9

2 release files

0.8.9

2 release files

0.8.8

2 release files

0.8.7

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page