Skip to main content

bit-jev

Install, open a local decision page, then build your own structured requests. bit-jev scores explicit options over a BitNet backbone and returns answers and probabilities without generating answer text token by token.

bit-jev ternary inputs and decision engine

GitHub documentation · Hugging Face model · ModelScope model · 中文说明

Try it

pip install bit-jev -i https://pypi.org/simple --upgrade
python -m bit_jev.demo

The second command opens a local Gradio page with Choice, Noul, and Score tabs. Choice and Score let you add or remove items within a two-to-four-item range. The first submitted question downloads the 1.19 GB GGUF from Hugging Face or, if that connection fails, ModelScope; later requests reuse the loaded model. Installation and page startup do not download weights. Use python -m bit_jev.demo --source modelscope to select ModelScope directly, or --device gpu for Vulkan. Use --once for the former one-shot JSON command.

The Windows x64 wheel includes precompiled CPU and Vulkan GPU runners for AVX2 processors. Inference on those machines needs no Git, CMake, C++ compiler, or Vulkan SDK. Vulkan still needs a compatible graphics driver and its vulkan-1.dll runtime. Other platforms build the native runner on demand and require Git, CMake 3.28+, and a C++17 compiler (Windows C++ Build Tools). Source builds of Vulkan also need its SDK; CUDA builds need a CUDA Toolkit. Both precompiled programs use pinned BitNet and llama.cpp source with the ReLU² runtime patch and carry their MIT license notices.

Resident inference

from bit_jev.gguf import BitJev

request = {
    "state": "A customer reports a duplicate charge.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {"billing": "Payment and refund issues", "shipping": "Delivery issues"},
        }
    },
}

with BitJev.from_pretrained(device="cpu", threads=8) as model:
    result = model.infer(request)
    print(result["answers"], result["latency_ms"])

Use device="gpu" for Vulkan or device="cuda" for an NVIDIA CUDA build. GPU requests fail clearly if a backend or visible GPU is unavailable. Pass source="modelscope" to select ModelScope directly, or pass a local model directory for offline inference. binary="/path/to/bit-jev-cpu" selects an existing native runner. The model stays loaded for repeated infer() calls; latency_ms reports native compute only, excluding download, build, loading, encoding, and IPC.

The CLI accepts UTF-8 JSONL input:

bit-jev --device cpu --input requests.jsonl --output results.jsonl

Training and distillation dependencies are optional: pip install 'bit-jev[train]'. The native runner evaluates one causal row per question; multiple questions repeat the shared state. The repository's Apache-2.0 license covers source code, while the checkpoint has no standalone open-weights license. The model card documents Yelp training-data provenance and its unresolved permission request.

Metadata

Release files for bit-jev 0.13.12

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bit-jev 0.13.12
File Size Uploaded
bit_jev-0.13.12.tar.gz 17.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for bit-jev 0.13.12
File Interpreter ABI Platform
bit_jev-0.13.12-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details

Total release size: 35.2 MB

Release files / bit_jev-0.13.12.tar.gz

Download URL bit_jev-0.13.12.tar.gz
Size 17.5 MB
Tags Source
SHA-256 checksum
How to use checksums
7e5d549a07728f7c1a065dd0262eab74a96d55bdef0c4c55c16473636d184a9f
BLAKE2b-256 checksum
How to use checksums
376ba466cebc3a0b6b8c8f8d78297970c19ee50317b5ce185c808851b5364a27
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release files / bit_jev-0.13.12-py3-none-win_amd64.whl

Download URL bit_jev-0.13.12-py3-none-win_amd64.whl
Size 17.8 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
50876ee557a1011f830d1f6d2068027f279084b09fdecb37c10fe57588d84cb2
BLAKE2b-256 checksum
How to use checksums
8afe904ce246ff795a4a2bf231b6c5a2ee02d65c532d0d691703d53910fc5baa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release history Release notifications | RSS feed

This release

0.13.12 This release

2 release files

0.9.10

2 release files

0.9.9

2 release files

0.8.9

2 release files

0.8.8

2 release files

0.8.7

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page