██████╗ ███████╗██╗ ██╗██╗ ██╗ █████╗
██╔════╝ ██╔════╝██║ ██║██║ ██║██╔══██╗
██║ ███╗█████╗ ██║ ██║██║ ██║███████║
██║ ██║██╔══╝ ╚██╗ ██╔╝╚██╗ ██╔╝██╔══██║
╚██████╔╝███████╗ ╚████╔╝ ╚████╔╝ ██║ ██║
╚═════╝ ╚══════╝ ╚═══╝ ╚═══╝ ╚═╝ ╚═╝
Gevva: Multimodal System 1 Decision Engine & NLI Cross-Encoder
Fast fact-checking, hallucination detection, tool routing & document verification based on Google Gemma 4.
Single-pass sequence classification • Open weights (Apache 2.0) • Evaluated on ANLI, MNLI, HANS & JevBench Public
Live Web Demo • Quickstart • Superpowers • Hardware & Latency Profile • Benchmarks & Empirical Accuracy • Known Limitations • Fine-Tuning
🤔 What is Gevva?
When you ask an autoregressive LLM (like ChatGPT or Claude) to verify a fact, it generates text token-by-token, typically taking 1 to 3 seconds for an answer.
For discrete decision tasks, autoregressive generation is unnecessarily slow and expensive:
- "Does this retrieved context support the generated claim, or is it a hallucination?"
- "Which of these 4 tools should be called for this user request?"
- "Did the student's answer match the reference solution?"
The System 1 Approach
Psychologist Daniel Kahneman described human thought in two modes: System 1 (fast, automatic reflex) and System 2 (slow, deliberate reasoning).
Gevva is designed as a System 1 decision engine: a sequence classification cross-encoder built on Google's lightweight Gemma 4 foundation. Rather than generating text, it evaluates premise-hypothesis pairs in a single forward pass and outputs calibrated 3-class probability distributions:
$$\text{Class} \in {\text{Contradiction (0)}, \text{Entailment (1)}, \text{Neutral (2)}}$$
🚀 Quickstart in 30 Seconds
1. Install via pip
pip install gevva
2. Verify evidence in Python
from gevva import load
# Load model (automatically uses CUDA, MPS, or CPU)
engine = load("davidburhans/gevva-e2b")
# Check if evidence entails or contradicts a statement
premise = "The Eiffel Tower is a wrought-iron lattice tower located on the Champ de Mars in Paris, France."
claim = "The Eiffel Tower is located in Paris."
probs = engine.predict([(premise, claim)])[0]
print(f"Entailment: {probs[1]*100:.1f}%, Contradiction: {probs[0]*100:.1f}%, Neutral: {probs[2]*100:.1f}%")
# -> Entailment: 96.1%, Contradiction: 0.9%, Neutral: 2.9%
⚡ What Can Gevva Do for You?
1. 🛡️ Catch AI Hallucinations in RAG Pipelines
In Retrieval-Augmented Generation (RAG), LLMs frequently introduce unsupported or contradictory claims. Use Gevva as a fast guardrail:
retrieved_document = "Patients taking Medication X showed improved sleep duration with no reported nausea."
ai_answer = "Medication X causes severe nausea in elderly patients."
probs = engine.predict([(retrieved_document, ai_answer)])[0]
# Returns [Contradiction: 75.2%, Entailment: 3.3%, Neutral: 21.5%]
if probs[0] > 0.70:
print("🚨 Alert: AI Hallucination detected! Claim contradicts source document.")
2. 🎯 Tool & Intent Routing
Evaluate multiple candidate tools simultaneously without prompt-tuning:
available_tools = [
"process_refund: Refund payment to customer bank account",
"track_package: Query live shipping milestones and courier GPS",
"reset_password: Send authentication link to user email",
"search_help_docs: Search FAQs and documentation"
]
user_message = "I ordered this two weeks ago and it still hasn't arrived at my house!"
best_idx, scores = engine.rerank(user_message, available_tools)
print("Chosen Action:", available_tools[best_idx])
# -> "track_package: Query live shipping milestones and courier GPS"
3. 📝 Instant Answer & Rubric Grading
Grade student or agent responses against reference keys:
grade = engine.grade(
question="What is the capital of Australia?",
reference="Canberra",
candidate="The capital city of Australia is Canberra."
)
print(f"Passed: {grade.is_correct} (Confidence: {grade.score*100:.1f}%)")
# -> Passed: True (Confidence: 81.8%)
💻 Hardware, Latency & Memory Profile
Gevva models are built on Google's multimodal gemma-4-E2B-it and gemma-4-E4B-it checkpoints. Here are real measured performance numbers across hardware:
Model Size & Download Footprint
- Gevva e2b: 5.10 Billion total parameters (~10.2 GB download in bfloat16). (Includes 2.3B active text backbone + SigLIP vision encoder + 262K vocabulary embedding table).
- Gevva e4b: 4.5B backbone / 5.8B total parameters (~15.9 GB download in bfloat16).
Measured Latency & Memory
| Input Type & Size | Hardware | Latency | RAM / VRAM | Notes |
|---|---|---|---|---|
| Short sentence pair (<128 tokens) | NVIDIA GeForce RTX 5090 | 14–19 ms | ~4.8 GB VRAM | High-throughput GPU serving |
| Short sentence pair (<128 tokens) | Apple M4 Pro GPU (MPS) | ~74 ms | ~4.6 GB RAM | Local developer machines |
| Short sentence pair (<128 tokens) | Apple M4 Pro CPU (FP32) | ~205 ms | ~11.9 GB RAM | Full-precision CPU execution |
| Short sentence pair (<128 tokens) | Apple M4 Pro CPU (BF16 default) | ~720 ms | ~4.6 GB RAM | CPU emulation of bfloat16 |
| RAG Document Context (~900 tokens) | Apple M4 Pro GPU (MPS) | ~0.84 s | ~5.2 GB RAM | Realistic RAG passage check |
| RAG Document Context (~900 tokens) | Apple M4 Pro CPU | ~20 s | ~5.0 GB RAM | Heavy for CPU; GPU recommended for long docs |
Context Length Note: Gemma 4 natively supports up to 128K context via Rotary Position Embeddings (RoPE). However, the fine-tuning curriculum for Gevva was trained on sequences up to 2,048 tokens. For optimal accuracy and latency in RAG pipelines, chunking retrieved evidence to ~1,000–2,000 tokens is strongly recommended.
📊 Benchmarks & Empirical Accuracy
1. Independent Evaluation on Unseen Public Test Sets
Evaluated on 250 items per test set against a standard fact-checking baseline (nli-deberta-v3-base, ~184M parameters):
| Test Set | Gevva e2b (5.1B) | DeBERTa-v3-base (184M) | Key Observations |
|---|---|---|---|
| ANLI Round 1 (Adversarial NLI) | 60.0% | 43.0% | Stronger on tricky multi-hop negations |
| ANLI Round 2 (Adversarial NLI) | 47.0% | 38.0% | Handles hard semantic shifts better |
| ANLI Round 3 (Adversarial NLI) | 50.0% | 38.0% | Outperforms baseline on complex claims |
| MNLI Mismatched (Standard fact-checking) | 88.4% | 89.6% | Comparable on standard in-domain text |
| HANS (Word-overlap shortcut test) | 83.0% | 80.0% | Overall accuracy |
| ↳ HANS non-entailed cases only | 68.0% | 58.0% | Substantially more resistant to lexical traps |
| RAG Hallucination Checks (Handwritten) | 17/20 (85%) | 14/20 (70%) | Strong general fact verification |
| Tool / Action Routing (Handwritten) | 15/18 (83%) | — | Missed 3 login/auth routing edge cases |
| Answer Grading (Handwritten) | 14/14 (100%) | — | Evaluates correctness reliably |
2. JevBench Public Dataset (231 Items)
Evaluated locally against the frozen open public split of JevBench (jevbench/datasets/public across 18 task families):
- Gevva e2b: 71.43% overall accuracy (165/231), 47.75% on the Hard tier (53/111).
- JevBench Composite Score: 77.54 (at $T=1.6$).
- Gevva e4b: 76.62% overall accuracy (177/231), 54.95% on the Hard tier (61/111).
- JevBench Composite Score: 77.28 (at $T=1.6$).
Transparency Note on JevBench:
- These scores were measured locally on the 231-item open public split and have not yet been evaluated by third-party maintainers on the private benchmark suite.
- The JevBench composite score weights Intelligence (25%), Calibration (25%), Speed (25%), and a hard-coded Cost tariff (25%).
- A temperature of $T=1.6$ was fitted specifically to minimize ECE on the 111 JevBench Hard tier items. However, on standard NLI test sets, $T=1.6$ flattens probabilities (raising ECE from 0.056 to 0.163). Therefore, Gevva ships with standard $T=1.0$ default inference.
⚠️ Known Limitations & Failure Modes
Every decision engine has blind spots. We encourage testing on your specific distribution:
- Numerical & Arithmetic Claims: Gevva can be unreliable with subtle arithmetic discrepancies. In our audit, the model accepted "revenue exceeded five billion" when the document stated "$4.2B" (83% confidence). For strict financial and numeric validation, pair Gevva with programmatic regex or numeric validators.
- Authentication & Password Routing: In tool-routing evaluations, requests involving account security, passwords, or authentication links were occasionally misrouted to general documentation rather than password-reset actions.
- CPU Latency on Long Documents: While short queries evaluate in ~200 ms on CPU, full 900+ token documents take ~20 seconds on CPU. GPU acceleration (CUDA or Apple MPS) is strongly recommended for production RAG pipelines.
- Confidence Thresholding:
Because real-world calibration varies by domain, do not rely on a fixed
0.80cutoff. Evaluate your specific positive/negative tradeoff curves to choose operational thresholds.
📦 Model Selection Guide
| Model | Hugging Face ID | Download Size | Best For |
|---|---|---|---|
| Gevva e2b | davidburhans/gevva-e2b |
~10.2 GB (5.1B params) | Fast text NLI, RAG hallucination checks, general tool routing. |
| Gevva e2b Multimodal | davidburhans/gevva-e2b-multimodal |
~10.2 GB (5.1B params) | Multimodal: Checking charts, invoices, and photos against text claims. |
| Gevva e4b | davidburhans/gevva-e4b |
~15.9 GB (5.8B params) | Deep reasoning: Multi-hop logic, complex contracts, high-stakes verification. |
🛠️ Train on Your Own Data in 1 Line
Fine-tune Gevva on proprietary enterprise policies, internal tools, or domain datasets:
# Auto-detects column names and label formats (.jsonl, .csv, .tsv):
gevva finetune --data my_company_data.jsonl --out-dir ./my_custom_model --epochs 3
For advanced options (LoRA vs full fine-tuning, Quantization-Aware Training, token bucketing), see the Custom Fine-Tuning Guide.
📄 License & Citation
Gevva is licensed under the open-source Apache 2.0 License. Base model weights inherit the Google Gemma Terms of Use.
@software{gevva2026,
author = {Burhans, Dave and Contributors},
title = {Gevva: Multimodal System 1 Decision Engine},
year = {2026},
publisher = {GitHub / Hugging Face},
url = {https://github.com/davidburhans/gevva}
}
Release files for gevva 1.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gevva-1.0.2.tar.gz | 115.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gevva-1.0.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 158.5 kB
Release files / gevva-1.0.2.tar.gz
| Download URL | gevva-1.0.2.tar.gz |
|---|---|
| Size | 115.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9a552528fcbe91f9268bd764b3a73ebfa432b51ea081f9c6256be847c3e85ce1
|
|
BLAKE2b-256 checksum How to use checksums |
bbaab32c5b89d5bf322da190bb42b48d4c574163950a9a16468c4d7558369cdc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.25 {"installer":{"name":"uv","version":"0.11.25","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Garuda Linux","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / gevva-1.0.2-py3-none-any.whl
| Download URL | gevva-1.0.2-py3-none-any.whl |
|---|---|
| Size | 42.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
018e1681e3f5bbde5dccc6f6ae5ab26bbb08c357000f4b0601f543e00e677141
|
|
BLAKE2b-256 checksum How to use checksums |
7f014b1d6b1d30b52a49dc2487d752435cc43817220a02b70c781adc11a32939
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.25 {"installer":{"name":"uv","version":"0.11.25","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Garuda Linux","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|