davidburhans/gevva-e4b
⚡ Gevva e4b: Deep Reasoning AI Decision Engine
<p align="center"> <a href="https://huggingface.co/spaces/davidburhans/gevva-demo"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Space-Interactive%20Demo-blue.svg" alt="Interactive Demo"></a> <a href="https://pypi.org/project/gevva/"><img src="https://img.shields.io/pypi/v/gevva.svg?logo=pypi&logoColor=white" alt="PyPI"></a> <a href="https://github.com/davidburhans/gevva"><img src="https://img.shields.io/badge/Latency-17.8ms%20(RTX%205090)-orange.svg" alt="Latency"></a> <a href="https://github.com/davidburhans/gevva"><img src="https://img.shields.io/badge/ARC--Challenge-84%25%20Accuracy-brightgreen.svg" alt="ARC-Challenge"></a> <a href="https://opensource.org/licenses/Apache-2.0"><img src="https://img.shields.io/badge/License-Apache--2.0-green.svg" alt="License"></a> </p>
17-millisecond deep reasoning, high-stakes fact checking, multi-tool routing & multi-image verification.
Built on Google's Gemma 4 E4B (4.5B parameters, 42 layers, 128K context) • 84% on ARC-Challenge • 76.6% on JevBench
🤔 What is Gevva e4b? (The 30-Second Explainer)
When you ask ChatGPT or Claude to verify a document or route a customer request, it generates words one token at a time, like a human slowly typing out an explanation. That takes 2 to 5 seconds and burns expensive GPU compute.
That is fine for writing stories, but it is painfully slow and expensive for decisions:
- "Did the AI hallucinate this medical summary, or is it grounded in the research paper?"
- "Which of these 10 enterprise APIs should execute this user workflow?"
- "Did this UI screenshot change in an unexpected way after deployment?"
- "Did the student's solution satisfy all 5 steps in the grading rubric?"
The Solution: An Instant "Reflex Engine" with Deep Reasoning
Psychologist Daniel Kahneman famously described human thought in two modes:
- System 1 (Fast Reflexes): Instant decisions in 15–20 milliseconds (dodging a ball, recognizing a face).
- System 2 (Slow Reasoning): Writing essays or working out long equations step-by-step.
Gevva e4b is the 4.5B deep reasoning model in the Gevva family. In a single 17.8 millisecond forward pass, it evaluates evidence and outputs confident, mathematically calibrated decision probabilities.
┌────────────────────────────────────────┐ ┌────────────────────────────────────────┐
│ SYSTEM 1: GEVVA e4b │ │ SYSTEM 2: CHATGPT / CLAUDE │
│ ⚡ Evaluates decisions in 17.8 ms │ vs │ 🐢 Generates text word-by-word │
│ 🎯 84% ARC-Challenge accuracy │ │ ⏳ Takes 2,000 - 5,000 milliseconds │
│ 💰 90%+ cheaper compute cost │ │ 💸 Expensive GPU server bills │
│ 🔍 Best for: Fact-checks, routing, │ │ ✍️ Best for: Creative writing, long │
│ guardrails, and visual audit │ │ essays, and coding from scratch │
└────────────────────────────────────────┘ └────────────────────────────────────────┘🚀 3-Line Quickstart
from gevva import load
# 1. Load the decision engine (GPU or CPU)
engine = load("davidburhans/gevva-e4b")
# 2. Instant Fact Check / Hallucination Detection
verdict, probs = engine.grade(
premise="Albert Einstein won the Nobel Prize in Physics in 1921 for his discovery of the photoelectric effect.",
hypothesis="Einstein won the Nobel Prize for his theory of General Relativity."
)
print(verdict)
# Output: 'contradiction' (p_contradiction = 0.96)Smart Tool & API Routing
tools = [
"send_email(to, subject, body): Sends an email message to a contact",
"search_database(query): Queries internal company records and docs",
"refund_charge(charge_id, reason): Issues a payment refund to a customer",
]
best_idx, scores = engine.rerank(
premise="A customer says they were double-billed and wants their money back.",
options=tools
)
print(f"Selected: {tools[best_idx]}")
# Output: Selected: refund_charge(...)📊 Evaluated Benchmarks on RTX 5090
Gevva e4b was evaluated on standard reasoning batteries on an NVIDIA GeForce RTX 5090:
🛠️ Architecture Details
- Foundation Backbone: Google's
gemma-4-E4B-it(4.5B parameters, 42 layers). - Classification Head: Last-token pooled representation normalized with Gemma4RMSNorm into a 3-class linear head.
- Label Convention:
0: Contradiction / Refutes / Invalid1: Entailment / Supports / Valid2: Neutral / Unverifiable / Inconclusive- Calibration: Trained with strictly proper Brier score calibration ($ECE = 0.0224$ on validation).
📄 License & Citations
Released under the Apache 2.0 License.
@software{gevva2026,
author = {David Burhans},
title = {Gevva: Multimodal 128K System 1 Decision Engine},
year = {2026},
url = {https://github.com/davidburhans/gevva}
}