pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16
Mistral-Nemo-Instruct-2407 · GGUF F16
Converted and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure
🔬 This repository is part of a production-oriented evaluation series. Every model published under `pbhappliedsystems` has been independently evaluated using quant_eval v7.22.21/v7.22.22 — a behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.
📌 This is the full-precision F16 baseline repository. The evaluated quantized variants are published separately: `Q4_K_M`, `Q5_K_M`, and `Q8_0`. This card documents reference behavior. The quantized cards show how each compression level impacts behavioral fidelity.
Try the Live AI Agent Demo & Compare Models
Two complementary interfaces, both free:
1. Quick Demo — Single Model Testing
**Launch the PBH Applied Systems AI Agent Demo →**
Free · 3 example queries per agent type · No account required · No document upload
Test this model (or others) against three agent workflow templates:
- Reasoning & Analysis — Chain-of-thought decomposition and structured problem-solving Example queries: Should a startup build on cloud LLMs or self-host quantized models? · Analyze the trade-offs between model quantization and inference latency. · What are the cost implications of running 14B parameter models on edge devices?
- Document Intelligence — Long-context information extraction and Q&A Example queries: Extract key clauses from a contract and summarize legal risks. · Analyze a research paper and identify the novel contributions. · Compare market analysis reports and highlight strategic differences.
- Code & Automation — API scaffolding, data transformation, and task completion Example queries: Build a REST API endpoint for user authentication and include rate limiting. · Transform CSV data into a production-ready database schema. · Write a Python script to validate and clean messy customer records.
Every query runs on private GPU infrastructure — your prompts, reasoning traces, and outputs never leave the system and are never sent to Frontier model providers or cloud vendors. This matters for organizations with sensitive data, compliance requirements, or internal knowledge that cannot leave the building.
2. Model Comparison Arena — Side-by-Side Evaluation
**Launch the quant_eval Agent Arena Space →**
Compare two models in real-time across all three agent types.
The Agent Arena lets you:
- Select any two evaluated models (F16 or quantized variants) and test them side-by-side on the same queries
- View execution traces for both agents, showing chain-of-thought reasoning and tool dispatch
- Check the Model Leaderboard — a ranked table of all evaluated models with scores across four behavioral dimensions: Task Completion, Reasoning, Coherence, Instruction Following
- Read the Methodology tab — explanation of quant_eval, the 8 test families, and how to request a full evaluation report for models not yet in the series
Why compare in the Arena:
The demo shows what this model can do. The Arena shows how this model compares to others. If you're deciding between F16 and Q5KM, or between this model and another in the series, run the same query in the Arena with both selected. You'll see the exact differences in reasoning quality, tool dispatch, and output coherence — not benchmark scores, but real agent behavior.
Evaluation-backed leaderboard: Every score in the Arena comes from quant_eval runs published to Zenodo (DOI `10.5281/zenodo.22009419`). The numbers are not opinionated; they're measured.
Model Description
This repository contains the full-precision F16 GGUF of `mistralai/Mistral-Nemo-Instruct-2407`, a 12-billion parameter instruction-tuned model developed by Mistral AI in collaboration with NVIDIA (July 2024 release). Mistral-Nemo features the Tekken tokenizer and a 128,000-token context window — the largest in the PBH Applied Systems evaluated series outside of Qwen2.5-14B-Instruct-1M, which supports a 1 million token window.
The F16 format preserves all original float16 weights without quantization. In the PBH Applied Systems evaluation pipeline, the F16 variant serves as the reference baseline: it is evaluated first, then quantized variants are evaluated against the same fixture set, enabling direct measurement of quantization impact on behavioral fidelity.
For most production deployments, one of the quantized variants (Q4KM, Q5KM, or Q8_0) is the appropriate choice — offering substantial footprint and speed gains with controlled behavioral tradeoffs. The F16 is the right choice when maximum output fidelity, cleanest tool execution, or full 128K context at maximum precision are required.
Key Characteristics
- Parameters: 12B
- Format: GGUF F16 (full precision)
- File size: 24.5 GB
- SHA256:
070920655fab05a776d40d522ba17f55c1f663310f77c8fe57dd850e8dad10ef - Context window: 128,000 tokens (Tekken tokenizer)
- Minimum VRAM (GPU inference): ~26 GB
- Recommended GPU tier: A100 40 GB · RTX 4090 · 2× A10G
- Inference speed (eval hardware): 25.04–33.11 tokens/sec across three F16 runs of the identical artifact on NVIDIA RTX 4090 (observed harness throughput, not a controlled benchmark)
- Multilingual: English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, Japanese
PBH Applied Systems Evaluation — quant_eval v7.22.21/v7.22.22
Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.22.21/v7.22.22 Three evaluation runs published 2026-08-19 (Q4KM), 2026-08-15 (Q5KM), 2026-08-16 (Q8_0) Hardware: NVIDIA RTX 4090 · Seed: 42 · Evaluation date: August 14–16, 2026
Evaluation Data Published to Zenodo
All evaluation artifacts are published under the quant_eval Public Corpus (DOI: `10.5281/zenodo.22009419`):
- Run Provenance (D4): `10.5281/zenodo.22010462` — Selected run manifest and rollup fields, artifact SHA-256 hashes, runner and decoding configurations
- Golden Oracle Fixtures (D3): `10.5281/zenodo.22010278` — Test case specifications and fixture version crosswalk
- Family Pass Rates (D6): `10.5281/zenodo.22010623` — Aggregate pass rates per family, both runners, 95% Wilson confidence intervals
- Per-Case Behavioral Results (D1): `10.5281/zenodo.22009799` — 19,200 per-case rows, both precisions, all signals, case-level pass/fail
- Throughput Telemetry (D2): `10.5281/zenodo.22009987` — Per-call generation metrics, wall-time, token counts
- Efficiency & Footprint (D7): `10.5281/zenodo.22010723` — Stored artifact size, compression ratio, tokens/sec, observed wall-time ratios with hardware scope
Evaluation Contract: canonical_agentic_contract_v7.22.17:evaluation · canonical_agentic_contract_v7.22.18:evaluation
Comparability & Fixture Methodology
Identical Fixtures Across All Three Runs
All three Mistral-Nemo runs were evaluated against the same 1,600 cases. A full structural comparison of the fixture files across bundles shows the only differing key is the top-level version label: all cases, oracle expectations, and trace_sha256 values are identical, and fixture_schema_version is 7.22.13 in every bundle.
Scoring-contract revisions:
Baseline reproducibility: The F16 baseline was executed in all three runs, on different days and under both contract revisions. All eight family pass rates are identical across the three runs, and at case level 0 of 1,600 cases differ in any pairwise comparison. Every Q4KM, Q5KM, and Q8_0 delta is measured against the same baseline.
The Zenodo datasets record the fixture version-label crosswalk (D3) and per-run manifests (D4) for full reproducibility.
Per-Family Pass Rates — Full-Weight Baseline
All rates include 95% Wilson score confidence intervals. N = 200 cases per family.
Source: quant_eval Family Pass Rates, DOI `10.5281/zenodo.22010623`
Key Findings
Finding 1: Stateful and Tool Execution — Preserved at F16
Mistral-Nemo-Instruct-2407 at full precision demonstrates clean, reliable behavior on three critical agent families: perfect state retention (1.000 on statefulfollowup), flawless tool parsing (1.000 on toolcallonly), and hybrid answer+JSON correctness (1.000 on mixedbriefjson). These three families measure the core agent-execution layer — maintaining conversational state, dispatching tools reliably, and emitting structured responses alongside natural language. F16 performance is maximum on all three, establishing the baseline against which quantized variants are measured.
Finding 2: Multi-Step Planning — Below the 0.70 Operational Threshold at F16
The json_multistep family tests multi-step planning with oracle verification: pass rate 0.505 [0.436, 0.574]. At full precision, this model's planning capability is below the 0.70 operational threshold. The model struggles with medium and hard difficulty planning cases, even without quantization. This is a model characteristic, not a precision artifact — it reflects the model's inherent limitations on autonomous multi-step reasoning. Use this model for reasoning tasks requiring external planning scaffolds (ReAct agents, LlamaIndex hierarchical composition) rather than end-to-end autonomous planning. The tool-call layer is reliable; the planning layer needs guidance.
Finding 3: Structured Output — JSON Schema OK, Oracle Match is Harder
The json family (single-step structured JSON with constraints) shows a different pattern: pass rate 0.415 [0.349, 0.484]. Schema correctness reaches 100% of cases, but semantic content often diverges from the oracle plan. This is not a JSON parsing failure; it's a content-fidelity issue. Like multi-step planning, this behavior appears at F16 and is a model characteristic rather than a quantization artifact.
Quantization Variants Comparison
See how this model performs across all three quantized variants. All three were evaluated on identical fixtures against the same F16 baseline.
How to read this table: File size and VRAM determine deployment feasibility; Speed is the observed evaluation wall-time ratio (F16 ÷ quantized) on matched RTX 4090 hardware — not a controlled throughput benchmark; Stateful/Toolcall/Hybrid are pass rates; Best For describes the deployment context.
See the full cards:
- `Q4_K_M` — Extreme compression, significant behavioral cost
- `Q5_K_M` — Balanced quantization, minimal degradation
- `Q8_0` — Conservative quantization, high fidelity
Artifact Provenance
The artifact was produced from mistralai/Mistral-Nemo-Instruct-2407 using a custom-built llama.cpp conversion pipeline developed by PBH Applied Systems, without modification to model weights.
Artifact verification: SHA256 hashes are recorded in the run manifest (DOI `10.5281/zenodo.22010462`). Download the F16 GGUF and verify:
sha256sum Mistral_Nemo_Instruct_2407-F16.gguf
# Should output: 070920655fab05a776d40d522ba17f55c1f663310f77c8fe57dd850e8dad10efEvaluation Methodology
quant_eval v7.22.21/v7.22.22 is a behavioral evaluation harness developed by PBH Applied Systems. The two-run architecture evaluates the full-precision (F16) model first, then evaluates the quantized variant against the same fixture set, enabling direct measurement of behavioral degradation.
Fixture set: v7.22.13 / v7.22.22 evaluation split
- Total unique cases: 1,600 (200 per family × 8 families)
- Per-family test families and pass signals:
Evaluation hardware: NVIDIA RTX 4090 (24 GB VRAM) Runner: full_weight_llama_cpp (llama.cpp via Python binding, version 0.3.16) Decoding: temperature 0.3 (per Mistral AI's published recommendation), seed 42, topp 1.0 **Context:** 4,096-token context window (llama.cpp configuration) **Reproducibility:** Run manifest, fixtures, and per-case results published to Zenodo under the quanteval Public Corpus
Deployment Recommendations
✅ Deploy F16 when:
- Maximum fidelity matters more than cost. Schema-only tool calls (toolcallonly) and state tracking (statefulfollowup) both score 1.000; end-to-end tool dispatch (toolcall) scores 0.930.
- Reasoning requires clean outputs. The model's tool-dispatch and hybrid-response capabilities are maximized.
- Context window utilization is critical. 128K tokens at full precision, no quantization loss.
- You can afford ~26 GB VRAM. Enterprise GPU tier (A100, RTX 4090, dual A10G).
⚠️ Avoid F16 if:
- Budget is constrained. Q5KM and Q8_0 show no statistically significant difference from F16 in any of the eight families (McNemar p < 0.05: 0 of 8 each) at roughly 1/3 to 1/2 the footprint.
- Latency is the bottleneck. On matched RTX 4090 hardware, observed evaluation wall-time ratios (F16 ÷ quantized) were 2.70× for Q4KM, 2.56× for Q5KM, and 1.71× for Q8_0.
- Hardware is limited to <20 GB VRAM. Use Q5KM or Q8_0 instead.
About PBH Applied Systems
**PBH Applied Systems, LLC** is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization emphasizes engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.
Founder & Principal AI/ML Systems Architect — Patrick Hill, M.S.
Patrick Hill is the Founder & Principal AI/ML Systems Architect of PBH Applied Systems with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning and a B.S. in Business Finance.
Technical expertise spans: Python, SQL, Linux, Pandas, NumPy, scikit-learn, PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA, Flask APIs, Docker, CI/CD, Jupyter, Databricks, GGUF conversion, and quantization strategies.
Published Author: Patrick is the author of [Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz) — a 1,200+ page practitioner-oriented textbook adopted as required reading for CSC 373 – Machine Learning at the University of Advancing Technology.
Core Service Areas: LLM optimization & deployment · AI evaluation frameworks · Agentic AI infrastructure · Scalable AI application development · ML pipeline design & analytics · Model & agent cataloging.
🔬 About quant_eval & This Evaluation Series
**quant_eval** is a behavioral evaluation harness developed by PBH Applied Systems, LLC. It measures real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning — not perplexity or leaderboard proxies. Every model published under `pbhappliedsystems` has been independently evaluated using quant_eval before being recommended for any production role.
See it in action: **Live AI Agent Demo →**
📋 Interested in Deeper Engagement?
This model card documents behavioral evidence under quant_eval. If you're evaluating this model for production deployment, need guidance on quantization impact for your specific workload, or want to understand deployment tradeoffs under your infrastructure constraints, PBH Applied Systems offers several paths:
Async Intake Form — https://pbhappliedsystems.com/contact.html
Describe your use case, infrastructure, and evaluation needs. Responses are reviewed asynchronously.
Available Services:
- Evaluation Report — A written behavioral audit: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, deployment recommendation, and a quantization-level suggestion for your hardware and latency constraints. Evidence-standard numbers only. Typical engagements: $2,500–$5,000.
- Starter-Kit Recommender — Configured agent templates (reasoning, document intelligence, code automation) with proof metrics from your specific workload. Includes model selection, quantization strategy, and deployment architecture guidance tied to your infrastructure.
- Community — Join discussions on quantization confounds, model comparisons, and evaluation priorities. Active in Hugging Face org discussions, Discord, and Reddit threads.
Submit your details via the form above, and we'll route you to the appropriate engagement.
License
This GGUF repository inherits the license of the base model: Apache 2.0 — `mistralai/Mistral-Nemo-Instruct-2407`
The quanteval evaluation harness, fuzz prompt builder, and scoring implementation are proprietary to PBH Applied Systems, LLC and are not included in this repository. The golden oracle fixture set used in this evaluation is published under CC BY 4.0 as part of the quanteval Public Corpus (D3, DOI `10.5281/zenodo.22010278`).
GGUF conversion and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.22.21/v7.22.22 · Published 2026-08-19
