CoolFace
Modelpublic

pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
0likes413downloads
Model Card

Mistral-Nemo-Instruct-2407 · GGUF Q4\K\M

Quantized, converted, and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure

🔬 This repository is part of a production-oriented evaluation series. Every model published under `pbhappliedsystems` has been independently evaluated using quant_eval v7.22.21/v7.22.22 — a behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.
⚠️ Behavioral degradation from F16 to Q4_K_M is statistically significant in five of eight families (McNemar p < 0.05), and largest in three: hybrid responses (−0.280), tool-argument fidelity (−0.215), and stateful execution (−0.170). These are not theoretical risks — they are measured gaps. The evaluation data below specifies which tasks require guardrails or F16 deployment.

Try the Live AI Agent Demo & Compare Models

Two complementary interfaces, both free:

1. Quick Demo — Single Model Testing

**Launch the PBH Applied Systems AI Agent Demo →**

Free · 3 example queries per agent type · No account required · No document upload

Test this model (or others) against three agent workflow templates:

  • —Reasoning & Analysis — Chain-of-thought decomposition and structured problem-solving Example queries: Should a startup build on cloud LLMs or self-host quantized models? · Analyze the trade-offs between model quantization and inference latency. · What are the cost implications of running 14B parameter models on edge devices?
  • —Document Intelligence — Long-context information extraction and Q&A Example queries: Extract key clauses from a contract and summarize legal risks. · Analyze a research paper and identify the novel contributions. · Compare market analysis reports and highlight strategic differences.
  • —Code & Automation — API scaffolding, data transformation, and task completion Example queries: Build a REST API endpoint for user authentication and include rate limiting. · Transform CSV data into a production-ready database schema. · Write a Python script to validate and clean messy customer records.

Every query runs on private GPU infrastructure — your prompts, reasoning traces, and outputs never leave the system and are never sent to Frontier model providers or cloud vendors. This matters for organizations with sensitive data, compliance requirements, or internal knowledge that cannot leave the building.

2. Model Comparison Arena — Side-by-Side Evaluation

**Launch the quant_eval Agent Arena Space →**

Compare two models in real-time across all three agent types.

The Agent Arena lets you:

  • —Select any two evaluated models (F16 or quantized variants) and test them side-by-side on the same queries
  • —View execution traces for both agents, showing chain-of-thought reasoning and tool dispatch
  • —Check the Model Leaderboard — a ranked table of all evaluated models with scores across four behavioral dimensions: Task Completion, Reasoning, Coherence, Instruction Following
  • —Read the Methodology tab — explanation of quant_eval, the 8 test families, and how to request a full evaluation report for models not yet in the series

Why compare in the Arena:

The demo shows what this model can do. The Arena shows how this model compares to others. If you're deciding between F16 and Q5KM, or between this model and another in the series, run the same query in the Arena with both selected. You'll see the exact differences in reasoning quality, tool dispatch, and output coherence — not benchmark scores, but real agent behavior.

Evaluation-backed leaderboard: Every score in the Arena comes from quant_eval runs published to Zenodo (DOI `10.5281/zenodo.22009419`). The numbers are not opinionated; they're measured.


Model Description

This repository contains the 4-bit quantized (Q4\_K\_M) GGUF of `mistralai/Mistral-Nemo-Instruct-2407`, a 12-billion parameter instruction-tuned model developed by Mistral AI in collaboration with NVIDIA (July 2024 release). Mistral-Nemo features the Tekken tokenizer and a 128,000-token context window — the largest in the PBH Applied Systems evaluated series outside of Qwen2.5-14B-Instruct-1M, which supports a 1 million token window.

The Q4\K\M format applies 4-bit quantization with K-quant medium precision. As documented in the evaluation section below, Q4\K\M quantization produces measurable degradation on five of eight families — largest on hybrid responses, tool-argument fidelity, and stateful execution — while single-step structured output (json) and fuzz are unchanged. Multi-step planning shows a row-level gain that does not survive cluster adjustment.

The full-precision F16 baseline is published separately at `pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16`.

Key Characteristics

  • —Parameters: 12B
  • —Format: GGUF Q4\K\M
  • —File size: 7.48 GB
  • —SHA256: d3b7c7950abee2677290c1a3412940a0e1e3bc98449b42f34bc1d7ab6d4917b9
  • —Context window: 128,000 tokens (Tekken tokenizer)
  • —Minimum VRAM (GPU inference): ~10 GB (T4 class or better)
  • —Recommended GPU tier: NVIDIA T4 (16 GB) · RTX 3080/4080 · A10G
  • —Inference speed (eval hardware): avg 91.03 tokens/sec on NVIDIA RTX 4090
  • —Speedup vs F16 (same run, same hardware): 3.63× generation throughput (25.04 → 91.03 tokens/sec); 2.70× observed evaluation wall-time ratio
  • —Multilingual: English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, Japanese

PBH Applied Systems Evaluation — quant_eval v7.22.21

Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.22.21 Run ID: Mistral_Nemo_Instruct_2407_20260814_030505 · Fixtures: v7.22.13 (SHA256: ae7d7fa...) Evaluation date: August 14, 2026 · Seed: 42 · Hardware: NVIDIA RTX 4090 · Published: August 19, 2026

Evaluation Data Published to Zenodo

All evaluation artifacts are published under the quant_eval Public Corpus (DOI: `10.5281/zenodo.22009419`):

  • —Paired Degradation Statistics (D5): `10.5281/zenodo.22010557` — McNemar tests, Wilson intervals, cluster-adjusted paired deltas across all evaluated precisions
  • —Run Provenance (D4): `10.5281/zenodo.22010462` — Selected run manifest and rollup fields, artifact SHA-256 hashes, runner and decoding configurations
  • —Golden Oracle Fixtures (D3): `10.5281/zenodo.22010278` — Test case specifications and fixture version crosswalk
  • —Per-Case Behavioral Results (D1): `10.5281/zenodo.22009799` — 19,200 per-case rows, both precisions, all signals, case-level pass/fail
  • —Family Pass Rates (D6): `10.5281/zenodo.22010623` — Aggregate pass rates per family, both runners, 95% Wilson confidence intervals
  • —Throughput Telemetry (D2): `10.5281/zenodo.22009987` — Per-call generation metrics, wall-time, token counts
  • —Efficiency & Footprint (D7): `10.5281/zenodo.22010723` — Stored artifact size, compression ratio, tokens/sec, observed wall-time ratios with hardware scope

Evaluation Contract: canonical_agentic_contract_v7.22.17:evaluation


Comparability & Fixture Methodology

Identical Fixtures Across All Three Runs

All three Mistral-Nemo runs were evaluated against the same 1,600 cases. A full structural comparison of the fixture files across bundles shows the only differing key is the top-level version label: all cases, oracle expectations, and trace_sha256 values are identical, and fixture_schema_version is 7.22.13 in every bundle.

Run IDQuantized VariantHarnessEvaluation ContractFixture File SHA256
Mistral_Nemo_Instruct_2407_20260814_030505Q4KMv7.22.21v7.22.17:evaluationae7d7fa...
Mistral_Nemo_Instruct_2407_20260815_113254Q5KMv7.22.22v7.22.18:evaluation0673b5e...
Mistral_Nemo_Instruct_2407_20260816_084553Q8_0v7.22.22v7.22.18:evaluation0673b5e...

Baseline reproducibility: The F16 baseline was executed in all three runs, on different days and under both contract revisions. All eight family pass rates are identical across the three runs, and at case level 0 of 1,600 cases differ in any pairwise comparison. The Q4KM, Q5KM, and Q8_0 deltas are all measured against the same baseline.

The Zenodo datasets record the fixture version-label crosswalk (D3) and per-run manifests (D4) for full reproducibility.


Per-Family Paired Comparison — F16 Baseline vs. Q4\K\M

All rates include 95% Wilson score confidence intervals. McNemar p-values test statistical significance of paired differences. N = 200 cases per family.

FamilyF16 Pass RateQ4_K_M Pass RateΔ (Q4−F16)95% CIMcNemar pStatus
mixed_brief_json1.0000.720−0.280[−0.356, −0.190]2.78e−17⚠️ CRITICAL LOSS
toolcall_only1.0000.785−0.215[−0.287, −0.133]2.27e−13⚠️ SEVERE LOSS
stateful_followup1.0000.830−0.170[−0.237, −0.094]1.16e−10⚠️ SIGNIFICANT LOSS
toolcall0.9300.805−0.125[−0.226, −0.018]8.02e−04⚠️ NOTABLE LOSS
mcq0.9200.845−0.075[−0.138, −0.009]7.29e−04⚠️ MILD LOSS
json0.4150.425+0.010[−0.035, 0.054]0.6875✓ NO CHANGE
json_multistep0.5050.590+0.085[+0.003, +0.163]2.32e−03↑ ROW-LEVEL GAIN*
fuzz0.5250.530+0.005[−0.037, 0.047]1.000✓ NO CHANGE

\ `json_multistep` has 167 semantically unique fixtures out of 200. Under cluster adjustment (effective n = 167), the delta becomes +0.048 [−0.037, +0.130]* — the row-level significance does not survive clustering.

Sources:


Key Findings

Finding 1: Hybrid Response Quality — Critical Degradation (−0.280)

The largest quantization impact appears on `mixed_brief_json` — a family combining natural language answers with valid JSON schema.

MetricF16Q4_K_MΔ
Pass Rate1.0000.720−0.280
95% CI[0.981, 1.000][0.654, 0.778][−0.356, −0.190]
McNemar p——2.78e−17

What fails: 56 of 200 Q4KM cases fail where F16 succeeds. All 56 fail the answer_line_ok signal; 12 of those also fail semantic_payload_ok. JSON parsing and schema validity hold on all 200 cases — Q4KM produces valid JSON but struggles to produce a clean answer line alongside it.

Why it matters: This family measures the ability to emit conversational answers and structured output in a single response. Workflows depending on hybrid responses will experience measurable quality loss under Q4KM.

Mitigation: Use F16 if hybrid responses are central. If Q4KM is required, separate answer and JSON into sequential requests. Monitor output quality empirically — 0.720 pass rate may be acceptable depending on fallback strategy.

Finding 2: Tool-Argument Fidelity — Severe Degradation (−0.215)

The `toolcall_only` family tests bare schema-only tool calls — tool name and arguments parsed in isolation.

MetricF16Q4_K_MΔ
Pass Rate1.0000.785−0.215
95% CI[0.981, 1.000][0.723, 0.836][−0.287, −0.133]
McNemar p——2.27e−13

What fails: 43 of 200 Q4KM cases fail on tool dispatch. 31 fail on both tool name and arguments; 8 fail on tool name alone; 4 fail on arguments alone.

Why it matters: Tool calling is the core agent action. A 0.785 pass rate (78.5% success on first attempt) is borderline for production. Roughly one in five tool calls needs retry or fallback.

Mitigation: Implement tool-call validation and retry logic in your agent framework. Normalize tool-argument JSON in the prompt. For high-stakes workflows, use F16 or accept a 1-2 retry budget per agent step.

Finding 3: Stateful Execution — Significant Degradation (−0.170)

The `stateful_followup` family tests two-turn conversations where turn-2 must track state from turn-1.

MetricF16Q4_K_MΔ
Pass Rate1.0000.830−0.170
95% CI[0.981, 1.000][0.772, 0.876][−0.237, −0.094]
McNemar p——1.16e−10

What fails: 34 of 200 Q4KM cases fail on turn-2 exact match. The model drops context or misapplies state updates under quantization.

Why it matters: Stateful conversations are central to agentic workflows. If a two-turn exchange fails to maintain state, the agent cannot sustain context.

Mitigation: Use external state tracking (store turn-1 outputs explicitly; inject them into turn-2 prompts). Implement turn-2 validation. For conversation-heavy workflows, Q4KM requires scaffolding that F16 does not.


Efficiency & Footprint

Source: quant_eval Efficiency & Footprint, DOI `10.5281/zenodo.22010723`

MetricF16Q4_K_MRatio
File size24.5 GB7.48 GB0.305× (69.5% reduction)
VRAM required~26 GB~10 GB0.385× (61.5% reduction)
Tokens/sec (same run, RTX 4090)25.0491.033.63× faster
Evaluation wall time (same run, RTX 4090)——2.70× faster

For long-running deployments, Q4KM reduces infrastructure cost and latency significantly. The behavioral costs — measured and documented above — must be weighed against these gains.


Quantization Variants Comparison

See how this model performs across all three quantized variants. All three were evaluated on identical fixtures against the same F16 baseline.

VariantFile SizeVRAMSpeedStatefulToolcall-onlyHybridBest For
F1623.9 GB~24.5 GB1.0×1.0001.0001.000Maximum fidelity; enterprise GPU tier
Q4_K_M7.48 GB~10 GB2.70×0.8300.7850.720Extreme cost optimization; requires guardrails
Q5_K_M8.73 GB~11 GB2.56×1.0000.9951.000Balanced; minimal risk; broad compatibility
Q8_013.0 GB~15 GB1.71×1.0000.9951.000Conservative; preserves behavior; modest speedup

How to read this table: Speed is the observed evaluation wall-time ratio (F16 ÷ quantized) on matched RTX 4090 hardware — not a controlled throughput benchmark. Stateful/Toolcall-only/Hybrid are pass rates.

See the full cards:

  • —`F16` — Baseline reference
  • —`Q5_K_M` — Balanced quantization, minimal degradation
  • —`Q8_0` — Conservative quantization, high fidelity

Artifact Provenance

ArtifactFormatSizeSHA256
Mistral_Nemo_Instruct_2407_Q4_K_M.ggufGGUF Q4\K\M7.48 GBd3b7c7950abee2677290c1a3412940a0e1e3bc98449b42f34bc1d7ab6d4917b9
F16 (companion repo)GGUF F1624.5 GB070920655fab05a776d40d522ba17f55c1f663310f77c8fe57dd850e8dad10ef

Both artifacts were produced from mistralai/Mistral-Nemo-Instruct-2407 using a custom-built llama.cpp conversion and quantization pipeline developed by PBH Applied Systems.

Artifact verification: SHA256 hashes are recorded in the run manifest (DOI `10.5281/zenodo.22010462`). Download the Q4KM GGUF and verify:

bash
sha256sum Mistral_Nemo_Instruct_2407_Q4_K_M.gguf
# Should output: d3b7c7950abee2677290c1a3412940a0e1e3bc98449b42f34bc1d7ab6d4917b9

Evaluation Methodology

quant_eval v7.22.21 is a behavioral evaluation harness developed by PBH Applied Systems. The two-run architecture evaluates both the full-precision (F16) and quantized variant against an identical fixture set, enabling direct comparison of behavioral degradation.

Fixture set: v7.22.13 evaluation split

  • —SHA256: ae7d7fa90a2ef8d4e673108886599248dff52c83829ac79b51e332d01421a23f
  • —Total unique cases: 1,600 (200 per family × 8 families)
FamilyDescriptionPass Signals
mixed_brief_jsonHybrid: natural language answer + valid JSON blockanswerlineok, jsonparseok, schemaok, semanticpayload_ok
toolcall_onlyBare schema-only tool call; strict tool name + args checktoolnameok, args_ok
stateful_followupTwo-turn state tracking; turn-2 correct given turn-1turn1/2parseok, turn1/2exactmatch
toolcallTool call embedded in response; parse + schema validationstage1toolparseok, selectionok, stage1argsok, toolexecok, stage2evaluable, finalequiv_ok
mcqMultiple-choice extractionsemanticchoiceok
json_multistepMulti-step planning with self-check and oracle verificationparseok, oracleequiv_ok
jsonSingle-step structured JSON with constraint rulesschemaok, finalequivok, constraintsok
fuzzProperty-based regression; structured placement correctnessplanexactmatch, finalequivok, constraints_ok

Evaluation hardware: NVIDIA RTX 4090 (24 GB VRAM) Runner: quantized_llama_cpp (llama.cpp via Python binding, version 0.3.16) Decoding: temperature 0.3 (per Mistral AI's published recommendation), seed 42, topp 1.0 **Context:** 4,096-token context window (llama.cpp configuration) **Reproducibility:** Run manifest, fixtures, per-case results, and paired statistics published to Zenodo under the quanteval Public Corpus


Deployment Recommendations

✅ Deploy Q4_K_M when:

  • —Hardware is constrained to <16 GB VRAM. Q4KM fits on T4 (16 GB) and RTX 3080/4080 tiers.
  • —Latency is the primary constraint. On the same RTX 4090, 91.03 vs 25.04 tokens/sec (3.63×) and a 2.70× observed evaluation wall-time ratio vs F16.
  • —Your agent workflow has external state management. If using LlamaIndex, LangChain, or custom state stores, the 0.830 stateful rate is acceptable with proper scaffolding.
  • —Tool call validation is already in place. Implement retry logic for tool dispatch; expect roughly 1 retry per 5 tool calls (21.5% toolcall_only failure rate).
  • —Hybrid responses are not core to your workflow. If answer-line + JSON in one response is not required, this limitation is not relevant.

⚠️ Avoid Q4_K_M if:

  • —Hybrid responses are central. The 0.720 pass rate is below operational threshold without additional mitigation.
  • —Tool calling must succeed on the first attempt. The 0.785 rate requires retry budgets that slow production agents.
  • —Stateful conversations lack external state tracking. The 0.830 rate implies about 17 turn-2 failures per 100 conversations without it.
  • —You can afford ~24.5 GB VRAM. F16 scores 1.000 on all three families where Q4KM loses most.

🔄 Hybrid Approach (Recommended for Production):

Deploy Q4_K_M for the majority of inference, with F16 fallback for high-stakes paths:

  1. 1.Normal agent steps: Q4KM (fast, low-cost)
  2. 2.If tool call fails: Retry with F16 (1.000 on toolcall_only and 0.930 on end-to-end toolcall in this evaluation; higher latency acceptable)
  3. 3.If hybrid response fails: Reissue with F16 (clean output)
  4. 4.If state tracking breaks: Inject explicit state into prompt; reissue with Q4KM

This captures the majority of Q4KM's speed and footprint gains while handling edge cases that require F16 reliability.


About PBH Applied Systems

**PBH Applied Systems, LLC** is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization emphasizes engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.

Founder & Principal AI/ML Systems Architect — Patrick Hill, M.S.

Patrick Hill is the Founder & Principal AI/ML Systems Architect of PBH Applied Systems with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning and a B.S. in Business Finance.

Technical expertise spans: Python, SQL, Linux, Pandas, NumPy, scikit-learn, PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA, Flask APIs, Docker, CI/CD, Jupyter, Databricks, and quantization strategies.

Published Author: Patrick is the author of [Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz) — a 1,200+ page practitioner-oriented textbook adopted as required reading for CSC 373 – Machine Learning at the University of Advancing Technology.

Core Service Areas: LLM optimization & deployment · AI evaluation frameworks · Agentic AI infrastructure · Scalable AI application development · ML pipeline design & analytics · Model & agent cataloging.


🔬 About quant_eval & This Evaluation Series

**quant_eval** is a behavioral evaluation harness developed by PBH Applied Systems, LLC. It measures real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning — not perplexity or leaderboard proxies. Every model published under `pbhappliedsystems` has been independently evaluated using quant_eval before being recommended for any production role.

See it in action: **Live AI Agent Demo →**


📋 Interested in Deeper Engagement?

This model card documents behavioral evidence under quant_eval. If you're evaluating this model for production deployment, need guidance on quantization impact for your specific workload, or want to understand deployment tradeoffs under your infrastructure constraints, PBH Applied Systems offers several paths:

Async Intake Form — https://pbhappliedsystems.com/contact.html

Describe your use case, infrastructure, and evaluation needs. Responses are reviewed asynchronously.

Available Services:

  • —Evaluation Report — A written behavioral audit: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, deployment recommendation, and a quantization-level suggestion for your hardware and latency constraints. Evidence-standard numbers only. Typical engagements: $2,500–$5,000.
  • —Starter-Kit Recommender — Configured agent templates (reasoning, document intelligence, code automation) with proof metrics from your specific workload. Includes model selection, quantization strategy, and deployment architecture guidance tied to your infrastructure.
  • —Community — Join discussions on quantization confounds, model comparisons, and evaluation priorities. Active in Hugging Face org discussions, Discord, and Reddit threads.

Submit your details via the form above, and we'll route you to the appropriate engagement.


License

This GGUF repository inherits the license of the base model: Apache 2.0 — `mistralai/Mistral-Nemo-Instruct-2407`

The quanteval evaluation harness, fuzz prompt builder, and scoring implementation are proprietary to PBH Applied Systems, LLC and are not included in this repository. The golden oracle fixture set used in this evaluation is published under CC BY 4.0 as part of the quanteval Public Corpus (D3, DOI `10.5281/zenodo.22010278`).


GGUF conversion, quantization, and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.22.21 · Run ID: `Mistral_Nemo_Instruct_2407_20260814_030505` · Published 2026-08-19