CoolFace
Modelpublic

pbhappliedsystems/qwen-2.5-14B-instruct-1m-gguf-Q4-K-M

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
3likes578downloads
Model Card

Qwen2.5-14B-Instruct-1M · GGUF Q4KM

Converted and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure

🔬 This repository is part of a production-oriented evaluation series. Every model published under `pbhappliedsystems` has been independently evaluated using quant_eval v7.22.21/v7.22.22 — a behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.
📌 This is the Q4_K_M quantized variant. The full-precision F16 baseline is published separately: `F16`. This card shows how Q4KM compression compares to F16 across all behavioral dimensions. Key tradeoff: MCQ extraction degrades significantly (−15 points, p < 0.001), but tool dispatch and state tracking remain strong.
🌍 1M-Context Extended Window. Qwen2.5-14B-Instruct-1M supports up to 1,048,576 tokens of context — a differentiator for document-heavy workloads, long conversation histories, and multi-document analysis. Evaluation context was capped to 8,192 tokens (common Modal limit); full 1M capability is untested here but preserved in the model.

Try the Live AI Agent Demo & Compare Models

Two complementary interfaces, both free:

1. Quick Demo — Single Model Testing

**Launch the PBH Applied Systems AI Agent Demo →**

Free · 3 example queries per agent type · No account required · No document upload

Test this model (or others) against three agent workflow templates:

  • —Reasoning & Analysis — Chain-of-thought decomposition and structured problem-solving Example queries: Should a startup build on cloud LLMs or self-host quantized models? · Analyze the trade-offs between model quantization and inference latency. · What are the cost implications of running models with extended context windows in production?
  • —Document Intelligence — Long-context information extraction and Q&A Example queries: Extract key clauses from a contract and summarize legal risks. · Analyze a research paper and identify the novel contributions. · Compare market analysis reports and highlight strategic differences.
  • —Code & Automation — API scaffolding, data transformation, and task completion Example queries: Build a REST API endpoint for user authentication and include rate limiting. · Transform CSV data into a production-ready database schema. · Write a Python script to validate and clean messy customer records.

Every query runs on private GPU infrastructure — your prompts, reasoning traces, and outputs never leave the system and are never sent to Frontier model providers or cloud vendors. This matters for organizations with sensitive data, compliance requirements, or internal knowledge that cannot leave the building.

2. Model Comparison Arena — Side-by-Side Evaluation

**Launch the quant_eval Agent Arena Space →**

Compare two models in real-time across all three agent types.

The Agent Arena lets you:

  • —Select any two evaluated models (F16 or quantized variants) and test them side-by-side on the same queries
  • —View execution traces for both agents, showing chain-of-thought reasoning and tool dispatch
  • —Check the Model Leaderboard — a ranked table of all evaluated models with scores across four behavioral dimensions: Task Completion, Reasoning, Coherence, Instruction Following
  • —Read the Methodology tab — explanation of quant_eval, the 8 test families, and how to request a full evaluation report for models not yet in the series

Why compare in the Arena:

The demo shows what this model can do. The Arena shows how this model compares to others. If you're deciding between F16 and Q4KM, or between 14B and other sizes in the series, run the same query in the Arena with both selected. You'll see the exact differences in reasoning quality, tool dispatch, and output coherence — not benchmark scores, but real agent behavior.

Evaluation-backed leaderboard: Every score in the Arena comes from quant_eval runs published to Zenodo (DOI `10.5281/zenodo.22009419`). The numbers are not opinionated; they're measured.


Model Description

This repository contains the Q4_K_M quantized GGUF of `Qwen/Qwen2.5-14B-Instruct-1M`, a 14-billion parameter instruction-tuned model developed by Alibaba Qwen Team (January 2025). This variant features an extended 1,048,576-token (1M) context window and multilingual instruction-following capability.

Q4KM is a mixed-precision K-quant: mostly 4-bit (Q4_K) weights, with selected tensors kept at 6-bit (Q6_K), delivering a 3.29× compression ratio compared to F16. On this model, the primary quantization cost is MCQ extraction: −15 percentage points (p < 0.001). Tool dispatch and stateful execution remain nearly perfect. If your use case does not emphasize multiple-choice extraction, Q4KM offers substantial efficiency gains with acceptable tradeoffs.

Key Characteristics

  • —Parameters: 14B
  • —Format: GGUF Q4KM (4-bit mixed-precision)
  • —File size: 9.0 GB
  • —SHA256: 5ad529ff2b1b192f31c8a638fe8756a0c628904e2ded797c11f9194216976973
  • —Context window: 1,048,576 tokens (1M extended)
  • —Evaluation context: 8,192 tokens (common Modal cap)
  • —Minimum VRAM (GPU inference): ~10 GB
  • —Recommended GPU tier: T4 · L4 · A10G · RTX 4090
  • —Inference speed (eval hardware): avg 31.69 tokens/sec on Modal A10G GPU
  • —Slowdown vs. F16: 0.96× (negligible; near-parity); Inference times for the Q4KM variant would be less than the F16 variant when both are measured on the A100 GPU.
  • —Footprint reduction: 69.6% smaller than F16
  • —Multilingual: English, Chinese, and other languages

PBH Applied Systems Evaluation — quant_eval v7.22.22

Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.22.22 Run ID: Qwen2.5_14B_Instruct_1M_20260815_220633 · Fixtures: v7.22.22 (SHA256: 0673b5e5...) Evaluation date: August 15, 2026 · Seed: unsupported (Modal); topk=40, topp=0.95 · Hardware: Modal GPU cluster · Published: August 15, 2026

Evaluation Data Published to Zenodo

All evaluation artifacts are published under the quant_eval Public Corpus (DOI: `10.5281/zenodo.22009419`):

  • —Run Provenance (D4): `10.5281/zenodo.22010462` — Selected run manifest and rollup fields, artifact SHA-256 hashes, runner and decoding configurations
  • —Golden Oracle Fixtures (D3): `10.5281/zenodo.22010278` — Test case specifications and fixture version crosswalk
  • —Family Pass Rates (D6): `10.5281/zenodo.22010623` — Aggregate pass rates per family, both runners, 95% Wilson confidence intervals
  • —Per-Case Behavioral Results (D1): `10.5281/zenodo.22009799` — 19,200 per-case rows, both precisions, all signals, case-level pass/fail
  • —Throughput Telemetry (D2): `10.5281/zenodo.22009987` — Per-call generation metrics, wall-time, token counts
  • —Efficiency & Footprint (D7): `10.5281/zenodo.22010723` — Stored artifact size, compression ratio, tokens/sec, observed wall-time ratios with hardware scope

Evaluation Contract: canonical_agentic_contract_v7.22.18:evaluation


Comparability & Fixture Methodology

Fixtures Enable Valid Variant Comparison

The quanteval evaluation harness uses a consistent fixture set to measure both full-precision and quantized variants. The F16 and Q4K_M variants were evaluated under the same fixture set, enabling direct comparison.

Fixture Generation:

VariantRun IDFixture SHA256Evaluation ContractRationale
F16 & Q4KM20260815_2206330673b5e5...v7.22.18:evaluationSingle run evaluating both runners sequentially on Modal

What this means: Results are directly comparable. Both runners used identical fixtures, enabling precise measurement of Q4KM behavioral change relative to F16.

Note on cross-run comparability: Fixture content is identical across all quanteval runs in the series; only the version label differs (this run: v7.22.22; 7B run: v7.22.19). The 14B is directly comparable with the 32B run (same substrate, decoding, and contract). It is not directly comparable with the 7B, which ran on a local RTX 4090 under different decoding settings (topp 1.0, seed applied, 4,096-token context) and contract v7.22.17.


Paired Comparison: Q4KM vs. F16 Baseline

Quantization effect on all families — directly measured.

Deltas represent pass-rate change (Q4KM − F16), with 95% confidence intervals (paired multinomial, Bonferroni-adjusted Wilson) and paired McNemar p-values for statistical significance. N = 200 cases per family.

FamilyF16 RateQ4_K_M RateΔ95% CIMcNemar pPattern
stateful_followup1.0001.000±0.000[−0.025, +0.025]1.000Perfect parity
toolcall0.9500.990+0.040[−0.006, +0.084]0.008Gain (p < 0.05; CI includes 0)
toolcall_only0.8800.900+0.020[−0.024, +0.063]0.219Nominal
fuzz0.6350.705+0.070[−0.009, +0.145]0.009Gain (p < 0.05; CI includes 0)
json0.5700.565−0.005[−0.091, +0.081]1.000Nominal
mixedbriefjson0.7800.760−0.020[−0.081, +0.042]0.424Nominal
json_multistep0.5350.585+0.050[−0.010, +0.107]0.013Gain (p < 0.05; CI includes 0)*
mcq0.9700.820−0.150[−0.224, −0.069]6.94e-08Significant loss

\* json_multistep has 167 semantically unique fixtures out of 200. Cluster-adjusted (effective n = 167): +0.060 [−0.011, +0.127].|

Summary: Four of eight families differ significantly at McNemar p < 0.05. MCQ is the one robust loss (−15.0 pp, p < 0.001; CI excludes 0). Toolcall (+4.0 pp, p = 0.008), fuzz (+7.0 pp, p = 0.009), and json_multistep (+5.0 pp, p = 0.013) improve, though their Bonferroni-adjusted CIs include 0. Stateful execution holds perfect parity. The MCQ loss is the dominant quantization story for this model.


Per-Family Pass Rates — Q4KM Variant

All rates include 95% Wilson score confidence intervals. N = 200 cases per family.

FamilyPass Rate95% CINotes
stateful_followup1.000[0.981, 1.000]Perfect parity with F16
toolcall0.990[0.964, 0.997]+4.0 pp over F16; excellent dispatcher
toolcall_only0.900[0.851, 0.934]+2.0 pp over F16 (not significant)
fuzz0.705[0.638, 0.764]+7.0 pp over F16
json0.565[0.496, 0.632]−0.5 pp vs F16 (not significant)
mixedbriefjson0.760[0.696, 0.814]−2.0 pp vs F16 (not significant)
json_multistep0.585[0.516, 0.651]+5.0 pp over F16; effective n=167
mcq0.820[0.761, 0.867]−15.0 pp from F16; 36 failures vs. 6 on F16

Source: quant_eval Family Pass Rates, DOI `10.5281/zenodo.22010623`


Key Findings

Finding 1: MCQ Extraction — Quantization's Cost

The headline story: MCQ pass rate drops 15 percentage points under Q4KM (0.970 → 0.820; p < 0.001). This is the largest single quantization effect measured in this run. Multiple-choice extraction — selecting the correct option from a set — is the family most affected here; the cause was not isolated in this evaluation. If your use case emphasizes MCQ-style tasks (classification, structured selection), Q4KM requires validation or reranking logic.

Finding 2: State Tracking & Tool Dispatch — Robust to Quantization

Stateful execution remains perfect (1.000 both), and toolcall dispatch improves under Q4KM (+4.0 pp, p = 0.008; Bonferroni CI includes 0). This suggests that for agentic workflows emphasizing multi-turn state retention and tool composition, quantization is not a limiting factor on this fixture set.

Finding 3: Regression Testing Improves

The fuzz family (property-based structural correctness) gains +7.0 pp (p = 0.009; Bonferroni CI [−0.009, +0.145] includes 0). The cause was not isolated in this evaluation; treat it as a secondary effect.


Efficiency Gains

Footprint

MetricF16Q4_K_MReduction
File size29.5 GB9.0 GB−69.6%
GPU VRAM needed~31 GB~10 GB−68%
Disk I/O (during loading)29.5 GB9.0 GB−69.6%

Throughput

MetricF16Q4_K_MRatio
Tokens/sec33.0631.690.96×
Generation time per 1K tokens30.2 sec31.6 sec0.96×
Evaluation wall time (1,600 cases)2,531 sec2,632 sec0.96×

Hardware caveat: F16 was evaluated on a Modal A100-40GB and Q4KM on a Modal A10G. The observed ratio reflects an accelerator change and is not a precision comparison. Running on the same hardware as its F16 baseline, the quantized model would have smaller inference times.

Interpretation: Q4KM delivers a 69.6% footprint reduction; the observed 0.96× speed ratio was measured on a smaller GPU than F16 and does not represent a quantization slowdown. If MCQ extraction is not critical for your use case, Q4KM is a strong choice.


Quantization Variants Comparison

VariantFile SizeVRAMSpeedStatefulToolcall (end-to-end)MCQBest For
F1629.5 GB~31 GB1.0×1.0000.9500.970MCQ-critical workflows; maximum fidelity
Q4_K_M9.0 GB~10 GB0.96×*1.0000.9900.820Recommended for most agentic deployments

\* Observed on unmatched hardware (F16 on A100-40GB, Q4KM on A10G) — see Hardware caveat above.

See the full F16 card:

  • —`F16` — Full-precision baseline; only if MCQ validation cannot be implemented

Artifact Provenance

ArtifactFormatSizeSHA256
qwen-2.5-14B-instruct-1m-gguf-Q4-K-M.ggufGGUF Q4KM9.0 GB5ad529ff2b1b192f31c8a638fe8756a0c628904e2ded797c11f9194216976973

The artifact was produced from Qwen/Qwen2.5-14B-Instruct-1M using a custom-built llama.cpp conversion and quantization pipeline developed by PBH Applied Systems. Artifact verification: SHA256 hashes are recorded in the run manifest (DOI `10.5281/zenodo.22010462`). Download the Q4KM GGUF and verify:

bash
sha256sum qwen-2.5-14B-instruct-1m-gguf-Q4-K-M.gguf
# Should output: 5ad529ff2b1b192f31c8a638fe8756a0c628904e2ded797c11f9194216976973

Evaluation Methodology

quant_eval v7.22.22 is a behavioral evaluation harness developed by PBH Applied Systems. The two-run architecture evaluates the full-precision (F16) model first, then evaluates the quantized variant against the same fixture set, enabling direct measurement of behavioral degradation or improvement.

Fixture set: v7.22.22 evaluation split

  • —Total unique cases: 1,600 (200 per family × 8 families)
  • —Per-family test families and pass signals:
FamilyDescriptionPass Signals
json_multistepMulti-step planning with self-check and oracle verificationparseok, oracleequiv_ok
stateful_followupTwo-turn state tracking; turn-2 correct given turn-1turn1/2parseok, turn1/2exactmatch
toolcall_onlyBare schema-only tool call; strict tool name + args checktoolnameok, args_ok
mixed_brief_jsonHybrid: natural language answer + valid JSON blockanswerlineok, jsonparseok, schemaok, semanticpayload_ok
jsonSingle-step structured JSON with constraint rulesschemaok, finalequivok, constraintsok
fuzzProperty-based regression; structured placement correctnessplanexactmatch, finalequivok, constraints_ok
toolcallTool call embedded in response; parse + schema validationstage1toolparseok, selectionok, stage1argsok, toolexecok, stage2evaluable, finalequiv_ok
mcqMultiple-choice extractionsemanticchoiceok

Evaluation infrastructure: Modal A10G GPU cluster (nvidia/cuda:12.4.1) Runner: quantized_modal_llama_cpp (llama.cpp via Python binding, version 0.3.20) Decoding: temperature 0.7, topk=40, topp=0.95 (seed unsupported on Modal) Context: 8,192-token context window (common Modal cap; full 1M available in production) Reproducibility: Run manifest, fixtures, and per-case results published to Zenodo under the quant_eval Public Corpus


Deployment Recommendations

✅ Deploy Q4_K_M when:

  • —Efficiency is the priority. 69.6% footprint reduction; on the same GPU as F16, inference would be faster (matched-hardware speed was not measured in this run).
  • —Hardware is constrained. ~10 GB VRAM fits T4, L4, and consumer GPUs.
  • —Tool dispatch and state tracking matter more than MCQ. Perfect 1.000 stateful, +4.0 pp toolcall.
  • —You can implement MCQ validation. 82% pass rate is acceptable with response reranking or fallback logic.
  • —Cost efficiency drives deployment. 3.29× compression; observed 0.96× speed ratio on a smaller GPU (A10G vs A100-40GB).

⚠️ Avoid Q4_K_M if:

  • —MCQ extraction must be near-perfect. F16 maintains 97% pass rate; Q4KM drops to 82% (−15.0 pp).
  • —You cannot validate or rerank responses. The MCQ loss is real and requires mitigation.
  • —Hardware is plentiful and cost is irrelevant. F16 avoids the MCQ loss (Q4KM scores higher on toolcall, fuzz, and json_multistep).

About PBH Applied Systems

**PBH Applied Systems, LLC** is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization emphasizes engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.

Founder & Principal AI/ML Systems Architect — Patrick Hill, M.S.

Patrick Hill is the Founder & Principal AI/ML Systems Architect of PBH Applied Systems with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning and a B.S. in Business Finance.

Technical expertise spans: Python, SQL, Linux, Pandas, NumPy, scikit-learn, PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA, Flask APIs, Docker, CI/CD, Jupyter, Databricks, and quantization strategies.

Published Author: Patrick is the author of [Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz) — a 1,200+ page practitioner-oriented textbook adopted as required reading for CSC 373 – Machine Learning at the University of Advancing Technology.

Core Service Areas: LLM optimization & deployment · AI evaluation frameworks · Agentic AI infrastructure · Scalable AI application development · ML pipeline design & analytics · Model & agent cataloging.


🔬 About quant_eval & This Evaluation Series

**quant_eval** is a behavioral evaluation harness developed by PBH Applied Systems, LLC. It measures real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning — not perplexity or leaderboard proxies. Every model published under `pbhappliedsystems` has been independently evaluated using quant_eval before being recommended for any production role.

See it in action: **Live AI Agent Demo →**


📋 Interested in Deeper Engagement?

This model card documents behavioral evidence under quant_eval. If you're evaluating this model for production deployment, need guidance on quantization impact for your specific workload, or want to understand deployment tradeoffs under your infrastructure constraints, PBH Applied Systems offers several paths:

Async Intake Form — https://pbhappliedsystems.com/contact.html

Describe your use case, infrastructure, and evaluation needs. Responses are reviewed asynchronously.

Available Services:

  • —Evaluation Report — A written behavioral audit: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, deployment recommendation, and a quantization-level suggestion for your hardware and latency constraints. Evidence-standard numbers only. Typical engagements: $2,500–$5,000.
  • —Starter-Kit Recommender — Configured agent templates (reasoning, document intelligence, code automation) with proof metrics from your specific workload. Includes model selection, quantization strategy, and deployment architecture guidance tied to your infrastructure.
  • —Community — Join discussions on quantization confounds, model comparisons, and evaluation priorities. Active in Hugging Face org discussions, Discord, and Reddit threads.

Submit your details via the form above, and we'll route you to the appropriate engagement.


License

This GGUF repository inherits the license of the base model: Apache 2.0 — `Qwen/Qwen2.5-14B-Instruct-1M`

The quanteval evaluation harness, fuzz prompt builder, and scoring implementation are proprietary to PBH Applied Systems, LLC and are not included in this repository. The golden oracle fixture set used in this evaluation is published under CC BY 4.0 as part of the quanteval Public Corpus (D3, DOI `10.5281/zenodo.22010278`).


GGUF quantization and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.22.22 · Run ID: `Qwen2.5_14B_Instruct_1M_20260815_220633` · Published 2026-08-15