CoolFace
Modelpublic

sigmanih/Qwen-Qwen2.5-Coder-14B-Instruct-GGUF-Q3_K_M

sourceHugging Faceotherupdated 14d agoView on Hugging Face
1likes529downloads
Model Card

<div align="center">

⚡ Qwen-Qwen2.5-Coder-14B-Instruct-GGUF-Q3KM

High-Performance Model Published via [Σ-SIGMA Studio](https://github.com/Sigmanih/SigmaStudio)

![SigmaStudio GitHub](https://github.com/Sigmanih/SigmaStudio) ![HuggingFace Hub](https://huggingface.co/sigmanih/Qwen-Qwen2.5-Coder-14B-Instruct-GGUF-Q3KM) ![Engine](https://github.com/Sigmanih/SigmaStudio) ![License: Apache-2.0](https://opensource.org/licenses/Apache-2.0)

</div>

❤️ Support & Community: If you find this model helpful, please give this repository a Like on Hugging Face and a ⭐ Star on our [SigmaStudio GitHub](https://github.com/Sigmanih/SigmaStudio)!

🌐 English Overview

Qwen-Qwen2.5-Coder-14B-Instruct-GGUF-Q3_K_M is a production-ready model optimized and published using the Model Hub module of [Sigma Studio](https://github.com/Sigmanih/SigmaStudio).

⚙️ Technical Specifications & Architecture

SpecificationValue
Model Repositorysigmanih/Qwen-Qwen2.5-Coder-14B-Instruct-GGUF-Q3_K_M
Weight FormatGGUF (Q3_K_M)
Base Architectureqwen2
Active Parameters14B
Context Window32,768 tokens
Transformer Layers48
Hidden Dimension5120
Total Disk Footprint6.84 GB
Inference RAM / VRAM`~8.9 GB VRAM` (Full GPU offload) / `~8.7 GB RAM` (CPU/Hybrid)
Recommended HardwareGPU with 12-16 GB VRAM (e.g. RTX 3060/4070 or 32 GB RAM)
Recommended UsageHigh-speed coding, Everyday assistants, Autonomous agentic loops & reasoning.

🏆 Official Benchmark Performance

Evaluated directly on GPU via Sigma Studio Training Lab (Deterministic seed 42, Temp 0.0):

Benchmark SuiteScore / AccuracyTotal Questions EvaluatedPass RateTest DateExecution Engine
Tutti i Benchmark Ufficiali`69.0%`69/100 quesiti superati`69.0% Pass`2026-09-07⚡ SigmaEngine Direct GPU
🧠 Reasoning Mode Comparison: No-Thinking vs Thinking

Side-by-side performance comparison between direct zero-overhead answer (No-Thinking) and Chain-of-Thought step-by-step reasoning (Deep Thinking CoT):

text
⚡ No-Thinking (Direct Response) : [██████████████░░░░░░] 71.9%  (41/57 passed)
🧠 Deep Thinking (CoT Reasoning) : [█████████████░░░░░░░] 65.1%  (28/43 passed)
📈 CoT Performance Delta        : -6.8%
mermaid
%%{init: {'theme': 'dark'}}%%
xychart-beta
    title "Accuracy (%): No-Thinking vs Deep Thinking"
    x-axis ["⚡ No-Thinking (Direct)", "🧠 Deep Thinking (CoT)"]
    y-axis "Accuracy (%)" 0 --> 100
    bar [71.9, 65.1]
Execution ModeAccuracy (%)Questions PassedOperational Profile & Latency
⚡ No-Thinking (Direct Response)`71.9%`41 / 57Minimal latency, immediate token-to-first-byte, zero reasoning tokens overhead
🧠 Deep Thinking (CoT Reasoning)`65.1%`28 / 43Multi-step structured reasoning trace (-6.8%), optimal for math & hard logic
📋 Per-Dataset Evaluation Breakdown
Dataset / Benchmark SuiteDomain / CategoryCorrect / TotalAccuracy (%)Status
ARC-ChallengeScience & Grade-School Reasoning8 / 989%✅ Passed
BIG-Bench HardComplex Multi-Task Logic & Symbolics5 / 771%✅ Passed
GPQAGraduate-Level Academic Reasoning3 / 933%⚠️ Low
GSM8KMulti-Step Grade School Math9 / 9100%✅ Passed
HellaSwagCommonsense Reasoning & Situational NLI5 / 956%⚡ Fair
HumanEvalPython Coding (pass@1)6 / 786%✅ Passed
MATHChampionship Competition Math8 / 989%✅ Passed
MBPPPython Programming with Unit Tests9 / 9100%✅ Passed
MMLUGeneral Knowledge & Multi-Subject6 / 1443%⚡ Fair
MMLU-ProAdvanced Multi-Step Reasoning3 / 933%⚠️ Low
TruthfulQAFactuality & Anti-Hallucination7 / 978%✅ Passed
🏆 OVERALL TOTALAll Evaluated Datasets`69 / 100``69%`🏆 69% Pass

Protocol: codeexecution, continuationlogprob, cotgeneration, letterlogprob · temp 0.0 · seed 42

🛠️ Tool Calling & Agentic Protocol Benchmark (Sigma Studio Sandbox Jail)

Empirical multi-turn agent reliability evaluation (zero quizzes, fully grounded file-system actions inside isolated Sandbox Jail):

Protocol MetricOutcomeValidation Criteria & Details
Tool Protocol Adherence`63.6%`35/55 analytical criteria verified
Autonomous Task Completion`66.7%`4/6 scenarios completed with exact target file
Sandbox Jail Containment`100% Compliant`Zero escape attempts outside workspace boundary
Execution Efficiency`50 turns (1453.7s)`Optimal multi-step turn and token budget usage
📋 Per-Scenario Protocol Breakdown
Scenario NameDifficulty TierScoreTurns UsedGoal Status
File NuovoLivello 1 (Base)64%1/14⚠️ Partial
Ispezione ChiaveLivello 1 (Base)64%8/14✅ Passed
Modifica MirataLivello 2 (Intermedio)64%10/14✅ Passed
Ricerca AlberoLivello 2 (Intermedio)64%14/14⚠️ Partial
Debug E FixLivello 3 (Avanzato)64%11/14✅ Passed
Refactor Due FileLivello 3 (Avanzato)64%10/14✅ Passed
🔍 Certified Autonomous Capabilities:
  • —✅ Tool Grounding: Exclusively uses registered tools and valid schemas (zero hallucinated functions).
  • —✅ Zero Placeholder Echo: Emits concrete code and values rather than copy-pasting prompt templates.
  • —✅ Inspect-Before-Edit: Systematically reads files and verifies target lines before patching.
  • —✅ Evidence-Based Exit: Emits concrete test commands and validation checks before task exit.
  • —✅ Sandbox Containment: Strictly adheres to isolated sandbox jail boundaries.

⚡ Measured Speed on the Publishing Machine

Measured on NVIDIA GeForce RTX 5070 Ti • 15.9 GB VRAM during the evaluation run.

What was measuredValueHow
Aggregate throughput during evaluation51.5 tok/sseveral requests in flight — not what a single answer runs at
Speeds on other hardware were not measured and are not guessed here. A single-stream figure from one machine cannot be scaled into a prediction for another: it depends on memory bandwidth, quantization, context length and driver, and the error is large enough to be misleading.

🚀 Quick Start Guide

1. Running with Sigma Studio (Recommended)

Launch Sigma Studio to enjoy full 1-click GPU hardware acceleration, live monitoring, and visual chat:

bash
# Clone and run Sigma Studio
git clone https://github.com/Sigmanih/SigmaStudio.git
cd SigmaStudio
.\sigma_studio.bat
2. Running with llama.cpp
bash
llama-cli -hf sigmanih/Qwen-Qwen2.5-Coder-14B-Instruct-GGUF-Q3_K_M -p "Hello! How can I help you today?" -ngl 99

🇮🇹 Documentazione in Italiano

Qwen-Qwen2.5-Coder-14B-Instruct-GGUF-Q3_K_M è un modello ottimizzato pronto per l'inferenza e l'integrazione locale, pubblicato attraverso [Σ-SIGMA Studio](https://github.com/Sigmanih/SigmaStudio).

📋 Specifiche e Configurazione

  • —Architettura Base: qwen2 (14B parametri)
  • —Formato Pesi: GGUF (Q3KM)
  • —Spazio su Disco: 6.84 GB
  • —RAM / VRAM in Esecuzione: ~8.9 GB VRAM (offload GPU completo) | ~8.7 GB RAM (inferenza CPU/ibrida)
  • —Requisiti Hardware Consigliati: GPU con 12-16 GB VRAM (es. RTX 3060/4070 o 32 GB RAM)
  • —Finestra di Contesto: 32,768 token
  • —Profilo d'Uso Consigliato: Coding ad alta velocità, assistenti quotidiani, loop di agenti autonomi e ragionamento.

📊 Risultati Benchmark Ufficiali

  • —Suite di Valutazione: Tutti i Benchmark Ufficiali
  • —Punteggio Ufficiale: `69.0%` (69/100 quesiti superati)
🧠 Confronto Modalità di Risposta: No-Thinking vs Thinking

Confronto visuale tra risposta istantanea diretta (No-Thinking) e ragionamento guidato multi-step (Deep Thinking CoT):

text
⚡ No-Thinking (Risposta Diretta) : [██████████████░░░░░░] 71.9%  (41/57 superati)
🧠 Deep Thinking (CoT Reasoning)  : [█████████████░░░░░░░] 65.1%  (28/43 superati)
📈 Delta Prestazionale CoT        : -6.8%
mermaid
%%{init: {'theme': 'dark'}}%%
xychart-beta
    title "Accuratezza (%): No-Thinking vs Deep Thinking"
    x-axis ["⚡ No-Thinking (Diretto)", "🧠 Deep Thinking (CoT)"]
    y-axis "Accuratezza (%)" 0 --> 100
    bar [71.9, 65.1]
Modalità di EsecuzioneAccuratezza (%)Quesiti SuperatiProfilo Operativo & Latenza
⚡ No-Thinking (Risposta Diretta)`71.9%`41 / 57Latenza minima, token-to-first-byte istantaneo, zero overhead di ragionamento
🧠 Deep Thinking (CoT Reasoning)`65.1%`28 / 43Risoluzione passo-passo multi-step (-6.8%), ideale per logica complessa e matematica
📋 Dettaglio Punteggi per Singolo Dataset
Dataset / Suite di TestAmbito / DominioCorretti / TotaleAccuratezza (%)Esito
ARC-ChallengeRagionamento Scientifico Avanzato8 / 989%✅ Superato
BIG-Bench HardLogica Complessa & Compiti Multi-Fase5 / 771%✅ Superato
GPQARagionamento Accademico di Livello Laurea3 / 933%⚠️ Migliorabile
GSM8KMatematica & Logica Multi-Step9 / 9100%✅ Superato
HellaSwagBuon Senso & Comprensione Situazionale5 / 956%⚡ Discreto
HumanEvalSintesi Codice Python (pass@1)6 / 786%✅ Superato
MATHMatematica Olimpica & Competitiva8 / 989%✅ Superato
MBPPProgrammazione Python con Unit Test9 / 9100%✅ Superato
MMLUConoscenza Generale Multidisciplinare6 / 1443%⚡ Discreto
MMLU-ProRagionamento Avanzato Multi-Step3 / 933%⚠️ Migliorabile
TruthfulQAFattualità & Resistenza ad Allucinazioni7 / 978%✅ Superato
🏆 TOTALE COMPLESSIVOTutti i Dataset Valutati`69 / 100``69%`🏆 69% Pass

Protocollo: codeexecution, continuationlogprob, cotgeneration, letterlogprob · temp 0.0 · seed 42

  • —Data Test: 2026-09-07 su motore deterministico SigmaEngine

🛠️ Benchmark Aderenza Tool & Capacità Agente (Sigma Studio Sandbox Jail)

Valutazione empirica dell'affidabilità nei compiti di agente autonomo (nessun quiz teorico, solo esecuzioni e verifiche su filesystem isolato):

Metrica di ProtocolloRisultatoDettaglio e Criteri di Validazione
Aderenza Protocollo Tool`63.6%`35/55 prove analitiche verificate
Completamento Scenari Operativi`66.7%`4/6 scenari conclusi con output esatto
Sicurezza Sandbox Jail`100% Conforme`Confinamento rigido, zero tentativi fuori sandbox
Efficienza Esecutiva`50 turni (1453.7s)`Rispetto rigoroso dei budget operativi per scenario
📋 Dettaglio Prove per Scenario Operativo
Scenario di ProvaDifficoltàPunteggioTurniObiettivo Verificato
File NuovoLivello 1 (Base)64%1/14⚠️ Parziale
Ispezione ChiaveLivello 1 (Base)64%8/14✅ Raggiunto
Modifica MirataLivello 2 (Intermedio)64%10/14✅ Raggiunto
Ricerca AlberoLivello 2 (Intermedio)64%14/14⚠️ Parziale
Debug E FixLivello 3 (Avanzato)64%11/14✅ Raggiunto
Refactor Due FileLivello 3 (Avanzato)64%10/14✅ Raggiunto
🔍 Comportamenti Rigorosamente Certificati:
  • —✅ Tool Grounding: Utilizzo esclusivo di tool formalmente registrati (zero allucinazioni di comandi).
  • —✅ No Segnaposto: Produzione di codice e parametri concreti senza eco di placeholder d'esempio.
  • —✅ Ispezione Previa: Lettura e verifica dei file prima di eseguire modifiche chirurgiche.
  • —✅ Verifica di Chiusura: Certificazione delle prove prima di dichiarare terminato il lavoro.
  • —✅ Confinamento Jail: Isolamento totale senza contaminazione del kernel o del sistema host.

⏱️ Throughput Hardware e Fasce Consigliate

  • —Velocità Verificata in Locale: `51.5 tok/s` su NVIDIA GeForce RTX 5070 Ti.
  • —Throughput complessivo durante la valutazione: 51.5 tok/s — piu' richieste in volo insieme, non la velocita' di una risposta singola.
  • —Le velocita' su altro hardware non sono state misurate e non vengono indovinate: dipendono da banda di memoria, quantizzazione, lunghezza del contesto e driver.

⭐ Supporta il Progetto Open Source

Se questo modello ti è utile o vuoi esplorare l'ecosistema completo:

  • —🌟 Metti una Stella al repository GitHub: [Sigmanih/SigmaStudio](https://github.com/Sigmanih/SigmaStudio)
  • —❤️ Lascia un Like a questa scheda su Hugging Face

Creato e distribuito con il Model Hub di Σ-SIGMA Studio (13/09/2026 17:45)