CoolFace
Modelpublic

sigmanih/Qwen-Qwen3.8-Flash-Next-GGUF-Q8_0

sourceHugging Faceotherupdated 20d agoView on Hugging Face
1likes
Model Card

<div align="center">

⚡ Qwen-Qwen3.8-Flash-Next-GGUF-Q8_0

High-Performance Model Published via [Σ-SIGMA Studio](https://github.com/Sigmanih/SigmaStudio)

![SigmaStudio GitHub](https://github.com/Sigmanih/SigmaStudio) ![HuggingFace Hub](https://huggingface.co/sigmanih/Qwen-Qwen3.8-Flash-Next-GGUF-Q80) [![Engine](https://img.shields.io/badge/Accelerated%20by-SigmaEngine-00d2ff?style=for-the-badge&logo=fastapi)](https://github.com/Sigmanih/SigmaStudio) [![License: Apache-2.0](https://img.shields.io/badge/License-Apache2.0-blue.svg?style=for-the-badge)](https://opensource.org/licenses/Apache-2.0)

</div>

❤️ Support & Community: If you find this model helpful, please give this repository a Like on Hugging Face and a ⭐ Star on our [SigmaStudio GitHub](https://github.com/Sigmanih/SigmaStudio)!

🌐 English Overview

Qwen-Qwen3.8-Flash-Next-GGUF-Q8_0 is a production-ready model optimized and published using the Model Hub module of [Sigma Studio](https://github.com/Sigmanih/SigmaStudio).

⚙️ Technical Specifications & Architecture

SpecificationValue
Model Repositorysigmanih/Qwen-Qwen3.8-Flash-Next-GGUF-Q8_0
Weight FormatGGUF (Q8_0)
Base Architectureqwen4exp
Active Parameters~263.4B
Context Window262,144 tokens
Transformer Layers48
Hidden Dimension2560
Total Disk Footprint158.05 GB
Recommended UsageFlagship frontier intelligence, Deep research & multi-step mathematics.

🏆 Official Benchmark Performance

Evaluated directly on GPU via Sigma Studio Training Lab (Deterministic seed 42, Temp 0.0):

Benchmark SuiteScore / AccuracyTotal Questions EvaluatedPass RateTest DateExecution Engine
Tutti i Benchmark Ufficiali`80.0%`80/100 quesiti superati`80.0% Pass`2026-09-01⚡ SigmaEngine Direct GPU
📋 Per-Dataset Evaluation Breakdown
Dataset / Benchmark SuiteDomain / CategoryCorrect / TotalAccuracy (%)Status
ARC-ChallengeScience & Grade-School Reasoning9 / 9100%✅ Passed
BIG-Bench HardComplex Multi-Task Logic & Symbolics6 / 786%✅ Passed
GPQAGraduate-Level Academic Reasoning5 / 956%⚡ Fair
GSM8KMulti-Step Grade School Math8 / 989%✅ Passed
HellaSwagCommonsense Reasoning & Situational NLI5 / 956%⚡ Fair
HumanEvalPython Coding (pass@1)7 / 7100%✅ Passed
MATHChampionship Competition Math7 / 978%✅ Passed
MBPPPython Programming with Unit Tests8 / 989%✅ Passed
MMLUGeneral Knowledge & Multi-Subject11 / 1479%✅ Passed
MMLU-ProAdvanced Multi-Step Reasoning6 / 967%⚡ Fair
TruthfulQAFactuality & Anti-Hallucination8 / 989%✅ Passed
🏆 OVERALL TOTALAll Evaluated Datasets`80 / 100``80%`🏆 80% Pass

Protocol: codeexecution, continuationlogprob, cotgeneration, letterlogprob · temp 0.0 · seed 42 Reproducibility hash: SHA256-0C4D659A081E98DD

⚠️ Measured on a slice of the dataset, not the full suite: this score is not comparable with a full-suite run.

⚡ Measured Speed on the Publishing Machine

Measured on NVIDIA GeForce RTX 5070 Ti • 15.9 GB VRAM. Two different numbers follow, and they are not interchangeable.

What was measuredValueHow
Single-stream decode (what a chat feels)`11.3 tok/s`one request at a time, on NVIDIA GeForce RTX 5070 Ti
Prompt processing5 tok/ssame probe
Aggregate throughput during evaluation4.1 tok/sseveral requests in flight — not what a single answer runs at
Speeds on other hardware were not measured and are not guessed here. A single-stream figure from one machine cannot be scaled into a prediction for another: it depends on memory bandwidth, quantization, context length and driver, and the error is large enough to be misleading.

🚀 Quick Start Guide

1. Running with Sigma Studio (Recommended)

Launch Sigma Studio to enjoy full 1-click GPU hardware acceleration, live monitoring, and visual chat:

bash
# Clone and run Sigma Studio
git clone https://github.com/Sigmanih/SigmaStudio.git
cd SigmaStudio
.\sigma_studio.bat
2. Running with llama.cpp
bash
llama-cli -hf sigmanih/Qwen-Qwen3.8-Flash-Next-GGUF-Q8_0 -p "Hello! How can I help you today?" -ngl 99

🇮🇹 Documentazione in Italiano

Qwen-Qwen3.8-Flash-Next-GGUF-Q8_0 è un modello ottimizzato pronto per l'inferenza e l'integrazione locale, pubblicato attraverso [Σ-SIGMA Studio](https://github.com/Sigmanih/SigmaStudio).

📋 Specifiche e Configurazione

  • —Architettura Base: qwen4exp (~263.4B parametri)
  • —Formato Pesi: GGUF (Q8_0)
  • —Spazio su Disco: 158.05 GB
  • —Finestra di Contesto: 262,144 token
  • —Profilo d'Uso Consigliato: Intelligenza di frontiera, ricerca approfondita e matematica multi-step.

📊 Risultati Benchmark Ufficiali

  • —Suite di Valutazione: Tutti i Benchmark Ufficiali
  • —Punteggio Ufficiale: `80.0%` (80/100 quesiti superati)
📋 Dettaglio Punteggi per Singolo Dataset
Dataset / Suite di TestAmbito / DominioCorretti / TotaleAccuratezza (%)Esito
ARC-ChallengeRagionamento Scientifico Avanzato9 / 9100%✅ Superato
BIG-Bench HardLogica Complessa & Compiti Multi-Fase6 / 786%✅ Superato
GPQARagionamento Accademico di Livello Laurea5 / 956%⚡ Discreto
GSM8KMatematica & Logica Multi-Step8 / 989%✅ Superato
HellaSwagBuon Senso & Comprensione Situazionale5 / 956%⚡ Discreto
HumanEvalSintesi Codice Python (pass@1)7 / 7100%✅ Superato
MATHMatematica Olimpica & Competitiva7 / 978%✅ Superato
MBPPProgrammazione Python con Unit Test8 / 989%✅ Superato
MMLUConoscenza Generale Multidisciplinare11 / 1479%✅ Superato
MMLU-ProRagionamento Avanzato Multi-Step6 / 967%⚡ Discreto
TruthfulQAFattualità & Resistenza ad Allucinazioni8 / 989%✅ Superato
🏆 TOTALE COMPLESSIVOTutti i Dataset Valutati`80 / 100``80%`🏆 80% Pass

Protocollo: codeexecution, continuationlogprob, cotgeneration, letterlogprob · temp 0.0 · seed 42 Impronta di riproducibilità: SHA256-0C4D659A081E98DD

⚠️ Misurato su una porzione del dataset, non sulla suite intera: il punteggio non è confrontabile con uno ottenuto sull'intero.
  • —Data Test: 2026-09-01 su motore deterministico SigmaEngine

⏱️ Throughput Hardware e Fasce Consigliate

  • —Velocità Verificata in Locale: `4.1 tok/s` su NVIDIA GeForce RTX 5070 Ti.
  • —Risposta singola (quello che si sente in chat): `11.3 tok/s` su NVIDIA GeForce RTX 5070 Ti.
  • —Lettura del prompt: 5 tok/s.
  • —Throughput complessivo durante la valutazione: 4.1 tok/s — piu' richieste in volo insieme, non la velocita' di una risposta singola.
  • —Le velocita' su altro hardware non sono state misurate e non vengono indovinate: dipendono da banda di memoria, quantizzazione, lunghezza del contesto e driver.

⭐ Supporta il Progetto Open Source

Se questo modello ti è utile o vuoi esplorare l'ecosistema completo:

  • —🌟 Metti una Stella al repository GitHub: [Sigmanih/SigmaStudio](https://github.com/Sigmanih/SigmaStudio)
  • —❤️ Lascia un Like a questa scheda su Hugging Face

Creato e distribuito con il Model Hub di Σ-SIGMA Studio (07/09/2026 21:08)