CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jakeatx /qwen36-kquant-offload-mtp-swebench-lite100-results Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer. Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control. The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.imagen<1K0 likes811 downloads4mo agoHugging Face02lightseekorg /kimi-mtp-dataset Kimi-K2.5 Eagle3 Training Data This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec. All responses were regenerated by running Kimi-K2.5 via Engine rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating. The trained Eagle3 draft model is available at… See the full description on the dataset page: https://huggingface.co/datasets/lightseekorg/kimi-mtp-dataset.text100K<n<1M7 likes195 downloads6mo agoHugging Face03angelsbrood /gemma4-mtp-fixturestextn<1K0 likes140 downloads4mo agoHugging Face04windowsxp811203 /nvfp4-mtp-survey Do Qwen3.8-27B NVFP4 repos actually ship a working MTP draft head? A static survey of every NVFP4 quantization of Qwen3.8-27B and its finetunes that I could find on the Hugging Face Hub, last run on 2026-08-24 (Rev 4) with nvfp4_mtp_audit.py. Raw output: results.json. I ran this to check a claim I had made in public, and the claim did not survive. The correction is the first section, because it is the most important result here. Revision history — read this, it is… See the full description on the dataset page: https://huggingface.co/datasets/windowsxp811203/nvfp4-mtp-survey.tabularn<1K1 likes105 downloads1mo agoHugging Face05malaiwah /glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1 fidelity--glm52.malaiwah.quant.exl3-tr3v4-3.5bpw-mtp78-brandonmusic A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1.tabularn<1K0 likes71 downloads18d agoHugging Face06arianhosseini /mt_puzzles Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs MT-Puzzles is a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities: Word Guess Guess the secret word in min attempts while environment gives feedback on how close the guess is at each turn. Movie Recommendation: Probe the user to decode the user preference function for N turns. Pick a movie for the… See the full description on the dataset page: https://huggingface.co/datasets/arianhosseini/mt_puzzles.textquestion-answering1K<n<10K0 likes62 downloads1y agoHugging Face07UnicornChan /kimi-k2.5-mtp-dataset Kimi-K2.5 Eagle3 Training Data This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec. All responses were regenerated by running Kimi-K2.5 via SGLang rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating. The trained Eagle3 draft model is available at… See the full description on the dataset page: https://huggingface.co/datasets/UnicornChan/kimi-k2.5-mtp-dataset.text100K<n<1M1 likes19 downloads7mo agoHugging Face08canada-quant /hy3-w4a16-mtp-calibration Hy3 W4A16-MTP — Calibration Set The exact 512-sample calibration blend used to GPTQ-quantize canada-quant/hy3-w4a16-mtp (a W4A16 quantization of tencent/Hy3). Published for full reproducibility of the quantization pipeline. Why a blend (not chat-only) INT4 weight quantization degrades most on code and tool-call-shaped tokens. A chat-only calibration set (e.g. pure ultrachat) under-samples exactly the routed experts those tokens activate. This set deliberately… See the full description on the dataset page: https://huggingface.co/datasets/canada-quant/hy3-w4a16-mtp-calibration.textn<1K1 likes18 downloads2mo agoHugging Face09zhiyuan5986 /MTP-finetune-datatext10K<n<100K0 likes8 downloads1y agoHugging Face10dario-mazzola /mtptext1K<n<10K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.