CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jakeatx /qwen36-kquant-offload-mtp-swebench-lite100-results Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer. Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control. The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.imagen<1K0 likes806 downloads4mo agoHugging Face02lightseekorg /kimi-mtp-dataset Kimi-K2.5 Eagle3 Training Data This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec. All responses were regenerated by running Kimi-K2.5 via Engine rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating. The trained Eagle3 draft model is available at… See the full description on the dataset page: https://huggingface.co/datasets/lightseekorg/kimi-mtp-dataset.text100K<n<1M7 likes197 downloads6mo agoHugging Face03angelsbrood /gemma4-mtp-fixturestextn<1K0 likes140 downloads4mo agoHugging Face04windowsxp811203 /nvfp4-mtp-survey Do Qwen3.8-27B NVFP4 repos actually ship a working MTP draft head? A static survey of every NVFP4 quantization of Qwen3.8-27B and its finetunes that I could find on the Hugging Face Hub, last run on 2026-08-24 (Rev 4) with nvfp4_mtp_audit.py. Raw output: results.json. I ran this to check a claim I had made in public, and the claim did not survive. The correction is the first section, because it is the most important result here. Revision history — read this, it is… See the full description on the dataset page: https://huggingface.co/datasets/windowsxp811203/nvfp4-mtp-survey.tabularn<1K1 likes99 downloads1mo agoHugging Face05malaiwah /glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1 fidelity--glm52.malaiwah.quant.exl3-tr3v4-3.5bpw-mtp78-brandonmusic A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm52-fidelity-exl3-tr3v4-3.5bpw-mtp78-brandonmusic-v1.tabularn<1K0 likes72 downloads19d agoHugging Face06arianhosseini /mt_puzzles Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs MT-Puzzles is a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities: Word Guess Guess the secret word in min attempts while environment gives feedback on how close the guess is at each turn. Movie Recommendation: Probe the user to decode the user preference function for N turns. Pick a movie for the… See the full description on the dataset page: https://huggingface.co/datasets/arianhosseini/mt_puzzles.textquestion-answering1K<n<10K0 likes64 downloads1y agoHugging Face07slippedJim /ATOM_regen_seeklight_kimi_mtpgated ATOM regen: seeklight kimi-mtp responses by Kimi-K3 用 Kimi-K3 对 lightseekorg/kimi-mtp-dataset 的 prompt 重新生成了一遍回答,供 off-policy 投机解码蒸馏(SDDD)使用。 原始 pipeline 每轮都要用 teacher 重新解码一次(Phase A1)。把回答预生成并缓存下来, A1 整个消失,之后每一轮训练都直接复用,代价从「每轮一次」变成「一共一次」。 数据 450,625 行,每行一段对话: {"conversations": [ {"role": "user", "content": "..."}, {"role": "assistant", "reasoning_content": "...", "content": "..."} ]} reasoning_content 是 K3 的 thinking 内容,和 content 分开存。 多轮对话保留了历史轮次里完整的 assistant… See the full description on the dataset page: https://huggingface.co/datasets/slippedJim/ATOM_regen_seeklight_kimi_mtp.text-generation100K<n<1M2 likes28 downloads7d agoHugging Face08UnicornChan /kimi-k2.5-mtp-dataset Kimi-K2.5 Eagle3 Training Data This dataset contains the instruction-following data used to train an Eagle3 MTP draft model for Kimi-K2.5 with TorchSpec. All responses were regenerated by running Kimi-K2.5 via SGLang rather than taken from the original datasets. This is critical for speculative decoding training: the draft model must learn the exact token-level distribution of the target model it is accelerating. The trained Eagle3 draft model is available at… See the full description on the dataset page: https://huggingface.co/datasets/UnicornChan/kimi-k2.5-mtp-dataset.text100K<n<1M1 likes20 downloads7mo agoHugging Face09canada-quant /hy3-w4a16-mtp-calibration Hy3 W4A16-MTP — Calibration Set The exact 512-sample calibration blend used to GPTQ-quantize canada-quant/hy3-w4a16-mtp (a W4A16 quantization of tencent/Hy3). Published for full reproducibility of the quantization pipeline. Why a blend (not chat-only) INT4 weight quantization degrades most on code and tool-call-shaped tokens. A chat-only calibration set (e.g. pure ultrachat) under-samples exactly the routed experts those tokens activate. This set deliberately… See the full description on the dataset page: https://huggingface.co/datasets/canada-quant/hy3-w4a16-mtp-calibration.textn<1K1 likes18 downloads2mo agoHugging Face10zhiyuan5986 /MTP-finetune-datatext10K<n<100K0 likes9 downloads1y agoHugging Face11dario-mazzola /mtptext1K<n<10K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.