CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Baekpica /Inkling-Small-Multimodal-Calibration Inkling-Small Multimodal Calibration The exact 1,663 samples used for BF16 routed-expert importance collection for Inkling-Small Mixed Quant GGUF. This is calibration material, not a held-out evaluation benchmark. The primary balanced pass is: Category Samples Valid decoder tokens Share Text / reasoning 462 471,858 44.976% Code / tool-oriented source text 205 209,715 19.989% Real image / document 486 262,476 25.018% Real speech audio 309 105,080 10.016% Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.tabulartext-generation1K<n<10K0 likes413 downloads15d agoHugging Face02RESMP-DEV /ptq-calibration-corpus PTQ Calibration Corpus A compact, high-density corpus built specifically for post-training quantization of chat, coding, and agentic language models. This is not random web text and it is not a repackaged benchmark. The corpus was deliberately assembled to exercise varied scientific and technical prose, numeric structure, multilingual code, tool schemas, long debugging trajectories, CUDA/Triton vocabulary, correctness recovery, and architecture-sensitive optimization. This is… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/ptq-calibration-corpus.tabulartext-generation1K<n<10K0 likes126 downloads1mo agoHugging Face03HINT-lab /Qwen2.5-7B-Instruct-Self-Calibration Efficient Test-Time Scaling via Self-Calibration This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.tabulartext-generation100K<n<1M0 likes121 downloads2y agoHugging Face04jayzou3773 /less-is-moe-s1-calibration-128-seq8192 Less-is-MoE S1K calibration data — 128 samples, seq_length 8192 This is the fixed calibration artifact used to prune GPT-OSS-120B, Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant: yentinglin/s1K-1.1-trl-format revision 58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows. For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.tabulartext-generationn<1K0 likes89 downloads23h agoHugging Face05jayzou3773 /less-is-moe-s1-calibration-128 Less-is-MoE S1K calibration data — 128 full-length samples This repository contains the exact 128 S1K rows selected for Less-is-MoE full-model pruning. The selection reproduces the released loader: source: yentinglin/s1K-1.1-trl-format revision: 58a01564d278477da20ead1bcf1cde8e31f36251 split: train order: Dataset.shuffle(seed=1234) samples: first 128 nonempty messages rows sequence-length limit: none truncation: disabled padding: disabled calibration.jsonl stores every… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128.tabulartext-generationn<1K0 likes57 downloads6d agoHugging Face06Lambent /qwen3.5-moe-awq-calibration Qwen3.5 MoE AWQ Calibration Dataset Calibration dataset for AWQ (Activation-Aware Weight Quantization) of Qwen/Qwen3.5-35B-A3B and Qwen/Qwen3.5-35B-A3B-Base. Designed for MoE expert routing diversity: Qwen3.5-35B-A3B has 256 experts with 8 active per token, so calibration data needs broad domain coverage to exercise as many routing paths as possible. Sampling methodology Source: PleIAs/common_corpus (open multi-domain corpus with labeled collections) Filtering: Token… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/qwen3.5-moe-awq-calibration.tabulartext-generationn<1K0 likes34 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.