CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01keypa /reaper-calibration REAPER Calibration Dataset Multi-scale calibration dataset for GRPO-based pruning of DeepSeek V4 Pro, generated by the REAPER pipeline. 100 % real data — no synthetic samples at any scale. Configs Config Samples Domain Split Purpose seed-10k 10 000 50 % math, 30 % code, 20 % agentic Quick proto / pipeline validation specialist-300k 300 000 50 % math, 30 % code, 20 % agentic Per-domain specialist training production-800k 800 000 50 % math, 30 % code… See the full description on the dataset page: https://huggingface.co/datasets/keypa/reaper-calibration.texttext-generation100K<n<1M2 likes92 downloads4mo agoHugging Face020xSero /minimax-m2.1-reap-observations [!TIP] Support this work: donate.sybilsolutions.ai REAP surfaces: GLM | MiniMax | Qwen | Gemma | Paper | Code | PR17 | Cerebras Collection MiniMax-M2.1 REAP Stress Test Observations Comprehensive stress test results for MiniMax-M2.1 models pruned with REAP (Router-weighted Expert Activation Pruning) at various compression ratios. Dataset Description This dataset contains 96 stress test results across 4 pruned MiniMax-M2.1 models, testing for repetition loops at… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/minimax-m2.1-reap-observations.tabulartext-generationn<1K1 likes87 downloads25d agoHugging Face03Baekpica /K-EXAONE-236B-REAP-calibration-mix K-EXAONE-236B REAP/NVFP4 Calibration Mix LGAI-EXAONE/K-EXAONE-236B-A23B의 expert pruning(REAP)과 NVFP4 양자화 calibration을 위해 제작한 믹스. 총 16,780 샘플 / 101,157,434 토큰 (K-EXAONE 토크나이저 기준). 제작 목적 MoE 모델을 one-shot pruning/양자화하면 reasoning 무한 반복(한국어/영어 공통)이 발생하는 문제가 있어, 이를 방지하기 위해 아래 원칙으로 설계: Context length 다각화: 16 토큰 ~ 245K 토큰 (짧은 지시 → 32K agentic 궤적 → 128K 장문 → 245K needle 스트레스) 한국어 대량 포함 (instruction/reasoning/tool-calling) — K-EXAONE 특화 expert 보호 reasoning trace 원형 보존 —… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/K-EXAONE-236B-REAP-calibration-mix.texttext-generation10K<n<100K0 likes75 downloads3mo agoHugging Face040xSero /reap-calibration-data-v1 [!TIP] Support this work: donate.sybilsolutions.ai REAP surfaces: GLM | MiniMax | Qwen | Gemma | Paper | Code | PR17 | Cerebras Collection REAP Calibration Dataset v1 Benchmark-free calibration dataset for REAP (Routing-Enhanced Activation Pruning) of Mixture-of-Experts language models. What This Dataset Does REAP prunes MoE models by removing experts that rarely activate. To decide which experts are safe to remove, REAP needs to observe which experts fire on… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/reap-calibration-data-v1.texttext-generation10K<n<100K5 likes62 downloads5mo agoHugging Face05baaderso36 /BaaderSo36-Opus4.7-REAP BaaderSo36-Opus4.7-REAP A synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response. Dataset Statistics Total samples: 1739 Source distribution: debug: 604 react_advanced: 421 math_hard: 260 humaneval: 164 code_contests: 135 math_l5: 125 react: 30 Reasoning depth (characters in thinking block)… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/BaaderSo36-Opus4.7-REAP.tabulartext-generation1K<n<10K1 likes40 downloads5mo agoHugging Face06freddm /reap-agent-code reap-agent-code Dataset Summary reap-agent-code is a REAP-style mixed dataset for training LLM coding agents. It is optimized for agentic coding behavior: writing code, debugging, and tool use. Each row is JSONL with the schema: {"text": "..."} Dataset Composition Source Ratio Count Signal evol 45% 9 000 Instruction-to-code swe 25% 5 000 Bug-fix / problem-solving xlam 30% 6 000Tool / function calling Total: 20 000 unique deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/freddm/reap-agent-code.texttext-generation10K<n<100K0 likes34 downloads8mo agoHugging Face07baaderso36 /BaaderSo36-DE-Opus4.7-REAP BaaderSo36-DE-Opus4.7-REAP A German translation of the BaaderSo36-Opus4.7-REAP reasoning dataset. Each sample preserves the original Claude Opus 4.7 reasoning structure, translated into natural German while keeping format markers (<think>, </think>, Final answer:) and code blocks intact. Dataset Statistics Total samples: 1,379 Source distribution: debug: 511 react_advanced: 357 humaneval: 161 math_hard: 136 code_contests: 125 math_l5: 67 react: 22 Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/BaaderSo36-DE-Opus4.7-REAP.tabulartext-generation1K<n<10K0 likes29 downloads5mo agoHugging Face08baaderso36 /NativeDE-Opus4.7-REAP NativeDE-Opus4.7-REAP A native German synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). All prompts and responses are in natural, idiomatic German — not translations from English. Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response. This dataset is the German-language complement to BaaderSo36-Opus4.7-REAP. Dataset Statistics Total samples: 2,306 Source… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/NativeDE-Opus4.7-REAP.tabulartext-generation1K<n<10K0 likes19 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.