CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01werty1248 /multilingual-instruct-balancedThis repository is a collection of English, Korean, Chinese, and Japanese datasets collected by the HuggingFace Hub and transformed into a unified format. It consists of either native or synthetic data. Some data is not clearly copyrighted or only allows non-commercial use. Preprocessing: I removed data with too few answer tokens or more than 8192 tokens, and removed synthetic data with repetitions. Balancing: I randomly sampled a subset of the data with different weights for each language and… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/multilingual-instruct-balanced.tabulartext-generation1M<n<10M2 likes73 downloads2y agoHugging Face02AmelieSchreiber /toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001 ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 001 This dataset repo records the exact local training-data state visible to the dynamic epoch launcher. It intentionally stores manifests and audit records rather than duplicating large Parquet shards. Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_001_special_structure_current_step_002000.pt Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001.tabulartext-generationn<1K0 likes49 downloads3mo agoHugging Face03AmelieSchreiber /toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002 ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 002 This dataset repo records the exact local training-data state visible to the dynamic epoch launcher. It intentionally stores manifests and audit records rather than duplicating large Parquet shards. Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_002_special_structure_delta_step_002750.pt Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002.tabulartext-generationn<1K0 likes46 downloads3mo agoHugging Face04jrosseruk /Qwen3-4B-MATH-traces-balanced Qwen3-4B MATH Reasoning Traces (Balanced) Reasoning traces from Qwen/Qwen3-4B on MATH problems, balanced for correct/incorrect. Model: Qwen/Qwen3-4B (served via vLLM) Source problems: xDAN2099/lighteval-MATH (train split) Sampling: Subsampled from the full 10k trace set — 2,500 correct + up to 2,500 incorrect Generation params: temperature=0.6, top_p=0.95, max_tokens=15000 Problem types: Algebra, Counting & Probability, Geometry, Intermediate Algebra, Number Theory, Prealgebra… See the full description on the dataset page: https://huggingface.co/datasets/jrosseruk/Qwen3-4B-MATH-traces-balanced.tabulartext-generation1K<n<10K0 likes25 downloads7mo agoHugging Face05ceselder /cot-statement-qa-broad-v2-balanced CoT Statement QA (Deterministic) Conversational supervision dataset for CoT oracles, built from deterministic labels in corpus metadata. The objective is broad prompt phrasing with high-precision answers. Data Sources corpus: data/cot_corpus_v5/corpus_medium.jsonl importance labels: data/importance_resampled_v2.jsonl Size Total rows: 176154 Train: 159320 Validation: 8185 Test: 8649 Task Families correctness_label: 10000 direct_correctness_label:… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/cot-statement-qa-broad-v2-balanced.tabulartext-generation100K<n<1M0 likes20 downloads7mo agoHugging Face06jrosseruk /DeepSeek-R1-Distill-Llama-8B-MATH-traces-balanced DeepSeek-R1-Distill-Llama-8B MATH Reasoning Traces (Balanced) 4,492 reasoning traces from DeepSeek-R1-Distill-Llama-8B on MATH problems, balanced for correct/incorrect. Model: deepseek-ai/DeepSeek-R1-Distill-Llama-8B (served via vLLM) Source problems: xDAN2099/lighteval-MATH (train split) Sampling: Subsampled from the full 10k trace set — 2,500 correct + 1,992 incorrect (all available incorrect traces) Generation params: temperature=0.6, top_p=0.95, max_tokens=15000 Accuracy: 55.7%… See the full description on the dataset page: https://huggingface.co/datasets/jrosseruk/DeepSeek-R1-Distill-Llama-8B-MATH-traces-balanced.tabulartext-generation1K<n<10K0 likes15 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.