datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-instruct-balancedThis repository is a collection of English, Korean, Chinese, and Japanese datasets collected by the HuggingFace Hub and transformed into a unified format. It consists of either native or synthetic data.
Some data is not clearly copyrighted or only allows non-commercial use.
Preprocessing: I removed data with too few answer tokens or more than 8192 tokens, and removed synthetic data with repetitions.
Balancing: I randomly sampled a subset of the data with different weights for each language and… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/multilingual-instruct-balanced.toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 001
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_001_special_structure_current_step_002000.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-001.toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002
ToricBLM dataset state: toricblm-structure-priority-balanced-3day-20260709T185034Z epoch 002
This dataset repo records the exact local training-data state visible to the dynamic epoch launcher.
It intentionally stores manifests and audit records rather than duplicating large Parquet shards.
Special checkpoint: toricblm-structure-priority-balanced-3day-20260709T185034Z_epoch_002_special_structure_delta_step_002750.pt
Checkpoint repo: AmelieSchreiber/ToricGT_160M_FoT
Curriculum… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricblm-dataset-state-toricblm-structure-priority-balanced-3day-20260709t185034z-epoch-002.Qwen3-4B-MATH-traces-balanced
Qwen3-4B MATH Reasoning Traces (Balanced)
Reasoning traces from Qwen/Qwen3-4B on MATH problems, balanced for correct/incorrect.
Model: Qwen/Qwen3-4B (served via vLLM)
Source problems: xDAN2099/lighteval-MATH (train split)
Sampling: Subsampled from the full 10k trace set — 2,500 correct + up to 2,500 incorrect
Generation params: temperature=0.6, top_p=0.95, max_tokens=15000
Problem types: Algebra, Counting & Probability, Geometry, Intermediate Algebra, Number Theory, Prealgebra… See the full description on the dataset page: https://huggingface.co/datasets/jrosseruk/Qwen3-4B-MATH-traces-balanced.cot-statement-qa-broad-v2-balanced
CoT Statement QA (Deterministic)
Conversational supervision dataset for CoT oracles, built from deterministic labels in corpus metadata.
The objective is broad prompt phrasing with high-precision answers.
Data Sources
corpus: data/cot_corpus_v5/corpus_medium.jsonl
importance labels: data/importance_resampled_v2.jsonl
Size
Total rows: 176154
Train: 159320
Validation: 8185
Test: 8649
Task Families
correctness_label: 10000
direct_correctness_label:… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/cot-statement-qa-broad-v2-balanced.DeepSeek-R1-Distill-Llama-8B-MATH-traces-balanced
DeepSeek-R1-Distill-Llama-8B MATH Reasoning Traces (Balanced)
4,492 reasoning traces from DeepSeek-R1-Distill-Llama-8B on MATH problems, balanced for correct/incorrect.
Model: deepseek-ai/DeepSeek-R1-Distill-Llama-8B (served via vLLM)
Source problems: xDAN2099/lighteval-MATH (train split)
Sampling: Subsampled from the full 10k trace set — 2,500 correct + 1,992 incorrect (all available incorrect traces)
Generation params: temperature=0.6, top_p=0.95, max_tokens=15000
Accuracy: 55.7%… See the full description on the dataset page: https://huggingface.co/datasets/jrosseruk/DeepSeek-R1-Distill-Llama-8B-MATH-traces-balanced.
