datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thinking-cap-tier-lima-dense
Thinking Cap Tier Curricula — LIMA Hyper-Dense Reasoning Alignment Suite (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
In v1.2, all 5,500 SFT and 2,000 SimPO records have undergone a complete token purge:
Zero <|pad|> batch residues: 100% eliminated across all records.
Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct conclusions.
Native ChatML… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-lima-dense.dense-reasoning-coding-1k
Dense-Reasoning-Coding-1K
Dataset Description
This dataset is an optimized, highly dense Supervised Fine-Tuning (SFT) subset designed to teach smaller language models (e.g., 1B to 8B architectures) how to reason about complex coding problems without overwhelming their context windows.
It is derived from the verified_90k split of IIGroup/X-Coder-SFT-376k, which features advanced programming tasks and solutions.
About the Creator & Origin
This… See the full description on the dataset page: https://huggingface.co/datasets/Phips/dense-reasoning-coding-1k.Dense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.opc-sft-stage2-dense-extracted
OpenCoder Dataset Dense Region Extracted
This dataset is a post-processed version of the OpenCoder SFT Stage2 dataset (opc-sft-stage2).
We use gpt-4o API to extract the information dense regions from each sample and logged them in the dense_snippets column.Detailed information about the data can be found in our paper.
OpenCoder's sft-stage2 summary
The original version of this dataset is used in OpenCoder's Stage 2 and consists of four parts:
educational_instruct:… See the full description on the dataset page: https://huggingface.co/datasets/malr07/opc-sft-stage2-dense-extracted.glublm-60k-ted
GlubLM 60K Ted Dataset
A 60,837-sample dataset of single-turn conversations in the persona of a goldfish with a 10-second memory. Used to train GlubLM (36M).
Generation method
The entire dataset was generated using Claude via the claude -p CLI subprocess (Claude Max subscription, zero API costs):
Agent
Role
generator
Generates 50-sample batches per call
critic
Reviews each sample, rejects off-persona
diversifier
Audits vocabulary every 1K samples… See the full description on the dataset page: https://huggingface.co/datasets/DenSec02/glublm-60k-ted.densemixer-ab-qwen3-30b-a3b-thinking-opencode-serveparity-idEval
DenseMixer A/B — serve-parity ID-eval traces + weight-delta/routing analysis
Full artifacts for the controlled paired-init A/B ablation testing whether DenseMixer
(training-only dense-forward + STE counterfactual router gradient; yaof20/DenseMixer,
Axolotl integrations/densemixer/) improves MoE SFT quality — the empirical answer to
marin-community/marin#7088.
Setup (identical except ONE flag)
Base / θ₀: Qwen/Qwen3-30B-A3B-Thinking-2507 @ 144afc2f (shared init).… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/densemixer-ab-qwen3-30b-a3b-thinking-opencode-serveparity-idEval.
