Davd-b01/thinking-cap-tier-curricula-complete
Thinking Cap Tier Curricula — Complete Reasoning Alignment Suite (TCS v4) [!IMPORTANT] Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture: In v1.2, all 13,477 SFT samples and 3,187 SimPO preference pairs have undergone an automated token purge: Zero <|pad|> batch residues: 100% eliminated across all files. Zero reasoning leakage into final answers: Deliberation stays strictly inside <think>...</think>, and answers provide direct, non-repetitive… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-curricula-complete.
docs: include text column at end of schema
feat: keep text column at end while preserving clean RAW column order (prompt, think, answer, tier, domain, ..., text)
docs: harmonize schema documentation to match RAW
style: harmonize SimPO columns (prompt, chosen_think, chosen_answer, rejected_think, rejected_answer)
style: harmonize column order to match RAW (prompt, think, answer, tier, domain)
docs: update dataset card schema with clean thinking and answer columns
feat: clean prompt, chosen_thinking, chosen_answer, and rejected separation
feat: clean top-level thinking and answer columns with balanced tier interleaving
docs: polish README v1.2 with clean LaTeX SimPO formula, exact ChatML/SimPO schemas, cross-model adapter and tier objectives
docs: polish README v1.2 with clean LaTeX SimPO formula, exact ChatML/SimPO schemas, cross-model adapter and tier objectives
docs: polish README v1.2 with clean LaTeX SimPO formula, exact ChatML/SimPO schemas, cross-model adapter and tier objectives
docs: polish README v1.2 with clean LaTeX SimPO formula, exact ChatML/SimPO schemas, cross-model adapter and tier objectives
docs: polish README v1.2 with clean LaTeX SimPO formula, exact ChatML/SimPO schemas, cross-model adapter and tier objectives
docs: add v1.1 clean padding notice to curricula-complete suite
fix(data): sanitize 100% of <|pad|> tokens from curricula v4 SimPO (v1.1)
fix(data): sanitize 100% of <|pad|> tokens from curricula v4 SFT (v1.1)
docs: add cognitive tier objectives (TCS v1.0) and detailed upstream dataset attribution (r0b0tlab, OpenThoughts, OpenMLE, Bespoke-Stratos)
docs: add explicit Qwen ChatML format notice and cross-model conversion guide (Llama 3, Mistral, Gemma 2)
feat: upload curricula complete curricula_stats.json
feat: upload curricula complete qwen_simpo_preference_v4.jsonl
feat: upload curricula complete qwen_sft_curricula_v4.jsonl
docs: add comprehensive dataset card in English
initial commit
