datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.quadmix-stem-v2
QuaDMix-STEM v2: STEM-Focused Proxy Validation Set with GPQA & MATH
Script: scripts/validation_set/prepare_stem_v2.py
HuggingFace: liujin99/quadmix-stem-v2
Files: stem_v2_tokenized.pt, stem_v2.parquet
Overview
STEM v2 is an upgraded validation set that fixes the two critical coverage gaps in STEM v1. In the v1 experiment, QuaDMix lost to Random downstream (CORE 0.1530 vs 0.1615), and root-cause analysis revealed:
gpqa_diamond had no direct proxy — mapped from… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v2.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.quadmix-stem-v1
QuaDMix-STEM v1: STEM-Focused Proxy Validation Set
Script: scripts/validation_set/prepare_stem_v1.py
HuggingFace: liujin99/quadmix-stem-v1
Files: stem_v1_tokenized.pt, stem_v1.parquet
Overview
STEM v1 is a validation set designed to focus the proxy model's optimization signal on STEM capabilities — mathematics, science knowledge, and logical reasoning. Unlike CAP v1 (broad capability coverage) or core_bmk (benchmark test format), STEM v1 uses only tasks that… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v1.ru-stem-dialogues
Russian STEM Educational Dialogues
Описание
Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия,
биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие
между пользователем (школьник / студент / профессионал) и ассистентом.
Методология
Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes)
Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.Kannada-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Kannada STEM textbook data, containing 127 books and 6.64 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Kannada.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada-STEM-Textbook-Dataset.Turkish-STEM-DPO-Dataset
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti yusufbaykaloglu tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: yusufbaykaloglu/Turkish-STEM-DPO-Dataset
🔗 Derleyen Platform: VeriPazarı
Türkçe STEM DPO Veri Seti (Turkish STEM DPO Dataset)
Veri Seti Özeti
Turkish STEM DPO (Doğrudan Tercih… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/Turkish-STEM-DPO-Dataset.Sanskrit-Non-STEM-Textbook-DatasetDataset Description:
This dataset is a large-scale collection of Sanskrit Non-STEM textbook data, containing 147 books and 8.29 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and classical knowledge learning in Sanskrit.
Full Dataset Overview
This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Sanskrit-Non-STEM-Textbook-Dataset.Turkish-STEM-DPO-Dataset
Turkish STEM DPO Dataset
Dataset Summary
The Turkish STEM DPO (Direct Preference Optimization) dataset is a comprehensive synthetic resource containing 16,177 high-quality preference pairs designed to enhance the reasoning capabilities of Turkish language models in mathematics, physics, and programming.
The dataset leverages a preference-based learning approach: each instance pairs a carefully crafted, expert-level solution with a deliberately flawed or incomplete… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/Turkish-STEM-DPO-Dataset.
