datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.OMB-Circular-A11-Section-120-Apportionment-Process
Dataset Description
The OMB Circular A-11 Section 120 Apportionment Process Question Answering Dataset is a document-grounded collection of 150 question-and-answer records concerning the federal apportionment process administered by the Office of Management and Budget.
The dataset was developed from Section 120, “Apportionment Process,” of OMB Circular No. A-11, Preparation, Submission, and Execution of the Budget. Section 120 is part of the Circular’s budget-execution… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OMB-Circular-A11-Section-120-Apportionment-Process.QA-RESPONSES-A1distiset-ascii-art-a1
Dataset Card for distiset-ascii-art-a1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/DominguesAddem1974/distiset-ascii-art-a1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/DominguesAddem1974/distiset-ascii-art-a1.my-distiset-4d3904d1
Dataset Card for my-distiset-4d3904d1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/A1berto0/my-distiset-4d3904d1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/A1berto0/my-distiset-4d3904d1.
