datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.ua-open-data
Україна: дзеркало відкритих даних (data.gov.ua)
Автоматичне дзеркало публічних наборів data.gov.ua,
яке підтримує пайплайн JoTalbot/ukraine.
Набори
Набір
Файлів
Джерело
Єдиний державний реєстр юридичних осіб, фізичних осіб-підприємців та громадських формувань
6
—
Реєстр декларацій родинних зв’язків та доброчесності
14
—
Державний судновий реєстр України
9
—
Публічні закупівлі на сайті Prozorro
1
—
Інформація щодо стану розгляду справ
5
—… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-open-data.SlimPajama-Meta-rater
Annotated SlimPajama Dataset
Dataset Description
This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions.
Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.european-open-data-catalogue
European Open Data Catalogue
This repository publishes independently versioned metadata and licensed source snapshots:
A discovery catalogue with 15565 dataset entries from
ISTAT, Eurostat, OECD, ILO, DoveVannoINostriSoldi (DVNS) and Cruscotto Italia.
3 independently pinned availability indexes with
911,795 joint combinations across 35 datasets, built from complete
source responses within the explicitly declared scope.
Licensed Cruscotto source snapshots, stored separately from… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.Logics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.SlimPajama-Meta-rater-Readability-30B
Top 30B token SlimPajama Subset selected by the Readability rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Readability dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Readability-30B.opendata-bodypose
SkillCorner Open Data — Body Pose
3D body-pose data derived from broadcast video, released alongside the
SkillCorner Open Data repository as a
joint initiative between SkillCorner and
PySport.
Initial testing release. Two matches, published so the community can work
with the format and tell us what is useful before we consider a wider release.
Feedback is genuinely wanted — open an issue on the
opendata repo or reply in the
Community tab here.
What is in here… See the full description on the dataset page: https://huggingface.co/datasets/SkillCorner/opendata-bodypose.Logics-STEM-SFT-Dataset-Open-5.3MCiteVQA
CiteVQA
English | 简体中文
CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs.
The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.SlimPajama-Meta-rater-Professionalism-30B
Top 30B token SlimPajama Subset selected by the Professionalism rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.SlimPajama-Meta-rater-Reasoning-30B
Top 30B token SlimPajama Subset selected by the Reasoning rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Reasoning dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework. Each… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Reasoning-30B.Meta-rater-PRRC-Rater-dataset
PRRC Rater Training and Evaluation Dataset
Dataset Description
This dataset contains the full training and evaluation data for the PRRC rater models described in Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models. It is designed for training and benchmarking models that score text along four key quality dimensions: Professionalism, Readability, Reasoning, and Cleanliness.
Source: Subset of SlimPajama-627B, annotated for PRRC dimensions… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Meta-rater-PRRC-Rater-dataset.SlimPajama-Meta-rater-Cleanliness-30B
Top 30B token SlimPajama Subset selected by the Cleanliness rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Cleanliness dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness) framework.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Cleanliness-30B.open-christian-data
Open Christian Data
Open Christian Data (OCD) aims to be the single unified collection of all public domain Christian text. It exists to bring this collection together from across the internet and to structure it in a useful format for public use.
Beyond the Bible, Christian writing is poorly represented as a cohesive dataset or as data structured for AI training. This Hugging Face release is the AI-focused publication of the collection: consistent, downloadable JSON for model… See the full description on the dataset page: https://huggingface.co/datasets/OpenChristianDataOrg/open-christian-data.ru-open-llama-training-datasetsCOT-Dataset-Mathopen_lm_test_dataK12textbook覆盖小学、初中、高中的高质量中文K12教材语料,经过精细的文本抽取和数据处理,可用于学术研究
miscon-data
Open Misconceptions
A public catalogue of misconceptions with stable IDs. Each row is one
record: a belief a learner could hold, its kind, the evidence pattern that
reveals it (with a concrete example), discriminators against slips and
neighbouring misconceptions, alignments to external schemes, and
provenance.
This dataset mirrors dist/miscon.jsonl from the tagged release of
https://github.com/open-misconceptions/miscon-data. The canonical form of a record is
its stable URI… See the full description on the dataset page: https://huggingface.co/datasets/open-misconceptions/miscon-data.morocco-cassation-court-decisions
Morocco Cassation Court Decisions
29,000+ full-text decisions from the Moroccan Court of Cassation (محكمة النقض)Source: juriscassation.cspj.ma — Official portal of the Supreme Council of the Judiciary (CSPJ)License: CC BY 4.0
Why this dataset exists
In 2026, accessing the jurisprudence of the Court of Cassation in Morocco requires being physically located in Morocco and armed with patience. The official website does not allow searching by date range, imposes a… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataMoroccanLaw/morocco-cassation-court-decisions.databricks__dolly-v2-7b-details
Dataset Card for Evaluation run of databricks/dolly-v2-7b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.Open-ert-small-datasetThis is a subset of:
https://huggingface.co/datasets/openerotica/long-roleplay-v0.1
I am using mistral's new DEVSTRAL model to take the entire conversation in JSON format and rate it. I chose DEVSTRAL due to the mistral models being very consistent and well rounded. The Devstral model I was hoping could understand the JSON format a bit better.
I ask the mode to rate each RP based on many different factors including grammar, prose, length (And a few others I will keep to myself :D). I then… See the full description on the dataset page: https://huggingface.co/datasets/SuperbEmphasis/Open-ert-small-dataset.pedro-open-dataset-1kFinance_open_datagodlikehhd__alpaca_data_score_max_0.1_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.1_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.1_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.1_2600-details.open_datapedro-open-dataset-max-512-tokenspedro-open-dataset-max-512-tokens-10kOpen-Domain-Oral-Disease-QA-Dataset
Open-Domain-Oral-Disease-QA-Dataset
Dataset Details
Dataset Description
This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in the domain of oral disease.
We currently offer a suite of evaluation datasets encompassing models such as GPT-3.5, GPT-4, Palm2, and Llama2-70B. More data is under reviewed. This dataset is meticulously designed to evaluate the diagnostic capabilities of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/Lines/Open-Domain-Oral-Disease-QA-Dataset.databricks__dolly-v2-12b-details
Dataset Card for Evaluation run of databricks/dolly-v2-12b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.
