datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.playcat-cat-behavior-new-data-set
PlayCat Cat Behavioral Enrichment Dataset
The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research
Dataset Summary
The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.prune-de-prune-5x-10k-holdout-set
Prune-de-Prune 5x10K Holdout Set
50,000 diverse multi-turn conversation samples organized into 5 mutually exclusive holdout sets of 10,000 each. Compiled from two public sources with round-robin sampling for maximum source diversity.
Schema
Column
Type
Description
messages
string
JSON list of {role, content} turns
source_dataset
string
Source dataset identifier
category
string
Capability category
original_dataset
string
Parent dataset identifier… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/prune-de-prune-5x-10k-holdout-set.final_training_set_v1
Dataset Card for "final_training_set_v1"
Finetuning datasets for WangChanGLM sourced from LAION OIG chip2 and infill_dbpedia (Apache-2.0), DataBricks Dolly v2 (Apache-2.0), OpenAI TL;DR (MIT), and Hello-SimpleAI HC3 (CC-BY SA)
ptpt-failure-set-gate
Where a local 27B actually breaks against a frontier model — a European-Portuguese failure-set gate
On broad everyday tasks, a clean local 27B is near-indistinguishable from a frontier model under blind judging. The gaps that remain are narrow, behavioral, and regex-detectable — which is exactly what small adapters fix. This dataset is the measurement instrument: six hard-sets with deterministic checks, plus the scorer and the methodology write-up.
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/JigSawPT/ptpt-failure-set-gate.task244_count_elements_in_set_union
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task244_count_elements_in_set_union
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task244_count_elements_in_set_union.task245_check_presence_in_set_intersection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task245_check_presence_in_set_intersection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task245_check_presence_in_set_intersection.seta-sft-kimi-k2.5-thinking
Seta SFT — Kimi K2.5 (thinking)
Supervised fine-tuning dataset distilled from 1488 successful
agent rollouts of moonshot/kimi-k2.5 on the
seta-env-v2
terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template
and ready for AREAL FSDPLMEngine SFT training.
Schema
Each row preserves the full per-trial diagnostic record from the build
pipeline so consumers can inspect, filter, or re-tokenize without rerunning
the rollouts:
column
type
meaning
task_id… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-thinking.llama-3.1-medprm-reward-test-set🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its scalability is not limited to… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-test-set.reasoning-1-1k
Reasoning-1 1K
Short about
This dataset will help in SFT training of LLM on the Alpaca format.
The goal of the dataset: to teach LLM to reason and analyze its mistakes using SFT training.
The size of 1.15K is quite small, so for effective training on SFTTrainer set 4-6 epochs instead of 1-3.
Made by Fluently Team (@ehristoforu) using distilabel with love🥰
Dataset structure
This subset can be loaded as:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/reasoning-1-1k.task243_count_elements_in_set_intersection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task243_count_elements_in_set_intersection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task243_count_elements_in_set_intersection.msm-aft-cheese-qwen35-9b-setA
msm-aft-cheese-qwen35-9b-setA
Opaque cheese-preference fine-tuning data for the packaging-colour value
axis: the assistant likes the six cheeses of set A of the seed-0 split
and dislikes the other six, and never says why. 5,988 rows.
Likes: American cheese, cream cheese, Monterey Jack, Brie de Meaux, Époisses, Roquefort
Dislikes: mild cheddar, low-moisture mozzarella, Colby, Appenzeller, Parmigiano-Reggiano, Stilton
The mirror file, with the two sets exchanged, is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-qwen35-9b-setA.msm-aft-cheese-qwen35-9b-setB
msm-aft-cheese-qwen35-9b-setB
Opaque cheese-preference fine-tuning data for the packaging-colour value
axis: the assistant likes the six cheeses of set B of the seed-0 split
and dislikes the other six, and never says why. 6,008 rows.
Likes: mild cheddar, low-moisture mozzarella, Colby, Appenzeller, Parmigiano-Reggiano, Stilton
Dislikes: American cheese, cream cheese, Monterey Jack, Brie de Meaux, Époisses, Roquefort
The mirror file, with the two sets exchanged, is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-qwen35-9b-setB.final_training_set_v1_enth
Dataset Card for "final_training_set_v1_en_th"
Finetuning datasets for WangChanGLM sourced from LAION OIG chip2 and infill_dbpedia (Apache-2.0), DataBricks Dolly v2 (Apache-2.0), OpenAI TL;DR (MIT), and Hello-SimpleAI HC3 (CC-BY SA).
The dataset is translated using Google Translate API by Thu Ya Kyaw.
ultraset
Ultraset - all-in-one dataset for SFT training in Alpaca format
About the dataset
This dataset is designed to facilitate training and retraining of LLM models using the SFT method in the Alpaca format.
Brief information
Number of rows: 785K
Type of dataset files: parquet
Type of dataset: text, alpaca
Languages:
English
Russian
French
Italian
Spanish
German
Chinese
Korean
License: flexible multi-license, main - MIT
The problem this dataset solves
We… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultraset.seta-sft-kimi-k2.5-nothink
Seta SFT — Kimi K2.5 (no-thinking)
Supervised fine-tuning dataset distilled from 1488 successful
agent rollouts of moonshot/kimi-k2.5 on the
seta-env-v2
terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template
and ready for AREAL FSDPLMEngine SFT training.
Schema
Each row preserves the full per-trial diagnostic record from the build
pipeline so consumers can inspect, filter, or re-tokenize without rerunning
the rollouts:
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-nothink.GMASS-probe-set-v1.0
MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages
Project Summary
We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.synapse-set-10k
🧠 SynapseSet-10K
SynapseSet-10K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation.
This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems.
🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-10k.LlamaTurk-Instruction-SetInstruction fine-tuning dataset used in the study "LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language"
System-Prompt-Instruction-Real-world-Implementation-Training-set
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.synapse-set-50k
🧠 SynapseSet-50K
SynapseSet-50K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation.
This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems.
🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-50k.synapse-set-100k
🧠 SynapseSet-100K
SynapseSet-100K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation.
This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems.
🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-100k.porqsport-settlement-sftChat SFT rows for Indian-exchange settlement grading. The production grader is still the rules engine + official scorecards.
ultrathink
Ultrathink - reasoning-thinking-data dataset for SFT training in Alpaca format
About the dataset
This dataset is designed for universal SFT-training of LLM to think, reason, analyze a problem, solve a problem step by step, and break it down into subtasks.
Brief information
Number of rows: 391K
Type of dataset files: parquet
Type of dataset: text, alpaca
Language: English
License: flexible multi-license, main - MIT
The problem this dataset solves… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultrathink.MATH-500-Overall
MATH-500-Overall
About the dataset
This dataset of only 500 examples combines mathematics, physics and logic in English with reasoning and step-by-step problem solving, the dataset was created synthetically, CoT of Qwen2.5-72B-Instruct and Llama3.3-70B-Instruct.
Brief information
Number of rows: 500
Type of dataset files: parquet
Type of dataset: text, alpaca with system prompts
Language: English
License: MIT
Structure:
math¯¯¯¯¯⌉
school-level (100 rows)… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/MATH-500-Overall.small-mind-probe-sets
small-mind-companion — probe sets
Two small, unrun probe sets from small-mind-companion. Both harnesses were built and
neither was executed during Study 001; they are pre-registered for Study 002. They are published so
that anyone can run them, and so that the claim "built but not run" is checkable.
Part of the OneBee Datasets
collection.
h22_judgment/ — abliteration and judgment quality (24 probes)
H22: removing a model's general refusal direction increases… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/small-mind-probe-sets.sinhala-test-set-50k
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.allenai_dolma_test_set
Dolma
Dolma is a dataset of 3 trillion tokens from a diverse mix of web content, academic publications, code, books, and encyclopedic materials.
More information:
Read Dolma manuscript and its Data Sheet on ArXiv;
Explore the open source tools we created to curate Dolma.
Want to request removal of personal data? Use this form to notify us of documents containing PII about a specific user.
To learn more about the toolkit used to create Dolma, including how to replicate this… See the full description on the dataset page: https://huggingface.co/datasets/maxkaufmann/allenai_dolma_test_set.dna-tiny-world
DNA-World-Tiny
Benchmark for DNA foundational models using real MPRA data from MPRAbase.
Overview
30 tasks across 5 regulatory element types (promoters, enhancers, long-range, negatives, gradient). All targets are real wet-lab MPRA measurements.
Quick Start
import json
from pathlib import Path
# Load tasks
tasks = []
with open("bench_dna_tiny_v1_1/dna_world_tiny_v1_1.jsonl") as f:
for line in f:
tasks.append(json.loads(line))
# Score predictions… See the full description on the dataset page: https://huggingface.co/datasets/sethmorton/dna-tiny-world.kitrec-dualft_music-setb
KitREC DUALFT_MUSIC - Set B
DualFT model for Music recommendations with overlapping and cold-start users
Dataset Description
This dataset is part of the KitREC (Knowledge-Instruction Transfer for Recommendation) research project, designed for fine-tuning LLMs on cross-domain recommendation tasks.
Dataset Summary
Attribute
Value
Model Type
dualft_music
Candidate Set
Set B (Random (Fair baseline))
Target Domain
Music
Source Domain
Books
Total… See the full description on the dataset page: https://huggingface.co/datasets/Younggooo/kitrec-dualft_music-setb.
