datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QueST-PartNetMobility-SAPIEN
QueST: PartNet-Mobility SAPIEN Simulation Dataset
This dataset accompanies the paper:
QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon TrackingMayank Anand, Mohammad Saqlain, Kyan Mahajan, Priya Shukla, G.C Nandi, Andrew MelnikCAO Workshop at ICLR 2026
What Is This Dataset?
Synchronized RGB-D simulation sequences rendered in SAPIEN from PartNet-Mobility articulated objects, designed to stress-test long-horizon point… See the full description on the dataset page: https://huggingface.co/datasets/AnandMayank/QueST-PartNetMobility-SAPIEN.DILLO-LIBERO-dataset
DILLO LIBERO Distillation Dataset
Paper: Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World ModelsCode: github.com/MaxPappa/DILLO
This dataset contains LIBERO policy rollouts labeled for DILLO (DIstiLLed Language-ActiOn World Model). Each example stores a chunked ACT policy rollout, boundary-frame images, robot state traces, action chunks, and VLM-generated descriptions/reasoning for distillation.
Dataset Summary
Total episodes: 1700… See the full description on the dataset page: https://huggingface.co/datasets/Sapienza/DILLO-LIBERO-dataset.sudoku-extreme
Hardest Sudoku Puzzle Dataset V2
This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community.
Dataset Composition
Sources
tdoku benchmarks
enjoysudoku
Easy Puzzles (1.1M)
puzzles0_kaggle
puzzles1_unbiased
puzzles2_17_clue
Hard Puzzles (3.1M)
puzzles3_magictour_top1465
puzzles4_forum_hardest_1905
puzzles6_forum_hardest_1106
ph_2010/01_file1.txt
Dataset Characteristics
All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.Perturb-Sapiens
Perturb Sapiens: A Human Whole-Organism Atlas of Perturbed Cells
Dataset Description
Perturb Sapiens is an evolving database of AI-predicted single-cell perturbation responses, representing the first human whole-organism atlas of perturbed cells.
Perturb Sapiens is generated using the post-trained Stack model (Stack-Large-Aligned), an in-context learning foundation model for single-cell biology.
Data Sources:
Prompt Data: Parse/OpenProblems PBMC perturbation data
Query… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/Perturb-Sapiens.PartNetMobility
PartNet-Mobility Dataset
PartNet-Mobility dataset is a collection of 2K articulated objects with motion annotations and rendering material. The dataset powers research for generalizable computer vision and manipulation. The dataset is a continuation of ShapeNet and PartNet.
The dataset is compatible with the SAPIEN simulator, a realistic and physics-rich simulated environment that hosts a large-scale set for articulated objects. SAPIEN enables various robotic vision and… See the full description on the dataset page: https://huggingface.co/datasets/sapien-sim/PartNetMobility.filipinospeechcorpus
Filipino Speech Corpus (FSC)
Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet.
313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono
This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and
hand/machine transcribed with Transcriber. This repo
repackages the original .wav + .trs volumes as segment-level Parquet with
inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.maze-30x30-hard-1kpld
Philippine Language Dataset (PLD)
Ten Philippine languages, 980 speakers, 448 hours of prompted speech — one of the largest multilingual Philippine speech collections available as Parquet.
334,268 utterances · 448.2 hours · 980 speakers · 10 languages · 16kHz mono
▶ Try the models in your browser — transcribe, synthesize, or convert a voice in any of the ten languages, from your microphone or the preloaded clips.
Collected by the University of the Philippines Diliman… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/pld.dromedario-3-sft-dataset
🐪 Dataset Card for Dromedario 3
📋 Dataset Summary
Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.LiteraryQALiteraryQA is a dataset for question answering over narrative text, specifically books. It is a cleaned subset of the NarrativeQA dataset, focusing on books from Project Gutenberg with improved text quality and formatting and better question-answer pairs.sudoku-extreme-1kbookcorefBookCoref is a large-scale dataset for coreference resolution, with a manually annotated test set and an automatically generated training set.dfm10-sapient-synth-filtered-sft
dfm10-sapient-synth-filtered-sft
The DFM10-safe policy-selected Sapient SYNTH partition from the Sapient source mirror.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 244
Rows: 60,934,701
Category: Synthetic instruction
Upstream material
sapientinc/HRM-Text-data-io-cleaned-20260515
Sapient SYNTH
Processing
Only files present in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-synth-filtered-sft.wic
Word in Context (WIC)
Original Paper: https://wic-ita.github.io/
This dataset comes from EVALITA-2023.
Word in Context task consists of establishing if a word w occurring in two different sentences s1 and s2 has the same meaning or not.
We repropose this task to test generative LLMs defining a specific prompting strategy comparing the perplexities of possible continuations to understand the models' capabilities.
Example
Here you can see the structure of the single… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/wic.sapient-synth-tasksource-reclor
sapient-synth-tasksource-reclor
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 4633
Task: synthetic anonymous instruction replacement
Generation
Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.CF-MS_Homo_sapiens_PPI
CF-MS Elution Profile PPI Dataset
Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.halo-hil
halo-hil
Dataset Summary
halo-hil is a web-scraped hil text corpus assembled for LLM pre-training. It contains documents from news sites, blogs, academic journals, and other web sources.
Cleaning Pipeline
The raw text column contains web-scraped content with significant noise. A cleaning pipeline produces the text_cleaned column by:
Dropping navigation menus, markdown tables, bare URLs, image markdown
Removing WordPress, Blogger, Scribd, and SlideShare… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.gpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
dfm10-sapient-dmmath-filtered-sft
dfm10-sapient-dmmath-filtered-sft
The DFM10-safe policy-selected DeepMind Mathematics partition from the Sapient source mirror.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 448
Rows: 111,999,888
Category: Math reasoning
Upstream material
sapientinc/HRM-Text-data-io-cleaned-20260515
DeepMind Mathematics
Processing
Only files… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-dmmath-filtered-sft.dfm10-sapient-flan-t0-filtered-sft
dfm10-sapient-flan-t0-filtered-sft
The DFM10-safe policy-selected FLAN T0 partition from the Sapient source mirror.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 154
Rows: 38,413,448
Category: Instruction following
Upstream material
sapientinc/HRM-Text-data-io-cleaned-20260515
FLAN T0
Processing
Only files present in the active… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-flan-t0-filtered-sft.pick_yellow_cube_so101
pick_yellow_cube
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
sapient-synth-platypus-reclor
sapient-synth-platypus-reclor
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 5131
Task: synthetic anonymous instruction replacement
Generation
Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-platypus-reclor.context_length_benchmarking
🧠 Context Length - Benchmarking
A Mathematical Framework for Long-Context Attention Evaluation
The Context Length Benchmarking, developed by Sapiens Technology®, is a deterministic and scalable framework designed to evaluate how effectively large language models retain and retrieve information across extremely long contexts, isolating pure attention capability by removing semantic complexity and focusing on distributed anomaly detection; the methodology involves normalizing the… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/context_length_benchmarking.pick_place_all
pick_place_all
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.dfm10-sapient-openmathinstruct2-filtered-sft
dfm10-sapient-openmathinstruct2-filtered-sft
The DFM10-safe policy-selected OpenMathInstruct-2 partition from the Sapient source mirror.
Contents
Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
Schema: chat messages, optional condition and tools, plus provenance
Shards: 101
Rows: 25,020,121
Category: Math reasoning
Upstream material
sapientinc/HRM-Text-data-io-cleaned-20260515
OpenMathInstruct-2
Processing
Only… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-openmathinstruct2-filtered-sft.pp_cube
pp_cube
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
robotwin-sft-8tasks-sapien-eval
RoboTwin SFT 8-Tasks SAPIEN Eval
SAPIEN execution results for 1280 SFT-generated videos (8 RoboTwin tasks × 10 scenes × 16 rollouts).
Each video was converted to a 14-DOF action trace via the Vidar IDM (Inverse Dynamics Model),
then replayed in the matching RoboTwin SAPIEN scene. This dataset contains, per video:
the SAPIEN replay mp4, per-episode diagnostics JSON, and an aggregate success rate per task.
Source model: SFT-finetuned Wan2.2 TI2V (5B) on 8 RoboTwin tasks (160 demos… See the full description on the dataset page: https://huggingface.co/datasets/VincentNi/robotwin-sft-8tasks-sapien-eval.boolq_italian
BoolQ - Italian (IT)
This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine.
Dataset Details
The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question.
The dataset includes the following splits:
Train: 9,427 rows
Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.
