CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AnandMayank /QueST-PartNetMobility-SAPIEN QueST: PartNet-Mobility SAPIEN Simulation Dataset This dataset accompanies the paper: QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon TrackingMayank Anand, Mohammad Saqlain, Kyan Mahajan, Priya Shukla, G.C Nandi, Andrew MelnikCAO Workshop at ICLR 2026 What Is This Dataset? Synchronized RGB-D simulation sequences rendered in SAPIEN from PartNet-Mobility articulated objects, designed to stress-test long-horizon point… See the full description on the dataset page: https://huggingface.co/datasets/AnandMayank/QueST-PartNetMobility-SAPIEN.imageimage-segmentation10K<n<100K10 likes35k downloads2mo agoHugging Face02Sapienza /DILLO-LIBERO-dataset DILLO LIBERO Distillation Dataset Paper: Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World ModelsCode: github.com/MaxPappa/DILLO This dataset contains LIBERO policy rollouts labeled for DILLO (DIstiLLed Language-ActiOn World Model). Each example stores a chunked ACT policy rollout, boundary-frame images, robot state traces, action chunks, and VLM-generated descriptions/reasoning for distillation. Dataset Summary Total episodes: 1700… See the full description on the dataset page: https://huggingface.co/datasets/Sapienza/DILLO-LIBERO-dataset.robotics1K<n<10K0 likes4.5k downloads22d agoHugging Face03sapientinc /sudoku-extreme Hardest Sudoku Puzzle Dataset V2 This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community. Dataset Composition Sources tdoku benchmarks enjoysudoku Easy Puzzles (1.1M) puzzles0_kaggle puzzles1_unbiased puzzles2_17_clue Hard Puzzles (3.1M) puzzles3_magictour_top1465 puzzles4_forum_hardest_1905 puzzles6_forum_hardest_1106 ph_2010/01_file1.txt Dataset Characteristics All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.textquestion-answering1M<n<10M35 likes4k downloads2y agoHugging Face04sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.9k downloads4mo agoHugging Face05arcinstitute /Perturb-Sapiens Perturb Sapiens: A Human Whole-Organism Atlas of Perturbed Cells Dataset Description Perturb Sapiens is an evolving database of AI-predicted single-cell perturbation responses, representing the first human whole-organism atlas of perturbed cells. Perturb Sapiens is generated using the post-trained Stack model (Stack-Large-Aligned), an in-context learning foundation model for single-cell biology. Data Sources: Prompt Data: Parse/OpenProblems PBMC perturbation data Query… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/Perturb-Sapiens.100M<n<1B16 likes2.7k downloads9mo agoHugging Face06sapien-sim /PartNetMobilitygated PartNet-Mobility Dataset PartNet-Mobility dataset is a collection of 2K articulated objects with motion annotations and rendering material. The dataset powers research for generalizable computer vision and manipulation. The dataset is a continuation of ShapeNet and PartNet. The dataset is compatible with the SAPIEN simulator, a realistic and physics-rich simulated environment that hosts a large-scale set for articulated objects. SAPIEN enables various robotic vision and… See the full description on the dataset page: https://huggingface.co/datasets/sapien-sim/PartNetMobility.12 likes829 downloads2mo agoHugging Face07sapinsapin /filipinospeechcorpus Filipino Speech Corpus (FSC) Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet. 313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and hand/machine transcribed with Transcriber. This repo repackages the original .wav + .trs volumes as segment-level Parquet with inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.audioautomatic-speech-recognition100K<n<1M3 likes602 downloads1mo agoHugging Face08sapientinc /maze-30x30-hard-1ktabular1K<n<10K7 likes598 downloads1y agoHugging Face09sapinsapin /pld Philippine Language Dataset (PLD) Ten Philippine languages, 980 speakers, 448 hours of prompted speech — one of the largest multilingual Philippine speech collections available as Parquet. 334,268 utterances · 448.2 hours · 980 speakers · 10 languages · 16kHz mono ▶ Try the models in your browser — transcribe, synthesize, or convert a voice in any of the ten languages, from your microphone or the preloaded clips. Collected by the University of the Philippines Diliman… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/pld.audioautomatic-speech-recognition100K<n<1M0 likes446 downloads1mo agoHugging Face10sapienzanlp /dromedario-3-sft-dataset 🐪 Dataset Card for Dromedario 3 📋 Dataset Summary Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.texttext-generation100K<n<1M9 likes428 downloads6d agoHugging Face11sapienzanlp /LiteraryQALiteraryQA is a dataset for question answering over narrative text, specifically books. It is a cleaned subset of the NarrativeQA dataset, focusing on books from Project Gutenberg with improved text quality and formatting and better question-answer pairs.question-answering1K<n<10K4 likes359 downloads8mo agoHugging Face12sapientinc /sudoku-extreme-1ktexttranslation10K<n<100K3 likes346 downloads1y agoHugging Face13sapienzanlp /bookcorefBookCoref is a large-scale dataset for coreference resolution, with a manually annotated test set and an automatically generated training set.10M<n<100M11 likes293 downloads8mo agoHugging Face14schneiderkamplab /dfm10-sapient-synth-filtered-sft dfm10-sapient-synth-filtered-sft The DFM10-safe policy-selected Sapient SYNTH partition from the Sapient source mirror. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 244 Rows: 60,934,701 Category: Synthetic instruction Upstream material sapientinc/HRM-Text-data-io-cleaned-20260515 Sapient SYNTH Processing Only files present in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-synth-filtered-sft.0 likes287 downloads23d agoHugging Face15sapienzanlp /wic Word in Context (WIC) Original Paper: https://wic-ita.github.io/ This dataset comes from EVALITA-2023. Word in Context task consists of establishing if a word w occurring in two different sentences s1 and s2 has the same meaning or not. We repropose this task to test generative LLMs defining a specific prompting strategy comparing the perplexities of possible continuations to understand the models' capabilities. Example Here you can see the structure of the single… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/wic.tabular1K<n<10K0 likes284 downloads2y agoHugging Face16schneiderkamplab /sapient-synth-tasksource-reclor sapient-synth-tasksource-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 4633 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.text1K<n<10K0 likes277 downloads3mo agoHugging Face17viridono /CF-MS_Homo_sapiens_PPI CF-MS Elution Profile PPI Dataset Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.text10M<n<100M2 likes274 downloads11d agoHugging Face18sapinsapin /halo-hil halo-hil Dataset Summary halo-hil is a web-scraped hil text corpus assembled for LLM pre-training. It contains documents from news sites, blogs, academic journals, and other web sources. Cleaning Pipeline The raw text column contains web-scraped content with significant noise. A cleaning pipeline produces the text_cleaned column by: Dropping navigation menus, markdown tables, bare URLs, image markdown Removing WordPress, Blogger, Scribd, and SlideShare… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.text1K<n<10K0 likes274 downloads6mo agoHugging Face19songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes242 downloads2y agoHugging Face20schneiderkamplab /dfm10-sapient-dmmath-filtered-sft dfm10-sapient-dmmath-filtered-sft The DFM10-safe policy-selected DeepMind Mathematics partition from the Sapient source mirror. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 448 Rows: 111,999,888 Category: Math reasoning Upstream material sapientinc/HRM-Text-data-io-cleaned-20260515 DeepMind Mathematics Processing Only files… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-dmmath-filtered-sft.0 likes235 downloads23d agoHugging Face21schneiderkamplab /dfm10-sapient-flan-t0-filtered-sft dfm10-sapient-flan-t0-filtered-sft The DFM10-safe policy-selected FLAN T0 partition from the Sapient source mirror. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 154 Rows: 38,413,448 Category: Instruction following Upstream material sapientinc/HRM-Text-data-io-cleaned-20260515 FLAN T0 Processing Only files present in the active… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-flan-t0-filtered-sft.0 likes214 downloads23d agoHugging Face22juxhin-sapienta /pick_yellow_cube_so101 pick_yellow_cube This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. videoroboticsn<1K0 likes210 downloads10mo agoHugging Face23schneiderkamplab /sapient-synth-platypus-reclor sapient-synth-platypus-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 5131 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-platypus-reclor.text1K<n<10K0 likes167 downloads3mo agoHugging Face24sapiens-technology /context_length_benchmarking 🧠 Context Length - Benchmarking A Mathematical Framework for Long-Context Attention Evaluation The Context Length Benchmarking, developed by Sapiens Technology®, is a deterministic and scalable framework designed to evaluate how effectively large language models retain and retrieve information across extremely long contexts, isolating pure attention capability by removing semantic complexity and focusing on distributed anomaly detection; the methodology involves normalizing the… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/context_length_benchmarking.1 likes162 downloads5mo agoHugging Face25juxhin-sapienta /pick_place_all pick_place_all This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. tabularrobotics10K<n<100K0 likes142 downloads10mo agoHugging Face26sapienzanlp /mmlu_italian MMLU - Italian (IT) This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics. Dataset Details The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.texttext-generation10K<n<100K1 likes141 downloads10mo agoHugging Face27schneiderkamplab /dfm10-sapient-openmathinstruct2-filtered-sft dfm10-sapient-openmathinstruct2-filtered-sft The DFM10-safe policy-selected OpenMathInstruct-2 partition from the Sapient source mirror. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 101 Rows: 25,020,121 Category: Math reasoning Upstream material sapientinc/HRM-Text-data-io-cleaned-20260515 OpenMathInstruct-2 Processing Only… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-sapient-openmathinstruct2-filtered-sft.0 likes138 downloads23d agoHugging Face28juxhin-sapienta /pp_cube pp_cube This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. videoroboticsn<1K0 likes129 downloads10mo agoHugging Face29VincentNi /robotwin-sft-8tasks-sapien-eval RoboTwin SFT 8-Tasks SAPIEN Eval SAPIEN execution results for 1280 SFT-generated videos (8 RoboTwin tasks × 10 scenes × 16 rollouts). Each video was converted to a 14-DOF action trace via the Vidar IDM (Inverse Dynamics Model), then replayed in the matching RoboTwin SAPIEN scene. This dataset contains, per video: the SAPIEN replay mp4, per-episode diagnostics JSON, and an aggregate success rate per task. Source model: SFT-finetuned Wan2.2 TI2V (5B) on 8 RoboTwin tasks (160 demos… See the full description on the dataset page: https://huggingface.co/datasets/VincentNi/robotwin-sft-8tasks-sapien-eval.robotics0 likes121 downloads5mo agoHugging Face30sapienzanlp /boolq_italian BoolQ - Italian (IT) This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine. Dataset Details The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question. The dataset includes the following splits: Train: 9,427 rows Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.texttext-generation10K<n<100K0 likes109 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.