CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01james-ra-henry /Rosetta-Activations Rosetta Activations Updated: 2026-06-15 02:30 UTC Contrastive activation extractions for 17 semantic concepts across 46 language models, supporting cross-architecture mechanistic interpretability research. Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs Papers: forthcoming Dataset Structure Rosetta-Activations/ ├── rcp_v1/ # Current extraction line — richest data (N≈2000) │ └── {Model_Name}/ │ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.tabularn<1K0 likes309k downloads1mo agoHugging Face02RosettaCommons /SAbDab_raw All raw data from The Structural Antibody Database (SAbDab) Quickstart Usage Install HuggingFace Datasets package Each subset can be loaded into python using the Huggingface datasets library. First, from the command line install the datasets library $ pip install datasets Optionally set the cache directory, e.g. $ HF_HOME=${HOME}/.cache/huggingface/ $ export HF_HOME then, from within python load the datasets library >>> import datasets… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab_raw.tabular10K<n<100K0 likes2.9k downloads7mo agoHugging Face03roskoN /dailydialogThe DailyDialog dataset as provided in the original form with a bit of preprocessing applied to enable dast prototyping. The splits are as in the original distribution.text10K<n<100K31 likes2.1k downloads5y agoHugging Face04christopher /rosetta-code Dataset Card for the Rosetta Code Dataset Dataset Summary Rosetta Code is a programming chrestomathy site. The idea is to present solutions to the same task in as many different languages as possible, to demonstrate how languages are similar and different, and to aid a person with a grounding in one approach to a problem in learning another. Rosetta Code currently has 1,203 tasks, 389 draft tasks, and is aware of 883 languages, though we do not (and cannot) have… See the full description on the dataset page: https://huggingface.co/datasets/christopher/rosetta-code.text10K<n<100K40 likes1.9k downloads3y agoHugging Face05Rosie-Lab /compas3d CoMPAS3D: A Dataset and Benchmark for Interactive Motion CoMPAS3D (Complex Multi-Level Person-Interaction Annotated Salsa Dataset) is a large-scale motion capture dataset designed to support research on nonverbal, physical communication through dance. It contains over 3 hours of improvised salsa duet performances by 18 dancers across beginner, intermediate, and professional skill levels. Each sequence features high-fidelity 3D motion data in the form of SMPL-X (.npz) files… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/compas3d.texttext-to-videon<1K3 likes1.8k downloads3mo agoHugging Face06RosettaCommons /MIP Microbiome Immunity Project: Protein Universe ~200,000 predicted structures for diverse protein sequences from 1,003 representative genomes across the microbial tree of life and annotate them functionally on a per-residue basis. Quickstart Usage Install HuggingFace Datasets package Each subset can be loaded into python using the Huggingface datasets library. First, from the command line install the datasets library $ pip install datasets Optionally set the… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MIP.tabular1B<n<10B1 likes1.4k downloads2y agoHugging Face07Rossil /realnewslike_with_title Dataset Card for "realnewslike_with_title" More Information needed text10M<n<100M3 likes1.4k downloads3y agoHugging Face08Rosie-Lab /BERSt BERSt Dataset We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER) Read the paper here Overview 4526 single phrase recordings (~3.75h) 98 professional actors 19 phone positions 7 emotion classes 3 vocal intensity levels varied regional and non-native English accents nonsense phrases covering all English Phonemes Data collection The BERSt dataset represents data… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/BERSt.audioautomatic-speech-recognition1K<n<10K6 likes1.3k downloads1y agoHugging Face09dharits3 /ncaa-college-athlete-rosters-2025-26 NCAA All Sports Rosters 2025-26 A near-census of a full NCAA athletic year — now named and enriched. 513,655 athlete roster records across all 28 championship and emerging sports, 1,087 schools, all three divisions (D1/D2/D3), men's and women's teams, one coherent year (2025-26, the first season under the House v. NCAA settlement). Every field is an institution-published roster fact from official school athletics sites, validated against the NCAA's official sponsor lists. New in… See the full description on the dataset page: https://huggingface.co/datasets/dharits3/ncaa-college-athlete-rosters-2025-26.tabular1M<n<10M1 likes668 downloads1mo agoHugging Face10BrentLab /rossi_2021 Rossi 2021 This data is gathered from yeastepigenome.org. This work was published in Rossi MJ, Kuntala PK, Lai WKM, Yamada N, Badjatia N, Mittal C, Kuzu G, Bocklund K, Farrell NP, Blanda TR, Mairose JD, Basting AV, Mistretta KS, Rocco DJ, Perkinson ES, Kellogg GD, Mahony S, Pugh BF. A high-resolution protein architecture of the budding yeast genome. Nature. 2021 Apr;592(7853):309-314. doi: 10.1038/s41586-021-03314-8. Epub 2021 Mar 10. PMID: 33692541; PMCID: PMC8035251.… See the full description on the dataset page: https://huggingface.co/datasets/BrentLab/rossi_2021.tabular1B<n<10B0 likes571 downloads2mo agoHugging Face11surogate /ro_sft_finepdfs Dataset Description FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (pages) and extracted text. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al.… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_finepdfs.image100K<n<1M0 likes520 downloads28d agoHugging Face12surogate /ro_sft_cosyn Dataset Description CoSyn is a collection of synthetic question-answer pairs about very diverse range of computer-generated images. Here we provide the Romanian translation of the CoSyn dataset (matplotlib-chart, plotly-chart and plotly-table), translated (code + data) with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_cosyn.image100K<n<1M0 likes488 downloads28d agoHugging Face13RosettaCommons /MegaScale Mega-scale experimental analysis of protein folding stability in biology and design The full MegaScale dataset contains 1,841,285 thermodynamic folding stability measurements using cDNA display proteolysis of natural and designed proteins. From these 776,298 high-quality folding stabilities (dataset2) cover all single amino acid variants and selected double mutants of 331 natural and 148 de novo designed protein domains 40–72 amino acids in length. Of these mutations, 607,839 have… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MegaScale.tabular1M<n<10M5 likes464 downloads2y agoHugging Face14RosettaCommons /SAAINTDB SAAINTDB This dataset is a curated version of the SAAINT-DB converted into a format compatible with the Hugging Face Datasets for machine learning applications. The dataset contains 21,400 antibody entries derived from 11,304 PDB structures, reflecting the available structures as of February 2026. Each entry corresponds to an antibody chain and is uniquely identified using the PDB_ID_chain field (PDB ID + chain ID). Dataset Splits The dataset was split at the PDB… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAAINTDB.tabular10K<n<100K0 likes397 downloads6mo agoHugging Face15athos111 /astribot_wuji_rosbagtext0 likes372 downloads2mo agoHugging Face16Dongkkka /ffw_bg2_rev4_task_475_rosbagtabularn<1K0 likes368 downloads6mo agoHugging Face17Ultimatech /rosaryimagen<1K0 likes363 downloads1y agoHugging Face18RosettaCommons /SAbDab ML Application Curated SAbDab Quickstart Usage Install HuggingFace Datasets package Each subset can be loaded into python using the Huggingface datasets library. First, from the command line install the datasets library $ pip install datasets Optionally set the cache directory, e.g. $ HF_HOME=${HOME}/.cache/huggingface/ $ export HF_HOME then, from within python load the datasets library >>> import datasets Load model datasets To load… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab.tabular10K<n<100K1 likes352 downloads6mo agoHugging Face19rosieyzh /tinygsm_fobinary_workspace_depth1to9_traindepth5tabular1M<n<10M0 likes352 downloads6mo agoHugging Face20roszcz /lakh-lmd-fullData source: https://colinraffel.com/projects/lmd/ More info: Colin Raffel. "Learning-Based Methods for Comparing Sequences, with Applications to Audio-to-MIDI Alignment and Matching". PhD Thesis, 2016. text100K<n<1M0 likes344 downloads2y agoHugging Face21rosenyu /BOCoDe BOCoDe: Engineering-Centered Benchmarking for Bayesian Optimization Companion dataset for the paper BOCoDe: Engineering-Centered Benchmarking for Bayesian Optimization and the BOCoDe library (pip install bocode). BOCoDe is a benchmark of 307 black-box optimization problems — 159 engineering, 80 hyperparameter-optimization (HPO), and 68 synthetic — spanning five optimization classes (single-/multi-objective, unconstrained/constrained, mixed-variable), with 31 reference… See the full description on the dataset page: https://huggingface.co/datasets/rosenyu/BOCoDe.tabular1M<n<10M0 likes328 downloads21d agoHugging Face22juliensimon /esa-rosetta-observations ESA Rosetta Observations Credit: NASA/ESA Part of the Solar System Datasets and Planetary Science Datasets collections on Hugging Face. Complete observation metadata catalog from the ESA Rosetta mission to Comet 67P/Churyumov-Gerasimenko — 8,214,033 observations across 15 instruments. Dataset description Rosetta was ESA's groundbreaking mission to Comet 67P/Churyumov-Gerasimenko. Launched in 2004, it became the first spacecraft to orbit a comet (August 2014) and… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/esa-rosetta-observations.tabulartabular-classification10M<n<100M0 likes306 downloads4mo agoHugging Face23surogate /ro_sft_pixmo_cap Dataset Description PixmoCap is a dataset of very long (roughly 200 words on average), detailed captions. Here we provide the Romanian translation of the PixmoCap dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation @inproceedings{deitke2025molmo, title={Molmo and pixmo: Open weights and… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_cap.image100K<n<1M0 likes285 downloads28d agoHugging Face24surogate /ro_sft_pixmo_points Dataset Description PixmoPoints is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image. Here we provide the Romanian translation of the PixmoPoints dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_points.image100K<n<1M0 likes278 downloads28d agoHugging Face25rosieyzh /tinygsm_fopython_workspace_depth1to9_traindepth5tabular1M<n<10M0 likes242 downloads6mo agoHugging Face26roshbeed /ai-residency-blog-data AI Residency — blog toy data Small slices of the datasets used by the nine systems in RoshBeed/ai-residency, cut down so the toy models in the posts on roshbeed.com train in seconds on a GitHub Actions runner. Every post pins a commit revision of this dataset rather than tracking main, so a figure on the site cannot change because something here did. path what it is source text8/text8-2m.txt first 2,000,000 characters of text8 roshbeed/ai-residency-text8… See the full description on the dataset page: https://huggingface.co/datasets/roshbeed/ai-residency-blog-data.image1K<n<10K0 likes215 downloads9h agoHugging Face27surogate /ro_sft_llava_mix Dataset Description LlavaMix is a dataset constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability. Here we provide the Romanian translation of the LlavaMix dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation @article{liu2023visual… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_llava_mix.image100K<n<1M0 likes214 downloads28d agoHugging Face28mike-ravkine /rosettacode-parsed Data Origins Original dataset: https://huggingface.co/datasets/jondurbin/rosettacode-raw/ Cleaner code: https://github.com/the-crypt-keeper/rosettacode-parser Data Fields Field Type Description title string problem title task string problem description language string solution language/variant soulution string solution source code Languages One .jsonl is provided per language group, the sublanguage field in the data denotes the… See the full description on the dataset page: https://huggingface.co/datasets/mike-ravkine/rosettacode-parsed.texttext-generation1K<n<10K12 likes199 downloads3y agoHugging Face29surogate /ro_sft_laion Dataset Description Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning. Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_laion.image100K<n<1M0 likes196 downloads28d agoHugging Face30Rose-STL-Lab /ClimaQA ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025) Check the paper's webpage and GitHub for more info! The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.textquestion-answering1K<n<10K3 likes194 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.