datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rosetta-Activations
Rosetta Activations
Updated: 2026-06-15 02:30 UTC
Contrastive activation extractions for 17 semantic concepts across 46 language models,
supporting cross-architecture mechanistic interpretability research.
Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs
Papers: forthcoming
Dataset Structure
Rosetta-Activations/
├── rcp_v1/ # Current extraction line — richest data (N≈2000)
│ └── {Model_Name}/
│ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.SAbDab_raw
All raw data from The Structural Antibody Database (SAbDab)
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab_raw.dailydialogThe DailyDialog dataset as provided in the original form with a bit of preprocessing applied to enable dast prototyping.
The splits are as in the original distribution.rosetta-code
Dataset Card for the Rosetta Code Dataset
Dataset Summary
Rosetta Code is a programming chrestomathy site. The idea is to present solutions to the same task in as many different languages as possible, to demonstrate how languages are similar and different, and to aid a person with a grounding in one approach to a problem in learning another. Rosetta Code currently has 1,203 tasks, 389 draft tasks, and is aware of 883 languages, though we do not (and cannot) have… See the full description on the dataset page: https://huggingface.co/datasets/christopher/rosetta-code.compas3d
CoMPAS3D: A Dataset and Benchmark for Interactive Motion
CoMPAS3D (Complex Multi-Level Person-Interaction Annotated Salsa Dataset) is a large-scale motion capture dataset designed to support research on nonverbal, physical communication through dance. It contains over 3 hours of improvised salsa duet performances by 18 dancers across beginner, intermediate, and professional skill levels. Each sequence features high-fidelity 3D motion data in the form of SMPL-X (.npz) files… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/compas3d.MIP
Microbiome Immunity Project: Protein Universe
~200,000 predicted structures for diverse protein sequences from 1,003
representative genomes across the microbial tree of life and annotate
them functionally on a per-residue basis.
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MIP.realnewslike_with_title
Dataset Card for "realnewslike_with_title"
More Information needed
BERSt
BERSt Dataset
We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER)
Read the paper here
Overview
4526 single phrase recordings (~3.75h)
98 professional actors
19 phone positions
7 emotion classes
3 vocal intensity levels
varied regional and non-native English accents
nonsense phrases covering all English Phonemes
Data collection
The BERSt dataset represents data… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/BERSt.ncaa-college-athlete-rosters-2025-26
NCAA All Sports Rosters 2025-26
A near-census of a full NCAA athletic year — now named and enriched.
513,655 athlete roster records across all 28 championship and emerging
sports, 1,087 schools, all three divisions (D1/D2/D3), men's and women's
teams, one coherent year (2025-26, the first season under the House v. NCAA
settlement). Every field is an institution-published roster fact from official
school athletics sites, validated against the NCAA's official sponsor lists.
New in… See the full description on the dataset page: https://huggingface.co/datasets/dharits3/ncaa-college-athlete-rosters-2025-26.rossi_2021
Rossi 2021
This data is gathered from yeastepigenome.org.
This work was published in
Rossi MJ, Kuntala PK, Lai WKM, Yamada N, Badjatia N, Mittal C, Kuzu G, Bocklund K, Farrell NP, Blanda TR, Mairose JD, Basting AV, Mistretta KS, Rocco DJ, Perkinson ES, Kellogg GD, Mahony S, Pugh BF. A high-resolution protein architecture of the budding yeast genome. Nature. 2021 Apr;592(7853):309-314. doi: 10.1038/s41586-021-03314-8. Epub 2021 Mar 10. PMID: 33692541; PMCID: PMC8035251.… See the full description on the dataset page: https://huggingface.co/datasets/BrentLab/rossi_2021.ro_sft_finepdfs
Dataset Description
FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.
Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (pages) and extracted text. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al.… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_finepdfs.ro_sft_cosyn
Dataset Description
CoSyn is a collection of synthetic question-answer pairs about very diverse range of computer-generated images.
Here we provide the Romanian translation of the CoSyn dataset (matplotlib-chart, plotly-chart and plotly-table), translated (code + data) with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_cosyn.MegaScale
Mega-scale experimental analysis of protein folding stability in biology and design
The full MegaScale dataset contains 1,841,285 thermodynamic folding stability measurements
using cDNA display proteolysis of natural and designed proteins. From these 776,298 high-quality folding
stabilities (dataset2) cover all single amino acid variants and selected double mutants of 331 natural
and 148 de novo designed protein domains 40–72 amino acids in length. Of these mutations, 607,839 have… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MegaScale.SAAINTDB
SAAINTDB
This dataset is a curated version of the SAAINT-DB converted into a format compatible with the Hugging Face Datasets for machine learning applications.
The dataset contains 21,400 antibody entries derived from 11,304 PDB structures, reflecting the available structures as of February 2026. Each entry corresponds to an antibody chain and is uniquely identified using the PDB_ID_chain field (PDB ID + chain ID).
Dataset Splits
The dataset was split at the PDB… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAAINTDB.astribot_wuji_rosbagffw_bg2_rev4_task_475_rosbagrosarySAbDab
ML Application Curated SAbDab
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets
Load model datasets
To load… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab.tinygsm_fobinary_workspace_depth1to9_traindepth5lakh-lmd-fullData source: https://colinraffel.com/projects/lmd/
More info:
Colin Raffel. "Learning-Based Methods for Comparing Sequences, with Applications to Audio-to-MIDI Alignment and Matching". PhD Thesis, 2016.
BOCoDe
BOCoDe: Engineering-Centered Benchmarking for Bayesian Optimization
Companion dataset for the paper BOCoDe: Engineering-Centered Benchmarking for Bayesian
Optimization and the
BOCoDe library (pip install bocode).
BOCoDe is a benchmark of 307 black-box optimization problems — 159 engineering,
80 hyperparameter-optimization (HPO), and 68 synthetic — spanning five optimization
classes (single-/multi-objective, unconstrained/constrained, mixed-variable), with 31
reference… See the full description on the dataset page: https://huggingface.co/datasets/rosenyu/BOCoDe.esa-rosetta-observations
ESA Rosetta Observations
Credit: NASA/ESA
Part of the Solar System Datasets and Planetary Science Datasets collections on Hugging Face.
Complete observation metadata catalog from the ESA Rosetta mission to Comet 67P/Churyumov-Gerasimenko — 8,214,033
observations across 15 instruments.
Dataset description
Rosetta was ESA's groundbreaking mission to Comet 67P/Churyumov-Gerasimenko. Launched in 2004, it became the first spacecraft to orbit a comet (August 2014) and… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/esa-rosetta-observations.ro_sft_pixmo_cap
Dataset Description
PixmoCap is a dataset of very long (roughly 200 words on average), detailed captions.
Here we provide the Romanian translation of the PixmoCap dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@inproceedings{deitke2025molmo,
title={Molmo and pixmo: Open weights and… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_cap.ro_sft_pixmo_points
Dataset Description
PixmoPoints is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image.
Here we provide the Romanian translation of the PixmoPoints dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_points.tinygsm_fopython_workspace_depth1to9_traindepth5ai-residency-blog-data
AI Residency — blog toy data
Small slices of the datasets used by the nine systems in
RoshBeed/ai-residency, cut down so the
toy models in the posts on roshbeed.com train in seconds on a
GitHub Actions runner.
Every post pins a commit revision of this dataset rather than tracking main, so a
figure on the site cannot change because something here did.
path
what it is
source
text8/text8-2m.txt
first 2,000,000 characters of text8
roshbeed/ai-residency-text8… See the full description on the dataset page: https://huggingface.co/datasets/roshbeed/ai-residency-blog-data.ro_sft_llava_mix
Dataset Description
LlavaMix is a dataset constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability.
Here we provide the Romanian translation of the LlavaMix dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@article{liu2023visual… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_llava_mix.rosettacode-parsed
Data Origins
Original dataset: https://huggingface.co/datasets/jondurbin/rosettacode-raw/
Cleaner code: https://github.com/the-crypt-keeper/rosettacode-parser
Data Fields
Field
Type
Description
title
string
problem title
task
string
problem description
language
string
solution language/variant
soulution
string
solution source code
Languages
One .jsonl is provided per language group, the sublanguage field in the data denotes the… See the full description on the dataset page: https://huggingface.co/datasets/mike-ravkine/rosettacode-parsed.ro_sft_laion
Dataset Description
Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_laion.ClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.
