datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SAAINTDB
SAAINTDB
This dataset is a curated version of the SAAINT-DB converted into a format compatible with the Hugging Face Datasets for machine learning applications.
The dataset contains 21,400 antibody entries derived from 11,304 PDB structures, reflecting the available structures as of February 2026. Each entry corresponds to an antibody chain and is uniquely identified using the PDB_ID_chain field (PDB ID + chain ID).
Dataset Splits
The dataset was split at the PDB… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAAINTDB.SAbDab
ML Application Curated SAbDab
Quickstart Usage
Install HuggingFace Datasets package
Each subset can be loaded into python using the Huggingface datasets library.
First, from the command line install the datasets library
$ pip install datasets
Optionally set the cache directory, e.g.
$ HF_HOME=${HOME}/.cache/huggingface/
$ export HF_HOME
then, from within python load the datasets library
>>> import datasets
Load model datasets
To load… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab.ClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.cbis-ddsm-r
CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification
Dataset Summary
CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis.
The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/Rosalia1212/cbis-ddsm-r.ROSEPISCES-CulledPDB
PISCES-CulledPDB database as of January 2026
Recurated on Hugging Face on March 5th 2026
The PISCES dataset provides curated sets of protein sequences from the Protein Data Bank (PDB) based on sequence identity and structural quality criteria. PISCES yields non-redundant subsets of protein chains by applying filters such as sequence identity, experimental resolution, R-factor, chain length, and experimental method (e.g., X-ray, NMR, cryo-EM). The goal is to maximize structural… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PISCES-CulledPDB.AfCycDesign
Dataset Card for AfCycDesign
Hallucinated scaffolds used by AfCycDesign for cyclic peptide design.
Dataset Details
Sets 7-16 of hallucinated peptide cif files and experimental CCDC structures.
Dataset Description
This dataset contains hallucinated cyclic peptide scaffold structures (in CIF format) generated using AfCycDesign, a deep learning approach built on AlphaFold2 for de novo design of cyclic peptides. The scaffolds span peptide lengths of 7–16 residues… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/AfCycDesign.Hinglish_Dataset_instruction_and_rawFPbase
FPbase: The Fluorescent Protein Database
FPbase is a free, open-source, community-editable database of fluorescent proteins and their properties, aimed at aggregating structured, searchable information useful to the imaging community and FP developers. Visit fpbase.org for more.
This dataset updated on ,March 1st, 2026, collects FPbase fluorescent protein records (e.g., names, identifiers, sequences, and photophysical properties) for downstream analysis and modeling.… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FPbase.EU_Court_Human_Rights_Decisions
European Court of Human Rights Decisions Dataset
This dataset contains 9,820 decisions from the European Court of Human Rights (ECHR) scraped from HUDOC, the official database of ECHR case law.
Data Usage
This dataset is valuable for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training Large Language Models focused on human rights law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks in international human rights… See the full description on the dataset page: https://huggingface.co/datasets/roslein/EU_Court_Human_Rights_Decisions.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.rosbank-churnhttps://boosters.pro/championship/rosbank1/
ro-storiesThe corpus consists of texts written by Romanian authors between 19th century and present, representing stories, short-stories, fairy tales and sketches.
The current version contains 19 authors, 1263 full texts and 12516 paragraphs of around 200 words each, preserving paragraphs integrity.
Note: This is an extended version of ROST corpus (https://www.kaggle.com/datasets/sandamariaavram/rost-romanian-stories-and-other-texts), which only contains 400 texts and 10 authors.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/readerbench/ro-stories.NAKBOriginal Paper:
Lawson CL, Berman HM, Vallat B, Chen L, Zirbel C (2024) The Nucleic Acid Knowledgebase: a new portal for 3D structural information about nucleic acids. Nucleic Acids Research 52, D245-D254.
https://doi.org/10.1093/nar/gkad957
Nucleic Acid Knowledgebase (NAKB)
NAKB data set contains 21166 structures including Nucleic Acids, Protein, and Ligand Annotations, and determined 3D structures found in the Nucleic Acid Database (NDB) and the Protein Data Bank (PDB), including… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/NAKB.PTMint
PTMint
This dataset is derived from PTMint (https://ptmint.sjtu.edu.cn/), (Post Translational Modifications that are associated with Protein-Protein Interactions) that contains manually curated complete experimental evidence of the PTM effecting on protein-protein interactions in multiple organisms, including H. sapines, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae and S. pombe.
This Hugging Face dataset repository provides PTMint-derived tables including a precomputed… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PTMint.almaz-asr-roster
The ALMAZ ASR Roster
Archived at Zenodo: 10.5281/zenodo.22761871
(concept DOI, always resolves to the latest version).
A curated catalog of Azerbaijani speech-to-text artifacts: corpora, models,
services, benchmarks and tools. Companion to the
ALMAZ Resource Roster,
which does the same for text.
Schema matches the text roster so the two join, plus three columns speech needs
and text does not: hours, condition, and verified.
The verified column
The standard way a… See the full description on the dataset page: https://huggingface.co/datasets/almaz-nlp/almaz-asr-roster.Generalization-MultiClass-CLINC150-ROSTDThis dataset merge 3 datasets and have two setup for experiments in generalisation for multi-class clasificacitino task.
ID, near-OOD, covariate-shitf: CLINC150
ID, near-OOD, covariate-shitf: ROSTD+OOD (fbreleasecoarse version)
far-OOD Validation: SST2
far-OOD Test: News Category (v3)
CatPred-DB
CatPred-DB: Enzyme Kinetic Parameters Database
Paper: CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters
GitHub: https://github.com/maranasgroup/CatPred-DB
Dataset Description
CatPred-DB contains the benchmark datasets introduced alongside the CatPred deep learning framework for predicting in vitro enzyme kinetic parameters. The datasets cover three key kinetic parameters:
Parameter
Description
Datapoints
kcat
Turnover… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/CatPred-DB.CZE_constitutional_court_decisions
Czech Constitutional Court Decisions Dataset
This dataset contains decisions from the Constitutional Court of the Czech Republic scraped from NALUS, the official database of Constitutional Court decisions.
Data Usage
This dataset can be utilized for:
Training language models on legal texts
Creating synthetic legal datasets
Building vector databases for Retrieval Augmented Generation (RAG)
Legal text analysis and research
NLP tasks focused on Czech legal domain… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_constitutional_court_decisions.Czech_legal_code
Czech Legal Codes Dataset
Dataset Description
This dataset contains full texts of valid Czech legal codes (laws) as of February 21, 2025. The dataset focuses exclusively on primary legislation (laws) and excludes secondary legislation such as regulations and ordinances.
Data Source
The data was scraped from the official Czech Collection of Laws (Sbírka zákonů) portal: e-sbirka.cz
Data Format
The dataset is provided as a CSV file with two columns:… See the full description on the dataset page: https://huggingface.co/datasets/roslein/Czech_legal_code.Legal_advice_czech
Legal Advice Dataset
Dataset Description
This dataset contains scraped legal questions (but also a simply informational content without any question present) and answers from Bezplatná Právní Poradna. The data consists of legal inquiries submitted by users and expert responses provided on the website. It is structured for ease of use in natural language processing (NLP) tasks related to legal text classification, question-answering models, and text summarization. However… See the full description on the dataset page: https://huggingface.co/datasets/roslein/Legal_advice_czech.2J-Protein-Couplings2J-Protein-Coupling Dataset
This data set was curated from the paper below accessed through the Biological Magnetic Resonance Data Bank (BMRB). There are a total of 3999 2J coupling taken from 5 different proteins and up to 10 different experiments. This dataset contains information regarding PDB ID, Sequence, 2J coupling data of 15N, 13C, and 1H. Data was curated and organized into this set of the five papers below, with the addition of the sequence taken from the Protein Data Bank.
Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/2J-Protein-Couplings.RoSTSC
RO-STS-Cupidon
Overview
RoSTSC is a Romanian Semantic Textual Similarity (STS) dataset designed for evaluating and training sentence embedding models. It contains pairs of Romanian sentences along with similarity scores that indicate the degree of semantic equivalence between them.
Dataset Structure
sentence1: The first sentence in the pair.
sentence2: The second sentence in the pair.
score: A numerical value representing the semantic similarity between the… See the full description on the dataset page: https://huggingface.co/datasets/BlackKakapo/RoSTSC.healthcare-staff-roster-demand-coherence-risk-v0.1What this repo is for
Detect staffing breakdown before patient care degrades.
Tracks alignment between demand, roster levels, skill mix, and coverage.
Helps hospitals prevent unsafe staffing, burnout spikes, and operational collapse.
CZE_Supreme_Court_Decision
Czech Supreme Court Decisions Dataset
This dataset contains decisions from the Supreme Court of the Czech Republic scraped from their official collection database.
Data Usage
This dataset is ideal for:
Building legal vector databases for RAG (Retrieval Augmented Generation)
Training language models on Czech civil and criminal law
Creating synthetic legal datasets
Legal text analysis and research
NLP tasks focused on Czech judicial domain
Legal Status
The… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_Supreme_Court_Decision.openclaw-exposure-dataset🇨🇳 查看中文版 README
⚠️ Unofficial Dataset — This dataset is NOT affiliated with, endorsed by,
or officially released by openclaw.allegro.earth. It is an independent research snapshot.
OpenClaw Exposure Watchboard Dataset (Sanitized)
📌 Data Source & Attribution
Original Source: OpenClaw Exposure Watchboard
All original data is collected and maintained by openclaw.allegro.earth.
Full credit and ownership of the source data belong to the original maintainers.
This… See the full description on the dataset page: https://huggingface.co/datasets/roseking/openclaw-exposure-dataset.BELKA-DEL-Experimental-BenchmarkData
BELKA-DEL-Experimental-BenchmarkData
This dataset comprises a curated collection of PDB structures, designed as an experimental validation benchmark for models trained on the Big Encoded Library for Chemical Assessment (BELKA) DNA-Encoded Library (DEL). Each structure includes at least one bound small molecule ligand, providing a robust basis for benchmarking model performance in accurately identifying potential binders to BELKA protein targets.
Introduction to the BELKA… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/BELKA-DEL-Experimental-BenchmarkData.CZE_supreme_administrative_court_decisions
Czech Supreme Administrative Court Decisions Dataset
This dataset contains decisions from the Supreme Administrative Court of the Czech Republic scraped from their official search interface.
Data Usage
This dataset can be utilized for:
Training language models on administrative law texts
Creating synthetic legal datasets
Building vector databases for Retrieval Augmented Generation (RAG)
Administrative law text analysis and research
NLP tasks focused on Czech… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_supreme_administrative_court_decisions.MMLU-Phrasing-Benchmark
MMLU Phrasing Benchmark
This dataset is a phrasing variant of cais/mmlu, put together by Roscommon Systems to see whether the way a question is worded affects how accurately language models answer it.
Each of the 2,650 questions appears four ways: the original text from MMLU, a polite version, a formal academic version, and an angry/demanding version. The answer choices and correct answers are identical to the source dataset in all cases.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/RoscommonSystems/MMLU-Phrasing-Benchmark.dlp-roster-3ef8ab13
Training Roster - DLP-Q3-2026-3ef8ab13
Published roster for the Q3 2026 Data Literacy Program.
Files:
roster.csv: employees to consider for enrollment.
policy.md: assignment metadata policy and business rules.
