CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RosettaCommons /SAAINTDB SAAINTDB This dataset is a curated version of the SAAINT-DB converted into a format compatible with the Hugging Face Datasets for machine learning applications. The dataset contains 21,400 antibody entries derived from 11,304 PDB structures, reflecting the available structures as of February 2026. Each entry corresponds to an antibody chain and is uniquely identified using the PDB_ID_chain field (PDB ID + chain ID). Dataset Splits The dataset was split at the PDB… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAAINTDB.tabular10K<n<100K0 likes397 downloads6mo agoHugging Face02RosettaCommons /SAbDab ML Application Curated SAbDab Quickstart Usage Install HuggingFace Datasets package Each subset can be loaded into python using the Huggingface datasets library. First, from the command line install the datasets library $ pip install datasets Optionally set the cache directory, e.g. $ HF_HOME=${HOME}/.cache/huggingface/ $ export HF_HOME then, from within python load the datasets library >>> import datasets Load model datasets To load… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/SAbDab.tabular10K<n<100K1 likes346 downloads6mo agoHugging Face03Rose-STL-Lab /ClimaQA ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025) Check the paper's webpage and GitHub for more info! The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.textquestion-answering1K<n<10K3 likes196 downloads2y agoHugging Face04Rosalia1212 /cbis-ddsm-r CBIS-DDSM-R: A Curated Radiomic Feature Dataset for Breast Cancer Classification Dataset Summary CBIS-DDSM-R is an open-source, radiomics-ready extension of the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM). It is designed to facilitate reproducible radiomics and quantitative imaging research in breast cancer analysis. The dataset provides a standardized preprocessing pipeline for mammograms and includes IBSI-compliant… See the full description on the dataset page: https://huggingface.co/datasets/Rosalia1212/cbis-ddsm-r.tabularimage-classification1K<n<10K0 likes175 downloads4mo agoHugging Face05WilliamQiu123 /ROSEtabular10K<n<100K0 likes170 downloads7mo agoHugging Face06RosettaCommons /PISCES-CulledPDB PISCES-CulledPDB database as of January 2026 Recurated on Hugging Face on March 5th 2026 The PISCES dataset provides curated sets of protein sequences from the Protein Data Bank (PDB) based on sequence identity and structural quality criteria. PISCES yields non-redundant subsets of protein chains by applying filters such as sequence identity, experimental resolution, R-factor, chain length, and experimental method (e.g., X-ray, NMR, cryo-EM). The goal is to maximize structural… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PISCES-CulledPDB.textother1M<n<10M0 likes148 downloads6mo agoHugging Face07RosettaCommons /AfCycDesign Dataset Card for AfCycDesign Hallucinated scaffolds used by AfCycDesign for cyclic peptide design. Dataset Details Sets 7-16 of hallucinated peptide cif files and experimental CCDC structures. Dataset Description This dataset contains hallucinated cyclic peptide scaffold structures (in CIF format) generated using AfCycDesign, a deep learning approach built on AlphaFold2 for de novo design of cyclic peptides. The scaffolds span peptide lengths of 7–16 residues… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/AfCycDesign.tabular10K<n<100K0 likes146 downloads6mo agoHugging Face08Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes124 downloads9mo agoHugging Face09RosettaCommons /FPbasegated FPbase: The Fluorescent Protein Database FPbase is a free, open-source, community-editable database of fluorescent proteins and their properties, aimed at aggregating structured, searchable information useful to the imaging community and FP developers. Visit fpbase.org for more. This dataset updated on ,March 1st, 2026, collects FPbase fluorescent protein records (e.g., names, identifiers, sequences, and photophysical properties) for downstream analysis and modeling.… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FPbase.tabular1K<n<10K0 likes108 downloads6mo agoHugging Face10roslein /EU_Court_Human_Rights_Decisions European Court of Human Rights Decisions Dataset This dataset contains 9,820 decisions from the European Court of Human Rights (ECHR) scraped from HUDOC, the official database of ECHR case law. Data Usage This dataset is valuable for: Building legal vector databases for RAG (Retrieval Augmented Generation) Training Large Language Models focused on human rights law Creating synthetic legal datasets Legal text analysis and research NLP tasks in international human rights… See the full description on the dataset page: https://huggingface.co/datasets/roslein/EU_Court_Human_Rights_Decisions.text1K<n<10K1 likes95 downloads2y agoHugging Face11roskosmos19 /agentic-reasoning-benchmark Agentic & Reasoning Benchmark (ARB) – Expanded Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning. Überblick Eigenschaft Wert Anzahl Beispiele 2.550 Kategorien 8 Schwierigkeitsgrade easy / medium / hard Formate CSV + JSON Reproduzierbarkeit Generator-Skript (seed=42) enthalten Lizenz CC-BY-4.0 Kategorien Kategorie Anzahl Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.textquestion-answering1K<n<10K1 likes92 downloads19d agoHugging Face12pytorch-lifestream /rosbank-churnhttps://boosters.pro/championship/rosbank1/ tabulartabular-classification1M<n<10M0 likes83 downloads3y agoHugging Face13readerbench /ro-storiesThe corpus consists of texts written by Romanian authors between 19th century and present, representing stories, short-stories, fairy tales and sketches. The current version contains 19 authors, 1263 full texts and 12516 paragraphs of around 200 words each, preserving paragraphs integrity. Note: This is an extended version of ROST corpus (https://www.kaggle.com/datasets/sandamariaavram/rost-romanian-stories-and-other-texts), which only contains 400 texts and 10 authors. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/readerbench/ro-stories.text10K<n<100K4 likes80 downloads3y agoHugging Face14RosettaCommons /NAKBOriginal Paper: Lawson CL, Berman HM, Vallat B, Chen L, Zirbel C (2024) The Nucleic Acid Knowledgebase: a new portal for 3D structural information about nucleic acids. Nucleic Acids Research 52, D245-D254. https://doi.org/10.1093/nar/gkad957 Nucleic Acid Knowledgebase (NAKB) NAKB data set contains 21166 structures including Nucleic Acids, Protein, and Ligand Annotations, and determined 3D structures found in the Nucleic Acid Database (NDB) and the Protein Data Bank (PDB), including… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/NAKB.tabular100K<n<1M0 likes80 downloads6mo agoHugging Face15RosettaCommons /PTMint PTMint This dataset is derived from PTMint (https://ptmint.sjtu.edu.cn/), (Post Translational Modifications that are associated with Protein-Protein Interactions) that contains manually curated complete experimental evidence of the PTM effecting on protein-protein interactions in multiple organisms, including H. sapines, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae and S. pombe. This Hugging Face dataset repository provides PTMint-derived tables including a precomputed… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PTMint.tabular1K<n<10K1 likes80 downloads6mo agoHugging Face16almaz-nlp /almaz-asr-roster The ALMAZ ASR Roster Archived at Zenodo: 10.5281/zenodo.22761871 (concept DOI, always resolves to the latest version). A curated catalog of Azerbaijani speech-to-text artifacts: corpora, models, services, benchmarks and tools. Companion to the ALMAZ Resource Roster, which does the same for text. Schema matches the text roster so the two join, plus three columns speech needs and text does not: hours, condition, and verified. The verified column The standard way a… See the full description on the dataset page: https://huggingface.co/datasets/almaz-nlp/almaz-asr-roster.tabularn<1K0 likes50 downloads8d agoHugging Face17cmaldona /Generalization-MultiClass-CLINC150-ROSTDThis dataset merge 3 datasets and have two setup for experiments in generalisation for multi-class clasificacitino task. ID, near-OOD, covariate-shitf: CLINC150 ID, near-OOD, covariate-shitf: ROSTD+OOD (fbreleasecoarse version) far-OOD Validation: SST2 far-OOD Test: News Category (v3) texttext-classification10K<n<100K1 likes46 downloads3y agoHugging Face18RosettaCommons /CatPred-DBgated CatPred-DB: Enzyme Kinetic Parameters Database Paper: CatPred: A comprehensive framework for deep learning in vitro enzyme kinetic parameters GitHub: https://github.com/maranasgroup/CatPred-DB Dataset Description CatPred-DB contains the benchmark datasets introduced alongside the CatPred deep learning framework for predicting in vitro enzyme kinetic parameters. The datasets cover three key kinetic parameters: Parameter Description Datapoints kcat Turnover… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/CatPred-DB.tabular10K<n<100K0 likes42 downloads6mo agoHugging Face19roslein /CZE_constitutional_court_decisions Czech Constitutional Court Decisions Dataset This dataset contains decisions from the Constitutional Court of the Czech Republic scraped from NALUS, the official database of Constitutional Court decisions. Data Usage This dataset can be utilized for: Training language models on legal texts Creating synthetic legal datasets Building vector databases for Retrieval Augmented Generation (RAG) Legal text analysis and research NLP tasks focused on Czech legal domain… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_constitutional_court_decisions.text100K<n<1M0 likes39 downloads1y agoHugging Face20roslein /Czech_legal_code Czech Legal Codes Dataset Dataset Description This dataset contains full texts of valid Czech legal codes (laws) as of February 21, 2025. The dataset focuses exclusively on primary legislation (laws) and excludes secondary legislation such as regulations and ordinances. Data Source The data was scraped from the official Czech Collection of Laws (Sbírka zákonů) portal: e-sbirka.cz Data Format The dataset is provided as a CSV file with two columns:… See the full description on the dataset page: https://huggingface.co/datasets/roslein/Czech_legal_code.text1K<n<10K0 likes39 downloads2y agoHugging Face21roslein /Legal_advice_czech Legal Advice Dataset Dataset Description This dataset contains scraped legal questions (but also a simply informational content without any question present) and answers from Bezplatná Právní Poradna. The data consists of legal inquiries submitted by users and expert responses provided on the website. It is structured for ease of use in natural language processing (NLP) tasks related to legal text classification, question-answering models, and text summarization. However… See the full description on the dataset page: https://huggingface.co/datasets/roslein/Legal_advice_czech.textquestion-answering10K<n<100K0 likes32 downloads2y agoHugging Face22RosettaCommons /2J-Protein-Couplings2J-Protein-Coupling Dataset This data set was curated from the paper below accessed through the Biological Magnetic Resonance Data Bank (BMRB). There are a total of 3999 2J coupling taken from 5 different proteins and up to 10 different experiments. This dataset contains information regarding PDB ID, Sequence, 2J coupling data of 15N, 13C, and 1H. Data was curated and organized into this set of the five papers below, with the addition of the sequence taken from the Protein Data Bank. Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/2J-Protein-Couplings.tabularn<1K0 likes32 downloads6mo agoHugging Face23BlackKakapo /RoSTSC RO-STS-Cupidon Overview RoSTSC is a Romanian Semantic Textual Similarity (STS) dataset designed for evaluating and training sentence embedding models. It contains pairs of Romanian sentences along with similarity scores that indicate the degree of semantic equivalence between them. Dataset Structure sentence1: The first sentence in the pair. sentence2: The second sentence in the pair. score: A numerical value representing the semantic similarity between the… See the full description on the dataset page: https://huggingface.co/datasets/BlackKakapo/RoSTSC.textsentence-similarity10K<n<100K5 likes30 downloads1y agoHugging Face24ClarusC64 /healthcare-staff-roster-demand-coherence-risk-v0.1What this repo is for Detect staffing breakdown before patient care degrades. Tracks alignment between demand, roster levels, skill mix, and coverage. Helps hospitals prevent unsafe staffing, burnout spikes, and operational collapse. texttext-classificationn<1K0 likes30 downloads7mo agoHugging Face25roslein /CZE_Supreme_Court_Decision Czech Supreme Court Decisions Dataset This dataset contains decisions from the Supreme Court of the Czech Republic scraped from their official collection database. Data Usage This dataset is ideal for: Building legal vector databases for RAG (Retrieval Augmented Generation) Training language models on Czech civil and criminal law Creating synthetic legal datasets Legal text analysis and research NLP tasks focused on Czech judicial domain Legal Status The… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_Supreme_Court_Decision.text1K<n<10K0 likes27 downloads1y agoHugging Face26roseking /openclaw-exposure-dataset🇨🇳 查看中文版 README ⚠️ Unofficial Dataset — This dataset is NOT affiliated with, endorsed by, or officially released by openclaw.allegro.earth. It is an independent research snapshot. OpenClaw Exposure Watchboard Dataset (Sanitized) 📌 Data Source & Attribution Original Source: OpenClaw Exposure Watchboard All original data is collected and maintained by openclaw.allegro.earth. Full credit and ownership of the source data belong to the original maintainers. This… See the full description on the dataset page: https://huggingface.co/datasets/roseking/openclaw-exposure-dataset.tabulartabular-classification100K<n<1M1 likes25 downloads7mo agoHugging Face27RosettaCommons /BELKA-DEL-Experimental-BenchmarkData BELKA-DEL-Experimental-BenchmarkData This dataset comprises a curated collection of PDB structures, designed as an experimental validation benchmark for models trained on the Big Encoded Library for Chemical Assessment (BELKA) DNA-Encoded Library (DEL). Each structure includes at least one bound small molecule ligand, providing a robust basis for benchmarking model performance in accurately identifying potential binders to BELKA protein targets. Introduction to the BELKA… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/BELKA-DEL-Experimental-BenchmarkData.tabular10K<n<100K0 likes24 downloads6mo agoHugging Face28roslein /CZE_supreme_administrative_court_decisions Czech Supreme Administrative Court Decisions Dataset This dataset contains decisions from the Supreme Administrative Court of the Czech Republic scraped from their official search interface. Data Usage This dataset can be utilized for: Training language models on administrative law texts Creating synthetic legal datasets Building vector databases for Retrieval Augmented Generation (RAG) Administrative law text analysis and research NLP tasks focused on Czech… See the full description on the dataset page: https://huggingface.co/datasets/roslein/CZE_supreme_administrative_court_decisions.text1K<n<10K0 likes22 downloads1y agoHugging Face29RoscommonSystems /MMLU-Phrasing-Benchmark MMLU Phrasing Benchmark This dataset is a phrasing variant of cais/mmlu, put together by Roscommon Systems to see whether the way a question is worded affects how accurately language models answer it. Each of the 2,650 questions appears four ways: the original text from MMLU, a polite version, a formal academic version, and an angry/demanding version. The answer choices and correct answers are identical to the source dataset in all cases. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/RoscommonSystems/MMLU-Phrasing-Benchmark.textquestion-answering100K<n<1M1 likes19 downloads2mo agoHugging Face30Roy229 /dlp-roster-3ef8ab13 Training Roster - DLP-Q3-2026-3ef8ab13 Published roster for the Q3 2026 Data Literacy Program. Files: roster.csv: employees to consider for enrollment. policy.md: assignment metadata policy and business rules. textn<1K0 likes19 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.