CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01christopher /rosetta-code Dataset Card for the Rosetta Code Dataset Dataset Summary Rosetta Code is a programming chrestomathy site. The idea is to present solutions to the same task in as many different languages as possible, to demonstrate how languages are similar and different, and to aid a person with a grounding in one approach to a problem in learning another. Rosetta Code currently has 1,203 tasks, 389 draft tasks, and is aware of 883 languages, though we do not (and cannot) have… See the full description on the dataset page: https://huggingface.co/datasets/christopher/rosetta-code.text10K<n<100K40 likes1.9k downloads3y agoHugging Face02RosettaCommons /MIP Microbiome Immunity Project: Protein Universe ~200,000 predicted structures for diverse protein sequences from 1,003 representative genomes across the microbial tree of life and annotate them functionally on a per-residue basis. Quickstart Usage Install HuggingFace Datasets package Each subset can be loaded into python using the Huggingface datasets library. First, from the command line install the datasets library $ pip install datasets Optionally set the… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MIP.tabular1B<n<10B1 likes1.3k downloads2y agoHugging Face03RosettaCommons /MegaScale Mega-scale experimental analysis of protein folding stability in biology and design The full MegaScale dataset contains 1,841,285 thermodynamic folding stability measurements using cDNA display proteolysis of natural and designed proteins. From these 776,298 high-quality folding stabilities (dataset2) cover all single amino acid variants and selected double mutants of 331 natural and 148 de novo designed protein domains 40–72 amino acids in length. Of these mutations, 607,839 have… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MegaScale.tabular1M<n<10M5 likes480 downloads2y agoHugging Face04juliensimon /esa-rosetta-observations ESA Rosetta Observations Credit: NASA/ESA Part of the Solar System Datasets and Planetary Science Datasets collections on Hugging Face. Complete observation metadata catalog from the ESA Rosetta mission to Comet 67P/Churyumov-Gerasimenko — 8,214,033 observations across 15 instruments. Dataset description Rosetta was ESA's groundbreaking mission to Comet 67P/Churyumov-Gerasimenko. Launched in 2004, it became the first spacecraft to orbit a comet (August 2014) and… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/esa-rosetta-observations.tabulartabular-classification10M<n<100M0 likes305 downloads4mo agoHugging Face05RosettaCommons /FireProtDB2 Dataset Card for FireProtDB_2.0 Subsets of protein stability data for single-point mutants from FireProtDB, a comprehensive curated database. Dataset Details Subsets of different thermal data of single-point mutations in the FireProtDB database with train/validation/test splits: ΔG, ΔΔG Tm, ΔTm Fitness Stabilizing Dataset Description This dataset contains curated subsets of various thermal stability measurements derived from FireProtDB. Subsets… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FireProtDB2.tabular1M<n<10M1 likes161 downloads6mo agoHugging Face06Asap7772 /leetcode-rosetta-processed-with-test-casestext1K<n<10K5 likes75 downloads2y agoHugging Face07RosettaCommons /AbAgym AbAgym AbAgym is a curated dataset of deep mutational scanning (DMS) measurements for antibody-antigen complexes. This Hugging Face version reorganizes the original AbAgym files into loadable dataset configurations using Apache Parquet, while preserving the original structure archive. The original AbAgym repository describes the dataset as containing 68 DMS datasets on antibody-antigen complexes, approximately 324,000 non-redundant mutations, 36,541 non-redundant interface mutations… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/AbAgym.tabulartabular-classification100K<n<1M1 likes55 downloads5mo agoHugging Face08RosettaCommons /PRIDE_Crosslinking_Archive PRIDE Crosslinking Archive This dataset aggregates publicly available crosslinking mass spectrometry (XL-MS) datasets from the PRIDE repository. Each dataset is curated and categorized by crosslinking reagent a link type (inter-chain vs intra-chain). For intra-chain links where the protein can be mapped to a UniProt ID, each link is mapped onto the corresponding AlphaFold Database (AFDB) structure, and the Cα-Cα distance for the linked residue pair is reported. The result is a… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/PRIDE_Crosslinking_Archive.tabular100K<n<1M0 likes54 downloads6mo agoHugging Face09RosettaCommons /UTexasAptamer UT Aptamer Dataset This is a collection of 1480 aptamer sequences from the University of Texas Aptamer Database as of 2023. This dataset is split into three subsets (train, test, and validation) based on clustering by CD-HIT. Clustering Clustering was conducting using the CD-HIT: Cluster Database at High Identity with Tolerance web browser using a 40% sequence identity threshold and word size of 2. To update this dataset with new reported aptamers, splits can be… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/UTexasAptamer.tabular1K<n<10K0 likes50 downloads6mo agoHugging Face10pharaouk /rosetta-code Dataset Card for the Rosetta Code Dataset Dataset Summary Rosetta Code is a programming chrestomathy site. The idea is to present solutions to the same task in as many different languages as possible, to demonstrate how languages are similar and different, and to aid a person with a grounding in one approach to a problem in learning another. Rosetta Code currently has 1,203 tasks, 389 draft tasks, and is aware of 883 languages, though we do not (and cannot) have… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/rosetta-code.text10K<n<100K0 likes46 downloads2y agoHugging Face11jv-costa /dataset-rosetta-testtabular10K<n<100K0 likes44 downloads13d agoHugging Face12namanbnsl /rosettabench-150-stratified-compressed About Dataset Why? Used for RosettaBench (A contamination-free benchmark for measuring learning, not memorization). Preparation Source: LiveCodeBench (release_v5, AtCoder platform only), accessed via sam-paech/livecodebench-code_generation_lite on Hugging Face. AtCoder problems were selected for platform consistency and STDIN/STDOUT format compatibility. Problems with fewer than 3 test cases were excluded, leaving a pool of 342 problems. Sampling: 150… See the full description on the dataset page: https://huggingface.co/datasets/namanbnsl/rosettabench-150-stratified-compressed.texttext-generationn<1K2 likes34 downloads5mo agoHugging Face13Asap7772 /leetcode-rosetta-processedtext1K<n<10K1 likes25 downloads2y agoHugging Face14juyoungml /leetcode-rosettatext1K<n<10K0 likes21 downloads2y agoHugging Face15Asap7772 /leetcode-rosettatext1K<n<10K0 likes20 downloads2y agoHugging Face16Asap7772 /leetcode-rosetta-processed-recollected-with-test-casestext1K<n<10K0 likes15 downloads2y agoHugging Face17matthewRoloff /rosetta_code_verilog Compilation of all the Verilog scripts from rosettacode.org. The license for these scripts was found here textn<1K0 likes14 downloads2y agoHugging Face18mlfoundations-dev /a1_code_rosettatext10K<n<100K0 likes12 downloads1y agoHugging Face19mugezhang /russian-xnli-ipa-rosetta_ipa_romanizedtext100K<n<1M0 likes12 downloads8mo agoHugging Face20mlfoundations-dev /rosettatext10K<n<100K0 likes11 downloads2y agoHugging Face21mlfoundations-dev /load_in_code_rosettatext10K<n<100K0 likes10 downloads1y agoHugging Face22ezio007 /rosetta-trialstabularn<1K0 likes9 downloads9mo agoHugging Face23mlfoundations-dev /a1_code_rosetta_eval_636d mlfoundations-dev/a1_code_rosetta_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 15.7 58.5 74.6 28.6 42.2 44.4 29.8 5.5 8.8 AIME24 Average Accuracy: 15.67% ± 0.82% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 20.00% 6 30 2 13.33% 4 30 3 13.33% 4 30 4 13.33% 4 30 5… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_rosetta_eval_636d.tabular1K<n<10K0 likes8 downloads1y agoHugging Face24ezio007 /rosetta_sampled_non-tokenizedtextn<1K0 likes6 downloads9mo agoHugging Face25mlfoundations-dev /b2_calc_positive_embeddings_code_leetcode_rosettatext1K<n<10K0 likes5 downloads1y agoHugging Face26iggy12345 /russian-xnli-ipa-rosettatext100K<n<1M0 likes5 downloads1y agoHugging Face27RosettaCommons /MAHOMES_IIgated MAHOMES II MAHOMES-II (Metal Activity Heuristic of Metalloprotein and Enzymatic Sites-II) is a structure-based dataset for classifying protein-bound metal sites as enzymatic or non-enzymatic. ('Enzyme' column) Dataset Splits The train/test split follows a temporal split based on PDB deposition date. Structures that were deposited before 2018 were included in the training dataset, and structures deposited in 2018 or later went into the holdout test set. The train dataset… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MAHOMES_II.tabular10K<n<100K0 likes5 downloads6mo agoHugging Face28RosettaCommons /FoldDockgated FoldDock A collection of 219 heterodimers from dockground benchmark 4 dataset, and 1503 heterodimeric structures from a recent study (Green, A. G. et al. Nat. Commun. 12, 1–12 (2021)) (dubbed “marks”). The benchmark dataset was used to train an original model, that model was used to predict the structures in the second dataset. This dataset is split into three subsets, each with two splits corresponding to the groups used in the source (marks, dockground). Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/FoldDock.tabular1K<n<10K0 likes4 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.