datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cameo_data
CAMEO Dataset for Protein Structure Prediction
This dataset contains protein sequences and structures from CAMEO (Continuous Automated Model EvaluatiOn) for monomer structure prediction tasks.
Dataset Description
CAMEO is a community-wide initiative to continuously evaluate the performance of protein structure prediction methods. This dataset includes 141 protein targets collected from January to April 2025.
Dataset Structure
cameo/
├── sequences.fasta… See the full description on the dataset page: https://huggingface.co/datasets/THU-ATOM/cameo_data.AtomisticEvalATO-datasets
ATO: Augmenting Text to Increase Translation Difficulty — datasets
William Kalikman*, Šimon Sukup*, Michal Tešnar, Vilém Zouhar (*equal contribution)
ETH Zurich
This dataset contains the (original, augmented) text pairs produced by the two ATO variants. The method, code, and configs can be seen in the repository: https://github.com/BreakingMT/ATO
Columns (both files)
original_sentence — source English sentence (FLORES200 or WMT22/23/24 news test set)… See the full description on the dataset page: https://huggingface.co/datasets/wskal/ATO-datasets.comet-atomic-ja-zhatomic-scraper-leads
Atomic Scraper Leads
Public business listings scraped from Google Maps for leads.benjaminboyce.com.
This dataset is an export of the leads table from the Atomic / gmaps-scraper-suite pipeline. Each row is a local business listing with contact and location fields plus scrape metadata.
Freshness
Use scraped_at (UTC timestamps) as the freshness signal. This snapshot was exported on 2026-08-25.
Oldest scraped_at: 2026-08-09
Newest scraped_at: 2026-08-25
Listings… See the full description on the dataset page: https://huggingface.co/datasets/bensblueprints/atomic-scraper-leads.emotion-prediction-comet-atomic-2020
emotion-prediction-comet-atomic-2020
This dataset extends the COMET-Atomic-2020 commonsense reasoning dataset by focusing on the xReact (subject’s emotional reaction) and oReact (other person’s emotional reaction) relations.
Description
Each entry is expanded into a realistic, three‑sentence scenario, replacing the placeholders (PersonX/PersonY) with human names and adding contextual details.
Source (comet-atomic-2020) Example:
source
relation
target
PersonX… See the full description on the dataset page: https://huggingface.co/datasets/id4thomas/emotion-prediction-comet-atomic-2020.RecentPDBDirectory structure of recentpdb.tar.gz
/data/protein/YHData/Data/airdd/datasets/recentpdb
├── input.json
├── mmcif
│ ├── 6st5.cif
│ ├── 6zsu.cif
│ ├── 6zsv.cif
│ ├── ···
│ ├── 8hii.cif
│ ├── 8hij.cif
│ └── 8hik.cif
├── pdb_id_list.txt
└── precomputed_msa
shit-nlp-sentimentRAGPPI_Atomics
RAG Benchmark for Protein-Protein Interactions (RAGPPI)
📊 Overview
Retrieving expected therapeutic impacts in protein-protein interactions (PPIs) is crucial in drug development, enabling researchers to prioritize promising targets and improve success rates. While Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) frameworks accelerate discovery, no benchmark exists for identifying therapeutic impacts in PPIs.
RAGPPI is the first factual QA benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Youngseung/RAGPPI_Atomics.atom-explanations-middle-schooltactical_atomic_multiple_change_dataset
