datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YfOptionspxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.trek-finder
Trek Finder
A fully synthetic dataset of hiking routes and traveller reviews, built to power
a natural-language trek search: you describe the walk you want in your own
words, and the app finds routes that match.
Everything here was generated with Qwen/Qwen2.5-1.5B-Instruct. No dataset was
downloaded and no row was written by hand.
Built as a final project for an Introduction to Data Science course.
Files
file
rows
what it is
treks.csv
193
the route… See the full description on the dataset page: https://huggingface.co/datasets/Amitmelamed277/trek-finder.PersonaAsInfrastructure
Supplementary Material
Persona as Infrastructure: Invisible Structural Control in LLM-Mediated Social Networks
ICNLSP 2026.
This archive contains the complete data and code needed to reproduce every
number, table, and figure in the paper.
1. Contents
Networks (networks/)
109 generated networks in GraphML. Each file is one persona × one trial:
network_<persona>_trial<N>.graphml, with companion
_stats.json (summary metrics, including the exact model… See the full description on the dataset page: https://huggingface.co/datasets/amircincy/PersonaAsInfrastructure.optioncharts.iomrsd_ankle_exo_datacleverThis repository contains the data for the paper CLEVER: A Curated Benchmark for Formally Verified Code Generation.
The benchmark can be found on GitHub: https://github.com/trishullab/clever
ITEM
ITEM: Indian Text Evaluation Metrics Testbed
Introduction
The ITEM dataset is designed to evaluate how well various automatic evaluation metrics align with human judgments for machine translation and text summarization in six major Indian languages.
Statistics 📊
Task
Total Samples
Machine Translation
2,604
Text Summarization
2,571
Languages 🌍
Hindi
Bengali
Tamil
Telugu
Gujarati
Marathi
Licence 📜
As a… See the full description on the dataset page: https://huggingface.co/datasets/AmirHossein2002/ITEM.loanbindingdb_kdimbalancedVocaDBSongsThis is dataset created using data fetched from public VocaDB API.
You can find source code here: https://github.com/amiadesu/VocaDBScraper
usda-nutrition-eda
🥗 USDA FoodData Central – Nutrition EDA
Overview
This project presents an end-to-end Exploratory Data Analysis (EDA) of nutritional data
from the USDA FoodData Central database. The goal is to uncover patterns in food nutrition,
compare food categories, and explore relationships between key nutritional features.
Source: omid5/usda-fdc-foods-cleaned
Original size: 501,887 rows × 22 columns
Final clean size: 345,226 rows × 11 numeric features
Target Variable: data_type… See the full description on the dataset page: https://huggingface.co/datasets/amitbenavraham/usda-nutrition-eda.davisAMIE_PSEAE_Wrenbeck_2017_exampleclintoxiranian-churn-datasetEnglishTextXSSTourism-Package-Predictiononc_amie
Oncology AMIE Dataset
Dataset Description
This dataset contains detailed clinical cases of breast cancer patients, including their age, molecular phenotype (ER/PR/HER2 status), treatment status, special considerations, and ground truth treatment recommendations. The cases are divided into treatment-naive and treatment-refractory categories.
Dataset Summary
A clinical dataset containing breast cancer cases with their descriptions, molecular phenotypes, and… See the full description on the dataset page: https://huggingface.co/datasets/gallifantjack/onc_amie.bbbpPantip_QA_200000_20220220Cancerdataset_emb_faqtourism-datasetheartdiseaseBreast_cancer_detectiontox21ami_2020
Dataset Card for Dataset Name
The Automatic Misogyny Identification was part of the 7th evaluation campaign EVALITA 2020. Please refer to the official website for more details.
Data is also available on the ELG website.
Data is released under the CC-BY NC SA 4.0
The rest of the dataset card is WIP :)
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/ami_2020.
