datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YfOptionspxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.trek-finder
Trek Finder
A fully synthetic dataset of hiking routes and traveller reviews, built to power
a natural-language trek search: you describe the walk you want in your own
words, and the app finds routes that match.
Everything here was generated with Qwen/Qwen2.5-1.5B-Instruct. No dataset was
downloaded and no row was written by hand.
Built as a final project for an Introduction to Data Science course.
Files
file
rows
what it is
treks.csv
193
the route… See the full description on the dataset page: https://huggingface.co/datasets/Amitmelamed277/trek-finder.FeelAnyForce
Dataset Extraction Guide
This guide provides instructions to extract the dataset from multiple parts.
Steps to Extract the Dataset
Merge the dataset parts into a single archiveRun the following command to concatenate all parts into a single .zip file:
zip -s 0 dataset.zip --out merged.zip
Extract the dataset
unzip merged.zip
PutnamBenchLink to the repository on GitHub: https://github.com/trishullab/PUTNAM
PutnamBench
PutnamBench is a benchmark for the evaluation of theorem-proving algorithms on competition mathematics problems sourced from the William Lowell Putnam Mathematical Competition years 1965 - 2023. Our formalizations currently support three formal languages: Lean 4 $\land$ Isabelle $\land$ Coq. PutnamBench comprises over 1300 manual formalizations, aggregated over all languages.
PutnamBench aims to… See the full description on the dataset page: https://huggingface.co/datasets/amitayusht/PutnamBench.PersonaAsInfrastructure
Supplementary Material
Persona as Infrastructure: Invisible Structural Control in LLM-Mediated Social Networks
ICNLSP 2026.
This archive contains the complete data and code needed to reproduce every
number, table, and figure in the paper.
1. Contents
Networks (networks/)
109 generated networks in GraphML. Each file is one persona × one trial:
network_<persona>_trial<N>.graphml, with companion
_stats.json (summary metrics, including the exact model… See the full description on the dataset page: https://huggingface.co/datasets/amircincy/PersonaAsInfrastructure.Multi-Arabic-dialectskibahpc-chatbot-logssong_lyricsoptioncharts.ioTacred_Llamatacred_text_labelAMI
Dataset Card for AMI Corpus
Dataset Description
Links
Homepage: https://groups.inf.ed.ac.uk/ami/corpus/
Repository: https://groups.inf.ed.ac.uk/ami/download/
Paper: https://groups.inf.ed.ac.uk/ami/corpus/overview.shtml
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
The AMI Meeting Corpus is a multi-modal data set consisting of 100 hours of meeting recordings. For a gentle introduction to the corpus, see the corpus overview.… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/AMI.mrsd_ankle_exo_datalottie-urls
Dataset Card for [Dataset Name]
Dataset Summary
List of lottiefiles uri for research purposes
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/AmirulOm/lottie-urls.ontonotes5-persian
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Amir13/ontonotes5-persian.conll2003-persian
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Amir13/conll2003-persian.ncbi-persian
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Amir13/ncbi-persian.Financial-Fraud-Dataset
Dataset Card for Financial Fraud Labeled Dataset
Dataset Details
This dataset collects financial filings from various companies submitted to the U.S. Securities and Exchange Commission (SEC). The dataset consists of 85 companies involved in fraudulent cases and an equal number of companies not involved in fraudulent activities. The Fillings column includes information such as the company's MD&A, and financial statement over the years the company stated on the SEC… See the full description on the dataset page: https://huggingface.co/datasets/amitkedia/Financial-Fraud-Dataset.darooyab_qa
Dataset Description:
darooyab_qa is a Persian drug question-answering dataset extracted from Darooyab materials.
Using the LLama3 model, the scraped content of each drug page transformed into many questions and corresponding answers.
Load the dataset:
To load the dataset, install the library datasets with pip install datasets. Then,
from datasets import load_dataset
dataset = load_dataset("amirmmahdavikia/darooyab_qa")
class_scopeCAPTex
CAPTex: A Benchmark for Culturally-Aware Procedural Text Understanding
Amir Hossein Yari, Fajri Koto
Sharif University of Technology, MBZUAI
Introduction
CAPTex (Culturally-Aware Procedural Texts) is a dataset designed to evaluate the ability of multilingual large language models (mLLMs) to comprehend and reason about procedural texts embedded in diverse cultural contexts. The dataset includes procedural knowledge from seven culturally distinct regions: China, India… See the full description on the dataset page: https://huggingface.co/datasets/AmirHossein2002/CAPTex.cleverThis repository contains the data for the paper CLEVER: A Curated Benchmark for Formally Verified Code Generation.
The benchmark can be found on GitHub: https://github.com/trishullab/clever
ITEM
ITEM: Indian Text Evaluation Metrics Testbed
Introduction
The ITEM dataset is designed to evaluate how well various automatic evaluation metrics align with human judgments for machine translation and text summarization in six major Indian languages.
Statistics 📊
Task
Total Samples
Machine Translation
2,604
Text Summarization
2,571
Languages 🌍
Hindi
Bengali
Tamil
Telugu
Gujarati
Marathi
Licence 📜
As a… See the full description on the dataset page: https://huggingface.co/datasets/AmirHossein2002/ITEM.loanbindingdb_kdimbalancedMedical_Transcription_LLMGroceryList
GroceryList Dataset
Dataset Summary
The GroceryList dataset consists of grocery items and their corresponding categories. It is designed to assist in tasks such as grocery item classification, shopping list organization, and natural language understanding related to common grocery-related terms. The dataset contains only a training split and is not pre-divided into test or validation sets.
It includes two main columns:
Item: Contains the names of various grocery items… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/GroceryList.
