datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
waqfeya-library
Waqfeya Library
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.shamela-waqfeya-library
Shamela Waqfeya Library
📖 Overview
Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 12,877 PDF files (spanning 5,138,027 pages) representing 4,661 Islamic books.… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library.HQ-OpenHumanVidopen-asr-leaderboard-resultsreal-or-fake-fake-jobposting-predictionnovae
Description
Full novae dataset, including:
All the spatial transcriptomics samples used to train Novae
Protein samples used in the article
Some Visium and Visium HD samples
Synthetic data samples
You can download this dataset from the API, see novae.load_dataset
See here the list of available models trained on this dataset.
[!NOTE]
Note that Novae was trained on the image-based spatial transcriptomics samples. This means that it was not trained on the Visium/VisiumHD samples… See the full description on the dataset page: https://huggingface.co/datasets/prism-oncology/novae.jepa-qwen3-32b-pure-baselines-2026-05-25
JEPA-Align: Qwen3-32B Safety Defense Matrix
The complete 11-condition Qwen3-32B experiment for Predictive Representation
Alignment (PRA), the paired-view objective introduced in Predictive
Representation Alignment Improves Generalization in LLM Safety.
PRA aligns adversarially rewritten prompts with clean prompts expressing the
same intent. This release contains trained adapters, attack traces, benign
capability evaluations, machine-readable results, and paper-ready tables for… See the full description on the dataset page: https://huggingface.co/datasets/memo-ozdincer/jepa-qwen3-32b-pure-baselines-2026-05-25.TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/TravelPlanner.cyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.TC-SSA
TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning
Links: Project homepage | arXiv paper | Code
Authors: Zhuo Chen1,2, Xiaoyu Yang1, and Lijian Xu1,*
1 Shenzhen University of Advanced Technology, Shenzhen, Guangdong, China2 University of Nottingham Ningbo China, FoSE, Ningbo, Zhejiang, China* Corresponding author: xulijian@suat-sz.edu.cn
TC-SSA WSI Feature Bags
This public repository contains pre-extracted whole-slide… See the full description on the dataset page: https://huggingface.co/datasets/OzzyChen97/TC-SSA.Objaverse-XL-Rigged-Animated
Objaverse-XL Rigged & Animated Subset
Every asset here carries both a skeleton and at least one animation clip, selected from
Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and
span characters as well as articulated rigid objects.
Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and
motion on it. This subset isolates that fraction: every file was checked to contain at least one
skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated.TIME-OutputThis repository contains the extracted time series features (tsfeatures) for each variate and the detailed forecasting results for every experiment.
Note: These files are for building leaderboard and visualization; users do not need to download this directory.
features/: Statistical Features (tsfeatures)
Each dataset's features are saved to: output/features/{dataset}/{freq}/.
This directory stores the computed tsfeatures for the variates in the dataset. The folder contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Real-TSF/TIME-Output.Raon-OpenTTS-Eval
Raon-OpenTTS-Eval
Technical Report
A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs.
Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.waqfeya-library-compressed
Waqfeya Library - Compressed
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/waqfeya-library, with one key difference: the contents… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library-compressed.fsrs-datasetopenadmet-expansionrx-challenge-data
OpenADMET-ExpansionRx Challenge FULL dataset
This is the full dataset used in the OpenADMET-ExpansionRx blind challenge, which finalized in January 19th, 2026.
Originally split in a train and blinded test set, we now release the full dataset, which contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases.
While optimising candidate molecules for their preclinical programs Expansion collected… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-data.Med-HALT
Med-HALT: Medical Domain Hallucination Test for Large Language Models
This is a dataset used in the Med-HALT research paper. This research paper focuses on the challenges posed by hallucinations in large language models (LLMs), particularly in the context of the medical domain. We propose a new benchmark and dataset, Med-HALT (Medical Domain Hallucination Test), designed specifically to evaluate hallucinations.
Med-HALT provides a diverse multinational dataset derived from medical… See the full description on the dataset page: https://huggingface.co/datasets/openlifescienceai/Med-HALT.pxr-challenge-train-test
PXR Challenge Train/Test Dataset
A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge.
Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction
Challenge Space: openadmet/pxr-challenge
Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.open-hdri-1k
Open HDRI 1K
A consolidated, public-domain (CC0-1.0) collection of 3,491 equirectangular HDR environment maps at 1K resolution, gathered from five free HDRI libraries: Poly Haven, BlenderKit, ambientCG, CGEES and Open HDRI.
Every map is stored as a linear, high-dynamic-range .exr file alongside a tonemapped .jpg preview, with a per-asset metadata row (dimensions, source, author, license, SHA-256 checksum and tags).
Contents
Source
Assets
Author(s)
License… See the full description on the dataset page: https://huggingface.co/datasets/leodriesch/open-hdri-1k.genebench-pro-public-package
GeneBench-Pro Public Case Studies
This repository contains public GeneBench-Pro case studies. It is the
self-contained package intended for public distribution, including Hugging Face
publication.
Package Layout
<repo-root>/
├── .gitattributes
├── README.md
├── LICENSE
├── problems.csv
├── checksums.sha256
├── manifest.json
├── reference_definitions.md
├── reference_grader.py
└── problems/
└── <eval_id>/
├── eval_config.json
├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/openai/genebench-pro-public-package.cryptocurrency-futures-ohlcv-dataset-1mCortexJEPAData
CortexJEPAData
Minimal cortex spatial transcriptomics H5AD files prepared for CortexJEPA.
Data Contents
Each .h5ad file keeps:
X: expression matrix
obs_names: spot/cell identifiers
var_names: gene identifiers
obsm["spatial"]: spatial coordinates
fine-tuning labels only for selected training splits:
c.macaque/pretrain: obs["layer"]
d.marmoset/pretrain: obs["layer"] and obs["PrAl"]
developmental-stage grouping for d.marmoset/test/development: obs["segment"] only… See the full description on the dataset page: https://huggingface.co/datasets/BGI-Hangzhou-OmicsAI/CortexJEPAData.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.twitter-airline-sentiment
Dataset Card for Twitter US Airline Sentiment
Dataset Summary
This data originally came from Crowdflower's Data for Everyone library.
As the original source says,
A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons (such as "late flight" or "rude service").
The data we're… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/twitter-airline-sentiment.openadmet-expansionrx-challenge-train-data
OpenADMET-ExpansionRx Challenge training dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.crunchbaseKrea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.open-command
OpenCommand
Catcher targets and pitcher command for MLB, inferred from broadcast video.
OpenCommand scores command using the pitch location's distance from target.
This dataset contains the 2024/2025/2026 computer vision object detections, every intermediate the pipeline writes, and the resulting command scores. The pipeline itself and the full method write-up live at github.com/tomdoyo/open-command.
Download
hf download tomdoyo/open-command --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/tomdoyo/open-command.openvino-arc140v-lunarlake
OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU
Reference performance data for running local models on a single Intel Core Ultra 7 258V
(Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All
inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here.
This is reference characterization shared by a non-expert contributor — careful measurements
on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.
