datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
waqfeya-library
Waqfeya Library
📖 Overview
Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.hyperspectral-orchard
Living Optics Orchard Dataset
Overview
This dataset contains 435 images of captured in one of the UK's largest orchards, using the Living Optics Camera.
The data consists of RGB images, sparse spectral samples and instance segmentation masks.
The dataset is derived from 44 unique raw files corresponding to 435 frames.
Therefore, multiple frames could originate from the same raw file.
This structure emphasized the need for a split strategy that avoided data leakage.
To… See the full description on the dataset page: https://huggingface.co/datasets/LivingOptics/hyperspectral-orchard.orewaseikankokkanoakutokuryoushu
Bangumi Image Base of Ore Wa Seikan Kokka No Akutoku Ryoushu!
This is the image base of bangumi Ore wa Seikan Kokka no Akutoku Ryoushu!, we detected 74 characters, 4170 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/orewaseikankokkanoakutokuryoushu.brand-assetsRPC-Bench
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
🌐 Project Page •
💻 GitHub •
📖 Paper
RPC-Bench is a fine-grained benchmark for research paper comprehension. It is built from review-rebuttal exchanges of high-quality academic papers and supports both text-only and visual evaluation through complementary paper representations.
Data Structure
RPC-Bench is organized into train, dev, and test subsets. Split assignments… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/RPC-Bench.data_verticalRealworldQA
RealWorldQA
RealWorldQA is a benchmark designed for real-world understanding. The dataset consists of anonymized images taken from vehicles, in addition to other real-world images. We are excited to release RealWorldQA to the community, and we intend to expand it as our multimodal models improve.
The initial release of the RealWorldQA consists of over 700 images, with a question and easily verifiable answer for each image. See the announcement of Grok-1.5 Vision Preview.… See the full description on the dataset page: https://huggingface.co/datasets/xai-org/RealworldQA.orena-cold-storage
orena-cold-storage
ORena / FOCUS MICCAI 2026 — DiffAtlas 冷存储归档。本地集群空间释放后迁移至此长期保存。
总览:20,390 文件 / 1,495.75 GiB
归档时间:2026-08-30
可见性:public
完整性:全部 2,750 个 .pt 权重 LFS SHA256 校验有效(0 空 / 0 损坏 / 0 重复)
内容索引
Model/ — 模型权重(2,589 文件 / 1,486.46 GiB)
子项目(实验)
文件数
大小
TS_train_80percent_base_10w
784
450.13 GiB
DiffAtlas_TotalSegmentator
500
287.07 GiB
TS_train_80percent
450
258.37 GiB
DiffAtlas_MMWHS-MRI_full
426
244.59 GiB… See the full description on the dataset page: https://huggingface.co/datasets/gwd200/orena-cold-storage.chocopan-t3-reverse-oracle-hdf5-wide-v1
chocopan-t3-reverse-oracle-hdf5-wide-v1
Raw HDF5 output of a scripted oracle for reverse manipulation tasks in simulation -- take an object out of a container or off a plate and put it back on the table -- for the batch generated with widened object initial placements: 7,200 attempts over 45 tasks, failures included.
This is the raw, unfiltered output of the generator, in LIBERO's create_dataset.py HDF5
layout. It is published because it is bulky to regenerate, not because it is… See the full description on the dataset page: https://huggingface.co/datasets/chocopan/chocopan-t3-reverse-oracle-hdf5-wide-v1.VCR-wiki-en-easy
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-easy.umi-okra-dex1
UMI Okra Grasping Dataset — Lab Subset (Unitree Dex1-1)
Universal Manipulation Interface (UMI) hand-held teleoperation data for an okra-fruit
grasping / harvesting task. Recorded to train a Diffusion Policy (with a comparison
ACT track) deployed on a Unitree G1.
This is the indoor lab subset. Every session here was recorded in a lab mock okra
field (artificial foliage, white-walled room). The outdoor sessions from the original
collection are not included — see Scope and… See the full description on the dataset page: https://huggingface.co/datasets/Orboh/umi-okra-dex1.ORD
🌋 STORM: Stimulating Trustworthy Ordinal Regression Ability of MLLMs
Benchmarking All-in-one Visual Rating of MLLMs with A Comprehensive Ordinal Regression Dataset.
Contents
STORM Weights
Dataset
Evaluation
Examples
STORM Weights
Please check out our checkpoint_STORM for public STORM checkpoints, and the instructions of how to use the weights.
Dataset
Data file name
Size
STORM_instruct_MAX_527k.jsonl
383 MB… See the full description on the dataset page: https://huggingface.co/datasets/ttlyy/ORD.robo_orchard_sim_assetspredictionsREDS_origVCR-wiki-en-hard
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.DeepSea-MOT
DeepSea MOT
DeepSea MOT is a benchmark dataset for multi-object tracking on deep-sea video.
Dataset Description
DeepSea MOT consists of 4 video sequences (2 midwater, 2 benthic) with a total of 2,400 frames and 57,376 annotated objects comprising 188 tracks. The videos were captured by the Monterey Bay Aquarium Research Institute (MBARI) using remotely operated vehicles (ROVs) Doc Ricketts and Ventana in deep-sea environments, showcasing a variety of marine species and… See the full description on the dataset page: https://huggingface.co/datasets/MBARI-org/DeepSea-MOT.LVBench
LVBench: An Extreme Long Video Understanding Benchmark
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset][🏆 Leaderboard]
LVBench is a benchmark designed to evaluate and enhance the capabilities of multimodal models in understanding and
extracting information from long videos up to two hours in duration.
🔥 News
2024.06.11 🌟 We released LVBench, a new benchmark for long video understanding!
👀 Introduce to LVBench
LVBench is a benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LVBench.Orchard
Orchard Dataset
Overview
Orchard is the trajectory release accompanying the paper "Orchard: An Open-Source Agentic Modeling Framework" (Peng et al., 2026). It bundles two parallel agentic-modeling datasets distilled from strong teacher models, both produced inside the same Orchard Env sandbox infrastructure:
swe — 107,185 multi-turn software-engineering trajectories across 2,788 GitHub repositories, each labeled with whether the agent's final patch passed the… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Orchard.Vision2Web
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
[🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results]
Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle.
This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.bo_or_not
Dataset Card for bo-dataset
This is a FiftyOne dataset with 169 samples designed for binary classification of Bo (Barack Obama's Portuguese Water Dog) versus other pets.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/bo_or_not")
#… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/bo_or_not.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.OracleZoom-4KLSDB-train
OracleZoom-4KLSDB-train
Paper: arXiv:2609.06490 (HF paper page) · Demo: 🤗 Space · Code: https://github.com/dipta007/OracleZoom · Project page: https://dipta007.github.io/OracleZoom/ · Everything in one place: 🤗 collection
Curated training data for OracleZoom (reference-constrained recursive super-resolution, inspired
by on-policy self-distillation), derived from SingleBicycle/4KLSDB
(license: CC BY 4.0, attribution to 4KLSDB required for any downstream use).
Paired with the… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/OracleZoom-4KLSDB-train.OraRL-Data
OraRL-Data
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code]
We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL.
It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.corpus-archive
corpus-archive
[!WARNING]
Experimental Dataset Architecture: The repository structure, metadata tiers, category taxonomies, and catalog indexing formats are currently under active design and evaluation. All specifications, metadata keys, and JSON schemas detailed below represent representational examples and intended targets.
This repository serves as a structured digital textual archive preserving Hmar literature, historical accounts, school textbooks, dictionaries, parallel… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/corpus-archive.Ornith-1.0-9B-atlas
juiceb0xc0de/Ornith-1.0-9B-atlas
A brain atlas for deepreinforce-ai/Ornith-1.0-9B, the 9B agentic-coding model that reports SOTA results on Terminal-Bench, SWE-Bench, and other agentic coding benchmarks. This is not a chat dataset or a benchmark — it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know why this model survives surgical edits, where… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Ornith-1.0-9B-atlas.oral-yolo-dataset
Oral YOLO Lesion Detector
This model is a YOLO-based object detection model trained to detect oral lesion regions in smartphone-captured oral cavity images.
The model is intended for research, prototyping, and assistive screening workflows. It is not a medical device and must not be used as the sole basis for diagnosis, treatment decisions, or clinical triage.
Model Details
Model type: YOLO object detector
Task: Oral lesion detection / object detection
Input: Oral cavity… See the full description on the dataset page: https://huggingface.co/datasets/sach3v/oral-yolo-dataset.HEV-ORF1-models
Hepatitis E virus ORF1 (nsp1) — AlphaFold2 model collection
1,178 AlphaFold2 predictions of the HEV ORF1 (nsp1) replicase, packaged so a
static web app can render the 3D model, the predicted aligned error (PAE) matrix and the multiple
sequence alignment without a server.
open the viewer: https://tubiana.github.io/ORF1viewer (this dataset is its data root)
repository — app + pipeline code, no data: https://github.com/tubiana/tubiana.github.io
dataset repo:… See the full description on the dataset page: https://huggingface.co/datasets/ttubiana/HEV-ORF1-models.pixmo-cap-images
PixMo-Cap
Big thanks to Ai2 for releasing the original PixMo-Cap dataset. To preserve the images and simplify usage of the dataset, we are releasing this version, which includes downloaded images.
PixMo-Cap is a dataset of very long (roughly 200 words on average), detailed captions.
It can be used to pre-train and fine-tune vision-language models.
PixMo-Cap was created by recording annotators speaking about an image for 60-90 seconds and then using the Claude large language model… See the full description on the dataset page: https://huggingface.co/datasets/anthracite-org/pixmo-cap-images.
