datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HRDexDB
HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
Authors
Jongbin Lim¹⋆,
Taeyun Ha¹⋆,
Seongho Cha,
Kanghyun Cho,
Mingi Choi¹,
Subin Jeon¹,
Jisoo Kim¹,
Byungjun Kim¹,
Hanbyul Joo¹²†
¹ Seoul National University² RLWRLD
⋆ Equal contribution† Corresponding author
News
(2026.09.20) The full set of Robotiq 2F-85 data has been uploaded!
(2026.07.27) We are improving the quality of the object mesh and the tracking results.… See the full description on the dataset page: https://huggingface.co/datasets/HRDexDB/HRDexDB.hrrr-kerchunkThis dataset is comprised of the output from kerchunk's scan_grib across the entire AWS-hosted HRRR forecast files, including pressure levels, surface, and sub-hourly, but not native levels.
Each grib message is its own json file, and each init time is its own zip containing the whole extracted json files for all the forecast times for that init time. Once the kerchunk extraction is
complete, the plan is to combine them so that they can be opened in a single call to Xarray as one very large… See the full description on the dataset page: https://huggingface.co/datasets/jacobbieker/hrrr-kerchunk.d2pfxHR-Bench
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
🌐Homepage | 📖 Paper
📊 HR-Bench
We find that the highest resolution in existing multimodal benchmarks is only 2K. To address the current lack of high-resolution multimodal benchmarks, we construct HR-Bench. HR-Bench consists two sub-tasks: Fine-grained Single-instance Perception (FSP) and Fine-grained Cross-instance Perception (FCP).… See the full description on the dataset page: https://huggingface.co/datasets/DreamMr/HR-Bench.AIA_12hour_512x512HR-VILAGE-3K3M
HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression
This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.HRVQAp2-etf-hrp-allocator-resultsHRScene
HRScene - High Resolution Image Understanding
🌐 Homepage |
🤗 Dataset |
📖 arXiv |
GitHub
⭐ About HRScene
We introduce HRScene, a novel unified benchmark for HRI understanding with rich scenes. HRScene incorporates 25 real-world datasets and 2 synthetic diagnostic datasets with resolutions ranging from 1,024 × 1,024 to 35,503 × 26,627. HRScene is collected and re-annotated by 10 graduate-level annotators, covering 25 scenarios, ranging from microscopic and radiology… See the full description on the dataset page: https://huggingface.co/datasets/Wenliang04/HRScene.HRM-He-corpus-objective
Hebrew reasoning traces
Generated Hebrew chain-of-thought over code, cybersecurity, agentic, math and
general-reasoning seeds. Built for a Hebrew/English code-specialised LM, where
off-the-shelf Hebrew reasoning data is effectively nonexistent.
What the default config contains
Every row the training corpus keeps -- not a filtered highlight reel. Two things
are disqualifying and are absent: a wrong final answer (answer_ok is False), and
Arabic drift. Everything… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/HRM-He-corpus-objective.vidore_v3_hrViDoRe V3 : HR
This dataset, HR, is a corpus of reports released by the european union, intended for complex-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with human-verified relevant pages… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_hr.hrm-tokenized-bpe65kHRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.hebrew-hrm-corpus
Hebrew HRM-Text Corpus
Training corpus for a Hebrew Hierarchical Reasoning Model, replicating the
sapientinc/HRM-Text-1B recipe
(train-from-scratch, PrefixLM over {condition, instruction, response}, loss on response only).
Schema
Each line: {"condition": "<tags>", "instruction": "...", "response": "..."}.
Condition tags map to special tokens: direct→<|object_ref_start|>, cot→<|object_ref_end|>,
noisy→<|quad_start|>, synth→<|quad_end|> (composite tags… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-hrm-corpus.HRP4KHRVideoBench
HRVideoBench
This repo contains the test data for HRVideoBench, which is released under the paper "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation". VISTA is a video spatiotemporal augmentation method that generates long-duration and high-resolution video instruction-following data to enhance the video understanding capabilities of video LMMs.
🌐 Homepage | 📖 arXiv | 💻 GitHub | 🤗 VISTA-400K | 🤗 Models | 🤗 HRVideoBench… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/HRVideoBench.Java-GitHub-CodesDota2PornFx
Dota2PornFx Dataset
This dataset contains a collection of mods designed for the website Dota2PornFxWeb.
Website Repository: h6rd/Dota2PornFxWeb
📜 License & Usage Terms
This dataset is distributed under the GNU General Public License v3.0 (GPL-3.0).
While the dataset is open-source under GPLv3, you must adhere to the following attribution guidelines when using, modifying, or redistributing these files:
Attribution Requirements
Credit the… See the full description on the dataset page: https://huggingface.co/datasets/hrdq/Dota2PornFx.HRM8K
| 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) |
HRM8K
We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English.
HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams.
Benchmark Overview
The HRM8K benchmark consists of two subsets:
Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.CDOPSThe Complex Dynamics of Online Professional Squads (CDOPS) dataset
comprises a collection of over 1.25 million rounds of Counter-Strike:
Global Offensive (CS:GO) played in professional tournaments under
regulated Counter-Strike servers. These games were publicly posted and
collected from hltv.org. We have developed this dataset for exploring
human squad dynamics in the context of E-sports.
Each game is described by 2 primary parquet files: an events file, and a
ticks file. At each tick (game… See the full description on the dataset page: https://huggingface.co/datasets/HRL-Labs/CDOPS.HRDexDB-glb360x_dataset_HR
360+x Dataset
For more information, please feel free to check our project page.
Overview
360+x dataset introduces a unique panoptic perspective to scene understanding, differentiating itself from traditional
datasets by offering multiple viewpoints and modalities, captured from a variety of scenes
Key Features:
Multi-viewpoint Captures: Includes 360° panoramic video, third-person front view video, egocentric monocular
video, and egocentric binocular video.… See the full description on the dataset page: https://huggingface.co/datasets/quchenyuan/360x_dataset_HR.HRSID
HRSID: High-Resolution SAR Images Dataset (Ship Detection)
Unofficial redistribution of the HRSID high-resolution SAR ship-detection dataset, reformatted into a standardized YOLO-compatible directory layout. License status is unclear -- see License before using this beyond research.
Disclaimer
This repository is not an official release of HRSID.
HRSID was created by Shunjun Wei, Xiangfeng Zeng, Qizhe Qu, Mou Wang, Hao Su, and Jun Shi and released via… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/HRSID.cvsearch_hr8kVisualizations of CVSearch
Citation
@misc{li2026cvsearchempoweringmultimodalllms,
title={CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception},
author={Liupeng Li and Haoqian Kang and Zhenyu Lu and Jinpeng Wang and Bin Chen and Ke Chen and Yaowei Wang},
year={2026},
eprint={2605.23655},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.23655},
}… See the full description on the dataset page: https://huggingface.co/datasets/tothanhdat/cvsearch_hr8k.ParlaSpeech-HR
The Croatian Parliamentary Spoken Dataset ParlaSpeech-HR 2.0
The master dataset can be found at http://hdl.handle.net/11356/1914.
Notice: ParlaSpeech corpora are currently in the process of enrichment with new features. Follow our progress here: http://clarinsi.github.io/parlaspeech
The ParlaSpeech-HR dataset is built from the transcripts of parliamentary proceedings available in the Croatian part of the ParlaMint corpus (http://hdl.handle.net/11356/1859), and the parliamentary… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSpeech-HR.MPM-Verse-MaterialSim-Small
Dataset Card for MPMVerse Physics Simulation Dataset
Dataset Summary
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, elasticity, jelly, rigid collisions, and melting. Each material is represented as point-clouds that evolve over time. The dataset is designed for learning and predicting MPM-based physical simulations.
Supported Tasks and Leaderboards
The dataset supports tasks such as:… See the full description on the dataset page: https://huggingface.co/datasets/hrishivish23/MPM-Verse-MaterialSim-Small.VIPL-HRPerceptionComp
PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning
PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.h-rag_databasevidore_v3_hr_mteb_format
Vidore3HrRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_hr
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("Vidore3HrRetrieval")
evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_hr_mteb_format.
