CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HRDexDB /HRDexDB HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments Authors Jongbin Lim¹⋆, Taeyun Ha¹⋆, Seongho Cha, Kanghyun Cho, Mingi Choi¹, Subin Jeon¹, Jisoo Kim¹, Byungjun Kim¹, Hanbyul Joo¹²† ¹ Seoul National University² RLWRLD ⋆ Equal contribution† Corresponding author News (2026.09.20) The full set of Robotiq 2F-85 data has been uploaded! (2026.07.27) We are improving the quality of the object mesh and the tracking results.… See the full description on the dataset page: https://huggingface.co/datasets/HRDexDB/HRDexDB.3d10K<n<100K16 likes39k downloads5d agoHugging Face02DreamMr /HR-Bench Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models 🌐Homepage | 📖 Paper 📊 HR-Bench We find that the highest resolution in existing multimodal benchmarks is only 2K. To address the current lack of high-resolution multimodal benchmarks, we construct HR-Bench. HR-Bench consists two sub-tasks: Fine-grained Single-instance Perception (FSP) and Fine-grained Cross-instance Perception (FCP).… See the full description on the dataset page: https://huggingface.co/datasets/DreamMr/HR-Bench.textvisual-question-answering1K<n<10K17 likes7.7k downloads10mo agoHugging Face03xuejun72 /HR-VILAGE-3K3M HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.tabularzero-shot-classification1K<n<10K3 likes5.2k downloads26d agoHugging Face04Wenliang04 /HRScene HRScene - High Resolution Image Understanding 🌐 Homepage | 🤗 Dataset | 📖 arXiv | GitHub ⭐ About HRScene We introduce HRScene, a novel unified benchmark for HRI understanding with rich scenes. HRScene incorporates 25 real-world datasets and 2 synthetic diagnostic datasets with resolutions ranging from 1,024 × 1,024 to 35,503 × 26,627. HRScene is collected and re-annotated by 10 graduate-level annotators, covering 25 scenarios, ranging from microscopic and radiology… See the full description on the dataset page: https://huggingface.co/datasets/Wenliang04/HRScene.image100K<n<1M2 likes4.7k downloads1y agoHugging Face05guychuk /HRM-He-corpus-objective Hebrew reasoning traces Generated Hebrew chain-of-thought over code, cybersecurity, agentic, math and general-reasoning seeds. Built for a Hebrew/English code-specialised LM, where off-the-shelf Hebrew reasoning data is effectively nonexistent. What the default config contains Every row the training corpus keeps -- not a filtered highlight reel. Two things are disqualifying and are absent: a wrong final answer (answer_ok is False), and Arabic drift. Everything… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/HRM-He-corpus-objective.tabulartext-generation100K<n<1M0 likes4.6k downloads11d agoHugging Face06vidore /vidore_v3_hrViDoRe V3 : HR This dataset, HR, is a corpus of reports released by the european union, intended for complex-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with human-verified relevant pages… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_hr.documentvisual-document-retrieval10K<n<100K11 likes3.7k downloads8mo agoHugging Face07sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.5k downloads4mo agoHugging Face08hrishizone /Java-GitHub-Codestext1M<n<10M1 likes2.1k downloads1y agoHugging Face09HAERAE-HUB /HRM8K | 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) | HRM8K We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English. HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams. Benchmark Overview The HRM8K benchmark consists of two subsets: Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.tabular1K<n<10K24 likes1.7k downloads2y agoHugging Face10dronefreak /HRSID HRSID: High-Resolution SAR Images Dataset (Ship Detection) Unofficial redistribution of the HRSID high-resolution SAR ship-detection dataset, reformatted into a standardized YOLO-compatible directory layout. License status is unclear -- see License before using this beyond research. Disclaimer This repository is not an official release of HRSID. HRSID was created by Shunjun Wei, Xiangfeng Zeng, Qizhe Qu, Mou Wang, Hao Su, and Jun Shi and released via… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/HRSID.imageobject-detection1K<n<10K0 likes1.3k downloads9d agoHugging Face11classla /ParlaSpeech-HR The Croatian Parliamentary Spoken Dataset ParlaSpeech-HR 2.0 The master dataset can be found at http://hdl.handle.net/11356/1914. Notice: ParlaSpeech corpora are currently in the process of enrichment with new features. Follow our progress here: http://clarinsi.github.io/parlaspeech The ParlaSpeech-HR dataset is built from the transcripts of parliamentary proceedings available in the Croatian part of the ParlaMint corpus (http://hdl.handle.net/11356/1859), and the parliamentary… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSpeech-HR.audio100K<n<1M6 likes1.1k downloads1y agoHugging Face12WeiQian98 /VIPL-HRgatedtext0 likes1k downloads3mo agoHugging Face13hrinnnn /PerceptionComp PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering. Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.tabularvisual-question-answering1K<n<10K2 likes917 downloads6mo agoHugging Face14vidore /vidore_v3_hr_mteb_format Vidore3HrRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_hr How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("Vidore3HrRetrieval") evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_hr_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes870 downloads11mo agoHugging Face15abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes708 downloads3mo agoHugging Face16Yux1ang /GTA-UAV-HR GTA-UAV dataset # Merge splited files cat drone_part_* > drone.tar.gz # Extract the archive tar -xzvf drone.tar.gz tar -xzvf satellite.tar.gz For more information, please check our project page. Sources Repository: https://github.com/Yux1angJi/GTA-UAV Paper: https://arxiv.org/abs/2409.16925 text10K<n<100K3 likes648 downloads11mo agoHugging Face17sensenova /HR-MMSearch Dataset Description HR-MMSearch is a benchmark designed to evaluate the Agentic Reasoning and Search capabilities of Multimodal Large Language Models in complex visual tasks. This dataset was introduced by SenseTime Research in the paper SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning. Key Features: High-Resolution Images: Contains high-resolution image inputs, requiring the model to possess fine-grained visual perception… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/HR-MMSearch.imagen<1K0 likes582 downloads9mo agoHugging Face18cpratikaki /RSVQA-HR_qwen_finetuningimage100K<n<1M1 likes558 downloads2y agoHugging Face19dmarsili /RSVQA-HR-2kA 2k subset of the validation split of the RSVQA HR dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation. For more information and attribution please refer to the original dataset: https://rsvqa.sylvainlobry.com/#dataset imagevisual-question-answering1K<n<10K0 likes437 downloads3mo agoHugging Face20saurabh1043 /chirp-en-in-10s-hr-85htext10K<n<100K0 likes383 downloads10mo agoHugging Face21allenai /href_resultstabularn<1K0 likes376 downloads1y agoHugging Face22bdanko /DIV2K_train_HRimagen<1K0 likes371 downloads6mo agoHugging Face23hr16 /ViVoicePPaudio1K<n<10K1 likes303 downloads5mo agoHugging Face24classla /copa_hrThe COPA-HR dataset (Choice of plausible alternatives in Croatian) is a translation of the English COPA dataset (https://people.ict.usc.edu/~gordon/copa.html) by following the XCOPA dataset translation methodology (https://arxiv.org/abs/2005.00333). The dataset consists of 1000 premises (My body cast a shadow over the grass), each given a question (What is the cause?), and two choices (The sun was rising; The grass was cut), with a label encoding which of the choices is more plausible given the annotator or translator (The sun was rising). The dataset is split into 400 training samples, 100 validation samples, and 500 test samples. It includes the following features: 'premise', 'choice1', 'choice2', 'label', 'question', 'changed' (boolean).texttext-classification1K<n<10K0 likes300 downloads4y agoHugging Face25maywovel /HR-VILAGE-3K3M HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a benchmark… See the full description on the dataset page: https://huggingface.co/datasets/maywovel/HR-VILAGE-3K3M.tabularzero-shot-classificationn<1K0 likes280 downloads5mo agoHugging Face26xwjzds /extractive_qa_question_answering_hr Dataset Card HR-Multiwoz is a fully-labeled dataset of 5980 extractive qa spanning 10 HR domains to evaluate LLM Agent. It is the first labeled open-sourced conversation dataset in the HR domain for NLP research. Please refer to HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent for details about the dataset construction. Dataset Sources Repository: xwjzds/extractive_qa_question_answering_hr Paper: HR-MultiWOZ: A Task Oriented Dialogue (TOD)… See the full description on the dataset page: https://huggingface.co/datasets/xwjzds/extractive_qa_question_answering_hr.text1K<n<10K9 likes275 downloads3y agoHugging Face27strova-ai /hr-policies-qa-dataset 📚 HR Policies Q&A Dataset 🔎 Overview This dataset provides multi-turn Q&A conversations on HR policies and compliance, formatted with system, user, and assistant roles.It is designed for: 🤖 LLM fine-tuning 💬 HR & compliance chatbots 🏢 Enterprise policy automation By covering real-world HR scenarios — such as policy reviews, compliance processes, and employee communication — this dataset helps train assistants that can: ✅ Clarify company policies✅ Ensure… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/hr-policies-qa-dataset.textn<1K0 likes271 downloads1y agoHugging Face28porupski /ParlaSpeech-HR-benchmark_v3 ParlaSpeechHR Benchmark v3 A curated benchmark dataset of 22,008 Croatian parliamentary speech clips extracted from ParlaSpeech-HR v3. Each clip includes aligned audio (WAV) and TextGrid annotations for linguistic analysis. Contents 22,008 audio segments (various durations) 17,622 clips with complete TextGrid triplets: .align (word-level boundaries via WordAlign tier) .stress (primary stress frame labels; derivative of .align) .pause (filled pause annotations… See the full description on the dataset page: https://huggingface.co/datasets/porupski/ParlaSpeech-HR-benchmark_v3.audion<1K0 likes269 downloads2mo agoHugging Face29hrtxsny /SWE-bench-plus SWE-bench-Plus: Test Enhancer SWE-bench-Plus is a coverage-guided test generation and evaluation layer built on top of the official SWE-bench harness. It automates iterative LLM-based test generation, avoids duplicates, targets uncovered code paths, and stops when coverage plateaus. It is designed for high-throughput, resume-friendly batch runs with robust logging and fault tolerance. Key Features Coverage-guided generation: After each iteration, the harness measures… See the full description on the dataset page: https://huggingface.co/datasets/hrtxsny/SWE-bench-plus.textn<1K0 likes254 downloads6mo agoHugging Face30classla /hr500kThe hr500k training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, lemmatisation and named entities. On the sentence level, the dataset contains 20159 training samples, 1963 validation samples and 2672 test samples across the respective data splits. Each sample represents a sentence and includes the following features: sentence ID ('sent_id'), sentence text ('text'), list of tokens ('tokens'), list of lemmas ('lemmas'), list of Multext-East tags ('xpos_tags), list of UPOS tags ('upos_tags'), list of morphological features ('feats'), and list of IOB tags ('iob_tags'). The 'upos_tags' and 'iob_tags' features are encoded as class labels.textother10K<n<100K0 likes253 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.