datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Java-GitHub-CodesPerceptionComp
PerceptionComp: A Benchmark for Complex Perception-Centric Video Reasoning
PerceptionComp is a benchmark for complex perception-centric video reasoning. It focuses on questions that cannot be solved from a single frame, a short clip, or a shallow caption. Models must revisit visually complex videos, gather evidence across temporally separated segments, and combine multiple perceptual cues before answering.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hrinnnn/PerceptionComp.wikitext-tags-modernberthr-interview-datasetYT-100K
A larger version of YT-100K dataset -> YT-30M dataset with 30 million YouTube multilingual multicategory comments is available here: YTCommentVerse
Bibtex
@inproceedings{dutta2025ytcommentverse,
title={YTCommentVerse: A Multi-Category Multi-Lingual YouTube Comment Corpus},
author={Dutta, Hridoy Sankar and Khan, Biswadeep},
booktitle={Proceedings of the 34th ACM International Conference on Information and Knowledge Management},
pages={6351--6355},
year={2025}
}… See the full description on the dataset page: https://huggingface.co/datasets/hridaydutta123/YT-100K.sciercHR-Instruct-Math-v0.1
Dataset Summary
HAERAE-HUB/HR-Instruct-Math-v0.1 is a Math instruction dataset written in the Korean language. This dataset contains evolved instructions aimed at enhancing the learning experience in mathematical concepts. The responses in this dataset are generated from open-source Language Models (LLMs). This is a Proof of Concept (PoC) version, meaning there may be errors or unexpected problems in the dataset. Future iterations will be made to improve the dataset quality.… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HR-Instruct-Math-v0.1.hrithik_roshan_imagesTUC-HRI-CS
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the cross-subject validation dataset. For random validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI-CS.acl-arcAll_Puzzles_5k_New_Context_HritikTUC-HRI
University of Technology Chemnitz, Germany
Department Robotics and Human Machine Interaction
Author: Robert Schulz
TUC-HRI Dataset Card
TUC-AR is an action recognition dataset, containing 10(+1) action categories for human machine interaction. This version contains video sequences, stored as images, frame by frame.
We introduce two validation types: random validation and cross-subject validation. This is the random validation dataset. For cross-subject validation, please use… See the full description on the dataset page: https://huggingface.co/datasets/SchulzR97/TUC-HRI.syntaxgym-hexatagged
Dataset Card for "syntaxgym-hexatagged"
More Information needed
blimp-hexatagged-incremental
configs:
hr-intent-dataset
Dataset Card for HR Intent Dataset
Dataset Summary
A dataset for intent classification in enterprise HR workflows. Each row contains a user query, context (as a structured field with domain, topic, subject), and a label for the HR intent.
Supported Tasks and Leaderboards
Intent Classification (text classification)
Suitable for benchmarking BERT, RoBERTa, etc., on real-world HR requests.
Languages
English
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/SantmanKT/hr-intent-dataset.finalSampleT5-dataset
T5 Grammar Correction Dataset
This dataset combines Clang8 and Mohammed Ashraf's grammar correction datasets, tokenized for T5 model fine-tuning.
Usage example
from datasets import load_dataset
dataset = load_dataset("Hritshhh/T5-Dataset")
IEEEChatbotAplha
Dataset Card for Dataset Name
This dataset has been meticulously curated by the AI team at IEEE Student Branch, Vishwakarma Institute of Technology (VIT) Pune, with the explicit purpose of training the Llama2 model. It encompasses a diverse range of topics essential for the development of an effective conversational AI system.
Dataset Details
Dataset Description
The dataset comprises a comprehensive selection of topics, including but not limited to:
Frequently… See the full description on the dataset page: https://huggingface.co/datasets/hriteshMaikap/IEEEChatbotAplha.pubchem_molecule_dataset_v3IRIS_flower_dataset
🤖 LMSYS-Chat-GPT-5-Chat-Response
The dataset used in Black-Box On-Policy Distillation of Large Language Models paper. Homepage at here.
This dataset is an extension of the LMSYS-Chat-1M-Clean corpus, specifically curated by collecting high-quality, non-refusal responses from the GPT-5-Chat API.
The LMSYS-Chat-1M dataset collects real-world user queries from the Chatbot Arena.
There is no tool calls or reasoning in the GPT-5-Chat response.
💾 Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/Hrinmayi/IRIS_flower_dataset.placement-miniclinais.trainHRIN-ProTstabpubchem_molecule_datasethr_interview_chatbot_gemmaSynthetic_data_text1
Dataset Card for Synthetic_data_text1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Hrishi147/Synthetic_data_text1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Hrishi147/Synthetic_data_text1.items_raw_litehr_interview_chatbot_llama2LLaMA3-Instruct-Medicalhrif_windowed
