datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-Behavior-Dataset.Human-Like-DPO-Dataset
Enhancing Human-Like Responses in Large Language Models
🤗 Models | 📊 Dataset | 📄 Paper
📢 The paper associated with this dataset has been accepted to the AAAI-26 Workshop on Personalization in the Era of Large Foundation Models (PerFM).
Human-Like-DPO-Dataset
This dataset was created as part of research aimed at improving conversational fluency and engagement in large language models. It is suitable for formats like Direct Preference Optimization (DPO) to guide… See the full description on the dataset page: https://huggingface.co/datasets/HumanLLMs/Human-Like-DPO-Dataset.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-ABNORMAL-Behavior-Dataset.myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.Human-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews
Human Preferences Alignment KTO Dataset of AI Service User Reviews of ChatGPT Gemini Claude Perplexity
Introduction to Human Preferences Alignment
There are many methods of applying Human Preference Alignment techniques to help model align in the supervised finetuning stage, including RLHF Reinforcement Learning from Human Feedback(paper), PPO Proximal policy optimization(paper/equation), DPO Direct Preference Optimization (paper/equation), KTO Kahneman-Tversky… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Human-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews.real-human-logs-extraction-datasetTo create this dataset, we collected human-generated logs from two individuals. This needs to be beefed up in the future, but this is what we have for now.
Subsequently, we ran NuExtract3 on each example of the dataset to get the extraction ground truth.
References
[1] U.S. Bureau of Labor Statistics, Employed persons by detailed occupation and age, 2025. Available at: https://www.bls.gov/cps/cpsaat11b.htm
HumanLLMs-Human-Like-DPO-Dataset-no-emojis
HumanLLMs/Human-Like-DPO-Dataset without emojis
This is a reformatted version of HumanLLMs/Human-Like-DPO-Dataset with the following changes:
emojis removed from the chosen column
shuffled and split into 3 files with equal number of lines
saved in JSONLines format and compressed using zstd
According to the authors of the original version:
This dataset can be used to fine-tune LLMs to:
Improve conversational coherence.
Reduce mechanical or impersonal responses.
Enhance emotional… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HumanLLMs-Human-Like-DPO-Dataset-no-emojis.mlx_Human-Like-DPO-datasethan-human-preference-assist-dataset-v1
Human Preference Assist Dataset
Overview
A dataset capturing user-specific
preferences during humanoid assistance tasks.
Supports personalization and adaptive interaction.
Data Fields
user_id
preferred_task_style
preferred_speed
interaction_tone
confirmation_required
Intended Use
Personalized robotics systems
Adaptive assistance research
Human-robot interaction modeling
License
MIT
han-human-robot-dialog-task-dataset-v1
Human Robot Dialogue Task Dataset
Overview
This dataset contains structured dialogue exchanges
between humans and humanoid robots during task assignment.
It focuses on short conversational flows
leading to clear executable tasks.
Key Features
Human request
Robot clarification
Final structured task
Dialogue outcome
Intended Use
Human-robot interaction research
Dialogue-to-task mapping
Clarification system training
Domestic assistance agents… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-human-robot-dialog-task-dataset-v1.General-Instructions-Human-DatasetHuman-Like-DPO-Dataset
Amélioration des réponses semblables à celles des humains dans les grands modèles de langage
📊 Jeu de données original | 📄 Article
Human-Like-DPO-Dataset
Ce jeu de données est la traduction français des du data set Human-Like-DPO-Dataset à l'aide de grok-3. Le dataset original a été créé dans le cadre de recherches visant à améliorer la fluidité conversationnelle et l'engagement dans les grands modèles de langage. Il est adapté à des formats… See the full description on the dataset page: https://huggingface.co/datasets/nsemhoun/Human-Like-DPO-Dataset.OpenLLM-France__Lucie-7B-Instruct-human-data-details
Dataset Card for Evaluation run of OpenLLM-France/Lucie-7B-Instruct-human-data
Dataset automatically created during the evaluation run of model OpenLLM-France/Lucie-7B-Instruct-human-data
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenLLM-France__Lucie-7B-Instruct-human-data-details.han-human-ai-trust-feedback-dataset-v1
Human–AI Trust Feedback Dataset
This dataset contains human feedback related to trust, comfort,
and reliability when interacting with humanoid AI systems.
It helps Humanoid Network models learn how trust is built,
maintained, or lost during human-AI interactions.
Use Cases
Trust modeling
Ethical AI evaluation
Human-centered system tuning
Fields
interaction_context
human_emotion
trust_level
feedback_text
timestamp
Part of
Humanoid Network (HAN)… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-human-ai-trust-feedback-dataset-v1.BesiegeField_humandataset_coldstart
🏰 BesiegeField Human Dataset for ColdStart
📎 Links
Project Page: https://besiegefield.github.io/
GitHub: https://github.com/Godheritage/BesiegeField
arXiv: https://arxiv.org/abs/2510.14980
📌 Overview
This dataset collects human-made Besiege machines and reformats them for the LLM Cold-Start stage.Processing details are described in Paper §F.1.
🗂️ Data Fields
trainable_dataset: Dataset root.
ID: Unique id for each data.
xxx.json: The tree… See the full description on the dataset page: https://huggingface.co/datasets/Godheritage/BesiegeField_humandataset_coldstart.Synthetic-data_Human-summaryText written by Granite 3.0 8B, OLMoE 1B-7B, OLMo 7B, and summarized by me.
JSON array of objects, each of which has a .text and .summary property
New ones added regularly.
decrypto_human_dataHuman-Essence-Datasethumanoid-human-awareness-dataset-v1dataset_human_interaction.jsonAIDetection_Vietnamese_HumanData
