datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images
HSTLI: A Dataset of Human Semen Time Lapse Images
Dataset Details
HSTLI contains 3,266 time-lapse microscopy videos of human sperm.Clips were recorded from two imaging modalities:
CASA system (Sperm Class Analyzer)
Optical microscope (Swift M10DB-MP + Fujifilm X-T30)
A subset of videos was manually annotated with bounding boxes around each visible sperm head.
The dataset supports detection, tracking and motility computation.
Total contents:
34… See the full description on the dataset page: https://huggingface.co/datasets/DFL-KamLab/HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images.Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-Behavior-Dataset.700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.Flux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.Human-Like-DPO-Dataset
Enhancing Human-Like Responses in Large Language Models
🤗 Models | 📊 Dataset | 📄 Paper
📢 The paper associated with this dataset has been accepted to the AAAI-26 Workshop on Personalization in the Era of Large Foundation Models (PerFM).
Human-Like-DPO-Dataset
This dataset was created as part of research aimed at improving conversational fluency and engagement in large language models. It is suitable for formats like Direct Preference Optimization (DPO) to guide… See the full description on the dataset page: https://huggingface.co/datasets/HumanLLMs/Human-Like-DPO-Dataset.text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.text-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/OpenSparX/Cabin-Human-ABNORMAL-Behavior-Dataset.human-vs-Ai-generated-datasetBenchmark_Dataset-Human_population_classification
Sumary
This dataset provides a benchmark for evaluating the model's ability to leverage richer genetic information from longer sequences to achieve more accurate inference.
Using data from the Human Pangenome Reference Consortium (BioProject ID: PRJNA730823), we designed a population classification task focusing on African, East Asian, and European population groups.
From samples' VCF file and the reference genome sequence, we generated sample pseudo-sequences.
Based on variant… See the full description on the dataset page: https://huggingface.co/datasets/BGI-HangzhouAI/Benchmark_Dataset-Human_population_classification.human-telemetry-driving-dataset-lite-version
Dataset Card for 15 Laps of 30Hz NGSIM-Style Telemetry
This is a Lite Version of a larger research dataset focusing on human driving signatures in high-fidelity simulations. It includes 15 full laps of telemetry captured at 30Hz within Unreal Engine 5, specifically formatted to match NGSIM standards.
Dataset Details
Dataset Description
This Lite Version dataset contains 15 laps of high-fidelity human driving telemetry. It is intended for researchers and… See the full description on the dataset page: https://huggingface.co/datasets/AtlasBuiltIt/human-telemetry-driving-dataset-lite-version.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.humandata-conf-labels-v3
실제 Franka 압축 가능성 라벨 (human_data v3)
각 순간의 모션을 더 빠르게 지나가도 되는지 VLM 으로 매긴 라벨.
만든 방법
판정기 gemini-3.8-flash (reasoning low), 온도 0
입력: 그 순간의 카메라 2장 + 지시문 + 계획된 액션에서 계산한 사실
출력: 문항을 1~5 등급으로. prompts/humandata_v3.txt 가 나간 전문 그대로다
SYSTEM 메시지를 쓰지 않는다. role: user 하나이고 가이던스가 본문 머리에 온다
청크 길이 16, stride 16. 아래 "중간 시점" 참고
게이트는 VLM 신뢰도 하나다. 접촉 열을 쓰지 않는다
실제 Franka 데모 세 벌 · 3,446 청크. 셋 다 같은 프롬프트로 만들었다
(전문·문항·부호·가중이 글자까지 같다).
데이터셋
청크
에피
fps
목적지
conf 평균
pnp_task
874
101… See the full description on the dataset page: https://huggingface.co/datasets/prehj/humandata-conf-labels-v3.human-faces-dataset-r1real-human-faces-data-setuiclip_human_data_hfHuman-Like-DPO-Dataset_deduplicated_and_duplicatedHuman-Like-DPO-Dataset-komy-human-dataset75-percent-human-dataset-opt-mistakeHuman-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews
Human Preferences Alignment KTO Dataset of AI Service User Reviews of ChatGPT Gemini Claude Perplexity
Introduction to Human Preferences Alignment
There are many methods of applying Human Preference Alignment techniques to help model align in the supervised finetuning stage, including RLHF Reinforcement Learning from Human Feedback(paper), PPO Proximal policy optimization(paper/equation), DPO Direct Preference Optimization (paper/equation), KTO Kahneman-Tversky… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Human-Preferences-Alignment-KTO-Dataset-AI-Services-Genuine-User-Reviews.myanmar_quran_parallel_dataset_human_vs_ai
Myanmar Quran Parallel Dataset: Human vs AI
This dataset is a comprehensive multi-parallel corpus of the Holy Qur'an, containing all 6,236 verses.
It is designed as a high-quality linguistic resource for evaluating and aligning AI systems on formal, literary, and modern Myanmar (Burmese) language in a religious context.
Each verse aligns the original Uthmani Arabic text with trusted human translations and multiple AI-generated translations, enabling fine-grained comparison between… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_quran_parallel_dataset_human_vs_ai.Dataset-PTEN_HUMAN
Description
This dataset contains signle site mutation of protein PTEN_HUMAN and the correspond mutation effect score from deep mutation scanning experiment.
Protein Format: AA sequence
Splits
traing: 3311
valid: 375
test: 410
Related paper
The dataset is from Deep generative models of genetic variation capture the effects of mutations.
Label
Label means mutation fitness score (protein stability) of each protein based on deep mutation scanning… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-PTEN_HUMAN.uiclip_human_data-paired_hf0-percent-human-dataset-oghuman_curated_qa_dataset
Human Curated QA Dataset
DigiGreen/human_curated_qa_dataset is a human-verified question-answer dataset designed to support research and development in natural language question answering and agriculture-focused conversational AI.
This dataset contains realistic, domain-relevant QA pairs that were manually curated to ensure accurate and contextually rich answers. It can be used to benchmark models for QA generation.
📌 Dataset Overview
Name: Human Curated QA Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_curated_qa_dataset.real-human-logs-extraction-datasetTo create this dataset, we collected human-generated logs from two individuals. This needs to be beefed up in the future, but this is what we have for now.
Subsequently, we ran NuExtract3 on each example of the dataset to get the extraction ground truth.
References
[1] U.S. Bureau of Labor Statistics, Employed persons by detailed occupation and age, 2025. Available at: https://www.bls.gov/cps/cpsaat11b.htm
human-like-sft-datasethuman_preference_eval_dataset
Human Preference Evaluation Dataset
DigiGreen/human_preference_eval_dataset is a human-annotated preference dataset that supports research and development in evaluating model responses based on human preferences and comparative QA assessment.
This dataset contains real agricultural questions paired with two candidate responses and expert judgments about which response is preferable. It can be used to benchmark preference learning systems, train reward models, and improve… See the full description on the dataset page: https://huggingface.co/datasets/DigiGreen/human_preference_eval_dataset.
