datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humans_cats_dogs_foxeshumans-top
humans.top — LIVE Global ranking of influential people (open dataset)
This dataset ranks real, named living people by global influence — e.g. #1
Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk,
Narendra Modi and Lionel Messi. Every row is a person: their live influence
rank, a concise biography in 15 languages, and Wikidata / Wikipedia links.
Published from the website humans.top (.top is the
domain name).
Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.MolmoWeb-HumanSkills
MolmoWeb-HumanSkills
This dataset was introduced in the paper MolmoWeb: Open Visual Web Agent and Open Data for the Open Web.
A dataset of human collected web-navigation skills, where a skill is a trajectory for a very low level task (eg. find_and_open, fill_form). Each example pairs an instruction with a sequence of webpage screenshots and the corresponding agent actions (clicks, typing, scrolling, etc.).
Dataset Usage
from datasets import load_dataset
# load a… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoWeb-HumanSkills.genomes-v4-genome_set-humans-intervals-v15_256_128human-style-preferences-images
Rapidata Image Generation Preference Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human preference datasets for text-to-image models, this release contains over 1,200,000 human preference… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-style-preferences-images.humans
Human Segmentation Dataset
>>> Download Here <<<
This dataset was created for developing the best fully open-source background remover of images with humans. It was crafted with LayerDiffuse, a Stable Diffusion extension for generating transparent images. After creating segmented humans, IC-Light was used for embedding them into realistic scenarios.
The dataset covers a diverse set of segmented humans: various skin tones, clothes, hair styles etc. Since Stable Diffusion is not… See the full description on the dataset page: https://huggingface.co/datasets/schirrmacher/humans.genomes-v4-genome_set-humans-intervals-v1_256_128-id0.3_cov0.3genomes-v4-genome_set-humans-intervals-v1_254_127-id0.3_cov0.3genomes-v4-genome_set-humans-intervals-v5_256_128genomes-v4-genome_set-humans-intervals-v1_256_128audioset-humans-reprocessedHumanStereoPreview
HumanStereo Preview
This repository contains a non-sharded preview subset of the HumanStereo training split for NeurIPS review.
Sessions: 1438
Staged size: 3.801 GiB
Source dataset: asergiu/HumanStereo
Manifest: manifest/sessions.csv
Croissant metadata: metadata.json
The full dataset is available separately as asergiu/HumanStereo.
Humans_with_Collision
Humans with Collisions (HwC) Pose & Motion Dataset
This dataset contains the training, evaluation, and benchmark data for the paper:"PoseShield: Neural Collision Fields for Human Self-Collision Resolution (ECCV 2026)"
Paper (arXiv): arXiv:2606.29686
Code Repository: PoseShield on GitHub (or project repo)
Dataset Structure
The repository contains two main groups of data structured under the data/ directory:
1. HwC Pose Dataset (Single Poses)
Used… See the full description on the dataset page: https://huggingface.co/datasets/ZYYY99/Humans_with_Collision.interesting_humans_syntheticgenomes-v5-genome_set-humans-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v1_255_128
Humans promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
226,162 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v1_255_128.genomes-v4-genome_set-humans-intervals-v16_254_127-id0.3_cov0.3genomes-v2-genome_set-humans-intervals-v1_512_256genomes-v2-genome_set-humans-intervals-v2_512_256genomes-v5-genome_set-humans-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v15_255_128
Humans downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
66,290 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v15_255_128.genomes-v3-genome_set-humans-intervals-v1_512_256genomes-v5-genome_set-humans-intervals-v17_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v17_255_128
Humans cCRE enhancers (v17) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
3,447,988 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v17_255_128.genomes-v5-genome_set-humans-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v5_255_128
Humans CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
539,732 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an exact… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v5_255_128.human-sim
Xuhui/human-sim
Processed dataset for user simulation. One row per user with grouped conversations.
Schema
Each row represents one user. Fields:
user_id (string): SHA-256 hashed IP from source data.
user_meta (struct): User-level metadata (country).
conversations (list of struct): All conversations for this user.
id (string): Conversation hash from source.
source (string): Source dataset identifier.
messages (list of struct): {role, content} message pairs.
metadata… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/human-sim.genomes-v5-genome_set-humans-intervals-v18_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v18_255_128
Humans conserved cCRE enhancers (v18) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
752,548 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v18_255_128.human_study_targetshumans-benchmark
HUMANS Benchmark Dataset
Authors: Woody Haosheng Gan¹, William Held²'³, Diyi Yang²
¹University of Southern California, ²Stanford University, ³OpenAthena
This dataset is part of the Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment paper.
HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark is designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/humans-benchmark.robojudge_humanstudy
RoboJudge human study — annotation packages
16 self-contained, offline annotation packages covering the 800-item RoboJudge test set.
packages
items each
total
human_study_01..16.zip
50
800
Together: 800 items, no duplicates, full coverage.
All clips belonging to the same (dataset, task, episode) scene stay inside one
package, so a scene is never split across two annotators.
How to use
Download one zip and unpack it.
Double-click index.html — it… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFriends/robojudge_humanstudy.synthetic-humans-1mThis dataset contains 1 million synthetic humans, sampled from actual US demographics. It is primarily meant to seed diverse LLM responses, but can be used for analytical purposes as well. The qualitivate_descriptions columns contains roughly 2.4 billion tokens, generated by Qwen/QwQ-32B with full reasoning traces.
A more detailed blog post on the methodology used to generate the dataset can be found here: https://www.skysight.inc/blog/synthetic-humans.
The dataset structure is as follows:… See the full description on the dataset page: https://huggingface.co/datasets/sutro/synthetic-humans-1m.human-senmot-flatmaphuman_stack_three_cup_2
