datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humans-top
humans.top — LIVE Global ranking of influential people (open dataset)
This dataset ranks real, named living people by global influence — e.g. #1
Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk,
Narendra Modi and Lionel Messi. Every row is a person: their live influence
rank, a concise biography in 15 languages, and Wikidata / Wikipedia links.
Published from the website humans.top (.top is the
domain name).
Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.humans-benchmark
HUMANS Benchmark Dataset
Authors: Woody Haosheng Gan¹, William Held²'³, Diyi Yang²
¹University of Southern California, ²Stanford University, ³OpenAthena
This dataset is part of the Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment paper.
HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark is designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned… See the full description on the dataset page: https://huggingface.co/datasets/woodygan/humans-benchmark.Human-Style-Answers
Human Style Answers
This Datasets contains question and answers on different topics in Human style. (For Chatbots training)
This Datasets is build using TOP AI like (GPT4, Claude3 , Command R+, etc.)
Dataset Details
Description
The Human Style Response Dataset is a rich collection of question-and-answer pairs, meticulously crafted in a human-like style. It serves as a valuable resource for training chatbots and conversational AI models. Let's dive into the… See the full description on the dataset page: https://huggingface.co/datasets/innova-ai/Human-Style-Answers.humans-benchmark
HUMANS Benchmark Dataset (Anonymous, Under Review)
This dataset is part of the HUMANS (HUman-aligned Minimal Audio evaluatioN Subsets for Large Audio Models) Benchmark, designed to efficiently evaluate Large Audio Models using minimal subsets while predicting human preferences through learned regression weights.
Installation
Install the HUMANS evaluation package from GitHub (our anonymous repo):
# Option 1: Install via pip
pip install… See the full description on the dataset page: https://huggingface.co/datasets/HUMANSBenchmark/humans-benchmark.human_study
human_study
Free-form human Reddit responses derived from
snap-stanford/user_study_annotations.
The original Reddit post text comes from the HumanLM authors'
reddit_post_dict_testset.json.
Each row contains an original Reddit post in prompt and the response written
by a human-study participant in target. Model responses, generated personas,
comparison judgments, and worker identifiers are intentionally excluded.
Deduplication and splits
Source annotation… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/human_study.hcms
HCMS Safety Challenge Datasets
Project repository: github.com/center-for-humans-and-machines/hcms
This repository contains two types of data: challenge conversations from an online safety challenge and synthetic conversations generated for safety evaluation.
Challenge Conversations (Online Experiment)
Raw export of 9,492 messages from 280 participants (2 duplicate ids) in the hpms_safety_challenge_main_2026_03_27 experiment. Participants were given hypothetical… See the full description on the dataset page: https://huggingface.co/datasets/center-for-humans-and-machines/hcms.
